Spaces:
Running
Download README.md from primerz/pixagram-search-v4: direct link, hf CLI and curl.
- Browser
- Download file 11 kB
-
https://huggingface.co/spaces/primerz/pixagram-search-v4/resolve/main/README.md
- Command line
-
hf download hf://spaces/primerz/pixagram-search-v4/README.md
-
curl -L -o README.md https://huggingface.co/spaces/primerz/pixagram-search-v4/resolve/main/README.md
A newer version of the Gradio SDK is available: 6.30.0
title: pixagram-search embeddings
emoji: 🟪
colorFrom: purple
colorTo: indigo
sdk: gradio
sdk_version: 6.28.0
app_file: app.py
pinned: false
license: apache-2.0
short_description: SigLIP image/text embeddings for the Pixagram search engine
SigLIP embedding Space for pixagram-search
One model, two towers. The Cloudflare Worker sends artwork PNGs here while indexing and the
query text here at search time; both land in the same vector space, which is what makes
text → image search work. app.py is the Space entry point; it serves:
| route | |
|---|---|
POST /embed |
{"inputs": {"images": ["<base64>", …], "texts": ["…", …]}} → {"model", "dim", "embeddings": [[…], …], "calibration", "max_num_patches"} (images first, then texts; L2-normalised; max_num_patches for NaFlex models) |
GET /health |
{"ok", "model", "dim", "ready", "stub", "calibration", "max_num_patches"} — ready flips to true once the weights are loaded |
GET / |
Gradio page: embed a text or an image, score an image against captions |
calibration is the model's learned {"logit_scale", "logit_bias"}:
sigmoid(logit_scale · cos + logit_bias) is SigLIP's own text↔image match probability. The v3
Worker stores it (KV calib:<model>) and uses it until it has a background sample of the
corpus; older Spaces without the field still work (the Worker knows the SigLIP 2 base values).
Deploy as a Space
Create a Space (SDK: Gradio, hardware: CPU basic is enough — a SigLIP-base pass on one artwork takes ~0.25 s on 2 vCPU, ~0.5 s for NaFlex at 576 patches). The Space computes one embedding per CPU at once and queues the rest; the Worker reads that number from
/health(concurrency) every ten minutes and sends as many per consumer invocation, so bigger hardware is used with nothing else to change: on CPU upgrade (8 vCPU) it computes 8 at once instead of 2.Upload this folder —
README.md(the frontmatter above is the Space config),app.py,siglip.py,requirements.txt;handler.pycan stay, it is only used by Inference Endpoints. From the repository root:pip install -U huggingface_hub hf auth login # older CLI: huggingface-cli login hf upload <owner>/<space> hf . --repo-type spacescripts/deploy.shdoes all of this for the v4 Space (primerz/pixagram-search-v4); the v3 repository did it forprimerz/pixagram-search-v3.Visibility:
- Private Space (recommended): every request must carry an HF token that can read the
Space — set that token as the Worker's
HF_TOKENsecret. LeaveAPI_TOKENunset. - Public Space: set a Space secret
API_TOKENto a long random string and use the same string as the Worker'sHF_TOKEN;/embedthen rejects anything else with 401.
- Private Space (recommended): every request must carry an HF token that can read the
Space — set that token as the Worker's
Worker:
npx wrangler secret put HF_TOKEN, and setHF_EMBED_URL = "https://<owner>-<space-name>.hf.space/embed"inwrangler.jsonc(the host is<owner>-<space-name>with dots in the owner name replaced by dashes; the Space settings page shows the exact "Direct URL").Deploy the Worker, then
scripts/admin.sh reindex-all embed,textto fill the vectors in.
Cold start: the SigLIP base checkpoints are ~1.5 GB; the Space answers /health
in seconds and /embed blocks until the weights are loaded (about 20 s once cached, a few
minutes on the first build). Free CPU Spaces sleep after 48 h without traffic; the Worker
treats the wake-up page as a transient error and retries with backoff, and search simply runs
without the semantic leg until the Space is back. If that gap matters, use paid CPU hardware
and set the sleep time to "never".
Concurrency
The Space computes MAX_CONCURRENCY embeddings at once (default: one per CPU), each with
CPUs / MAX_CONCURRENCY PyTorch threads (TORCH_THREADS overrides), on threads that live as long
as the server. Search queries (texts only) have threads of their own and never wait behind
images being indexed. /health answers at once even under load.
Measured on 2 vCPU (Oct 2026): 24 NaFlex images at 576 patches from 4 clients, alone, then with a search text every second:
| images alone | images with searches | a search text meanwhile | /health meanwhile |
|
|---|---|---|---|---|
| before: one request at a time, 2 threads | 1.81 / s | 1.76 / s | median 1.25 s, worst 1.9 s | blocked (seconds) |
| one per CPU, 1 thread each | 1.89 / s | 1.77 / s | median 0.22 s, worst 0.29 s | 3 ms |
| one at a time, 2 threads | 1.55 / s |
Two images at once with one thread each beat one image with two threads (1.89 against 1.55
images/s here), which is why the default is one per CPU. Fewer at once with more threads each
(MAX_CONCURRENCY=1) answers a single query sooner (~70 ms instead of ~125 ms for a text on 2
vCPU) at the cost of throughput.
python3 hf/loadtest.py https://<owner>-<space>.hf.space <token> measures a running Space the
same way (the token is the Worker's HF_TOKEN, SPACE_API_TOKEN in ~/.pixagram-search-v4.json):
run it before and after changing the hardware.
Troubleshooting
Text embeddings taking 8-10 s on a CPU Space:
os.cpu_count()inside the container reports the host's cores (64-96) while the cgroup allows 2-8, and PyTorch spin-waits on the oversubscribed thread pool.siglip.pysizes everything from the cgroup quota (effective_cpus());/healthreportscpus,concurrency(embeddings at once) andthreads(PyTorch threads per embedding) so you can see what it picked.[Errno 98] error while attempting to bind on address ('0.0.0.0', 7860)in the build/run log: Spaces setGRADIO_SSR_MODE=True, which makes Gradio start a Node SSR server on 7860 before uvicorn.app.pymounts Gradio withssr_mode=Falsefor that reason — keep it.WARNING: Running pip as the 'root' user…during the build is harmless; the Space image installs requirements as root by design.Your space is in erroron the direct URL: open the Space page → Logs (Build, then Container);/healthonly answers once the container is running.Intermittent
502HTML pages (fast, ~0.15 s, nox-proxied-host/x-proxied-replicaresponse header) while the Space shows Running: Hugging Face's edge fails before the request reaches the container, so the Container log shows nothing. On 2026-09-28 this came in windows of 10-50 s (up to 34 failures in a row). A quick retry does not bridge that. Search falls back to full text, and the queue retries with backoff. Restart the Space. If it persists, report it to HF with thex-request-idof a failed call, or move to paid hardware or an Inference Endpoint (handler.py).
Test
curl -s https://<owner>-<space>.hf.space/health
curl -s https://<owner>-<space>.hf.space/embed -H "Authorization: Bearer $HF_TOKEN" \
-H "Content-Type: application/json" -d '{"inputs":{"texts":["a swan on a lake at sunset"]}}' | jq '.dim'
Local run (any machine with Python 3.10+): pip install -r requirements.txt && python app.py,
then point the Worker's .dev.vars at HF_EMBED_URL=http://127.0.0.1:7860/embed.
EMBED_STUB=1 python app.py starts without weights and returns deterministic pseudo-vectors —
handy for wiring tests, never for production (/health reports "stub": true).
Changing the model
The code default is google/siglip-base-patch16-256-multilingual, which is what the
production Space runs. Choose a model per Space with the MODEL_ID variable, so that
uploading hf/ never changes production's model:
- the v2 Space (
primerz/pixagram-siglip2) setsMODEL_ID=google/siglip2-base-patch16-256; - the v3 Space (
primerz/pixagram-search-v3) and the v4 Space (primerz/pixagram-search-v4, created by this repository'sscripts/deploy.sh) setMODEL_ID=google/siglip2-base-patch16-naflexandMAX_NUM_PATCHES=576: v4 embeds exactly as v3 does.
NaFlex
NaFlex checkpoints keep each image's aspect ratio. The processor resizes the image to the
largest size that fits MAX_NUM_PATCHES patches of 16×16 pixels (default 256), and batches mix
shapes through an attention mask. Every reply reports max_num_patches. The Worker refuses
image vectors whose budget differs from its EMBED_PATCHES (text vectors do not depend on it),
and re-embeds when EMBED_PATCHES changes. On the Pixagram corpus, 256 patches did worse than
the fixed 256 px model and 576 slightly better (README-V3.md, "SigLIP 2 NaFlex").
Any SigLIP / SigLIP 2 checkpoint works unchanged: fixed-resolution ones load as SiglipModel,
NaFlex ones as Siglip2Model. Stick to multilingual checkpoints, because queries arrive
in any language: every SigLIP 2 model, google/siglip-base-patch16-256-multilingual or
google/siglip-so400m-patch16-256-i18n. The other SigLIP 1 checkpoints have an English text
tower.
Measured on the 106 Pixagram artworks (+500 pixel-art distractors, 157 queries in EN/FR/DE/JA, 2 vCPU, Sept 2026):
| model | dim | ms / image | ms / query | notes |
|---|---|---|---|---|
| siglip-base-patch16-256-multilingual | 768 | 310 | 90 | production |
| siglip2-base-patch16-256 | 768 | 315 | 87 | v2 Space; same cost; EN/FR on par, DE/JA better |
| siglip2-base-patch16-384 / -naflex (256 patches) | 768 | 650 / 305 | 85 | no gain over base-256 |
| siglip2-base-patch16-naflex, 576 patches | 768 | ~530 | 85 | v3 Space: small gain, measured in Oct 2026 on 133 artworks (README-V3.md) |
| siglip2-large-patch16-256 | 1024 | 980 | 290 | no gain |
| siglip2-so400m-patch16-256 | 1152 | 1270 | 390 | small gain (JA) at 4× the CPU |
| siglip-so400m-patch16-256-i18n | 1152 | 1280 | 395 | best measured, 4× the CPU |
Switching a stack in place, same dimension (768 → 768): keep the Vectorize index. Set the
Space's MODEL_ID, then set EMBED_MODEL to the same id in that stack's wrangler config
and deploy. The Worker refuses vectors whose model differs from EMBED_MODEL, and
re-embeds artworks whose embed_model differs (the 10-minute sweeper picks them up; to do it
at once: scripts/admin.sh reindex-all embed,text). Then refresh the background sample
(scripts/admin.sh background), since z-scores are per model.
A different dimension also needs new Vectorize indexes (with the metadata indexes from
scripts/lib.sh) that VEC and VEC_TEXT point at, and a new EMBED_DIM. Vectors from
different models must never be compared.
Inference Endpoint instead of a Space
handler.py implements the EndpointHandler contract for a dedicated Inference Endpoint
(task custom, same siglip.py, same request/response shape). Use it when you want
autoscaling, private networking, or GPU without the Space UI. Then HF_EMBED_URL is the
endpoint URL itself and the Worker's X-Scale-Up-Timeout: 600 header makes a scaled-to-zero
replica wake up within the request instead of returning 503.