FashionCLIP ONNX (split vision + text encoders)

ONNX export of patrickjohncyh/fashion-clip split into separate vision and text encoders. Hosted for use by @looped-app/garment-classifier as a third backbone option alongside the standard OpenAI CLIP variants from Xenova.

Files

File Size Purpose
vision_model.onnx ~353 MB Image encoder. Input pixel_values: float32 [B, 3, 224, 224] โ†’ output image_embeds: float32 [B, 512] (not L2-normalized).
text_model.onnx ~255 MB Text encoder. Input input_ids: int64 [B, sequence_length] โ†’ output text_embeds: float32 [B, 512] (not L2-normalized).
tokenizer.json 3.6 MB OpenAI CLIP BPE tokenizer (49408 vocab + SOT 49406 / EOT 49407). Compatible with @huggingface/tokenizers / @huggingface/transformers AutoTokenizer.
config.json, vocab.json, merges.txt, special_tokens_map.json, tokenizer_config.json, preprocessor_config.json misc Standard transformers metadata.

Caveats

  • Text encoder batch size must be 1 at runtime due to a torch 2.12 export-time dynamic-axis baking quirk. Run prompts sequentially (one forward per prompt). Our garment-classifier package does this naturally in its prompt-embedding precompute loop.
  • Outputs are NOT L2-normalized. Normalize in the consumer code before cosine similarity (CLIP convention).
  • Image preprocessing: requires standard CLIP-style normalization: resize to 224ร—224 + mean=[0.48145466, 0.4578275, 0.40821073] / std=[0.26862954, 0.26130258, 0.27577711] applied per-channel.

Provenance

  • Base model: patrickjohncyh/fashion-clip (282 likes, 3M downloads on HF as of 2026-05-24)
  • Export: torch 2.12.0 + onnx 1.21.0 via torch.onnx.export on separate CLIPVisionModelWithProjection / CLIPTextModelWithProjection wrappers.
  • Verification: real-image cosine vs PyTorch reference. Output is functional (no NaN, sensible top-1 predictions on apparel imagery).
  • Export script: /tmp/fashionclip-split-export.py in the original session (orchestrator-direct).
  • License: MIT (inherits from patrickjohncyh/fashion-clip).
Downloads last month
5
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Frapic/fashion-clip-onnx

Quantized
(3)
this model