SigLIP base-patch16-224 β€” image and text towers (ONNX)

ONNX exports of both towers of google/siglip-base-patch16-224 (Apache-2.0), as served by VisionServe β€” crop rescoring in the rfdetr-gdino-siglip* open-vocabulary router.

visionserve pull siglip-image
visionserve pull siglip-text
visionserve pull rfdetr-gdino-siglip       # pulls both towers as dependencies

Files

file contents
image/model.onnx + image/model.onnx.data vision tower, opset 17, fp32 (external weights; keep the pair together)
text/model.onnx + text/model.onnx.data text tower, opset 17, fp32
text/tokenizer.json the checkpoint's SentencePiece Unigram tokenizer (32 000 pieces)

I/O contract

tower input output
image pixel_values f32 [N, 3, 224, 224] image_embeds f32 [N, 768], not L2-normalised
text input_ids i64 [N, 64] text_embeds f32 [N, 768], not L2-normalised
  • Image preprocessing: resize to exactly 224Γ—224 (squash, bicubic, no centre crop), /255, mean = std = 0.5 per channel β†’ [-1, 1]. CLIP's constants are wrong here and fail silently.
  • Text: pad to 64 tokens with </s> (id 1), not <pad> (id 0). SigLIP pools with attention and reads the padding; padding with id 0 moves embeddings 0.706 cosine away from the reference.

Both exports were verified against the PyTorch checkpoint before being written (image: 32 real crops; text: per-prompt cosine floor 0.9999).

License

Apache-2.0, inherited from google/siglip-base-patch16-224.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for mtbui2010/siglip-base-patch16-224-ONNX

Quantized
(9)
this model