SigLIP base-patch16-224 β image and text towers (ONNX)
ONNX exports of both towers of google/siglip-base-patch16-224
(Apache-2.0), as served by VisionServe β crop
rescoring in the rfdetr-gdino-siglip* open-vocabulary router.
visionserve pull siglip-image
visionserve pull siglip-text
visionserve pull rfdetr-gdino-siglip # pulls both towers as dependencies
Files
| file | contents |
|---|---|
image/model.onnx + image/model.onnx.data |
vision tower, opset 17, fp32 (external weights; keep the pair together) |
text/model.onnx + text/model.onnx.data |
text tower, opset 17, fp32 |
text/tokenizer.json |
the checkpoint's SentencePiece Unigram tokenizer (32 000 pieces) |
I/O contract
| tower | input | output |
|---|---|---|
| image | pixel_values f32 [N, 3, 224, 224] |
image_embeds f32 [N, 768], not L2-normalised |
| text | input_ids i64 [N, 64] |
text_embeds f32 [N, 768], not L2-normalised |
- Image preprocessing: resize to exactly 224Γ224 (squash, bicubic, no centre crop), /255, mean = std = 0.5 per channel β [-1, 1]. CLIP's constants are wrong here and fail silently.
- Text: pad to 64 tokens with
</s>(id 1), not<pad>(id 0). SigLIP pools with attention and reads the padding; padding with id 0 moves embeddings 0.706 cosine away from the reference.
Both exports were verified against the PyTorch checkpoint before being written (image: 32 real crops; text: per-prompt cosine floor 0.9999).
License
Apache-2.0, inherited from google/siglip-base-patch16-224.
Model tree for mtbui2010/siglip-base-patch16-224-ONNX
Base model
google/siglip-base-patch16-224