decider-2b-vision β€” option probabilities from an image in LiteRT

Give it one image and lettered questions, and it returns a probability for each option. For the published Pong frame game_pong_atari_up, asked "What should you do right now?" with the options move paddle up / move paddle down / stay, the upstream model in fp32 gives 0.99986 to "move paddle up" on the same 256Γ—256 input. The fp16 and fp16-int8vocab files return the same value to five digits through the reference readout on a Mac CPU.

This is a third-party conversion of Mapika/decider-2b-vision at revision 863e290863655f1d6b69324d77d09ac972d21609. Mapika built it from the Qwen3.5-2B vision-language model with decider-2b v5 text weights (Qwen/Qwen3.5-2B-Base). The upstream model reads the letter logits at one answer slot per question and turns them into option probabilities; it does not generate answer text. Its training, results and limitations are on the upstream card. This repository changes the file format only.

The decoder has 24 layers: 18 gated-delta linear-attention layers and 6 full-attention layers. The vision encoder has 24 blocks and a 2Γ—2 merger, so a 256Γ—256 image becomes 64 tokens. The files were tested on LiteRT-LM 0.17.1 on a Mac (CPU and GPU); the int8 file was also tested on a Galaxy S26 with LiteRT-LM v0.16.0 (CPU and GPU). The reference readout needs ai-edge-litert 2.2.0, numpy, Pillow and tokenizers; it does not use torch or transformers.

Files

File Weights Bytes SHA-256 Use
decider-2b-vision_fp16.litertlm fp16 casting of every fully connected and embedding weight, float compute; vision encoder and adapter fp16 5,509,356,416 75b226796c43405b900399551487903dbe4f87d1a2c40aea422b2a9c33c8f62a Desktop CPU or GPU; closest to upstream. Not for phones: XNNPACK expands its weights to 7.5 GB, and it was not tried on the S26
decider-2b-vision_fp16-int8vocab.litertlm decoder fully connected weights fp16, output head dynamic int8, token embedding int8 (channelwise); vision encoder and adapter fp16 4,498,201,472 72f6836051f54ad32f7804e06bc0dae2e022a22f97d864d480f39710f17357bc Desktop GPU at about half the memory of fp16 (Mac GPU peak RSS 7.0 GB against 15.6 GB); probabilities within 0.0085 of upstream on the published rows. Did not fit the Galaxy S26 GPU (see Performance)
decider-2b-vision_int8.litertlm dynamic int8 fully connected weights, int8 token embedding (channelwise); vision encoder and adapter fp16 3,171,081,088 5fb2e19aa2066d2e3955431366bc572edb7abbc6037db53d4eaf4a4a4d0e7bd9 Phones. Tested on a Galaxy S26 CPU and GPU. Use the GPU when the probabilities matter (see below)

All three bundles hold the same tokenizer and prompt metadata: an identity template (no role markers), no start token, stop token 248044, one 256Γ—256 image, a 4096-token cache, ExecutorMetadata for the 48 state buffers and an fp32 activation preference for the decoder. XNNPACK unpacks fp16 weights to fp32, so LiteRT-LM's CPU cache folder for one bundle holds 7,540,052,280 B (fp16 decoder), 6,016,360,768 B (fp16-int8vocab decoder) or 1,896,902,288 B (int8 decoder), plus 1,312,687,576 B for the vision encoder and adapter. With the GPU backend, the folder gets a decoder program cache that grew by 270,663,680 B at every engine creation in our runs (821,634,912 B after three), plus a 508,563,920 B weight cache for the int8 head of the fp16-int8vocab file. File checksums are in SHA256SUMS.

The int8 file quantizes activations on the fly on the CPU, and there its probabilities move: on the published image questions 51/53 answers match upstream, max |Ξ”p| 0.429, p95 0.164, and the values also depend on how the prompt is split into prefill chunks (max |Ξ”p| 0.233 with one padded chunk). On the Mac GPU (Metal, one padded chunk per answer slot) it computes the int8 weights in float: 52/53, max |Ξ”p| 0.032, p95 0.011. Its probabilities on the phone were not measured, because the runtime returns text only; on the Galaxy S26 its first token was upstream's answer letter on 6/6 rows on both backends.

Use it

Option probabilities: the reference readout

python -m venv .venv
.venv/bin/pip install ai-edge-litert==2.2.0 numpy pillow tokenizers
.venv/bin/python -B reference/decider_litert.py --bundle decider-2b-vision_fp16-int8vocab.litertlm \
  --request reference/example_request.json

The example request is the Pong frame above. Output of that command on an Apple M4 Max:

{
 "model": "decider-2b-vision_fp16-int8vocab.litertlm",
 "backend": "cpu",
 "answers": [
  {
   "choice": "move paddle up",
   "confidence": 0.9998610273009231,
   "probs": {
    "move paddle up": 0.9998610273009231,
    "move paddle down": 1.4132675638724978e-05,
    "stay": 0.00012484002343821706
   }
  }
 ]
}

From Python:

import sys; sys.path.insert(0, 'reference')
from decider_litert import DeciderLiteRT

d = DeciderLiteRT('decider-2b-vision_fp16-int8vocab.litertlm')
d.decide('frame.png', 'You play Pong (Atari) and control the right paddle. ...',
         [{'question': 'What should you do right now?', 'options': ['move paddle up', 'move paddle down', 'stay']}])
# -> [{'choice': 'move paddle up', 'confidence': 0.99986..., 'probs': {...}}]

The first run copies the bundle's sections into .cache/readout/<bundle sha256>/ beside the reference/ folder (or --cache-dir) and writes the decoder's CPU weight cache there: 13,065,400,405 B in all for fp16 and 10,530,549,449 B for fp16-int8vocab. Pass image=None for a text-only request and several questions in one call; both are supported here and not through the runtime path below. backend='gpu' runs the decoder on the GPU with one padded prefill chunk per answer slot, for slots within the first 1024 tokens.

The answer letter: the LiteRT-LM runtime

pip install litert-lm==0.17.1 pillow tokenizers
python -B reference/runtime_example.py --bundle decider-2b-vision_fp16-int8vocab.litertlm \
  --image reference/fixtures/images/game_pong_atari_up_160x210.png \
  --context "You play Pong (Atari) and control the right paddle. Move the paddle so the ball hits it; the ball bounces off paddles and walls. Missing the ball loses a point. The image shows the current game screen." \
  --question "What should you do right now?" --options "move paddle up" "move paddle down" "stay"

runtime_example.py builds the prompt with the upstream prompt code, resizes the image to 256Γ—256 with PIL, sends [image, text] as one user message and reads the first generated token with greedy decoding. This path covers one image first, one question per request, and the argmax only. The runtime returns text, not option probabilities, and its greedy token is the most likely token over the whole vocabulary, not only over the option letters. The two agree on every single-question published row. In one two-question row, the most likely token for a two-option question is "C".

Several questions about one image go into one conversation: the image with the first question, then each further question as a text-only message that starts with ) (it closes the previous letter) and carries the next Question: … Options: … Answer: ( block. Do not open a new conversation per request on the same engine: LiteRT-LM keeps this model's linear-attention state across conversations (LiteRT-LM #3165), and the answers drift. Measured on a Galaxy S26 through the Android library (litertlm-android 0.16.1, the int8 file, 2026-09-29): one conversation with five questions per bar chart gave the answer of a fresh engine on 18 of 20 questions on the GPU and 5 of 5 on the CPU; a new conversation per request gave it on 33 of 60 (GPU) and 34 of 60 (CPU) with pixel-identical inputs. An app that does this, with its records: hfmodels-android samples/ask.

Examples

Published fixture images, one question each. The letter is the first token LiteRT-LM 0.17.1 generated on the Mac (fp16-int8vocab file, CPU and GPU backends, greedy); the probability is that option's, from the reference readout (fp16-int8vocab, CPU) and from upstream fp32 on the same 256Γ—256 input.

Image Question and options LiteRT-LM (CPU / GPU) Reference readout Upstream fp32
What should you do right now? (A) move paddle up (B) move paddle down (C) stay A / A 0.9999 0.9999
What should you do right now? (A) move paddle left (B) move paddle right (C) stay (D) launch ball D / D 0.9972 0.9972
How many dots are in the image? (A) 5 (B) 6 (C) 7 (D) 8 (E) 9 B / B 0.7814 0.7780

What matches upstream

The files reproduce decider/vision.py prepare() + slot_logits() from the checkpoint under this contract:

  1. One image, first in the prompt, resized to 256Γ—256 with PIL bicubic. The vision encoder takes [0, 1] pixels and normalizes them itself; 64 image tokens follow <|vision_start|>.
  2. The text is the upstream build() output with the option order kept, decoded and encoded again as upstream does. No BOS, no role markers.
  3. Answer slots follow the upstream rule: a " (" token after ":", with "Answer" within the 5 tokens before it. The letter logits A–J at each slot are cut to the option count and softmaxed at T = 1.
  4. Positions: the checkpoint uses 3-channel M-RoPE positions (image tokens on an 8Γ—8 grid). LiteRT-LM passes one 1-D position, so the decoder derives the three channels from it inside its graph. On the 53 published image questions, plain 1-D positions flipped 2 answers and moved one probability by 0.198 in the upstream fp32 model; the files follow the M-RoPE values instead.
  5. Text-only requests start at position 65, where the derived positions keep upstream's relative positions. Started at 0, the same 9 text-only questions moved by up to 0.0497 and one answer flipped in the fp32 LiteRT graph. The runtime starts text-only prompts at 0, so text-only requests go through the reference readout.

Agreement with upstream on the published rows

The fixtures hold 42 requests (37 with an image, 5 text-only; 62 answer slots) with the upstream fp32 letter logits and probabilities. Every image is a synthetic drawing made for this check: colour cards, shapes, text and simple Pong- and Breakout-style frames. check_fixtures.py rebuilds each request from the original image and text, asserts that the 256Γ—256 pixels, token ids and answer slots equal upstream's, and compares the probabilities. |Ξ”p| is the largest absolute probability difference over a slot's options. No upstream top-two gap is within the 1e-4 tie threshold.

File Backend / feeding Image slots: argmax equal max |Ξ”p| p95 |Ξ”p| Text-only slots: argmax equal max |Ξ”p|
fp16 CPU / exact-fit prefill 53/53 4.06e-05 3.04e-05 9/9 2.41e-06
fp16-int8vocab CPU / exact-fit prefill 52/53 0.00846 0.00636 9/9 0.00306
fp16 Mac GPU (Metal) / one padded chunk per slot 53/53 4.07e-05 3.31e-05 9/9 1.10e-06
fp16-int8vocab Mac GPU (Metal) / one padded chunk per slot 53/53 0.00621 0.00290 9/9 0.00228

The fp16-int8vocab CPU flip is v2_game_breakout_atari_fewbricks question 1, where upstream's top two options are 0.0073 apart. CPU rows are the included reference (ai-edge-litert 2.2.0 CompiledModel, 8 threads); its letter logits are bit-identical to the conversion's own graph readout on all 62 slots of all three files. GPU rows are ai-edge-litert 2.2.0 CompiledModel with GpuOptions(enforce_f32=True). Through LiteRT-LM 0.17.1 on the Mac, all three files on CPU and on GPU gave upstream's answer letter as the first token on 12/12 single-question image rows, with prefill counts equal to upstream's token counts and no Validation error lines.

The price of the fixed 256Γ—256 input

Upstream picks a resolution per image; these files take 256Γ—256. For 224Γ—224, 256Γ—240 and 256Γ—256 originals, upstream's own processing produces the same pixels, so the probabilities are identical (13 published requests, 21 slots, |Ξ”p| 0). Every other size gets fewer image tokens than upstream would use: 70 for the 160Γ—210 frames and 144 to 475 for the larger synthetic images. On those published requests (24 requests, 32 slots), upstream fp32 on the 256Γ—256 input matches upstream fp32 on the original image on 30/32 answers, with max |Ξ”p| 0.373 and p95 0.318. On photographs and documents kept out of this repository, the largest move was 0.58. Both sides of these comparisons are upstream fp32, so the moves come from the input size, not from the conversion.

Performance

One decision per fresh process on an Apple M4 Max (Mac16,9, 128 GB, macOS 27.0) through LiteRT-LM 0.17.1 (PyPI, Python API): the published Pong frame (game_pong_atari_up, 256Γ—256) with its one question, 150 prompt tokens including 64 image tokens, greedy, one output token. TTFT is the wall time from sending the message to the first token, measured in the process. Engine creation is the Engine(...) constructor. Peak RSS and peak footprint come from /usr/bin/time -l; RSS also counts the mapped model and cache files. CPU runs use 8 threads; GPU runs use Backend.GPU() (the WebGPU delegate on Metal) for the decoder and the vision encoder, after a rest of at least 120 s since the previous GPU run. Each value is the median of three runs, with the runs in parentheses. "none" is cache_dir=":nocache"; "yes, warm" runs after a first run that wrote the cache folder. Runs taken while another process used more than 120 % CPU were repeated or dropped and are not in the table. Every run gave the letter A, and the GPU runs printed no Validation error line. The int8 file was not timed on the Mac; its phone rows follow.

File Backend Runtime cache folder TTFT s Engine creation s Prefill tok/s Peak RSS GB Peak footprint GB
fp16 CPU none 1.68 (1.52, 2.05, 1.68) 16.8 132 53.7 47.9
fp16 CPU yes, warm 1.72 (1.26, 7.59, 1.72) 13.7 115 21.2 6.6
fp16 GPU none 0.29 (0.35, 0.28, 0.29) 28.5 826 15.6 19.3
fp16 GPU yes, warm 0.32 (0.28, 0.32, 0.41) 29.2 811 15.8 19.6
fp16-int8vocab CPU none 1.54 (1.54, 1.50, 1.65) 16.3 140 45.5 41.3
fp16-int8vocab CPU yes, warm 4.42 (1.16, 4.42, 4.43) 13.4 49 12.5 1.5
fp16-int8vocab GPU none 0.26 (0.35, 0.26, 0.26) 26.8 843 7.0 10.6
fp16-int8vocab GPU yes, warm 0.47 (0.28, 0.66) 30.2 688 6.7 10.2

Median TTFT is 0.26–0.47 s on the GPU and 1.54–1.72 s on the CPU, apart from the fp16-int8vocab warm-folder CPU row explained below. GPU engine creation takes 27–30 s against 13–17 s on the CPU. Without a cache folder, the CPU process's peak footprint is 47.9 GB (fp16) and 41.3 GB (fp16-int8vocab); with a warm folder it is 6.6 GB and 1.5 GB. With a warm folder, CPU TTFT depends on whether the cache file is still in memory: 1.16–1.72 s when it was, 4.4–7.6 s when about 0.46–0.56 million page faults read it back from disk. A warm folder did not shorten GPU engine creation. The fp16-int8vocab GPU warm row has two runs; its third run overlapped another process and was dropped.

The reference readout on the same machine (ai-edge-litert 2.2.0 CompiledModel CPU, 8 threads, its section folder and weight cache already on disk), one process per run. The first decision in a process includes the first use of each graph; the second decision in the same process is the steady cost.

File Constructor s First decision s Second decision s Peak RSS GB
fp16 12.8 1.28 (11.69, 1.28, 1.28) 0.90 22.1
fp16-int8vocab 12.9 2.85 (3.89, 2.85, 1.00) 1.02 13.5

Galaxy S26 (SM-S942Q, MemTotal 11,389,756 kB), LiteRT-LM v0.16.0 litert_lm_advanced_main, a different runtime version from the Mac rows. One cold run per row (n = 1): the published row game_pong_atari_level (150 prompt tokens, one question), greedy, up to 3 output tokens, runtime default cache folder, thermal status 0 and no frequency cap before the run. Engine creation is the runtime's Init Total; the conversation set-up after it is in parentheses. Peak memory is the process VmHWM.

File Backend Result TTFT s Prefill tok/s Engine creation s Peak VmHWM GB
int8 CPU (4 threads) first token C = upstream; 6/6 rows 1.22 136.68 17.7 (+2.0) 4.54
int8 GPU (OpenCL) first token C = upstream; 6/6 rows 0.56 340.76 64.4 (+7.0) 4.32
fp16-int8vocab GPU (OpenCL) engine not created: stopped at 6.55 GB VmHWM during OpenCL set-up of prefill_1024 (our 6.5 GB guard) β€” β€” β€” β€”
fp16-int8vocab CPU not run: its XNNPACK cache alone is 6.0 GB β€” β€” β€” β€”
fp16 CPU / GPU not run: XNNPACK expands its weights to 7.5 GB β€” β€” β€” β€”

On the GPU all 7 decoder signatures were fully delegated to OpenCL (one partition each); the vision encoder ran on the GPU and the adapter on the CPU (the cache files each wrote). VmHWM does not include GPU memory: during the int8 GPU run the phone's MemAvailable fell to 0.77 GB, so the int8 GPU run is close to the limit of a phone of this RAM class. The other five int8 CPU rows peaked at 4.82–4.85 GB VmHWM. Every int8 row on the S26 printed no Validation error line and processed the same number of prefill tokens as upstream.

Limitations

  • The runtime path gives the argmax, with one image first; several questions about that image go as turns of one conversation (above), not as one conversation per question. The option probabilities cannot be read through LiteRT-LM: its text-scoring call returns scores that disagree with its own greedy decode on 0.17.1 (LiteRT-LM #3561). Use the reference readout for probabilities and for text-only requests.
  • The runtime's own image resize was not exercised: every runtime check handed it a 256Γ—256 PNG resized with PIL. The examples here do the same.
  • One image per request. English only. The text behaviour is the upstream model's (decider-2b v5 text weights); see the upstream limitations.
  • Rows longer than 4096 tokens do not fit the cache. Every published row is at most 225 tokens; prompts above 2048 tokens were not exercised.
  • Phones: only the int8 file was tested on a phone (Galaxy S26, one cold run per row). The fp16-int8vocab file did not fit its GPU, and the fp16 file was not tried. Probabilities on the phone were not measured.

Other ports

As of 2026-09-29, the Hugging Face search for decider-2b-vision lists two GGUF ports, mradermacher/decider-2b-vision-GGUF and mindchain/decider-2b-vision-GGUF, and no other LiteRT port.

License and changes

Apache-2.0, as declared by the upstream checkpoint; the vendored prompt code is Apache-2.0 too (LICENSE). Source revision: 863e290863655f1d6b69324d77d09ac972d21609.

Changes: converted the text decoder, the vision encoder and the merger to LiteRT flatbuffers; the decoder derives the upstream M-RoPE positions from a 1-D position inside its graph; fixed the image input at 256Γ—256; cast the weights as listed under Files; replaced the chat template with an identity template; repackaged the tokenizer unchanged; added ExecutorMetadata and an fp32 activation preference. The multi-token-prediction head is not included. This community conversion is not an official Mapika or Qwen release.

decider_vendored/prompt.py is the checkpoint's own decider/prompt.py, unchanged (SHA-256 c2fadbe0e4703a381011007f52103ab29142a55a441302d62c1dfb2056acce16). decider_vendored/infer.py keeps only its Q and Example dataclasses. The LICENSE is from github.com/Mapika/decider at commit 23579f7a7e8f10e1045be492af3c1c05a005d67c.

Reproduction (conversion scripts, gates, fixtures): hf-to-litertlm decider2bv_work/.

Downloads last month
47
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for litert-community/decider-2b-vision-LiteRT

Finetuned
(1)
this model