Instructions to use litert-community/decider-2b-vision-LiteRT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT-LM
How to use litert-community/decider-2b-vision-LiteRT with LiteRT-LM:
# LiteRT-LM runs on various platforms (Android, iOS, Windows, Linux, macOS, IoT, Web/WASM) # and supports many APIs (C++, Python, Kotlin, Swift, JavaScript, Flutter). # For platform-specific integration guides, please refer to the official developer website: # https://ai.google.dev/edge/litert-lm # To try LiteRT-LM, the easiest way is to use our CLI tool. # 1. Install the LiteRT-LM CLI tool: pip install -U litert-lm # 2. Download and run this model locally: # See: https://ai.google.dev/edge/litert-lm/cli litert-lm run \ --from-huggingface-repo=litert-community/decider-2b-vision-LiteRT \ --prompt="Write me a poem"
- LiteRT
How to use litert-community/decider-2b-vision-LiteRT with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
decider-2b-vision β option probabilities from an image in LiteRT
Give it one image and lettered questions, and it returns a probability for each option. For the published Pong frame game_pong_atari_up, asked "What should you do right now?" with the options move paddle up / move paddle down / stay, the upstream model in fp32 gives 0.99986 to "move paddle up" on the same 256Γ256 input. The fp16 and fp16-int8vocab files return the same value to five digits through the reference readout on a Mac CPU.
This is a third-party conversion of Mapika/decider-2b-vision at revision 863e290863655f1d6b69324d77d09ac972d21609. Mapika built it from the Qwen3.5-2B vision-language model with decider-2b v5 text weights (Qwen/Qwen3.5-2B-Base). The upstream model reads the letter logits at one answer slot per question and turns them into option probabilities; it does not generate answer text. Its training, results and limitations are on the upstream card. This repository changes the file format only.
The decoder has 24 layers: 18 gated-delta linear-attention layers and 6 full-attention layers. The vision encoder has 24 blocks and a 2Γ2 merger, so a 256Γ256 image becomes 64 tokens. The files were tested on LiteRT-LM 0.17.1 on a Mac (CPU and GPU); the int8 file was also tested on a Galaxy S26 with LiteRT-LM v0.16.0 (CPU and GPU). The reference readout needs ai-edge-litert 2.2.0, numpy, Pillow and tokenizers; it does not use torch or transformers.
Files
| File | Weights | Bytes | SHA-256 | Use |
|---|---|---|---|---|
| decider-2b-vision_fp16.litertlm | fp16 casting of every fully connected and embedding weight, float compute; vision encoder and adapter fp16 | 5,509,356,416 | 75b226796c43405b900399551487903dbe4f87d1a2c40aea422b2a9c33c8f62a |
Desktop CPU or GPU; closest to upstream. Not for phones: XNNPACK expands its weights to 7.5 GB, and it was not tried on the S26 |
| decider-2b-vision_fp16-int8vocab.litertlm | decoder fully connected weights fp16, output head dynamic int8, token embedding int8 (channelwise); vision encoder and adapter fp16 | 4,498,201,472 | 72f6836051f54ad32f7804e06bc0dae2e022a22f97d864d480f39710f17357bc |
Desktop GPU at about half the memory of fp16 (Mac GPU peak RSS 7.0 GB against 15.6 GB); probabilities within 0.0085 of upstream on the published rows. Did not fit the Galaxy S26 GPU (see Performance) |
| decider-2b-vision_int8.litertlm | dynamic int8 fully connected weights, int8 token embedding (channelwise); vision encoder and adapter fp16 | 3,171,081,088 | 5fb2e19aa2066d2e3955431366bc572edb7abbc6037db53d4eaf4a4a4d0e7bd9 |
Phones. Tested on a Galaxy S26 CPU and GPU. Use the GPU when the probabilities matter (see below) |
All three bundles hold the same tokenizer and prompt metadata: an identity template (no role markers), no start token, stop token 248044, one 256Γ256 image, a 4096-token cache, ExecutorMetadata for the 48 state buffers and an fp32 activation preference for the decoder. XNNPACK unpacks fp16 weights to fp32, so LiteRT-LM's CPU cache folder for one bundle holds 7,540,052,280 B (fp16 decoder), 6,016,360,768 B (fp16-int8vocab decoder) or 1,896,902,288 B (int8 decoder), plus 1,312,687,576 B for the vision encoder and adapter. With the GPU backend, the folder gets a decoder program cache that grew by 270,663,680 B at every engine creation in our runs (821,634,912 B after three), plus a 508,563,920 B weight cache for the int8 head of the fp16-int8vocab file. File checksums are in SHA256SUMS.
The int8 file quantizes activations on the fly on the CPU, and there its probabilities move: on the published image questions 51/53 answers match upstream, max |Ξp| 0.429, p95 0.164, and the values also depend on how the prompt is split into prefill chunks (max |Ξp| 0.233 with one padded chunk). On the Mac GPU (Metal, one padded chunk per answer slot) it computes the int8 weights in float: 52/53, max |Ξp| 0.032, p95 0.011. Its probabilities on the phone were not measured, because the runtime returns text only; on the Galaxy S26 its first token was upstream's answer letter on 6/6 rows on both backends.
Use it
Option probabilities: the reference readout
python -m venv .venv
.venv/bin/pip install ai-edge-litert==2.2.0 numpy pillow tokenizers
.venv/bin/python -B reference/decider_litert.py --bundle decider-2b-vision_fp16-int8vocab.litertlm \
--request reference/example_request.json
The example request is the Pong frame above. Output of that command on an Apple M4 Max:
{
"model": "decider-2b-vision_fp16-int8vocab.litertlm",
"backend": "cpu",
"answers": [
{
"choice": "move paddle up",
"confidence": 0.9998610273009231,
"probs": {
"move paddle up": 0.9998610273009231,
"move paddle down": 1.4132675638724978e-05,
"stay": 0.00012484002343821706
}
}
]
}
From Python:
import sys; sys.path.insert(0, 'reference')
from decider_litert import DeciderLiteRT
d = DeciderLiteRT('decider-2b-vision_fp16-int8vocab.litertlm')
d.decide('frame.png', 'You play Pong (Atari) and control the right paddle. ...',
[{'question': 'What should you do right now?', 'options': ['move paddle up', 'move paddle down', 'stay']}])
# -> [{'choice': 'move paddle up', 'confidence': 0.99986..., 'probs': {...}}]
The first run copies the bundle's sections into .cache/readout/<bundle sha256>/ beside the reference/ folder (or --cache-dir) and writes the decoder's CPU weight cache there: 13,065,400,405 B in all for fp16 and 10,530,549,449 B for fp16-int8vocab. Pass image=None for a text-only request and several questions in one call; both are supported here and not through the runtime path below. backend='gpu' runs the decoder on the GPU with one padded prefill chunk per answer slot, for slots within the first 1024 tokens.
The answer letter: the LiteRT-LM runtime
pip install litert-lm==0.17.1 pillow tokenizers
python -B reference/runtime_example.py --bundle decider-2b-vision_fp16-int8vocab.litertlm \
--image reference/fixtures/images/game_pong_atari_up_160x210.png \
--context "You play Pong (Atari) and control the right paddle. Move the paddle so the ball hits it; the ball bounces off paddles and walls. Missing the ball loses a point. The image shows the current game screen." \
--question "What should you do right now?" --options "move paddle up" "move paddle down" "stay"
runtime_example.py builds the prompt with the upstream prompt code, resizes the image to 256Γ256 with PIL, sends [image, text] as one user message and reads the first generated token with greedy decoding. This path covers one image first, one question per request, and the argmax only. The runtime returns text, not option probabilities, and its greedy token is the most likely token over the whole vocabulary, not only over the option letters. The two agree on every single-question published row. In one two-question row, the most likely token for a two-option question is "C".
Several questions about one image go into one conversation: the image with the first question, then each further question as a text-only message that starts with ) (it closes the previous letter) and carries the next Question: β¦ Options: β¦ Answer: ( block. Do not open a new conversation per request on the same engine: LiteRT-LM keeps this model's linear-attention state across conversations (LiteRT-LM #3165), and the answers drift. Measured on a Galaxy S26 through the Android library (litertlm-android 0.16.1, the int8 file, 2026-09-29): one conversation with five questions per bar chart gave the answer of a fresh engine on 18 of 20 questions on the GPU and 5 of 5 on the CPU; a new conversation per request gave it on 33 of 60 (GPU) and 34 of 60 (CPU) with pixel-identical inputs. An app that does this, with its records: hfmodels-android samples/ask.
Examples
Published fixture images, one question each. The letter is the first token LiteRT-LM 0.17.1 generated on the Mac (fp16-int8vocab file, CPU and GPU backends, greedy); the probability is that option's, from the reference readout (fp16-int8vocab, CPU) and from upstream fp32 on the same 256Γ256 input.
What matches upstream
The files reproduce decider/vision.py prepare() + slot_logits() from the checkpoint under this contract:
- One image, first in the prompt, resized to 256Γ256 with PIL bicubic. The vision encoder takes [0, 1] pixels and normalizes them itself; 64 image tokens follow
<|vision_start|>. - The text is the upstream
build()output with the option order kept, decoded and encoded again as upstream does. No BOS, no role markers. - Answer slots follow the upstream rule: a
" ("token after":", with"Answer"within the 5 tokens before it. The letter logits AβJ at each slot are cut to the option count and softmaxed at T = 1. - Positions: the checkpoint uses 3-channel M-RoPE positions (image tokens on an 8Γ8 grid). LiteRT-LM passes one 1-D position, so the decoder derives the three channels from it inside its graph. On the 53 published image questions, plain 1-D positions flipped 2 answers and moved one probability by 0.198 in the upstream fp32 model; the files follow the M-RoPE values instead.
- Text-only requests start at position 65, where the derived positions keep upstream's relative positions. Started at 0, the same 9 text-only questions moved by up to 0.0497 and one answer flipped in the fp32 LiteRT graph. The runtime starts text-only prompts at 0, so text-only requests go through the reference readout.
Agreement with upstream on the published rows
The fixtures hold 42 requests (37 with an image, 5 text-only; 62 answer slots) with the upstream fp32 letter logits and probabilities. Every image is a synthetic drawing made for this check: colour cards, shapes, text and simple Pong- and Breakout-style frames. check_fixtures.py rebuilds each request from the original image and text, asserts that the 256Γ256 pixels, token ids and answer slots equal upstream's, and compares the probabilities. |Ξp| is the largest absolute probability difference over a slot's options. No upstream top-two gap is within the 1e-4 tie threshold.
| File | Backend / feeding | Image slots: argmax equal | max |Ξp| | p95 |Ξp| | Text-only slots: argmax equal | max |Ξp| |
|---|---|---|---|---|---|---|
| fp16 | CPU / exact-fit prefill | 53/53 | 4.06e-05 | 3.04e-05 | 9/9 | 2.41e-06 |
| fp16-int8vocab | CPU / exact-fit prefill | 52/53 | 0.00846 | 0.00636 | 9/9 | 0.00306 |
| fp16 | Mac GPU (Metal) / one padded chunk per slot | 53/53 | 4.07e-05 | 3.31e-05 | 9/9 | 1.10e-06 |
| fp16-int8vocab | Mac GPU (Metal) / one padded chunk per slot | 53/53 | 0.00621 | 0.00290 | 9/9 | 0.00228 |
The fp16-int8vocab CPU flip is v2_game_breakout_atari_fewbricks question 1, where upstream's top two options are 0.0073 apart. CPU rows are the included reference (ai-edge-litert 2.2.0 CompiledModel, 8 threads); its letter logits are bit-identical to the conversion's own graph readout on all 62 slots of all three files. GPU rows are ai-edge-litert 2.2.0 CompiledModel with GpuOptions(enforce_f32=True). Through LiteRT-LM 0.17.1 on the Mac, all three files on CPU and on GPU gave upstream's answer letter as the first token on 12/12 single-question image rows, with prefill counts equal to upstream's token counts and no Validation error lines.
The price of the fixed 256Γ256 input
Upstream picks a resolution per image; these files take 256Γ256. For 224Γ224, 256Γ240 and 256Γ256 originals, upstream's own processing produces the same pixels, so the probabilities are identical (13 published requests, 21 slots, |Ξp| 0). Every other size gets fewer image tokens than upstream would use: 70 for the 160Γ210 frames and 144 to 475 for the larger synthetic images. On those published requests (24 requests, 32 slots), upstream fp32 on the 256Γ256 input matches upstream fp32 on the original image on 30/32 answers, with max |Ξp| 0.373 and p95 0.318. On photographs and documents kept out of this repository, the largest move was 0.58. Both sides of these comparisons are upstream fp32, so the moves come from the input size, not from the conversion.
Performance
One decision per fresh process on an Apple M4 Max (Mac16,9, 128 GB, macOS 27.0) through LiteRT-LM 0.17.1 (PyPI, Python API): the published Pong frame (game_pong_atari_up, 256Γ256) with its one question, 150 prompt tokens including 64 image tokens, greedy, one output token. TTFT is the wall time from sending the message to the first token, measured in the process. Engine creation is the Engine(...) constructor. Peak RSS and peak footprint come from /usr/bin/time -l; RSS also counts the mapped model and cache files. CPU runs use 8 threads; GPU runs use Backend.GPU() (the WebGPU delegate on Metal) for the decoder and the vision encoder, after a rest of at least 120 s since the previous GPU run. Each value is the median of three runs, with the runs in parentheses. "none" is cache_dir=":nocache"; "yes, warm" runs after a first run that wrote the cache folder. Runs taken while another process used more than 120 % CPU were repeated or dropped and are not in the table. Every run gave the letter A, and the GPU runs printed no Validation error line. The int8 file was not timed on the Mac; its phone rows follow.
| File | Backend | Runtime cache folder | TTFT s | Engine creation s | Prefill tok/s | Peak RSS GB | Peak footprint GB |
|---|---|---|---|---|---|---|---|
| fp16 | CPU | none | 1.68 (1.52, 2.05, 1.68) | 16.8 | 132 | 53.7 | 47.9 |
| fp16 | CPU | yes, warm | 1.72 (1.26, 7.59, 1.72) | 13.7 | 115 | 21.2 | 6.6 |
| fp16 | GPU | none | 0.29 (0.35, 0.28, 0.29) | 28.5 | 826 | 15.6 | 19.3 |
| fp16 | GPU | yes, warm | 0.32 (0.28, 0.32, 0.41) | 29.2 | 811 | 15.8 | 19.6 |
| fp16-int8vocab | CPU | none | 1.54 (1.54, 1.50, 1.65) | 16.3 | 140 | 45.5 | 41.3 |
| fp16-int8vocab | CPU | yes, warm | 4.42 (1.16, 4.42, 4.43) | 13.4 | 49 | 12.5 | 1.5 |
| fp16-int8vocab | GPU | none | 0.26 (0.35, 0.26, 0.26) | 26.8 | 843 | 7.0 | 10.6 |
| fp16-int8vocab | GPU | yes, warm | 0.47 (0.28, 0.66) | 30.2 | 688 | 6.7 | 10.2 |
Median TTFT is 0.26β0.47 s on the GPU and 1.54β1.72 s on the CPU, apart from the fp16-int8vocab warm-folder CPU row explained below. GPU engine creation takes 27β30 s against 13β17 s on the CPU. Without a cache folder, the CPU process's peak footprint is 47.9 GB (fp16) and 41.3 GB (fp16-int8vocab); with a warm folder it is 6.6 GB and 1.5 GB. With a warm folder, CPU TTFT depends on whether the cache file is still in memory: 1.16β1.72 s when it was, 4.4β7.6 s when about 0.46β0.56 million page faults read it back from disk. A warm folder did not shorten GPU engine creation. The fp16-int8vocab GPU warm row has two runs; its third run overlapped another process and was dropped.
The reference readout on the same machine (ai-edge-litert 2.2.0 CompiledModel CPU, 8 threads, its section folder and weight cache already on disk), one process per run. The first decision in a process includes the first use of each graph; the second decision in the same process is the steady cost.
| File | Constructor s | First decision s | Second decision s | Peak RSS GB |
|---|---|---|---|---|
| fp16 | 12.8 | 1.28 (11.69, 1.28, 1.28) | 0.90 | 22.1 |
| fp16-int8vocab | 12.9 | 2.85 (3.89, 2.85, 1.00) | 1.02 | 13.5 |
Galaxy S26 (SM-S942Q, MemTotal 11,389,756 kB), LiteRT-LM v0.16.0 litert_lm_advanced_main, a different runtime version from the Mac rows. One cold run per row (n = 1): the published row game_pong_atari_level (150 prompt tokens, one question), greedy, up to 3 output tokens, runtime default cache folder, thermal status 0 and no frequency cap before the run. Engine creation is the runtime's Init Total; the conversation set-up after it is in parentheses. Peak memory is the process VmHWM.
| File | Backend | Result | TTFT s | Prefill tok/s | Engine creation s | Peak VmHWM GB |
|---|---|---|---|---|---|---|
| int8 | CPU (4 threads) | first token C = upstream; 6/6 rows | 1.22 | 136.68 | 17.7 (+2.0) | 4.54 |
| int8 | GPU (OpenCL) | first token C = upstream; 6/6 rows | 0.56 | 340.76 | 64.4 (+7.0) | 4.32 |
| fp16-int8vocab | GPU (OpenCL) | engine not created: stopped at 6.55 GB VmHWM during OpenCL set-up of prefill_1024 (our 6.5 GB guard) |
β | β | β | β |
| fp16-int8vocab | CPU | not run: its XNNPACK cache alone is 6.0 GB | β | β | β | β |
| fp16 | CPU / GPU | not run: XNNPACK expands its weights to 7.5 GB | β | β | β | β |
On the GPU all 7 decoder signatures were fully delegated to OpenCL (one partition each); the vision encoder ran on the GPU and the adapter on the CPU (the cache files each wrote). VmHWM does not include GPU memory: during the int8 GPU run the phone's MemAvailable fell to 0.77 GB, so the int8 GPU run is close to the limit of a phone of this RAM class. The other five int8 CPU rows peaked at 4.82β4.85 GB VmHWM. Every int8 row on the S26 printed no Validation error line and processed the same number of prefill tokens as upstream.
Limitations
- The runtime path gives the argmax, with one image first; several questions about that image go as turns of one conversation (above), not as one conversation per question. The option probabilities cannot be read through LiteRT-LM: its text-scoring call returns scores that disagree with its own greedy decode on 0.17.1 (LiteRT-LM #3561). Use the reference readout for probabilities and for text-only requests.
- The runtime's own image resize was not exercised: every runtime check handed it a 256Γ256 PNG resized with PIL. The examples here do the same.
- One image per request. English only. The text behaviour is the upstream model's (decider-2b v5 text weights); see the upstream limitations.
- Rows longer than 4096 tokens do not fit the cache. Every published row is at most 225 tokens; prompts above 2048 tokens were not exercised.
- Phones: only the int8 file was tested on a phone (Galaxy S26, one cold run per row). The fp16-int8vocab file did not fit its GPU, and the fp16 file was not tried. Probabilities on the phone were not measured.
Other ports
As of 2026-09-29, the Hugging Face search for decider-2b-vision lists two GGUF ports, mradermacher/decider-2b-vision-GGUF and mindchain/decider-2b-vision-GGUF, and no other LiteRT port.
License and changes
Apache-2.0, as declared by the upstream checkpoint; the vendored prompt code is Apache-2.0 too (LICENSE). Source revision: 863e290863655f1d6b69324d77d09ac972d21609.
Changes: converted the text decoder, the vision encoder and the merger to LiteRT flatbuffers; the decoder derives the upstream M-RoPE positions from a 1-D position inside its graph; fixed the image input at 256Γ256; cast the weights as listed under Files; replaced the chat template with an identity template; repackaged the tokenizer unchanged; added ExecutorMetadata and an fp32 activation preference. The multi-token-prediction head is not included. This community conversion is not an official Mapika or Qwen release.
decider_vendored/prompt.py is the checkpoint's own decider/prompt.py, unchanged (SHA-256 c2fadbe0e4703a381011007f52103ab29142a55a441302d62c1dfb2056acce16). decider_vendored/infer.py keeps only its Q and Example dataclasses. The LICENSE is from github.com/Mapika/decider at commit 23579f7a7e8f10e1045be492af3c1c05a005d67c.
Reproduction (conversion scripts, gates, fixtures): hf-to-litertlm decider2bv_work/.
- Downloads last month
- 47


