Instructions to use litert-community/Nemotron-3-Nano-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT-LM
How to use litert-community/Nemotron-3-Nano-4B with LiteRT-LM:
# LiteRT-LM runs on various platforms (Android, iOS, Windows, Linux, macOS, IoT, Web/WASM) # and supports many APIs (C++, Python, Kotlin, Swift, JavaScript, Flutter). # For platform-specific integration guides, please refer to the official developer website: # https://ai.google.dev/edge/litert-lm # To try LiteRT-LM, the easiest way is to use our CLI tool. # 1. Install the LiteRT-LM CLI tool: pip install -U litert-lm # 2. Download and run this model locally: # See: https://ai.google.dev/edge/litert-lm/cli litert-lm run \ --from-huggingface-repo=litert-community/Nemotron-3-Nano-4B \ --prompt="Write me a poem"
- LiteRT
How to use litert-community/Nemotron-3-Nano-4B with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Measured on device (edge-compat): Galaxy S26 · LiteRT-LM 0.16.0 · GPU · did not run: run_error (2026-09-05); Mac Studio M4 Max · LiteRT-LM 0.16.0 · GPU · decode 83.3 tok/s · prefill 803 tok/s · TTFT 331 ms (2026-08-27); Galaxy S26 · LiteRT-LM 0.16.0 · CPU · decode 12.9 tok/s · prefill 53 tok/s · TTFT 4.06 s (2026-09-05). Record: https://github.com/john-rocky/edge-compat/blob/main/cards/nemotron-3-nano-4b-int8/CARD.md
Nemotron-3-Nano-4B — LiteRT-LM
nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16 converted to the LiteRT-LM (.litertlm) format for on-device inference with Google's LiteRT-LM runtime. Requires litert-lm ≥ 0.15. To our knowledge this is the first Nemotron-3-Nano in LiteRT form.
A reasoning model (<think>, ChatML turns) on a three-kind hybrid stack: 21 Mamba2 selective-scan layers + 17 plain MLP layers + 4 grouped-query attention layers (42 total). Only the 4 attention layers keep KV (4096-token budget here), the mamba layers carry constant-size conv + SSM state, and the MLP layers carry no state at all — 50 state buffers in total (42 mamba, 8 KV), so memory stays nearly flat with context length.
| File | Recipe | Size |
|---|---|---|
Nemotron-3-Nano-4B_int8.litertlm |
int8 dynamic on linears + embedding (convs and the scan stay float); fp32 activations declared | 4.13 GB |
2026-09-21: chat template updated to accept the 0.18 content-parts form (string form unchanged); weights, tokenizer and executor metadata byte-identical.
Geometry: hidden 3136, 40 query / 8 KV heads, mamba 96 heads × 80 dim (state 128, conv 4, 8 groups), vocab 131,072, untied embeddings.
Correctness
All numbers below were measured on this exact file (litert-lm 0.16.0, Apple M4 Max).
- 8-question sanity gate: 7/8 on CPU, 8/8 on GPU, non-degenerate on both. The single CPU miss is the rhyme-completion item — it answers "violets are purple" where the gate wants "blue"; the GPU run answers "blue". Every arithmetic, factual, and translation item is correct on both backends.
- Chat template is byte-equal to the source: the embedded Jinja matches the repo's
chat_template.jinjaexactly (10,504 / 10,504 bytes). Note the source repo'stokenizer_config.jsoncarries a different 10,497-byte copy; the bundle embeds the oneAutoTokenizeractually resolves. - Turn-end stop tokens are
<|im_end|>(id 11) alongside the exported id 2. - No spurious start token. The source tokenizer sets
add_bos_token: Falseand the template never renders a leading BOS, so the<s>the bundler would otherwise prepend is dropped — the on-device token stream matches the training stream. Honest note: at this scale the model is robust either way (the gate scores 7/8 with the token and 7/8 without, and greedy decoding in PyTorch is byte-identical on 2 of 3 probes), so this is a correctness-of-convention fix rather than a rescue.
Usage
# CPU
litert-lm run ./Nemotron-3-Nano-4B_int8.litertlm --prompt "What is the capital of France? Answer in one word."
# GPU — pass --cache no (see the honest note below)
litert-lm run ./Nemotron-3-Nano-4B_int8.litertlm --backend gpu --cache no --prompt "..."
The bundle carries the tokenizer and the stock Nemotron-3-Nano chat template. Seven prefill signatures (1024/256/64/16/4/1 + decode) are exported so the runtime picks tight chunks.
Performance
litert-lm benchmark <file> -p 256 -d 256 --runs 3 --cache no, litert-lm 0.16.0, Apple M4 Max. Two independent runs:
| Backend | Prefill (256) | Decode | TTFT |
|---|---|---|---|
GPU (--cache no) |
803 / 792 tok/s | 83.3 / 82.4 tok/s | 0.33 s |
| CPU | 99.6 / 113.3 tok/s | 22.7 / 22.5 tok/s | 2.64 / 2.30 s |
Both figures per cell are the two runs, not a range estimate. GPU repeats within ~1.4%; CPU prefill spreads ~13%, partly because the host was not idle during these runs (another export was using the machine) — read the CPU column as an order of magnitude, not a precise figure. --cache no matters for more than tidiness here: with the compiled-graph cache the benchmark reports a much faster CPU prefill because it is not doing the same work.
Honest notes
- GPU requires
--cache noon this bundle. With the compiled-graph cache enabled,litert-lm run --backend gpufails with WebGPUInvalid BindGroupvalidation errors, and an 8-question sweep through the Mac verify harness returns token soup (0/8). The same file with--cache noanswers 8/8. Measured as a one-variable comparison — same runner, same file, cache flag flipped — so the cache path is where it goes wrong; the root cause is not isolated further here. - Not measured on a phone yet. The desktop numbers above are Mac-only. A 4B of this shape did not fit an 8 GB Android phone when the sibling Nemotron-H-4B was measured, so expect to need a higher-RAM device; that is an expectation carried over from a different bundle, not a measurement of this one.
- It is a reasoning model. Answers arrive after a
<think>block, so give it a token budget that fits the thought (the gate above used 3200). - int8 is applied to linears and the embedding only; the convolutions and the selective scan stay float, which is what keeps the hybrid state numerically sane.
Conversion notes
Converted with litert-torch plus a hybrid-cache patch. One command, no per-model work — the reproduction script, the patch, and the full measurement record are in hf-to-litertlm:
python scripts/convert.py nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16
Two things that route this model correctly and are worth knowing if you convert your own:
auto_mapin a config is not proof of remote code. This repo declaresauto_map, but transformers registersnemotron_hnatively, so withouttrust_remote_codethe library implementation loads and the repo's Python is never imported. A converter that refuses onauto_mapalone will refuse this model for no reason.- ≥3B exports use a reduced 7-signature prefill ladder. Every exported signature costs engine RAM whether or not it is called, and a 4B hybrid with the full 11-signature ladder is exactly the shape that trips memory limits at GPU program init.
Conversion took 1645 s on an M4 Max. See REPRODUCE.md for the Nemotron-H family recipe and the measurements behind every claim on this card.
2026-08-31 — thought channel declared (metadata only, weights unchanged)
Nemotron-3-Nano-4B_int8.litertlm now declares the reasoning channel in its metadata (LlmMetadata.channels: channel name thought, markers <think>…</think> exactly as this model emits them). Without the declaration the runtime has no way to tell the reasoning apart from the answer: the raw thinking streamed inline into the visible text, and a thinking_token_budget was silently ignored (the API returns OK and only logs a warning). With the channel declared, LiteRT-LM returns the reasoning separated in channels["thought"] and the thinking budget takes effect.
Metadata-only change: every section of the bundle except LlmMetadata is byte-identical to the previous file (verified per section, tokenizer included), so the weights, the graph, the tokenizer and the chat template are unchanged and the speed and accuracy numbers on this card still describe this file — only the file's own sha256 differs. Verified on the LiteRT-LM runtime (litert-lm-api 0.16.1): the visible answer stays clean, the reasoning lands in channels["thought"], and on one bundle of this batch thinking_token_budget=16 was confirmed to truncate the reasoning at exactly 16 tokens where it was a no-op before. If you downloaded before 2026-08-31, re-download to get the channel-aware file.
Raspberry Pi 5 (CPU)
Measured on a Raspberry Pi 5 Model B Rev 1.1 (8 GB, Raspberry Pi OS 64-bit) with litert-lm benchmark 0.16.1: CPU backend, 4 threads, 256 prefill + 256 decode tokens, --cache memory (the compile cache lives and dies with the process, so every invocation compiles the model from scratch; nothing is reused between runs), one warm-up plus one timed iteration per invocation, 3 invocations per file with cooldown in between. Values are the median across invocations (min–max in parentheses). No thermal throttling occurred during these runs (vcgencmd get_throttled stayed 0x0). Every file listed produced coherent text in a real generation on this backend before its numbers were recorded.
| File | Prefill (tok/s) | Decode (tok/s) | TTFT | Peak RSS |
|---|---|---|---|---|
Nemotron-3-Nano-4B_int8.litertlm |
22.7 (22.6–22.8) | 2.2 (2.2–2.2) | 11.7 s | 5.0 GB |
- Downloads last month
- 273
Model tree for litert-community/Nemotron-3-Nano-4B
Base model
nvidia/NVIDIA-Nemotron-Nano-12B-v2-Base