parakeet-ultra-onnx-int8 β most accurate Parakeet for offline dictation
Offline transducer bundle of moondream/parakeet-ultra
for sherpa-onnx (nemo_transducer, TDT).
Drop-in replacement for parakeet-tdt-0.6b-v3-int8: same 4-file layout, same ~600 MB,
measurably lower word error rate.
Download
from huggingface_hub import snapshot_download
snapshot_download("mldecode/parakeet-ultra-onnx-int8", local_dir="./parakeet-ultra-onnx-int8")
Or fetch individual files:
https://huggingface.co/mldecode/parakeet-ultra-onnx-int8/resolve/main/<encoder.int8.onnx|decoder.int8.onnx|joiner.int8.onnx|tokens.txt>
Why this model
Parakeet Ultra keeps the full-precision parakeet-tdt-0.6b-v3 architecture (FastConformer
24Γ1024, TDT decoder, 25 European languages) and improves it with further training.
This repo converts those weights to sherpa-onnx int8 so CPU apps (e.g. DictFlow) get the
accuracy without a GPU stack.
Word error rate %, lower is better β upstream evaluations on identical pipelines (model card):
| Benchmark | parakeet-tdt-0.6b-v3 | parakeet-redux (ternary) | ultra (this repo) |
|---|---|---|---|
| English, 7 sets (LibriSpeech, AMI, Earnings-22, GigaSpeech, SPGISpeech, VoxPopuli) | 6.26 | 6.55 | 5.80 |
| FLEURS, 25 languages | 11.62 | 10.56 | 9.55 (β18% vs v3) |
| Business speech (AA-WER style) | 6.15 | 6.96 | 5.79 |
| Background noise (MUSAN, 9 conditions) | 6.72 | 9.04 | 5.82 |
| Long-form TED-LIUM (11 talks, 10β20 min) | 2.71 | 2.51 | 1.94 (β28% vs v3) |
Biggest win over the ternary redux build: background noise (5.82 vs 9.04) β the case
that matters for real-world dictation with room mics and office noise.
Speed / size (CPU)
| Runtime (8Γ x86 cores) | Weights | Throughput | WER (LibriSpeech clean) |
|---|---|---|---|
| Photon + parakeet-redux (ternary kernels) | 178 MB | ~113Γ realtime | 1.94 |
| sherpa-onnx + ultra-int8 (this repo, expected β v3-int8) | ~600 MB | ~40Γ realtime | β |
| sherpa-onnx + v3-int8 (reference) | 670 MB | ~42Γ realtime | 1.97 |
Apple M2 CPU: sherpa int8 β 28Γ realtime. (Single-thread greedy sanity check on a Windows
dev machine: RTF β 0.12, i.e. ~8Γ realtime; multi-threaded sherpa is much faster.)
If you need maximum speed on CPU and dictate in quiet rooms, parakeet-redux + Photon is
faster and smaller; if you need the best transcript on CPU, this is it.
Files
| file | size | notes |
|---|---|---|
encoder.int8.onnx |
~584 MB | FastConformer encoder + projector, dynamic QUInt8 |
decoder.int8.onnx |
~12 MB | 2-layer LSTM prediction network, QInt8 |
joiner.int8.onnx |
~5 MB | TDT joint: 8192 BPE + blank + 5 durations, QInt8 |
tokens.txt |
8 k | BPE vocab + <blk>, token-identical to parakeet-tdt-0.6b-v3 |
feat_dim=128, subsampling_factor=8, normalize_type=per_feature,
vocab_size=8192 (+1 blank). Greedy and modified-beam-search TDT decoding supported.
Note: the encoder emits 640-dim frames (HF layout β projector folded into the encoder), whereas NVIDIA's NeMo export emits 1024-dim. The three files are self-consistent; do not mix halves with the v3 bundle.
Validation
- I/O contract matches
offline-transducer-nemo-model.cc(positional inputs, int32 decoder tokens,[B,1,1,8198]joiner logits with duration head). sherpa_onnx.OfflineRecognizer(model_type="nemo_transducer")greedy decode of the reference EN sample: "Ask not what your country can do for you. Ask what you can do for your country." β exact match.- ONNXβPyTorch cosine similarity: decoder 0.997, joiner 0.999 (int8 noise only).
Use (Python)
import sherpa_onnx, soundfile as sf
rec = sherpa_onnx.OfflineRecognizer.from_transducer(
encoder="encoder.int8.onnx", decoder="decoder.int8.onnx",
joiner="joiner.int8.onnx", tokens="tokens.txt",
model_type="nemo_transducer", feature_dim=128,
decoding_method="greedy_search", num_threads=4)
audio, sr = sf.read("speech.wav", dtype="float32", always_2d=True)
s = rec.create_stream(); s.accept_waveform(sr, audio[:, 0]); rec.decode_stream(s)
print(s.result.text)
Provenance / license
- Source weights: moondream/parakeet-ultra
(retrain of NVIDIA
parakeet-tdt-0.6b-v3), exported from the HF transformers checkpoint withtorch.onnx(opset 17) + ONNX Runtime dynamic quantization. Benchmarks above are upstream figures; int8 parity spot-checked (see Validation). - License: CC-BY-4.0 β attribution to NVIDIA and Moondream.
Model tree for mldecode/parakeet-ultra-onnx-int8
Base model
moondream/parakeet-ultra