Instructions to use WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-NInfer-v3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NInfer
How to use WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-NInfer-v3 with NInfer:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Qwen3.8-Flash-Next GSQ-RCO Q2_0, NInfer v3 artifact
A NInfer v3 artifact of ISTA-DASLab's
Qwen3.8-Flash-Next GSQ-RCO Q2_0
(2.40 bits per weight over the transformer weights) for NInfer-all, the master branch of
iamwavecut/ninfer-all. Qwen3.8-Flash-Next is a
125B-parameter mixture of experts: 48 layers of 512 routed experts, 10 of them active per token,
about 6B active parameters. The GGUF assigns a ggml quantization type to every tensor; this artifact
keeps each tensor's ggml blocks byte for byte (the hyper-connection projections excepted: they are re-quantized from BF16 to Q8_0), and the engine multiplies
the blocks in place. Q2_0 is the release built for speed: it avoids the lookup-table formats.
Stock NInfer builds refuse these files: the model family and the gguf_* formats exist only in that
line; build master from commit
5395ec30b on. NInfer serves it
with up to eight concurrent requests, prompt-prefix reuse, structured output and its Vision tower;
This file includes the MTP block. Enable drafting with --spec mtp --draft-tokens 4.
What is inside
| component | representation |
|---|---|
| routed experts, 48 layers × 512 | Q2_0 gate, up and down banks, as in the GGUF; expert-major, so one expert is one contiguous range of bytes |
| shared experts | the GGUF's own blocks, one type per tensor: Q2_0, Q3_K, Q4_0, Q4_K, Q5_0, Q5_K, Q6_K, Q8_0, IQ4_XS, IQ4_NL |
| 36 Gated DeltaNet and 12 sparse-attention (QSA) layers | the GGUF's own blocks, Q2_0 to Q6_K per tensor |
| output head, token table | Q5_K, Q3_K |
| router | BF16, as in the GGUF |
| hyper-connection projections | ggml Q8_0: the GGUF's BF16 matrices re-quantized (October 9, 2026), the MTP block's as stored in its GGUF |
GDN controls, norms, convolution, A_log, dt_bias |
BF16/FP32, restored from llama.cpp's exporter conventions (grouped value heads, w instead of 1 + w, A_log from -exp(A_log), the head-interleaved query and gate) |
| Vision tower | the release's mmproj-Qwen3.8-Flash-Next-BF16.gguf in BF16 (0.9 GB): the Qwen3.5/3.6 tower, 27 blocks of width 1152, merging 2×2 patches onto the text model's 2,560, with the checkpoint's image and video preprocessor configurations |
| n-gram table | not in this file: the model's ngram component names its table (the hash constants, IQ4_NL rows, SHA-256 dd55c289…6cc3 of the rows), the same table as every Flash-Next release, published once as WaveCut/Qwen3.8-Flash-Next-ngram-table-NInfer-v3 |
| chat template | Qwen/Qwen3.8-Flash-Next's chat_template.jinja |
One file, Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-ninfer-v3.ninfer, of 40,709,159,936 bytes
(37.91 GiB), next to its conversion report, SHA256SUMS, NOTICE and LICENSE. Original text/Vision conversion
command, before the MTP attachment:
python3 -m tools.convert --model Qwen3.8-Flash-Next \
--recipe qwen3_8_flash_next_gguf --components text,vision \
--source gguf=Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-00001-of-00002.gguf \
--source ngram=Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-00002-of-00002.gguf \
--source vision=mmproj-Qwen3.8-Flash-Next-BF16.gguf \
--device cpu --rows-per-chunk 65536 \
--name qwen3.8-flash-next --out Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-ninfer-v3.ninfer
The converter reads the table shard to record its digest; the rows themselves are in the table
repository. --model needs the configuration, tokenizer and preprocessor files of
Qwen/Qwen3.8-Flash-Next (config.json,
tokenizer.json, tokenizer_config.json, chat_template.jinja, generation_config.json,
preprocessor_config.json, video_preprocessor_config.json).
Running
Download the artifact and the n-gram table, then serve them from a build of master
(its README; CMAKE_CUDA_ARCHITECTURES is
86, 89 or 120a, and 120a needs CUDA 13.1 or newer):
hf download WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-NInfer-v3 \
Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-ninfer-v3.ninfer --local-dir models
hf download WaveCut/Qwen3.8-Flash-Next-ngram-table-NInfer-v3 \
Qwen3.8-Flash-Next-ngram-table-IQ4_NL-ninfer-v3.ninfer --local-dir models
M=models/Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-ninfer-v3.ninfer
T=models/Qwen3.8-Flash-Next-ngram-table-IQ4_NL-ninfer-v3.ninfer
# One GPU of 48 GB or more (L40S, RTX PRO 6000): every expert in device memory.
ninfer-serve $M --ngram-table $T --model-id qwen3.8-flash-next --max-context 32768
# Two 24 GB GPUs (Linux): every expert in device memory, one pipeline stage per GPU.
ninfer-serve $M --ngram-table $T --model-id qwen3.8-flash-next --devices 0,1 --max-context 32768
# One 24 GB GPU: the experts in page-locked host memory (34 GB), the most used of them cached on the GPU.
ninfer-serve $M --ngram-table $T --model-id qwen3.8-flash-next --expert-residency host --max-context 32768
# One 24 GB GPU and little RAM: the experts stay in the file and stream into a GPU cache.
ninfer-serve $M --ngram-table $T --model-id qwen3.8-flash-next --expert-residency disk --max-context 32768
The Docker image (an NVIDIA driver of the CUDA
13 branch, 580 or newer, and the NVIDIA Container Toolkit) runs the same commands as serve with
the files under /models, and answers on http://localhost:8080/v1:
docker run --rm --gpus all -p 8080:8080 --ulimit memlock=-1 \
-v "$PWD/models:/models" -v ninfer-cache:/cache \
ghcr.io/iamwavecut/ninfer-all serve /models/Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-ninfer-v3.ninfer \
--ngram-table /models/Qwen3.8-Flash-Next-ngram-table-IQ4_NL-ninfer-v3.ninfer \
--model-id qwen3.8-flash-next --expert-residency host --max-context 32768
With host experts the GPU holds the 3.1 GB of dense weights, the request state and an expert cache
that takes what is free after startup (--expert-cache-mib sets it). With disk experts the host keeps
no copy of the banks: the cache reads the experts each layer routes to from the file, through the OS
page cache. The n-gram rows are read from the table's file the same way (16 per token) unless
--ngram-ram loads the 28.8 GB table into RAM. A model started without its table is refused;
--no-ngram-table runs it without, a non-standard experimental mode with no practical use (see the
table's card). --vision adds the Vision tower (0.9 GB on the GPU) for images and video. The KV cache of the 12
sparse-attention layers is BF16.
Qwen3.8-Flash-Next
describes the placements and the execution.
October 9, 2026: Q8_0 hyper-connection projections
The hyper-connection projections were the largest read of a decoded token, 1.27 GB in BF16. This
update stores every one of them as ggml Q8_0 (32 inputs of a row share a binary16 scale): the text
model's, BF16 in the GGUF, quantized with ggml's reference rounding, and the MTP block's as the
Unsloth MTP GGUF stores them (this file had decoded them to BF16). The file is 618,159,104 bytes
smaller; every other object is byte for byte the previous one, which the rewrite checked on
readback. hc-q8-requantization.json records the transformation and its error: over the text
matrices the values differ from the BF16 ones by at most 0.0334 (4.79e-4 RMS). Builds before
commit 944811b5c refuse this file.
Perplexity over the repository's quick corpora (4,096-token windows advancing by 2,048, int8 KV, host experts) moved from 3.93783 to 3.93842 on perplexity-1m (+0.015%) and from 4.69573 to 4.69673 on perplexity-heldout-2026-09 (+0.021%); domain by domain from -0.11% to +0.36%. The AIME 2025 and GPQA-Diamond results below were measured with the BF16 projections and were not rerun.
On one RTX 3090 with host experts, int8 KV and 256 greedy tokens per answer (three prompts, each run twice), plain decode went from 81.3-93.6 to 84.5-97.7 tokens/s (+3.6% to +6.0%) and decode with three MTP drafts from 84.5-114.4 to 83.0-120.4 tokens/s (+0.4% to +10.0% per prompt on average); a 4,958-token prompt prefilled at 1,268 instead of 1,234 tokens/s. The dense weights take 2.83 instead of 3.38 GiB of the GPU, which the host expert cache uses. Calls of nine or more tokens a step (prompts, batched verification) spend up to 25 µs more on each hyper-connection read. The speed table below predates this update.
The minimum build above also fixes host experts' prompts: from commit 5118e069d (October 9) on, a
prompt call of 256 tokens or more with --expert-residency host could multiply some layers by
other layers' experts. Over 40 KB of WikiText the Q2_0 model read a perplexity of about 4.0
instead of 2.54 with host experts; disk and device experts were not affected.
MTP attachment and qualification
The October 8, 2026 update adds 28 MTP objects (2,791,415,296 bytes) and 1,567 bindings. Every previous text weight, Vision weight, tokenizer resource and IQ4 table descriptor passed byte-preservation checks. The separate 28.80 GB IQ4 table is unchanged.
MTP comes from the Unsloth shared-Q8_0 release, through the pinned donor recorded in mtp-attachment.json. The original text/Vision conversion report remains in the repository. The base quantization labels describe the text model, not the added MTP block.
The public Engine passed host-to-GPU and disk-to-GPU MTP requests on one RTX 3090, with Vision enabled, int8 KV, 4,096 context capacity, 128-token prefill chunks and an 8 GiB device expert cache. All expert arithmetic ran on the GPU. Each mode used two fixed prompts, three greedy repetitions and 64 generated tokens per request, with prefix reuse disabled. Each fixed mode repeated exactly and accepted MTP drafts. Image requests identified the red test image with thinking disabled. Flash-Next disables MTP for media requests and uses plain decoding. A fresh text request resumed MTP after each image check. An initial check incorrectly required MTP for media and failed. Source inspection confirmed this existing limitation; the engine behavior was not changed.
Samples include the first request. Cache contents are not reset between repetitions. Other artifacts were being prepared on the same host during this campaign, so these timings do not isolate the performance effect of MTP.
The measurements below cover only those short requests. The OS page cache was not controlled. They do not establish cold-disk speed, low-RAM operation or RTX 3080 support. Different draft widths or modes can produce different tokens because their arithmetic rounds differently. Prior quality results below were measured without this MTP attachment.
Prompt 1 requests a Python merge function after a repeated prose prefix. Prompt 2 requests a simple explanation of the blue sky. Each row summarizes the three repetitions of that prompt.
| prompt | expert placement | draft tokens | minimum draft probability | median decode tok/s | range tok/s |
|---|---|---|---|---|---|
| 1 | host | 4 | 0 | 38.43 | 25.90–41.98 |
| 2 | host | 4 | 0 | 32.00 | 26.02–35.37 |
| 1 | disk | 4 | 0 | 13.64 | 10.61–13.82 |
| 2 | disk | 4 | 0 | 11.72 | 9.88–11.88 |
| 1 | host | 0 | 0 | 50.68 | 31.28–55.58 |
| 2 | host | 0 | 0 | 45.30 | 35.36–49.92 |
| 1 | host | 4 | 0.3 | 38.20 | 26.59–42.29 |
| 2 | host | 4 | 0.3 | 31.81 | 26.35–35.37 |
The probability-floor experiment uses the tested development snapshot. The normal MTP command above uses the default floor. mtp-qualification.json records every sample and the tested source snapshot; the minimum commit above identifies artifact support.
Quality
NInfer computes what llama.cpp computes from the same GGUF. Over the first 72 KB of the WikiText
sample in the repository's eval/corpora/perplexity-1m, in 2,560-token windows that advance by 1,024
targets (llama.cpp's --ppl-stride 1024 -c 2048), the fourteen windows both programs score the same
way (14,336 targets) give:
| perplexity | |
|---|---|
| this artifact in NInfer, BF16 KV | 2.6579 |
| the GGUF in llama.cpp, BF16 KV | 2.6502 |
Window by window the two differ by -0.014 to +0.016 nats. Deep into a long context they agree as
closely: over four PG-19 streams of the same corpus joined (261,412 tokens), in 65,536-token windows
that advance by 32,768 targets (--ppl-stride 32768 -c 49152), so that every scored position lies
past the first 32,768 of its window, the five windows both score the same way give 8.0834 in NInfer
and 8.0802 in llama.cpp, window by window within 0.0011 nats (two RTX 5090s, BF16 KV).
The GSQ-RCO card evaluates this quantization in llama.cpp at AIME25 96.67, GPQA-Diamond 89.39 and LiveCodeBench v6 81.14 (BF16: 100.00, 91.92, 87.43), without stating its sampling, output limit or number of runs. In NInfer, EvalScope ran AIME 2025 and GPQA-Diamond against this artifact with its table on two RTX 5090s, every expert in device memory and six requests at a time: thinking on, temperature 1.0, top-p 0.95, top-k 20, one sampled run, at most 106,000 output tokens (October 2026). The answers cut there were then continued from where they stopped to 122,880 (AIME) and 245,760 (GPQA) output tokens:
| AIME 2025 | GPQA-Diamond | |
|---|---|---|
| this artifact in NInfer, the cut answers continued | 93.33 (28/30) | 86.36 (171/198) |
| this artifact in NInfer, at most 106,000 output tokens | 93.33 | 84.34 |
| the GGUF in llama.cpp (GSQ-RCO card) | 96.67 | 89.39 |
One AIME answer and ten GPQA answers reached 106,000 tokens without answering. Continued, four of the GPQA answers came out right and four wrong; one GPQA answer was still reasoning at 245,760 tokens, one was repeating a codon of its question's DNA sequence, and the AIME answer was still calculating at 122,880. One run of GPQA-Diamond's 198 questions has a standard error of about 2.5 points. The answers are long, 23,000 to 26,000 output tokens on average; the run took 7 h 44 min at 193 tokens per second across the six requests, and the continuations 58 min more. LiveCodeBench needs a code sandbox and was not run. Measurements has the details.
Speed
October 2026, greedy decoding, the repository's generate test (ninfer_qwen4_exp_generate_real).
Decode is measured over the 78 tokens of a short answer to a 20-token prompt and over the first five
tokens after a 4,463-token prompt; prefill is that prompt in 512-token chunks. Host memory is the
process's peak resident set. These are the weights of this file (the table and the Vision tower are
not read here).
The RTX PRO 6000 and RTX 5090 rows ran a 120a build with CUDA 13.1 and the separate table
artifact, single-GPU rows pinned to the GPU's NUMA node:
- RTX PRO 6000 Workstation Edition: 600 W, a PCIe 4.0 x16 host link.
- RTX 5090s: boards with a 600 W default limit, PCIe 5.0 x16.
With host or disk experts the device expert cache takes the free memory, about 94 GB on the PRO 6000. The rows measure it while it is still filling.
The RTX 4090 rows ran this file with the separate table artifact on RTX 4090s at 450 W (PCIe 4.0 x16) in a cloud VM with two EPYC 7543 sockets, pinned to the CPUs of the GPUs' NUMA node, two runs each where a range is given; left unpinned, host experts decode there at 40 tok/s and disk experts at 32 to 33. The RTX 3090 Ti and first three RTX 3090 rows were measured before the weights moved into one file, the last in a single file with the table, on another RTX 3090 (310 W power limit, 62 GB of RAM, 23 cores of an AMD EPYC 7663) next to the IQ3_S artifact.
| hardware and placement | device memory | host memory | decode, short answer | decode after 4,463 tokens | prefill |
|---|---|---|---|---|---|
| RTX PRO 6000, experts on the GPU | 37.7 GiB | 0.8 GiB | 136.0-139.7 tok/s | 160.7-162.6 tok/s | 2,925-2,952 tok/s |
| RTX PRO 6000, experts in pinned host memory | 38.5 GiB | 34.0 GB pinned | 54.8-55.0 tok/s | 67.7-70.0 tok/s | 1,296-1,297 tok/s |
| RTX PRO 6000, experts on disk, artifact in the page cache | 93.6 GiB | 1.0 GiB + page cache | 75.9 tok/s | 115.8 tok/s | 2,024 tok/s |
| RTX PRO 6000, experts on disk, the files' pages evicted every second | 93.5 GiB | 1.0 GiB | 35.1 tok/s | 87.1 tok/s | 1,000 tok/s |
| 2× RTX 5090 (PCIe), experts on the GPUs | 18.8 + 19.6 GiB | 0.9 GiB | 136.1-136.3 tok/s | 157.0-157.2 tok/s | 3,807-3,811 tok/s |
| RTX 5090, experts in pinned host memory | 30.0 GiB | 34.0 GB pinned | 74.1-74.2 tok/s | 93.9-94.2 tok/s | 1,676-1,678 tok/s |
| RTX 5090, experts on disk, artifact in the page cache | 29.9 GiB | 1.0 GiB + page cache | 67.5 tok/s | 76.7 tok/s | 1,739 tok/s |
| RTX 5090, experts on disk, the files' pages evicted every second | 29.9 GiB | 1.0 GiB | 25.3 tok/s | 28.7 tok/s | 694 tok/s |
| 2× RTX 4090 (PCIe, no P2P), experts on the GPUs | 18.5 + 19.4 GiB | 0.8 GiB | 92.6-100.7 tok/s | 116-120 tok/s | 3,482-3,734 tok/s |
| 2× RTX 3090 Ti (PCIe, no P2P), experts on the GPUs | 18.4 + 19.4 GB | 0.7 GB | 90.2 tok/s | 107 tok/s | 1,504 tok/s |
| RTX 4090, experts in pinned host memory | 20.7 GiB | 34.0 GB pinned | 52.6-52.7 tok/s | 44.3-44.6 tok/s | 1,137-1,139 tok/s |
| RTX 4090, experts on disk, artifact in the page cache | 20.7 GiB | 1.0 GiB + page cache | 42.0 tok/s | 45.6 tok/s | 723 tok/s |
| RTX 4090, experts on disk, the file's pages evicted every second | 20.5 GiB | 1.0 GiB | 17.7 tok/s | 12.2 tok/s | 160 tok/s |
| RTX 3090, experts in pinned host memory | 22.6 GB | 34 GB pinned | 48.9 tok/s | 42.1 tok/s | 842 tok/s |
| RTX 3090, experts on disk, artifact in the page cache | 22.6 GB | 0.9 GB + page cache | 47.0 tok/s | 39.1 tok/s | 630 tok/s |
| RTX 3090, experts on disk, page cache dropped every second (NVMe) | 22.6 GB | 0.9 GB | 17.2 tok/s | 11.3 tok/s | 195 tok/s |
| RTX 3090, experts in pinned host memory, this file | 22.1 GiB | 34.0 GB pinned | 50.2 tok/s | 43.7 tok/s | 847 tok/s |
llama.cpp runs the GGUF with the experts on the CPU (--n-cpu-moe 48) at 11.2 tok/s decode (tg128)
and 267 tok/s prefill (pp512) on the first RTX 3090's machine (32 threads), and at 12.7 and 368 tok/s
on the second (23 threads).
Credits and license
Quantized weights: ISTA-DASLab,
Qwen3.8-Flash-Next GSQ-RCO GGUFs,
produced with GSQ and RCO
from Qwen/Qwen3.8-Flash-Next. The model is
released under the Qwen Community License 1.0, reproduced in LICENSE, and its conditions
apply to these files; ISTA-DASLab publish their GGUFs under Apache-2.0. Attribution notices are
collected in NOTICE. This repository only re-packs those weights into NInfer's container.
- Downloads last month
- 244
Model tree for WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-NInfer-v3
Base model
Qwen/Qwen3.8-Flash-Next