Qwen3.8-Flash-Next GSQ-RCO Q2_0, NInfer v3 artifact

A NInfer v3 artifact of ISTA-DASLab's Qwen3.8-Flash-Next GSQ-RCO Q2_0 (2.40 bits per weight over the transformer weights) for NInfer-all, the master branch of iamwavecut/ninfer-all. Qwen3.8-Flash-Next is a 125B-parameter mixture of experts: 48 layers of 512 routed experts, 10 of them active per token, about 6B active parameters. The GGUF assigns a ggml quantization type to every tensor; this artifact keeps each tensor's ggml blocks byte for byte (the hyper-connection projections excepted: they are re-quantized from BF16 to Q8_0), and the engine multiplies the blocks in place. Q2_0 is the release built for speed: it avoids the lookup-table formats.

Stock NInfer builds refuse these files: the model family and the gguf_* formats exist only in that line; build master from commit 5395ec30b on. NInfer serves it with up to eight concurrent requests, prompt-prefix reuse, structured output and its Vision tower; This file includes the MTP block. Enable drafting with --spec mtp --draft-tokens 4.

What is inside

component representation
routed experts, 48 layers × 512 Q2_0 gate, up and down banks, as in the GGUF; expert-major, so one expert is one contiguous range of bytes
shared experts the GGUF's own blocks, one type per tensor: Q2_0, Q3_K, Q4_0, Q4_K, Q5_0, Q5_K, Q6_K, Q8_0, IQ4_XS, IQ4_NL
36 Gated DeltaNet and 12 sparse-attention (QSA) layers the GGUF's own blocks, Q2_0 to Q6_K per tensor
output head, token table Q5_K, Q3_K
router BF16, as in the GGUF
hyper-connection projections ggml Q8_0: the GGUF's BF16 matrices re-quantized (October 9, 2026), the MTP block's as stored in its GGUF
GDN controls, norms, convolution, A_log, dt_bias BF16/FP32, restored from llama.cpp's exporter conventions (grouped value heads, w instead of 1 + w, A_log from -exp(A_log), the head-interleaved query and gate)
Vision tower the release's mmproj-Qwen3.8-Flash-Next-BF16.gguf in BF16 (0.9 GB): the Qwen3.5/3.6 tower, 27 blocks of width 1152, merging 2×2 patches onto the text model's 2,560, with the checkpoint's image and video preprocessor configurations
n-gram table not in this file: the model's ngram component names its table (the hash constants, IQ4_NL rows, SHA-256 dd55c289…6cc3 of the rows), the same table as every Flash-Next release, published once as WaveCut/Qwen3.8-Flash-Next-ngram-table-NInfer-v3
chat template Qwen/Qwen3.8-Flash-Next's chat_template.jinja

One file, Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-ninfer-v3.ninfer, of 40,709,159,936 bytes (37.91 GiB), next to its conversion report, SHA256SUMS, NOTICE and LICENSE. Original text/Vision conversion command, before the MTP attachment:

python3 -m tools.convert --model Qwen3.8-Flash-Next \
  --recipe qwen3_8_flash_next_gguf --components text,vision \
  --source gguf=Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-00001-of-00002.gguf \
  --source ngram=Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-00002-of-00002.gguf \
  --source vision=mmproj-Qwen3.8-Flash-Next-BF16.gguf \
  --device cpu --rows-per-chunk 65536 \
  --name qwen3.8-flash-next --out Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-ninfer-v3.ninfer

The converter reads the table shard to record its digest; the rows themselves are in the table repository. --model needs the configuration, tokenizer and preprocessor files of Qwen/Qwen3.8-Flash-Next (config.json, tokenizer.json, tokenizer_config.json, chat_template.jinja, generation_config.json, preprocessor_config.json, video_preprocessor_config.json).

Running

Download the artifact and the n-gram table, then serve them from a build of master (its README; CMAKE_CUDA_ARCHITECTURES is 86, 89 or 120a, and 120a needs CUDA 13.1 or newer):

hf download WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-NInfer-v3 \
  Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-ninfer-v3.ninfer --local-dir models
hf download WaveCut/Qwen3.8-Flash-Next-ngram-table-NInfer-v3 \
  Qwen3.8-Flash-Next-ngram-table-IQ4_NL-ninfer-v3.ninfer --local-dir models
M=models/Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-ninfer-v3.ninfer
T=models/Qwen3.8-Flash-Next-ngram-table-IQ4_NL-ninfer-v3.ninfer

# One GPU of 48 GB or more (L40S, RTX PRO 6000): every expert in device memory.
ninfer-serve $M --ngram-table $T --model-id qwen3.8-flash-next --max-context 32768

# Two 24 GB GPUs (Linux): every expert in device memory, one pipeline stage per GPU.
ninfer-serve $M --ngram-table $T --model-id qwen3.8-flash-next --devices 0,1 --max-context 32768

# One 24 GB GPU: the experts in page-locked host memory (34 GB), the most used of them cached on the GPU.
ninfer-serve $M --ngram-table $T --model-id qwen3.8-flash-next --expert-residency host --max-context 32768

# One 24 GB GPU and little RAM: the experts stay in the file and stream into a GPU cache.
ninfer-serve $M --ngram-table $T --model-id qwen3.8-flash-next --expert-residency disk --max-context 32768

The Docker image (an NVIDIA driver of the CUDA 13 branch, 580 or newer, and the NVIDIA Container Toolkit) runs the same commands as serve with the files under /models, and answers on http://localhost:8080/v1:

docker run --rm --gpus all -p 8080:8080 --ulimit memlock=-1 \
  -v "$PWD/models:/models" -v ninfer-cache:/cache \
  ghcr.io/iamwavecut/ninfer-all serve /models/Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-ninfer-v3.ninfer \
  --ngram-table /models/Qwen3.8-Flash-Next-ngram-table-IQ4_NL-ninfer-v3.ninfer \
  --model-id qwen3.8-flash-next --expert-residency host --max-context 32768

With host experts the GPU holds the 3.1 GB of dense weights, the request state and an expert cache that takes what is free after startup (--expert-cache-mib sets it). With disk experts the host keeps no copy of the banks: the cache reads the experts each layer routes to from the file, through the OS page cache. The n-gram rows are read from the table's file the same way (16 per token) unless --ngram-ram loads the 28.8 GB table into RAM. A model started without its table is refused; --no-ngram-table runs it without, a non-standard experimental mode with no practical use (see the table's card). --vision adds the Vision tower (0.9 GB on the GPU) for images and video. The KV cache of the 12 sparse-attention layers is BF16. Qwen3.8-Flash-Next describes the placements and the execution.

October 9, 2026: Q8_0 hyper-connection projections

The hyper-connection projections were the largest read of a decoded token, 1.27 GB in BF16. This update stores every one of them as ggml Q8_0 (32 inputs of a row share a binary16 scale): the text model's, BF16 in the GGUF, quantized with ggml's reference rounding, and the MTP block's as the Unsloth MTP GGUF stores them (this file had decoded them to BF16). The file is 618,159,104 bytes smaller; every other object is byte for byte the previous one, which the rewrite checked on readback. hc-q8-requantization.json records the transformation and its error: over the text matrices the values differ from the BF16 ones by at most 0.0334 (4.79e-4 RMS). Builds before commit 944811b5c refuse this file.

Perplexity over the repository's quick corpora (4,096-token windows advancing by 2,048, int8 KV, host experts) moved from 3.93783 to 3.93842 on perplexity-1m (+0.015%) and from 4.69573 to 4.69673 on perplexity-heldout-2026-09 (+0.021%); domain by domain from -0.11% to +0.36%. The AIME 2025 and GPQA-Diamond results below were measured with the BF16 projections and were not rerun.

On one RTX 3090 with host experts, int8 KV and 256 greedy tokens per answer (three prompts, each run twice), plain decode went from 81.3-93.6 to 84.5-97.7 tokens/s (+3.6% to +6.0%) and decode with three MTP drafts from 84.5-114.4 to 83.0-120.4 tokens/s (+0.4% to +10.0% per prompt on average); a 4,958-token prompt prefilled at 1,268 instead of 1,234 tokens/s. The dense weights take 2.83 instead of 3.38 GiB of the GPU, which the host expert cache uses. Calls of nine or more tokens a step (prompts, batched verification) spend up to 25 µs more on each hyper-connection read. The speed table below predates this update.

The minimum build above also fixes host experts' prompts: from commit 5118e069d (October 9) on, a prompt call of 256 tokens or more with --expert-residency host could multiply some layers by other layers' experts. Over 40 KB of WikiText the Q2_0 model read a perplexity of about 4.0 instead of 2.54 with host experts; disk and device experts were not affected.

MTP attachment and qualification

The October 8, 2026 update adds 28 MTP objects (2,791,415,296 bytes) and 1,567 bindings. Every previous text weight, Vision weight, tokenizer resource and IQ4 table descriptor passed byte-preservation checks. The separate 28.80 GB IQ4 table is unchanged.

MTP comes from the Unsloth shared-Q8_0 release, through the pinned donor recorded in mtp-attachment.json. The original text/Vision conversion report remains in the repository. The base quantization labels describe the text model, not the added MTP block.

The public Engine passed host-to-GPU and disk-to-GPU MTP requests on one RTX 3090, with Vision enabled, int8 KV, 4,096 context capacity, 128-token prefill chunks and an 8 GiB device expert cache. All expert arithmetic ran on the GPU. Each mode used two fixed prompts, three greedy repetitions and 64 generated tokens per request, with prefix reuse disabled. Each fixed mode repeated exactly and accepted MTP drafts. Image requests identified the red test image with thinking disabled. Flash-Next disables MTP for media requests and uses plain decoding. A fresh text request resumed MTP after each image check. An initial check incorrectly required MTP for media and failed. Source inspection confirmed this existing limitation; the engine behavior was not changed.

Samples include the first request. Cache contents are not reset between repetitions. Other artifacts were being prepared on the same host during this campaign, so these timings do not isolate the performance effect of MTP.

The measurements below cover only those short requests. The OS page cache was not controlled. They do not establish cold-disk speed, low-RAM operation or RTX 3080 support. Different draft widths or modes can produce different tokens because their arithmetic rounds differently. Prior quality results below were measured without this MTP attachment.

Prompt 1 requests a Python merge function after a repeated prose prefix. Prompt 2 requests a simple explanation of the blue sky. Each row summarizes the three repetitions of that prompt.

prompt expert placement draft tokens minimum draft probability median decode tok/s range tok/s
1 host 4 0 38.43 25.90–41.98
2 host 4 0 32.00 26.02–35.37
1 disk 4 0 13.64 10.61–13.82
2 disk 4 0 11.72 9.88–11.88
1 host 0 0 50.68 31.28–55.58
2 host 0 0 45.30 35.36–49.92
1 host 4 0.3 38.20 26.59–42.29
2 host 4 0.3 31.81 26.35–35.37

The probability-floor experiment uses the tested development snapshot. The normal MTP command above uses the default floor. mtp-qualification.json records every sample and the tested source snapshot; the minimum commit above identifies artifact support.

Quality

NInfer computes what llama.cpp computes from the same GGUF. Over the first 72 KB of the WikiText sample in the repository's eval/corpora/perplexity-1m, in 2,560-token windows that advance by 1,024 targets (llama.cpp's --ppl-stride 1024 -c 2048), the fourteen windows both programs score the same way (14,336 targets) give:

perplexity
this artifact in NInfer, BF16 KV 2.6579
the GGUF in llama.cpp, BF16 KV 2.6502

Window by window the two differ by -0.014 to +0.016 nats. Deep into a long context they agree as closely: over four PG-19 streams of the same corpus joined (261,412 tokens), in 65,536-token windows that advance by 32,768 targets (--ppl-stride 32768 -c 49152), so that every scored position lies past the first 32,768 of its window, the five windows both score the same way give 8.0834 in NInfer and 8.0802 in llama.cpp, window by window within 0.0011 nats (two RTX 5090s, BF16 KV).

The GSQ-RCO card evaluates this quantization in llama.cpp at AIME25 96.67, GPQA-Diamond 89.39 and LiveCodeBench v6 81.14 (BF16: 100.00, 91.92, 87.43), without stating its sampling, output limit or number of runs. In NInfer, EvalScope ran AIME 2025 and GPQA-Diamond against this artifact with its table on two RTX 5090s, every expert in device memory and six requests at a time: thinking on, temperature 1.0, top-p 0.95, top-k 20, one sampled run, at most 106,000 output tokens (October 2026). The answers cut there were then continued from where they stopped to 122,880 (AIME) and 245,760 (GPQA) output tokens:

AIME 2025 GPQA-Diamond
this artifact in NInfer, the cut answers continued 93.33 (28/30) 86.36 (171/198)
this artifact in NInfer, at most 106,000 output tokens 93.33 84.34
the GGUF in llama.cpp (GSQ-RCO card) 96.67 89.39

One AIME answer and ten GPQA answers reached 106,000 tokens without answering. Continued, four of the GPQA answers came out right and four wrong; one GPQA answer was still reasoning at 245,760 tokens, one was repeating a codon of its question's DNA sequence, and the AIME answer was still calculating at 122,880. One run of GPQA-Diamond's 198 questions has a standard error of about 2.5 points. The answers are long, 23,000 to 26,000 output tokens on average; the run took 7 h 44 min at 193 tokens per second across the six requests, and the continuations 58 min more. LiveCodeBench needs a code sandbox and was not run. Measurements has the details.

Speed

October 2026, greedy decoding, the repository's generate test (ninfer_qwen4_exp_generate_real). Decode is measured over the 78 tokens of a short answer to a 20-token prompt and over the first five tokens after a 4,463-token prompt; prefill is that prompt in 512-token chunks. Host memory is the process's peak resident set. These are the weights of this file (the table and the Vision tower are not read here).

The RTX PRO 6000 and RTX 5090 rows ran a 120a build with CUDA 13.1 and the separate table artifact, single-GPU rows pinned to the GPU's NUMA node:

  • RTX PRO 6000 Workstation Edition: 600 W, a PCIe 4.0 x16 host link.
  • RTX 5090s: boards with a 600 W default limit, PCIe 5.0 x16.

With host or disk experts the device expert cache takes the free memory, about 94 GB on the PRO 6000. The rows measure it while it is still filling.

The RTX 4090 rows ran this file with the separate table artifact on RTX 4090s at 450 W (PCIe 4.0 x16) in a cloud VM with two EPYC 7543 sockets, pinned to the CPUs of the GPUs' NUMA node, two runs each where a range is given; left unpinned, host experts decode there at 40 tok/s and disk experts at 32 to 33. The RTX 3090 Ti and first three RTX 3090 rows were measured before the weights moved into one file, the last in a single file with the table, on another RTX 3090 (310 W power limit, 62 GB of RAM, 23 cores of an AMD EPYC 7663) next to the IQ3_S artifact.

hardware and placement device memory host memory decode, short answer decode after 4,463 tokens prefill
RTX PRO 6000, experts on the GPU 37.7 GiB 0.8 GiB 136.0-139.7 tok/s 160.7-162.6 tok/s 2,925-2,952 tok/s
RTX PRO 6000, experts in pinned host memory 38.5 GiB 34.0 GB pinned 54.8-55.0 tok/s 67.7-70.0 tok/s 1,296-1,297 tok/s
RTX PRO 6000, experts on disk, artifact in the page cache 93.6 GiB 1.0 GiB + page cache 75.9 tok/s 115.8 tok/s 2,024 tok/s
RTX PRO 6000, experts on disk, the files' pages evicted every second 93.5 GiB 1.0 GiB 35.1 tok/s 87.1 tok/s 1,000 tok/s
2× RTX 5090 (PCIe), experts on the GPUs 18.8 + 19.6 GiB 0.9 GiB 136.1-136.3 tok/s 157.0-157.2 tok/s 3,807-3,811 tok/s
RTX 5090, experts in pinned host memory 30.0 GiB 34.0 GB pinned 74.1-74.2 tok/s 93.9-94.2 tok/s 1,676-1,678 tok/s
RTX 5090, experts on disk, artifact in the page cache 29.9 GiB 1.0 GiB + page cache 67.5 tok/s 76.7 tok/s 1,739 tok/s
RTX 5090, experts on disk, the files' pages evicted every second 29.9 GiB 1.0 GiB 25.3 tok/s 28.7 tok/s 694 tok/s
2× RTX 4090 (PCIe, no P2P), experts on the GPUs 18.5 + 19.4 GiB 0.8 GiB 92.6-100.7 tok/s 116-120 tok/s 3,482-3,734 tok/s
2× RTX 3090 Ti (PCIe, no P2P), experts on the GPUs 18.4 + 19.4 GB 0.7 GB 90.2 tok/s 107 tok/s 1,504 tok/s
RTX 4090, experts in pinned host memory 20.7 GiB 34.0 GB pinned 52.6-52.7 tok/s 44.3-44.6 tok/s 1,137-1,139 tok/s
RTX 4090, experts on disk, artifact in the page cache 20.7 GiB 1.0 GiB + page cache 42.0 tok/s 45.6 tok/s 723 tok/s
RTX 4090, experts on disk, the file's pages evicted every second 20.5 GiB 1.0 GiB 17.7 tok/s 12.2 tok/s 160 tok/s
RTX 3090, experts in pinned host memory 22.6 GB 34 GB pinned 48.9 tok/s 42.1 tok/s 842 tok/s
RTX 3090, experts on disk, artifact in the page cache 22.6 GB 0.9 GB + page cache 47.0 tok/s 39.1 tok/s 630 tok/s
RTX 3090, experts on disk, page cache dropped every second (NVMe) 22.6 GB 0.9 GB 17.2 tok/s 11.3 tok/s 195 tok/s
RTX 3090, experts in pinned host memory, this file 22.1 GiB 34.0 GB pinned 50.2 tok/s 43.7 tok/s 847 tok/s

llama.cpp runs the GGUF with the experts on the CPU (--n-cpu-moe 48) at 11.2 tok/s decode (tg128) and 267 tok/s prefill (pp512) on the first RTX 3090's machine (32 threads), and at 12.7 and 368 tok/s on the second (23 threads).

Credits and license

Quantized weights: ISTA-DASLab, Qwen3.8-Flash-Next GSQ-RCO GGUFs, produced with GSQ and RCO from Qwen/Qwen3.8-Flash-Next. The model is released under the Qwen Community License 1.0, reproduced in LICENSE, and its conditions apply to these files; ISTA-DASLab publish their GGUFs under Apache-2.0. Attribution notices are collected in NOTICE. This repository only re-packs those weights into NInfer's container.

Downloads last month
244
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-NInfer-v3

Quantized
(16)
this model

Papers for WaveCut/Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-NInfer-v3