Configuration Parsing Warning:In UNKNOWN_FILENAME: "quantization_config.config_groups.group_0.format" must be a string

gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound

This checkpoint is an AutoRound model-free MXFP8 RTN quantization of google/gemma-4-26B-A4B, exported in llm_compressor / compressed-tensors format. Routed experts and the self-attention projections present in the source are stored as F8_E4M3; sensitive/shared and multimodal weights remain BF16. Static FP8 KV scales were calibrated with AutoRound using the text dataset NeelNanda/pile-10k; the vision tower was not quantized.

The comparisons below use a BF16 reference with BF16 KV cache and this MXFP8 checkpoint with FP8 KV cache. The four standard tasks and the separate 13-task RULER suite have separate tables and separate unweighted means; their scores are never pooled. Scores and MXFP8/BF16 ratios are shown as percentages to two decimal places. The runs used different parallelism and batching settings, so the comparisons are descriptive and do not isolate the effect of weight quantization.

Read this first: The four standard benchmark prompts were not a 128K long-context stress test. A separate RULER section reports results at the 131,072-token context bucket, with 100 evaluated samples per task. The evaluations covered text tasks only; vision quality, multimodal prompts, thinking-on, throughput and latency were not measured. The tokenizer configuration used by the standard-task run had no chat_template, so GSM8K used a plain-prompt configuration.

1. Model summary

Base model This checkpoint
Name / architecture / parameters Gemma 4 26B-A4B; Gemma4ForConditionalGeneration; nominal count 26,544,131,376 (26.544B; tied-weight counting convention) gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound; 25,805,936,206 unique serialized weight elements
Layer structure 30 text layers; 128 routed experts/layer, top-8; 25 sliding-attention + 5 full-attention layers; vision tower Routed-expert and source-present text self-attention projection weights MXFP8; protected modules BF16
Weight tensors 1,013 source index entries, BF16 24,222 exported tensor entries: 11,635 F8_E4M3 weights, 11,635 U8 block scales, 838 BF16 weights and 114 F32 static KV scales
Size on disk 51,611,872,412 bytes, source safetensors index total_size 28,412,087,908 bytes, exported safetensors index total_size
Compression / effective bits per parameter BF16 source 1.8165× smaller; 8.8079 effective bits per unique serialized weight element (8 × export bytes / 25,805,936,206)
Quantization / format / context / license Gemma 4; Apache 2.0 MXFP8 weights + static FP8 KV; compressed-tensors; evaluated with a 131,072-token maximum length; Apache 2.0 terms apply

The model configuration sets tie_word_embeddings: true. The nominal 26.544B model count and the index-derived unique serialized element count are different counting bases; the size and dtype accounting below uses the latter. The A4B designation is the upstream model's active-parameter naming, not a measurement made by this quantization run.

Serving was validated through the vLLM 0.29.0 backend on three RTX 5090 GPUs with TP=1 / PP=3. That is the tested topology, not a claim that other hardware or parallel layouts are equivalent.

2. Precision plan

Percentages below are relative to the 25,805,936,206 unique stored model-weight elements. Scale tensors are metadata and are not included in this denominator.

Bit width / stored dtype Module (layers) Weight tensors Weight elements Share
8-bit float (F8_E4M3) Routed expert gate/up/down weights (30 × 128 experts) 11,520 22,837,985,280 88.4990%
8-bit float (F8_E4M3) Existing text self-attention projections (q/k/o in 30 layers; v in 25) 115 1,110,179,840 4.3020%
16-bit (BF16) Shared MLP branch 90 535,265,280 2.0742%
16-bit (BF16) Router 90 10,901,760 0.0422%
16-bit (BF16) Vision tower 355 569,550,384 2.2071%
16-bit (BF16) Vision embedding projection 1 3,244,032 0.0126%
16-bit (BF16) Token embedding 1 738,197,504 2.8606%
16-bit (BF16) Other retained tensors, including norms and patch/position tensors 301 612,126 0.0024%
Total model weights F8_E4M3 + BF16 12,473 25,805,936,206 100.0000%

Shares are rounded independently; the unrounded category shares sum to 100%. The exported weight configuration is Linear, 8-bit float, symmetric, group-wise, group size 32, static weights, with U8 scales and memoryless_minmax observer. The config also records dynamic 8-bit float input-activation quantization metadata. Static KV is 8-bit float, tensor-wise and dynamic=false.

The exported checkpoint also has 11,635 U8 block-scale tensors and 114 F32 static KV scales. Of the latter, the 60 text K/V scales (30 layers × K/V) are finite and nonzero; the 54 vision scales are zero because calibration and evaluation were text-only.

Ignore groups

Group Reason / evidence
router Preserve routing decisions; retained as BF16 in the header audit.
mlp This is the always-on shared branch added to routed-expert output; retained as BF16.
vision_tower, embed_vision Preserve the multimodal tower and projection; not part of this text-only evaluation.
embed_tokens Embedding is not a supported AutoRound Linear target; source/export key is present and BF16. The logged unmatched-module warning is benign, not a missing checkpoint key.

Norms and other non-target tensors remain BF16 by the model-free defaults. The checkpoint contains no multi_modal_projector key, so it is not included in the ignore list.

2.1 Quantization plan and artifact deviations

Audited plan Export observation Difference / consequence
Quantize all routed experts 11,520 F8_E4M3 tensors across 30 layers × 128 experts Matches plan; 60 fused source expert keys were expanded into per-expert tensors.
Quantize source-present text self-attention projections 115 F8_E4M3 projection tensors Matches source index: q/k/o exist in all 30 layers, while v exists only in 25 sliding-attention layers. The five full-attention source layers have no v_proj.
Keep shared MLP, router, vision, embeddings and other non-target tensors BF16 838 BF16 tensors, 1,857,771,086 elements Matches plan; no non-fused source keys were absent from the exported index.
Add static FP8 KV scales 114 F32 scales; all finite; 60 text scales nonzero Text scales cover the evaluated workload. Vision scales are zero after text-only calibration and were not exercised because language_model_only=true.

There was no unplanned weight-precision change. The embed_tokens warning reflects that embeddings are not Linear quantization targets; the key remains present in BF16.

2.2 Size accounting

Artifact Header/index bytes Comparison
BF16 source checkpoint 51,611,872,412 Source index total_size; 1,013 BF16 entries
MXFP8 + BF16 export 28,412,087,908 Export index total_size; 24,222 weight/scale entries

The ratio is 51,611,872,412 / 28,412,087,908 = 1.8165×. Effective bits per unique serialized model-weight element are 28,412,087,908 × 8 / 25,805,936,206 = 8.8079; this includes block/KV scale storage represented in the exported index and is not the nominal MXFP8 weight bit width.

3. Evaluation results

Measured with lm_eval 0.4.13 and vLLM 0.29.0, using seed 42 and the full repository task sets. The BF16 reference was evaluated on 2026-10-09 with BF16 compute and kv_cache_dtype=auto, which vLLM/FlashInfer resolved to BF16. The MXFP8 checkpoint was evaluated on 2026-09-28 with BF16 compute and FP8 KV cache. N is the evaluated sample count. Scores are the primary task metrics; the average is the unweighted arithmetic mean of the four scores. MXFP8 / BF16 is the score ratio (MXFP8 score ÷ BF16 score × 100%), computed from the unrounded values. Because KV-cache precision, tensor/pipeline parallelism and batching differ between runs, these comparisons are descriptive and are not a controlled estimate of the effect of weight quantization alone.

Task / primary metric N BF16 weights + BF16 KV MXFP8 weights + FP8 KV Δ (pp) MXFP8 / BF16
PIQA (acc) 1,838 82.32% 82.86% +0.54 100.66%
MMLU (acc) 14,042 74.40% 74.41% +0.01 100.02%
HellaSwag (acc) 10,042 63.32% 63.47% +0.15 100.24%
GSM8K (exact_match,strict-match) 1,319 74.37% 73.24% −1.14 98.47%
Unweighted average — 73.60% 73.50% −0.11 99.85%

Scores, differences and ratios are rounded to two decimal places; calculations use the full-precision results. A ratio above 100% means the MXFP8 score is higher for that task; below 100% means it is lower. The average ratio is the ratio of the two unrounded arithmetic means, not the mean of the task ratios.

3.2 Evaluation protocols

Both runs used vLLM 0.29.0, BF16 compute, max_model_len=131072, gpu_memory_utilization=0.85, attention_backend=FLASHINFER, language_model_only=true, trust_remote_code=true, add_bos_token=true, prefix caching disabled, thinking disabled, seed 42 and the full task sets. The cache dtype and execution settings were not the same:

Run KV cache Parallelism Evaluation batching
BF16 reference kv_cache_dtype=auto, resolved to BF16 in runtime logs TP=4 / PP=1 batch size 64; max_num_seqs=64, max_num_batched_tokens=32768; FlashInfer startup autotune disabled
MXFP8 checkpoint kv_cache_dtype=fp8, matching its static FP8 KV scheme TP=1 / PP=3 PIQA/MMLU/HellaSwag batch size auto; GSM8K batch size 64 and 5-shot

Both used the standard 2,048-token generation budget for GSM8K. Since KV-cache precision and parallelism/batching differ, do not interpret the score ratio as a weight-only comparison.

3.3 Score ratio interpretation

The four-task table reports the MXFP8/BF16 ratio for each primary metric and the ratio of the two four-task means. The separate RULER table below reports its own per-task ratios and 13-task mean. These are point-score comparisons across the stated protocols, not paired significance tests or evidence that MXFP8 improves accuracy.

3.4 Functional evidence and provenance

Direct evidence: the exported checkpoint loaded in vLLM 0.29.0 with FP8 KV cache and completed the four standard tasks and the separate RULER evaluation. Its quantization_config.json records compressed-tensors MXFP8 weights and a static 8-bit float KV scheme.

The RULER results below cover the native 131,072-token bucket with 100 evaluated samples per task; they are not a general full-window saturation test across arbitrary prompts. Not measured: multimodal/vision quality, thinking-on, agentic tasks, throughput or latency. Environment support probes and sibling model runs are not substituted for this checkpoint's measured evidence.

3.5 RULER 128K evaluation (separate suite)

The table reports the 13 lm_eval RULER tasks at the 131072,none metric bucket, with N=100 evaluated samples per task. Scores are percentages. MXFP8 / BF16 is the task score ratio (MXFP8 score ÷ BF16 score × 100%), calculated from full-precision results. The RULER mean is the unweighted arithmetic mean of these 13 task scores only; it is separate from, and is not combined with, the four-task mean above.

Task / metric N BF16 weights + BF16 KV MXFP8 weights + FP8 KV Δ (pp) MXFP8 / BF16
niah_single_1 (131072,none) 100 100.00% 100.00% +0.00 100.00%
niah_single_2 (131072,none) 100 100.00% 100.00% +0.00 100.00%
niah_single_3 (131072,none) 100 100.00% 100.00% +0.00 100.00%
niah_multikey_1 (131072,none) 100 97.00% 95.00% −2.00 97.94%
niah_multikey_2 (131072,none) 100 99.00% 98.00% −1.00 98.99%
niah_multikey_3 (131072,none) 100 100.00% 98.00% −2.00 98.00%
niah_multiquery (131072,none) 100 97.75% 96.00% −1.75 98.21%
niah_multivalue (131072,none) 100 98.75% 97.50% −1.25 98.73%
ruler_vt (131072,none) 100 98.80% 99.20% +0.40 100.40%
ruler_cwe (131072,none) 100 82.80% 80.10% −2.70 96.74%
ruler_fwe (131072,none) 100 98.00% 94.33% −3.67 96.26%
ruler_qa_squad (131072,none) 100 51.83% 51.50% −0.33 99.36%
ruler_qa_hotpot (131072,none) 100 55.00% 54.00% −1.00 98.18%
RULER unweighted mean (13 tasks only) — 90.69% 89.51% −1.18 98.70%

The mean ratio is the ratio of the two unrounded RULER means, not the mean of the task ratios. BF16 used kv_cache_dtype=auto (resolved to BF16); MXFP8 used its calibrated static FP8 KV cache. Consequently, these RULER results compare the evaluated serving configurations and are not a weights-only ablation. The 100-sample cap is not the full RULER test set.

4. Usage

4.1 vLLM evaluation

The checkpoint was evaluated with vLLM through lm_eval. The following reproduces the text benchmark invocation using the published model ID:

MODEL_ID=INCModel4/gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound
MODEL_ARGS="pretrained=${MODEL_ID},tensor_parallel_size=1,pipeline_parallel_size=3,max_model_len=131072,gpu_memory_utilization=0.85,dtype=bfloat16,trust_remote_code=True,add_bos_token=True,enable_prefix_caching=False,max_gen_toks=2048,attention_backend=FLASHINFER,language_model_only=True,kv_cache_dtype=fp8,max_num_seqs=1,enable_thinking=False"

# Select three idle GPUs before running; this example assumes physical GPUs 1-3 are idle.
export CUDA_VISIBLE_DEVICES=1,2,3
lm_eval --model vllm --model_args "$MODEL_ARGS" \
  --tasks piqa,mmlu,hellaswag --batch_size auto --log_samples --seed 42
lm_eval --model vllm --model_args "$MODEL_ARGS" \
  --tasks gsm8k --batch_size 64 --log_samples --seed 42 --num_fewshot 5

Hard constraints and evidence:

Setting Requirement / evidence
KV cache Keep kv_cache_dtype=fp8 to match the evaluated quantized model protocol and the exported KV scheme.
Context max_model_len=131072 is the measured configuration. The four standard tasks are not a long-context test; the separate RULER table covers its 131,072-token bucket with 100 samples per task. Neither establishes general 128K throughput or quality.
Parallel layout TP=1 / PP=3 on three GPUs is the tested layout. A two-GPU layout encountered memory pressure during CUDA graph warmup; other layouts are unverified.
Backend vLLM 0.29.0 with attention_backend=FLASHINFER was exercised by the four evaluation tasks. For head dimension 512, the XQA decode path logged a fallback to native FlashInfer decode; KV remained FP8 and evaluation completed.
Prompt formatting The run's local/upstream tokenizer configuration had no chat_template. GSM8K therefore used apply_chat_template=false and fewshot_as_multiturn=false; multimodal/chat formatting was not evaluated.

For text inference, supply a plain prompt compatible with the task rather than assuming a chat template. A direct production vllm serve deployment was not separately benchmarked by this run.

4.2 Transformers inspection path

The following reads model configuration only; it is not a weight-loading or serving validation path:

from transformers import AutoConfig

model_id = "INCModel4/gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound"
config = AutoConfig.from_pretrained(model_id, trust_remote_code=True)
print(config.model_type)

The checkpoint contains about 28.41 GB of indexed tensors. Loading full weights through an inspection-oriented Transformers path can use substantially more host/GPU memory; use the tested vLLM backend for inference.

5. Reproducibility

5.1 Hardware / OS

  • GPUs: 3 × NVIDIA GeForce RTX 5090 used (compute capability 12.0; physical devices 1, 2 and 3); 32,607 MiB total memory reported per device.
  • Host probe on 2026-09-28: NVIDIA driver 610.57.04; Intel Xeon 6767P (256 logical CPUs); 188.4 GiB RAM.
  • OS: Ubuntu 24.04.4 LTS, kernel 6.8.0-139-generic, x86_64.
  • CUDA: PyTorch build CUDA 13.0; runtime 13.0.88.
  • Storage: local filesystem; filesystem type was not recorded.

5.2 Environment

Package Version
Python 3.12.14
auto-round 0.16.0.dev193+g22d6e28c
transformers 5.17.0
torch 2.13.0+cu130
vllm 0.29.0
flashinfer-python 0.7.0
compressed-tensors 0.17.0
lm-eval 0.4.13

These are the versions used for the measured quantization and evaluation, not a statement of minimum compatible versions.

5.3 Version selection

No packages were installed or upgraded for this run. These exact versions were used for the successful quantization and evaluation. Compatibility with other AutoRound, Transformers, vLLM or FlashInfer versions was not established here.

5.4 Artifact-level dependency warning

AutoRound 0.16.0.dev193 with Transformers 5.17.0 initially failed when its fused-MoE meta skeleton could not materialize the new self_attn.k_scale and self_attn.v_scale parameters. The successful run set AR_DISABLE_META_LOAD=1, which disabled that meta-load path and allowed CPU loading/initialization. Peak host RAM was 64.14 GiB. No source-code patch was applied.

5.5 Environment variables

For quantization, set the tested workaround and map three verified idle physical devices:

export AR_DISABLE_META_LOAD=1
export CUDA_VISIBLE_DEVICES=<three-idle-physical-GPU-IDs>

AR_DISABLE_META_LOAD is a quantization workaround; it is not required for inference. Do not put Hub tokens in this README or in configuration files.

6. Reproduce the artifact

Use a fresh output directory and select three idle GPUs before starting. model_free uses the RTN weight path, but static KV quantization still loads the checkpoint and runs calibration forwards. The validated run used the standard text calibration dataset NeelNanda/pile-10k; the vision tower was not quantized.

6.1 Quantize weights and static KV scales

export AR_DISABLE_META_LOAD=1
export CUDA_VISIBLE_DEVICES=<three-idle-physical-GPU-IDs>

auto-round --model_name /path/to/gemma-4-26B-A4B --scheme MXFP8 --model_free \
  --static_kv_dtype fp8 \
  --ignore_layers router,vision_tower,embed_vision,embed_tokens,mlp \
  --device_map 0,1,2 --format llm_compressor \
  --output_dir /path/to/output

6.2 Evaluate the published checkpoint and BF16 baseline

Use separate settings for the BF16 reference and MXFP8 checkpoint so the BF16 run uses BF16 KV rather than inheriting the candidate's FP8 KV setting. Select the requested number of idle physical GPUs before each run. Both commands use the measured context length, task definitions, seed and generation budget. GSM8K has no chat-template flags because the tokenizer configuration used during evaluation had no chat_template.

# BF16 reference: use four idle GPUs; auto resolves KV cache to BF16.
export CUDA_VISIBLE_DEVICES=<four-idle-physical-GPU-IDs>
MODEL_ARGS_BF16="pretrained=google/gemma-4-26B-A4B,tensor_parallel_size=4,pipeline_parallel_size=1,max_model_len=131072,gpu_memory_utilization=0.85,dtype=bfloat16,trust_remote_code=True,add_bos_token=True,enable_prefix_caching=False,max_gen_toks=2048,attention_backend=FLASHINFER,language_model_only=True,kv_cache_dtype=auto,max_num_seqs=64,max_num_batched_tokens=32768,enable_thinking=False,enable_flashinfer_autotune=False"
lm_eval --model vllm --model_args "$MODEL_ARGS_BF16" \
  --tasks piqa,mmlu,hellaswag --batch_size 64 --log_samples --seed 42
lm_eval --model vllm --model_args "$MODEL_ARGS_BF16" \
  --tasks gsm8k --batch_size 64 --log_samples --seed 42 --num_fewshot 5

# MXFP8 checkpoint: use three idle GPUs and its static FP8 KV cache.
export CUDA_VISIBLE_DEVICES=<three-idle-physical-GPU-IDs>
MODEL_ARGS_MXFP8="pretrained=INCModel4/gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound,tensor_parallel_size=1,pipeline_parallel_size=3,max_model_len=131072,gpu_memory_utilization=0.85,dtype=bfloat16,trust_remote_code=True,add_bos_token=True,enable_prefix_caching=False,max_gen_toks=2048,attention_backend=FLASHINFER,language_model_only=True,kv_cache_dtype=fp8,max_num_seqs=1,enable_thinking=False"
lm_eval --model vllm --model_args "$MODEL_ARGS_MXFP8" \
  --tasks piqa,mmlu,hellaswag --batch_size auto --log_samples --seed 42
lm_eval --model vllm --model_args "$MODEL_ARGS_MXFP8" \
  --tasks gsm8k --batch_size 64 --log_samples --seed 42 --num_fewshot 5

The validated export contains six safetensors shards and has total_size=28,412,087,908 bytes in its safetensors index.

7. Known issues and caveats

  • The four standard benchmarks do not test long context. The separate RULER table measures selected tasks at the 131,072-token bucket (N=100 each); do not extrapolate beyond those tasks or treat it as arbitrary full-window saturation testing.
  • Do not infer vision/multimodal quality from the text-only results. The vision weights remain BF16, but their quality was not evaluated.
  • PP=2 at the same length hit a CUDA graph warmup OOM on this host; the tested topology is TP=1 / PP=3 on three RTX 5090 GPUs.
  • FlashInfer XQA did not support this model's 512 head dimension in the tested path; vLLM fell back to native FlashInfer decode. FP8 KV evaluation still completed.
  • The token embedding warning is benign: it is not an AutoRound Linear target, its key is present, and it remains BF16.
  • The BF16 reference in the table uses BF16 KV; the MXFP8 checkpoint uses FP8 KV. Keep these settings explicit when interpreting or reproducing the comparison.
  • The score table reports only each task's primary metric; secondary metrics and their uncertainties are not part of this comparison.

8. License and attribution

The base model metadata identifies Apache 2.0 and links to the Gemma 4 license information. Users must comply with the base-model license and preserve its attribution. This quantized checkpoint is not an official Google release.

Quantization: Intel AutoRound. Inference validation: vLLM and FlashInfer. Evaluation: lm-evaluation-harness. Text calibration used NeelNanda/pile-10k; calibration samples are not included in this repository. This repository contains the quantized weights, tokenizer/processor configuration and model configuration; it does not redistribute benchmark datasets.

Downloads last month
75
Safetensors
Model size
26B params
Tensor type
F32
·
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for INCModel4/gemma-4-26B-A4B-MXFP8-FP8KV-CT-RTN-AutoRound

Quantized
(36)
this model