Instructions to use intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound") model = AutoModelForCausalLM.from_pretrained("intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound
- SGLang
How to use intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound with Docker Model Runner:
docker model run hf.co/intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound
DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound
MXFP4 mixed-precision re-quantization of
DeepSeek-V4-Flash-0731 — a 304 B
total / ~33 B-active sparse MoE (43 layers, 256 routed + 1 shared experts/layer, hybrid compressed
attention + DSA indexer, 1M context, attached DSpark/MTP draft module) — produced with
Intel AutoRound in model-free RTN mode (no calibration
dataset), exported as compressed-tensors mixed-precision:
- Routed experts (main model + MTP) → MXFP4 W4A4 (E2M1, group_size 32, E8M0 scales; activations MXFP4 dynamic)
- All other quantized Linear layers → MXFP8 W8A8 (E4M3, group_size 32, E8M0 scales)
Accuracy vs the official checkpoint (same protocol): AVG 0.8305 vs 0.8348 — −0.43 pp.
Completeness note. This checkpoint is a from-scratch re-export fixing a defect in the published community re-quant (INCModel2/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound), which silently omits the entire
layers.29.*(3.57 GB) while itsconfig.jsonstill declares 43 layers — models served from it effectively run 42 layers deep and lose ~4.5 pp on gsm8k. This checkpoint carries all 72 317 tensors of the official layout (43/43 layers + full MTP/DSpark).
About a "BF16 baseline". DeepSeek has never released a BF16 version of V4-Flash — the official checkpoint is natively FP4+FP8 quantized. The strongest available reference is therefore the official checkpoint itself, used as the baseline below.
1. Model summary
| Official | This checkpoint | |
|---|---|---|
| Architecture | DeepseekV4ForCausalLM (deepseek_v4): 43 MoE layers, hidden 4096, 256 routed experts (top-6) + 1 shared, head_dim 512 with per-layer KV compression (ratio 4/128 alternating) + DSA indexer on 21 layers, mHC hyper-connections, 3-layer DSpark/MTP draft module, 1M context |
identical |
| Parameters | ~304.2 B total (277.0 B main routed experts + 19.8 B MTP, 91 % in experts) | identical (weights re-encoded) |
| Quantization | FP8 E4M3 128×128 block (linear) + MXFP4 experts, expert_dtype=fp4 |
MXFP4 W4A4 (experts, incl. MTP) + MXFP8 W8A8 (other Linear), group_size 32, E8M0 scales |
| Format | native (quant_method: fp8 + expert_dtype) |
compressed-tensors mixed-precision |
| Size | 166.89 GB | 167.08 GB / 155.61 GiB, 48 shards |
| Tensors | 72 317 | 72 317 (complete) |
| License | MIT | same |
2. Quantization scope
Quantized:
| Modules | Format | Count |
|---|---|---|
layers.0-42.ffn.experts.*.w{1,2,3} (routed experts, 43 × 256 × 3) |
MXFP4 W4A4 | 33 024 |
mtp.0-2.ffn.experts.*.w{1,2,3} (MTP/DSpark draft experts, 3 × 256 × 3) |
MXFP4 W4A4 | 2 304 |
attn.{wq_a,wq_b,wkv,wo_a,wo_b} · attn.indexer.wq_b · ffn.shared_experts.w{1,2,3} · mtp.* attention/main_proj |
MXFP8 W8A8 | 390 |
Not quantized (BF16/F32; ignore, 196 entries): embed / head (0.53 B each), per-layer KV
compressor.{wgate,wkv} (41×2) and indexer.compressor.* / indexer.weights_proj (21×3 — the
sparse-attention routing path), MoE routers ffn.gate (43+3), DSpark markov_head /
confidence_head, all hyper-connection (hc_*) and norm parameters.
3. Evaluation results
Same protocol for both columns: lm-eval 0.4.13 + vLLM, TP=2 on 2× B300, kv_cache_dtype=fp8,
max_model_len=8192, seed=42, gsm8k 5-shot, piqa/mmlu/hellaswag 0-shot, full sample counts
(1 319 / 1 838 / 14 042 / 10 042).
| GSM8K (strict) | MMLU | PIQA | HellaSwag | AVG | |
|---|---|---|---|---|---|
| Official DeepSeek-V4-Flash-0731 (native FP4+FP8) | 0.9583 | 0.8649 | 0.8341 | 0.6818 | 0.8348 |
| MXFP4-Mixed-CT (this checkpoint) | 0.9568 | 0.8631 | 0.8232 | 0.6788 | 0.8305 |
| Δ | −0.15 pp | −0.18 pp | −1.09 pp | −0.30 pp | −0.43 pp |
Reading:
- Near-parity with the official checkpoint (99.5 % of AVG) — the ~4 pp gap previously reported
for CT re-quants of this model is not a quantization effect; it traces to the missing
layers.29.*in that release (see the completeness note above). Re-running the community-format checkpoint through the Marlin W4A16 MoE path recovered ~1.3 pp of it (0.9257), consistent with the W4A4 activation cost; the remaining gap was the missing layer. - The largest single-task delta here is PIQA (−1.09 pp, ≈1.2σ of the ±0.89 pp standard error). GSM8K (−0.15 pp, ±0.53 pp) is inside noise.
4. Usage (vLLM)
vllm serve DeepSeek-V4-Flash-0731-MXFP4-Mixed-0929 \
--trust-remote-code \
--tensor-parallel-size 2 \
--block-size 256 \
--kv-cache-dtype fp8 \
--max-model-len 8192 --max-num-seqs 256 \
--gpu-memory-utilization 0.90 --dtype bfloat16 \
--enforce-eager \
--no-enable-flashinfer-autotune
Serving constraints (each verified by measurement):
- KV cache dtype is backend-dependent: the default selected backends (FlashMLA / fp8_ds_mla
layout,
deepseek_v4/attention.py:103-106) accept fp8 only —--kv-cache-dtype autofails at start-up withDeepseekV4 fp8_ds_mla layout only supports fp8 kv-cacheand the requestedfp8is rewritten to the packedfp8_ds_mlaformat (576 B/token slot). bf16/auto KV is legal only on the FlashInfer sparse backend withuse_fp4_indexer_cache(attention.use_fp8_ds_mla_layout=False). All numbers in §3 usefp8— identical for both compared columns, so the comparison is apples to apples. - Requires the DSv4 compressed-tensors o_proj path — the stock fused-attention output projection
(
deep_gemm fp8_einsum) only accepts the native 128×128 block-FP8 weight scales and dies on this checkpoint's per-32 MXFP8 scales (t.dim() == N). Use a vLLM build whose DSv4o_projdispatches compressed-tensors MXFP8wo_ato an MXFP8 GEMM (this checkpoint was validated on such a build). - MoE: the CUTLASS MXFP4 W4A4 path is correct at TP=1/2/4 here (per-rank expert scale columns
2048/TP/32are always multiples of 4). Do not use the flashinfertrtllm_mxfp4auto-pick on builds that JIT-compiletrtllm_moe_sm100at first use (long stalls);--no-enable-flashinfer-autotuneand disabling JIT/cuteDSL warmup are recommended. - TP must divide 64 (attention and indexer heads). DSpark speculative decoding is available with
--speculative-config '{"method":"dspark",...}'(draft weights ship in this checkpoint).
5. Reproduce
5.1 Quantization
auto-round \
--model_name deepseek-ai/DeepSeek-V4-Flash-0731 \
--model_free \
--scheme MXFP8 \
--ignore_layers compressor,indexer.weights_proj \
--layer_config '{ffn.experts:{bits:4,data_type:mx_fp}}' \
--format llm_compressor \
--output_dir ./DeepSeek-V4-Flash-0731-MXFP4-Mixed
Always verify layer completeness after a model-free run — the community re-quant shipped without
layers.29.* and no tool warned:
python3 -c "
import json, re
wm = json.load(open('<output>/model.safetensors.index.json'))['weight_map']
layers = sorted({int(m.group(1)) for k in wm if (m := re.match(r'layers\.(\d+)\.', k))})
print('layers:', len(layers), 'missing:', [i for i in range(43) if i not in layers])"
5.2 Evaluation
MODEL_ARGS='{"pretrained": "DeepSeek-V4-Flash-0731-MXFP4-Mixed-0929",
"tensor_parallel_size": 2, "max_model_len": 8192, "max_num_seqs": 256, "block_size": 256,
"gpu_memory_utilization": 0.9, "dtype": "bfloat16", "trust_remote_code": true,
"kv_cache_dtype": "fp8", "enable_prefix_caching": false, "max_gen_toks": 2048,
"add_bos_token": true, "enforce_eager": true,
"kernel_config": {"enable_flashinfer_autotune": false, "enable_cutedsl_warmup": false, "enable_jit_warmup": false}}'
# gsm8k: 5-shot generative, chat template + multi-turn few-shot
lm_eval --model vllm --model_args "$MODEL_ARGS" --tasks gsm8k \
--batch_size 32 --seed 42 --apply_chat_template --fewshot_as_multiturn --output_path ./results
# piqa / mmlu / hellaswag: 0-shot log-likelihood, no chat template
lm_eval --model vllm --model_args "$MODEL_ARGS" --tasks piqa,mmlu,hellaswag \
--batch_size 32 --seed 42 --output_path ./results
6. Known limitations
- Four academic benchmarks only; no long-context / agentic / code / thinking-on evaluation; no throughput numbers.
kv_cache_dtype=fp8is forced by the architecture, so all published numbers carry fp8-KV effects (same for the official baseline — apples to apples).- Serving requires a vLLM build with the compressed-tensors DSv4 o_proj path (§4).
- MXFP4 experts here are W4A4 (4-bit activations); the official native path runs FP4 weights with FP8 activations. PIQA (−1.09 pp) is the visible cost of that difference.
7. License and attribution
Base model DeepSeek-V4-Flash-0731 is
MIT (DeepSeek-AI); the LICENSE in this repository applies to this derivative unchanged.
Quantization with Intel AutoRound (Apache-2.0), model-free
RTN — no calibration corpus used. MXFP4/MXFP8 follow the
OCP Microscaling Formats (MX) specification. Serving via
vLLM; evaluation via
lm-evaluation-harness.
- Downloads last month
- 242
Model tree for intel-ai/DeepSeek-V4-Flash-0731-MXFP4-Mixed-CT-AutoRound
Base model
deepseek-ai/DeepSeek-V4-Flash-0731