Instructions to use Akicou/DeepSeek-V4-Flash-NE50-42-OP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Akicou/DeepSeek-V4-Flash-NE50-42-OP with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Akicou/DeepSeek-V4-Flash-NE50-42-OP")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Akicou/DeepSeek-V4-Flash-NE50-42-OP") model = AutoModelForCausalLM.from_pretrained("Akicou/DeepSeek-V4-Flash-NE50-42-OP", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Akicou/DeepSeek-V4-Flash-NE50-42-OP with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Akicou/DeepSeek-V4-Flash-NE50-42-OP" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Akicou/DeepSeek-V4-Flash-NE50-42-OP", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Akicou/DeepSeek-V4-Flash-NE50-42-OP
- SGLang
How to use Akicou/DeepSeek-V4-Flash-NE50-42-OP with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Akicou/DeepSeek-V4-Flash-NE50-42-OP" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Akicou/DeepSeek-V4-Flash-NE50-42-OP", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Akicou/DeepSeek-V4-Flash-NE50-42-OP" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Akicou/DeepSeek-V4-Flash-NE50-42-OP", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Akicou/DeepSeek-V4-Flash-NE50-42-OP with Docker Model Runner:
docker model run hf.co/Akicou/DeepSeek-V4-Flash-NE50-42-OP
DeepSeek-V4-Flash-NE50-42-OP
This is an offline seed-pruned version of deepseek-ai/DeepSeek-V4-Flash, produced with Akicou/ream.
It is not an official DeepSeek release.
Name decoding
DeepSeek-V4-Flash-NE50-42-OP means:
NE50: created with--n-experts 5042: seed42OP: offline pruning
Important: in this pruning script, --n-experts 50 means 50 routed experts were removed from each MoE layer, not retained. The resulting model keeps 206 / 256 routed experts per MoE layer.
What was changed
The base DeepSeek-V4-Flash checkpoint was pruned directly at the safetensors level without loading the full Transformers model into memory and without GPU calibration.
Pruning details:
| Item | Value |
|---|---|
| Base model | deepseek-ai/DeepSeek-V4-Flash |
| Method | Offline random seed expert pruning |
| Command flag | --n-experts 50 |
| Seed | 42 |
| Original routed experts per MoE layer | 256 |
| Experts removed per MoE layer | 50 |
| Routed experts retained per MoE layer | 206 |
| MoE layers processed | 43 |
| Config value | n_routed_experts: 206 |
| Experts per token | 6 |
| Shared experts | 1 |
| Output model safetensors payload | ~`130.85 GB` |
Only routed MoE experts and their router tensors were pruned/remapped. Shared experts, non-MoE weights, tokenizer files, inference files, and model metadata were copied from the base model.
How it was created
python examples/compress_model.py \
--model deepseek-ai/DeepSeek-V4-Flash \
--output ./DeepSeek-V4-Flash-NE50-42-OP \
--offline-seed-prune \
--n-experts 50 \
--seed 42
Important notes
- This is random seed pruning, not calibrated saliency pruning.
- No benchmark evaluation is claimed here.
- Quality may be worse than the original model.
- The model may require a recent
torch/transformersstack with DeepSeek-V4 support.
Basic text usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Akicou/DeepSeek-V4-Flash-NE50-42-OP"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
low_cpu_mem_usage=True,
)
prompt = "What is a reaper?"
inputs = tokenizer(prompt, return_tensors="pt").to(next(model.parameters()).device)
with torch.no_grad():
output = model.generate(**inputs, max_new_tokens=128, do_sample=False)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Attribution
All architecture, tokenizer, inference files, and original weights are from deepseek-ai/DeepSeek-V4-Flash. This repository only contains an offline-pruned derivative checkpoint.
- Downloads last month
- 31
Model tree for Akicou/DeepSeek-V4-Flash-NE50-42-OP
Base model
deepseek-ai/DeepSeek-V4-Flash