Instructions to use pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder") model = AutoModelForCausalLM.from_pretrained("pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder
- SGLang
How to use pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder with Docker Model Runner:
docker model run hf.co/pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder
Most sub-10B coding models crumble the moment they enter real-world agentic workflows: they either produce clean code but loop endlessly when a shell command fails, or handle tool calls reasonably well while hallucinating obscure API syntax.
Triumvirate is a merge designed to solve that dilemma. It combines three of the most capable specialized fine-tunes of Qwen 3.5 9B and fuses their task vectors directly into the base backbone:
- Algorithmic & Syntax Precision from Qwopus
- SWE-bench Problem Decomposition & Tool Calling from MiMo-V2.6
- Loop-Termination & Error-Recovery Discipline from Ornith-1.5
The result is a lean, blisteringly fast 9B pure-text causal engine with a native 256k context window that runs comfortably on consumer GPUs.
Contents
- Architectural Specifications
- Composition & Donor Weighting
- Merge Methodology & Mathematical Formulation
- Layer-Stratified Component Policies
- Agentic Chat Template
- Recommended Generation Parameters
- How to Use
- Citation & References
Architectural Specifications
| Parameter | Specification |
|---|---|
| Total Parameters | 8.8B (Text Backbone) |
| Architecture Type | Dense Causal Language Model (qwen3_5_text) |
| Hidden Dimension (dmodel) | 4096 |
| Intermediate Dimension (dmlp) | 12288 (SwiGLU) |
| Decoder Layers | 32 |
| Attention Mechanism | Hybrid Gated DeltaNet (3 Linear Attention : 1 Full Attention) |
| Full Attention Layers | Layers 3, 7, 11, 15, 19, 23, 27, 31 |
| Linear Attention Heads | 16 Key Heads / 32 Value Heads (dk = dv = 128) |
| Full Attention Heads | 16 Query / 4 Key-Value (GQA, dh = 256) |
| Rotary Position Embedding (RoPE) | 1D Partial RoPE (θ = 10⁷, Factor = 0.25) |
| Maximum Sequence Length | 262,144 tokens (256k) |
| Native Precision | bfloat16 |
Composition & Donor Weighting
The foundation checkpoint serves as the structural base (W₀). Three donor models contribute directional task vectors weighted continuously across network depth:
| Model | Role | Specialization Focus | Depth Target |
|---|---|---|---|
| Qwen/Qwen3.5-9B | Base Anchor (W₀) | Structural anchor & GDN linear attention state | Global |
| Jackrong/Qwopus3.5-9B-Coder | Donor 1 (D₁) | Claude 3.5 Opus distillation; typing, syntax, algorithms | Lower Layers (x ≤ 0.35) |
| ornith-ai/Ornith-1.5-9B | Donor 2 (D₂) | Agentic RL; loop-termination & error-pivot discipline | Mid Layers (0.35 < x < 0.70) |
| XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B | Donor 3 (D₃) | 77.4B tokens SFT; SWE-bench Pro, multi-turn tool logic | Top Layers (x ≥ 0.70) |
Merge Methodology & Mathematical Formulation
The merge combines TIES-DELLA saliency trimming, consensus sign election, Gated DeltaNet norm stabilization, and continuous sinusoidal depth modulation.
1. Task Vector Formulation
For each donor checkpoint k ∈ {1, 2, 3}, the parameter update delta is isolated relative to the base anchor W₀:
2. Asymmetric Sinusoidal Depth Modulation
Task vector mixing coefficients are continuously parameterized over normalized network depth x = l / (L - 1), where l ∈ {0, 1, ..., 31} and L = 32:
The donor weights αk(l) are normalized to form a partition of unity across all layers:
- Lower Layers (x → 0): Qwopus dominates with α₁(0) ≈ 0.67, ensuring foundational language representations and syntax heads are grounded in Claude 3.5 Opus traces.
- Middle Layers (x ≈ 0.5): The sub-linear exponent (x0.85) accelerates Ornith's activation to peak across middle transformer blocks with α₂(16) ≈ 0.354, reinforcing state-space continuity and execution discipline.
- Top Layers (x → 1): The super-linear exponent (x1.20) concentrates MiMo's task vector with α₃(31) ≈ 0.652 into the upper decoders, governing semantic reasoning, multi-turn planning, and final token synthesis.
3. Saliency Trimming (TIES-DELLA Pruning)
To eliminate parameter interference and cross-talk, task vectors are pruned based on parameter energy. Given density parameter ρ = 0.70, an update threshold γk is computed per tensor:
Updates below the top 70% magnitude are zeroed out via a saliency mask:
4. Consensus Sign Election & Disjoint Averaging
Surviving task vectors often conflict in directional signs, causing mutual cancellation when averaged naively. A directional consensus sign vector Γ is elected:
A binary agreement mask Ak discards parameter updates that oppose the elected consensus sign:
The merged task delta is reconstructed using only parameters aligned with the majority direction:
The dense layer weights are restored onto the base foundation:
5. Gated DeltaNet (GDN) Gate Norm Stabilization
In linear attention layers, gate matrices control state retention and output gating via non-linear sigmoid activations. Direct delta merging shifts the operator norm, causing activation saturation or exploding outputs. To guarantee numerical stability, the merged gate weight Wgate, unscaled = W₀ + ∑k αk τk is projected onto the base tensor's Frobenius norm:
6. Log-Decay and Normalization Parameter Convexity
For state-space logarithmic decay tensors (Alog ∈ (-∞, 0]), biases, and layer normalization parameters, delta blending can violate mathematical boundary constraints. These tensors are merged strictly via convex interpolation:
Because ∑k αk(l) = 1.0, αk(l) ≥ 0, and Dk, ij ≤ 0 for all decay parameters:
This guarantees Bounded-Input Bounded-Output (BIBO) stability and prevents exponential divergence in recurrent linear attention states.
Layer-Stratified Component Policies
| Parameter Group | Target Identifiers | Applied Policy | Density (ρ) | Mathematical Invariant |
|---|---|---|---|---|
| Embeddings & LM Head | embed_tokens, lm_head |
Convex Blend | — | Fixed weights: 50% Qwopus, 30% MiMo, 20% Ornith. |
| Dense MLPs & Self-Attention | self_attn, mlp.gate_proj, up_proj, down_proj |
TIES-DELLA | 0.70 | Saliency pruning + consensus sign election. |
| Recurrent Linear Attention | linear_attn.in_proj_*, out_proj, conv1d |
Recurrent Delta | — | Unpruned linear delta accumulation. |
| DeltaNet Attention Gates | attn_output_gate |
Norm-Stabilized | — | Projected onto base Frobenius norm ||W₀||F. |
| Decay Rates & Normalizations | A_log, norm, bias |
Convex Blend | — | Enforces Alog ≤ 0 to preserve recurrent stability. |
Agentic Chat Template
This model uses the Improved Chat Template for Qwen 3.x by Olivia Rossi to support multi-tier Chain-of-Thought (CoT) reasoning, dual-format agentic tool execution, automatic error-recovery heuristics, and strict token-waste elimination.
Recommended Generation Parameters
The following parameters are optimal for code synthesis, terminal agent execution, and complex reasoning:
| Parameter | Recommended Setting | Operational Function |
|---|---|---|
| Temperature | 0.6 |
Balances deterministic syntax structure with creative algorithmic pathing. |
| Top-P | 0.95 |
Nucleus sampling cutoff to discard degenerate token tails. |
| Top-K | 20 |
Restricts sampling pool to top candidates, preventing syntactic drift. |
| Min-P | 0.0 (Off) |
Disables relative thresholding in favor of Top-K / Top-P governance. |
| Repetition Penalty | Off (1.0) |
Disabled to prevent penalty distortion on repeated syntax (braces, boilerplate). |
| Presence Penalty | Off (0.0) |
Preserves deterministic variable and function naming across long contexts. |
How to Use
Serving via vLLM
vllm serve pragmaticcs/Triumvirate \
--dtype bfloat16 \
--max-model-len 65536 \
--gpu-memory-utilization 0.95 \
--enable-prefix-caching \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--enable-reasoning \
--reasoning-parser qwen3
Inference via Transformers
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "pragmaticcs/Triumvirate"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
messages = [
{
"role": "system",
"content": "You are a principal software engineer. Think carefully before outputting production-grade code."
},
{
"role": "user",
"content": "Implement an asynchronous token-bucket rate limiter in Python supporting burst handling and thread-safe redis synchronization."
}
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
enable_thinking=True,
return_tensors="pt"
).to(model.device)
outputs = model.generate(
inputs,
max_new_tokens=2048,
temperature=0.6,
top_p=0.95,
top_k=20,
do_sample=True,
)
response = tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True)
print(response)
Citation & References
- Qwen/Qwen3.5-9B
- XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
- ornith-ai/Ornith-1.5-9B
- Jackrong/Qwopus3.5-9B-Coder
- Improved Chat Template for Qwen 3.x
@inproceedings{yadav2023ties,
title={Resolving Interference When Merging Models},
author={Yadav, Prateek and Tam, Derek and Choshen, Leshem and Raffel, Colin and Bansal, Mohit},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
volume={36},
pages={7093--7115},
year={2023}
}
@article{deep2024della,
title={DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling},
author={Deep, Pala Tej and Bhardwaj, Rishabh and Poria, Soujanya},
journal={arXiv preprint arXiv:2406.11617},
year={2024}
}
@inproceedings{yu2024dare,
title={Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch},
author={Yu, Le and Yu, Bowen and Yu, Haiyang and Huang, Fei and Li, Yongbin},
booktitle={International Conference on Machine Learning (ICML)},
year={2024}
}
- Downloads last month
- 439
