Triumvirate-Qwopus-MiMo-Ornith-9B-Coder

License Library Merge Method Architecture Attention Context

Most sub-10B coding models crumble the moment they enter real-world agentic workflows: they either produce clean code but loop endlessly when a shell command fails, or handle tool calls reasonably well while hallucinating obscure API syntax.

Triumvirate is a merge designed to solve that dilemma. It combines three of the most capable specialized fine-tunes of Qwen 3.5 9B and fuses their task vectors directly into the base backbone:

  • Algorithmic & Syntax Precision from Qwopus
  • SWE-bench Problem Decomposition & Tool Calling from MiMo-V2.6
  • Loop-Termination & Error-Recovery Discipline from Ornith-1.5

The result is a lean, blisteringly fast 9B pure-text causal engine with a native 256k context window that runs comfortably on consumer GPUs.


Contents


Architectural Specifications

Parameter Specification
Total Parameters 8.8B (Text Backbone)
Architecture Type Dense Causal Language Model (qwen3_5_text)
Hidden Dimension (dmodel) 4096
Intermediate Dimension (dmlp) 12288 (SwiGLU)
Decoder Layers 32
Attention Mechanism Hybrid Gated DeltaNet (3 Linear Attention : 1 Full Attention)
Full Attention Layers Layers 3, 7, 11, 15, 19, 23, 27, 31
Linear Attention Heads 16 Key Heads / 32 Value Heads (dk = dv = 128)
Full Attention Heads 16 Query / 4 Key-Value (GQA, dh = 256)
Rotary Position Embedding (RoPE) 1D Partial RoPE (θ = 10⁷, Factor = 0.25)
Maximum Sequence Length 262,144 tokens (256k)
Native Precision bfloat16

Composition & Donor Weighting

The foundation checkpoint serves as the structural base (W₀). Three donor models contribute directional task vectors weighted continuously across network depth:

Model Role Specialization Focus Depth Target
Qwen/Qwen3.5-9B Base Anchor (W₀) Structural anchor & GDN linear attention state Global
Jackrong/Qwopus3.5-9B-Coder Donor 1 (D₁) Claude 3.5 Opus distillation; typing, syntax, algorithms Lower Layers (x ≤ 0.35)
ornith-ai/Ornith-1.5-9B Donor 2 (D₂) Agentic RL; loop-termination & error-pivot discipline Mid Layers (0.35 < x < 0.70)
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B Donor 3 (D₃) 77.4B tokens SFT; SWE-bench Pro, multi-turn tool logic Top Layers (x ≥ 0.70)

Merge Methodology & Mathematical Formulation

The merge combines TIES-DELLA saliency trimming, consensus sign election, Gated DeltaNet norm stabilization, and continuous sinusoidal depth modulation.

1. Task Vector Formulation

For each donor checkpoint k ∈ {1, 2, 3}, the parameter update delta is isolated relative to the base anchor W₀:

τk=Dk−W0,k∈{MiMo,Ornith,Qwopus} \tau_k = D_k - W_0, \quad k \in \{\text{MiMo}, \text{Ornith}, \text{Qwopus}\}

2. Asymmetric Sinusoidal Depth Modulation

Task vector mixing coefficients are continuously parameterized over normalized network depth x = l / (L - 1), where l ∈ {0, 1, ..., 31} and L = 32:

uqwopus(x)=0.45cos⁡2(π2x)+0.25 u_{\text{qwopus}}(x) = 0.45 \cos^2\left(\frac{\pi}{2} x\right) + 0.25

uornith(x)=0.35sin⁡2(πx0.85)+0.15 u_{\text{ornith}}(x) = 0.35 \sin^2\left(\pi x^{0.85}\right) + 0.15

umimo(x)=0.55sin⁡2(π2x1.20)+0.20 u_{\text{mimo}}(x) = 0.55 \sin^2\left(\frac{\pi}{2} x^{1.20}\right) + 0.20

The donor weights αk(l) are normalized to form a partition of unity across all layers:

αk(l)=uk(x)∑j=13uj(x),∑k=13αk(l)=1.0 \alpha_k(l) = \frac{u_k(x)}{\sum_{j=1}^3 u_j(x)}, \quad \sum_{k=1}^3 \alpha_k(l) = 1.0

  • Lower Layers (x → 0): Qwopus dominates with α₁(0) ≈ 0.67, ensuring foundational language representations and syntax heads are grounded in Claude 3.5 Opus traces.
  • Middle Layers (x ≈ 0.5): The sub-linear exponent (x0.85) accelerates Ornith's activation to peak across middle transformer blocks with α₂(16) ≈ 0.354, reinforcing state-space continuity and execution discipline.
  • Top Layers (x → 1): The super-linear exponent (x1.20) concentrates MiMo's task vector with α₃(31) ≈ 0.652 into the upper decoders, governing semantic reasoning, multi-turn planning, and final token synthesis.

3. Saliency Trimming (TIES-DELLA Pruning)

To eliminate parameter interference and cross-talk, task vectors are pruned based on parameter energy. Given density parameter ρ = 0.70, an update threshold γk is computed per tensor:

γk=Quantile1−ρ({∣τk,ij∣}) \gamma_k = \text{Quantile}_{1 - \rho}\left(\{|\tau_{k, ij}|\}\right)

Updates below the top 70% magnitude are zeroed out via a saliency mask:

Mk=I(∣τk∣≥γk) M_k = \mathbb{I}\left(|\tau_k| \ge \gamma_k\right)

τ^k=τk⊙Mk \hat{\tau}_k = \tau_k \odot M_k

4. Consensus Sign Election & Disjoint Averaging

Surviving task vectors often conflict in directional signs, causing mutual cancellation when averaged naively. A directional consensus sign vector Γ is elected:

Γ=sgn⁡(∑k=13αk(l)τ^k) \Gamma = \operatorname{sgn}\left(\sum_{k=1}^3 \alpha_k(l) \hat{\tau}_k\right)

A binary agreement mask Ak discards parameter updates that oppose the elected consensus sign:

Ak=I(sgn⁡(τ^k)=Γ)⊙I(τ^k≠0) A_k = \mathbb{I}\left(\operatorname{sgn}(\hat{\tau}_k) = \Gamma\right) \odot \mathbb{I}\left(\hat{\tau}_k \neq 0\right)

The merged task delta is reconstructed using only parameters aligned with the majority direction:

ΔTIES={∑k=13αk(l)(τ^k⊙Ak)∑k=13αk(l)Akif ∑k=13αk(l)Ak>00otherwise \Delta_{\text{TIES}} = \begin{cases} \frac{\sum_{k=1}^3 \alpha_k(l) \left(\hat{\tau}_k \odot A_k\right)}{\sum_{k=1}^3 \alpha_k(l) A_k} & \text{if } \sum_{k=1}^3 \alpha_k(l) A_k > 0 \\ 0 & \text{otherwise} \end{cases}

The dense layer weights are restored onto the base foundation:

Wdense=W0+ΔTIES W_{\text{dense}} = W_0 + \Delta_{\text{TIES}}

5. Gated DeltaNet (GDN) Gate Norm Stabilization

In linear attention layers, gate matrices control state retention and output gating via non-linear sigmoid activations. Direct delta merging shifts the operator norm, causing activation saturation or exploding outputs. To guarantee numerical stability, the merged gate weight Wgate, unscaled = W₀ + ∑k αk τk is projected onto the base tensor's Frobenius norm:

Wgate=Wgate, unscaled⋅∥W0∥F∥Wgate, unscaled∥F W_{\text{gate}} = W_{\text{gate, unscaled}} \cdot \frac{\|W_0\|_F}{\|W_{\text{gate, unscaled}}\|_F}

6. Log-Decay and Normalization Parameter Convexity

For state-space logarithmic decay tensors (Alog ∈ (-∞, 0]), biases, and layer normalization parameters, delta blending can violate mathematical boundary constraints. These tensors are merged strictly via convex interpolation:

Wconvex=∑k=13αk(l)Dk W_{\text{convex}} = \sum_{k=1}^3 \alpha_k(l) D_k

Because ∑k αk(l) = 1.0, αk(l) ≥ 0, and Dk, ij ≤ 0 for all decay parameters:

∑k=13αk(l)Dk,ij≤max⁡k(Dk,ij)≤0  ⟹  exp⁡(Wconvex,ij)∈(0,1] \sum_{k=1}^3 \alpha_k(l) D_{k, ij} \le \max_k(D_{k, ij}) \le 0 \implies \exp\left(W_{\text{convex}, ij}\right) \in (0, 1]

This guarantees Bounded-Input Bounded-Output (BIBO) stability and prevents exponential divergence in recurrent linear attention states.


Layer-Stratified Component Policies

Parameter Group Target Identifiers Applied Policy Density (ρ) Mathematical Invariant
Embeddings & LM Head embed_tokens, lm_head Convex Blend — Fixed weights: 50% Qwopus, 30% MiMo, 20% Ornith.
Dense MLPs & Self-Attention self_attn, mlp.gate_proj, up_proj, down_proj TIES-DELLA 0.70 Saliency pruning + consensus sign election.
Recurrent Linear Attention linear_attn.in_proj_*, out_proj, conv1d Recurrent Delta — Unpruned linear delta accumulation.
DeltaNet Attention Gates attn_output_gate Norm-Stabilized — Projected onto base Frobenius norm ||W₀||F.
Decay Rates & Normalizations A_log, norm, bias Convex Blend — Enforces Alog ≤ 0 to preserve recurrent stability.

Agentic Chat Template

This model uses the Improved Chat Template for Qwen 3.x by Olivia Rossi to support multi-tier Chain-of-Thought (CoT) reasoning, dual-format agentic tool execution, automatic error-recovery heuristics, and strict token-waste elimination.


Recommended Generation Parameters

The following parameters are optimal for code synthesis, terminal agent execution, and complex reasoning:

Parameter Recommended Setting Operational Function
Temperature 0.6 Balances deterministic syntax structure with creative algorithmic pathing.
Top-P 0.95 Nucleus sampling cutoff to discard degenerate token tails.
Top-K 20 Restricts sampling pool to top candidates, preventing syntactic drift.
Min-P 0.0 (Off) Disables relative thresholding in favor of Top-K / Top-P governance.
Repetition Penalty Off (1.0) Disabled to prevent penalty distortion on repeated syntax (braces, boilerplate).
Presence Penalty Off (0.0) Preserves deterministic variable and function naming across long contexts.

How to Use

Serving via vLLM

vllm serve pragmaticcs/Triumvirate \
  --dtype bfloat16 \
  --max-model-len 65536 \
  --gpu-memory-utilization 0.95 \
  --enable-prefix-caching \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --enable-reasoning \
  --reasoning-parser qwen3

Inference via Transformers

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "pragmaticcs/Triumvirate"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

messages = [
    {
        "role": "system",
        "content": "You are a principal software engineer. Think carefully before outputting production-grade code."
    },
    {
        "role": "user",
        "content": "Implement an asynchronous token-bucket rate limiter in Python supporting burst handling and thread-safe redis synchronization."
    }
]

inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    enable_thinking=True,
    return_tensors="pt"
).to(model.device)

outputs = model.generate(
    inputs,
    max_new_tokens=2048,
    temperature=0.6,
    top_p=0.95,
    top_k=20,
    do_sample=True,
)

response = tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True)
print(response)

Citation & References

@inproceedings{yadav2023ties,
  title={Resolving Interference When Merging Models},
  author={Yadav, Prateek and Tam, Derek and Choshen, Leshem and Raffel, Colin and Bansal, Mohit},
  booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
  volume={36},
  pages={7093--7115},
  year={2023}
}

@article{deep2024della,
  title={DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling},
  author={Deep, Pala Tej and Bhardwaj, Rishabh and Poria, Soujanya},
  journal={arXiv preprint arXiv:2406.11617},
  year={2024}
}

@inproceedings{yu2024dare,
  title={Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch},
  author={Yu, Le and Yu, Bowen and Yu, Haiyang and Huang, Fei and Li, Yongbin},
  booktitle={International Conference on Machine Learning (ICML)},
  year={2024}
}
Downloads last month
439
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder

Collection including pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder

Paper for pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder