ruos-foundry-cu-qwen3-30b-a3b-e64

ruos-foundry-cu-qwen3-30b-a3b-e64 is a ruOS specialist cut from Qwen/Qwen3-30B-A3B-Instruct-2507 with MoE-Foundry (ADR-064): the parent's routed experts were traced on the ruOS ruos-cu calibration set, the 64 most-used experts per routed layer were kept, and a complete smaller Mixture-of-Experts model was exported โ€” same backbone, same tokenizer, same token top-k, fewer experts per layer. No weight was changed; experts were removed and renumbered and the router rows sliced to match.

Status: unevaluated and disabled

Per MoE-Foundry's rule an exported specialist starts quality_status: unevaluated, enabled: false. A structural export proves tensor integrity, not retained capability. This checkpoint is published so that slim-eval (slim/eval: run, redteam, regression, verdict) can measure it against the parent on the frozen ruOS test split; until that verdict is recorded here it must not be routed to. The full parent remains the fallback.

Parent

repo Qwen/Qwen3-30B-A3B-Instruct-2507
revision 0d7cf23991f47feeb3a57ecb4c9cee8ea4a17bfe
licence Apache-2.0
MoE-Foundry parent_id 44c2b6971fe6979409ac72586db5fa4abfd73746afc52607368a59740848149d
experts per routed layer (parent โ†’ this) 128 โ†’ 64
routed layers 48
token top-k (unchanged) 8

Calibration

Domain ruos-cu: 148 rows (148 from the val split, 0 topped up from train; the test split is never used for calibration), ~63523 tokens, families {"browser_step":142,"cu_step":6}, licences {"project-owned":148}. Calibration file sha256 94a5527b141269ab5078a14e3d129b387aa69412c65334160715fd9a58dbfef5. Texts are the ChatML prompt plus the reference answer.

Router traces: 412 tasks, 148144 tokens, 7110912 rows on NVIDIA A100-SXM4-80GB (bfloat16, transformers 4.51.3); trace sha256 7aabb05cbf715156437393fc5a5b0d756951ecee35eb4542a06fc812f11f9403.

Selection method: mass (accumulated routing probability per expert per layer) โ€” a usage proxy, not causal importance.

Receipt

specialist_id 1409f5fef6a8bfbef4be7f1f2e91700f7c3e7d15690ec8a01a37ed2a1c6ca1ea
checkpoint_id 5a74281b02e586e0c6a9e79323f61c90884e188c6b4719fd91ad600b94d2c0da
mask_sha256 d8ea421921a28f6a9c4f0ae3f1bc01efc159fad83163d73c8a2629a3f0b9d721
parent_id 44c2b6971fe6979409ac72586db5fa4abfd73746afc52607368a59740848149d
input tensor bytes 61064245248
output tensor bytes 32060633088
reduction 47.5%

separator_receipt.json in this repo is the full MoE-Foundry receipt including the per-layer expert mask.

Files

path bytes sha256
LICENSE 11343 05cab46843576551502bfdf712f84e93e6e9590d9997306ed4f6635ef82811d9
SHA256SUMS 737 7761e07f29ce824f3893518837e285cc4d1a4c60ba1f3e295d00945c7618a10a
config.json 963 f0bf60e7da89b14b5763344441e4f60eccda09f62c0111c5f178283c5f3d81ec
generation_config.json 239 19d306dd769db12a9d710b44cf7f83b635efbe5166b84fb4358a08fb7d88bb53
merges.txt 1671839 599bab54075088774b1733fde865d5bd747cbcc7a547c5bc12610e874e26f5e3
model.safetensors 32061839304 54ae831f3df008547d5734fab1a2bd3a50349f2dbcb481ad418a57e318a18742
separator_receipt.json 32964 dd6ec404f4400ee8e7e26cb40a2d6e4c3581c4c165232042ff89c379f17b9b56
tokenizer.json 11422654 aeb13307a71acd8fe81861d94ad54ab689df773318809eed3cbe794b4492dae4
tokenizer_config.json 9377 a62ff0a2472a0fa1b8eaabcb57c59b58afa42a22831dc141400b6e0cf2b65ce3
vocab.json 2776833 ca10d7e9fb3ed18575dd1e277a2579c16d108e32f27439684afa0e10b1440910

Run

vllm serve ruvnet/ruos-foundry-cu-qwen3-30b-a3b-e64 --max-model-len 8192

Loads with transformers as a standard qwen3_moe checkpoint (single safetensors file, num_experts reduced in config.json).

Limitations

  • Unevaluated: no capability, memory or latency claim is made here.
  • Retained experts were chosen by routing mass on ruOS calibration prompts; requests outside that domain should go to the parent.
  • Memory: fewer experts means a smaller checkpoint; loading several specialists next to the parent can use more total memory than the parent alone.

Provenance

  • MoE-Foundry 6677a25 (moe-separator inspect โ†’ profile_hf โ†’ select โ†’ export โ†’ mixture)
  • run foundry-20260907T172935Z-qwen3-30b-a3b on a single vast.ai GPU; ruos-desktop slim/foundry + slim/scripts/foundry-e2e.sh
  • authorisation: rUv, "implement this using ruvnet/MoE-Foundry using vast.ai in a worktree, implement e2e and push models to repo"
Downloads last month
296
Safetensors
Model size
16B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ruvnet/ruos-foundry-cu-qwen3-30b-a3b-e64

Finetuned
(98)
this model