POCKET-Darwin-180B: a clean R0 vs R3 test on held-out SuperGPQA (+2.45 points)

#1
by SeaWolf-AI - opened
FINAL_Bench org

We released POCKET-Darwin-180B, a 4-bit GGUF build of Darwin-180B-RSI (R3) that runs on a laptop or a CPU-only server.
https://huggingface.co/FINAL-Bench/POCKET-Darwin-180B-GGUF

How it was built makes it a useful test of RSI itself. We took the public Unsloth UD-Q4_K_XL build of the parent model (Qwen3.8-Flash-Next) and replaced only the 300 tensors that our self-improvement training changed. Everything else is byte-identical. So comparing POCKET with the parent 4-bit build isolates the effect of RSI: same quantized experts, same everything, except the 2.78% of bytes RSI touched.

Held-out SuperGPQA, 1,000 questions never used in training or selection, 4 samples each, 16K generation budget:

  • Parent 4-bit: 59.10%
  • POCKET (R3): 61.55%
  • Paired difference: +2.45 points, 95% CI [+1.54, +3.36]

Where the gain comes from: POCKET reasons more concisely, so more of its answers finish within the same budget (675 vs 798 of 4,000 samples reached the 16K cap). It uses about 13% fewer tokens per question (geometric mean). The largest gains are on the longest questions, where the parent often runs out of budget before answering.

On MMLU-Pro (2,000 questions) the two builds score the same (87.65% vs 87.95%, within noise), while POCKET's correct answers are shorter: short answers are unchanged, and answers past roughly 700 tokens get 11 to 26% shorter.

These numbers were checked in public. A community member, Dipankar Sarkar, verified the graft byte by byte and reconciled our tables in the discussion here:
https://huggingface.co/posts/SeaWolf-AI/446226611129845

Next: a greedy-decoding comparison to locate where in the reasoning the saving happens (before or after the model first commits to an answer). We will post the results in that thread.

Sign up or log in to comment