Jebadiah 9B v2 MLX

An MLX build of Jebadiah 9B v2 for Apple silicon, in the folder 8bit/ (group size 64). Jebadiah answers a typed question (choice, noul or score) with a probability for every option, read from one forward pass. Nothing is generated. Code, trainer and evals: getainode/jebadiah.

Results and docs

  • Project site, with every result and how to run the models: jebadiah.ai.
  • JevBench v1.4.2: on its 231 public items, run through its own harness, Jebadiah 9B v2 scores 0.818 (Jebadiah 27B: 0.866, the same as Jev 1.13.0). This is my own run on the public items, not the official board, which also uses sealed items. Details and caveats.
  • JDE blind test (as of 2026-09-26): on 290 real decisions from Titanium Computing's production decision engine, Jebadiah 9B v2 got 277 against Jev's 282, with 0 flips across 5,800 repeat calls. Jev still leads on coverage checks.
  • For the family: Decision Index 0.2.1: the 27B scores 54.67, #5 of 67 open models (as of 2026-09-26). This is the board's own number; the maintainer validated my run and put it on the leaderboard. Run record.
  • All sizes: the Hugging Face collection, mirrored on ModelScope.

Which one should I use? For local use, start with Jebadiah 9B v2 GGUF. On Apple silicon, use an MLX build: 27B, 9B v2 or 4B v2. For vLLM or fine-tuning, use the full weights: 27B, 9B v2 or 4B v2. Setup for every runtime, and for JDE: Run Jeb locally.

Files

Every file was checked on the 260 held-out questions the merged weights were checked on, and compared with the merged bf16 weights and with the training run's own eval records.

File Size Same answer as bf16 Same as the run choice + noul score Prob. diff median / max
8bit/ 9.5 GB 256 / 260 258 / 260 173 / 173 85 / 87 0.003 / 0.062
bf16 weights 257 / 260 172 / 173 85 / 87 0.002 / 0.029

8bit needs about its own size in unified memory, plus about 1 GB for a 2k-token prompt. On a Mac with less memory, run the GGUF build with llama.cpp instead.

A 4bit build was measured and not published: 224 of 260 answers the same as bf16, and 154 of 173 choice and noul answers the same as the run's records (under 90%).

Run it

The script renders the prompt exactly as AINode does, runs one forward pass with mlx-lm, multiplies the last hidden state by the option labels' output-head rows in fp32 and applies temperatures.json (choice 1.1863, noul 1.0903, score 0.8329). Needs mlx-lm 0.31 or newer (qwen3_5 support).

pip install "mlx-lm>=0.31"
hf download frontier-infra/jebadiah-9b-v2-MLX --include "8bit/*" --include "scripts/*" --include temperatures.json --local-dir jebadiah-9b-v2-MLX
cd jebadiah-9b-v2-MLX
python scripts/decide_mlx.py --model 8bit --request scripts/example-request.json

--no-temperatures returns the raw probabilities.

HelpSteer2-like traffic. A held-out HelpSteer2 check (2026-09-29): on the 418 nvidia/HelpSteer2 validation rows that no reported evaluation set uses (natural label mix, never in the training pool), this model's score answers want a temperature of 1.29, not 0.83, and still 1.11 after reweighting to a flat label mix. If your traffic looks like HelpSteer2, keep the older score temperature, 1.22 (the train fit in temperatures.json; every script here takes --temperatures with a file of your own), or refit on your own labels. v3's calibration split will draw HelpSteer2 from held-out data.

On example-request.json (8bit/):

{
 "route": {"type": "choice", "choice": "billing", "confidence": 0.460831, "probabilities": {"billing": 0.640554, "support": 0.037693, "sales": 0.321753}},
 "urgent": {"type": "noul", "noul": 0.165308}
}

How it was measured

Jevals PubMedQA, Banking77 (77 options) and HelpSteer2, plus Nimble: the merge check's fixed sample (seed 20260925), the run's option order and the run's temperatures, so the probability differences compare like with like (the shipped temperatures.json since 2026-09-29 uses 0.8329 for score; top picks do not depend on it). "Same answer" is the top option; "prob. diff" is the largest change on any option against the run's CUDA record. Records: eval/agreement-*.json.

License

Apache-2.0, as the base model. Made in Texas.

Downloads last month

-

Downloads are not tracked for this model. How to track
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for frontier-infra/jebadiah-9b-v2-MLX

Finetuned
Qwen/Qwen3.5-9B
Quantized
(4)
this model

Collection including frontier-infra/jebadiah-9b-v2-MLX