Instructions to use frontier-infra/jebadiah-9b-v2-MLX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use frontier-infra/jebadiah-9b-v2-MLX with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir jebadiah-9b-v2-MLX frontier-infra/jebadiah-9b-v2-MLX
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Jebadiah 9B v2 MLX
An MLX build of Jebadiah 9B v2 for Apple silicon, in the folder 8bit/ (group size 64).
Jebadiah answers a typed question (choice, noul or score) with a probability for every option, read from one
forward pass. Nothing is generated. Code, trainer and evals: getainode/jebadiah.
Results and docs
- Project site, with every result and how to run the models: jebadiah.ai.
- JevBench v1.4.2: on its 231 public items, run through its own harness, Jebadiah 9B v2 scores 0.818 (Jebadiah 27B: 0.866, the same as Jev 1.13.0). This is my own run on the public items, not the official board, which also uses sealed items. Details and caveats.
- JDE blind test (as of 2026-09-26): on 290 real decisions from Titanium Computing's production decision engine, Jebadiah 9B v2 got 277 against Jev's 282, with 0 flips across 5,800 repeat calls. Jev still leads on coverage checks.
- For the family: Decision Index 0.2.1: the 27B scores 54.67, #5 of 67 open models (as of 2026-09-26). This is the board's own number; the maintainer validated my run and put it on the leaderboard. Run record.
- All sizes: the Hugging Face collection, mirrored on ModelScope.
Which one should I use? For local use, start with Jebadiah 9B v2 GGUF. On Apple silicon, use an MLX build: 27B, 9B v2 or 4B v2. For vLLM or fine-tuning, use the full weights: 27B, 9B v2 or 4B v2. Setup for every runtime, and for JDE: Run Jeb locally.
Files
Every file was checked on the 260 held-out questions the merged weights were checked on, and compared with the merged bf16 weights and with the training run's own eval records.
| File | Size | Same answer as bf16 | Same as the run | choice + noul | score | Prob. diff median / max |
|---|---|---|---|---|---|---|
8bit/ |
9.5 GB | 256 / 260 | 258 / 260 | 173 / 173 | 85 / 87 | 0.003 / 0.062 |
| bf16 weights | 257 / 260 | 172 / 173 | 85 / 87 | 0.002 / 0.029 |
8bit needs about its own size in unified memory, plus about 1 GB for a 2k-token prompt. On a Mac with
less memory, run the GGUF build with llama.cpp instead.
A 4bit build was measured and not published: 224 of 260 answers the same as bf16, and 154 of 173 choice and noul answers the same as the run's records (under 90%).
Run it
The script renders the prompt exactly as AINode does, runs one forward pass with mlx-lm, multiplies the
last hidden state by the option labels' output-head rows in fp32 and applies temperatures.json
(choice 1.1863, noul 1.0903, score 0.8329). Needs mlx-lm 0.31 or newer (qwen3_5 support).
pip install "mlx-lm>=0.31"
hf download frontier-infra/jebadiah-9b-v2-MLX --include "8bit/*" --include "scripts/*" --include temperatures.json --local-dir jebadiah-9b-v2-MLX
cd jebadiah-9b-v2-MLX
python scripts/decide_mlx.py --model 8bit --request scripts/example-request.json
--no-temperatures returns the raw probabilities.
HelpSteer2-like traffic. A held-out HelpSteer2 check (2026-09-29): on the 418 nvidia/HelpSteer2 validation rows that no reported evaluation set uses (natural label mix, never in the training pool), this model's score answers want a temperature of 1.29, not 0.83, and still 1.11 after reweighting to a flat label mix. If your traffic looks like HelpSteer2, keep the older score temperature, 1.22 (the train fit in temperatures.json; every script here takes --temperatures with a file of your own), or refit on your own labels. v3's calibration split will draw HelpSteer2 from held-out data.
On example-request.json (8bit/):
{
"route": {"type": "choice", "choice": "billing", "confidence": 0.460831, "probabilities": {"billing": 0.640554, "support": 0.037693, "sales": 0.321753}},
"urgent": {"type": "noul", "noul": 0.165308}
}
How it was measured
Jevals PubMedQA, Banking77 (77 options) and HelpSteer2, plus Nimble: the merge check's fixed sample (seed
20260925), the run's option order and the run's temperatures, so the probability differences compare like with like (the shipped temperatures.json since 2026-09-29 uses 0.8329 for score; top picks do not depend on it). "Same answer" is the top option; "prob. diff" is the
largest change on any option against the run's CUDA record. Records: eval/agreement-*.json.
License
Apache-2.0, as the base model. Made in Texas.
Quantized