Model Card: cloudyu/GPT-OSS-120B-MLX-q4-Claude-4.6-Opus-Reasoning-Distilled

This is a of only 60G size LoRA‑fine‑tuned version of openai/gpt-oss-120b (converted to MLX 4‑bit format) that has been specialised for step‑by‑step reasoning on mathematical, coding and algorithmic problems.
The adapter was trained using the mlx‑tune framework on Apple Silicon (M‑series).


Model Description

  • Base model : openai/gpt-oss-120b (MLX quantised, 4‑bit)
  • Fine‑tuning method : LoRA (rank 16, α 32) on attention + MoE projection layers, plus full‑precision fine‑tuning of the router weights
  • Training dataset : nohurry/Opus-4.6-Reasoning-3000x-filtered, filtered to prioritise code & mathematics samples (target 80%)
  • Optimiser : Muon (for LoRA matrices) + AdamW (for router and 1‑D parameters)
  • Context length : 4 096 tokens
  • Training hardware : Apple M‑series (≥64 GB unified memory)
  • Training steps : 200 steps (≈ 1 epoch)

Chat Template – Harmony

The model uses the Harmony token system (not ChatML).
Message format:

<|start|>role<|channel|>channel<|message|>content<|end|>
  • role : user | assistant | system
  • channel :
    • query – user question
    • analysis – assistant reasoning chain
    • final – assistant final answer
    • system – system prompt

Example:

<|start|>system<|channel|>system<|message|>You are a helpful assistant.<|end|>
<|start|>user<|channel|>query<|message|>What is 2+2?<|end|>
<|start|>assistant<|channel|>analysis<|message|>I need to add 2 and 2.<|end|>
<|start|>assistant<|channel|>final<|message|>4<|end|>

Using LM Studio with OpenAI Codex on Mac (MLX Models)

A step-by-step guide to running cloudyu/gpt-oss-120b-Sonnet-Reasoning-Distilled locally on Apple Silicon Mac and connecting it to OpenAI Codex CLI.


Prerequisites

  • Apple Silicon Mac (M1 / M2 / M3 / M4)
  • macOS 13 or later
  • At least 64 GB unified memory (the 120B model requires ~60 GB)
  • Homebrew installed
  • Node.js 18+ (for Codex CLI)

Part 1 — Install LM Studio

1.1 Download and install the app

Go to https://lmstudio.ai/download and download the macOS (Apple Silicon) installer.
Drag LM Studio.app into your /Applications folder and open it at least once — this bootstraps the lms CLI.

1.2 Add lms to your PATH

Open a terminal and run:

npx lmstudio install-cli

Open a new terminal window, then verify:

lms --version

Part 2 — Download the Model

You have two options: download via the LM Studio app GUI, or use the CLI.

Option A — GUI (easiest)

  1. Open LM Studio.
  2. Press ⌘ + Shift + M to open the model search.
  3. Search for gpt-oss-120b-Sonnet-Reasoning-Distilled.
  4. Select the MLX variant and click Download.

Option B — CLI

lms get cloudyu/gpt-oss-120b-Sonnet-Reasoning-Distilled --mlx

Option C — Use a locally downloaded model

If you already have the model folder on disk, create a symlink into LM Studio's model directory:

# Create the target directory
mkdir -p ~/Documents/LM\ Studio/models/cloudyu/

# Symlink (no file copying, saves disk space)
ln -s /path/to/your/gpt-oss-120b-Sonnet-Reasoning-Distilled \
  ~/Documents/LM\ Studio/models/cloudyu/gpt-oss-120b-Sonnet-Reasoning-Distilled

Verify LM Studio detected the model:

lms ls

You should see something like:

LLM                                        PARAMS    ARCH       SIZE        DEVICE
gpt-oss-120b-sonnet-reasoning-distilled    120B      gpt_oss    63.42 GB    Local

Part 3 — Load the Model and Start the Server

3.1 (Optional) Estimate memory usage first

lms load --estimate-only gpt-oss-120b-sonnet-reasoning-distilled

3.2 Load the model

lms load gpt-oss-120b-sonnet-reasoning-distilled \
  --context-length 32768 \
  --gpu max
  • --context-length 32768 — Codex needs a large context window.
  • --gpu max — offloads all layers to Apple Metal GPU (recommended for MLX models).

Wait for: Model loaded successfully.

3.3 Start the local server

lms server start --port 1234

3.4 Verify the API is working

curl http://127.0.0.1:1234/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer lm-studio" \
  -d '{
    "model": "gpt-oss-120b-sonnet-reasoning-distilled",
    "messages": [{"role": "user", "content": "hello"}],
    "max_tokens": 50
  }'

You should receive a JSON response with the model's reply.


Part 4 — Install and Configure OpenAI Codex CLI

4.1 Install Codex

npm install -g @openai/codex

Verify:

codex --version

4.2 Configure Codex

Create (or overwrite) the Codex config file:

cat > ~/.codex/config.toml << 'EOF'
# Set this profile as default — note: the key is "profile", not "default_profile"
profile = "gpt-oss-local"

[model_providers.my-lmstudio]
name = "LM Studio Local"
base_url = "http://127.0.0.1:1234/v1"
api_key = "lm-studio"

[profiles.gpt-oss-local]
model_provider = "my-lmstudio"
model = "gpt-oss-120b-sonnet-reasoning-distilled"
context_window = 32000
wire_api = "responses"

[projects."/Users/YOUR_USERNAME/your-project"]
trust_level = "trusted"
EOF

Important notes:

  • The correct key for a default profile is profile, not default_profile.
  • my-lmstudio is a custom name. Do not use reserved names: openai, ollama, or lmstudio.
  • Replace /Users/YOUR_USERNAME/your-project with your actual project path.
  • wire_api = "responses" tells Codex to use the /v1/responses endpoint, which LM Studio supports natively.

4.3 Run Codex

codex

You should see:

╭──────────────────────────────────────────────────────────╮
│ >_ OpenAI Codex (v0.128.0)                               │
│                                                          │
│ model:     gpt-oss-120b-sonnet-reasoning-distilled       │
│ directory: ~/your-project                                │
╰──────────────────────────────────────────────────────────╯

If you see gpt-5.5 in the model field, your profile is not loading. Run with an explicit flag instead:

codex --profile gpt-oss-local

Part 5 — Everyday Workflow

Each time you start a new terminal session, run these two commands before launching Codex:

# 1. Load the model (skip if already loaded)
lms load gpt-oss-120b-sonnet-reasoning-distilled --context-length 32768 --gpu max

# 2. Start the server (skip if already running)
lms server start --port 1234

# 3. Launch Codex
codex

To check whether the model is already loaded:

lms ps

To stop the server:

lms server stop

Troubleshooting

Problem Solution
lms command not found Run npx lmstudio install-cli, then open a new terminal
Model not detected by lms ls Check that your model folder is inside ~/Documents/LM Studio/models/<author>/
Codex shows gpt-5.5 as model Use codex --profile gpt-oss-local or verify profile = in config.toml
Codex request times out Make sure proxy env vars are unset: unset http_proxy https_proxy HTTP_PROXY HTTPS_PROXY
model_providers contains reserved built-in provider IDs Rename your provider — avoid openai, ollama, lmstudio
Out of memory when loading Reduce --context-length (e.g. 16384) or close other apps


在 Mac 上使用 LM Studio 配合 OpenAI Codex(MLX 模型)

本文以 cloudyu/gpt-oss-120b-Sonnet-Reasoning-Distilled 为例,手把手介绍如何在 Apple Silicon Mac 上本地运行大模型并接入 OpenAI Codex CLI。


前提条件

  • Apple Silicon Mac(M1 / M2 / M3 / M4)
  • macOS 13 或更高版本
  • 至少 64 GB 统一内存(120B 模型约需 60 GB)
  • 已安装 Homebrew
  • Node.js 18+(用于 Codex CLI)

第一步 — 安装 LM Studio

1.1 下载并安装 App

前往 https://lmstudio.ai/download,下载 macOS(Apple Silicon)版本安装包。
将 LM Studio.app 拖入 /Applications 文件夹,并至少打开一次,这一步会初始化 lms 命令行工具。

1.2 将 lms 加入 PATH

打开终端,运行:

npx lmstudio install-cli

重新打开一个新终端窗口,验证安装:

lms --version

第二步 — 下载模型

有两种方式,根据你的情况选择:

方式 A — 图形界面(最简单)

  1. 打开 LM Studio。
  2. 按 ⌘ + Shift + M 打开模型搜索。
  3. 搜索 gpt-oss-120b-Sonnet-Reasoning-Distilled。
  4. 选择 MLX 格式,点击 Download。

方式 B — 命令行下载

lms get cloudyu/gpt-oss-120b-Sonnet-Reasoning-Distilled --mlx

方式 C — 使用本地已有模型

如果模型文件夹已经在磁盘上,用软链接导入,无需复制文件:

# 创建目录
mkdir -p ~/Documents/LM\ Studio/models/cloudyu/

# 创建软链接(不占用额外磁盘空间)
ln -s /path/to/your/gpt-oss-120b-Sonnet-Reasoning-Distilled \
  ~/Documents/LM\ Studio/models/cloudyu/gpt-oss-120b-Sonnet-Reasoning-Distilled

验证 LM Studio 已识别到模型:

lms ls

正常输出如下:

LLM                                        PARAMS    ARCH       SIZE        DEVICE
gpt-oss-120b-sonnet-reasoning-distilled    120B      gpt_oss    63.42 GB    Local

第三步 — 加载模型并启动服务器

3.1 (可选)加载前预估内存

lms load --estimate-only gpt-oss-120b-sonnet-reasoning-distilled

3.2 加载模型

lms load gpt-oss-120b-sonnet-reasoning-distilled \
  --context-length 32768 \
  --gpu max
  • --context-length 32768:Codex 需要较大的上下文窗口。
  • --gpu max:将所有层卸载到 Apple Metal GPU,MLX 模型推荐此设置。

等待出现 Model loaded successfully 即加载完成。

3.3 启动本地服务器

lms server start --port 1234

3.4 验证 API 是否正常

curl http://127.0.0.1:1234/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer lm-studio" \
  -d '{
    "model": "gpt-oss-120b-sonnet-reasoning-distilled",
    "messages": [{"role": "user", "content": "你好"}],
    "max_tokens": 50
  }'

收到包含模型回复的 JSON 响应即为成功。


第四步 — 安装并配置 OpenAI Codex CLI

4.1 安装 Codex

npm install -g @openai/codex

验证安装:

codex --version

4.2 配置 Codex

创建(或覆盖)Codex 配置文件:

cat > ~/.codex/config.toml << 'EOF'
# 设置默认 profile,注意字段名是 "profile" 而不是 "default_profile"
profile = "gpt-oss-local"

[model_providers.my-lmstudio]
name = "LM Studio Local"
base_url = "http://127.0.0.1:1234/v1"
api_key = "lm-studio"

[profiles.gpt-oss-local]
model_provider = "my-lmstudio"
model = "gpt-oss-120b-sonnet-reasoning-distilled"
context_window = 32000
wire_api = "responses"

[projects."/Users/你的用户名/你的项目路径"]
trust_level = "trusted"
EOF

重要说明:

  • 默认 profile 的字段名是 profile,不是 default_profile。
  • my-lmstudio 是自定义名称,不能使用保留名称:openai、ollama、lmstudio。
  • 将 /Users/你的用户名/你的项目路径 替换为你实际的项目目录。
  • wire_api = "responses" 指定 Codex 使用 /v1/responses 接口,LM Studio 原生支持此接口。

4.3 启动 Codex

codex

正常启动后应显示:

╭──────────────────────────────────────────────────────────╮
│ >_ OpenAI Codex (v0.128.0)                               │
│                                                          │
│ model:     gpt-oss-120b-sonnet-reasoning-distilled       │
│ directory: ~/你的项目                                     │
╰──────────────────────────────────────────────────────────╯

如果 model 显示的是 gpt-5.5,说明 profile 没有生效,可以手动指定:

codex --profile gpt-oss-local

第五步 — 日常使用流程

每次打开新终端后,按以下顺序执行:

# 1. 加载模型(已加载则跳过)
lms load gpt-oss-120b-sonnet-reasoning-distilled --context-length 32768 --gpu max

# 2. 启动服务器(已运行则跳过)
lms server start --port 1234

# 3. 启动 Codex
codex

检查模型是否已加载:

lms ps

停止服务器:

lms server stop

常见问题排查

问题 解决方法
lms 命令找不到 运行 npx lmstudio install-cli,然后重新打开终端
lms ls 看不到模型 确认模型文件夹位于 ~/Documents/LM Studio/models/<作者名>/ 下
Codex 显示 gpt-5.5 使用 codex --profile gpt-oss-local 或检查 config.toml 里 profile = 字段
Codex 请求超时 清除代理环境变量:unset http_proxy https_proxy HTTP_PROXY HTTPS_PROXY
reserved built-in provider IDs 错误 重命名自定义 provider,避免使用 openai、ollama、lmstudio
加载模型时内存不足 减小 --context-length(如改为 16384),或关闭其他占内存的应用

Intended Use

This model is designed for reasoning‑intensive tasks such as:

  • Solving mathematics competition problems (AMC, Olympiad style)
  • Answering coding / algorithm questions
  • Logical puzzles and formal reasoning
  • Short scientific simulations (e.g., three‑body orbit, optimisation)

It is not intended for general conversation, creative writing, or factual retrieval outside its training distribution.


Evaluation Results

The fine‑tuned model was tested on a comprehensive benchmark of 14 tasks (code execution, auto‑graded).
Passed 12 out of 14 tasks – see full log below.

Task Verdict Notes
AMC 2025 – n‑Norwegian Number PASS 50% keywords (deep number theory)
Euler Totient Sum (last 6 digits) FAIL Correct algorithm, output mismatch (format)
Lattice Paths Avoiding Anti‑Diagonal PASS
Segmented Sieve [10¹², 10¹²+10⁶] FAIL Prime count off by ~500
Median of Two Sorted Arrays (O(log n)) PASS
Thread‑Safe LRU Cache with TTL PASS
Persistent Segment Tree – K‑th Smallest PASS
Multi‑Head Attention + RoPE (NumPy only) PASS
Dijkstra vs A* on Large Random Graph PASS
Knights & Knaves – Exhaustive Solver PASS
Verify Three Mathematical Claims PASS
Figure‑8 Three‑Body Orbit – Energy Conservation PASS
Metropolis‑Hastings vs HMC Comparison PASS
Optimizer Comparison on Rosenbrock PASS

The model shows strong reasoning, algorithmic implementation, and numerical stability.
The two failures are due to output formatting mismatches or edge‑case off‑by‑one issues, not a lack of understanding.


HydraNode Design Benchmark: GPT-OSS-120B-MLX-q4-Claude-4.6-Opus-Reasoning-Distilled at Temperature 0.9

Summary

We evaluated multiple responses generated by the same model (cloudyu/GPT-OSS-120B-MLX-q4-Claude-4.6-Opus-Reasoning-Distilled) on a complex distributed‑systems design task: HydraNode – a P2P real‑time risk hedging system with strict latency (<10 ms), custom Gossipsub layer, Vector Clock + MMR consistency, and parallel Delta/Gamma computation in Rust.

Among all tested variants, the response generated at temperature 0.9 was judged optimal across all key dimensions:

Dimension 0.9 Temperature Performance
System architecture & topology Clear Mermaid diagram with DHT, control plane, and exchange adapters
Consistency mechanism Full implementation of Vector Clock + Merkle Mountain Range + Bloom filter with probabilistic error bound
Rust code quality Concurrency‑safe risk matrix using AtomicCell + Rayon parallel updates; Fixed‑Point arithmetic
Latency formula Detailed breakdown (L_net, L_bloom, L_mmr, L_sync, L_compute) with numerical validation (<10 ms)
Network partition handling AP vs CP trade‑off analysed via game theory (payoff matrix, Nash equilibrium)
Attack survival model Quantitative analysis of 51% and Eclipse attacks with survival probability ≈0.995
Readability & structure Clear separation of <think> reasoning chain and final answer (Harmony template)

Why Temperature 0.9?

The model is a LoRA‑finetuned reasoning distil of GPT‑OSS‑120B, specialised for mathematical, coding, and algorithmic problems. At 0.9 temperature:

  • It maintains logical coherence while introducing healthy diversity – avoids repetitive or overly conservative outputs.
  • It generates multiple reasoning branches internally (thanks to the Harmony <analysis> channel) and merges the best path into the final answer.
  • It produces detailed, non‑trivial components (e.g., game‑theoretic payoff matrices, MMR + Bloom hybrid verification) that lower temperatures (0.0–0.5) tend to oversimplify.

In contrast, temperatures below 0.7 often yield “textbook” solutions lacking depth, while temperatures above 1.0 introduce excessive randomness and reduce code correctness.

Recommendation

Use temperature = 0.9 for this model on:

  • Complex system design (distributed protocols, risk engines, consensus mechanisms)
  • Mathematical proofs and algorithmic reasoning
  • Multi‑step code generation with concurrency and fixed‑point arithmetic

Always follow the Harmony chat template:

How to Use

1. Install dependencies

pip install mlx-tune datasets huggingface_hub
# or (old package name)
pip install unsloth-mlx

2. Load the LoRA adapter

from mlx_tune import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    "cloudyu/GPT-OSS-120B-MLX-q4-Claude-4.6-Opus-Reasoning-Distilled",
    max_seq_length=4096,
    load_in_4bit=True,
)

command line

mlx_lm.chat --model cloudyu/GPT-OSS-120B-MLX-q4-Claude-4.6-Opus-Reasoning-Distilled --max-tokens 10000 --temp 0.9

OpenAI Local Service

mlx-openai-server launch --model-path cloudyu/GPT-OSS-120B-MLX-q4-Claude-4.6-Opus-Reasoning-Distilled --model-type lm --host 127.0.0.1 --temperature 0.9

3. Format prompts with Harmony

def harmony_msg(role, channel, content):
    return f"<|start|>{role}<|channel|>{channel}<|message|>{content}<|end|>"

prompt = (
    harmony_msg("system", "system", "You are a helpful assistant.")
    + harmony_msg("user", "query", "What is the 3rd smallest prime?")
    + harmony_msg("assistant", "analysis", "")   # let model generate analysis
)

inputs = tokenizer(prompt, return_tensors="np")
outputs = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0]))

Important – Always use the exact Harmony tokens shown above.
Do not use <|im_start|> / <|im_end|> (ChatML) – the model will not understand them.


Training Details

Full training script can be found in train_0404.py.
Key hyperparameters:

Parameter Value
LoRA rank (r) 16
LoRA alpha 32
Target modules q_proj,k_proj,v_proj,o_proj,gate_proj,up_proj,down_proj
Router fine‑tuning full‑precision, lr=1e-5 (AdamW)
Muon LR (LoRA) 1e-5
AdamW LR (1‑D / bias) 1e-5
Batch size 4 (×8 gradient accumulation)
Warmup ratio 0.05
Scheduler cosine
Max steps 200 (≈ 1 epoch)

Limitations & Biases

  • Quantisation : The base model is loaded in 4‑bit MLX format – some precision loss is inherent.
  • Domain : Primarily trained on reasoning problems; performance on general knowledge, creative writing, or non‑English text may be poor.
  • Output length : The model may sometimes generate very long analysis chains. Use max_new_tokens to bound generation.
  • Harmony token sensitivity : If the tokenizer does not recognise the special Harmony tokens as single tokens, the model will not work correctly. The provided adapter expects the exact vocabulary of openai/gpt-oss-120b.

Citation

If you use this model in your work, please cite:

@misc{gpt-oss-120b-lora-reasoning-distilled,
  author       = {Yuhai},
  title        = {GPT-OSS-120B LoRA Fine‑tune for Reasoning (Claude-4.6-Opus Distilled)},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/cloudyu/GPT-OSS-120B-MLX-q4-Claude-4.6-Opus-Reasoning-Distilled}}
}

Contact

For questions or issues, please open an issue on the Hugging Face repository or contact the author via the training script repository.

Downloads last month
230
Safetensors
Model size
117B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support