Instructions to use PleIAs/baguettotron-600m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PleIAs/baguettotron-600m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="PleIAs/baguettotron-600m") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("PleIAs/baguettotron-600m") model = AutoModelForCausalLM.from_pretrained("PleIAs/baguettotron-600m", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use PleIAs/baguettotron-600m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "PleIAs/baguettotron-600m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PleIAs/baguettotron-600m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/PleIAs/baguettotron-600m
- SGLang
How to use PleIAs/baguettotron-600m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "PleIAs/baguettotron-600m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PleIAs/baguettotron-600m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "PleIAs/baguettotron-600m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PleIAs/baguettotron-600m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use PleIAs/baguettotron-600m with Docker Model Runner:
docker model run hf.co/PleIAs/baguettotron-600m
🥖 Baguettotron-600M
Paper (NeurIPS 2026) · SYNTH dataset · Baguettotron-350M · Baguettotron-MoE
Baguettotron-600M is a 0.6B-parameter reasoning model trained from scratch on 158B tokens of SYNTH, a fully synthetic corpus amplified from 58,698 Wikipedia articles. It is the main reference model of our NeurIPS 2026 paper, It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs.
- Single training stage: no separate mid-training, SFT or RLHF. The model follows instructions, reasons in
<think>traces and recalls facts because SYNTH trains those capabilities directly. - Token-efficient: within 4.7 points of Qwen3-0.6B on 22 multiple-choice tasks and 0.7 points on 8 open-ended tasks, trained on ~230× fewer tokens.
- Factually precise: highest FActScore in its parameter tier (41.7% macro vs. 31.6% for Qwen3-0.6B).
Model design
Baguettotron-600M is a standard dense decoder (Llama/Qwen-style) and loads natively in transformers and vLLM (LlamaForCausalLM, no remote code).
| Parameters | 608M (incl. 67M tied embeddings) |
| Layers | 48 |
| Hidden size | 1024 |
| Attention | GQA, 16 query / 4 KV heads, head dim 64 |
| MLP | SwiGLU, intermediate size 2816 |
| Position encoding | RoPE, θ = 10,000 |
| Norm | RMSNorm (pre-norm), ε = 1e-6 |
| Embeddings | Tied input/output |
| Vocabulary | 65,536 (Pleias tokenizer) |
| Context length | 2,048 tokens |
Training
- Data: SYNTH, 158.3B tokens (~2 passes over the ~75B-token corpus)
- Steps: 151,000 at global batch 512 × 2,048 tokens (~1.05M tokens/step)
- Optimizer: AdamW, peak LR 1.5e-3, 10k-step warmup, linear decay over the final 16.6% of steps to 0.2% of peak; weight decay 0.01, grad clip 1.0
- Hardware: 16× H100 (4 nodes × 4 GPUs) on MareNostrum 5 (BSC),
torchtitanwith FSDP, ~52.6 h, ~30% MFU
Evaluation
All scores are from the paper. Each model is prompted in its native format; Baguettotron models get no system prompt and have <think>\n seeded after the assistant turn.
| Model | Training tokens | MCQ (22 tasks) | Open-ended (8 tasks) | FActScore (macro) |
|---|---|---|---|---|
| Baguettotron-600M | 158B | 42.2 | 24.3 | 41.7 |
| Baguettotron-MoE (13B / 1B active) | 50B | 39.6 | 23.8 | 46.3 |
| Baguettotron-350M | 200B | 36.3 | 16.0 | 32.4 |
| Qwen3-0.6B | ~36T | 46.9 | 25.0 | 31.6 |
| LFM2.5-350M | 28T | 42.3 | 20.8 | 22.0 |
| SmolLM2-360M-Instruct | 4T | 22.8 | 21.6 | 26.6 |
| Gemma-3-270M-IT | 2T | 22.5 | 14.9 | 19.4 |
- Ties or beats Qwen3-0.6B on TruthfulQA (42.6 vs. 32.7), ESGenius (62.2 vs. 55.1) and FormationEval (62.8 vs. 58.0)
- Weak: benchmarks that reward broad web knowledge (ARC-Challenge, GeoBench), since SYNTH covers only its seed articles
- Data, not architecture: the same architecture, tokenizer and step count trained on FineWiki or FinePDFs-Edu reaches only 26.1 / 25.1 on MCQ, even after SmolTalk post-training
- Reasoning traces matter: retraining without traces costs 9.1 points on TruthfulQA and 10.0 on NuclearQA
Prompt format
The model was trained on ChatML with a <think> block and no system prompt. The bundled chat template opens the assistant turn with <think>\n:
<|im_start|>user
What do you know about the Treaty of Westphalia?<|im_end|>
<|im_start|>assistant
<think>
The model writes its reasoning, closes it with </think>, answers, and ends the turn with <|im_end|>.
RAG. Pass sources inside the user turn. The answer then cites them with <ref>[quote]</ref>:
<|im_start|>user
{question}
<source_1>[…]</source_1>
<source_2>[…]</source_2><|im_end|>
<|im_start|>assistant
<think>
Reasoning notation. Traces use SYNTH's compact stenographic style (→ derivation, ↺ backtracking, ∴ conclusion, …). The confidence markers ● (high), ◐ (partial) and ○ (low) are informative: on traces dominated by uncertain markers, FActScore drops from 0.44 to 0.40, and the model commits to ~15–20% fewer facts. See the Baguettotron-350M card for the full notation.
Inference
vLLM
vllm serve PleIAs/baguettotron-600m --reasoning-parser deepseek_r1
--reasoning-parser deepseek_r1 moves the <think> trace into a separate reasoning field. Tested with vLLM 0.24.0.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
model="PleIAs/baguettotron-600m",
messages=[{"role": "user", "content": "What do you know about the Treaty of Westphalia?"}],
temperature=0.1,
top_p=0.95,
presence_penalty=0.1,
max_tokens=1536,
)
print(resp.choices[0].message.reasoning) # <think> trace
print(resp.choices[0].message.content) # answer
Recommended sampling: temperature=0.1, top_p=0.95, presence_penalty=0.1. Prompt and generation share the 2,048-token context, so keep max_tokens below 2,048 minus the prompt length.
Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "PleIAs/baguettotron-600m"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype="bfloat16", device_map="auto")
messages = [{"role": "user", "content": "What do you know about the Treaty of Westphalia?"}]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt"
).to(model.device)
out = model.generate(inputs, max_new_tokens=1536, do_sample=True, temperature=0.1, top_p=0.95)
print(tokenizer.decode(out[0, inputs.shape[1]:]))
Limitations
- 2,048-token context shared by prompt, reasoning and answer
- Knowledge is bounded by the seed corpus (~58.7k Wikipedia articles): recall is precise on seed topics and weaker outside them
- Reasoning traces are English-only, even when the prompt and answer are in another language
- No preference tuning or safety alignment; not intended for high-stakes use without further evaluation
Citation
@inproceedings{langlais2026itsalltraining,
title = {It's All Training: A Fully Synthetic Single-Stage Recipe for {LLMs}},
author = {Langlais, Pierre-Carl and Delobelle, Pieter and Detrois, Yannick and Chizhov, Pavel and Rosas-Hinostroza, Carlos and Si Smail, Neil and Burtin, Benjamin and Shcharbakova, Hanna and Yamshchikov, Ivan and Stasenko, Anastasia},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026},
eprint = {2609.37891},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2609.37891}
}
Acknowledgements
Trained on MareNostrum 5 (BSC) through the EuroHPC Extreme Scale Access call (EHPC-EXT-2025E01-092, JULIP: Powerful LLMs made in Europe). Parts of this research received funding from SPRIN-D, the German Federal Agency for Breakthrough Innovation.
- Downloads last month
- 144