super interesting new paper from Microsoft "Agent Lightning v1.0: Towards Harnessed Agentic RL" by Zhiyuan He et al.
same idea we've seen already several times: you train the agent inside the real harness it ships with, instead of a reimplementation of it
now that recipe has a name → harnessed agentic RL
paper: huggingface.co/papers/2608.17528
the tricky bit they nail down: one rollout is not one training sample
the harness calls the model many times, so a single episode → a variable number of (prompt, response) rows
you don't even know the batch size until the episode finishes running
its real contribution is being first to systematically map the four problems that fall out of that:
> retokenization + sample merging > advantage calculation over a variable sample count > loss normalization at the rollout level, not per sample > backend scheduling when the batch size is dynamic
and it actually works → plain RL inside the real harness, no reimplementation
Qwen3.5-9B on SWE-bench Verified 41.8 → 56.4 (+14.6), with only ~6k examples
the whole thing is ~3,500 lines, any harness, self-hosted k8s
from our side, we've shared some materials on the same line you may want to check out :)
Introducing Unsloth Desktop 🦥 The first desktop app to run and train models locally.
• Open-source. Runs on Mac, Windows and Linux • Supports MLX, diffusion image/video, audio, GGUF • Connect Claude Code and Codex to local LLMs • 50% more accurate, self-healing tool calls + sandboxed code exec • Works for CPU + multiGPU setups - NVIDIA, AMD, Intel, Mac • Train models 2× faster with 70% less VRAM • Private web search, deep research, RAG, MCP and exports (NVFP4, GGUF) • Use Unsloth’s OpenAI-compatible API and cloud models • Securely deploy LLMs remotely and access anywhere
Something I really like when I study a subject is understanding its history, how it reached the point where it is today
I did that exercise for RL in post-training: from RLHF and PPO, to verifiable rewards, to the GRPO family of variants, to agents acting in environments. Everything is backed by what the labs themselves say in their public reports (DeepSeek, Qwen, Kimi, GLM-5, Nemotron, Mistral and more), in their own words
This is the companion piece to Class 3 of our Training Agents series with @burtenshaw. The class explains how GRPO works, with three hands-on experiments. The article shows where the same ideas appear at frontier scale
we just released a new blog "Training a coding agent using the OpenCode harness in remote HF sandboxes with TRL and OpenEnv"
you can take a real coding agent (OpenCode), let it run its own tool loop against real coding problems, and train it with RL on the exact tokens it produced
and every rollout runs in its own remote HF sandbox, so rollouts scale out beyond one machine
the loop: - OpenCode owns its tool loop inside an OpenEnv sandbox - an in-sandbox proxy records the real token ids + logprobs, per turn - a hidden-test verifier scores the result, and that is the reward - TRL trains with AsyncGRPO, weights sync back to vLLM over NCCL
Unlike agents that depend on cloud APIs, local agents give you free inference, low latency, and real privacy.
Removing the per-token cost changes how developers build: agents can now be massively parallelized on local hardware, running background tasks that burn through millions of tokens at no marginal cost!
and the @liquidai blog comes with some nice details about the training procedure, so let's analyze it.
basically, a full agent training pipeline but compressed into 2.6B
base model → SFT → specialized teachers per domain (SFT + RLVR) → on-policy distillation back into one student → agentic RL
the two most interesting stages
→ MOPD: the student generates, each prompt routes to its domain teacher for token-level feedback. teachers branch from the same SFT checkpoint, so their signal stays close to the student's distribution
→ agentic RL: multi-turn GRPO inside real harnesses (OpenClaw, Hermes Agent), one sandbox per rollout, a proxy captures token-level trajectories while the harness stays a black box
this makes a 2.6B that beats much larger models on instruction following and tool use
SFT, distillation, RL, RL envs: exactly what we're covering in our Training Agents livestream series (next one coming soon!)
tomorrow (Tuesday, July 28), we're back with Class 3 of the Training Agents live series
🧠 what: reinforcement learning for training agents (GRPO): how it works, how to implement it in TRL, and end-to-end examples 🗓️ when: Tuesday, July 28 - 🕔 5:00 PM CEST / 8:30 PM IST 📍 where: Live on @huggingface's X, YouTube, and LinkedIn
you can now train your own coding agents with trl + openenv, starting with opencode
we just added end-to-end support for training agent harnesses:
> TRL: a loop-owning training path (AsyncGRPOTrainer + HarnessRolloutWorker) that launches the agent in an OpenEnv session, reads back its trace, reconstructs the training samples, and trains with AsyncGRPO > OpenEnv: the OpenCode harness environment plus a transparent proxy that forwards the agent's model calls and records each turn's token ids and logprobs
you train the actual opencode agent as is, it runs its own loop and tools and the policy learns from the exact tokens it produced
we're shipping a self-contained example: local subprocess sandbox, DeepCoder problems, validated on Qwen3-8B.
join us next Tuesday, July 28, for Class 3 of the Training Agents live series!
we'll dive into reinforcement learning for agent training, covering the intuition behind GRPO, how it works, and how to implement it in TRL with practical, e2e examples