Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL
Abstract
Reinforcement learning (RL) for code agents often uses executable tests to provide binary rewards. With these rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to test-passing trajectories within each rollout group, overlooking differences in implementation quality and adherence to task requirements. This leaves the policy without a learning signal that favors clean, targeted implementations over those containing unnecessary or out-of-scope changes. We introduce GAGAR, a framework for quality-aware credit redistribution in code agent RL. Built on dynamic sampling that retains groups containing both passing and failing trajectories, GAGAR places all trajectories from each group in a shared workspace, where an SFT-trained agentic grader jointly inspects them and ranks the test-passing candidates. Based on this ranking, we downweight lower-ranked trajectories and proportionally rescale the advantages of all test-passing trajectories to restore their original sum. This sum-preserving redistribution retains the relative weights established by quality-based downweighting while shifting credit toward higher-quality implementations. We evaluate GAGAR at industrial scale using pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters). Controlled code-only Flash experiments show improved code agent performance, reduced trajectory-length growth, and more stable training. We further apply GAGAR in large-scale mixed-task RL with both Flash and Pro. Our results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.
Community
At Xiaomi, we developed GAGAR, a new approach for code agent reinforcement learning (RL) that introduces a groupwise agentic grader to perform online advantage redistribution. Instead of only optimizing for test passing, GAGAR enables code agents to learn how to produce higher-quality, more maintainable implementations. This technology has been integrated into the large-scale RL training pipeline of the recently released MiMo-V2.6 series.
The motivation behind this work came from feedback from a colleague who tested our first-generation model: "The task is solved, but the model leaves 💩 everywhere in the code."
This comment captured a fundamental limitation of test-based evaluation: passing tests does not necessarily mean producing a high-quality implementation. Among successful solutions, some are precise and minimal, while others introduce unnecessary complexity. Some directly address the core issue, while others take long and convoluted paths. A binary test reward is insufficient to capture these differences.
GAGAR addresses this challenge by allowing a grader to jointly evaluate multiple trajectories solving the same task within a group, and then redistribute credit from lower-quality successful solutions toward higher-quality ones. Through this process, the model learns not only how to pass tests, but also how to choose better solution paths and write cleaner, more precise, merge-ready code.
Making this work required three key designs:
- Groupwise: Differences become visible through comparison
Quality differences are often difficult to identify when examining a single implementation in isolation. Many unnecessary complexities may appear reasonable on their own. However, when multiple solutions are presented together, the grader can better identify which approaches are unnecessarily complicated and which solutions are more elegant and effective. - Agentic: Evaluation requires active investigation
Evaluating code trajectories also requires an agentic process. For tasks involving hundreds of interaction turns, the grader needs to actively inspect relevant information, including intermediate trajectories, patches, repositories, and execution evidence, rather than making judgments based only on a single input or a compressed summary. - Zero-sum: Redistribute credit instead of simply penalizing
The credit removed from lower-quality solutions should be transferred to higher-quality ones. Rather than simply reducing the advantage of lower-quality positive examples, GAGAR preserves the overall positive reward signal while adjusting relative preferences among successful trajectories. This avoids imbalance between positive and negative advantages during RL training.
We integrated GAGAR into RL training for industrial-scale models with 310B and 1.02T parameters, covering all code training tasks. We observed improvements in pass rates and implementation quality, along with reductions in interaction turns and output token consumption, as well as more stable training dynamics.
More broadly, GAGAR provides a new way to scale RL training: instead of only scaling model size or rollout volume, we can allocate additional computation to graders to better understand which behaviors are actually worth learning.
With GAGAR, the MiMo-V2.6 series achieves code capabilities competitive with frontier models.
For code agents operating in environments with hundreds of interaction turns, where repositories and execution states continuously branch, agentic graders are particularly suitable compared with credit assignment methods that rely on shared cross-trajectory states. We are also continuing to explore turn-level credit assignment and investigating how zero-sum redistribution can be applied to different graders, penalty mechanisms, and advantage adjustment strategies.
Finally, we would like to share the release of the MiMo-V2.6 series:
- Blog: https://mimo.xiaomi.com/mimo-v2-6
- Technical Report: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf
- Hugging Face Weights: https://huggingface.co/collections/XiaomiMiMo/mimo-v26
Get this paper in your agent:
hf papers read 2609.32577 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper