Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Paper • 2607.20911 • Published 5 days ago • 24
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Paper • 2607.13705 • Published 13 days ago • 44
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 12 days ago • 103
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published 14 days ago • 226
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation Paper • 2607.05382 • Published 19 days ago • 87
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published 17 days ago • 70
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published 15 days ago • 84
AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification Paper • 2607.11849 • Published 15 days ago • 33
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published 20 days ago • 138
WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence Paper • 2607.06838 • Published 21 days ago • 14
PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception Paper • 2606.28322 • Published Jun 26 • 43
Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning Paper • 2607.01191 • Published 27 days ago • 19
Dockerless: Environment-Free Program Verifier for Coding Agents Paper • 2606.28436 • Published Jun 26 • 113
GEAR: Guided End-to-End AutoRegression for Image Synthesis Paper • 2606.32039 • Published 28 days ago • 34
Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation Paper • 2606.26907 • Published Jun 25 • 49
Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence Paper • 2606.15932 • Published Jun 16 • 38