JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Paper • 2607.23588 • Published 3 days ago • 112
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Paper • 2607.24720 • Published 2 days ago • 22
StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents Paper • 2607.22798 • Published 5 days ago • 55
From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search Paper • 2607.24280 • Published 2 days ago • 73
Multi-Head Latent Control: A Unified Interface for LLM Agent Decision Making Paper • 2607.14277 • Published 14 days ago • 8
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems Paper • 2607.21503 • Published 6 days ago • 22
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Paper • 2607.21653 • Published 7 days ago • 29
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills Paper • 2607.22529 • Published 5 days ago • 38
OpenForgeRL: Train Harness-native Agents in Any Environment Paper • 2607.21557 • Published 6 days ago • 7
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Paper • 2607.20911 • Published 6 days ago • 25
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents Paper • 2607.20709 • Published 7 days ago • 31
AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper • 2607.21461 • Published 6 days ago • 147
Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing Paper • 2607.07953 • Published 21 days ago • 16
Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents Paper • 2607.08716 • Published 20 days ago • 15
Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code Paper • 2607.13921 • Published 14 days ago • 11
Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness Paper • 2607.19322 • Published 8 days ago • 9