The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published 27 days ago • 171
MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism Paper • 2606.07512 • Published Jun 5 • 40
CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Paper • 2605.24786 • Published May 24 • 9
gradients-io-tournaments/tournament-tourn_33c30659e3c920d5_20260601-28242151-75bf-473a-9acf-d0b4b0ed337d-5FpdSckw Text Generation • Updated Jun 2 • 1
Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion Paper • 2605.31170 • Published May 29 • 12
RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion Models Paper • 2605.26632 • Published May 26 • 16