-
Cosmos 3: Omnimodal World Models for Physical AI
Paper • 2606.02800 • Published • 142 -
Robots Need More than VLA and World Models
Paper • 2606.06556 • Published • 31 -
VLANeXt: Recipes for Building Strong VLA Models
Paper • 2602.18532 • Published • 52 -
Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models
Paper • 2606.11025 • Published • 42
Collections
Discover the best community collections!
Collections including paper arxiv:2606.11025
-
Tencent-Hunyuan-Multimodal-RL/SD3.5-GenEval2-Single-Reward
Text-to-Image • Updated • 6 -
Tencent-Hunyuan-Multimodal-RL/SD3.5-GenEval2-Multi-Reward
Text-to-Image • Updated • 6 -
Tencent-Hunyuan-Multimodal-RL/FLUX2-klein-base-9b-GenEval2-Single-Reward
Text-to-Image • Updated • 21 • • 1 -
Tencent-Hunyuan-Multimodal-RL/FLUX2-klein-base-9b-GenEval2-Multi-Reward
Text-to-Image • Updated • 11 • • 3
-
E-GRPO: High Entropy Steps Drive Effective Reinforcement Learning for Flow Models
Paper • 2601.00423 • Published • 11 -
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Paper • 2601.05242 • Published • 235 -
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning
Paper • 2601.18150 • Published • 11 -
DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment
Paper • 2601.20218 • Published • 16
-
Multi-Agent Computer Use
Paper • 2606.01533 • Published • 6 -
OpenSkill: Open-World Self-Evolution for LLM Agents
Paper • 2606.06741 • Published • 29 -
Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills
Paper • 2606.07412 • Published • 13 -
Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnesses
Paper • 2606.08348 • Published • 16
-
lusxvr/nanoVLM-222M
Image-Text-to-Text • 0.2B • Updated • 472 • 103 -
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Paper • 2503.09516 • Published • 41 -
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time
Paper • 2505.24863 • Published • 98 -
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Paper • 2505.17667 • Published • 89
-
Cosmos 3: Omnimodal World Models for Physical AI
Paper • 2606.02800 • Published • 142 -
Robots Need More than VLA and World Models
Paper • 2606.06556 • Published • 31 -
VLANeXt: Recipes for Building Strong VLA Models
Paper • 2602.18532 • Published • 52 -
Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models
Paper • 2606.11025 • Published • 42
-
Tencent-Hunyuan-Multimodal-RL/SD3.5-GenEval2-Single-Reward
Text-to-Image • Updated • 6 -
Tencent-Hunyuan-Multimodal-RL/SD3.5-GenEval2-Multi-Reward
Text-to-Image • Updated • 6 -
Tencent-Hunyuan-Multimodal-RL/FLUX2-klein-base-9b-GenEval2-Single-Reward
Text-to-Image • Updated • 21 • • 1 -
Tencent-Hunyuan-Multimodal-RL/FLUX2-klein-base-9b-GenEval2-Multi-Reward
Text-to-Image • Updated • 11 • • 3
-
Multi-Agent Computer Use
Paper • 2606.01533 • Published • 6 -
OpenSkill: Open-World Self-Evolution for LLM Agents
Paper • 2606.06741 • Published • 29 -
Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills
Paper • 2606.07412 • Published • 13 -
Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnesses
Paper • 2606.08348 • Published • 16
-
E-GRPO: High Entropy Steps Drive Effective Reinforcement Learning for Flow Models
Paper • 2601.00423 • Published • 11 -
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Paper • 2601.05242 • Published • 235 -
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning
Paper • 2601.18150 • Published • 11 -
DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment
Paper • 2601.20218 • Published • 16
-
lusxvr/nanoVLM-222M
Image-Text-to-Text • 0.2B • Updated • 472 • 103 -
Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Paper • 2503.09516 • Published • 41 -
AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time
Paper • 2505.24863 • Published • 98 -
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
Paper • 2505.17667 • Published • 89