VLAct Collection Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models • 10 items • Updated about 12 hours ago
VLAct Collection Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models • 10 items • Updated about 12 hours ago
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published Jul 21 • 77
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published Jul 27 • 37
XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment Paper • 2608.09892 • Published 17 days ago
StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding Paper • 2608.16320 • Published 11 days ago • 9
Mage Collection A family of lightweight multimodal models, including understanding and generation. • 8 items • Updated Jul 26 • 29
StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Paper • 2608.05703 • Published 22 days ago • 16
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published Jul 27 • 37
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published Jul 27 • 37
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems Paper • 2604.11757 • Published Apr 13
SciForma: Structure-Faithful Generation of Scientific Diagrams Paper • 2607.18091 • Published Jul 20 • 24
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published Jul 21 • 77