VLX-Go: Vision-Language Short-Horizon Waypoint Prediction for Embodied Navigation
• 12
Multimodal AI, VLM, VLA, VAM, etc
TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models
Open Agent Leaderboard
Mark regions in images based on text descriptions
Process and answer questions about webpage videos
VLM-R1 model for Open-Vocabulary Object Detection