LTX-2.5 Collection LTX-2.5 base models, quantized models and accompanying LoRAs and IC-LoRAs • 4 items • Updated 13 days ago • 46
KVAE: Family of Tokenizers for Multimodal Generative Models Paper • 2608.05798 • Published 18 days ago • 29
Kandinsky WM 1.0 Collection Image-to-Video models for Physical AI: autonomous driving, robotics, general physics. • 3 items • Updated 20 days ago • 4
Laguna S 2.1 Collection Our most capable model to date, designed for long-horizon work. • 13 items • Updated 21 days ago • 45
PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation Paper • 2607.02515 • Published Jul 2 • 20
Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders Paper • 2606.10029 • Published Jun 8 • 12
A Geometric Account of Activation Steering through Angle-Norm Decomposition Paper • 2606.06735 • Published Jun 4 • 27
Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders Paper • 2606.07473 • Published Jun 5 • 15
view article Article Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action nvidia • Jun 1 • 88
KVAE 2.0 Collection KVAE 2.0 is a family of image and video tokenizers with a time compression ratio of 4 and spacial compression ratio of 8 and 16 • 3 items • Updated 18 days ago • 5
Interpreting CLIP with Hierarchical Sparse Autoencoders Paper • 2502.20578 • Published Feb 27, 2025 • 1
SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models Paper • 2511.08379 • Published Nov 11, 2025 • 5
AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders Paper • 2602.05027 • Published Feb 4 • 63
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Paper • 2601.05242 • Published Jan 8 • 234
Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models Paper • 2506.09229 • Published Jun 10, 2025 • 7