Mage Collection A family of lightweight multimodal models, including understanding and generation. • 8 items • Updated 27 days ago • 28
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 26 days ago • 37
nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Viewer • Updated Jun 9 • 156M • 28.7k • 37
view article Article Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action nvidia • Jun 1 • 88
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Paper • 2506.05414 • Published Jun 4, 2025 • 4