Span-based Localizing Network for Natural Language Video Localization Paper • 2004.13931 • Published Jun 14, 2020
Video-KTR: Reinforcing Video Reasoning via Key Token Attribution Paper • 2601.19686 • Published Jan 27
Temporal Sentence Grounding in Videos: A Survey and Future Directions Paper • 2201.08071 • Published Jan 20, 2022
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization Paper • 2607.14614 • Published Jul 16 • 12
Scaling Language-Centric Omnimodal Representation Learning Paper • 2510.11693 • Published Oct 13, 2025 • 109
MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources Paper • 2509.21268 • Published Sep 25, 2025 • 104
MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources Paper • 2509.21268 • Published Sep 25, 2025 • 104
MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources Paper • 2509.21268 • Published Sep 25, 2025 • 104
GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning Paper • 2509.17437 • Published Sep 22, 2025 • 17
RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation Paper • 2509.15212 • Published Sep 18, 2025 • 22