Flux-OPD: On-Policy Distillation with Evolving Contexts Paper • 2607.28022 • Published 18 days ago • 44
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Paper • 2607.20465 • Published May 19 • 56
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation Paper • 2607.14202 • Published Jul 15 • 42
CapRiCorn-1K: A Comprehensive Benchmark for Video Captioning and Subject Referential Consistency Across Temporal Scales Paper • 2606.21949 • Published Jun 20
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Paper • 2607.20465 • Published May 19 • 56
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Paper • 2607.28509 • Published 18 days ago • 30
Beacon: Knowing When and How to Perform Agentic Visual Reasoning Paper • 2607.28595 • Published 18 days ago • 55
Flux-OPD: On-Policy Distillation with Evolving Contexts Paper • 2607.28022 • Published 18 days ago • 44
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published about 1 month ago • 143
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation Paper • 2607.14202 • Published Jul 15 • 42
MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation Paper • 2607.14189 • Published Jul 15 • 34
OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data Paper • 2606.13432 • Published Jun 11 • 113
Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos Paper • 2605.18984 • Published May 18 • 22
Towards Next-Generation LLM Training: From the Data-Centric Perspective Paper • 2603.14712 • Published Mar 16
VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction Paper • 2605.15186 • Published May 14 • 26
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos Paper • 2605.07593 • Published May 8 • 1
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning Paper • 2605.22012 • Published May 21 • 46
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV Paper • 2605.26244 • Published May 25 • 38