Supervise What Survives: Geometry-Guided VLA Adaptation from Synthetic Robot Videos Paper • 2606.24448 • Published Jun 23
Where Success Breaks: Failure-Boundary Learning for Robust Vision-Language-Action Models Paper • 2609.06114 • Published Sep 5
Where Success Breaks: Failure-Boundary Learning for Robust Vision-Language-Action Models Paper • 2609.06114 • Published Sep 5
Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance Paper • 2603.02175 • Published Mar 2 • 24
UniAPO: Unified Multimodal Automated Prompt Optimization Paper • 2508.17890 • Published Aug 25, 2025
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning Paper • 2607.02963 • Published Jul 3 • 28
Show-o2: Improved Native Unified Multimodal Models Paper • 2506.15564 • Published Jun 18, 2025 • 31
MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities Paper • 2504.02478 • Published Apr 3, 2025
OpenFACADES: An Open Framework for Architectural Caption and Attribute Data Enrichment via Street View Imagery Paper • 2504.02866 • Published Apr 1, 2025