Decoding Looped Transformers Better for (Almost) Free Paper • 2610.02185 • Published 10 days ago • 46
Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes Paper • 2610.02117 • Published 10 days ago • 25
HiRAE: Hierarchical Representation Autoencoding with Residual Budgets Paper • 2609.37775 • Published 12 days ago • 26
EVO-WAM: Evolving World Action Models through Video-Action Verification Paper • 2609.38057 • Published 12 days ago • 49
What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling Paper • 2609.34981 • Published 12 days ago • 138
The Past Frames the Future: Memory for Autoregressive Video Generation Paper • 2609.28466 • Published 18 days ago • 65
GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation Paper • 2609.24981 • Published 20 days ago • 76
Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms Paper • 2609.23658 • Published 21 days ago • 30
Uncertainty-Aware World Model for Aerial Image-Goal Navigation Paper • 2608.05597 • Published Aug 6 • 12
SimWAM: A Simple World Action Model for End-to-End Autonomous Driving Paper • 2608.07468 • Published Aug 7 • 41
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report Paper • 2607.18367 • Published Jul 20 • 63
Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model Paper • 2607.03509 • Published Jul 3 • 15
Multiplayer Interactive World Models with Representation Autoencoders Paper • 2607.05352 • Published Jul 6 • 27
WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory Paper • 2607.02517 • Published Jul 2 • 31