FilmBench: A Film-Grade Benchmark for Cinematic Video Generation Paper • 2607.24241 • Published 4 days ago • 6
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation Paper • 2607.24720 • Published 4 days ago • 25
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification Paper • 2607.24027 • Published 4 days ago • 34
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Paper • 2607.23855 • Published 5 days ago • 26
GraphVid: Interactive Graph-Controllable Video Generation Paper • 2607.21580 • Published 8 days ago • 8
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Paper • 2607.21553 • Published 8 days ago • 39
Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers Paper • 2607.21594 • Published 8 days ago • 16
AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper • 2607.21461 • Published 8 days ago • 149
ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion Paper • 2607.20417 • Published 9 days ago • 9
SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction Paper • 2607.15550 • Published 14 days ago • 27
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning Paper • 2607.14183 • Published 16 days ago • 68
Hierarchical Denoising For Multi-Step Visual Reasoning Paper • 2607.15278 • Published 15 days ago • 6
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published 9 days ago • 304
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report Paper • 2607.18367 • Published 10 days ago • 57
Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Paper • 2607.16401 • Published 14 days ago • 43
FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications Paper • 2607.18171 • Published 11 days ago • 6
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published 11 days ago • 59