DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation Paper • 2608.13489 • Published Aug 13 • 101
InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis Paper • 2608.02437 • Published Aug 3 • 75
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published Jul 29 • 69
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 83