Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling Paper • 2608.30821 • Published 12 days ago • 119
SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models Paper • 2602.04208 • Published Feb 4 • 21
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published Jul 29 • 140
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Paper • 2607.13429 • Published Jul 15 • 16
GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction Paper • 2605.23888 • Published May 22 • 18
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published Jul 13 • 87
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 90
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space Paper • 2607.05373 • Published Jul 6 • 67
Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling Paper • 2607.01642 • Published Jul 2 • 41
Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views Paper • 2606.29513 • Published Jun 28 • 54
TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction Paper • 2605.26115 • Published May 25 • 55
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments Paper • 2605.30280 • Published May 28 • 148
Geometry-Aware Representation Denoising for Robust Multi-view 3D Reconstruction Paper • 2605.26230 • Published May 25 • 41
SpatialBench: Is Your Spatial Foundation Model an All-Round Player? Paper • 2605.27367 • Published May 26 • 74