SPW-Nav: A Streaming Panoramic World Model for Language-Guided Navigation
Abstract
Language-guided panoramic video generation benefits various downstream applications, such as interactive 3D scene exploration, virtual reality experiences, and embodied agent training. Existing panoramic generators follow predefined trajectories, and interactive world models act through low-level actions in perspective views. We propose SPW-Nav, a streaming panoramic world model that understands movement instructions and streams one minute of 2K 360-degree video in real time from a single panorama. SPW-Nav interprets each instruction in the previously generated panorama as camera motion. Spherical rotation decoupling applies rotation exactly on the sphere, pose-aligned conditioning keeps translation inputs bounded over long streams, and a multi-term memory with a few-step generator continues the scene as instructions change. We also build SPW-NavSet, panoramic videos with camera trajectories and verified instructions. Driven by language, SPW-Nav outperforms prior panoramic generators in camera-following accuracy and video quality, and supports on-the-fly instruction switching.
Community
TL;DR: SPW-Nav is a real-time, language-guided panoramic world model that generates minute-long 2K 360° videos from a single panorama, enabling interactive navigation through spherical rotation decoupling, pose-aligned conditioning, and streaming memory.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video (2026)
- WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory (2026)
- InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter (2026)
- MUGEN: Interactive Panoramic World Exploration via Camera Control (2026)
- Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control (2026)
- 4DStreamCtrl: Interactive Video Generation with Online 4D Control (2026)
- EchoWM: Open and Enterable Omnimodal World Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper