StreamRender-H3 runtime assets
The complete inference asset bundle for the code-driven StreamRender-H3 runtime: controllable game engine โ semantic reference frames โ streaming TAE encode โ two full H3 render evaluations โ causal ViT24 decode โ browser feedback.
Runtime source and deployment instructions:
NVlabs/Sana, sol-engine/models/streamrender_h3,
validated integration commit fcef8d4287d5f45e1a5e839c940b7e14c91c71fc.
Included components
| Directory | Component |
|---|---|
backbone/ |
Latest 1008b step224 merged EMA H3 backbone, all 50 layers, standard safetensors shards |
tae/ |
Official TAE weights used by the online reference encoder |
causal_decoder/ |
Fast causal ViT24 decoder step6000, release 20261005-v1 |
qwen_model/ |
Qwen3-VL text/visual conditioning model, original safetensors shards and configuration |
qwen_dcp/ |
Corresponding distributed checkpoint used by the runtime loader, including .metadata |
tokenizer/, qwen_processor/ |
Exact tokenizer/processor assets |
native_vae/ |
Native H3 video VAE checkpoint used to construct the decoder architecture |
audio_vae/ |
H3 audio codec checkpoint required by the model contract |
reference_image/ |
Reference image for the validated opening racing scene |
configs/ |
Runtime inference profile and tested environment versions |
This is the current deployment bundle, not a collection of historical training
checkpoints. Optimizer states are excluded. Both Qwen formats are included to
preserve the tested loader setup. Total payload is approximately 217 GB.
The original backbone contains 535 tensors, 66,280,488,840 bytes, SHA256:
56c3ee3270db40354ef70fcbea16248809914cfa8f51acb57dae7d30fa4bb597.
Sharding preserves every tensor and permits byte-identical reconstruction.
Download and run
Keep model caches and downloads on a sufficiently large storage volume.
from huggingface_hub import snapshot_download
snapshot_download("yitongl/StreamRender-H3", local_dir="/your/code/models/StreamRender-H3")
The current runtime expects one merged backbone file. Reconstruct it first (requires an additional 66.28 GB; the script checks the original SHA256):
python /your/code/models/StreamRender-H3/download_assets.py
cd Sana/models/streamrender_h3
python scripts/preflight.py --assets /your/code/models/StreamRender-H3/assets.json
GPUS=8 PYTHON_BIN=python bash scripts/launch_local.sh \
--config /your/code/models/StreamRender-H3/configs/runtime.json \
--assets /your/code/models/StreamRender-H3/assets.json
Use the runtime's documented Python/CUDA/FA4 environment and install its
frontend dependencies. The tested deployment uses eight GB200 GPUs within
one NVLink domain, two denoising evaluations, CFG=1, S=2/W=2/C=2, 832ร480
reference encoding and 1344ร768 output. Qwen consumes prompt and first image
once per session. The included assets.json uses relative paths.
The actual browser-to-H3-to-browser integration was verified in a same-rack eight-GPU run that exited successfully. The longer validation measured a 295.58 ms median GPU-resident RGB-to-RGB-ready interval and approximately 1.03 s browser publication intervals. These are different timing boundaries; the browser test is not a claim of 24-fps end-to-end feedback.
License and attribution
This bundle is not blanket Apache/MIT licensed. H3 weights and derivatives
are governed by the included MiniMax H3 Community License, including its
territorial and redistribution terms. Qwen components retain Apache-2.0
terms. TAE source attribution is preserved separately. See LICENSE,
LICENSE-Qwen.txt, LICENSE-TAE.txt, NOTICE and SOURCE_SNAPSHOT.json.
Consult your applicable authorization when using or redistributing the bundle.