StreamRender-H3 runtime assets

The complete inference asset bundle for the code-driven StreamRender-H3 runtime: controllable game engine โ†’ semantic reference frames โ†’ streaming TAE encode โ†’ two full H3 render evaluations โ†’ causal ViT24 decode โ†’ browser feedback.

Runtime source and deployment instructions: NVlabs/Sana, sol-engine/models/streamrender_h3, validated integration commit fcef8d4287d5f45e1a5e839c940b7e14c91c71fc.

Included components

Directory Component
backbone/ Latest 1008b step224 merged EMA H3 backbone, all 50 layers, standard safetensors shards
tae/ Official TAE weights used by the online reference encoder
causal_decoder/ Fast causal ViT24 decoder step6000, release 20261005-v1
qwen_model/ Qwen3-VL text/visual conditioning model, original safetensors shards and configuration
qwen_dcp/ Corresponding distributed checkpoint used by the runtime loader, including .metadata
tokenizer/, qwen_processor/ Exact tokenizer/processor assets
native_vae/ Native H3 video VAE checkpoint used to construct the decoder architecture
audio_vae/ H3 audio codec checkpoint required by the model contract
reference_image/ Reference image for the validated opening racing scene
configs/ Runtime inference profile and tested environment versions

This is the current deployment bundle, not a collection of historical training checkpoints. Optimizer states are excluded. Both Qwen formats are included to preserve the tested loader setup. Total payload is approximately 217 GB. The original backbone contains 535 tensors, 66,280,488,840 bytes, SHA256: 56c3ee3270db40354ef70fcbea16248809914cfa8f51acb57dae7d30fa4bb597. Sharding preserves every tensor and permits byte-identical reconstruction.

Download and run

Keep model caches and downloads on a sufficiently large storage volume.

from huggingface_hub import snapshot_download
snapshot_download("yitongl/StreamRender-H3", local_dir="/your/code/models/StreamRender-H3")

The current runtime expects one merged backbone file. Reconstruct it first (requires an additional 66.28 GB; the script checks the original SHA256):

python /your/code/models/StreamRender-H3/download_assets.py
cd Sana/models/streamrender_h3
python scripts/preflight.py --assets /your/code/models/StreamRender-H3/assets.json
GPUS=8 PYTHON_BIN=python bash scripts/launch_local.sh \
  --config /your/code/models/StreamRender-H3/configs/runtime.json \
  --assets /your/code/models/StreamRender-H3/assets.json

Use the runtime's documented Python/CUDA/FA4 environment and install its frontend dependencies. The tested deployment uses eight GB200 GPUs within one NVLink domain, two denoising evaluations, CFG=1, S=2/W=2/C=2, 832ร—480 reference encoding and 1344ร—768 output. Qwen consumes prompt and first image once per session. The included assets.json uses relative paths.

The actual browser-to-H3-to-browser integration was verified in a same-rack eight-GPU run that exited successfully. The longer validation measured a 295.58 ms median GPU-resident RGB-to-RGB-ready interval and approximately 1.03 s browser publication intervals. These are different timing boundaries; the browser test is not a claim of 24-fps end-to-end feedback.

License and attribution

This bundle is not blanket Apache/MIT licensed. H3 weights and derivatives are governed by the included MiniMax H3 Community License, including its territorial and redistribution terms. Qwen components retain Apache-2.0 terms. TAE source attribution is preserved separately. See LICENSE, LICENSE-Qwen.txt, LICENSE-TAE.txt, NOTICE and SOURCE_SNAPSHOT.json. Consult your applicable authorization when using or redistributing the bundle.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support