Different inputs. A common context.
- Context-IR is a separate hosted interpretation stage.
- FL2VA and Ref2VA are distinct released Base families.
text / visual / audio → matching encoders → conditioningExplore the architecture of H3-Base: language conditioning, spatial and temporal compression, and joint audio–video generation. Bring video inference into a deliberate production workflow.
Compress space and time. Mix audio and video. Decode them into one coherent result.
H3-Base uses Qwen3-VL-32B hidden states from layer 50. Visual references also enter VisualVAE; audio references use AudioVAE. The encoded paths condition the packed generation sequence.
text / visual / audio → matching encoders → conditioningH3-Base schematic. Context-IR and Regenerate-2K belong to separate hosted orchestration; the solid graph depicts the Base generation path. The animation is explanatory; no model inference is executed.
Two token refiners and 50 Omni Transformer blocks. This inventory does not include the separate context encoder or VAEs.
Configuration, memory and computation—each with a specific role in the model.
text / visual / audio → matching encoders → conditioningpacked tokens → shared attention / FFN → velocity headsvideo → f16t4d24 → 1×2×2 patches / stereo audio latentstimestep schedule → AdaLN tables → block modulationDeveloper papers, released configurations and reference implementations. Reviewed 10 September 2026.
H3-Encoder, H3-VAE, Omni Transformer and system boundary
Model architecture, Base variants and VAE
50 transformer layers, two refiners and attention dimensions
Precomputed cache and online rebuild; runtime-specific behavior
H3-Base produces 768p output. Context-IR and the complete 2K regeneration workflow are separate components.
Native sparse attention was not included in the initial Base inference release. Approximate sparse backends and FastH3 derivatives are distinct.
Develop moving concepts and short-form product narratives.
Iterate on scenes and direction inside an agreed generation workflow.
Plan capacity around video length, resolution, batch size and production demand.
Perchy provides managed AI inference through authenticated APIs. Customers do not receive SSH, root, operating-system or bare-metal access. The service does not provide model-training or fine-tuning functionality. Server administration remains with Perchy.
Perchy’s managed inference scope covers the open H3-Base path. Context-IR and 2K regeneration are separate hosted components and are not implied here. Shape calculations are illustrative; conditioning, audio, padding and runtime buffers also consume memory. Cost scenarios are not performance or price guarantees.
Read the model developer’s documentationOne conversation.
A whole new trajectory.