Models/MiniMax H3
VIDEO / MINIMAX H3

Stories that move.
On your terms.

Explore the architecture of H3-Base: language conditioning, spatial and temporal compression, and joint audio–video generation. Bring video inference into a deliberate production workflow.

MINIMAX H3
TRANSFORMER33B dense
API MODEL SCOPEH3-Base
BASE OUTPUT768p
AUDIO32 kHz stereo
INSIDE THE MODEL / 3D ARCHITECTURE ATLAS

A shared timeline.
Sound and motion.

Compress space and time. Mix audio and video. Decode them into one coherent result.

PC / MINIMAX / SYSTEMINTERACTIVE SCHEMATIC
Drag to rotate · Arrow keys supported
MODULE 01 / MULTIMODAL CONDITIONING
50th
QWEN ENCODER HIDDEN-STATE LAYER

Different inputs. A common context.

H3-Base uses Qwen3-VL-32B hidden states from layer 50. Visual references also enter VisualVAE; audio references use AudioVAE. The encoded paths condition the packed generation sequence.

text / visual / audio → matching encoders → conditioning
READING THE DIAGRAM

H3-Base schematic. Context-IR and Regenerate-2K belong to separate hosted orchestration; the solid graph depicts the Base generation path. The animation is explanatory; no model inference is executed.

BACKBONE MAP

52 layers. A deliberate pattern.

50 Omni block2 Token refiner

Two token refiners and 50 Omni Transformer blocks. This inventory does not include the separate context encoder or VAEs.

A CLOSER LOOK / MINIMAX

The details behind
the diagram.

Configuration, memory and computation—each with a specific role in the model.

01 / MULTIMODAL CONDITIONING

Different inputs. A common context.

  • Context-IR is a separate hosted interpretation stage.
  • FL2VA and Ref2VA are distinct released Base families.
text / visual / audio → matching encoders → conditioning
02 / JOINT AUDIO–VIDEO TRANSFORMER

Let modalities communicate inside the model.

  • 33B describes the transformer, excluding Qwen and the VAEs.
  • Initial Base inference uses full, noncausal attention.
packed tokens → shared attention / FFN → velocity heads
03 / SPATIAL AND TEMPORAL COMPRESSION

A video becomes a volume of patches.

  • The visual latent has 24 channels; effective spatial token factor is 32.
  • Audio is represented at 40 latent steps/second/channel.
video → f16t4d24 → 1×2×2 patches / stereo audio latents
04 / REUSABLE MODULATION

Cache the modulation, not yesterday’s frames.

  • The cache holds scale, shift and gate values; Q/K/V still follow current latents.
  • Precomputed AdaLN and approximate Cache-DiT skipping are different mechanisms.
timestep schedule → AdaLN tables → block modulation
FOLLOW THE RESEARCH

Go straight
to the source.

Developer papers, released configurations and reference implementations. Reviewed 10 September 2026.

WHAT THIS GUIDE COVERS

H3-Base produces 768p output. Context-IR and the complete 2K regeneration workflow are separate components.

Native sparse attention was not included in the initial Base inference release. Approximate sparse backends and FastH3 derivatives are distinct.

MADE MEANINGFUL IN YOUR BUSINESS

From capability
to possibility.

01

Product stories

Develop moving concepts and short-form product narratives.

02

Creative exploration

Iterate on scenes and direction inside an agreed generation workflow.

03

Content pipelines

Plan capacity around video length, resolution, batch size and production demand.

Know your inference service scope.

Perchy provides managed AI inference through authenticated APIs. Customers do not receive SSH, root, operating-system or bare-metal access. The service does not provide model-training or fine-tuning functionality. Server administration remains with Perchy.

Perchy’s managed inference scope covers the open H3-Base path. Context-IR and 2K regeneration are separate hosted components and are not implied here. Shape calculations are illustrative; conditioning, audio, padding and runtime buffers also consume memory. Cost scenarios are not performance or price guarantees.

Read the model developer’s documentation
KEEP EXPLORING

Another model.
Another way to compute.

All models
LANGUAGE

Qwen3.8-Flash-Next

Explore architecture
LANGUAGE

DeepSeek-V4.1-Flash

Explore architecture
LANGUAGE

GLM-5.3-Flash

Explore architecture
IMAGE

Ideogram 4.0

Explore architecture
THE NEXT MOVE IS YOURS

Find your
unfair advantage.

Let’s build your direction

One conversation.
A whole new trajectory.