MiniMax has launched H3, a general-purpose multimodal generation model that breaks the boundaries between tasks and modalities.
One Model, Every Modality
H3 understands unified context across text, images, video, and audio. A prompt can reference multiple inputs across modalities — “use the Hitchcock camera movement from Video 1, have the character in Image 2 sing, with the vocals matching Audio 3” — and H3 handles the full-modality understanding on its own. All audio output is native stereo; multi-shot modeling is native.
The design philosophy: previous generations split generation into siloed expert tasks (T2I, editing, subject reference, motion reference, and so on for each modality). H3 unifies and generalizes across tasks from the pretraining stage, using language as the generalizable bridge.
2K Video at a Fraction of the Price
H3 generates up to 15 seconds of 2K-resolution video with synchronized stereo audio. The pricing is aggressive: at 2K, the per-second price is less than a third of mainstream models; at 768p, it’s less than half the price of rivals’ 720p. Early testing shows commercial readiness across advertising, branding, e-commerce, product design, UI/UX, and gaming — with accurate text and brand rendering plus V2V motion transfer.
Open Weights Incoming
Closed-source models have long dominated video generation. MiniMax says it plans to open the model weights within days of launch, subject to applicable laws — to support the open-source community, broaden AI hardware compatibility, and let users build customized versions.
The technical stack includes H3-VAE (a completely overhauled tokenizer enabling native 2K via high compression), the H3-Omni Transformer (an architecture built for task generalization that lifted training throughput ~30%), and In-Context Regeneration (the base model regenerates its own low-res output in context for 2K, recovering detail traditional super-resolution can only guess at).
Try it at hailuoai.video; read the full announcement at minimax.io/blog/minimax-h3.