ModelRadar
/

MiniMax-H3

MiniMax🇨🇳 CNMajor lab· Jul 28, 2026

Videoopen weights

MiniMax H3 is a general-purpose video generation model with native stereo audio support. It takes text, images, videos, and audio as inputs and outputs synchronized audio-video clips between 4 and 15 seconds long at up to 2K resolution and 24 FPS. The input side is highly flexible, supporting first-and-last-frame mode, multi-reference inputs with up to nine images, three video clips, and three audio clips combined. A dedicated preprocessing component called H3-Context-IR parses and relates the multimodal context before generation. The model communicates fluently in eleven languages. Performance details beyond the specification sheet are not covered in the provided sources.

Summary of the model card

License
other
Modalities
text → video

1 variant

StandardHF

Sources

HuggingFace
Deutsch

MiniMax H3 ist ein allgemeiner Videomodell mit nativem Stereo-Audio. Es nimmt Text, Bilder, Videos und Audio als Eingabe und erzeugt synchronisierte Audio-Video-Clips zwischen vier und sechzehn Sekunden Dauer in bis zu zwei K Aufloesung und 24 FPS. Die Eingabe ist sehr flexibel, unterstuetzt First-and-Last-Frame, mehrere Referenzbilder sowie gemischte Inputs aus bis zu neun Bildern, drei Video- und drei Audioclips. Eine Vorsystem namens H3-Context-IR versteht und verarbeitet die multimodalen Angaben vor der Generierung. Das Modell beherrscht elf Sprachen fluessig. Genauere Leistungskennzahlen enthalten die Quellen nicht.

Provenance

Detected
Sep 27, 2026
First source
huggingface
Description
Summary of the model card
Editorially reviewed
—