MiniMax H3
MiniMax H3 creates high-quality videos with sound at a fixed framerate of 24 fps with a duration of up to ~15s. The FL2VA variant supports pure text prompting, as well as text prompting plus conditioning with a first frame and/or last frame image. The Ref2VA variant supports pure text prompting as well as up to 9 reference images and 3 reference videos. The reference videos can be used to guide the generated media’s video, audio, or both video and audio.
The H3 Models
Section titled “The H3 Models”MiniMax H3 ships as two task transformers that share every other component. FL2VA — first/last-frame to video with audio — covers text-to-video, first frame, last frame, both, and extend. Ref2VA — reference to video with audio — conditions a new clip on reference media: up to 3 videos and 9 images, in an order you choose (order changes the result). Every generation from either includes a jointly generated stereo soundtrack, muxed into the MP4 as AAC.
The panel switches task by transformer: select the Ref2VA single-file transformer as the Model and the conditioning sections become the ordered References list; select the FL2VA transformer to get the five frame/extend modes back. A video reference can contribute its image track, its soundtrack, or both — soundtrack-only references must be paired with at least one visual reference. Image references default to a high-detail 2048 px encoding (“Max”); the “Match generation size” option is several times faster at some fidelity cost.
What makes it different from Wan in practice:
-
Guidance-distilled — there is no CFG and no negative prompt. Prompt adherence is baked in; the panel hides both controls.
-
Fixed 24 fps and a fixed frame grid: 90–345 frames in steps of 17 (about 3.8 to 14.4 seconds), default 124.
-
Two fixed resolutions “768 highres” pins the short edge to 768 (native quality); “768 lowres” pins the long edge to 768 for fast previews.
-
Aspect ratios from 1:4 to 4:1 are supported, wider than the preset list.
-
Turbo fast path — step-distillation LoRAs render in ~6–8 steps instead of ~50. The panel’s Turbo toggle defaults to on when one is installed (8 steps).
-
Conditioning text encoder is a truncated Qwen3-VL-32B vision-language model, which is why H3 follows detailed scene descriptions well.
The hybrid: FL2VA quality with Ref2VA conditioning
Section titled “The hybrid: FL2VA quality with Ref2VA conditioning”The two transformers are not equal. FL2VA produces the cleaner output; Ref2VA is the only one that takes references, but a known training defect degrades what it generates — video and audio artifacts that FL2VA does not show. A tensor-by-tensor comparison of the two checkpoints shows they differ almost only in their per-block AdaLN modulation projections — the small time-conditioning layers that tell each block how far along the denoise it is. Attention, MLPs, and output heads are near-identical.
The hybrid exploits that: it runs the FL2VA transformer with Ref2VA’s AdaLN projections swapped in for the upper part of the network. References still route (the task is decided by those projections), while everything else keeps FL2VA’s quality.
To use it:
- Select the Ref2VA transformer as the Model, as for any reference generation.
- Under Model Components, pick the FL2VA transformer in Hybrid quality base (FL2VA). The slot lists only files of the same kind as the model — pruned with pruned, full with full — because the projections have different shapes otherwise.
- Leave Ref2VA AdaLN from block at its default of 25 to start. H3 has 50 blocks (0–49); blocks from the chosen one through 49 keep Ref2VA’s projections, earlier blocks use FL2VA’s. Lower values follow the references more closely at more of Ref2VA’s quality cost; higher values are cleaner but weaker on adherence.
Nothing is merged or written to disk: at generation time the loader loads the FL2VA base and the selected AdaLN tensors are read from the Ref2VA file and swapped in for the duration of the run. On the pruned starter repacks that costs about 40 MB of RAM; full-size pairs cost far more (roughly 520 MB per selected block). The Turbo toggle keys off the model, so it picks the Ref2VA Turbo LoRA (8 steps with the LightX2V release); the FL2VA Turbo can be added by hand under Concepts as an experiment, since neither distillation saw the hybrid’s mixed path. Your own LoRAs apply on top of the hybrid as usual.
The hybrid is an experiment adopted from the community — see the acknowledgements below. Early renders with the defaults were free of the artifacts Ref2VA alone produced, but judge it on your own results rather than treating it as a fixed improvement.
Ref2VA Prompt Enhancement
Section titled “Ref2VA Prompt Enhancement”Unlike FL2VA, Ref2VA was trained with a specific structured prompt format. This format is particularly important to follow when you have multiple character reference images and videos and you wish to assign specific references to different characters in your video.
To help you with this, InvokeAI provides an optional prompt enhancer. To use it, you will need to install one of the optional text LLMs, such as “Qwen2.5-1.5B-instruct.” After installing the model, enter your unenhanced prompt and click on the “sparkle pencil” icon. Select the LLM you installed, then select “MiniMax H3 Ref2VA Structured Prompt” from the System prompt menu. This will fill in a template structured prompt which you can then tweak.
When you use reference images and videos in a Ref2VA generation, you may refer to them within the prompt as <Picture 1>, <Video 1> and so forth. These labels appear adjacent to the reference image or video.
License
Section titled “License”MiniMax H3 is released under the MiniMax H3 Community License. Commercial use is allowed, with conditions: products earning more than $20M a year need written authorization from MiniMax, commercial products must show “MiniMax H3” in their interface, use is restricted in some territories (see the note under Installing MiniMax H3), and outputs may not be used to improve other AI models. The downloads are not gated.
This is a summary, not legal advice. Read the full license on the model’s HuggingFace page before using a model commercially.
Installing MiniMax H3
Section titled “Installing MiniMax H3”The MiniMax H3 starter bundle (~85 GB) installs the working set:
- MiniMax H3 FL2VA Transformer (int8, pruned) and MiniMax H3 Ref2VA Transformer (int8, pruned) (~21 GB each) — single-file transformer builds. One of these is the model you select in the Video panel, and the one you pick decides the task: FL2VA for the frame/extend modes, Ref2VA for reference conditioning (with FL2VA optionally in the hybrid base slot).
- MiniMax H3 Components (~11 GB) — the Diffusers-format install of the tokenizer, processor, and video + audio VAEs, which goes in the panel’s Model components slot — plus MiniMax H3 Text Encoder (int8) (~27 GB), the single-file build for its Text encoder slot (the Components install carries no text-encoder weights). The bundled templates wire them the same way.
- MiniMax H3 Turbo LoRA, MiniMax H3 LightX2V Turbo LoRA, and MiniMax H3 LightX2V Ref2V Turbo LoRA — independent step-distillation LoRAs; the Turbo toggle picks the one matching the selected task (the Ref2V build is trained against the Ref2VA transformer only). The earlier 4-step MiniMax H3 Ref2V Turbo LoRA (v0.1) is no longer a starter: it pans the camera rightward whenever a video reference is used. If you still have it installed, the toggle prefers the LightX2V release once that is installed too.
Reference conditioning adds rows to every denoising step roughly linearly: a few references typically cost ~1–2 GiB of extra activation memory, but a full-detail 2048×2048 image reference alone adds ~4096 rows (about 1 GiB), and three full-length video references can triple the packed sequence. Prefer “Match generation size” for image references and trimmed clips for video references on smaller cards.
H3 is the heaviest local video option — With 16 GB VRAM, enable_partial_loading enabled and 32 GB of system RAM it can create a full 345 frame video (about 14s) at 768x768 resolution, but for larger or longer videos a high-VRAM GPU (24 GB+) is strongly recommended.