Example input
News report
Open-weight video model by MiniMax
Direct scenes with text, frames, images, video, and audio references—then plan the whole sequence before you render.
Start from a shot briefPreview
Example input
News report
Loading history…
Enter to generate · Shift + Enter for a new line
0/5000
Story Workshop
Decide how you want AI to help
Choose a conversation starter, review the brief, then open it in your preferred AI chat.
Choose a starting prompt.
The AI will guide the discussion and finish with one copy-ready video prompt.
Open in your preferred AI chat.
Your prepared prompt is sent to the service you choose and may appear in its URL, browser or account history, and logs. That provider's terms and privacy policy apply.
Watch the model work
See the range MiniMax chose to introduce H3: character detail, title design, interactive concepts, and vertical motion graphics.

A branded motion study with synchronized visual beats.
Cinematic opening titles with controlled typography and mood.
A polished interface concept presented as a moving product shot.
A portrait-format animated poster built for social feeds.
Media is sourced from the model publisher's official showcase.
Why this model matters
H3 accepts more than a text description. Its unified context can combine frames, images, video, and audio so direction lives in one brief.
Direct dialogue, ambience, sound effects, and music alongside the picture with native stereo output.
Use first and last frames to define where a shot begins and where its motion must land.
Guide character, object, style, motion, or sound with image, video, and audio references in one context.
Plan for 24 FPS output and an up-to-2K regeneration workflow when final detail matters.
Start with direction
These starters are shaped around H3's multimodal strengths. Load one into the current tool, change the fruit, setting, sound, or camera, and make it yours.
Dialogue, ambience, and two deliberate camera beats.
A beginning and ending composition for controlled motion.
Character, movement, and camera direction in one brief.
A compact three-shot sequence with readable scene text.
Same foundation, different priorities
Choose H3 when broad reference control and the 2K regeneration path matter most. Choose H3 Max when fast iteration and instruction following matter more than maximum resolution.
Best for
MiniMax H3
Omni-modal control and a higher-resolution finish
H3 Max by fal
Fast, prompt-faithful iteration
Resolution
MiniMax H3
Up to 2K through regeneration
H3 Max by fal
480p or 768p
Direction
MiniMax H3
Text, frames, images, video, and audio
H3 Max by fal
Text, start/end images, and references
Core strengths
MiniMax H3
Reference control, native stereo, and editing
H3 Max by fal
Speed, adherence, aesthetics, and visual coherence
Prompt recipe
A strong video brief reads like compact direction, not a bag of style words. Give the model an ordered scene it can execute.
Straight answers
MiniMax released H3 as an open-weight omni-modal video model. fal is one of the platforms that serves the base model through an API.
H3 is MiniMax's base model, with broad multimodal reference controls and an up-to-2K regeneration path. H3 Max is a separate fal Research post-trained variant focused on faster generation, prompt adherence, and aesthetics at 480p or 768p.
Yes. MiniMax documents native 32 kHz stereo audio, so a prompt can direct dialogue, ambience, sound effects, and music together with the picture.
H3 can combine text, first and last frames, images, video, and audio references in one context. MiniMax also documents editing and an up-to-2K regeneration workflow.