MiniMax H3 is a browser-based AI video generator for creating 2K video from a clear creative brief. Pick the input that matches your starting material—text, a still image, a first-and-last frame pair, or a subject-reference face—then describe the motion, camera, and atmosphere you want. This guide walks through each mode, when to use it, and how to review the result before you download. For prompt patterns and examples by mode, see the MiniMax H3 Prompt Guide.
MiniMax H3 modes at a glance
| Capability | Quick draft (Seedance 2.0 Mini / Fast) | MiniMax H3 |
|---|---|---|
| Primary job | Fast iterations at lower cost | Directed 2K creative output |
| Text to video | Supported | Supported |
| Image to video | Supported | Supported |
| First & last frame | Limited / workflow-dependent | Supported |
| Subject reference | Limited / workflow-dependent | Supported |
| Output focus | Speed and exploration | Clarity of prompt and input control |
MiniMax H3 workflow highlights
1. Text to video for a written scene
Describe the subject, setting, action, lighting, and camera move in natural language. Use this mode when you want to explore a concept before you commit to reference media.
2. Image to video for a still that should move
Upload a starting image and prompt the motion that should unfold from it. Keep the visual idea of the plate while adding atmosphere, camera travel, or character action.
3. First and last frame for controlled transitions
Provide an opening and closing image so the model bridges the two states. Best for before-and-after beats, reveals, and transformations where endpoints matter more than free improvisation.
4. Subject reference for a recognizable face
Use a clear face or character plate when identity consistency is the priority. Keep the reference simple and pair it with a prompt that states pose, wardrobe, and environment separately.
5. Prompt-level camera direction
Call out pan, push-in, static framing, orbit, or handheld feel in the prompt. Clear camera language usually improves continuity more than stacking many vague adjectives.
MiniMax H3 generation modes
1. Text to Video
Start with a written brief only. Name the subject, place, action, lighting, and camera move in one coherent paragraph. Keep the shot list short so the model can commit to a single readable beat.
- Best for: Concept clips, storyboard exploration, and social drafts when you do not yet have locked artwork.
- Prompt tip: Lead with subject + action, then add setting, then camera. Avoid packing multiple unrelated scenes into one prompt.
2. Image to Video
Upload a still and describe how the scene should evolve. The image sets look and composition; the prompt sets motion and atmosphere.
- Best for: Product plates, portraits, key art, and campaign stills that need a short motion pass.
- Prompt tip: Say what should stay stable (face, product, logo) and what should move (camera, cloth, light, weather).
3. First & Last Frame
Provide a start frame and an end frame, then describe the transition between them. The prompt should focus on how the change happens, not on inventing a third unrelated scene.
- Best for: Before-and-after concepts, controlled reveals, and visual story beats with known endpoints.
- Prompt tip: Name the transformation clearly—dissolve, camera push, wardrobe change, day-to-night—so the bridge between frames stays intentional.
4. Subject Reference
Use a subject-reference image when a face or character must stay recognizable. Pair the plate with a prompt that covers action, wardrobe, and environment without asking the model to invent a new identity.
- Best for: Character-led concepts, UGC-style creator spots, and campaigns that reuse the same person across variants.
- Prompt tip: Prefer one clean face plate over many similar crops. Describe pose and setting in text instead of stacking near-duplicate references.
Frequently Asked Questions
What inputs does MiniMax H3 accept?
Text input: Natural-language scene, motion, lighting, and camera direction.
Image input: Starting image, end frame, or subject-reference face (jpeg, png, webp).
Frame pair: First and last frames to guide a transition or reveal.
Subject reference: A face or character plate for identity consistency.
Output: 2K AI video, with duration and resolution selected in the generator before render.
Which generation mode should I pick first?
Use Text to Video for a written concept with no locked artwork. Use Image to Video when you already have a strong still. Use First & Last Frame when the start and end states are known. Use Subject Reference when a recognizable face is the priority.
How do I keep a subject recognizable across variants?
Upload one clear subject-reference plate, keep wardrobe and identity instructions consistent, and change only the action or setting in the prompt. Avoid mixing many near-duplicate face crops in the same job.
When should I use Seedance 2.0 Mini or Fast instead?
Use Mini or Fast for cheaper, faster prompt experiments while you explore composition. Move to MiniMax H3 when you need clearer direction across text, image, frame-pair, or subject-reference workflows for a more finished 2K concept.
How do I refine a clip after the first render?
Tighten the prompt, swap or simplify reference images, and re-render. Change one variable at a time—camera, action, or lighting—so you can tell what improved. Save winning prompts in the Prompt Gallery workflow for reuse.
How many credits does a typical render use?
Credits scale with duration, resolution, and model choice. The generator shows the exact total before you submit. New accounts receive starter credits to try the workflow. Check our Pricing Page for detailed credit packages.
Create 2K video with MiniMax H3
Start from text, an image, a frame pair, or a subject reference—then generate, review, and download in your browser.