MiniMax H3 LogoMiniMax H3

User Guide

MiniMax H3 Prompt Guide

Choose the right prompt format, then direct shots, dialogue, references, ambience, and music without mixing their roles.

Choose the format before you write

The uploaded asset does not decide the task by itself. Decide what role the asset plays in the target video, then use the matching structure.

Base format: T2VA

Text to Video

Start from a written audiovisual idea with no frame anchor.

Begin with the three core fields. No image-alignment instruction is required.

Base format: I2VA

Image to Video

Use one image as the exact first frame, then develop forward from it.

Place the first-frame alignment instruction before the three core fields.

Base format: FL2VA

First and Last Frame

Use two images as exact opening and ending frames.

Describe the observable motion path between the frames. A single continuous shot is usually the clearest fit.

Base format: L2VA

Last Frame to Video

Use one image as the exact final frame while allowing the model to infer a compatible opening state.

Align the image to the exact end time and make the action, object state, framing, and lighting progressively converge on it.

Reference format: six sections

Full-Reference

Reuse or transform subjects, scenes, motion, editing, audio, or other attributes from reference assets.

Define each reference role first, then track its use and retention through the target timeline.

Base prompt format

T2VA, I2VA, FL2VA, and L2VA share one audiovisual body. Image-based tasks add a precise alignment line before that body.

The three core fields

integrated_multimodal_description
Write the visual timeline here: style, composition, subjects, actions, shots, speakers, dialogue, singing, and synchronized in-scene sound.
overall_soundscape
Summarize ambience, physical action sounds, and non-verbal human sounds across the full clip. Do not repeat dialogue or music here.
non_diegetic_music
Describe audience-only background music by instrumentation, tempo, rhythm, and dynamic change. Write N/A when no score is wanted.

How each Base task begins

T2VA

No alignment line

Start directly with integrated_multimodal_description and build the complete audiovisual timeline from text.

I2VA

For the target video, at 0.00 seconds, <Picture 1> from [Shot 1] is fully referenced.

Anchor style, identity, clothing, composition, and space to the image, then describe continuous development.

FL2VA

Picture 1 from [Shot 1] aligns to 0.00 seconds; Picture 2 from the actual final [Shot N] aligns to the exact S.SS-second end time.

Write the changing pose, object state, composition, lighting, and camera path that progressively lands on Picture 2.

L2VA

<Picture 1> from the actual final [Shot N] aligns to the exact S.SS-second end time.

Infer a plausible earlier state, then make every change converge on the final image.

Build one observable timeline

The main field is not a plot summary. Every sentence should describe something visible or audible at a specific point in the clip.

  • Start [Shot 1] with the visual style, initial framing, subject appearance and position, environment, lighting, and the first action.
  • Do not put a timestamp on [Shot 1]. Start every later shot with [Shot N] At MM:SS.mmm, using strictly increasing cut times inside the requested duration.
  • A cut must introduce new subject, space, state, viewpoint, or time information. Use camera motion for a smaller distance or angle change.
  • For I2VA, develop forward from Picture 1. For FL2VA, describe the visible path between both anchors. For L2VA, infer a plausible earlier state and converge on the final picture.
  • Keep identity, clothing, colors, key objects, and spatial relationships stable unless the prompt explicitly describes their change.
[Shot 2] At 00:03.500, the shot cuts to a close-up of the biscuit filling as the two halves separate and crumbs scatter outward.

Write camera motion as an action

A useful camera instruction names the motion type, then adds amplitude and speed only when they affect the result.

IntentVocabularyWhat changes
Lens changeZoom In / Zoom OutThe camera stays in place while focal length changes.
Depth travelPush In / Pull OutThe camera body moves toward or away from the subject.
Horizontal travelTruck Left / Truck RightThe whole camera moves sideways without merely rotating.
Vertical travelPedestal Up / Pedestal DownThe whole camera rises or lowers through space.
Fixed rotationPan Left / Pan Right / Tilt Up / Tilt DownThe camera body stays in place while the lens rotates horizontally or vertically.
Around the subjectArc ShotThe camera travels on a curved path around the subject.
Subject followTracking ShotThe camera moves with a moving subject and maintains the relationship.
Subject viewpointPOVThe framing adopts what a subject sees.
Axis rotationRoll Clockwise / Roll CounterclockwiseThe camera rotates around the lens axis.
InstabilityShake Slightly / Shake StronglyControlled camera shake changes the stability of the frame.
No movementStatic ShotThe position and lens remain still.

with small amplitude

Use for a restrained framing change that keeps the original composition readable.

with large amplitude

Use when the camera must reveal substantially more space or transform the composition.

at slow speed

Use for a gradual move with enough time to inspect a subject or product detail.

at fast speed

Use for a rapid reveal or energetic transition that is still physically legible.

The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.

Bind every spoken line to a stable speaker

  • Assign speakers in vocal-event order as (S1), (S2), and so on. Reuse the same ID in later shots.
  • Put the speaker identity, delivery, and action outside the dialogue tag.
  • Inside <d>, keep only the language tag and the exact spoken words.
  • For voiceover, state that it is off-screen and keep the visible character lips closed.
  • Use <scenetrans> when one line continues across a cut and <cutoff> when the video ends mid-line.
The baker with a calm, slightly raspy voice (S1) says: <d>[English] First batch of the morning.</d>

Preserve visible text exactly

Put text that must appear on a sign, label, subtitle, banner, interface, or product frame inside English double quotation marks. Keep the original language, words, capitalization, and punctuation. Do not translate it or paraphrase it.

A red neon sign above the door reads "OPEN" and remains unobstructed.

Keep the three audio layers separate

Shot-level sound
Put a synchronized door slam, spoken line, radio cue, or instrument played on screen in the timeline body.
Overall soundscape
Summarize continuing ambience and physical sounds. Use N/A only when complete silence is explicitly required.
Non-diegetic music
Describe only the score that characters cannot hear. Do not use abstract mood words in place of instrumentation and tempo.

Base format template

For I2VA, FL2VA, or L2VA, add the exact image-alignment instruction as the first line. T2VA starts at the first field.

integrated_multimodal_description: [Shot 1] [Style and initial composition]. [Subject identity, position, action, and camera motion]. [Speaker (S1) and <d>[Language] exact dialogue</d> when needed]. [Shot 2] At 00:SS.mmm, the camera cuts to [new information].

overall_soundscape: [Ambience, physical action sounds, and non-verbal human sounds across the clip.]

non_diegetic_music: [Instrumentation, tempo, rhythm, and dynamic change, or N/A.]

Full-Reference prompt format

Full-Reference mode is not a longer Base prompt. It starts by defining what each asset contributes, then records how that contribution is used in the target video.

subject_definitions

Define every reusable subject, concrete frame anchor, whole-video source, and audio source.

summary

State the task type and the main reference relationships in one short paragraph.

retention_analysis

Record how each defined visual or audio reference is preserved, transferred, copied, or loosely followed.

detailed_description

Write the target video shot by shot in playback order, inserting reference labels where their roles take effect.

overall_soundscape

Summarize ambience and physical sounds, including any applicable audio-copy or audio-reference relationship.

non_diegetic_music

Describe the audience-only score and its relationship to referenced audio, or write N/A.

Reference labels

<Subject N>
Reusable visible content extracted from one or more assets: a person, object, scene, clothing, prop, interface, style, action, pose, or effect. It describes the content unit used in the target, not the uploaded file itself.
<Picture N>
The image itself is a concrete first frame, keyframe, last frame, edit keyframe, composition anchor, or storyboard plan that must be cited as a frame.
<Video N>
The whole source video is directly edited, continued, or used for global camera, cut, rhythm, pacing, or timing structure. A person or action extracted from it is still a Subject.
<Audio N>
An actively used audio signal: fully or partly copied, or referenced for voice timbre, music, words, effects, beat, rhythm, or continuity. A video containing sound does not automatically require this label.

Task types

keyframe completion
Use when a picture is a concrete target first frame, keyframe, last frame, edit keyframe, or other exact frame anchor.
reference generation
Use when an asset guides identity, scene, style, action, camera, storyboard, sound, pacing, or rhythm without being directly edited or continued.
video editing
Use only when the existing source video itself is directly modified. Editing or interpolating between still images is not video editing.
video continuation
Use only when new content continues, extends, resumes, or transitions from an existing source video.
audio reuse
Use when the same audio signal is copied in full or in part. If an edited source video keeps audible original audio, include this type too.
audio reference
Use when the signal is not copied and only timbre, delivery, music style, words, effects, beat, rhythm, or continuity is followed.

Decide whether the asset needs its own label

  • If an image only defines a character, product, scene, clothing, or style, cite the image inside the Subject definition. Do not create a separate Picture line.
  • Create a Picture line only when the image itself anchors a specific frame or provides explicit storyboard and composition planning.
  • Create a Video line only when the source video has a whole-video role such as editing, continuation, cut structure, camera path, pacing, or timing.
  • Visual content extracted from a video remains a Subject even when the same file also has a Video label for its structural role.
  • Video and Audio numbering are independent. Define Audio only when a sound signal has an active copy or reference role.
  • One Subject may combine appearance from a picture, motion from a video, and voice timbre from audio. State what each source contributes.

Retention vocabulary

fully_preserved
Visual: retain the full defining role and identity of the reference content. Name the concrete features that remain unchanged.
partially_preserved
Visual: keep the reference content but alter or omit part of its defining appearance, state, structure, or use. State exactly what changes.
attribute_transfer
Visual: transfer a recognizable property such as motion, style, pose, composition, or effect to a different target subject.
weak_reference
Visual: follow only a broad category, atmosphere, composition family, or stylistic resemblance. Do not imply identity-level fidelity.
fully_copy
Audio: reuse the entire source audio signal as the complete final target track.
partially_copy
Audio: copy only selected time ranges or layers, or modify the copied signal by adding, removing, or replacing other sounds.
reference
Audio: follow timbre, delivery, beat, music style, dialogue content, or sound texture without copying the source signal.
weak_reference (audio)
Audio: preserve only a broad sound category or atmosphere with no close signal or performance match.

Write detailed_description as the production plan

This is the main body of a Full-Reference prompt. It must explain the target video in playback order and show exactly where each reference becomes visible or audible.

  • Write one or two sentences of overall visual style before [Shot 1]. In Base mode, style begins after [Shot 1]; this ordering is different.
  • For each shot, establish current framing, subject appearance and position, environment, lighting, visible action, state change, camera behavior, synchronized sound, and the event that leads to the next shot.
  • Insert a Subject, Picture, Video, or Audio label at the point where its defined role actually takes effect. Reuse the same label later without redefining it.
  • Use natural frame language: the shot begins from <Picture 1>, the keyframe corresponds to <Picture 2>, or the shot ends on <Picture 3>.
  • When a referenced Subject speaks, combine the visual and vocal identities as <Subject N> (Sx). Keep (Sx) out of retention_analysis.
  • If referenced speech or lyrics are reused, preserve the exact words inside <d>. Mark unintelligible source words as [unclear] instead of guessing.

For a typical generation task, 350 to 500 English words is a useful detail range, not a quota. Dialogue-heavy clips should prioritize the complete spoken timeline. Editing prompts may be shorter or longer according to the source video complexity, and a single shot can still require substantial detail.

Rules that prevent reference drift

  • Write all six section bodies in English. Keep only exact dialogue, lyrics, and text visibly present in the scene in their original language.
  • A source file and the content extracted from it are different things. A person from Video 1 is still a Subject, while Video 1 names the whole-video source or structure.
  • Do not create an Audio label only because a reference video contains sound. Define it only when that audio has an active role.
  • One subject may combine identity from a picture, movement from a video, and voice timbre from audio.
  • Combine every applicable summary task type with +, without duplicates. The presence of a video or audio file alone does not justify video editing, continuation, reuse, or reference.
  • Keep every label meaning stable across all six sections. Do not introduce new labels in the summary.
  • Generation tasks normally need a detailed_description rich enough to specify composition, action, camera, sound, and reference activation, not a plot summary.

Full-Reference template

Complete the sections in this order. Remove unused labels, but do not remove required sections.

subject_definitions:
<Subject 1> is [reusable visible content] from <Picture 1>, preserving [identity-defining features].
<Video 1> is [editing, continuation, or structural role] when applicable.
<Audio 1> is [copy or reference role] for <Subject 1> (S1) when applicable.

summary:
[task type + task type] [Target video and main reference relationships.]

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - [features retained].
<Audio 1>: reference - [audio characteristic followed without copying the signal].

detailed_description:
[Overall visual style in one or two sentences.]
[Shot 1] [Composition, references, subject action, camera, synchronized sound, and dialogue.]
[Shot 2] At 00:SS.mmm, [cut and new information].

overall_soundscape: [Ambience, physical sounds, and audio relationship.]

non_diegetic_music: [Score and audio relationship, or N/A.]

Worked video case: cafe character plus motion and audio references

This case combines one image, one short motion clip, and one audio file. Study how the prompt gives each asset a narrow job, then compare the final output with the three references.

What the prompt is trying to produce

A short cinematic cafe scene. The woman and street-cafe composition come from the image, the close-up body rhythm and tense reaction come from the video, and the audio file guides vocal pacing and sound texture.

Image 1 owns identity and scene

The prompt preserves the woman, clothing, cup, street cafe, warm daylight, shallow depth of field, and background customers from the image. These details are not left for the video reference to decide.

Video 1 transfers only motion

The beach setting and original people in the source video are excluded. The useful reference is the compact close-up framing, slight backward tension, forward recovery, and uneasy facial rhythm.

Audio 1 is labeled as audio reference

The prompt does not pretend the audio is a visual subject. It marks the file as an audio reference for cadence, tonal texture, and timing, then keeps ambient cafe sound in overall_soundscape.

The output goal is measurable

Success can be checked frame by frame: the woman stays recognizable, the cafe remains the setting, the reaction follows the motion clip, and the spoken or vocal rhythm follows the audio without adding unrelated plot.

A reliable writing workflow

Build the prompt in passes. Each pass resolves one production decision and prevents contradictions from accumulating.

  1. Classify the task

    Decide whether each asset is a frame anchor, reusable subject, source video, structural reference, copied audio, or audio reference.

  2. Lock the timeline

    Choose the clip duration, shot count, cut times, and the observable state at the beginning and end of every shot.

  3. Write visible change

    Describe what moves, what stays fixed, and how pose, object state, composition, lighting, and camera evolve.

  4. Bind speech and sound

    Assign stable speaker IDs, preserve exact dialogue, and separate synchronized sound, ambience, and audience-only music.

  5. Audit references

    Check that every label has one stable meaning and that every claimed preservation or copy relationship is visible in the timeline.

  6. Remove contradictions

    Delete camera moves, cuts, identity changes, or audio instructions that compete with the chosen frame and reference constraints.

Final output checklist

Run this audit after writing. A structurally complete prompt can still fail when timing, labels, dialogue, or sound roles contradict one another.

All modes

  • The correct image-alignment instruction is the first line when required, followed by one blank line. T2VA has no alignment line.
  • S.SS uses exactly two decimal places, and [Shot N] is the actual final shot index.
  • [Shot 1] has no timestamp. Every later cut time is strictly increasing and remains inside the requested duration.
  • Every cut introduces new subject, space, state, viewpoint, or time information. Minor reframing uses camera motion instead.
  • Speaker IDs stay stable across shots. Non-speaking characters do not receive an ID.
  • Each <d> block contains only the language tag and exact spoken words. Visible text stays in double quotation marks.
  • overall_soundscape contains ambience and physical sounds, not repeated dialogue, singing, or on-screen music.
  • non_diegetic_music describes instrumentation, tempo, rhythm, and dynamics, or uses N/A for no audience-only score.

Full-Reference only

  • All six sections are present and remain in the required order.
  • Every reference label is defined once and keeps the same meaning everywhere.
  • The summary task prefix matches the actual role of every asset and introduces no new labels.
  • retention_analysis uses the correct visual or audio relationship marker and contains no (Sx) speaker IDs.
  • The overall style appears before [Shot 1], and each label is cited where its role becomes active in the timeline.
  • Copied audio and referenced audio are distinguished. Merely uploading a video with sound does not create an Audio label.

Common failures and precise fixes

Open a failure that matches your result. Each diagnosis explains what you will see, why the prompt causes it, and the exact rewrite pattern to use.

Using one generic formula for every mode

Typical output

The output ignores a frame anchor, treats a reference as loose inspiration, or follows the scene but loses the intended identity and timing.

Why it happens

A Base task describes one audiovisual timeline. Full-Reference must first define what every asset contributes and how strongly that contribution is retained. One generic paragraph cannot express both jobs reliably.

How to fix it

Use the Base three-field format for T2VA, I2VA, FL2VA, and L2VA. Use the six ordered sections for Full-Reference. Decide the task before writing any visual detail.

Rewrite pattern

Full-Reference: define <Subject 1> from @Image1, state [reference generation], record fully_preserved, then place <Subject 1> in the required shots.

Treating every uploaded file as a Subject

Typical output

The model copies the wrong property from a file, confuses a person with the entire source video, or tries to reproduce a reference frame that was only meant to define identity.

Why it happens

A file is a source container, while a Subject is reusable visible content extracted from that source. Picture, Video, and Audio labels describe different relationships and cannot be exchanged freely.

How to fix it

Use Subject for a person, object, scene, style, action, or effect. Use Picture only for a concrete frame anchor, Video for editing, continuation, or whole-video structure, and Audio only when sound is copied or referenced.

Rewrite pattern

<Subject 1> is the woman from @Image1. <Video 1> is @Video1 and provides only the target cut and pacing structure.

Describing first and last frames as two static images

Typical output

The clip holds on the opening state, jumps abruptly near the end, or reaches the last image through an implausible morph.

Why it happens

FL2VA already knows the two endpoints. Repeating their appearance does not tell the model how pose, object state, composition, light, and camera should change between them.

How to fix it

Describe a continuous chain of observable intermediate changes. Keep the differences narrowing until the final pose, object arrangement, lighting, and framing match Picture 2.

Rewrite pattern

She releases the handle, raises the umbrella, opens the canopy, rotates it into the final angle, and settles into the spacing shown in Picture 2.

Adding cuts for minor framing changes

Typical output

The video feels fragmented, invents new backgrounds, changes subject details, or loses continuity even though the intended action is simple.

Why it happens

Every cut invites the model to establish a new subject state, viewpoint, space, or time. A closer view alone does not require that reset.

How to fix it

Use push, pull, pan, tilt, truck, arc, tracking, or zoom for a distance or angle change. Add a cut only when the next shot contributes genuinely new information.

Rewrite pattern

Keep [Shot 1] and write: The camera pushes in with small amplitude at slow speed toward the filling, instead of creating [Shot 2].

Dialogue comes from the wrong character

Typical output

A line is spoken off-screen, assigned to another face, delivered by both characters, or accompanied by incorrect lip movement.

Why it happens

A character description alone does not create a stable vocal identity across shots. The dialogue tag also cannot carry speaker identity or delivery instructions.

How to fix it

Assign each actual vocal source one stable (Sx) ID in vocal-event order. Put identity, voice, action, and delivery outside <d>. Keep only the language and exact words inside <d>.

Rewrite pattern

The young woman with a quiet voice (S1) turns to camera and says: <d>[English] I get off at the next station.</d>

Music, ambience, and dialogue are mixed together

Typical output

Background music behaves like an object in the scene, dialogue is repeated, sound effects arrive at the wrong moment, or the clip adds music when silence was intended.

Why it happens

MiniMax H3 separates synchronized shot events, full-clip ambience, and audience-only score. Repeating one sound across multiple fields creates competing instructions.

How to fix it

Place dialogue, a door slam, a radio cue, or an on-screen instrument in the shot timeline. Summarize continuing ambience in overall_soundscape. Put only audience-only score in non_diegetic_music.

Rewrite pattern

overall_soundscape: Rain and train-wheel rhythm continue. non_diegetic_music: N/A.

Reference fidelity is asserted but not described

Typical output

The output is broadly similar to the reference but changes the face, clothing, product construction, color, or movement that actually matters.

Why it happens

Words such as keep consistent do not identify which features define success. The model also needs to know where each reference becomes active in the timeline.

How to fix it

Name the identity-defining or product-defining features in subject_definitions, choose the correct retention marker, and cite every shot where the reference appears or transfers its attributes.

Rewrite pattern

<Subject 1>: fully_preserved - retain the biscuit shape, baked surface, pink filling, and clean product appearance in [Shot 1] through [Shot 6].

Prompt reads like a plot summary

Typical output

The clip captures the general topic but invents staging, skips important actions, produces weak camera direction, or fails to land on the intended ending.

Why it happens

A plot summary explains meaning but not production. The model can only render concrete visual and audible events along a timeline.

How to fix it

For every shot, specify composition, subject position, visible action, state change, camera behavior, synchronized sound, and the condition that ends the shot.

Rewrite pattern

Replace "the product becomes exciting" with "the biscuit snaps at center frame; filling expands; crumbs scatter; the camera holds close until the fragments settle."

MiniMax H3 prompt FAQ

Answers follow the official Base and Full-Reference prompt-writing guides.

Do all MiniMax H3 modes use the same prompt structure?

No. T2VA, I2VA, FL2VA, and L2VA use the three Base fields, with an image-alignment instruction added for image-anchored tasks. Full-Reference uses six ordered sections beginning with subject_definitions.

What belongs in integrated_multimodal_description?

Write the visible and audible timeline: style, composition, subject identity and position, actions, shot changes, camera motion, speakers, exact dialogue, singing, and synchronized in-scene sound.

How should I prompt First and Last Frame video?

Align Picture 1 to 0.00 seconds and Picture 2 to the exact final time. Then describe the continuous changes in pose, objects, composition, scene, lighting, and camera that land on the second image.

What is the difference between Subject, Picture, Video, and Audio labels?

Subject names reusable visible content. Picture names a concrete frame or planning anchor. Video names a whole-video editing, continuation, or structural source. Audio names a copied or referenced audio signal with an active role.

How do I stop dialogue from being assigned to the wrong person?

Give each actual vocal source a stable speaker ID such as (S1). Keep the same ID across shots, place identity and delivery outside <d>, and put only the language tag plus exact spoken text inside <d>.

Where should sound effects and music go?

Put synchronized events in the shot timeline. Summarize ambience and physical sounds in overall_soundscape. Put only audience-only background score in non_diegetic_music. Diegetic radio or instruments stay in the timeline.

Should I add multiple shots to every prompt?

No. Add a cut only when it introduces new subject, space, state, viewpoint, or time information. For a small distance or angle change, use camera motion. First and Last Frame tasks often work best as one continuous shot.

Can one reference asset provide more than one role?

Yes. One picture may define multiple Subjects, and one Subject may combine appearance from a picture with motion from a video. A reference video may also provide a whole-video structure and a separately enabled audio role.

Turn the structure into a shot

Open the generator with the format that matches your source material and production intent.