Base format: T2VA
Text to Video
Start from a written audiovisual idea with no frame anchor.
Begin with the three core fields. No image-alignment instruction is required.
User Guide
Choose the right prompt format, then direct shots, dialogue, references, ambience, and music without mixing their roles.
The uploaded asset does not decide the task by itself. Decide what role the asset plays in the target video, then use the matching structure.
Base format: T2VA
Start from a written audiovisual idea with no frame anchor.
Begin with the three core fields. No image-alignment instruction is required.
Base format: I2VA
Use one image as the exact first frame, then develop forward from it.
Place the first-frame alignment instruction before the three core fields.
Base format: FL2VA
Use two images as exact opening and ending frames.
Describe the observable motion path between the frames. A single continuous shot is usually the clearest fit.
Base format: L2VA
Use one image as the exact final frame while allowing the model to infer a compatible opening state.
Align the image to the exact end time and make the action, object state, framing, and lighting progressively converge on it.
Reference format: six sections
Reuse or transform subjects, scenes, motion, editing, audio, or other attributes from reference assets.
Define each reference role first, then track its use and retention through the target timeline.
T2VA, I2VA, FL2VA, and L2VA share one audiovisual body. Image-based tasks add a precise alignment line before that body.
The three core fields
No alignment line
Start directly with integrated_multimodal_description and build the complete audiovisual timeline from text.
For the target video, at 0.00 seconds, <Picture 1> from [Shot 1] is fully referenced.
Anchor style, identity, clothing, composition, and space to the image, then describe continuous development.
Picture 1 from [Shot 1] aligns to 0.00 seconds; Picture 2 from the actual final [Shot N] aligns to the exact S.SS-second end time.
Write the changing pose, object state, composition, lighting, and camera path that progressively lands on Picture 2.
<Picture 1> from the actual final [Shot N] aligns to the exact S.SS-second end time.
Infer a plausible earlier state, then make every change converge on the final image.
The main field is not a plot summary. Every sentence should describe something visible or audible at a specific point in the clip.
[Shot 2] At 00:03.500, the shot cuts to a close-up of the biscuit filling as the two halves separate and crumbs scatter outward.
A useful camera instruction names the motion type, then adds amplitude and speed only when they affect the result.
| Intent | Vocabulary | What changes |
|---|---|---|
| Lens change | Zoom In / Zoom Out | The camera stays in place while focal length changes. |
| Depth travel | Push In / Pull Out | The camera body moves toward or away from the subject. |
| Horizontal travel | Truck Left / Truck Right | The whole camera moves sideways without merely rotating. |
| Vertical travel | Pedestal Up / Pedestal Down | The whole camera rises or lowers through space. |
| Fixed rotation | Pan Left / Pan Right / Tilt Up / Tilt Down | The camera body stays in place while the lens rotates horizontally or vertically. |
| Around the subject | Arc Shot | The camera travels on a curved path around the subject. |
| Subject follow | Tracking Shot | The camera moves with a moving subject and maintains the relationship. |
| Subject viewpoint | POV | The framing adopts what a subject sees. |
| Axis rotation | Roll Clockwise / Roll Counterclockwise | The camera rotates around the lens axis. |
| Instability | Shake Slightly / Shake Strongly | Controlled camera shake changes the stability of the frame. |
| No movement | Static Shot | The position and lens remain still. |
with small amplitude
Use for a restrained framing change that keeps the original composition readable.
with large amplitude
Use when the camera must reveal substantially more space or transform the composition.
at slow speed
Use for a gradual move with enough time to inspect a subject or product detail.
at fast speed
Use for a rapid reveal or energetic transition that is still physically legible.
The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.
The baker with a calm, slightly raspy voice (S1) says: <d>[English] First batch of the morning.</d>
Put text that must appear on a sign, label, subtitle, banner, interface, or product frame inside English double quotation marks. Keep the original language, words, capitalization, and punctuation. Do not translate it or paraphrase it.
A red neon sign above the door reads "OPEN" and remains unobstructed.
For I2VA, FL2VA, or L2VA, add the exact image-alignment instruction as the first line. T2VA starts at the first field.
integrated_multimodal_description: [Shot 1] [Style and initial composition]. [Subject identity, position, action, and camera motion]. [Speaker (S1) and <d>[Language] exact dialogue</d> when needed]. [Shot 2] At 00:SS.mmm, the camera cuts to [new information].
overall_soundscape: [Ambience, physical action sounds, and non-verbal human sounds across the clip.]
non_diegetic_music: [Instrumentation, tempo, rhythm, and dynamic change, or N/A.]Full-Reference mode is not a longer Base prompt. It starts by defining what each asset contributes, then records how that contribution is used in the target video.
Define every reusable subject, concrete frame anchor, whole-video source, and audio source.
State the task type and the main reference relationships in one short paragraph.
Record how each defined visual or audio reference is preserved, transferred, copied, or loosely followed.
Write the target video shot by shot in playback order, inserting reference labels where their roles take effect.
Summarize ambience and physical sounds, including any applicable audio-copy or audio-reference relationship.
Describe the audience-only score and its relationship to referenced audio, or write N/A.
This is the main body of a Full-Reference prompt. It must explain the target video in playback order and show exactly where each reference becomes visible or audible.
For a typical generation task, 350 to 500 English words is a useful detail range, not a quota. Dialogue-heavy clips should prioritize the complete spoken timeline. Editing prompts may be shorter or longer according to the source video complexity, and a single shot can still require substantial detail.
Complete the sections in this order. Remove unused labels, but do not remove required sections.
subject_definitions:
<Subject 1> is [reusable visible content] from <Picture 1>, preserving [identity-defining features].
<Video 1> is [editing, continuation, or structural role] when applicable.
<Audio 1> is [copy or reference role] for <Subject 1> (S1) when applicable.
summary:
[task type + task type] [Target video and main reference relationships.]
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - [features retained].
<Audio 1>: reference - [audio characteristic followed without copying the signal].
detailed_description:
[Overall visual style in one or two sentences.]
[Shot 1] [Composition, references, subject action, camera, synchronized sound, and dialogue.]
[Shot 2] At 00:SS.mmm, [cut and new information].
overall_soundscape: [Ambience, physical sounds, and audio relationship.]
non_diegetic_music: [Score and audio relationship, or N/A.]This case combines one image, one short motion clip, and one audio file. Study how the prompt gives each asset a narrow job, then compare the final output with the three references.
A short cinematic cafe scene. The woman and street-cafe composition come from the image, the close-up body rhythm and tense reaction come from the video, and the audio file guides vocal pacing and sound texture.
The prompt preserves the woman, clothing, cup, street cafe, warm daylight, shallow depth of field, and background customers from the image. These details are not left for the video reference to decide.
The beach setting and original people in the source video are excluded. The useful reference is the compact close-up framing, slight backward tension, forward recovery, and uneasy facial rhythm.
The prompt does not pretend the audio is a visual subject. It marks the file as an audio reference for cadence, tonal texture, and timing, then keeps ambient cafe sound in overall_soundscape.
Success can be checked frame by frame: the woman stays recognizable, the cafe remains the setting, the reaction follows the motion clip, and the spoken or vocal rhythm follows the audio without adding unrelated plot.
Build the prompt in passes. Each pass resolves one production decision and prevents contradictions from accumulating.
Decide whether each asset is a frame anchor, reusable subject, source video, structural reference, copied audio, or audio reference.
Choose the clip duration, shot count, cut times, and the observable state at the beginning and end of every shot.
Describe what moves, what stays fixed, and how pose, object state, composition, lighting, and camera evolve.
Assign stable speaker IDs, preserve exact dialogue, and separate synchronized sound, ambience, and audience-only music.
Check that every label has one stable meaning and that every claimed preservation or copy relationship is visible in the timeline.
Delete camera moves, cuts, identity changes, or audio instructions that compete with the chosen frame and reference constraints.
Run this audit after writing. A structurally complete prompt can still fail when timing, labels, dialogue, or sound roles contradict one another.
Open a failure that matches your result. Each diagnosis explains what you will see, why the prompt causes it, and the exact rewrite pattern to use.
Typical output
The output ignores a frame anchor, treats a reference as loose inspiration, or follows the scene but loses the intended identity and timing.
Why it happens
A Base task describes one audiovisual timeline. Full-Reference must first define what every asset contributes and how strongly that contribution is retained. One generic paragraph cannot express both jobs reliably.
How to fix it
Use the Base three-field format for T2VA, I2VA, FL2VA, and L2VA. Use the six ordered sections for Full-Reference. Decide the task before writing any visual detail.
Rewrite pattern
Full-Reference: define <Subject 1> from @Image1, state [reference generation], record fully_preserved, then place <Subject 1> in the required shots.
Typical output
The model copies the wrong property from a file, confuses a person with the entire source video, or tries to reproduce a reference frame that was only meant to define identity.
Why it happens
A file is a source container, while a Subject is reusable visible content extracted from that source. Picture, Video, and Audio labels describe different relationships and cannot be exchanged freely.
How to fix it
Use Subject for a person, object, scene, style, action, or effect. Use Picture only for a concrete frame anchor, Video for editing, continuation, or whole-video structure, and Audio only when sound is copied or referenced.
Rewrite pattern
<Subject 1> is the woman from @Image1. <Video 1> is @Video1 and provides only the target cut and pacing structure.
Typical output
The clip holds on the opening state, jumps abruptly near the end, or reaches the last image through an implausible morph.
Why it happens
FL2VA already knows the two endpoints. Repeating their appearance does not tell the model how pose, object state, composition, light, and camera should change between them.
How to fix it
Describe a continuous chain of observable intermediate changes. Keep the differences narrowing until the final pose, object arrangement, lighting, and framing match Picture 2.
Rewrite pattern
She releases the handle, raises the umbrella, opens the canopy, rotates it into the final angle, and settles into the spacing shown in Picture 2.
Typical output
The video feels fragmented, invents new backgrounds, changes subject details, or loses continuity even though the intended action is simple.
Why it happens
Every cut invites the model to establish a new subject state, viewpoint, space, or time. A closer view alone does not require that reset.
How to fix it
Use push, pull, pan, tilt, truck, arc, tracking, or zoom for a distance or angle change. Add a cut only when the next shot contributes genuinely new information.
Rewrite pattern
Keep [Shot 1] and write: The camera pushes in with small amplitude at slow speed toward the filling, instead of creating [Shot 2].
Typical output
A line is spoken off-screen, assigned to another face, delivered by both characters, or accompanied by incorrect lip movement.
Why it happens
A character description alone does not create a stable vocal identity across shots. The dialogue tag also cannot carry speaker identity or delivery instructions.
How to fix it
Assign each actual vocal source one stable (Sx) ID in vocal-event order. Put identity, voice, action, and delivery outside <d>. Keep only the language and exact words inside <d>.
Rewrite pattern
The young woman with a quiet voice (S1) turns to camera and says: <d>[English] I get off at the next station.</d>
Typical output
Background music behaves like an object in the scene, dialogue is repeated, sound effects arrive at the wrong moment, or the clip adds music when silence was intended.
Why it happens
MiniMax H3 separates synchronized shot events, full-clip ambience, and audience-only score. Repeating one sound across multiple fields creates competing instructions.
How to fix it
Place dialogue, a door slam, a radio cue, or an on-screen instrument in the shot timeline. Summarize continuing ambience in overall_soundscape. Put only audience-only score in non_diegetic_music.
Rewrite pattern
overall_soundscape: Rain and train-wheel rhythm continue. non_diegetic_music: N/A.
Typical output
The output is broadly similar to the reference but changes the face, clothing, product construction, color, or movement that actually matters.
Why it happens
Words such as keep consistent do not identify which features define success. The model also needs to know where each reference becomes active in the timeline.
How to fix it
Name the identity-defining or product-defining features in subject_definitions, choose the correct retention marker, and cite every shot where the reference appears or transfers its attributes.
Rewrite pattern
<Subject 1>: fully_preserved - retain the biscuit shape, baked surface, pink filling, and clean product appearance in [Shot 1] through [Shot 6].
Typical output
The clip captures the general topic but invents staging, skips important actions, produces weak camera direction, or fails to land on the intended ending.
Why it happens
A plot summary explains meaning but not production. The model can only render concrete visual and audible events along a timeline.
How to fix it
For every shot, specify composition, subject position, visible action, state change, camera behavior, synchronized sound, and the condition that ends the shot.
Rewrite pattern
Replace "the product becomes exciting" with "the biscuit snaps at center frame; filling expands; crumbs scatter; the camera holds close until the fragments settle."
Answers follow the official Base and Full-Reference prompt-writing guides.
No. T2VA, I2VA, FL2VA, and L2VA use the three Base fields, with an image-alignment instruction added for image-anchored tasks. Full-Reference uses six ordered sections beginning with subject_definitions.
Write the visible and audible timeline: style, composition, subject identity and position, actions, shot changes, camera motion, speakers, exact dialogue, singing, and synchronized in-scene sound.
Align Picture 1 to 0.00 seconds and Picture 2 to the exact final time. Then describe the continuous changes in pose, objects, composition, scene, lighting, and camera that land on the second image.
Subject names reusable visible content. Picture names a concrete frame or planning anchor. Video names a whole-video editing, continuation, or structural source. Audio names a copied or referenced audio signal with an active role.
Give each actual vocal source a stable speaker ID such as (S1). Keep the same ID across shots, place identity and delivery outside <d>, and put only the language tag plus exact spoken text inside <d>.
Put synchronized events in the shot timeline. Summarize ambience and physical sounds in overall_soundscape. Put only audience-only background score in non_diegetic_music. Diegetic radio or instruments stay in the timeline.
No. Add a cut only when it introduces new subject, space, state, viewpoint, or time information. For a small distance or angle change, use camera motion. First and Last Frame tasks often work best as one continuous shot.
Yes. One picture may define multiple Subjects, and one Subject may combine appearance from a picture with motion from a video. A reference video may also provide a whole-video structure and a separately enabled audio role.
Open the generator with the format that matches your source material and production intent.