MiniMax H3 Reference to Video is the multimodal mode built for creators who need more control than a text prompt alone can give. Instead of describing every visual detail, movement, camera path, or voice from scratch, you feed the model reference images, reference videos, or reference audio and tell it what each asset should contribute to the final clip.
A portrait can define a character. A short clip can supply motion or camera movement. An audio file can guide voice or sound. Multiple references can work together in a single generation, which is what separates this mode from a standard image-to-video workflow.
MiniMax describes H3 as a general-purpose multimodal video model that reads text, image, video, and audio in one unified context, generating clips up to 15 seconds at 2K resolution with native stereo sound. This guide covers how the mode works, how to choose the right reference materials, how to write stronger prompts, and how to apply it to motion transfer, character consistency, product videos, and branded content.
What Is MiniMax H3 Reference to Video?
This mode creates a new video from a text prompt plus one or more reference assets — images, video clips, audio clips, or any combination of the three. The important word is reference: an image doesn’t have to become the opening frame. It can instead tell the model what a character, product, logo, or visual identity should look like, while a reference video supplies movement, camera motion, or editing rhythm, and a reference audio clip supplies voice or sound.
MiniMax’s documentation treats this as a separate workflow from standard first-frame or first-and-last-frame generation — the two modes can’t be combined in one request. That distinction is what makes reference generation useful whenever a result is easier to show than to describe.
Instead of writing out every detail of a complex movement, you could provide a video containing the motion and say:
Use the character from Image 1. Follow the body movement and performance timing from Video 1. Keep the character’s facial features, outfit, and proportions consistent. Place the scene in a futuristic underground station with cinematic lighting.
The prompt explains the relationship between references rather than trying to recreate every visual detail in words.
What Can You Control With Reference Images, Videos, and Audio?
Reference | What It Guides | Example |
|---|---|---|
Image | Character, product, outfit, visual identity | Keep the same person consistent across a fashion clip |
Video | Motion, performance, camera movement, rhythm | Transfer a dance or a camera orbit |
Audio | Voice or sound characteristics | Guide a speaking or singing performance |
Image + video | Appearance + motion | Put a new character into an existing motion pattern |
Image + video + audio | Appearance + motion + sound | Build a full multimodal performance |
MiniMax demonstrates this idea directly: one video supplies a camera movement, an image supplies the character, and an audio clip supplies the vocal reference — all in a single generation. The useful question isn’t “how many references can I upload,” but “what information should each reference contribute.” That framing is why the mode shouldn’t be treated as just another image-to-video feature; you’re defining what information comes from each input.
Reference to Video vs. Image to Video
Image to video typically starts with an image as the first frame, and the model animates it forward. This works well when you already know exactly how the shot should begin.
Reference generation uses an image as contextual information rather than a fixed starting frame. A full-body photo could define a character’s face, hairstyle, clothing, and proportions, while the generated video begins from a completely different camera angle — giving the model more freedom to construct the new scene.
A simple rule: use image-to-video when the image should be the shot. Use MiniMax H3 Reference to Video when the image should describe something inside the shot.
How to Use MiniMax H3 Reference to Video
Step 1 — Decide what to preserve. Ask one question: what must stay recognizable in the output? It might be a person’s identity, a specific outfit, a product design, a logo, a motion sequence, a camera move, a voice, or several of these together. Don’t add references just because more slots are available — each one should have a job.
Step 2 — Choose the right reference type. For character appearance, use clean images that clearly show the face, hairstyle, and proportions. For motion, choose a clip where the movement is easy to read. For camera reference, pick a clip with an obvious camera path. For products, use images that reveal shape, materials, and branding clearly. For audio, use a clean clip without competing background sound.
Step 3 — Assign each reference a role in the prompt. A useful pattern is: subject + reference relationship + action + environment + camera + style + audio. For example:
Use Image 1 as the character reference. Preserve her face, long dark hair, red coat, and body proportions. Follow the walking motion from Video 1. Place her inside a minimalist gallery with soft directional lighting. The camera slowly tracks backward as she walks toward it.
Step 4 — Start simple. Test with one character, one action, one reference video, and one camera instruction before layering in more. This makes it far easier to tell whether a problem comes from the identity reference, the motion reference, or the prompt itself.
MiniMax H3 Reference to Video Prompt Formula
A reusable template:
Use Image 1 as the main subject reference. Preserve [identity / clothing / product design / colors]. Follow the [movement / timing] shown in Video 1. Reference the [camera movement] from Video 2. Place the subject in [environment]. Use [lighting and style]. The final video should show [main action]. Keep [key elements] consistent throughout.
Vague prompts like “make a cool cinematic video based on these references” leave too much open to interpretation. A better version explicitly assigns roles: Image 1 defines the character. Video 1 defines movement. Video 2 defines camera motion. That kind of labeling dramatically reduces ambiguity.
MiniMax H3 Reference to Video Prompt Examples
Character Motion Transfer
Goal: transfer a performer’s movement onto a different character without reproducing the performer.
Use Image 1 as the character reference. Preserve the futuristic armor design, helmet shape, and body proportions. Follow the body movement, timing, and combat performance from Video 1. Place the character on a platform at night with dramatic cinematic lighting. Keep the armor design consistent throughout.
Useful for game characters, robots, virtual influencers, and stylized avatars — MiniMax specifically highlights motion transfer among H3’s commercial use cases.
Consistent Fashion Character
Goal: preserve identity and wardrobe while generating an entirely new composition, rather than simply animating the original pose.
Use Image 1 as the character and wardrobe reference. Preserve the facial appearance, hairstyle, red outfit, and silhouette. Generate a cinematic fashion shot in a modern gallery. Start with a medium full-body shot, then transition into a side tracking shot. Soft editorial lighting, premium fashion-film aesthetic.
This pattern targets searches like “consistent character,” “reference image to video,” and “how to keep characters consistent.”
Logo and Brand Animation
References don’t need to be human subjects. A logo or icon can anchor an entire motion-design sequence.
Use Image 1 as the exact logo reference. Preserve the silhouette and brand colors. Transform the flat logo into a metallic 3D emblem without changing its shape. Slow cinematic push-in, controlled reflections, clean minimal background.
For brand work, stating what must not change — silhouette, colors, proportions, packaging layout — is often more useful than describing what should. H3’s launch materials specifically call out text and brand rendering as focus areas.
Product Commercial
Combine several product angles with a motion reference for more complete coverage:
Use Images 1–3 as product references. Preserve exact proportions, color scheme, and surface materials. Follow the camera orbit from Video 1. Begin with a close-up of surface texture, pull back to reveal the full product, then finish on a three-quarter hero angle.
Multiple angles give the model more information than asking it to infer hidden geometry from one image alone.
Animated Typography
Graphic and text elements work as references too — useful for title cards and motion graphics:
Use Image 1 as the typography reference. Preserve the exact text. Begin centered on a dark surface as glowing particles drift through the scene. Slowly push the camera in as light builds behind the text.
Shorter, high-impact text is generally easier to keep stable than dense paragraphs or complex layouts.
How to Write Better MiniMax H3 Reference Prompts
Assign every reference a role. Instead of “use these references,” write “Image 1 defines the character, Video 1 defines the movement, Video 2 defines camera motion.”
Separate identity from action. Keep who or what the subject is from the image, but take what it does from the video — for example: “Preserve the robot design from Image 1 while following the running motion from Video 1.”
Separate motion from camera movement. A subject can walk forward while the camera tracks backward, orbits, or pushes in — these aren’t the same instruction, and conflating them is a common source of unpredictable results.
State non-negotiable details up front. “Keep the same face, hairstyle, and body proportions” or “do not change the product silhouette or logo placement” tells the model what to prioritize when instructions compete.
Troubleshooting MiniMax H3 Reference to Video
Even strong references don’t guarantee a perfect result on the first try. The fastest way to fix a MiniMax H3 Reference to Video generation that isn’t working is to identify which type of control is failing.
Problem | Likely Cause | Fix |
|---|---|---|
Character face keeps changing | Weak or conflicting identity reference | Use a clearer image and explicitly list features to preserve |
Motion gets ignored | Prompt text conflicts with the reference | Simplify the action description and clearly assign the motion source |
Camera doesn’t match | Subject motion and camera motion are mixed together | Describe them as two separate instructions |
Product changes shape | Too little visual information | Add clean reference views from more angles |
Output feels too literal | Reference relationship is unclear | State explicitly what to borrow vs. what should be newly generated |
Details disappear | Too many active constraints at once | Reduce the number of references and simplify the request |
The general rule: one reference, one job. Add complexity only after a simple version works.
MiniMax H3 Reference to Video Input Limits
The current API supports up to nine reference images, three reference videos, and three reference audio clips, with 12 mixed files total per request. Reference video and audio clips can run 2–15 seconds each, generated output ranges from 4–15 seconds, and available resolutions are 768P and 2K. A non-empty text prompt is required for every generation. More files doesn’t automatically mean a better result — three well-chosen references usually outperform twelve ambiguous ones.
When to Use MiniMax H3 Reference to Video vs. Alternatives
Reach for MiniMax H3 Reference to Video when the goal sounds like “make this character perform that motion,” “keep this product consistent across a new commercial,” or “combine visual, motion, and audio references.” Use standard first-frame image-to-video instead when the goal is simply to animate one exact image, or first-and-last-frame generation when you need to start on one frame and end on another. The distinction is small in principle but changes which workflow gives you the most control in practice.
Try MiniMax H3 Reference to Video on MiniMax H3.
FAQ
What is MiniMax H3 Reference to Video?
A multimodal workflow that generates a new video from a prompt plus reference images, videos, audio, or a combination of them, guiding character appearance, motion, camera behavior, style, or voice.
Can a reference video be used for motion transfer?
Yes. Motion reference and V2V motion transfer are listed among H3’s core capabilities — a reference video can guide movement while a separate image supplies the target character.
Can images, video, and audio be combined in one generation?
Yes, up to 12 mixed reference files total, subject to the per-type limits above.
How many reference images are supported?
Up to nine per request in the current API.
What’s the difference between this mode and standard image-to-video?
Image-to-video generally uses an image as the first or last frame. Reference generation uses images, video, or audio as contextual guidance for identity, motion, camera style, or voice — the two are separate modes and can’t be combined in one request.
Does it support 2K output?
Yes, alongside 768P, with durations from 4 to 15 seconds.
Are audio references supported?
Yes, alongside image and video references, for guiding voice or sound characteristics.
What’s a good prompt structure to start with?
Subject reference + details to preserve + motion reference + camera reference + environment + style + constraints — the key is stating clearly what each reference should control.
Can it be used for character consistency?
Yes — assign a clear character image as the identity reference and explicitly ask the model to preserve facial appearance, hairstyle, clothing, and proportions across the new generation.
Can it be used for product videos?
Yes — multiple product images can supply shape, color, and branding information, while a separate reference video guides camera motion or presentation style.
Final Thoughts
MiniMax H3 Reference to Video earns its place whenever text alone can’t carry the full creative intent — a character, a camera move, a voice, and a visual style all at once. The most reliable approach stays simple: give each reference one clear job, then write a prompt that explains how the pieces should interact. For character consistency, motion transfer, product ads, and branded animation, that relationship-based approach is what makes this mode meaningfully more capable than a basic image-animation workflow.