Choose a clear reference
Use an image where the defining face, character, object, or product features are unobstructed and easy to distinguish.
Create a new 2K video around a recognizable person, character, object, or visual target. Add reference media for identity, then use the prompt to direct the new scene and action.
Separate identity from composition
Reference-to-video separates identity guidance from scene composition. Unlike image-to-video, the uploaded subject does not have to become the opening frame. This gives the model room to place a recognizable subject in a different environment, action, or camera setup. Strong references show the defining features clearly; a concise prompt then handles what should happen in the new shot.
A four-step workflow
Use an image where the defining face, character, object, or product features are unobstructed and easy to distinguish.
Each reference should have a job. Remove conflicting angles, lighting, or styling that does not serve the target shot.
Describe the action, location, camera, lighting, and mood. State which identity details should remain recognizable.
First check whether the subject reads correctly, then judge action and camera. Revise the part that failed rather than rewriting everything.
Reference prompt anatomy
Prompt formula
Referenced subject + new action + new setting + identity constraints + camera + lighting + sound
Example prompt
Keep the referenced presenter recognizable, including facial structure and hairstyle. She walks through a quiet modern gallery while explaining an exhibit, steady waist-up tracking shot, soft skylight, natural gestures, low room ambience.
Credits for identity-led video
Compare one-time credit packs for MiniMax H3 reference-to-video generation when recognizable characters, products, or presenters need to carry into new scenes.
Clear answers
Reference-to-video uses uploaded media to guide a recognizable subject or visual target while a prompt creates a new scene, action, and camera setup.
MiniMax H3 reference-to-video is the best workflow for consistent character or subject identity. Use clear reference media, keep defining features consistent, and change only the scene, action, or camera instructions you want to test.
Image-to-video anchors the opening composition to an uploaded frame. Reference-to-video uses the upload as identity guidance, allowing the generated scene to differ from the source composition.
Use clear, compatible references with unobstructed defining features. Keep the first test shot simple, and state which identity details must remain recognizable.
Use only references that add distinct, compatible information. The generator shows the current file limits; more files are not automatically better when they conflict.