Most AI video workflows treat vertical as an afterthought: generate a wide shot, crop the sides, then spend twenty minutes adding music that has nothing to do with what is on screen.
A MiniMax H3 TikTok video can be built the other way around. Here the 9:16 frame is a generation setting rather than a crop, and stereo sound is produced in the same pass as the picture. That changes what you write in the prompt and how many tools sit between an idea and a postable file.
This guide covers what the MiniMax H3 generator on this site exposes for vertical work: which ratios each mode offers, how long a clip can run, and how to write a vertical brief that reads on a phone. Every specification below was checked against the MiniMaxH3.co generator interface, because full model documentation and what a given generator exposes do not always match.
Start your MiniMax H3 TikTok video in 9:16
Can MiniMax H3 create TikTok videos?
Yes. MiniMax H3 produces vertical short-form clips suitable for TikTok-style publishing, solving several production problems inside one generation.
TikTok requirement | How MiniMax H3 handles it |
|---|---|
Vertical canvas | Generate directly at 9:16 |
Short runtime | 4–15 second clips in one-second steps |
Spoken hook | Dialogue directed in the same prompt as the picture |
Environmental sound and music | Ambience, effects, and score generated natively |
Product or creator consistency | Image supplied as a MiniMax H3 reference |
Trend-like movement | Video reference guiding motion and camera rhythm |
The point is not simply selecting 9:16. A strong MiniMax H3 TikTok video is designed around how short-form content is actually watched.
MiniMax H3 TikTok video specs: ratio, duration, resolution
A lot of published information about MiniMax H3 aspect ratios is unreliable, so it is worth being exact about what this generator offers.
How each MiniMax H3 mode reaches 9:16
The three MiniMax H3 modes do not offer the same ratio controls, so the route to a vertical frame differs.
MiniMax H3 mode | Ratio control on this generator | Getting to 9:16 |
|---|---|---|
Text to Video | 16:9, 4:3, 1:1, 3:4, 9:16 | Select 9:16 before generating |
First & Last Frame | Selector defaults to 16:9 | Set 9:16 and upload vertical frames |
Reference | Adaptive, or 21:9 through 9:16 | Select 9:16, or let adaptive follow vertical references |
All three MiniMax H3 workflows can produce vertical video, but they get there differently. Text to Video and Reference let you target 9:16 directly. First & Last Frame is the one to watch: its selector defaults to 16:9, so set the ratio and supply vertical frames before generating. A vertical keyframe paired with a leftover widescreen setting is the most common way to waste a MiniMax H3 generation.
One difference worth knowing if you also cut widescreen versions from the same brief: cinematic 21:9 and the adaptive setting appear only in MiniMax H3 Reference mode on this generator.
MiniMax H3 duration
The MiniMax H3 duration selector runs from 4 to 15 seconds in one-second increments, defaulting to 5 seconds.
Length should follow the action, not the platform. Five seconds covers one gesture, one product reveal, or one spoken line. Longer durations only pay off when something keeps developing — a continuous camera move, a transformation, an exchange of dialogue. Longer TikTok concepts are cleaner built from several focused MiniMax H3 clips assembled in editing.
MiniMax H3 resolution
Two options: 2K and 768P. 2K is the MiniMax H3 default and the better choice when final output detail matters; 768P suits iteration on composition and wording. Vertical platforms re-encode on upload, so 2K is not automatically required for every workflow.
Think vertically before writing the prompt
The most common vertical mistake is writing a landscape brief and switching the ratio to 9:16. MiniMax H3 will produce a tall frame, but the staging fights it. Compare these two instructions.
Weak:
Create a TikTok product video in 9:16.
Better:
Vertical 9:16 creator-style framing. A woman fills the center of the frame from the waist up, holding the skincare bottle beside her face. Keep the bottle large and readable throughout the shot. The camera makes a subtle handheld push-in while she speaks directly to camera.
The second version gives MiniMax H3 a production decision rather than an aspect-ratio label. Four rules translate straight into prompt wording.
Stack depth instead of spreading width
Two important subjects side by side feel cramped in a tall frame. Arrange them through depth instead: put the creator slightly behind the product, let a second character enter from the background, or move a foreground object toward the lens.
Keep important subjects near the vertical center
Captions, usernames, and buttons occupy the top and bottom of the feed. Do not rely on the extreme edges of the generated frame for the only important face, label, or visual detail.
Use movement that fits the frame
A subtle push-in, tilt, rise, drop, or forward track makes better use of a tall composition than a long horizontal pan. Camera direction should support the subject rather than reveal empty side space.
Review the result at phone size
A MiniMax H3 generation can look balanced in a large preview and still feel empty on a phone. Shrink it down first: can a viewer immediately recognize the face, product, or action? If not, the composition needs revision, not more adjectives.
Best MiniMax H3 settings for TikTok
Setting | Practical starting point |
|---|---|
Aspect ratio | 9:16 |
Duration | 6–12 seconds for most single-concept clips |
Resolution | 768P while testing, 2K for the selected concept |
Shot count | 1–3 clear shots |
Opening | Visible action or result immediately |
Dialogue | One short idea, not a paragraph |
Camera | One primary movement per shot |
Audio | Voice plus one supporting layer is often enough |
References | Only references with a defined role |
Final frame | Product, reaction, payoff, or visual CTA |
These are workflow recommendations, not mandatory MiniMax H3 settings. The right duration and shot structure depend on what the clip is trying to accomplish.
The MiniMax H3 TikTok prompt formula
A good MiniMax H3 TikTok video prompt behaves like a compact production brief, not a pile of adjectives:
Format + opening hook + subject + action + vertical composition + camera + dialogue + synchronized sound + music + ending state
Adding words like “viral,” “cinematic,” or “high quality” does far less than naming what happens in the first second.
Five MiniMax H3 TikTok video prompts
1. UGC product recommendation
Vertical 9:16 creator-style TikTok video, 10 seconds. A creator stands in a bright apartment kitchen holding the referenced insulated cup. Start mid-action as she adds ice, closes the lid, lifts the cup toward the camera, and says with amused surprise: “Okay, this actually stayed cold all afternoon.” Natural handheld phone-camera feel with slight human movement, no dramatic cinematic camera motion. Preserve the product color, lid shape, logo placement, and proportions from the reference. Crisp ice sounds, lid click, soft room ambience, and subtle upbeat background music. End on a close product shot in her hand.
Judge this on product fidelity, believable handling, natural delivery, and whether the first second contains visible action.
2. Talking creator hook
Vertical 9:16 talking-head TikTok, 8 seconds. Medium close-up of a creator looking directly into the camera in a simple home office. She raises one finger at the opening and says in a clear, energetic voice: “If your AI videos still feel slow, change the first three seconds.” Slight handheld movement, natural blinking and small facial gestures, soft window light, realistic casual social-video appearance. Quiet room tone only, no background music. Keep the framing centered and consistent until the final frame.
Talking clips do not need a complicated scene — over-directing the background tends to destabilize the performance.
3. Motion reference clip
Create a vertical 9:16 short-form fashion clip using the referenced character for identity and clothing and the reference video for the main body movement and camera rhythm. Keep the character’s face, hairstyle, jacket, and proportions recognizable. Reproduce the confident two-step turn and forward movement while adapting it naturally to the new character. The camera tracks backward, vertically framed at three-quarter length. Street ambience, footsteps synchronized to movement, and a restrained electronic beat. End as the character stops close to camera and looks directly into the lens.
Use only material you have the right to use. A motion reference should guide movement or camera behavior, not reproduce another creator’s protected work or likeness.
4. Cinematic three-second hook
Vertical 9:16 cinematic TikTok hook, 8 seconds. Begin with an extreme macro shot of a red sneaker suspended above rain-covered city pavement at night. At once, the sneaker drops into the puddle in slow motion and water explodes upward around the frame. The camera rapidly pulls back to reveal the wearer standing under neon reflections. Strong physical splash sound synchronized to impact, distant city ambience, then a short bass hit as the full shoe is revealed. Keep the sneaker design stable. End with the shoe filling the lower center of the vertical frame.
The payoff starts immediately, rather than spending three seconds establishing an empty street.
5. Product reveal
Vertical 9:16 premium product reveal, 7 seconds. Start with the referenced perfume bottle almost completely hidden in cold white mist at the center of the frame. The mist parts immediately as the camera makes a slow push-in. Light travels across the glass and reveals the bottle silhouette, cap, and label without changing their design. A soft glass resonance and moving air are synchronized with the reveal. Minimal ambient electronic tone, no dialogue. Finish with a stable front-facing product frame with clean negative space above the bottle.
One clear transformation survives a short vertical runtime better than a sequence of effects.
When a MiniMax H3 TikTok video needs references
Reach for MiniMax H3 Reference mode when something specific has to survive the generation. If nothing must stay recognizable, Text to Video is faster. Give each reference one clear job:
TikTok goal | Useful reference |
|---|---|
Keep the same creator across clips | Face or character image |
Preserve a product’s appearance | Product image |
Reproduce a movement pattern | Video reference |
Guide camera behavior | Video reference |
Guide voice or delivery | Audio reference |
The mechanism worth learning is the reference tag. Each file you supply is addressed in the prompt with @Image, @Video, or @Audio, and that syntax is how you tell MiniMax H3 what a file is for. A face reference and a location reference are not interchangeable, and naming the role removes ambiguity the model would otherwise resolve on its own.
More references do not mean more control. If two images supply conflicting versions of the same object, extra context becomes extra ambiguity.
Learn the complete multimodal workflow in MiniMax H3 Reference to Video
Prompting sound in three layers
Because MiniMax H3 generates audio with the picture, sound is a prompt instruction rather than an editing step. Think in layers so their roles do not collide.
Audio layer | Example |
|---|---|
Foreground event | Spoken hook, lid click, splash, footsteps |
Environment | Room tone, café ambience, traffic, wind |
Music | Light beat, bass hit, restrained electronic rhythm |
Not every clip needs all three. A creator review may need speech and room tone; a movement clip, footsteps and music; a product reveal, one synchronized effect and a minimal score. The more layers a short MiniMax H3 prompt demands, the harder it is to identify which one caused a weak result.
Describe when the sound happens
Instead of:
Add realistic city sounds.
Try:
A bus passes behind her as she finishes the sentence, followed by a brief tire hiss on the wet road.
The second version gives MiniMax H3 both a sound and a position in the clip.
Test the result on a phone speaker
A mix that feels spacious on headphones can turn muddy through a small driver. Dialogue and the primary physical sound should stay legible even on poor playback.
Make the first three seconds stronger
“Viral” is not a generation setting, and no prompt guarantees distribution. What you can design is opening information density.
A weak prompt begins with atmosphere: a stylish woman stands in a beautiful modern apartment with cinematic lighting. Nothing has happened yet.
A stronger one begins with an event: a creator thrusts the product toward the lens before the first spoken word, or the character is already running when the video begins.
The result, transformation, or reaction should arrive before the scene explains itself. This is where MiniMax H3 TikTok video prompting parts ways with cinematic prompting: a film shot may earn a slow reveal, while a short-form shot usually needs the reveal to be the opening.
Choosing a MiniMax H3 mode
Use the simplest mode that protects what matters.
Goal | Better starting point |
|---|---|
Original visual hook | Text to Video |
Animate a still, or control exact opening and ending composition | First & Last Frame |
Keep a specific product | Reference (product image) |
Preserve a character across clips | Reference (face image) |
Recreate movement or camera rhythm | Reference (video) |
Guide voice or delivery | Reference (audio, with image or video) |
On the MiniMaxH3.co generator, First & Last Frame uploads currently accept JPEG, PNG, WebP, BMP, TIFF, and GIF, with a single-file limit under 10MB. Other platforms running MiniMax H3 may apply different input limits.
For TikTok the useful question is not which mode has the most features. It is: which mode preserves the one thing this video cannot afford to lose? If that is the product, use a product reference. If it is a person’s identity, prioritize the character reference. If it is a dance or a camera pattern, use a motion reference. If the idea is original and nothing needs preserving, Text to Video is enough.
Open the MiniMax H3 text to video generator
Common MiniMax H3 vertical problems
The subject looks too small. Repeating “vertical” does not fix it. Describe the composition: medium close-up, waist-up creator, product beside the face, subject filling the center of the tall frame.
The opening feels slow. Remove establishment that does not communicate the idea. Begin with action, speech, transformation, or the result.
The dialogue is too long. An eight-second clip has no room for a paragraph plus several visual events. Shorten the line before adding timing instructions.
The product changes during motion. Use a clean product reference and name the features that must stay fixed. Do not simultaneously demand a dramatic transformation of the same object.
The audio feels crowded. Reduce layers rather than adding more direction. More audio instructions are not always better.
A motion reference overwrites the character. Separate the roles: one asset defines the subject, another defines the movement or camera behavior.
It looks like a polished commercial instead of a TikTok. Replace “luxury commercial, flawless studio lighting, cinematic advertising shot” with handheld creator framing, soft window light, natural pauses, and casual direct-to-camera delivery. Specific production language works better than asking MiniMax H3 to “make it authentic.”
The ratio was left at 16:9. Especially in First & Last Frame mode, where 16:9 is the default on this generator.
Keeping one character across a series
Short-form rewards recurring characters, so identity has to survive separate generations. A face reference carried into each clip and addressed as @Image is the mechanism, with the fixed features named explicitly. What holds and what drifts is covered in the MiniMax H3 consistent character guide.
Adapting a MiniMax H3 TikTok Video for Reels and Shorts
The same vertical-first principles transfer well to Instagram Reels and YouTube Shorts. A 9:16 MiniMax H3 source clip can be adapted across these platforms, although overlays, safe areas, captions, and publishing conventions differ. Preview the final edit inside the destination platform before publishing.
MiniMax H3 TikTok video FAQ
Does MiniMax H3 support 9:16 TikTok videos?
Yes. All three MiniMax H3 modes can produce vertical output. Text to Video and Reference let you select 9:16 directly; First & Last Frame needs the ratio set and vertical frames supplied. It is a generation setting, not a crop applied afterward.
How long can a MiniMax H3 TikTok video be?
The MiniMax H3 duration selector runs from 4 to 15 seconds in one-second steps, defaulting to 5 seconds. Longer concepts are better assembled from several clips in editing.
Why can’t I select 21:9 in Text to Video mode?
On this generator, 21:9 and the adaptive option appear only in MiniMax H3 Reference mode. Text to Video offers 16:9, 4:3, 1:1, 3:4, and 9:16. Ratio controls vary between platforms running MiniMax H3, so check the mode you are actually working in.
What MiniMax H3 resolution should I use for TikTok?
2K is the default and a good choice when final source detail matters. 768P is useful for faster iteration while testing composition and wording.
Can a MiniMax H3 TikTok video include sound?
Yes. Native stereo sound — dialogue, ambience, effects, and music — is generated in the same pass as the picture. Describe it in the same prompt as the visuals.
How many reference files can one MiniMax H3 generation use?
Up to 9 images, 3 video clips, and 3 audio files, capped at 12 files total. An audio reference must be accompanied by at least one image or video reference.
How do I tell MiniMax H3 what each reference is for?
Address them in the prompt using @Image, @Video, and @Audio, and name the role of each file — identity, location, camera movement, voice.
Does MiniMax H3 guarantee viral TikTok videos?
No. MiniMax H3 generates short-form creative assets, but distribution depends on concept, audience, timing, account history, editing, captioning, and platform response. Treat “viral” as a creative goal, not a model feature.
Start creating
Vertical output, native sound, and a duration that fits a short-form hook — from one MiniMax H3 generation.