MiniMax H3 LogoMiniMax H3

Conclusion

MiniMax-H3 looks noticeably stronger than the earlier MiniMax video workflow I used, especially for atmosphere, lighting, layered environments, and cinematic framing. It is productive for storyboards, mood reels, concept pitches, background plates, and rough cuts. It is not yet reliable for unattended final delivery: speech, lip sync, sports rules, hand topology, object trajectories, text, and product structure still require review and repair.

The evaluation covers MiniMax-H3 on one RTX 4090 using Hugging Face MiniMaxAI/MiniMax-H3 weights and an FP8 workflow. My verdict is deliberately two-sided. H3 looks stronger than the earlier MiniMax video workflow I tested: scenes have richer lighting, deeper atmosphere, more convincing materials, and a stronger sense of a complete cinematic shot. It can also produce long clips with an audio stream on a single consumer GPU. At the same time, it is not yet a dependable “type a sentence and publish” production system. Its most reliable use is supervised visual ideation, storyboarding, pitch material, background plates, mood reels, and rough cuts. Its most dangerous areas are spoken language, lip synchronization, sports rules, hand topology, object identity, text rendering, and physical continuity.

This review separates technical success from creative usability. The formal outputs were valid MP4 files with readable video and audio streams, but that does not mean the content was correct. Several clips looked as though a person was speaking while the audio became repeated syllables, wrong-language fragments, subtitle-like material, or unrelated short phrases. Sports clips often looked energetic but violated the rules. Cards, hands, balls, lettering, and product structure drifted in ways that matter if the clip is instructional or factual. I therefore recommend H3 as a supervised draft generator rather than an unattended final-video generator.

The evidence combines generated outputs, media inspection, GPU telemetry, and local Whisper Large-v3 transcription. Whisper text is an intelligibility signal, not a perfect word-error-rate benchmark. When it reports long repetitions, unrelated languages, obvious nonsense, or music as words, I treat the audio as requiring human review.

Test Summary

SuiteClipsResolutionDurationPassVRAM
Broad motion20608×35220.75 sec20/2045,130–46,250 MiB
Long coverage20608×35220.75 sec20/2046,250 MiB
Resolution4640×384–960×54420.75 sec4/443,306–48,490 MiB
Duration3608×3525.17–20.75 sec3/341,546–46,250 MiB
Queue3608×3525.17 sec3/341,546 MiB
Control/audio4608×3525.17 sec4/441,546 MiB

Evaluation coverage and production meaning

The evaluation covers landscapes, city scenes, Chinese settings, product text, calligraphy, portraits, hands, animals, pets, sports, camera moves, ambient audio, and miniature style. This coverage demonstrates that H3 can produce a broad visual range, but it was biased toward wide shots, slow movement, and camera travel. Those clips are useful for showing atmosphere, yet they do not answer the harder production questions: Can the model make a person say a complete sentence? Can it preserve two people taking turns? Can it maintain a ball, a hand, a card, a product, or a track across a sequence of fast actions?

The production-focused coverage includes a talking close-up, profile speech, two-person dialogue, a group meeting, singing, fast card handling, cooking with several hand actions, boxing, basketball, soccer, skateboarding, cycling, dog actions, horse racing, foreground occlusion, whip pans, a chase, rapid montage, camera assembly, and emotional speech. Coverage also includes resolution, duration, queue, and repeatability/audio comparisons. The result is a more honest picture of H3 as a tool rather than a gallery of attractive frames.

The visual impression remains encouraging. H3 can turn a short idea into a discussable shot quickly. A rainy street, a mountain valley, a bamboo temple, a flower market, a sailboat, or a cardboard space station can communicate tone and direction even when small details drift. That speed is real productivity: a director can compare moods, a designer can test a concept, and a product team can show a rough visual before investing in a shoot. The key is to ask H3 for a visual hypothesis, not a legally or physically exact record.

Speech and audio: the largest production risk

The talking close-up is the clearest warning. The prompt asked for a young woman to say a short Mandarin sentence. The first words were recognizable, but the rest turned into a long run of repeated and semantically disconnected Chinese syllables. This was not a subtle accent issue; it was a failure of linguistic continuity. The profile test included the requested English sentence, but malformed words surrounded it and the sentence was repeated. The two-person dialogue produced Japanese-like fragments that did not form a reliable exchange. The group meeting produced disconnected Chinese fragments rather than meaningful turn-taking.

The audio sweep found more examples. The singer repeated “yeah” for much of the clip. Some scenes were transcribed as “Thank you,” “you,” or “The” even when the visual prompt was a boxing match, horse race, basketball action, or landscape. Other tracks looked like music, subtitles, or closing credits to the ASR system. This does not prove that every syllable is objectively wrong, because ASR can hallucinate text from noise. It does prove that the soundtracks cannot be accepted without listening and script comparison. A production pipeline should treat the existence of an AAC track as a container fact, not a quality guarantee.

The Chinese control clip was a useful counterexample: its intended sentence was transcribed correctly. The emotional speech clip was also substantially understandable, although the main sentence was repeated. These successes show that H3 can sometimes produce intelligible speech. They do not cancel the broader pattern. When a user needs exact dialogue, I would generate the visual with H3 and replace the native audio using a dedicated TTS system or a human recording. If native audio is kept, an automatic transcription check should compare the result with the intended script, and low-confidence or high-repetition clips should be rejected.

Sports, physics, and temporal continuity

Sports scenes are an important production test because viewers know the rules. The boxing clip showed energetic movement, but I observed moments of audio-video mismatch. The basketball clip contained action and a ball, yet the ball path and body mechanics did not consistently follow a real game. The soccer clip had a similar problem: the sequence looked like an attack in broad terms, but the actions did not remain coordinated. The queue fast-action clip and all four resolution probes used a running or athletic scene; all technically succeeded, but the track, field logic, and competition rules were not reliable. More pixels did not fix the reasoning problem.

This matters for more than sports. A product demo, repair tutorial, cooking lesson, laboratory procedure, or safety video depends on objects staying where they should be. The fast-card test produced malformed cards and unstable relationships between cards and hands. The camera assembly ended with a frame that conflicted with the camera’s structure. The hand-focused cooking clip produced a third hand. The dog-and-ball clip lost a believable ball trajectory, and the basketball coverage clip also lost motion continuity. These are not defects that a viewer can always see in the first second, but they become obvious when paused or used as instruction.

H3 is safer when the important information is composition, light, mood, or broad movement. It is riskier when the viewer must count objects, read text, verify a rule, follow a procedure, or trust a precise trajectory. A practical operator should define the “must be true” facts in a shot before generation. If the must-be-true facts include the exact spoken sentence, the number of hands, a scoreboard, a product assembly order, or a ball landing at a specific point, H3 should be treated as a draft source and the final shot should be repaired or replaced.

Text, signs, and product detail

Text rendering is another independent failure mode. The product-text prompt asked for readable English, while the calligraphy prompt asked for Chinese characters. The calligraphy video suddenly introduced extra writing, and the whip-pan street scene showed unreliable signage and text. In the product assembly clip, the final camera structure conflicted with the preceding assembly. These results suggest that text and exact product geometry should be composited in post rather than delegated to the generative model.

For a creator, this is not necessarily a deal breaker. It simply changes the workflow. Generate the camera movement, light, hands, and background with H3; add the logo, label, subtitle, UI panel, scoreboard, or legal notice later in an editor. This is usually faster than generating dozens of candidates while hoping that every character remains stable. It also protects the team from publishing a clip where a fabricated word looks like an official label.

Resolution, duration, and runtime cost

The baseline workflow used 608×352, 24 fps, 20 diffusion steps, the res_multistep sampler, the simple scheduler, and 484 frames for a 20.75-second clip. The FP8 diffusion weight and NVFP4/AWQ text encoder allowed the workflow to run on one RTX 4090 with low-VRAM offloading. Long clips peaked at about 46,250 MiB, leaving limited room for another large GPU workload. The safe operating rule is one H3 job at a time, with other requests waiting in a queue.

The duration probes were useful for setting expectations. A five-second request produced 5.167 seconds and took 55.55 seconds. A ten-second request produced 10.834 seconds and took 136.24 seconds. A twenty-second request produced 20.75 seconds and took 353.03 seconds. The relationship is not perfectly linear, so a UI should expose tested presets instead of letting users combine arbitrary length and resolution values without a warning.

The resolution probes all passed technically. 640×384 took 90.88 seconds and peaked at 43,306 MiB. 720×416 took 610.15 seconds and peaked at 48,394 MiB. 768×432 took 731.05 seconds and peaked at 47,466 MiB. 960×544 took 1,451.97 seconds and peaked at 48,490 MiB. The largest setting is therefore an experimental hero-shot mode, not a normal iteration mode. The fact that a larger output fits does not mean it is economical to use repeatedly.

All formal clips were readable as H.264 video with AAC audio, and the long clips were 24 fps with 32 kHz stereo audio. That makes the files convenient for preview and editing. It does not solve semantic defects. A player being able to open the file is the first gate, not the last gate.

Stability, queue behavior, and reproducibility

The three-job queue result was They all completed in submission order, each passed the media checks, and the peak VRAM stayed around 41,546 MiB. There was no sign of VRAM growth, task loss, or out-of-memory failure in that run. This supports a low-concurrency service design: accept requests, serialize them, show a queue position, and make the user wait for a verified result. It does not support running multiple H3 samplers concurrently on the same 4090.

The same-prompt control test used the same seed twice and a different seed once. The same-seed video-frame and audio-frame hashes matched. The second request completed in only 3.13 seconds because ComfyUI reused cached computation, so it was not a fresh inference timing sample. For production, the cache is valuable because repeated previews can be cheap. For benchmarking or regression testing, the cache must be cleared or bypassed. The different-seed case completed as a normal generation and produced a different file.

What I would build around H3

I would put H3 inside a supervised four-stage workflow. Stage one is ideation: generate several short candidates and select composition, lighting, mood, and camera language. Stage two is previsualization: use the selected candidate to discuss timing, blocking, and editorial intent. Stage three is repair: replace native audio when the words matter, add all important text in post, and correct or cut around bad hands, objects, or trajectories. Stage four is delivery validation: check the transcript, listen to the audio, inspect the first, middle, and last third of the clip, and verify that the output meets the intended use.

The stable production preset should be 608×352 at five to ten seconds, one job at a time, with native audio treated as optional. A clarity preset can use 640×384 when extra detail is worth a moderate increase in wait time. 720×416 and 768×432 should be reserved for selected hero shots. 960×544 passed this test but took about 24 minutes for one 20.75-second clip and ran near the card’s memory ceiling. It should not be the default for iterative work.

I would not use the current workflow as an unattended final-video service for sports explanations, product assembly instructions, hand-action tutorials, legal or financial narration, customer-facing dialogue, or any clip where every word must match a script. A dedicated TTS or voice-recording system can solve part of the audio problem, but it will not repair extra hands, incorrect ball paths, unstable text, or product topology. Those require visual review, editing, or regeneration.

Final verdict

MiniMax-H3 is a visually capable local video model with a meaningful productivity advantage for supervised concept generation. In my tests it produced more layered and cinematic visuals than the earlier MiniMax video workflow, and it ran on one RTX 4090 with quantized components. The strongest results are atmospheric environments, broad camera movement, visual mood, and short concept shots.

The qualification is essential: H3 is more reliable at making a shot look finished than at making every fact inside the shot correct. Speech can become gibberish, sports rules can be broken, hands and cards can change topology, text can drift, audio can be unrelated, and sound can fall out of sync. I would recommend H3 for storyboards, mood reels, pitch visuals, background plates, transitions, and rough cuts. I would not recommend direct publication without a review-and-repair stage.

The deployment questions that matter in practice

Before adopting MiniMax-H3, a practical operator needs more than a visual demo. The important questions are whether it runs inside ComfyUI, whether a quantized or GGUF-style route is available, what local deployment looks like on 5090, 4090, or 3060 hardware, how open weights affect usage, and how regional or license conditions should be checked. Public evaluation pages such as Artificial Analysis can provide comparison context, but the model should ultimately be judged as a complete local toolchain: weights, quantization, workflow, memory behavior, output quality, audio behavior, and the operator’s ability to validate the result.

ComfyUI deployment is part of the product experience

For a local video model, installation is not a footnote. The user encounters the model through a workflow graph, node versions, model directories, encoders, decoders, attention backends, and memory-management settings. A model can be theoretically compatible with a GPU and still be practically unusable if the workflow expects a different ComfyUI revision or if a custom node silently changes an input. The tested ComfyUI route made the model accessible as a repeatable visual tool, but it also made queueing and cache behavior visible. A production user therefore needs a pinned environment, a known workflow template, a clear model inventory, and an output validation step.

The low-VRAM configuration was important on the RTX 4090. H3 did not behave like a small image checkpoint that can be launched casually beside several other GPU services. Long clips reached roughly 46 GiB of peak allocation, while the largest resolution probe approached the card’s memory ceiling. CPU offloading and quantized components made execution possible, but they did not make memory free. The practical consequence is simple: reserve the GPU for one H3 job, keep the queue explicit, and do not interpret a successful launch as proof that concurrent generation is safe.

Quantization, GGUF discussion, and what was actually measured

Quantization is central to local H3 deployment because it changes the difference between “interesting repository” and “model that a creator can run.” The measured workflow used FP8 diffusion weights together with an NVFP4/AWQ text-encoder variant. This is a tested low-precision path on one RTX 4090. It is not the same claim as saying that every community GGUF package, every text-encoder conversion, or every 30-series configuration will work. GGUF can be useful for local deployment, but a file format alone does not determine whether a video workflow has the right kernels, tensor layout, precision behavior, or memory schedule.

The safe way to evaluate a quantized variant is to separate four gates. First, the files must load without missing tensors or incompatible metadata. Second, the graph must complete without out-of-memory failure. Third, the resulting container must contain valid video and audio streams. Fourth, the content must be judged for language, motion, object identity, and continuity. H3 passed the first three gates on the tested path, while the fourth gate remained mixed. That distinction prevents a common mistake in local-model discussions: treating “it runs” as equivalent to “it is production ready.”

The 4090 results also provide a useful baseline for interpreting reports from other cards. A 5090 may offer a more comfortable memory and speed envelope, while a 3060 may require a different compromise involving shorter clips, smaller frames, stronger offloading, or a community conversion. Those are reasonable hypotheses, not measurements in this report. The only hardware result claimed here is the observed RTX 4090 path. Anyone comparing a GGUF build or another GPU should repeat the same four gates and record peak memory, elapsed time, actual duration, and media validity.

What H3 adds to a creator’s workflow

The strongest productivity benefit is the speed at which an ambiguous idea becomes a discussable visual. A director can test whether a rainy street should feel noir or documentary. A designer can see whether a product launch needs a macro shot, a tracking move, or a wider environment. A game team can explore a mood reel before building assets. A marketing team can compare several visual directions before paying for a shoot. In these cases, exact text, exact physics, and exact dialogue are not the first decision. Composition, lighting, atmosphere, material response, camera language, and pacing are the decision.

H3 is especially useful for storyboards and previsualization because a slightly wrong object can still communicate editorial intent. The model can suggest a sequence, a transition, a background plate, or a rough camera move. It can help a team reject weak ideas early. It can also make a pitch more concrete: stakeholders react more effectively to a rough moving image than to a paragraph describing one. The value is not that every generated frame becomes a final shot. The value is that the cost of exploring visual alternatives falls.

There is a second, less obvious benefit: local operation gives the team control over iteration and retention. A creator can keep prompts, seeds, workflow files, generated clips, and review notes in one environment. A studio can decide which material may leave its network. A product team can build a queue and a review dashboard around a known service. These advantages matter when the work contains unreleased products, internal locations, customer material, or confidential creative direction. Local control does not remove licensing obligations or quality risks, but it changes the operational boundary.

Where H3 should not be trusted without repair

The evaluation shows a consistent boundary between visual suggestion and factual demonstration. The model can make a sports scene feel fast, but it may not preserve the rules of basketball, soccer, running, or boxing. It can show a hand interacting with a card or a product, but the hand count, card shape, and object relationship may change. It can make a person appear to speak, but the soundtrack may contain repeated syllables, unrelated words, or a sentence that does not survive the full clip. It can place lettering in a scene, but the characters may drift or extra writing may appear.

These failures are dangerous when the viewer is expected to learn something. A product assembly video must preserve the order and geometry of each component. A safety video must not invent a physically impossible action. A sports explainer must not show a rule-breaking move as a correct example. A customer-facing spokesperson must deliver the intended sentence, not merely move their lips. A legal, financial, medical, or educational clip has an even higher burden because an attractive but incorrect video can create false confidence.

The repair strategy should be chosen by failure type. Native audio can be replaced with a dedicated TTS track or human recording. Logos, labels, subtitles, scoreboards, and legal text should be composited after generation. Bad hands or objects can be hidden by reframing, cutting, or replacing a short segment. A physically important trajectory should be regenerated or built with a more controllable method. No single post-processing step repairs all of these defects. The important production habit is to identify the must-be-true facts before generation and assign each fact to the right tool.

Audio deserves its own acceptance gate

Video teams often check whether an MP4 opens and stop there. That is insufficient for H3 because a valid AAC track can still contain unusable speech. The evaluation found that visual speech and linguistic speech can diverge. A face can move naturally while the audio becomes repetition, a wrong-language fragment, a music-like pattern, or a short phrase unrelated to the scene. Conversely, the Chinese control case shows that intelligible output is possible. The correct conclusion is not “audio never works”; it is “audio must be accepted clip by clip.”

An appropriate production gate has three parts. A human listens to the clip at normal speed and checks whether the voice sounds intentional. A transcription or script comparison checks whether the words match the intended language and sentence. A visual review checks mouth motion, timing, speaker identity, and any audio-video offset. If the application does not need native speech, disabling or replacing the soundtrack may be safer than asking the model to solve both visual generation and exact narration at once.

Interpreting speed, resolution, and cost together

The runtime numbers reveal why a creator should use presets rather than arbitrary combinations. A 608×352 clip at roughly twenty seconds was practical enough for evaluation, while 960×544 reached about twenty-four minutes for a single clip. The higher-resolution result was technically successful, but it was not an economical default for creative iteration. A larger frame may improve detail in a selected hero shot, yet it also increases waiting time, memory pressure, and the cost of discovering a semantic defect late in the process.

The better workflow is hierarchical. Use the smaller stable preset to decide whether the idea works. Use the short-duration preset to explore alternate prompts and seeds. Promote only the selected concept to a higher resolution or longer duration. Then perform the same media and content checks again, because a successful low-resolution preview does not guarantee that a larger generation will preserve hands, text, trajectories, or speech. This approach treats GPU time as a design resource rather than an unlimited background service.

How to run MiniMax-H3 in ComfyUI

The following is the practical deployment path used for the local RTX 4090 service. It assumes a clean Linux host, an NVIDIA driver and CUDA runtime that already work with PyTorch, and a separate Python environment for ComfyUI. Do not install into the system Python if the machine also hosts other AI services.

1. Prepare or update ComfyUI

Clone the official ComfyUI repository into a dedicated application directory. If it already exists, update it with a fast-forward pull and keep a copy of the working workflow before changing dependencies.

git clone https://github.com/Comfy-Org/ComfyUI.git
cd ComfyUI
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip setuptools wheel
python -m pip install -r requirements.txt

For an existing installation, the equivalent maintenance step is:

cd ComfyUI
git pull --ff-only
source .venv/bin/activate
python -m pip install --upgrade -r requirements.txt

After updating, restart ComfyUI and verify that the custom nodes required by the MiniMax-H3 workflow load without red errors. Keep the ComfyUI commit, Python package versions, PyTorch version, and CUDA version in the deployment record so a working environment can be reproduced.

2. Download the model from Hugging Face

Authenticate with Hugging Face only if the repository requires it, then download the model into a persistent local model directory. The repository page and the workflow loader together determine the final subdirectory layout.

python -m pip install -U huggingface_hub
huggingface-cli download MiniMaxAI/MiniMax-H3   --local-dir ./models/MiniMax-H3

If the workflow expects separate diffusion, text-encoder, VAE, or tokenizer folders, place or link each component into the corresponding ComfyUI model directory. Do not rename files merely to make a dropdown entry appear: first confirm the loader type and expected directory in the workflow. Keep the download revision and SHA256 values with the deployment record.

3. Import the H3 workflow

Open the supplied MiniMax-H3 workflow JSON in the ComfyUI browser, or load the current workflow template that accompanies the model package. In each model loader, select the downloaded H3 diffusion weight, the quantized text encoder, and the matching VAE or tokenizer. If a node shows a missing model, fix the model path or refresh the model list before pressing Queue.

The tested low-precision route used an FP8 diffusion weight and an NVFP4/AWQ text-encoder variant. The graph should expose at least the prompt, seed, frame count, width, height, sampler, scheduler, and output path. For a first usable preset on an RTX 4090, use 608×352, 24 fps, 20 steps, res_multistep, simple, and 124 frames for a short preview. Promote a selected shot to 484 frames only after the short preview is visually acceptable.

4. Start ComfyUI with conservative memory settings

On the tested 4090 path, start the service in low-VRAM mode and keep the port private to the trusted network or protect it with a reverse proxy. A generic launch command is:

source .venv/bin/activate
python main.py --listen 0.0.0.0 --port 8198 --lowvram

The exact port is a deployment choice. Do not stop or restart unrelated GPU services when changing this service. Before queuing H3, confirm that the GPU has enough free memory and that no second H3 process is already running. Long clips can peak around 46 GiB, so serialize jobs rather than attempting parallel generation.

5. Run and validate a generation

Enter a short prompt, choose a fixed seed, and queue one preview. Watch the ComfyUI progress state and GPU telemetry until the job completes. Then validate the output as a media file: it should contain readable H.264 video and AAC audio, and the actual duration should be close to the requested frame count at 24 fps. A successful queue response is not enough; a player must open the result and the content must still pass human review.

For production use, record the workflow revision, model revision, precision, resolution, frame count, seed, steps, sampler, scheduler, elapsed time, peak VRAM, actual duration, and output path. Keep native audio optional when exact dialogue matters. Check hands, text, object identity, sports logic, trajectories, and audio-video sync before delivering a clip.

6. Troubleshoot common failures

If ComfyUI reports a missing model, check the loader’s expected directory and refresh the model list. If the process reaches out-of-memory, reduce frame count or resolution first, then confirm low-VRAM mode and CPU offloading. If generation completes but the video has no audio, inspect the audio node and the output container separately. If repeated runs become unexpectedly fast, check whether ComfyUI returned a cache hit; cached previews are useful for iteration but should not be reported as fresh inference timings. If a custom node fails after an update, compare the recorded ComfyUI commit and package versions before changing multiple dependencies at once.

MiniMax-H3 setup and RTX 4090 video generation appendix

Tested environment

ItemTested value
GPUOne NVIDIA RTX 4090
DiffusionFP8 scaled weight
Text encoderNVFP4/AWQ Qwen3-VL 32B variant
Baseline608×352, 24 fps
Samplerres_multistep
Schedulersimple
Steps20
OutputH.264 video, AAC audio, 32 kHz stereo

Model download

The tested weights came from the Hugging Face MiniMax-H3 repository. A reproducible download pattern is:

huggingface-cli download MiniMaxAI/MiniMax-H3 --local-dir ./models/MiniMax-H3

The FP8 diffusion file used in this review has SHA256 12944c1f7791637e7de12208aef04da82bd26b95271b1b47d817364315ade993. Keep the Python environment isolated and use the current ComfyUI workflow template.

Comparison context

MiniMax’s official Video-01 announcement describes 720p, 25 fps generation up to 6 seconds, while this local H3 workflow produced 20.75-second clips with an audio stream. These are not an equivalent benchmark; they represent different workflows and model generations. Video-01-style short high-definition generation is attractive for quick polished shots, while H3’s local path is attractive when local control, privacy, and longer drafts matter. See the official Video-01 announcement and the MiniMax official GitHub organization.

FAQ

Can H3 run on one RTX 4090?

Yes. The quantized FP8/NVFP4 workflow ran on one RTX 4090, but long clips peaked around 46 GiB, so keep one job on the GPU.

Which resolution should I use?

Use 608×352 for iteration, 640×384 for a modest clarity increase, and 720×416 or 768×432 for selected hero shots. 960×544 passed but took about 24 minutes for one long clip.

Is native audio ready for final dialogue?

No. Several clips contained repeated phrases, wrong-language fragments, music-like hallucinations, or unrelated words. Transcribe and listen to every dialogue clip, or replace the soundtrack.

Is H3 suitable for sports or instruction?

Only with human review. Basketball, soccer, running, cards, hands, and product assembly exposed physical or topology errors.

What should I use H3 for?

Use it for storyboards, mood reels, concept pitches, background plates, transitions, and rough cuts. Treat speech, text, sports rules, hands, and exact object relationships as review-required areas.

Evidence and previews

The following entries contain the measured settings, runtime, VRAM, media validation, audio transcript, and a playable local video element for every formal test clip.

Motion and action examples

SceneGeneration settingsRuntime / VRAMOutputPreview
Close-up speech608×352 / 484 frames / seed 2026082001 / 20 steps352.93 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The target sentence was partly clear at the start, followed by repeated and meaningless word strings.
Side-profile speech608×352 / 484 frames / seed 2026082002 / 20 steps352.91 s / 46250 MiB20.75 s, H.264 + AAC
Test point: Unintelligible English fragments appeared and the target sentence was repeated.
Two-person dialogue608×352 / 484 frames / seed 2026082003 / 20 steps358.07 s / 45130 MiB20.75 s, H.264 + AAC
Test point: The transcript resembled meaningless Japanese-style speech and did not form reliable dialogue.
Multi-person meeting608×352 / 484 frames / seed 2026082004 / 20 steps358.07 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The meeting audio was incoherent and could not serve as normal dialogue.
Singing and performance608×352 / 484 frames / seed 2026082005 / 20 steps357.99 s / 46250 MiB20.75 s, H.264 + AAC
Test point: “Yeah” was repeated for a long time without recognizable lyrics.
Fast hand gestures608×352 / 484 frames / seed 2026082006 / 20 steps358.01 s / 46218 MiB20.75 s, H.264 + AAC
Test point: The playing cards had errors in shape, count, and relationships.
Multi-step cooking608×352 / 484 frames / seed 2026082007 / 20 steps358.13 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The audio transcript was incoherent English.
Boxing combination608×352 / 484 frames / seed 2026082008 / 20 steps358.02 s / 46250 MiB20.75 s, H.264 + AAC
Test point: Audio and video were not synchronized.
Basketball dribble, pass, and shot608×352 / 484 frames / seed 2026082009 / 20 steps358.01 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The movement and ball path did not follow basketball logic.
High-speed soccer action608×352 / 484 frames / seed 2026082010 / 20 steps357.96 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The action transitions were incoherent and the audio did not form normal sentences.
Skateboard jump608×352 / 484 frames / seed 2026082011 / 20 steps358.02 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The audio sounded more like an end-credit music cue.
Sharp bicycle turn608×352 / 484 frames / seed 2026082012 / 20 steps358.14 s / 46250 MiB20.75 s, H.264 + AAC
Test point: Subtitle-like content appeared and the language did not match the scene.
Continuous pet interaction608×352 / 484 frames / seed 2026082013 / 20 steps358.04 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The audio was transcribed as a continuous sequence of musical symbols.
Group running and overtaking608×352 / 484 frames / seed 2026082014 / 20 steps358.01 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The only recognized phrase was “Thank you,” unrelated to the race.
Crowd occlusion and foreground motion608×352 / 484 frames / seed 2026082015 / 20 steps358.00 s / 46218 MiB20.75 s, H.264 + AAC
Test point: A Japanese-style end-credit thank-you appeared.
Fast whip-pan transition608×352 / 484 frames / seed 2026082016 / 20 steps357.99 s / 46250 MiB20.75 s, H.264 + AAC
Test point: Traffic-light rules, text rendering, and dialogue were all unreliable.
Chase and camera cuts608×352 / 484 frames / seed 2026082017 / 20 steps357.96 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The audio appeared as subtitle-like content.
Fast multi-scene montage608×352 / 484 frames / seed 2026082018 / 20 steps357.99 s / 46250 MiB20.75 s, H.264 + AAC
Test point: Only the word “you” was recognized and it could not serve as narrative dialogue.
Multi-step product assembly608×352 / 484 frames / seed 2026082019 / 20 steps358.00 s / 45130 MiB20.75 s, H.264 + AAC
Test point: The final frame conflicted with the camera structure, and the audio was unrelated to assembly.
Emotional speech and movement608×352 / 484 frames / seed 2026082020 / 20 steps358.08 s / 46218 MiB20.75 s, H.264 + AAC
Test point: The main sentence was mostly recognizable but repeated phrasing appeared.

Long-form scene examples

SceneGeneration settingsRuntime / VRAMOutputPreview
Natural landscape608×352 / 484 frames / seed 2026081001 / 20 steps352.92 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The transcript was too short or unrelated to the image and could not serve as reliable dialogue.
City at night608×352 / 484 frames / seed 2026081002 / 20 steps352.85 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The transcript was too short or unrelated to the image and could not serve as reliable dialogue.
Chinese scene608×352 / 484 frames / seed 2026081003 / 20 steps352.92 s / 46250 MiB20.75 s, H.264 + AAC
Test point: Manual review recommended.
Complex Chinese scene608×352 / 484 frames / seed 2026081004 / 20 steps352.79 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The transcript was too short or unrelated to the image and could not serve as reliable dialogue.
Text generation608×352 / 484 frames / seed 2026081005 / 20 steps357.93 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The transcript was too short or unrelated to the image and could not serve as reliable dialogue.
Chinese text608×352 / 484 frames / seed 2026081006 / 20 steps352.99 s / 46250 MiB20.75 s, H.264 + AAC
Test point: A large amount of extra text suddenly appeared in the image.
Portrait608×352 / 484 frames / seed 2026081007 / 20 steps352.94 s / 46250 MiB20.75 s, H.264 + AAC
Test point: Manual review recommended.
Hand movement608×352 / 484 frames / seed 2026081008 / 20 steps353.01 s / 46250 MiB20.75 s, H.264 + AAC
Test point: A third hand appeared.
Realistic animal608×352 / 484 frames / seed 2026081009 / 20 steps353.04 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The transcript was too short or unrelated to the image and could not serve as reliable dialogue.
Animal group608×352 / 484 frames / seed 2026081010 / 20 steps352.94 s / 46250 MiB20.75 s, H.264 + AAC
Test point: Manual review recommended.
Pet interaction608×352 / 484 frames / seed 2026081011 / 20 steps352.94 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The ball trajectory was discontinuous.
High-speed motion608×352 / 484 frames / seed 2026081012 / 20 steps353.03 s / 46250 MiB20.75 s, H.264 + AAC
Test point: Manual review recommended.
Ball sport608×352 / 484 frames / seed 2026081013 / 20 steps352.92 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The basketball trajectory was unstable.
Water sport608×352 / 484 frames / seed 2026081014 / 20 steps353.01 s / 46250 MiB20.75 s, H.264 + AAC
Test point: Manual review recommended.
Long take608×352 / 484 frames / seed 2026081015 / 20 steps352.95 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The transcript was too short or unrelated to the image and could not serve as reliable dialogue.
Orbiting camera608×352 / 484 frames / seed 2026081016 / 20 steps352.95 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The transcript was too short or unrelated to the image and could not serve as reliable dialogue.
Zoom shot608×352 / 484 frames / seed 2026081017 / 20 steps353.00 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The transcript was too short or unrelated to the image and could not serve as reliable dialogue.
Ambient audio608×352 / 484 frames / seed 2026081018 / 20 steps353.00 s / 46250 MiB20.75 s, H.264 + AAC
Test point: Manual review recommended.
Audio-video sync608×352 / 484 frames / seed 2026081019 / 20 steps352.93 s / 46250 MiB20.75 s, H.264 + AAC
Test point: Manual review recommended.
Stylized608×352 / 484 frames / seed 2026081020 / 20 steps353.03 s / 46250 MiB20.75 s, H.264 + AAC
Test point: The transcript was too short or unrelated to the image and could not serve as reliable dialogue.

Resolution comparison

SceneGeneration settingsRuntime / VRAMOutputPreview
R01_640x384640×384 / 484 frames / seed not recorded / 20 steps90.88 s / 43306 MiB20.75 s, H.264 + AAC
Test point: The run did not follow the race rules, the track changed between frames, and the speech was unstable.
R02_720x416720×416 / 484 frames / seed not recorded / 20 steps610.15 s / 48394 MiB20.75 s, H.264 + AAC
Test point: The run did not follow the race rules and the audio transcript was incoherent English.
R03_768x432768×432 / 484 frames / seed not recorded / 20 steps731.05 s / 47466 MiB20.75 s, H.264 + AAC
Test point: The run did not follow the race rules and the audio transcript was incoherent English.
R04_960x544960×544 / 484 frames / seed not recorded / 20 steps1451.97 s / 48490 MiB20.75 s, H.264 + AAC
Test point: The run did not follow the race rules and the audio drifted semantically.

Duration comparison

SceneGeneration settingsRuntime / VRAMOutputPreview
T01_5sec608×352 / 124 frames / seed not recorded / 20 steps55.55 s / 41546 MiB5.17 s, H.264 + AAC
Test point: Manual review recommended.
T02_10sec608×352 / 244 frames / seed not recorded / 20 steps136.24 s / 42730 MiB10.83 s, H.264 + AAC
Test point: Manual review recommended.
T03_20sec608×352 / 484 frames / seed not recorded / 20 steps353.03 s / 46250 MiB20.75 s, H.264 + AAC
Test point: Manual review recommended.

Queue stability

SceneGeneration settingsRuntime / VRAMOutputPreview
Q01_dialogue608×352 / 124 frames / seed not recorded / 20 steps51.71 s / 41546 MiB5.17 s, H.264 + AAC
Test point: This was not normal dialogue; the audio was incoherent.
Q02_fast_action608×352 / 124 frames / seed not recorded / 20 steps103.47 s / 41546 MiB5.17 s, H.264 + AAC
Test point: The movement did not follow the rules of the sport.
Q03_product608×352 / 124 frames / seed not recorded / 20 steps155.17 s / 41546 MiB5.17 s, H.264 + AAC
Test point: Manual review recommended.

Speech and audio consistency

SceneGeneration settingsRuntime / VRAMOutputPreview
P01_repeat_seed1608×352 / 124 frames / seed 2026086001 / 20 steps51.72 s / 41546 MiB5.17 s, H.264 + AAC
Test point: Manual review recommended.
P02_repeat_seed1608×352 / 124 frames / seed 2026086001 / 20 steps3.13 s / 40042 MiB / GPU 0%5.17 s, H.264 + AAC
Test point: Manual review recommended.
P03_repeat_seed2608×352 / 124 frames / seed 2026086002 / 20 steps51.79 s / 41546 MiB5.17 s, H.264 + AAC
Test point: Manual review recommended.
P04_chinese_dialogue608×352 / 124 frames / seed 2026086004 / 20 steps51.76 s / 41546 MiB5.17 s, H.264 + AAC
Test point: Manual review recommended.

External references