Skip to content
AI Primer
workflow

MiniMax H3 Max workflow uses staged prompts to control timing

Glenn Has A Beard found that MiniMax H3 Max can miss object orientation and delayed cause-and-effect. His tests lock the camera, specify a final state, and stage timing constraints one at a time.

4 min read
MiniMax H3 Max workflow uses staged prompts to control timing
MiniMax H3 Max workflow uses staged prompts to control timing

TL;DR

  • Delayed cause and effect becomes more controllable when the time window is written into the shot: GlennHasABeard's fan test keeps papers flat until the final two seconds.
  • Physical motion benefits from a pictured endpoint and measured duration. GlennHasABeard's cork test gives the cork a hand-height destination, while his feather test assigns its fall twelve seconds across one window.
  • The reveal itself can be the animation, as CharaspowerAI's example has particles assemble into a portrait rather than cutting to a finished image.
  • Local generation changes the cost of experimentation, even when it is slow: GlennHasABeard's first-day report describes free prompt testing, while his render-time report puts one 15-second clip at 50 minutes to an hour.

A community ComfyUI Prompt Builder turns H3 requests into named sections, timings, and reference-media tags. FastVideo's eight-step checkpoint uses a different schedule from base H3, and a Comfy-Org template limits that speed-focused branch to text-to-video/audio.

One constraint at a time

GlennHasABeard's knight test bundles a hard geometry problem into 15 seconds: the piece stays upright on its base, traces an L-shaped glide, and remains under a locked camera. The first take went diagonal. A later wording change made the knight lie down instead, which is why he describes the process as closing each new failure mode separately.

The Prompt Builder project describes the same broader prompt grammar: named sections, shot timing, speaker IDs, and tags for reference media. Its maintainer presents it as a hand-authored replacement for the H3-Context-IR rewriting model, which was not open-sourced.

Two clocks

The fan experiment splits one shot into two schedules: the fan spins up while papers stay flat, then only the top three sheets lift from seconds 13 to 15.

His other tests use the same arrangement:

End states

The cork test swaps a loose action verb for a visible destination: the cork rises straight out of the bottle, rotates once, and finishes a hand's height above it.

A twelve-second fall over the height of one window supplies a rate for the feather in GlennHasABeard's timed-feather test. The domino scene applies the same literalism to cast size, specifying one piece and “no others ever appear” in GlennHasABeard's domino test; in GlennHasABeard's domino reply, he joked that the clip existed to bug a commenter.

Reference passes

GlennHasABeard says H3 is strongest for his live-action text-to-video work and text effects, while reference-to-video is less reliable. He still gives Seedance the edge at preserving an animation's established look in GlennHasABeard's model comparison.

Koldo2k uses reference video in several passes instead of asking for every change at once. His prompt locks the camera and composition, then changes only production design.

  • Opening framing, angle, height, distance, rotation, speed, and duration stay identical.
  • The citadel and two armies keep their screen positions.
  • Plastic materials become weathered limestone, steel, hide, iron, grass, and mud.
  • The pass bars interface elements, text, studs, seams, and plastic.

The community reports match that split. an r/StableDiffusion test found a 3D-reference workflow good for animation but plasticky on realistic people, while one r/StableDiffusion report described approximate resemblance in one reference path and mushy, pixelated output in another.

A local overnight queue

GlennHasABeard runs H3 locally and controls the setup from his phone.

His posted machine uses an i7-8700K, 32 GB of RAM, and a 12 GB RTX 3060.

He frames the wait as queue time, saying renders can run while he sleeps or works in GlennHasABeard's overnight-render reply. He avoids turbo because he has not found the quality trade worthwhile in GlennHasABeard's turbo comment, and says a higher-quality run is next in his quality-test reply.

FastH3 8-Step V2

The FastVideo variant is a separate speed branch. Its model card says it generates synchronized video and audio with eight transformer forwards, using a scheduler shift of 10 rather than base H3's 12.

r/StableDiffusion

Open Weight FastVideo FastH3 V2!

0 comments

Comfy-Org's text-to-video template says the distilled checkpoint supports text-to-video/audio only. First/last-frame and multi-reference conditioning remain tasks for base MiniMax H3.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 5 threads
TL;DR4 posts
Two clocks2 posts
End states3 posts
Reference passes1 post
A local overnight queue4 posts
Share on X