MiniMax H3 supports audio and image references for lip sync and motion control
Creators fed recorded dialogue, Seed Audio, and images into Hailuo H3 for lip sync, UI motion, outfit changes, and reverse-time shots. Hailuo is also pitching H3 for ads, UI/UX, and precise edits.

TL;DR
- MiniMax launched H3 as a general multimodal video model for text, image, video, and audio context, with 15-second 2K output and native stereo sound, according to MiniMax's official H3 post and Hailuo_AI's launch post.
- Creator tests clustered around audio control: Uncanny_Harry's dialogue test fed recorded speech and an image reference into H3, while bennash's music-video test used 15-second song slices to drive dancing and lip sync.
- H3's creator sweet spot is motion graphics, UI, and text inside the scene: CharaspowerAI's interface test pushed transitions and micro-interactions, and koldo2k's gameplay test kept a HUD intact while the game scene moved.
- The rollout spread quickly across creator platforms: runwayml's post said H3 was available on Runway, Magnific's launch post listed 2K clips with up to 9 images and 3 videos, and Pika's launch post put it inside Pika MCP.
- The main caveat from hands-on tests is action clarity: 0xInk_'s comparison said Seedance 2.0 was still better overall because H3 suffered from motion blur in action scenes.
MiniMax's official H3 post says the model uses text, images, video, and audio as one context, then generates video with native stereo sound. The API guide lists text-to-video, first/last-frame image-to-video, and reference generation, with 768p or 2K output, 4 to 15 second durations, up to 9 reference images, up to 3 videos, up to 3 audio clips, and 12 total files. MatanCohenGrumi's Pika thread showed a beat-grid workflow, while Basic Slice appeared almost immediately as a browser tool for chopping audio into H3-friendly 15-second chunks.
Omni Reference
MiniMax frames H3 as a unified-context model, not a pile of separate text-to-video, image-to-video, lip-sync, and editing tools. The docs describe reference generation as prompt plus images, videos, or audio to control character, motion, camera, style, voice, or editing rhythm.
The creator-facing interface made that model visible: upload images, videos, or audio, then combine characters, scenes, actions, and music. A Hailuo web screenshot shared by DavidmComfort showed “Refs (0/12),” Omni Reference, MiniMax H3, 2K, and 15 seconds in the same creation panel.
Hailuo kept nudging testers toward heavier reference stacks. Hailuo_AI told one creator to try “more complicated omni ref,” and Hailuo_AI later called Omni Reference “highly recommend.”
Audio as the edit timeline
The fastest creator discovery was that H3 could treat audio as more than a vibe reference. For voice actors, it appears to preserve the uploaded performance closely enough to matter.
Uncanny_Harry recorded dialogue, fed that track in with an image reference of the actor, and used the text prompt for performance direction like a look to camera. He wrote that, unlike Seedance 2.0, he did not need a black video or the dialogue typed into the prompt to get the words right.
Other tests converged on the same workflow:
- DavidmComfort generated dialogue and ambient sound in Seed Audio, then used it in Omni MiniMax H3, with “pretty good” lip sync and weaker source audio.
- ozansihay showed H3 generating Turkish spoken video directly from a prompt.
- bennash said H3 made a character dance to the beat and lip-sync, then bennash's prompt showed the setup: one source video, 15-second song clips, and an instruction to dance and lip sync to Audio1.
- bennash built Basic Slice because H3-style lip-sync references made quick 15-second audio slicing useful.
The audio feature turned H3 into Christmas morning for AI music-video makers. Timing, mouth shapes, and character motion moved into the same prompt surface.
Text and UI motion
Hailuo is openly pitching H3 at UI/UX, ads, gaming, and precise control. Hailuo_AI said H3 excels at UI/UX and commercial scenarios, while Hailuo_AI called it good at precise editing and control.
The best UI examples were not generic sci-fi dashboards. They tested whether interface elements stayed legible while the camera, transitions, and micro-interactions moved.
Creator tests split into a few repeatable UI/text patterns:
- Interface animation: CharaspowerAI highlighted transitions, micro-interactions, camera movement, and text integration.
- Text VFX: CharaspowerAI prompted ocean spray to compress into “ASCEND,” then explode into light.
- Ad typography: 0xInk_ used H3 for a phone-scroll shot because Seedance 2 struggled with clean moving text.
- Game HUDs: AllaAisling prompted a deck-builder turn with screen-locked cards, energy, damage numbers, and an END TURN button.
- Chalkboard writing: venturetwins said a long-running video-model test finally worked when H3 wrote “Hi” on a chalkboard.
Occlusion cuts
AllaAisling's fashion shot is the cleanest proof-of-work for continuity control: one continuous shot, one model, eight outfit changes, every change hidden behind a real object passing the lens.
The prompt structure is worth stealing as a format, because it turned the sequence into a timing sheet:
- Lock identity from Image 1.
- Pull eight complete garments from Image 2.
- Pull street, crowd density, light, and grade from Image 3.
- Hide each outfit change behind a total occluder for 0.3 to 0.5 seconds.
- Preserve walk rhythm across every change.
- Make fabric behavior match material: satin specular, wool matte, linen translucent, silk trailing, denim structured.
- Keep street behavior continuous, including pedestrians, traffic flow, sun direction, and shadows.
That same continuity obsession showed up in reverse-time and physics tests. AllaAisling's “Undo” prompt asked for wine climbing back into a glass, smoke retracting into a wick, torn photos knitting together, and reversed foley that stayed legible.
Beat-grid music videos
Pika's H3 launch thread turned audio control into a production recipe. The parcel was the repeating visual object, each scene used its own 15-second audio slice, and the edit was cut to the song's beat grid.
The mechanics were simple and sharp:
- Same parcel reference image in every generation, so the object carried continuity from scene to scene.
- Each scene got a 15-second slice of the track as an audio reference.
- The prompt told bodies to dance on beat while faces stayed deadpan.
- One laundromat scene used 9 image references: parcel, outfits, faces, shoes, empty location, and a dryer.
- A cucumber shot was fed back as input with one change per pass, swapping it for a carrot while preserving the original video.
ai_artworkgen's music-video workflow landed on the same segmented pattern: Midjourney for original images, GPT plus Nano Banana for character sheets and expansion, Suno for the track, H3 for 11 clips at 15 seconds each in 2K, then CapCut for the edit, according to ai_artworkgen's process post.
Ads and product shots
H3's commercial lane is already obvious: product motion, beauty macro shots, UGC, typography, and brand-safe composition in one 15-second clip.
hasantoxr called H3 “the closest thing we've got to After Effects for AI video,” then listed the useful bits for commercial work: baked-in text, subtitles, brand elements, flexible visual and audio inputs, in-model compositing, and ad-ready aesthetics, in hasantoxr's launch read.
Creator prompts pushed H3 into formats agencies actually bill for:
- Beauty: AIwithSynthia combined a consistent character, marble staircase, velvet sofa, macro applicator shots, glossy lips, and a floating product hero shot.
- Travel UGC: dustinhollywood prompted a 15-shot smartphone-style day-in-the-city sequence with selfies, subway, ramen, mirror shots, and neon streets.
- Product-food cinematography: HalimAlrasihi used Omni Reference for commercial and cinematic pancake shots.
- Platform montage: magnific said the clips and effects in its H3 reel were done with the model, with music added behind them.
Price, open weights, and access
MiniMax's official post says H3's 2K per-second price is less than a third of mainstream models and says the company plans to open the weights “in the coming days,” subject to laws and regulations.
A pricing screenshot shared by DavidmComfort put H3 at $7.80 per minute with audio at 2K, below Dreamina Seedance 2.0 1080p, HappyHorse-1.1, and Kling 3.0 1080p in that comparison. dustinhollywood did the same math as $0.13 per second, $1.95 for a 15-second generation, and $19,500 for 10,000 generations.
Access spread beyond Hailuo fast:
- runwayml said MiniMax H3 was available on Runway.
- Magnific listed H3 with up to 9 images, 3 videos, 2K output, and 5 to 15 seconds.
- Pika announced H3 on the Pika MCP, with the Pika MCP page positioning it as an agent-connected creative workflow.
- MayorKingAI said Topview was offering native 2K, 15-second generations, and multimodal control at 30 percent of Seedance 2.0's cost.
- PurzBeats showed H3 reference-to-video on ComfyUI Cloud at 2K, 8 seconds, and roughly 300 credits.
Rough edges
The strongest caveat came from side-by-side action testing. 0xInk_ said Seedance 2.0 was still better overall because MiniMax H3 had motion blur during action scenes, even though H3 was much cheaper and produced solid results.
Other rough edges were narrower:
- underwoodxie96 said skin texture still needed improvement after the giant-koi park test.
- kaigani tried a 20-frame burst prompt and said H3 captured 19 frames in 5 seconds.
- ozansihay agreed that one comparison output was skipping frames.
- PurzBeats warned that announcing open weights and dropping weights are separate events, after some users treated the weights as already downloadable.