MiniMax H3 creators test audio references for lip sync and dialogue
Hailuo H3 tests use recorded audio to guide lip sync, emotion, and dialogue without black-video workarounds. Other demos show stronger text-in-motion, preserved game HUDs, and reference-heavy ad shots.

TL;DR
- H3's sharpest early workflow is audio as a performance reference: Uncanny_Harry's test used a recorded dialogue track plus one image reference, while bennash's music-video test drove dance and lip sync with 15-second song clips.
- Text-in-motion is the other creator magnet: 0xInk_ said H3 beat Seedance 2.5 after several text tests, and 0xInk_'s client ad used H3 for a scrolling-phone shot Seedance could not keep clean.
- Interface work held up better than usual in early demos: koldo2k's gameplay test kept a game HUD intact, and AllaAisling's deck-builder prompt asked H3 to keep UI screen-locked during camera moves.
- H3 is already spread across creator platforms: Runway announced it, Pika Labs shipped it through Pika MCP, and Magnific listed 2K clips from 5 to 15 seconds with image and video references.
- The big caveat is motion clarity: 0xInk_'s Seedance comparison said Seedance 2.0 was still better overall because H3 blurred during action scenes.
MiniMax's official H3 post describes a model that understands text, images, video, and audio together, then outputs up to 15 seconds at 2K with native stereo sound. The same post gives the cleanest mental model for the new workflow: “Reference the Hitchcock camera movement from Video 1, have the character in Image 2 sing, with the vocals matching Audio 3.” DavidmComfort's pricing screenshot puts H3 at $7.80 per minute with audio at 2K, and kaigani's screenshot caught the open-weights promise: “in the coming days.”
Audio references
Uncanny_Harry recorded dialogue as an audio track, fed it to H3 with an image reference of the man, then used the text prompt only for acting direction, including a look to camera at the end. His comparison said he did not need the Seedance workaround of feeding black video or writing the spoken words into the prompt.
Magnific framed the same capability as native per-character voice-to-lips sync in a four-shot scene generated from one prompt in its H3 demo. Hailuo's own account replied to one creator with the phrase “precise lip-sync” in a short response.
The music workflows were more mechanical:
- bennash's prompt used one source video for the character and 15-second song clips as
Audio1, asking H3 to make the character dance and lip-sync in sync with the track. - bennash then built Basic Slice, a free browser tool for cutting audio into 15-second or custom-length chunks; the linked tool processes audio client-side.
- MatanCohenGrumi said Pika's launch film gave every scene its own 15-second track slice, with bodies dancing to the waveform while faces stayed deadpan.
- ai_artworkgen used Suno for a song, cut it into eleven 15-second segments, then used character sheets as references for an H3 music video.
Audio reference is doing more than lip movement here. It is becoming the timing track for bodies, cuts, facial performance, and scene rhythm.
Text and UI surfaces
H3's text tests were not only subtitles. CharaspowerAI prompted a water-vortex reveal where droplets compress into the word “ASCEND,” and gizakdag posted large text animations made from a simple prompt.
0xInk_ pushed the comparison harder. After several tests, 0xInk_ said H3 handled text in motion better than Seedance 2.5, then a client ad example showed H3 replacing Seedance 2 for a first shot with clean scrolling-phone text.
The UI demos were unusually specific:
- venturetwins said a long-running chalkboard prompt finally produced a clear handwritten “Hi.”
- koldo2k recreated gameplay and said the HUD stayed intact while the interface reacted to context.
- CharaspowerAI tested interface motion design, with a mobile UI moving from a music player into market-style screens.
- AllaAisling's deck-builder prompt specified card shapes, filigree, iconography, legible text, hover states, energy counters, turn banners, and a “no UI drift or warping” negative prompt.
MiniMax's official launch post names “accurate text and brand rendering” as an H3 strength, which lines up with the creator pile-on around moving typography and screen graphics.
Omni Reference stacks
Magnific's H3 rollout listed up to 9 images for character and style, 3 videos for motion and camera control, 2K output, and clips from 5 to 15 seconds in its launch post. Fal's MiniMax H3 page lists a similar reference-heavy pattern: up to 9 images, 3 video clips, and 3 audio tracks in one generation.
Creators immediately treated the model like a reference mixer:
- hellorob compared H3 and Seedance 2.0 with 9 reference images, saying only one model used all 9 and produced the better output at one-third the cost.
- Uncanny_Harry made a short with 5 image references, with most of it coming from two 15-second generations.
- Artedeingenio uploaded several illustrations and a song to Omni Reference for an illustrated music test.
- techhalla used a 3x3 storyboard grid as the visual reference for a hard sci-fi desert sequence.
AllaAisling's wardrobe-change prompt shows how structured the references can get. The prompt split the job into a model identity image, a garment grid, and a location-grade image, then asked for eight outfits in one 15-second continuous shot.
The mechanics were explicit:
- One locked model identity.
- Eight complete outfits from a flat-lay grid.
- One street location and grade reference.
- Seven outfit changes hidden behind physical occluders.
- Each occlusion lasting 0.3 to 0.5 seconds.
- Continuous walk cycle, footsteps, street ambience, crowd flow, sun direction, and fabric behavior.
Previs, grids, and beat edits
The strongest H3 workflows look less like prompt-only generation and more like cheap preproduction. bennash first made a previs animatic with audio and still images as placeholders, then used H3 to turn it into a polished futuristic city shot.
Pika's launch-film thread adds another repeatable structure. MatanCohenGrumi made a parcel stand in for the model itself, kept the same parcel reference image in every generation so it would not drift across scenes, and sent a cucumber shot back into H3 with one change per pass to create seamless food-cut edits.
techhalla's 3x3 workflow put a storyboard before H3. The prompt first generated a nine-panel hard-sci-fi grid, then mapped H3's 15 seconds into timed shots: wide desert walk, face close-up, tool insert, industrial tower, sandstorm advance, rocky outcrop, dune charge, macro sand particles, and cave close-up.
Access, pricing, and open weights
H3 was not confined to Hailuo. Runway added it to its frontier-model lineup, Pika Labs put it on Pika MCP, Magnific shipped it with reference controls, TopviewAI advertised native 2K and 15-second videos, and hellorob said it was live in Comfy and coming soon to ComfyCloud.
The creator pricing math clustered around the same range. 0xInk_ put H3 at about $1.50 for 15 seconds at 2K versus about $6 for 15 seconds of Seedance 2.0 at 1080p, and dustinhollywood calculated $0.13 per second as $1.95 per 15-second generation.
The open-weights language stayed future tense. kaigani highlighted MiniMax's plan to open model weights “in the coming days,” while PurzBeats warned that announcing weights and dropping weights were different events.
Motion blur and weird failures
H3 did not clean-sweep the creator tests. 0xInk_ said Seedance 2.0 was still better overall on the same prompt and character references because H3 suffered from motion blur during action scenes, even though H3 was much cheaper.
Other rough edges were stranger. kaigani's café outtake said H3 read out the instructions during a comedy scene, and a burst-frame test missed one of 20 requested frames in five seconds. mrjonfinger described a recurring look across H3 clips as “slightly uncanny smoothed motion and intense sharpness.”