MiniMax H3 tests in Hailuo early access with Omni Reference and native 2K video
Creators are posting MiniMax H3 clips from Hailuo early access. They cite Omni Reference, native 2K with audio, 15-second clips, motion transfer, voice cloning, micro edits, and fast minimal-prompt runs.

TL;DR
- MiniMax H3 is still moving through early access, according to BLVCKLIGHTai's reply, while Hailuo_AI's reply confirmed the model name as H3.
- The core feature is Omni Reference, with ozansihay's feature list describing video, image, text, and audio references in one generation.
- The demos are already pushing production controls, from native 2K with audio in chrisfirst's clip to 24 prompted cuts in dustinhollywood's first-try cut test.
- Reference continuity is the headline skill, with Uncanny_Harry's five-reference test and AIwithSynthia's one-character fashion walk both built around character retention.
- The rough edges are visible too, with kaigani's Café outtake reading instructions aloud and BLVCKLIGHTai's alternate outcome arguing that multiples still matter.
Daily Economic News saw H3 at WAIC under a “Coming Soon” sign, framed as a model that treats text, images, video, and sound as one multimodal context. The public MiniMax API docs still list Hailuo 2.3 as the current video model, at 1080p 6s or 768p 6s/10s, while a third-party Muapi model page already lists H3 workflows for text-to-video, image-to-video, and frame-controlled video. Hailuo_AI's own replies nudged creators toward motion transfer, commercial generation, and film teasers Hailuo_AI's motion-transfer reply Hailuo_AI's commercial-generation reply Hailuo_AI's film-teaser reply.
Omni Reference
The most useful H3 description came from ozansihay, who said MiniMax opened the new model for early access inside Hailuo and described it as an “omni” model rather than a text-only video generator.
The feature stack he listed:
- Video, image, text, and audio can be used as references in the same generation.
- Character, camera movement, and scene aesthetic can be guided separately.
- MiniMax is emphasizing multi-character and camera references.
- Motion transfer is part of the announced feature set.
- Voice cloning is part of the announced feature set.
- Micro editing covers character, environment, motion, and audio.
- Strict instruction following is a stated focus.
That tracks with the WAIC framing: Daily Economic News reported that H3 is meant to move beyond separate image, video, and sound generation tasks into a shared multimodal creative context.
AllaAisling's first H3 post explicitly called the workflow “Omni Reference,” while AIwithSynthia used the same feature name for a single-character fashion transformation AIwithSynthia's fashion test.
Character and style locks
Uncanny_Harry said his short used five image references, kept the art style intact, changed camera angles, and followed direction across mostly two 15-second generations. The reference setup was simple: four character references plus one troll reference Uncanny_Harry's character references Uncanny_Harry's troll reference.
AIwithSynthia's fashion test turned the same character through six outfit changes in one uninterrupted shot, using Omni Reference to keep the face and body stable while the styling changed AIwithSynthia's fashion test.
In a reply, Uncanny_Harry summed up the early win as retention: H3 was “really good at retaining character and art style” Uncanny_Harry's retention reply.
Native 2K, audio, and 15 seconds
chrisfirst posted “Synthetic Daydreams” as native 2K with audio. Uncanny_Harry also called the model “2k native” in a reply Uncanny_Harry's 2K reply.
The 15-second generation length showed up repeatedly. techhalla answered a duration question with “15s” techhalla's duration reply, and the main prompt examples in the thread were built as 0 to 15 second timelines.
dustinhollywood pushed long prompt adherence hardest: two clips generated at 2K, with prompts of 6,219 and 6,788 characters, plus 24 cuts in sequence. He said the fast cuts were generated by H3, not added in editing dustinhollywood's first-try cut test.
He later claimed H3 had 7,000-character context adherence, native 2K out of the box, lower pricing than Seedance 2, and no character or face restrictions dustinhollywood's Seedance comparison. techhalla added one useful correction on “8K” language: the setting he was discussing was an optional quality booster, not literal 8K output techhalla's quality-booster reply.
Storyboard prompts
techhalla's clearest workflow was not “write a prompt and hope.” It was storyboard-first:
- Generate a cinematic 3x3 grid in Nano Banana Pro.
- Use that grid as the only visual reference in MiniMax H3.
- Lock the same character, face, scars, eyes, headscarf, armor, cape, and environments across panels.
- Split the video into timed beats from 0 to 15 seconds.
- Specify camera language for each beat, including wide establishing, extreme close-up, detail insert, hero wide, dynamic tracking, macro atmospheric, and intimate close-up.
The grid method turns H3 into a continuity engine. The prompt gives it shot order, identity constraints, palette, lens language, and texture targets instead of a single mood board.
ozansihay's prompt used the same production grammar for a different problem: a single continuous shot that starts as 1930s rubber-hose animation and becomes live-action photorealism. The prompt separately defined style, setting, dialogue, transformation mechanics, camera behavior, lighting, color, and audio, including English lip sync ozansihay's cartoon-to-live-action prompt.
Cuts, multicam, and comedy
BLVCKLIGHTai tested multicamera shots, back-and-forth dialogue, and a small comedy setup around a glowing orb. He also said his earlier H3 clips were the fastest generations he had seen at that quality with minimal prompting BLVCKLIGHTai's minimal-prompt examples.
Comedy still has weird failure modes. kaigani said a Café benchmark outtake was meant to be comedy, read out the instructions, and still landed as some of the better comedy these models have produced kaigani's Café outtake.
The strongest cut-heavy demo came from dustinhollywood, who said he generated only two clips and matched them to the music progression after H3 produced the 24-cut sequence dustinhollywood's first-try cut test.
Early access, price, and Hailuo's hints
Access is still uneven. BLVCKLIGHTai said H3 was early access and “should be releasing soon,” while techhalla also said he had early access techhalla's early-access reply. ozansihay told one user the test was inside Hailuo's own platform ozansihay's Hailuo-platform reply.
Pricing claims are coming from creators, not an H3 pricing page. techhalla said H3 was “like 5 times cheaper than Seedance 2.0” if he was not wrong techhalla's pricing reply, BLVCKLIGHTai called the pricing “really good” BLVCKLIGHTai's pricing reply, and dustinhollywood claimed it was cheaper with no character or face restrictions dustinhollywood's Seedance comparison.
Hailuo_AI's replies point at the launch lanes MiniMax wants creators to test: motion transfer, film teasers, complicated references, and commercial generation for e-commerce and branding videos Hailuo_AI's motion-transfer reply Hailuo_AI's film-teaser reply Hailuo_AI's complicated-reference reply Hailuo_AI's commercial-generation reply. The account also replied “H3!” when asked what model was being used Hailuo_AI's H3 reply.