MiniMax H3 ships in Hailuo, Runway, Pika, and Magnific creator tests
Hailuo, Runway, Pika, and Magnific now offer MiniMax H3 or early access to creators. Tests cite multi-reference control, 15-second 2K/4K outputs, audio-aware motion, and some blur in action shots.

TL;DR
- MiniMax H3 is live as an omni-modal video model, with Hailuo’s H3 post and MiniMax’s official blog framing it around text, image, video, and audio context rather than a single prompt-to-video lane.
- The rollout hit creator surfaces fast: Hailuo listed editing, text-to-video, and image-to-video modes, while Runway’s announcement, Pika’s MCP post, and the Magnific rollout all put H3 in front of users.
- Omni Reference is the practical hook: Magnific’s launch note lists up to 9 images, 3 videos, 2K output, and 5 to 15 second generations.
- The best early workflows use audio as structure, not decoration: MatanCohenGrumi’s Pika breakdown used 15-second track slices per scene, and bennash’s prompt used a source video plus song clips for beat-matched dancing and lip-sync.
- The Seedance comparison is price versus motion clarity: 0xInk_’s side-by-side found Seedance 2.0 better overall in action scenes because H3 still blurred motion, while H3 cost about $1.50 for 15 seconds at 2K versus roughly $6 for 15 seconds at 1080p on Seedance 2.0, depending on platform.
The official MiniMax H3 post hides the nerd candy in the architecture section: 100K-token multimodal captioning runs distilled to about 4K tokens, a 4x effective sequence-length gain from H3-VAE, and nearly 30% higher training throughput. Pika MCP makes the rollout stranger, because H3 is already showing up inside an agent-style creative toolchain, not only in a video app. kaigani’s animation video points to a 97-style animation test, while kaigani’s café outtake caught the model reading instructions aloud and accidentally making decent comedy.
What shipped
MiniMax’s official article says H3 generates video with native stereo audio at up to 2K resolution and 15 seconds, and positions it for advertising, branding, e-commerce, product design, UI/UX, and gaming.
The public surfaces split like this:
- Hailuo: Hailuo’s mode list names video editing with audio, text-to-video with audio, and image-to-video without audio.
- Runway: Runway’s announcement says H3 is available alongside its other frontier models.
- Pika: Pika’s post says MiniMax H3 is available on the Pika MCP, with Pika MCP described as an agent-connected creative toolchain.
- Magnific: the Magnific rollout lists character, motion, and beat controls with up to 9 images, 3 videos, 2K output, and 5 to 15 second clips.
Christmas came early for creators who live inside model switchers. H3 did not land as a single website demo, it landed as a model people could immediately compare inside their existing production wrappers.
Omni Reference
The control story is Omni Reference. In ozansihay’s early-access notes, H3 can use video, image, text, and audio in the same generation, with separate control over character, camera movement, and scene aesthetic.
The same post says MiniMax highlighted:
- multi-character references
- camera references
- motion transfer
- voice cloning
- micro-editing for character, environment, motion, and sound
- strict instruction following
Magnific exposed a product-shaped version of that stack: its rollout post says creators can bring up to 9 images for character and style plus 3 videos for motion and camera control. CharaspowerAI’s spec note described a broader Omni setup with up to 12 mixed image, video, and audio references.
Matan Cohen Grumi’s laundromat scene used 9 image references:
- the parcel
- his outfit
- her outfit
- both faces
- both pairs of shoes
- the empty location
- one dryer
The parcel then became the continuity spine, because MatanCohenGrumi’s follow-up says the same reference image traveled through every generation without drifting.
Audio-locked motion
Pika’s H3 launch film treated audio like choreography data. MatanCohenGrumi’s breakdown says every scene received its own 15-second slice of the track, and the prompt locked the body to the beat while keeping the face deadpan.
Bennash used one source video for the character and 15-second song clips for each generated video. In the shared prompt, the instruction was simple: make the character from @Video1 dance to the beat from @Audio1 and lip-sync in sync with @Audio1.
David Comfort tested generated dialogue and ambient sound from Seed Audio with H3. The lip sync worked, while his note blamed the weaker speech quality on Seed Audio rather than H3.
Hailuo’s own mode list still matters here: Hailuo’s post says text-to-video and video editing support audio, while image-to-video is listed as no audio.
Multi-shot prompting
Early testers are pushing H3 with long timeline prompts instead of single scene descriptions. dustinhollywood’s test used two 2K clips with 6,219 and 6,788 character prompts, asked for 24 cuts in sequence, and said the fast cuts came from H3 before he matched them to the music progression.
underwoodxie96’s koi scene used a 15-second continuous single take with explicit time blocks:
- 0 to 3s: handheld medium shot of a suited man feeding koi
- 3 to 7s: a giant koi grabs the bread bag
- 7 to 11s: the fish pulls him toward the pond
- 11 to 15s: the bag tears, he falls in, and the koi swims past
The prompt also pinned continuity constraints: keep the man’s face, suit, tie, paper bag, koi markings, inertia, traction, and water displacement consistent, according to underwoodxie96’s full prompt.
Techhalla’s storyboard workflow went one step earlier in the pipeline. The 3x3 grid prompt first generated a 9-panel hard sci-fi board, then used that grid as the only visual reference for a 15-second H3 animation with shot-by-shot timing.
Commercial, UI, and game shots
The commercial tests are already polished enough to annoy production people. AIwithSynthia’s lip gloss campaign asked H3 to preserve a model’s facial identity, hairstyle, eye color, makeup, skin tone, body proportions, jewelry, and facial consistency across macro product shots and a floating hero shot.
Hailuo is steering users toward that lane too. In one Hailuo reply, the account said H3 is good at commercial generation for e-commerce and branding videos.
The UI tests are more interesting than the glossy ads:
- Interface motion: CharaspowerAI’s UI test shows a mobile mockup shifting from a music player into a stock tracking screen with synchronized transitions.
- Game UI: AllaAisling’s deck-builder prompt uses one image for the board, one for the UI kit, then demands screen-locked cards, legible text, no UI drift, no invented card names, and no warping.
- Fashion continuity: AIwithSynthia’s fashion test keeps one character consistent through six outfit changes in one uninterrupted shot.
Price and Seedance
The sharpest comparison came from 0xInk_, who used the same prompt and same character references against Seedance 2.0. 0xInk_’s post says Seedance 2.0 was still better overall, mainly because MiniMax H3 showed motion blur during action scenes.
The same comparison put a rough price on the trade:
- Seedance 2.0: about $6 for a 15-second 1080p clip, depending on platform
- MiniMax H3: about $1.50 for a 15-second 2K clip
MiniMax’s official blog makes the same claim at vendor scale, saying H3’s 2K per-second price is less than one third of mainstream models and its 768p price is less than half of mainstream 720p.
hellorob’s split-screen test framed the comparison around reference handling. The 9-reference comparison says only one model used all 9 reference images and only one generated the better output at one third the cost.
Open weights
MiniMax says it plans to open H3’s model weights in the coming days, subject to applicable laws and regulations. That is the sleeper detail in MiniMax’s official H3 post, because frontier video models have mostly stayed closed.
The architecture notes name four core pieces:
- Contextual Omni Representation: multimodal captioning describes the target video, the context, and relationships among context elements. MiniMax says most source material needs about 100K tokens of inference, then gets distilled to an average of roughly 4K tokens.
- H3-VAE: a rebuilt tokenizer gives a 4x effective sequence-length gain and helps make native 2K generation economical.
- H3-Omni Transformer: MiniMax says separating understanding and generation workloads improved end-to-end training throughput by nearly 30%.
- In-Context Regeneration: H3 uses the base model to regenerate low-resolution outputs in context instead of relying on a dedicated super-resolution module.
MiniMax also names its own gaps: stronger multimodal understanding, larger scaling, and better visual detail in some scenes are all listed as future work.