Posts say MiniMax H3 reaches TopviewAI, Magnific, Hailuo, and Comfy workflows
Posts place MiniMax H3 in TopviewAI, Magnific, Hailuo, and Comfy workflows with native 2K, 15-second clips, lip sync, and multimodal references. Creators cite lower 2K pricing, while others note the weights were announced but not released.

TL;DR
- MiniMax H3 shipped as a 2K, multimodal video model, with Hailuo AI's announcement pointing to the launch and MiniMax's video docs listing
MiniMax-H3, 4 to 15 second outputs, and text, image, video, and audio inputs. - The price headline is real: David Comfort's pricing screenshot puts H3 at $7.80 per minute, while MiniMax's pay-as-you-go page lists 2K generation at $0.13 per second.
- Audio is becoming the control surface: Uncanny_Harry's test used a dialogue track plus an image reference, and bennash's music-video post used 15 second song clips to drive dance and lip sync.
- The rollout already reaches multiple creator surfaces, with Magnific's launch post, Runway's post, Pika's post, and hellorob's Comfy reply all placing H3 outside MiniMax's own Hailuo site.
- The open-weights story is still pending: kaigani's screenshot highlights MiniMax's plan to open weights "in the coming days," and PurzBeats's reply notes that announcing weights is not the same as dropping weights.
The official launch post says H3 uses Contextual Omni Representation, H3-VAE, H3-Omni Transformer, and In-Context Regeneration, which is the kind of model-stack naming that usually hides the creator-facing bit: one prompt can bind a camera reference, character image, and audio track. MiniMax's API docs cap mixed reference inputs at 12 files, while Matan Cohen Grumi's Pika breakdown shows a launch film built around a single parcel reference, 15 second audio slices, and one-change video edits. The ComfyUI changelog has already added MiniMax H3 model support, so this is moving through node workflows too.
What shipped
MiniMax's launch post describes H3 as a general-purpose multimodal generation model that understands text, images, video, and audio, then generates video with native stereo sound up to 15 seconds at 2K.
The docs make the usable shape clearer:
- Model name:
MiniMax-H3 - Output: 2K
- Duration: 4 to 15 seconds, integer values only
- Modes: text-to-video, first/last-frame image-to-video, and reference generation
- Reference inputs: up to 9 images, up to 3 video clips, up to 3 audio clips, capped at 12 files total
- Prompt limit: 7,000 characters
MiniMax positions H3 for advertising, branding, e-commerce, product design, UI/UX, gaming, and film work, and says early testing showed strength in instruction following, text and brand rendering, and V2V motion transfer.
The price math
MiniMax's pricing page lists H3 at $0.13 per second for 2K, or $1.95 for a 15 second generation. The same page lists 768P at $0.09 per second, with 768P marked as closed beta.
Input billing matters for reference-heavy workflows:
- Audio input: free
- Images: first 5 free, then $0.04 per additional image
- Video input: billed by input duration and output resolution
Several creator comparisons framed the number against Seedance. 0xInk_'s comparison said a 15 second Seedance 2.0 clip at 1080p cost around $6 depending on platform, while 15 seconds at 2K on H3 was around $1.50 in that test. 0xInk_'s reply separately put H3 around $1.50 for 15 seconds at 2K and Seedance 2.5 around $6 for 15 seconds at 720p upscaled to 2K.
Omni Reference
Magnific's launch post describes the H3 wrapper as supporting up to 9 images for character and style, 3 videos for motion and camera control, 2K output, and 5 to 15 second clips. That matches MiniMax's own reference model: the prompt can combine character, motion, camera, style, voice, or editing rhythm as separate inputs.
Ozan Sihay wrote in Turkish that H3 can use text, video, image, and sound as references in the same generation, with separate control over character, camera movement, and scene aesthetic. He also listed multi-character and camera references, motion transfer, voice cloning, micro editing of character, environment, motion, and sound, plus strict instruction following.
Creators immediately turned that into repeatable recipes. hellorob's comparison used 9 reference images against Seedance 2.0, while Uncanny_Harry's early-access post used 5 image references and said most of the short came from two 15 second generations.
Audio as direction
Uncanny_Harry recorded dialogue as an audio track, fed it with an image reference of the man, and used text for performance notes such as the look to camera. He said H3 appeared to use the exact provided audio, without needing a black video or the words written into the prompt as he had done with Seedance 2.0.
Bennash used one source video for the character, then 15 second song clips as the audio reference for each generated video. His prompt was compact: make the character from @Video1 dance to the beat from @Audio1 and lip sync with @Audio1 bennash's prompt.
The same pattern showed up in longer workflows. ai_artworkgen's music-video process used character sheets, a Suno track, 11 separate 15 second audio segments, H3 for video generation, and CapCut for the edit.
Text and interface motion
0xInk said several tests left H3 ahead of Seedance 2.5 for text in motion. In a client ad, 0xInk_'s French-client example said Seedance 2 handled most shots, but H3 produced the cleaner phone-scroll text motion for the first shot.
UI and game motion were a separate mini-thread. AllaAisling's deck-builder prompt specified a board reference, a UI kit reference, card dealing, energy spend, enemy turns, and crisp screen-locked UI text. koldo2k's gameplay test said H3 kept a HUD intact while the interface reacted to the scene context.
The little miracle test was handwriting. venturetwins' chalkboard post said H3 succeeded at a long-running prompt that asks a model to write "Hi" on a chalkboard.
Long prompts and cut plans
Dustin Hollywood said two 2K clips used prompts of 6,219 and 6,788 characters, with 24 cuts prompted in sequence and no added effects or speed changes. That sits just under the 7,000 character prompt cap in MiniMax's docs.
AllaAisling posted a fashion workflow built around one continuous shot, one model, and eight outfit changes hidden behind real occluding objects. The prompt thread specified the identity reference, garment grid, street reference, occluder timings, fabric behavior, unbroken footsteps, and negative constraints against cuts, morphing, pace changes, identity drift, and shadow resets AllaAisling's full prompt.
Techhalla used a different planning trick: generate a 3x3 storyboard grid first, then animate it with H3 as the only visual reference techhalla's 3x3 workflow. The follow-up prompt maps nine beats across a 15 second sci-fi desert sequence, with each beat assigned a time range techhalla's prompt.
Where it shows up
Runway said H3 is available alongside other frontier models on its platform. Pika said H3 is available through the Pika MCP, and the Pika MCP page frames that surface as an agent-connected creative workflow.
Matan Cohen Grumi's Pika breakdown is the best workflow note in the rollout. He said the launch film used the same parcel reference image across generations, gave each scene its own 15 second music slice, used 9 image references for one laundromat scene, and sent a cucumber-cutting shot back into H3 for one-change edits Matan Cohen Grumi's edit note.
Magnific said H3 had landed with 2K, 5 to 15 second generations and up to 9 images plus 3 videos as references Magnific's launch post. TopviewAI promoted native 2K, 15 second videos, multimodal generation, and a price at 30% of Seedance 2.0 TopviewAI's launch post.
Comfy is already in the loop. hellorob's Comfy reply said H3 was live in Comfy and coming soon to ComfyCloud, and the ComfyUI changelog lists MiniMax H3 under partner node updates.
Open weights
MiniMax's launch post says the company plans to open up H3 model weights "in the coming days," subject to applicable laws and regulations. The same paragraph says hardware compatibility was considered from the earliest stages of H3's design.
That wording produced a predictable timeline mess. Kaigani flagged the open-weights line as news, while PurzBeats replied that announcing weights and dropping weights are different events, because some posts were already telling people the weights were downloadable.
The clean status at publication: H3 is live as a product and API model, while downloadable weights remain promised rather than shipped.