Skip to content
AI Primer
update

MiniMax H3 adds native 2K, 15-second generation to more creator tools

TopviewAI posts cite native 2K, 15-second generations, multimodal control, and claimed pricing at 30% of Seedance 2.0. Creators also tested H3 on ComfyUI Cloud, with ComfyUI support for open weights reported.

8 min read
MiniMax H3 adds native 2K, 15-second generation to more creator tools
MiniMax H3 adds native 2K, 15-second generation to more creator tools

TL;DR

  • MiniMax H3 is now a 2K, 15-second creator tool rather than a single launch-page demo, with Topview claiming native 2K, multimodal control, and pricing at 30% of Seedance 2.0 through MayorKingAI's Topview post.
  • The real workflow is reference-heavy: Magnific lists up to 9 images, 3 videos, 2K output, and 5 to 15 seconds, while MiniMax's docs cap mixed reference input at 12 files.
  • Cost is driving the timeline: dustinhollywood priced a 15-second 2K H3 generation at $1.95 from the API rate, and 0xInk_'s comparison put Seedance 2.0 around $6 for 15 seconds at 1080p on one platform.
  • The creator wins are audio sync, text motion, UI/HUD stability, and long prompt adherence, with Uncanny_Harry showing audio-reference dialogue without a black-video workaround and koldo2k testing a game HUD that stayed intact.
  • Open weights are promised, not downloadable yet: kaigani surfaced MiniMax's plan to open weights “in the coming days,” while PurzBeats called out posts claiming people could already download them.

MiniMax's official H3 post buries the best technical detail: 2K output comes from in-context regeneration, where H3 regenerates its own low-res output using the original multimodal context instead of a separate super-resolution pass. The video-generation docs put hard limits on the workflow: 7,000-character prompts, 4 to 15 second outputs, up to 9 images, 3 videos, 3 audio clips, and 12 mixed files total. Artificial Analysis ranks H3 first for video editing with audio, while MatanCohenGrumi showed a Pika launch film built around 15-second audio slices.

What shipped

MiniMax describes H3 as a general-purpose multimodal generation model that understands text, images, video, and audio together, then outputs video with native stereo sound at up to 2K and 15 seconds in the official launch post.

The MiniMax API docs define three main modes:

  • Text-to-video from a prompt.
  • First/last-frame image-to-video from text plus one or two frame controls.
  • Reference generation from text plus images, video, or audio.

The same docs list MiniMax-H3 output at 768P or 2K, 4 to 15 seconds, with image, video, and audio reference inputs. MiniMax's model overview lists H3 at 24 fps.

Hailuo positioned the model for UI/UX, ads, and commercial work in one Hailuo_AI post, then followed with another Hailuo_AI post saying H3 is good at precise editing and control.

Omni Reference

Omni Reference is the part creators immediately grabbed: multiple character sheets, style refs, motion refs, and audio refs in one generation.

The hard limits from MiniMax's docs are concrete:

  • Reference images: up to 9.
  • Reference videos: up to 3 clips, 2 to 15 seconds each, 15 seconds total.
  • Reference audio: up to 3 clips, 2 to 15 seconds each, 15 seconds total, paired with an image or video.
  • Mixed input: 12 files total.
  • Prompt length: up to 7,000 characters.

The demos moved fast. In hellorob's Seedance comparison, H3 was the model that used all 9 reference images and produced the cheaper result. In egeberkina's fashion-film prompt, the generation used separate references for the woman, man, and horse, then layered choreography, kinetic type, art-pop audio, and motion graphics into a 15-second scene.

Pika's launch film used the same idea as a production spine: MatanCohenGrumi built one laundromat scene from 9 image references, including the parcel, outfits, faces, shoes, location, and dryer.

Cost and leaderboard pressure

MiniMax's pay-as-you-go pricing lists H3 at $0.13 per second for 2K and $0.09 per second for 768P, with 768P still in closed beta. The same page says audio input is free, the first 5 images are free, extra images cost $0.04 each, and input video is billed by duration at the output-resolution rate.

That makes a 15-second 2K clip $1.95 at list price, matching dustinhollywood's calculation. Topview's rollout claimed H3 was priced at 30% of Seedance 2.0 in MayorKingAI's post, and Artedeingenio's Topview post repeated the native 2K, 15-second, 30% framing.

Independent leaderboard data mostly matches the pressure creators felt on the timeline:

  • Video editing with audio: H3 ranked #1 with 1,130 Elo from 5,096 samples on Artificial Analysis.
  • Text-to-video with audio: H3 ranked #2 with 1,238 Elo from 5,740 samples on Artificial Analysis.
  • Image-to-video with audio: H3 ranked #3 with 1,184 Elo from 5,039 samples on Artificial Analysis.

The quality argument stayed split. In 0xInk_'s H3 versus Seedance 2.0 test, Seedance was still better overall for action scenes because H3 showed motion blur, but H3 was “significantly cheaper.” bilawalsidhu called Seedance 2.5 the priciest and “clearly the leader of the pack,” while BLVCKLIGHTai argued H3 was winning timeline volume through cost, inference speed, and working custom audio uploads.

Audio reference workflows

The audio feature landed as a workflow fix, not a novelty. Uncanny_Harry recorded dialogue, fed it as an audio reference with an image reference, and said H3 used the exact audio without requiring black video or prompt-written words, which he contrasted with Seedance 2.0.

Music-video workflows turned into 15-second slicing jobs. bennash's prompt used one source video for the character and 15-second song clips as Audio1, asking H3 to make the character dance and lip-sync in sync with the audio.

Pika's announcement film used the same beat-grid logic at launch. MatanCohenGrumi wrote that each scene got its own 15-second slice of the track, with the parcel acting as the repeating trigger and the body dancing “dead on the beat.”

The side tooling appeared immediately: bennash built Basic Slice, a free browser tool that cuts audio into 15-second or custom-length files client-side for H3-style lip-sync workflows.

Text, UI, and HUDs

Text rendering and interface motion became the most commercial-looking early lane. CharaspowerAI called out transitions, micro-interactions, camera movement, and text integration, while another CharaspowerAI test used a prompt where ocean spray compresses into the word “ASCEND.”

In 0xInk_'s ad example, Seedance 2 struggled with clean scrolling phone text on the first shot, so H3 handled that shot instead. 0xInk_'s reply said Seedance turned letters into Russian-looking characters in that use case.

Game UI also held up better than usual. koldo2k recreated gameplay and said H3 kept the HUD intact while reacting to the scene context, and AllaAisling's deck-builder prompt specified screen-locked cards, energy changes, damage numbers, an end-turn button, and legible text across a 15-second tactics sequence.

Production prompts

The strongest creator prompts read like shot lists, not image captions. AllaAisling's wardrobe prompt timecoded seven outfit changes from 0.0 to 15.0 seconds, specified the exact occluder for every change, and constrained fabric behavior, foot continuity, shadows, and crowd continuity.

Reverse physics got the same treatment. AllaAisling's “Undo” prompt breaks the film into five reverse-action shots, with bowl fragments reassembling, wine climbing back into a glass, smoke retracting into a wick, and reversed foley that resolves into one forward door latch.

Long prompt adherence was a public stress test. dustinhollywood said two H3 clips used prompts of 6,219 and 6,788 characters, with 24 prompted cuts in sequence and no added effects. The official 7,000-character prompt limit explains why these prompts fit inside one request.

Where it is live

H3 hit creator surfaces immediately:

  • Hailuo and MiniMax API: MiniMax's docs list MiniMax-H3 in the v2 video-generation API.
  • Topview: MayorKingAI posted H3 on Topview with native 2K, 15-second generations, and multimodal control, with the try link in the follow-up.
  • ComfyUI Cloud: PurzBeats tested reference-to-video at 2K, 8 seconds, for roughly 300 credits.
  • Magnific: Magnific launched H3 with up to 9 images, 3 videos, 2K output, and 5 to 15 seconds, with Magnific's follow-up pointing to availability.
  • Runway: H3 became available inside Runway's model picker alongside other frontier models.
  • Pika MCP: Pika announced H3 on the Pika MCP, and MatanCohenGrumi posted a prompt-tip thread for the launch film.
  • fal: fal's H3 page lists text-to-video, image-to-video, and reference-to-video endpoints.

The ComfyUI Wiki says H3 is available in ComfyUI through the official MiniMax partner node as an API-based integration, with native local support expected after weights ship.

Open weights caveat

MiniMax's official wording is future tense: the company says it plans to open the model weights “in the coming days,” subject to applicable laws and regulations. kaigani surfaced that line from the announcement.

The distinction mattered because creators started saying the weights had already dropped. PurzBeats wrote that announcing weights and dropping weights are different things, and that people were incorrectly telling others they could download H3 that day.

Local excitement is still real. hellorob called H3 a major step for open-source video users and said ComfyUI had optimized the model heavily, but hellorob's reply also described it as heavy while praising the quality.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 9 threads
TL;DR3 posts
What shipped2 posts
Omni Reference3 posts
Cost and leaderboard pressure5 posts
Audio reference workflows3 posts
Text, UI, and HUDs4 posts
Production prompts3 posts
Where it is live6 posts
Open weights caveat3 posts
Share on X