Skip to content
AI Primer
release

ComfyUI ships MiniMax H3 templates for 5 local video workflows

ComfyUI added MiniMax H3 templates for text-to-video, image-to-video, reference-to-video, first-and-last-frame, and in-place editing. Early Reddit tests reported errors, weak defaults, 66s generation on an RTX 5090, and license questions.

7 min read
ComfyUI ships MiniMax H3 templates for 5 local video workflows
ComfyUI ships MiniMax H3 templates for 5 local video workflows

TL;DR

  • ComfyUI shipped real MiniMax H3 support: the day-0 r/ComfyUI post lists text-to-video, image-to-video, first-and-last-frame, reference-to-video, in-place editing, 2K output, 15-second clips, and stereo audio generated with the video.
  • The local story has a 2K gotcha: MiniMax's Hugging Face card says H3-Context-IR and H3-Regenerate-2K are not included in the initial open-source release, while Comfy's native workflow docs describe local templates around H3-Base workflows.
  • Consumer hardware runs are already messy: scooglecops reported 167 seconds for a 608x352 clip on a 4070 and 64 GB RAM, while mamypokopants reported 66 seconds for a 5-second 864x480 clip on an RTX 5090.
  • The standout creator workflow is audio-conditioned reference video: PurzBeats ran 2K reference-to-video on Comfy Cloud, and hellorob said one generation used two reference images plus input audio.
  • H3 is being sold against Seedance on price: hellorob's comparison claimed H3 used nine references at one-third the cost, while 0xInk_'s comparison still put Seedance 2.0 ahead overall because H3 blurred more in action scenes.

MiniMax's launch post frames H3 as one omni-modal model for text, images, video, and audio, with 15-second 2K output and native stereo sound. Comfy's blog post hides the nerdiest bit in the middle: about 40% of the parameters were modulation weights that Comfy says it pruned into a lookup table. Comfy's docs put hard limits on reference work, up to nine images, three videos, and three standalone audio clips, while Basic Slice appeared almost immediately as a tiny browser utility for cutting reference audio into 15-second chunks.

ComfyUI templates

r/ComfyUI

Day 0 MiniMax Support for ComfyUI

0 comments

ComfyUI 0.30.0 adds MiniMax H3 to the template library, with model files hosted in Comfy-Org/MiniMax-H3. The three native example templates cover more than three practical modes:

  • Text-to-video: prompt-only video with native stereo audio.
  • Image-to-video: animate one input image.
  • First-and-last-frame: use the MiniMaxH3ImageToVideo node with first frame, last frame, or both.
  • Reference-to-video: condition on reference images, videos, or audio.
  • Shot editing and motion transfer: Comfy's launch post says a reference video can supply motion, camera movement, performance, or rhythm while subject and style come from other inputs.

The API version has its own ComfyUI partner-node docs: it runs on MiniMax servers, needs no local GPU, and bills each generated second to a Comfy API account.

Open weights, 2K caveat

MiniMax said on July 31 that it planned to open the weights "in the coming days," subject to laws and regulations. Comfy's August 3 post says the weights had dropped, but the full system is split across local and hosted pieces.

MiniMax's Hugging Face card lists three modules:

  • H3-Context-IR: hosted preprocessing for complex multimodal instructions, not included in the open-source release.
  • H3-Base: local audio-video generation at 768p.
  • H3-Regenerate-2K: in-context 2K regeneration, not yet open-sourced.

That timing gap created the first round of creator whiplash. PurzBeats' clarification separated announcing weights from dropping weights after people started telling followers they could download them immediately.

Local inference cuts

r/ComfyUI

Day 0 MiniMax Support for ComfyUI

0 comments

Comfy says it cut H3's local footprint with three moves:

  • Pruned modulation weights, about 40% of total parameters, into a functionally equivalent lookup table.
  • Shipped int8 convrot quantization.
  • Added custom kernels to reduce peak VRAM.

The result, according to Comfy's blog, is a 66% memory reduction: 123.6 GB in full precision to 42.5 GB for the smallest variants. Comfy also says dynamic VRAM offloading lets the model run on a GPU like an RTX 3060.

First local runs

r/ComfyUI

Minimax h3 12gb vram + 64ram

0 comments

Early local reports looked like day-zero open video usually looks: exciting, slow, and fragile.

  • One r/ComfyUI user hit an error while trying the MiniMax H3 template workflow.
  • Another first-run post said the default workflow produced a weak result without any prompt changes.
  • scooglecops generated 608x352 on a 4070 with 64 GB RAM in 167 seconds.
  • mamypokopants generated a 5-second 864x480 clip on an RTX 5090 in 66 seconds.
r/ComfyUI

MiniMax H3 feat Will Smith's Spaghetti

0 comments

The useful early read: H3 can run locally, but the first public timings are preview-resolution timings, not polished 2K production timings.

Reference audio

PurzBeats ran MiniMax H3 reference-to-video at 2K for 8 seconds and estimated the Comfy Cloud cost at about 300 credits. The follow-up prompt treated the attached audio track as an edit timeline:

  • Every cut lands exactly on a beat.
  • The visual style is a 90s anime title sequence.
  • Freeze-frame holds are specified as full-beat stops.
  • The title text must read "PURZ" and nothing else.

hellorob described another one-shot generation using two reference images and an input audio track. The prompt in hellorob's follow-up assigns <Picture 2>, <Picture 1>, and <Audio 1> to a two-cut comic-book sequence.

Audio references look like the breakout control surface. PurzBeats said attaching audio is optional but helps, and Uncanny_Harry said H3 used the exact recorded dialogue from an audio reference without needing a black video workaround or the spoken words written into the prompt.

Text, HUDs, brand elements

Creators immediately started testing the parts video models usually mangle: text, UI, and game interfaces.

  • hasantoxr called H3 the closest thing to After Effects for AI video, citing text, subtitles, brand elements, multimodal input mixing, and in-model compositing.
  • CharaspowerAI said one reference image was enough to push a 15-second shot and called out in-scene text handling.
  • 0xInk_ said several tests showed H3 beating Seedance 2.5 at handling text in motion.
  • koldo2k recreated gameplay and said H3 kept the HUD intact while the interface reacted to context.

The low-prompt counterexample came from Magnific: its prompt post used only "Nostalgic montage of different clips. Graphic motion layouts of friends having fun in a theme park." Hailuo_AI's reply amplified the text-in-motion angle, telling creators to try H3 and bring text to life in motion.

Cost and hosted surfaces

H3 also arrived through hosted creator tools before most people had time to configure local weights.

  • Artedeingenio said Topview had MiniMax H3 with native 2K, 15-second videos, multimodal generation, and pricing at 30% of Seedance 2.0.
  • hellorob compared H3 with Seedance 2.0 using nine reference images and said only one model used all nine at one-third the cost.
  • Magnific said H3 was already available on Magnific.
  • Magnific's demo said all clips and effects in its montage were made with the model, with only music added afterward.
  • Magnific's follow-up pointed users to try it on Magnific.

The price claims were not universal quality claims. 0xInk_ estimated 15 seconds at 1080p on Seedance 2.0 around $6 and 15 seconds at 2K on H3 around $1.50, but still found Seedance 2.0 better overall because H3 showed more motion blur in action scenes.

License questions

r/ComfyUI

MiniMax H3 License, Am I reading this right?

0 comments

The license thread started immediately because open weights do not automatically mean frictionless commercial use. MiniMax's Hugging Face card lists the MiniMax H3 Community License, says user submissions and enhanced prompts may be moderated, and names unlawful, pornographic, or third-party-rights-infringing content as categories that may be blocked.

The same card says the guardrails do not change licensee obligations under the community license. That is the part creators were already trying to parse before the first local workflow errors were solved.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 5 threads
TL;DR1 post
Open weights, 2K caveat1 post
Reference audio4 posts
Text, HUDs, brand elements4 posts
Cost and hosted surfaces5 posts
Share on X