ComfyUI users tune MiniMax H3 for RTX laptop video generation
Reddit users reported MiniMax H3 running on RTX laptops after fp16 and Spectrum mods, with tests for reference video, audio, and 25-second outputs. Others warned that turbo LoRA and attention settings can hurt motion or speed.

TL;DR
- Local H3 reached a 7-year-old RTX 2060 laptop: the laptop report says fp16 and Spectrum mods brought 480p, 5-second clips to about 12 minutes.
- The ComfyUI speed stack is moving fast: Diabolicor's post points to merged comfy-kitchen attention, while local workflow packs now stack INT8, Sol-Attn, Spectrum, Lightx2v, Turbo-LoRA, Motion Context, and latent upscaling.
- Turbo LoRA has a motion tradeoff: DifficultAd5938's warning found 4 to 6 steps more dynamic than 8 steps, while one 3060 Ref2V run skipped Turbo because quality degraded.
- Reference control is the creative unlock: the Castlevania TV test preserved a reference clip as in-scene screen content with copied audio, and ai_artworkgen's process used Hailuo's 9 reference slots for a music video.
- H3's floor and ceiling are both visible: the RTX 5090 dino run reported 7.3 seconds at 1920x832 in 12 minutes, while the RTX 3060 editing test hit almost 100% RAM usage.
The MiniMax H3 model card defines H3 as an omni-modal system: text, images, video, and audio in, 24 FPS video with 32 kHz stereo audio out, 4 to 15 seconds, with a 2K regenerate path. Comfy's MiniMax page frames it as open weights plus partner nodes. The useful local bits live in the stack: Comfy-Org's repack prefers int8_convrot with PyTorch/cu130, Amduraznak's fp16 fix targets pre-BF16 GPUs, and Spectrum skips selected transformer evaluations by forecasting post-transformer features. The web app also changed names mid-rollout: Hailuo_AI's post says MiniMax Hub is now MiniMax Design, with free credits and annual membership discounts extended to Aug. 15.
RTX laptop floor
I'm very floored. Minimax H3 actually runs on a 2060 laptop. Community appreciation!
0 comments
One Redditor expected H3 to be impossible on a 7-year-old RTX 2060 laptop, then got the default Ref2VA workflow down to about 12 minutes for a 480p, 5-second clip after adding the fp16 fix and Spectrum.
The linked fp16 repo says H3 declares only bf16 and fp32 inference dtypes, so GPUs without BF16 fall back to fp32. Amduraznak's fp16 fix patches the overflow points that made forced fp16 produce black frames and reports about 11x speedup over fp32 fallback on a V100 test.
Local performance reports now read like workstation recipes, not a single minimum spec:
- RTX 2060 laptop: 480p, 5 seconds, about 12 minutes after fp16 and Spectrum mods, per the laptop report.
- RTX 3060 12GB with 16GB RAM: 0.9MP, 10 seconds, about 1 hour, with Turbo LoRA disabled for quality, per irmemon225's Ref2V post.
- RTX 3060 with 64GB RAM: 0.35MP, 5-second video editing at about 100 seconds per iteration, with almost 100% RAM usage, per the 3060 editing test.
- RTX 5060 Ti 16GB: 9:16, 0.6MP, 5 seconds, Turbo v4 600 LoRA, Euler beta, 8 steps, about 2 minutes per clip, per the character-knowledge test.
- RTX 5090: 7.3 seconds at 1920x832 in 12 minutes, per the dino test.
A DEV Community comparison reached the same boring but necessary conclusion: VRAM numbers alone are not comparable unless precision, workflow type, resolution, frame count, steps, cache settings, audio, offloading, and dependency versions line up.
ComfyUI speed knobs
Comfyui comfy-kitchen Attention Speed UP
0 comments
ComfyUI's own H3 path changed on Aug. 11. The merged comfy-kitchen attention commit added a ModelAttentionBackend node and a --use-ck-attention flag, with the commit warning that default CK attention might break some models.
The community instruction in Diabolicor's post was blunt: disable Sage Attention, Minimax Mem Eff Sage Attention, and Sol Attention before testing CK attention, because only one attention backend should be active.
Javawock's workflow pack shows how quickly the local H3 graph turned into a tuning board:
- INT8: lower VRAM use and faster inference.
- Sol-Attn: Triton attention acceleration.
- Spectrum: sampling and inference acceleration.
- Motion Context: extra temporal motion conditioning for FLF and R2V.
- Lightx2v LoRA: reduced sampling steps.
- Turbo-LoRA: experimental low-step acceleration.
- Latent Upscaler: post-generation latent video upscaling.
Spectrum is more than a magic speed button. Its GitHub README says it forecasts post-transformer hidden features with Chebyshev ridge regression on selected future solver steps, while the native H3 output heads, video reconstruction, audio reconstruction, sigma mapping, and return structure still run every step.
Fragility showed up fast. one Spectrum user said Spectrum worked the day before, then returned to normal rendering times after a ComfyUI update.
Turbo LoRA local minima
Don't add too many extra steps to 4-step turbo lora
0 comments
Turbo LoRA's weirdest finding is that more denoising can mean less motion. With the 4-step Turbo LoRA, DifficultAd5938's test found the 4 to 6 step preview had more motion dynamics and closer prompt adherence than the cleaner 8-step result.
The same post said 6-step versus 8-step runs could fall into a bad local minimum where motion collapsed and outputs repeated despite prompt and seed changes.
Other users hit the same speed-quality boundary from different angles. one Ref2VA question said 4-step Turbo worked for image-to-video but turned reference-to-video into a mess; irmemon225 skipped Turbo on a 3060 Ref2V run because it degraded quality.
Reference video and copied audio
Every time I wonder if Minimax can do something, it can. You can have a character watch a full video clip with audio.
0 comments
The official H3 model card says Ref2VA supports up to 9 images, 3 video clips, and 3 audio clips, with 12 total files across input types. Audio has to be paired with image or video input, according to MiniMax's model card, not used alone.
The Castlevania test used that input structure for a compositing job that used to require a timeline:
subject_definitionstext-defined the viewer on the couch.- A reference clip was constrained to the TV screen, not the full frame.
- The reference audio was reused as the complete final soundtrack.
retention_analysisseparated character appearance, TV-screen video preservation, and audio copying.- The shot block locked an over-the-shoulder camera while preserving the TV content through the clip.
The open problem is voice cloning from default workflows. one reference-audio question and another voice-cloning question both came from users who could not make generated characters sound like the supplied reference.
15-second music-video chunks
The cleanest production recipe came from ai_artworkgen's gothic trap video. The final stack was Midjourney and Magnific for images, Suno v5.5 for the track, Hailuo H3 for video, and CapCut for assembly, according to the wrap post.
The Hailuo part was built around 15-second audio chunks:
- Split the song into 15-second segments.
- Use Hailuo's 9 total reference slots across image, video, and audio.
- Include 1 character sheet.
- Include 1 audio clip, max 15 seconds.
- Add up to 7 environment reference images.
- Tag each asset with
@at the start so the prompt can refer to it later. - Direct camera cuts and performance traits in the prompt.
- Edit the best parts of generations in CapCut.
- Recut lyrics when a 15-second split lands in the wrong place, then rerun the clean segment.
That is the most practical creator workflow in the evidence pile: model generation as shot harvesting, not one perfect pass.
VFX tests
Creators immediately used H3 as a VFX toy, not just a talking-head generator.
The tests clustered around physical gags, creature reveals, gravity failures, teleports, and stylized character motion. That mix explains why H3 discourse spread beyond local-diffusion setup threads so quickly.
Seedance comparisons
MiniMax H3 kept getting compared to Seedance because creators were testing reference control, audio, cost, and motion at the same time.
fabianstelzer's first H3 notes in the heyglif test put quality roughly on par with Seedance 2, praised first-frame and last-frame handling, and called H3 a first choice for image-to-video alongside Grok and Flux 3 because it had no issue using AI-generated faces.
Cost was the sharper pain point. BLVCKLIGHTai's Seedance post put a 30-second Seedance 2.5 prompt at 720p at $14.90, while BLVCKLIGHTai's follow-up called H3 the clear winner for that specific pain point.
gokayfem's reply in a model-selection thread split the tool map this way: H3 for open weights, control, and sub-30-second 768p generations on fal; Flux 3 for realistic scenes; Seedance 2.5 for longer generations with many references.
Motion and continuity failures
H3 looks great in still frames, but the motion seemed falling apart (how do u think
0 comments
[H3] 1MP - 25stest - Rock Painting
0 comments
Failure reports clustered around motion, continuity, and over-literal props.
Best_Candidate_9060's test said H3 produced solid character, material, and static-shot quality, but fire behaved unnaturally, creature motion felt off, and transitions between shots lost continuity.
A longer local experiment exposed the editing problem. kaigani's follow-up said the 7m47s sci-fi short was running locally on a 4090 at 412p, but position continuity still needed work and a floating head escaped Claude's review.
The 25-second rock-painting test added a different kind of stubbornness: Moarkush said H3 got a Bob Ross-like voice close, failed the image, and would not get rid of a house-painting brush, then upscaled with Topaz Starlight Precise 2.5 and Apollo to 60p.