MiniMax H3 gallery reports 15-second 768p render in 4:03 on RTX 5090
A MiniMax H3 workflow gallery reports a 15-second 768p render in 4:03 on an RTX 5090, while a ComfyUI implementation targets 24GB RTX 3090 Ti cards. Community tools centralize controls and preserve latent data for stitching longer runs.

TL;DR
- A community gallery reports a 15-second, 768p MiniMax H3 render in 4:03 on an RTX 5090, and BluePointDigital's post links that result to a growing controlled test archive.
- A 24GB-focused ComfyUI build uses INT8 ConvRot VDN weights, with NoMouse9610's 3090 Ti test reporting 10 seconds at 0.4 MP and eight steps in 2:08.
- A single ComfyUI graph now centralizes up to six image, three audio, and two video references, according to roychodraws's workflow post.
- H3 generations can be carried forward as
.mmh3packets for latent stitching and upscale passes, as EinhornArt's media-bundle post demonstrates with six 0.4 MP clips. - One environment workflow converts an equirectangular panorama into a short video reference, which jimtonyk's test says avoids the warped projection seen with the raw panorama.
BluePointDigital's benchmark gallery keeps a canonical prompt, seed, timing, and human review beside its cards, while its comparison methodology only accepts speed ratios for closely matched workloads. In dkackman11's MCP experiment, Claude assembled a 30-second sequence from a Z-Image reference, six H3 Ref2VA clips, MiniMax Music, and an edit pass without a prebuilt template.
RTX 5090 benchmark cards
A gallery & data of every Minimax H3 Workflow I've Tried. Best results are 15s @ 768p in 4:03 on a 5090.
0 comments
BluePointDigital reports its best current result as a 15-second, 768p render in 4:03 on an RTX 5090. The collection now presents 24 measured test cards, separating short compatibility checks from full-length workloads.
Its September 7 benchmark data says it added 15 distinct configurations and records Comfy execution totals. In a corrected VDN comparison, VDN was 5.04% faster, while the operator review preferred PDD plus Sage attention.
The project keeps timing separate from review of motion, identity, lip sync, artifacts, and transcript. That makes the 4:03 result a record for a specific graph, model state, and machine configuration.
24GB VDN stack
VDN-H3 on RTX 3090 Ti 24GB
0 comments
NoMouse9610 built a ComfyUI implementation around INT8 ConvRot VDN weights for an RTX 3090 Ti with 24GB of VRAM. The author tested both FL2VA and Ref2VA INT8 ConvRot models from 0.4 to 0.8 MP, and reports that 30-second videos run too, at a substantially slower pace.
The stated reference point is 10 seconds at 0.4 MP and eight steps in 2:08. The published VDN-H3 repository packages the 24GB-targeted implementation and example workflows.
One control panel
MiniMax Workflow Designed to be User Friendly for the Inexperienced User
0 comments
The community-maintained vLLM-Omni H3 recipe separates FL2VA, for text and first-frame generation, from Ref2VA, for image, audio, and multi-video conditions. roychodraws's graph turns those branches into one control surface:
- Up to six image references, three audio references, and two video references, with mixed types in one generation.
- First-frame and last-frame conditioning.
- Video continuation with overlap stitching, plus continuation-aware handling for source and generated audio.
- Forced audio from an audio reference or embedded video audio, with automatic duration trimming to the audio length.
- A final latent upscale and refinement pass that can run after the continuation has been stitched.
- Built-in RIFE interpolation, multiple LoRAs, sparse-attention settings, chunking, and other low-VRAM controls.
- Automatic reference routing based on the enabled image, audio, and video counts.
The downloadable all-in-one graph is designed to eliminate repeated rewiring when switching among these branches.
Latent packets
MiniMax H3 media bundle
0 comments
EinhornArt's package saves a generation's latent and context resources into an .mmh3 file, then uses them for stitching or latent upscaling. Its example joins six 0.4 MP generations before an upscale to 0.8 MP.
The mmh3_media repository describes each packet as holding the joint audiovisual latent alongside final video and audio, keyframes, references, masks, metadata, creator notes, and a disposable preview cache. The author lists lip sync, ControlNet, inpainting, and audio among work still in development, and says the package needs further polishing.
Four-step interactive scenes
Training a character LoRA for MiniMax H3 (4-step FastH3) for a realtime interactive scene - does a base-H3 LoRA transfer to the distilled checkpoint?
0 comments
One ComfyUI builder is using H3 for an interactive traffic-stop scene: viewers type dialogue, a character answers inside generated video, and five-second clips stay slightly ahead of playback through buffering. The post cites FastH3's four-step distillation at roughly six seconds for a five-second clip on four B200 GPUs, with a LoRA hook already present in the stack.
The builder also found that a stock ComfyUI conversion silently drops roughly 50 to_gate_compress tensors from VSA weights, producing noise. Dense weights converted cleanly in that test.
Four production questions remain open in the post:
- Whether a character LoRA trained on base H3 transfers to the four-step distilled checkpoint.
- Whether a character LoRA can stack with the distillation LoRA.
- Whether still-image training data holds identity during speaking motion, or T2VA clips are required.
- Whether anyone has run H3 with sequence parallelism across multiple datacenter GPUs.
Panoramas as reference video
Conistent(ish) environments with Minimax H3
0 comments
jimtonyk found that raw equirectangular panoramas only appeared coherent when the H3 shot was zoomed in, then warped under more complex direction. Environment stills also overpowered later prompts for scene changes in the author's tests.
The workaround renders the panorama into a conventional two-second reference video first. The post's conversion command was:
The author reports that 48 reference frames held the apartment without the projection warp and let characters interact with named objects in it. Later shots still required state to be re-described, including details such as a previously broken mirror.
Scene and shot interfaces
bennash is wrapping H3 Max in Basic Slop, a browser-based video studio. In the project's feature list, the interface is organized around projects, scenes, shots, references, generations, and local export.
- Scene-wide prompts establish shared direction, while each shot can add its own prompt.
- Images, video, and audio can be named with
@references, and each shot can inherit a selected set of scene references. - The UI can generate one shot, a scene, or a project, keep multiple takes, and export or import the project with its input media and generated video.
- The announced version uses a bring-your-own MiniMax H3 Max key, has no login, and saves projects locally in the browser.
The surrounding demos need close labeling: gokayfem's reply separately flagged one circulating project as not H3 Max focused.
Clip length and frame grids
Advise on how to do H3minimax longer static no cut clips?
0 comments
A creator attempting 30 to 60 seconds of a static dialogue scene reported that latent continuation preserved continuity while background colors, skin texture, and artifacts degraded by 30 seconds. First-and-last-frame generation and bridge clips retained the subject but introduced color shifts.
VisualRecording4960 took the editorial route of 27 H3 reference-to-video clips for a 3:30 music-video intro, then upscaled from 1.0 MP in DaVinci Resolve Studio in the music-video post.
bennash attributes Hailuo and fal duration drift to a 24 fps frame grid, frames % 17 == 5. The reported mapping is:
- 5 seconds becomes 124 frames, or 5.167 seconds.
- 10 seconds becomes 243 frames, or 10.125 seconds.
- 15 seconds becomes 362 frames, or 15.083 seconds.
The same breakdown describes reference audio as conditioning for voice and timing while H3 generates picture and sound together. Bit-exact audio passthrough remains a later post-production operation.