Creators report MiniMax H3 Max makes 15-second clips in under 15 seconds
Creators report that MiniMax H3 Max can produce a 15-second clip in under 15 seconds. They say stitching video and audio becomes the bottleneck; one Fal-agent workflow generated a 67-second H3 Max and Nano Banana Pro clip in two minutes.

TL;DR
- H3 Max is fal’s post-trained MiniMax H3 variant, and egeberkina’s demo reported a 15-second clip generated in under 15 seconds.
- The resolution trade-off is sharp: ozansihay’s matched-prompt test logged 5.17 seconds and $0.7125 for 15 seconds at 480p, versus 17.32 seconds and $0.855 at 720p.
- Video rendering has become faster than the surrounding assembly work, as gokayfem’s workflow put Krea 2 Turbo, ElevenLabs music, and H3 Max together and identified video stitching and audio merging as the bottleneck.
- Start and end frames are becoming the edit points: techhalla’s MCP workflow generated two 15-second H3 Max scenes, then used the first scene’s last frame to begin the next.
fal’s launch post calls H3 Max a post-trained MiniMax H3 optimized with its inference stack, and claims a five-second video in under three seconds. The original MiniMax H3 model card describes a multimodal base model that can produce up to 15 seconds of video with native stereo audio. An independent local comparison reported a 2.5-second 768p five-second render, while finding the quality winner less clear-cut.
480p and 720p
The new speed is most legible when wall time falls below screen time.
Ozan Sihay’s identical-prompt tests produced these figures:
- 15 seconds at 480p: 5.17 seconds to generate, $0.7125.
- 15 seconds at 720p: 17.32 seconds to generate, $0.855.
Sihay also reported visible detail loss at 480p, plus artifacts in complex human movement, hands, fast camera moves, and small scene details. He found Turkish dialogue less stable than English. A separate 720p 15-second sequence in ozansihay’s follow-up took 16.24 seconds and cost $0.855.
The reports span several timing bands: egeberkina posted a 15-second clip in under 15 seconds, chrisfirst logged 6.75 seconds for nine seconds of video, and techhalla’s 768p demo claimed 30 seconds in one minute. At the low-resolution extreme, gokayfem’s 480p demo said its clip arrived in one second.
Assembly
Gokay Fem combined Krea 2 Turbo for visuals, ElevenLabs for music, and H3 Max for video. The wait moved downstream to joining video and audio.
A 67-second clip in gokayfem’s Fal-agent post used H3 Max and Nano Banana Pro from a single prompt, with the whole clip reportedly generated in two minutes. Techhalla’s separate Magnific MCP run had Grok create Seedream 5 Pro stills, generate two H3 Max scenes, and pass the first scene’s last frame into the next scene’s start frame. techhalla’s setup put 15-second 768p renders at about 30 seconds and the complete sequence under five minutes.
Agents are already wrapping that pipeline. fabianstelzer’s Glif skill turns a script, idea, or prompt into an explainer-motion-graphics video, while the attached demo uses a Vox-like cutout style for an LLM explainer.
Shotlists
Magnific published five H3 Max examples with their full prompts. Its prompt notes distilled the multi-shot structure to four moves:
- Number the shots and timestamp them.
- Hold one element constant through the sequence.
- Escalate one variable rather than several.
- Put the reveal in the final two seconds.
The fighter example maps a 15-second sequence into five three-second beats: rising from a crate, a fist flurry, a locker kick, a haymaker into a hanging light, then a tight face shot before a hard cut. Every beat specifies framing, movement, sound, and the ending condition. Techhalla summed up the altered labor split in a follow-up: more time went into the prompt than generation.
Start and end frames
Magnific set a start frame and an end frame before rendering its fish-market-to-flooded-city sequence, then said every scene change came from one generation without opening an editor. The accompanying prompt uses five timestamped three-second shots, from a fish vendor to an upside-down flooded boulevard.
This sits alongside three creation modes in ozansihay’s test: start-end frame, image-to-video, and text-to-video. The clip tests each with Turkish dialogue, placing frame-pair generation alongside direct prompting rather than treating it as a separate finishing step.
Native audio
Techhalla’s paper-cutout music workflow used Nano Banana 2 for stills and H3 for animation, then uploaded and trimmed audio through Magnific’s media extractor to lock the video to music. The thread includes character references, scene references, an animation specification, and the audio-node setup.
H3’s video-to-video route can preserve source performance audio. fabianstelzer’s demo moved a person from a room into a car and then a castle while retaining the original audio and lip sync. A different 2.5-minute music video from ai_artworkgen combined Nano Banana Pro and GPT-2 images, Hailuo H3 video, Suno sound, and a CapCut edit.