Skip to content
AI Primer
workflow

Creators report using MiniMax H3 with up to six image references

Creators report using up to six image references and short clips in MiniMax H3 for reference-led scenes. One workflow used the setup to make a Lego-style sitcom sequence for a presentation.

5 min read
Creators report using MiniMax H3 with up to six image references
Creators report using MiniMax H3 with up to six image references

TL;DR

  • A creator made a Lego-style Seinfeld sequence for a work presentation with 10-second H3 clips, six eBay reference images and a laugh track, according to thrownblown's Lego Seinfeld post.
  • Reference-led H3 work is arriving as a timed shot plan: techhalla's CRT template assigns five source images to five consecutive three-second beats.
  • One MCP workflow generated two 15-second scenes from prepared stills, then used the first clip's final frame as the next clip's start, as techhalla's two-clip workflow describes.
  • Start and end frames can hold a multi-shot sequence together in one generation, according to the frame-anchored workflow.
  • Native picture, narration and score can arrive in one render, as the octopus demo claims of a 12-second nature short.

On Runway's MiniMax H3 page, the model takes a text brief of up to 6,000 characters, plus up to nine images, three video clips and three audio takes, or a first and last keyframe. Glif's product page says project references persist across assets, while fal's H3 Max page offers five free 768p, native-audio generations per rolling 24-hour period in its sandbox.

The six-image Lego test

r/StableDiffusion

moderate effort H3 Lego Seinfeld Episode

0 comments

The creator said the afternoon project used a seed-hunter workflow, 10-second clips and as many as six images from a Lego Seinfeld set listed on eBay. They added the laugh track from soundboard sites.

The reported production limits were equally useful:

  • No scene took more than four generations.
  • Shots containing all four characters were more hit-or-miss.
  • A future version would use Qwedit to set the start image and reduce continuity errors.

The six images were a reference pack for a tiny cast and set, not a claim about H3's maximum input allowance.

Five-image style packs

In techhalla's CRT template, five Midjourney images define five RGB-split CRT characters. The H3 prompt assigns each character a three-second interval, preserves its crop, and restricts the transitions to scanline tears and reformations.

r/StableDiffusion

H3 - mystique skin is alive I2VA Prompt included

0 comments

SIR_NVAX_A_LOT's transformation prompt uses an exact opening frame, one slow push-in and a four-part 15-second body transformation. It locks the face and salon while the scales, clothing, coils and wings change on schedule.

The paper-character workflow follows the same division of labor. techhalla's paper-animation workflow used Nano Banana 2 for stills and H3 for animation, while techhalla's paper-character post describes the result as a deliberately simple stills-then-motion setup. The references establish material and character design; the timed prompt supplies the movement.

15-second continuations

techhalla first asked Grok to connect to Magnific through MCP, a route shown in techhalla's MCP demo and explained in techhalla's workflow note. Seedream 5 Pro generated the character stills, then H3 Max generated two 15-second scenes.

After the first scene, the creator fed its final frame into the next scene as the start frame. They reported the full sequence took under five minutes, while techhalla's 768p timing note puts a 15-second H3 Max clip at about 30 seconds at 768p.

A separate workflow took the editorial route: ozansihay generated four 15-second image-to-video clips, joined them in Premiere Pro with Morph Cut, then added color, halation and grain before upscaling to 1080p at 60 fps.

Start and end frames

Magnific said it set both the start and end frame before generation, with every scene change coming from a single output and no editor opened.

Its escalation notes reduce a multi-shot prompt to four rules:

  1. Number the shots and timestamp them.
  2. Hold one element constant across them.
  3. Escalate one variable, not five.
  4. Put the reveal in the final two seconds.

Video-to-video restaging

fabianstelzer's Glif workflow accepts a video up to 15 seconds and an edit description. In fabianstelzer's car-and-castle demo, the creator said H3 video-to-video moved a person into a car and then a castle while retaining the original audio and lip sync.

fabianstelzer's reply says that demo used no LoRA. Asked about render time, fabianstelzer's timing reply put it at a few minutes.

Text-only set pieces

References are optional. AllaAisling's H3 prompt turns one old match into a 15-second museum mystery: the flame reveals a vase on an empty shelf, then a missing painting, before both vanish as the match burns out.

A separate H3 Max demo from ozansihay shows a robotic arm picking up a small object and placing it into a container. The two examples use the same time-boxed format for opposite jobs, one a narrative reveal and one a controlled physical action.

Native sound

Magnific's octopus short pairs generated nature footage with a British narrator and score in a claimed 12-second pass. Its the lighthouse example instead combines a weathered character's spoken line with wind, room tone, a foghorn and a chair creak in 10 seconds.

the time-lapse example adds another format: a three-shot, 15-second coastal-city time lapse with narration written into the same prompt and piano beneath it.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 5 threads
Five-image style packs2 posts
15-second continuations3 posts
Start and end frames1 post
Video-to-video restaging3 posts
Native sound2 posts
Share on X