MiniMax Design turns two references into an FPV-style transition
A creator used MiniMax Design to analyze two meme references and plan an FPV-style transition between them. The agent produced an H3 prompt specifying opening and closing frames for one continuous move.

TL;DR
- Two supplied meme references become a single FPV-style transition in CharaspowerAI's experiment, where MiniMax Design analyzed the sequence and wrote an H3 prompt.
- The brief can stay rough, covering references, movement, vibe, and transition, while CharaspowerAI says the agent handles the H3 prompt engineering.
- The starting and ending images become exact worlds, and the shared H3 prompt scripts a physical route through the space between them.
- The run occupies H3's 15-second ceiling: the generated prompt specifies that duration, while the H3 API reference lists a 4-to-15-second range for first-and-last-frame generation.
MiniMax Design's workflow page says its main agent reads a brief, decomposes the task, and selects a model. MiniMax's H3 launch post says H3 takes text, image, video, and audio in a unified context; here, the prompt turns a pointing gesture toward a television into the route between two still-image worlds.
Two reference frames
CharaspowerAI said they gave MiniMax Design two meme references and requested one seamless cinematic FPV move. The agent analyzed the idea, broke down the sequence, and produced the H3 text.
That creator's follow-up describes four elements supplied in the brief:
- The reference images
- The intended movement
- The visual vibe
- The transition logic
Four timecoded camera beats
The agent's 15-second sequence maps to H3's named first-and-last-frame input mode in the official API documentation. It assigns the middle of the shot four jobs:
- 0 to 3 seconds: Hold close to the first image, then drift forward with a slight side arc and align to the seated figure's pointing finger.
- 3 to 7 seconds: Curve around that figure, retain foreground-to-background parallax, and reveal the television.
- 7 to 11 seconds: Accelerate toward the screen while preserving its reflections, glass depth, and depth-of-field falloff.
- 11 to 15 seconds: Cross the television plane, match the screen image's perspective and lighting, then decelerate into the second frame.
Continuity constraints
Alongside the camera timestamps, the prompt locks down five continuity systems:
- First-frame identity: Clothing, pose, smoke, props, room geometry, textures, and warm lighting remain fixed.
- Spatial motivation: The pointing finger becomes the camera's guiding vector, making the television the target inside the first scene.
- A live destination: Before entry, the tuxedoed figure inside the screen is assigned breathing, blinking, eye shifts, small hand adjustments, glass movement, and sliding reflections.
- Physical surface detail: The television needs room reflections, bezel contrast, and perspective distortion as it fills the frame.
- Failure conditions: The text explicitly excludes cuts, glitches, teleportation, jumpy motion, wobble, abrupt zooms, and transition artifacts.
The result reads like a shot specification, with camera route, set continuity, and character performance in the same prompt.
MiniMax Design's desktop canvas
The product page describes a broader production pipeline beyond this single prompt: understand creative goals, decompose the task, launch copy, image, video, and audio agents in parallel, then merge, edit, and export through one canvas. It also says assets auto-save locally and can export to professional tools from the app.
The MiniMax Design download page lists macOS 13 or later, including Apple Silicon and Intel Macs, plus Windows 10 or later on x64. Its current desktop interface surfaces H3 and preset categories including cinematic intros, music videos, brand advertising, and UI motion.