Creators use 3D and image references to guide AI video shots
Creators are using stills, simple Blender previs, and reference clips to steer composition, character appearance, materials, and camera framing in video models. The examples span Seedance 2.5, Hailuo, and MiniMax H3.

TL;DR
- One reference can now carry a specific production job, as koldo2k’s H3 workflow freezes the camera while changing only materials.
- Simple 3D previs now doubles as a camera plan for video generation, according to HalimAlrasihi’s Seedance setup.
- Character sheets are carrying complete fashion-editorial layouts, as magnific’s fashion workflow turns four looks into moving reference-driven clips.
- Strict sequence logic remains a separate challenge: GlennHasABeard’s cat test reported a missing cat when an exact five-part order had to hold.
MiniMax’s H3 guide explicitly lets a reference set carry character, motion, camera, style, voice, or edit-rhythm information. Runway’s Seedance guide gives image, video, and audio inputs distinct roles, while Magnific’s guide calls image-to-video its controlled route for character animation and product shots.
3D previs and shot lock
koldo2k built a Lego version of a scene in GPT-6 Astra to test shots and camera moves, then fed that reference video to MiniMax H3. The instruction preserves opening framing, angle, height, distance, rotation, speed, duration, and screen positions, while replacing toy surfaces with limestone, steel, hide, grass, and mud over repeated passes.
The reference is functioning as an animatic, not merely a look reference. A Runway walkthrough starts with three-panel character sheets, an annotated blocking reference, and 20 static coverage shots before binding selected assets into a Seedance 2.5 prompt.
Reference assignments
MayorKingAI’s 30-second demon fight assigns each image a different responsibility. The prompt locks four layers:
- Cast and scale: the hunter, the oni, their faces, wardrobe, weapons, and a 2.5-to-1 height relationship.
- Place: Kyoto architecture, perspective, palette, pagoda, and cherry blossoms.
- Render: a painterly concept-art look held through movement, effects, and camera changes.
- Action continuity: timed choreography, readable impacts, persistent environmental damage, and exactly two characters.
The distinction between identity, location, rendering, and motion is the useful part. The model receives named constraints for each production variable instead of one general image prompt.
Character sheets and episodic casts
magnific’s fashion experiment starts with four-view character sheets: front, side, back, and portrait. Those sheets feed five-second editorial collages that keep a whole-body model in the center while the accessories and garment details move around it.
The close-up prompt assigns four synchronized inset views, face, upper-body fabric, hand or accessory, and shoes or hem. Colored borders and connector lines turn the motion piece into a product-detail layout rather than a conventional fashion film.
The same company applied the system to Sunny Six, an animated band designed as a three-chapter mini-series. magnific’s Sunny Six thread begins with a character sheet and group portrait, then preserves the six-member cast across a garage rehearsal, a sunflower-field performance, and a busking episode.
Counts, order, and visible proof
GlennHasABeard’s five-cat test gives each reference a name and position, then asks for one cat at a time, two seconds each, from left to right. He wrote that the model can lose an animal when five references must also appear in a stated order.
His box test makes the quantity visible over time: one cat joins every three seconds, and four heads must show above the rim at 12 seconds. Early takes stopped at two, GlennHasABeard’s box test reported. In both experiments, the count becomes a timed visual condition rather than an abstract cast description.