Reference frames preserve continuity across AI video shots
Creators are using character sheets, prior clip endings, and anchor frames to maintain identity, lighting, and shot continuity across generated video. Reported workflows pair reference assets with Seedance and game-camera setups.

TL;DR
- Continuity is moving upstream into a reference package: techhalla's workflow builds character sheets, then carries the previous clip's ending into the next generation to retain raccord, lighting, and atmosphere.
- A source video can be a selective anchor: magnific's camera-angle prompt keeps a subject, scene, and performance while explicitly replacing the original framing.
- Game-footage illusion comes from persistent camera and interface rules, as _VVSVS's game-style short demonstrates with a rigid follow camera and screen-locked HUD.
- Voice is now part of the continuity brief: Artedeingenio's Thermopylae experiment combined three generated clips with an unedited voice-over.
DomoAI's August 19 announcement says its Omni Reference accepts up to 30 images, 10 video clips, 10 audio files, and first and last frames. A Kapwing storyboard tutorial takes a character reference through a 15-panel storyboard before video generation.
Identity contracts
The reference pack in techhalla's workflow has three jobs:
- Character sheets, generated in Seedream 5 Pro, define every creature before animation.
- Character and optional environment references enter Seedance 2.5 with the scene prompt.
- The final seconds of the prior clip become the continuity reference for a bridging shot.
A single-image version can be this strict. In underwoodxie96's selfie prompt, one reference governs identity, wardrobe, lighting, and composition across four timed micro-sequences, while the negative constraints forbid camera shifts, facial drift, flickering skin detail, and changing accessories.
The same sequence appears in magnific's animal workflow, which builds character sheets in Seedream 5 Pro before using them as Seedance 2.5 references. starks_arq's short-film report also claimed five minutes of coherent material from two frames, and described organizing the project around characters, script, tags, voice, and consistent expressions.
Anchor video
Magnific's re-angle template assigns the incoming video a precise scope: it defines the woman, Tokyo crossing, performance, timing, wardrobe, and weather. The prompt then says not to inherit camera position, framing, or shot size, leaving the new camera description as the variable.
That division turns a video reference into a continuity anchor with one deliberate point of change. It can preserve a moment while producing a different coverage angle.
The control remains imperfect. mrjonfinger said Seedance 2.5 tends to alter input camera motion, parallax, and performance, and later said the team was working on possible solutions.
Gameplay grammar
_VVSVS built a game-looking short with Midjourney 8.2 for character and world images, Seedance 2.5 for shots, and Topaz Astra for upscaling. The prompt turns a generated sequence into gameplay by specifying a system rather than merely naming a genre.
_VVSVS's detailed game-camera prompt divides that system into fixed rules:
- Medium: screen-recorded gameplay, flat exposure, deep focus, no film grain or cinematic depth of field.
- Camera: rigid follow distance, heavy bob, snap turns, and space ahead of the player.
- HUD: health, ammo, minimap, crosshair, ability icons, and hit markers remain locked to the screen.
- Asset roles: separate image references define the player, enemy soldiers, and skyscraper-scale machine.
- Time: bullet time is confined to a named window in shot two, while the other shots stay at real-time speed.
- Continuity: the coat, helmet, rifle, enemy design, palette, night setting, and HUD positions stay unchanged through six hard-cut shots.
In _VVSVS's cost reply, the creator put the project at about $100 including upscaling, excluding hourly labor.
Voice across cuts
Visual reference packs do not automatically settle narration. Artedeingenio's Thermopylae experiment said its three separately generated clips retained one voice-over without an edit.
A broader out-of-the-box comparison points to model choice as another variable: fabianstelzer's 20-model test named Kling 3 Pro and SD2, but not SD2 Mini, as strongest for cross-scene character voice consistency, with Flux 3 close behind.