Skip to content
AI Primer
workflow

Reference-led video workflows place consistent characters across scenes

Creators are using image references to place prepared characters in generated scenes and guide image-to-video shots. Higgsfield's Seedance 2.5 workflow also swaps people, backgrounds, or effects from supplied photos.

3 min read
Reference-led video workflows place consistent characters across scenes
Reference-led video workflows place consistent characters across scenes

TL;DR

Reference images are starting to behave like a small production department, handling casting, set dressing, and continuity in separate files. Higgsfield's Ad Multiplier skill says it preserves a source video's motion, framing, timing, aspect ratio, default audio, and untargeted text during variations. Its BITT Animation case study adds a useful production detail: generative VFX returns as a pass, then gets composited against the original plate.

The reference packet

techhalla's paper-cutout experiment starts by generating stylized character stills in Nano Banana 2, places those assets into scenes, then sends them to H3 for animation.

The examples divide a reference packet into distinct controls:

  • Character: Wan 3.0 receives face, build, and material cues from one sheet in a character-sheet test.
  • Location: An empty kitchen was generated once, then reused as reference for scenes set across three timelines in magnific's short-film breakdown.
  • Shot action: The prompt supplies camera movement, staging, and the action happening within the locked character and location.

Video-to-video swaps

Higgsfield's Seedance 2.5 V2V recipe begins with a reference video in its walkthrough:

  1. Upload the source video.
  2. Add photos of the elements to swap.
  3. Describe the change in a prompt.
  4. Generate.

The public demo frames that as a way to replace a person, alter a background, or reproduce an effect from one photo. A Kapwing explainer describes the wider Seedance 2.5 reference pool as up to 30 images, 10 video clips, and 10 audio clips in one generation.

Timed keyframes

magnific set start and end frames before generating a 15-second fish-market-to-flood sequence, with every scene change coming from one generation rather than an edit.

The formula has four constraints, according to the accompanying prompt notes:

  • Number the shots and timestamp them.
  • Hold one element constant across every shot.
  • Escalate one variable, rather than five.
  • Put the reveal in the final two seconds.

Runway's Hailuo 3 product page exposes the same choice in its interface: attach a reference bundle, or set a first frame and optional last frame for H3 to animate between.

Audio-preserving edits

Video-to-video workflows can retain an existing performance alongside the visual change. fabianstelzer's hybrid-editing tutorial morphs real footage in Glif while retaining its full audio.

In a separate car-to-castle experiment, fabianstelzer said MiniMax H3 preserved audio and lip-sync, while his follow-up identified the method as Glif without a LoRA. Runway allows H3 reference bundles of up to nine images, three video clips, and three audio takes on its model page.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 4 threads
TL;DR2 posts
The reference packet2 posts
Timed keyframes1 post
Audio-preserving edits2 posts
Share on X