Skip to content
AI Primer
release

LTX-2.5 generates multi-shot video from audio and a start frame

A creator used LTX-2.5 with uploaded audio, a start frame, and an action prompt to generate a continuous sequence across several shots. They report stronger prompting and continuity across wide, medium, and close views.

3 min read
LTX-2.5 generates multi-shot video from audio and a start frame
LTX-2.5 generates multi-shot video from audio and a start frame

TL;DR

  • The demo's input stack was audio, a start frame and a written action prompt, as egeberkina's tutorial post documents.
  • One generated sequence can move through wide, over-the-shoulder, medium and close framing, egeberkina's workflow post says.
  • LTX-2.5 pairs open weights with a smaller distilled model intended for local runs, according to egeberkina's follow-up.
  • Speed claims span radically different setups: gokayfem's speed post says 15 seconds of video took 5.66 seconds, while KaisarasAR's Reddit post reports 10 to 20 minutes per 1080p generation on one RTX 5060 Ti.

The official LTX documentation divides the release into Fast, which reaches 4K, and Pro, which tops out at 1080p. Its model card describes multishot as continuity of character, environment, lighting, voice and visual style across cuts, while ComfyUI's workflow guide adds first/last-frame generation to the available routes.

Audio, a start frame, and an action prompt

egeberkina described a console workflow with four actions:

  1. Upload the audio.
  2. Add a start frame.
  3. Write what should happen.
  4. Generate.

Native multishot editing

egeberkina said the same generation carried a deliberate sequence of camera sizes rather than requiring separate clips.

  1. Wide shot
  2. Over-the-shoulder shot
  3. Medium shot
  4. Close-up

The LTX-2.5 model card makes the matching technical claim: native multishot generation keeps connected scenes in one pass.

Prompt load and fast action

egeberkina said in their follow-up that prompts can carry more detail without repeated rewrites, and that the smaller distilled model retains much of the quality.

They also reported stronger close-ups, more durable text and fewer breakdowns in fast movement in the earlier workflow post.

Thirty-second cut points

DavizCF7777 described a longer-form workflow in a reply: write the script, have an agent time it, then split it into 30-second blocks at safe cut points. Their rule was that no block ends mid-line or mid-action.

MatanCohenGrumi said in their post that 30-second launch videos can now be generated in one shot, without identifying the model or input setup.

Hosted speed and a local baseline

gokayfem's speed post claims a 15-second video in 5.66 seconds, but the post identifies neither the model nor the hardware.

KaisarasAR's Reddit post supplies a reproducible local counterpoint for LTX 2.5 Image-to-Video:

  • 1080p, 16:9 output
  • One RTX 5060 Ti with 16GB VRAM and 32GB system RAM
  • 10 to 20 minutes per generation
  • Frames from the original footage as start images, then interpolation in an editor
  • SeedVC for voice consistency
r/StableDiffusion

Realistic Breaking Bad | LTX 2.5 I2V

0 comments

Share on X