Skip to content
AI Primer
workflow

Seedance 2.5 supports 30-second Dreamina clips with reference-heavy workflows

Dreamina tests show restaurant, war, and aerial-action clips built from longer prompts and many references. Users cite better continuity, 720p output, higher prices, and occasional unwanted music.

8 min read
Seedance 2.5 supports 30-second Dreamina clips with reference-heavy workflows
Seedance 2.5 supports 30-second Dreamina clips with reference-heavy workflows

TL;DR

ByteDance's official Seedance 2.5 post says the model accepts 30 images, 10 videos, and 10 audio clips in one pass. Dreamina's launch page adds timestamp-specific editing and a regional subscriber rollout. The fun part is in the field tests: a restaurant meltdown, a rubber-chicken raid, a dragon film pipeline, and a 42-reference survival show.

What shipped

ByteDance says Seedance 2.5 extends single-pass generation from 15 to 30 seconds and supports multi-round extension for longer videos. The same official post says the model improves shot transitions, scene changes, image quality, audio, and motion.

The official feature list is short and creator-shaped:

  • 30-second audio-video clips in one pass, according to the Seed team post.
  • Multi-round extensions that append new 30-second sections while preserving characters, environments, and pacing, according to the same Seed team post.
  • Up to 30 images, 10 video clips, and 10 audio clips as references in one pass, per ByteDance's official announcement.
  • Timestamp-level editing for targeted audio and video changes, plus green screen, camera perspective, and reference-based editing, per the Seed team post.
  • Dreamina's own launch page says Seedance 2.5 is rolling out to Dreamina subscribers aged 16+ in Europe, Asia, the Middle East, and South America.

A third-party PiAPI spec lists text_to_video, first_last_frames, and omni_reference modes, with 1 to 30 second outputs at 480p or 720p. PiAPI's API page prices 720p at $0.60 per output second, which puts a 30-second 720p API run at $18.

Shot-list prompts

The 30-second demos that traveled fastest read less like prompts and more like miniature shooting scripts. The useful leap is temporal control, because creators can now specify emotional turns, camera moves, sound design, and reversals inside one generation.

The clearest prompt structures in the evidence use timed beats:

  • Restaurant acting test: ozansihay's prompt breaks an 11-second Obsession inspired scene into four connected shots, including vacant eye beats, panic escalation, lip-sync, ambient restaurant sound, and a failed recovery smile.
  • Aerial war spectacle: Artedeingenio's prompt maps a 30-second fighter-jet sequence into six five-second blocks, with cockpit inserts, wing-mounted perspectives, orbiting moves, and a hard cut to black.
  • Roller-coaster gag: techhalla's prompt freezes time after a wig flies off, rewinds the accident, and resolves the scene with chewing gum holding the wig down.
  • Tactical raid: ozansihay's test says the model missed the peephole instruction and added an extra door, but preserved the camera path, slow-motion flashbang beat, action rhythm, and sound-design structure.
  • Anime fight: Artedeingenio's anime prompt uses 0 to 5 second blocks for clash setup, mid-air impacts, speed ramps, and a final slow-motion strike.

kaigani pushed density instead of drama. kaigani's café test counted about 39 shots in 30 seconds, then said the result was still far from a 120-frame burst.

Reference-heavy workflows

techhalla's CapCut workflow starts with three assets: the environment, the main character, and the bike with trailer. The second step adds the last 15 seconds of the previous generation as a reference, plus the character sheet and bike, to preserve continuity.

The workflow pattern is becoming visible:

  • Start with a small reference pack for the first 30-second generation, as techhalla's first-clip step describes.
  • Add the end of the previous clip back into the next generation, as techhalla's continuity step shows.
  • Cut 15-second chunks that contain the elements needed to lock continuity, then reload those references, according to techhalla's repeatable step.
  • Use all the reference slots when the model needs a world, not just a character; techhalla's survival-show prompt uses 42 images across islands, boats, gear, maps, parachutes, clothing, faces, hands, and boots.

PJaccetturo's Nexus workflow is closer to indie production. The thread starts with a story meeting, turns the transcript into screenplay form, builds a shot list, collects look-development images, organizes character and scene ingredients, then runs Seedance 2.5 for 30 seconds of coverage.

PJaccetturo's money-saving move is anchor frames: PJaccetturo's anchor-image step says Theo screenshots useful frames from the first generation, uploads them back, and makes the next sequence revolve around those anchors. That is the most production-brained trick in the launch pile.

Smart Edit and video-to-video

HalimAlrasihi's Smart Edit test uses an existing dragon-fire clip and replaces the fire with three alternatives: wasps, ice and frost, and flowers. HalimAlrasihi's prompt reply says the ice version used a short instruction that kept the original motion and physics.

The UI evidence puts Smart Edit inside the same menu as Omni reference, first and last frame, multiframes, and Long video beta. HalimAlrasihi's menu screenshot is the clearest view of that surface.

venturetwins pointed to a Timberland-style ad where the shoes were swapped while the rest of the shot stayed aligned. The model edit surface may be more valuable for ads than yet another pretty clip generator.

Acting and voice

Seedance 2.5's acting demos focus on micro-expressions and dead-air timing. ozansihay's restaurant test asks whether AI acting can replace real acting after recreating the restaurant-breakdown cadence from Obsession.

venturetwins tested a different version of performance capture: one photo, one talking clip, and office photos produced a talking a16z office walkthrough. In venturetwins' voice-reference reply, the prompt was just: “use the voice from the video as a reference.”

The language results were mixed. ozansihay's Turkish speech test says Seedance 2.5 failed to directly generate Turkish speech like the previous model, while ozansihay's later Turkish test got a Turkish-speaking character with one word wrong.

Limits, cost, and competition

The launch has the classic AI-video tradeoff: better continuity, higher burn rate. techhalla's early-use reply says the model felt as good as 2.0 but longer, useful with up to 23 references, limited to 720p, and more expensive.

Reported constraints clustered around five points:

Hailuo H3 became the immediate comparison point. 0xInk_'s text-in-motion note says MiniMax H3 handled moving text better, and 0xInk_'s cost comparison puts H3 at about $1.50 for 15 seconds at 2K versus roughly $6 for 15 seconds on Seedance 2.5 at 720p upscaled to 2K.

Several creators still preferred Seedance 2.0 for specific work. icreatelife's model preference says 2.0 remained their favorite, and icreatelife's animation reply says 2.0 handled their animation style better.

Access routes

Seedance 2.5 is not arriving through one surface. ByteDance's Seed team post says it is rolling out on Jimeng AI, Doubao Pro, and other platforms, with API access coming soon through BytePlus ModelArk.

The ecosystem rollout visible in the evidence includes:

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 8 threads
TL;DR3 posts
What shipped6 posts
Shot-list prompts6 posts
Reference-heavy workflows6 posts
Smart Edit and video-to-video3 posts
Acting and voice3 posts
Limits, cost, and competition11 posts
Access routes8 posts
Share on X