Skip to content
AI Primer
workflow

Seedance 2.0 creators add depth maps to steer framing and motion

Seedance 2.0 workflows now pair style references with depth maps from images or footage to control framing and motion. Examples combine ComfyUI, Midjourney SREFs, character sheets, and Runway prompts.

8 min read
Seedance 2.0 creators add depth maps to steer framing and motion
Seedance 2.0 creators add depth maps to steer framing and motion

TL;DR

Dreamina's official Seedance 2.0 page says the model can take images, videos, audio, and text, with up to 9 images, 3 videos, and 3 audio clips per project. Magnific's Seedance page is advertising 4K, image/video/audio references, camera control, and multi-shot character consistency. ComfyUI has a Depth Anything V2 workflow for turning video into temporally stable depth maps, and Curious Refuge tested the same idea: replace raw reference footage with grayscale depth so motion survives and distracting visual style drops away.

Depth-map storyboards

_OAK200's move was clean: use the reference image for tone and final look, then let a depth-map storyboard control composition and camera framing in the demo thread. The split is the whole trick.

A normal storyboard carries style data. According to _OAK200's explanation, the depth-only version keeps only:

  • Camera placement
  • Subject scale
  • Foreground
  • Midground
  • Background
  • Silhouettes
  • Spatial relationships

His system prompt generates a 3×3 depth-only storyboard with exactly nine panels, a coherent visual event, fixed panel order, continuity checks, and a rule that brightness represents distance rather than lighting in the full prompt. The extraction prompt asks Nano Banana or GPT Image 2 for a physically accurate grayscale linear depth map where white is nearest and black is farthest in the depth-map step.

The finishing stack has three inputs: the original reference image for visual style, the depth-map storyboard for shot composition, and character sheets for identity consistency in _OAK200's final step. That is composition-first AI filmmaking, with the prettiness pushed to a separate layer.

Reference-video depth control

PurzBeats showed the same control idea with motion footage instead of a storyboard. His six-step workflow is short enough to steal as a production note:

  1. Reference video to depth map
  2. Depth map to Seedance 2.0 as the reference video
  3. Add reference image
  4. Set duration, resolution, and aspect
  5. Pass audio from the reference video to the final
  6. Run it

He later said he used Depth Anything 2 in ComfyUI in one reply, and that the input footage was driving the control in another reply. The Depth Anything V2 GitHub repo includes video depth scripts, and its notes say Video Depth Anything can generate consistent depth maps for long videos.

PurzBeats also linked a Comfy Cloud version of the workflow for hosted ComfyUI use. The ComfyUI workflow page describes the same target output: a temporally stable depth map converted from video.

Prompt timelines and shot grammar

Creators are writing Seedance prompts like shot lists, not like image prompts. Artedeingenio's aviation example starts in Midjourney with aeroplane --raw --sref 3263528706 --v 7, then gives Seedance a 15-second sequence split into 0-3s, 3-6s, 6-9s, 9-12s, and 12-15s beats, plus radial engine startup, cockpit vibration, wind, radio static, and orchestral sound design in the full prompt.

AllaAisling's Runway prompt uses the same grammar for a swamp chase: camera moves, obstacles, physics, and escalating danger as a beat list in the airboat prompt. Her orbital collapse prompt makes the structure even sharper, ending with “Rhythm collapse = danger” after a sequence of collision, cascade, near-crush, timing gamble, and last-second slip beats in the orbital prompt.

Common pattern across the stronger prompts:

  • Medium lock: live broadcast, 35mm film, smartphone footage, IMAX, anime, clay, or 1980s video.
  • Identity lock: one reference image, character sheet, or explicit “sole visual reference.”
  • Time lock: second-by-second shot blocks.
  • Camera lock: tracking, rack focus, orbit, handheld, wide, insert, or top-down.
  • Audio lock: foley, vocals, ambient sound, music, or native dialogue.

tranmautritam's jungle-temple prompt compresses that into “15s / 6 shots / 16:9,” with a specific character trio, dialogue per shot, and a final treasure gag in the single-prompt example.

Character consistency stacks

WordTrafficker described the low-randomness version as three inputs and one prompt: create characters, create a location in the same style, ask ChatGPT or Claude for a 15-second Seedance prompt using both images, generate 2-3 variations, then cut them together in the Viking girls workflow. “The less random your inputs, the less random your video” is the closest this thread gets to a rule.

Artedeingenio is using Midjourney style references as the look layer, especially sketch-like SREF outputs that animate well in Seedance 2.0 in the Topview example. He also tested a new Midjourney illustration style in Seedance 2.0 and said the result worked well in the style test.

0xInk_ pushed the character-lock idea into a 360-degree loop. The Seedance prompt tells the model to read the werewolf's appearance solely from Image1, keep the camera at a fixed radius, complete one clockwise orbit in 10 seconds, return to the exact first frame, and expose only slow idle performance details like breathing, ear motion, tail sweeps, chain movement, and growls in the full Seedance prompt.

AIwithSynthia's lifestyle clips show the same consistency stack in a different genre: GPT Image 2 for the still, Seedance for motion, and prompts that repeat clothing, face, body proportions, hair, ambient audio, and no-text constraints across a haircut story in the salon prompt and a laundry routine in the 4K laundry prompt.

Lip sync and audio passthrough

techhalla's lipsync workflow starts with an audio file and a main character, then connects both into a Seedance 2.0 video generation node inside Magnific Spaces in the node graph. He recommends keeping the audio under 15 seconds and using isolated vocals where possible in the setup step.

The reusable part comes after the first run: swap the reference audio, tweak the prompt, and generate filler clips where the main character walks, moves, or links key moments together in techhalla's follow-up. Magnific's video node docs list prompt, start frame, end frame, references, and audio as video-generator inputs, with audio used for lip-sync models.

PurzBeats' depth-map workflow uses audio differently, passing the audio from the reference video through to the final in the six-step workflow. carolletta said InVideo Agent One also asked ElevenLabs to make music for exact seconds in a sci-fi teaser workflow in the music reply.

4K, platform wrappers, and cost

Seedance 2.0 is showing up through wrappers more than a single canonical app. techhalla named Magnific, Dreamina, CapCut, OpenArt, and InVideo as places to use it in a platform reply, while other examples in the evidence ran through Topview, PixVerse, Higgsfield, Runway, Dreamina, and ComfyUI.

BytePlus says the Dreamina Seedance 2.0 API is fully available, Seedance 2.0 Mini is available, and Seedance 2.0 now supports 4K. AIgateway lists bytedance/seedance-2.0/reference-to-video at $0.014 per second and describes support for up to 9 images, 3 videos, and 3 audio clips.

On creator pricing, techhalla said a 15-second, 1080p Seedance 2.0 generation cost around $2 on his annual plan in a reply. Topview's pricing page lists Seedance 2.0 Mini unlimited 720p access for 365 days under its Ultra annual plan, matching Artedeingenio's reply that the Ultra Annual Plan gives unlimited Seedance 2.0 Mini generations for a full year in the Topview plan reply.

Open projects and loop tests

Higgsfield is turning some polished Seedance work into breakdown material. The Originals post says creators can see the exact prompt, references, generation settings, and complete workflow behind 4K Seedance 2.0 productions in the announcement.

The linked project page gives the open version of that pipeline: the full Higgsfield project includes prompts, references, and settings for the blockbuster breakdown. One sample prompt in the thread locks active references for a boy, machete, spider, and jungle, then specifies first frame, blocking, hard cuts at 4.0s and 9.5s, camera angles, physics, lighting, audio, and positive locks in the prompt dump.

0xInk_'s werewolf loop is a smaller reproducibility test: a 10-second full-body orbit where the first and final frames must coincide exactly, with no zoom, no displacement, and one continuous camera move in the Seedance prompt. That kind of constraint language is becoming the real craft layer around Seedance 2.0.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 8 threads
TL;DR2 posts
Depth-map storyboards4 posts
Reference-video depth control3 posts
Prompt timelines and shot grammar3 posts
Character consistency stacks5 posts
Lip sync and audio passthrough1 post
4K, platform wrappers, and cost1 post
Open projects and loop tests2 posts
Share on X