Skip to content
AI Primer
workflow

Seedance 2.0 creators use depth maps for AI video motion control

Creators shared Seedance 2.0 workflows that convert footage or storyboards into depth-map references. The workflows aim for tighter motion transfer and composition control with fewer prompt tokens.

6 min read
Seedance 2.0 creators use depth maps for AI video motion control
Seedance 2.0 creators use depth maps for AI video motion control

TL;DR

  • Depth maps became the control layer for Seedance motion transfer: CuriousRefuge tested reference footage to depth map to Seedance 2.0 Omni and said depth produced the cleanest motion transfer among variations.
  • Storyboards got a depth-only variant: _OAK200 used a reference image for tone and look while the depth-map storyboard carried composition and camera framing.
  • The revised storyboard workflow now starts with a normal storyboard, then converts it into depth: _OAK200 said that worked better than asking the image model to generate the whole storyboard directly as a depth map.
  • A Comfy-style version is already compact: PurzBeats posted a six-step chain with reference video, depth map, reference image, settings, audio passthrough, and run.
  • Character consistency is being handled with reference chaining too: techhalla used the previous video as the next Seedance reference after building a 15+ shot Magnific workflow.

ByteDance's launch post says Seedance 2.0 can take text, image, audio, and video together, including up to 9 images, 3 video clips, and 3 audio clips. Fal's Depth Anything Video docs describe per-frame depth maps with temporal consistency, grayscale output, and raw depth export. Comfy.org's Depth Motion Capture workflow calls the extracted depth sequence a motion skeleton. Magnific's changelog added Seedance 2.0 Pro and Fast endpoints for text-to-video and image-to-video.

Depth as motion skeleton

CuriousRefuge's tested motion-transfer workflow has five parts:

  1. Start with reference footage and create character and environment references.
  2. Convert the footage into a depth map with Depth Anything Video on Fal.
  3. Optionally add colorized character masks with SAM2Matting and skeleton-pose guidance.
  4. Feed the depth map into Seedance 2.0 Omni as a motion reference, alongside character and environment references.
  5. Generate the final video.

Fal's docs say Depth Anything Video produces temporally consistent per-frame depth estimates, with Small, Base, and Large model sizes plus grayscale visualization. That makes it a clean fit for movement, camera, and spatial layout instead of identity or texture.

PurzBeats compressed the same idea into a six-step Comfy run:

  1. Reference video -> convert to depth map.
  2. Depth map -> Seedance 2.0 reference video.
  3. Add reference image.
  4. Set duration, resolution, and aspect ratio.
  5. Pass audio from the reference video to the final output.
  6. Hit run.

PurzBeats later said he was using Depth Anything 2 in ComfyUI in a reply. The audio passthrough is the quietly useful part for dance, lipsync, and music-video tests.

Depth-map storyboards

The clean trick is separation: _OAK200 used a reference image to define tone and final look, while the depth-map storyboard defined composition and camera framing.

_OAK200's breakdown says the depth-map storyboard strips away color, lighting, texture, and style, leaving only:

  • Camera placement
  • Subject scale
  • Foreground
  • Midground
  • Background
  • Silhouettes
  • Spatial relationships

The final Seedance pass used three references: the original image for style, the depth-map storyboard for shot composition, and character sheets for identity consistency _OAK200's final step.

Normal storyboard first

The revised workflow adds one production-minded change: _OAK200 now generates the normal storyboard first, then converts it into a depth-map storyboard with Nano Banana or GPT Image 2.

The conversion prompt is strict about what survives the translation:

  • Preserve exact canvas size, aspect ratio, panel layout, dividers, framing, perspective, composition, and silhouettes.
  • Estimate depth independently for each panel.
  • Use white for closest surfaces, gray for middle distance, near-black for the farthest sky or horizon, and black for dividers.
  • Keep crisp silhouettes and fine details such as fingers, hair, vegetation, wires, and object edges.
  • Treat brightness as distance only, not lighting, texture, fog, shadows, reflections, or motion blur.
  • Output only the neutral grayscale depth map, with no labels, added objects, missing objects, or layout changes.

The workflow reduces prompt load because the depth map carries camera and blocking information that would otherwise need to be described in words, according to _OAK200's note.

The 3×3 depth grid

_OAK200's longer system prompt gives depth storyboarding a full previz grammar. The output must be exactly nine panels in a 3×3 grid, reading left to right across each row.

The prompt breaks the job into eight phases:

  1. Analyze the visual and depth references.
  2. Define a simple story beat.
  3. Plan the nine shots.
  4. Maintain continuity.
  5. Design cinematic depth.
  6. Generate depth maps only.
  7. Build the 3×3 storyboard.
  8. Run quality control.

The specified shot order is concrete: establishing wide, movement or intention, discovery or POV, reaction, preparation, insert detail, main action, consequence, and resolution wide.

Reference chaining

techhalla's Magnific workflow attacked consistency from another angle: one starting image, two models, and three prompts for 15+ different shots.

The chain looked like this:

techhalla's separate fantasy-fight prompt shows the other control style: a 12-second clip broken into timed beats, from 0 to 2 seconds for a slow push-in to 10.5 to 12 seconds for a slow-motion final impact techhalla's time-coded prompt.

Runway, Magnific, Comfy Cloud

The workflows are spreading across creator surfaces, not one native app. Cloudflare's Seedance 2.0 docs list bytedance/seedance-2.0 as a multimodal model with synchronized audio, reference images, reference videos, and reference audio.

AIgateway's reference-to-video page lists the model slug bytedance/seedance-2.0/reference-to-video at $0.014 per second. Magnific's changelog lists Seedance 2.0 Pro endpoints at 480p, 720p, and 1080p, with Seedance 2.0 Fast at 480p and 720p.

Creator posts show three practical access paths:

  • techhalla stayed inside Magnific, using GPT-2 for the starting image and Seedance 2.0 to animate it techhalla.
  • awesome_visuals made a Seedance 2.0 clip on Runway, then used Topaz Labs Astra for upscaling and FPS increase awesome_visuals.
  • PurzBeats pointed people to a hosted Comfy Cloud version of the workflow PurzBeats.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 6 threads
TL;DR1 post
Depth as motion skeleton1 post
Depth-map storyboards3 posts
Normal storyboard first1 post
Reference chaining4 posts
Runway, Magnific, Comfy Cloud1 post
Share on X