Creator reports making a 15-second SaaS ad with 74 Claude Opus 5.5 agents
A creator reports producing a 15-second motion-graphics ad without manual edits using Claude Opus 5.5 agents. The run involved 5,313 tool calls and about 12 hours to build the scenes, soundtrack and render engine.

TL;DR
- One prompt reportedly produced a 15-second SaaS motion-graphics ad with zero manual edits, according to tomaldertweets' post.
- The overnight run used 74 Opus 5.5 agents, 5,313 tool calls, 19.2 million tokens, and about 12 hours, per tomaldertweets' run breakdown.
- The reported production stack built vector components, scenes, soundtrack, and a render engine without After Effects, MCPs, extensions, or plugins, as tomaldertweets described it.
- A separate Claude Code workflow turns code-based motion into six repeatable steps, from timed states to ffmpeg export, in aakashgupta's motion-graphics guide.
Anthropic's launch post describes Opus 5.5 as a model for long-running agentic work, while the platform docs list a 1M-token context window and 128K maximum output. A technical breakdown explains the less obvious mechanism behind these clips: the model writes a time-indexed drawing program, then a browser and ffmpeg produce the video. A 63-video analysis found that launch-week showreels clustered around 15 seconds, dark backgrounds, and music without voiceover.
One prompt, an overnight render
The creator asked Claude to make a dynamic 15-second motion-graphics video for a linked SaaS product, styled as a showreel for a motion designer's resume. The post says the prompt was the only brief and the creator made zero edits.
The reported agent responsibilities were:
- Assemble a panel of creative directors to pitch and judge storyboards.
- Pull the product codebase and turn it into vector components.
- Build every scene.
- Compose the soundtrack.
- Write a video render engine from scratch.
The post says the run used Claude Desktop with Opus 5.5 at “Ultracode effort.” In a reply, tomaldertweets described the human intervention as sending the prompt, going to bed, and waking up to the finished result.
The 74-agent production line
The phrase “one shot” describes the handoff, not the amount of machine work. A follow-up from tomaldertweets counts 74 Opus 5.5 agents across a roughly 12-hour production run.
- 74 agents
- 5,313 tool calls
- 19.2 million tokens
- About 12 hours
- 128 BPM soundtrack, composed entirely in code
- 1080p output, 900 frames
- Each frame blended from 16 to 64 renders
Nine hundred frames over 15 seconds implies a 60-frame-per-second render. The soundtrack was synchronized to the beat rather than added as a separate finishing step, according to the same breakdown.
The render engine
Opus 5.5's official model documentation describes text and image input with text output. The video workflow documented by Pasquale Pillitteri therefore uses the model to write a program that draws each frame, rather than returning an MP4 directly.
The pipeline is compact:
- Claude writes a
draw(t)orseek(t)function that describes the frame at timet. - A headless browser calls that function at each frame timestamp and captures screenshots.
- ffmpeg encodes the screenshots into a video file.
The result can be regenerated at any frame because the scene is a function of time. koldo2k described a stickman animation as “animated and scored entirely in code by Opus 5.5,” while techhalla's example reported a separate coded motion piece made in 14 minutes.
The six-step motion workflow
A separate post about Sonnet 5.5 in Claude Code lays out a practical version of the same code-first approach.
- Start from a reference: attach a screenshot and specify a single HTML file, inline SVG, and an exact loop length.
- Write the states: describe what happens and when, with timestamps for each interaction.
- Make time the only input: define every element as a pure function of
t, remove timers and CSS keyframes, and make frame 0 equal frame 6 for a closed loop. - Specify the visual rules: provide hex colors and fonts, use springs for arrivals, eases for exits, and morph related shapes between states.
- Make the agent inspect itself: render eight moments side by side, then identify defects by timestamp, such as an overlap at 2.0 seconds.
- Export deliberately: save PNG frames, use ffmpeg's
palettegenandpaletteusewith no dithering, and export at 15 frames per second.
The method turns motion design into a set of inspectable states, timing functions, and render checks. It also leaves an editable source file behind instead of only a flattened clip.
The hybrid After Effects stack
The all-code report is one branch of the launch-week workflow. Another keeps a conventional compositor in the loop.
[higgsfield_ai's workflow] assigns concept and scene planning to Claude Opus 5.5, styleframes and visual assets to Higgsfield, and animation, camera moves, and compositing to After Effects. A related higgsfield_ai project post says the output includes an After Effects project that can be edited further.
A separate post from higgsfield_ai calls the Opus, Higgsfield, and After Effects combination a product-launch workflow. minchoi's post describes the broader shift as one person chaining several models and publishing the same day.
The examples end in different artifacts: a custom code renderer in the first report and an editable After Effects project in the Higgsfield workflow. Both assign planning and production decisions to an agent before final motion is rendered.
Access and the pixel gap
Anthropic's Opus 5.5 documentation lists the model ID claude-opus-5-5, a 1M-token context window, a 128K maximum output, and pricing of $4 per million input tokens and $20 per million output tokens. The API is listed on Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS.
The creator used Claude Desktop, while a follow-up from tomaldertweets said 1,600 people were already on the waitlist linked from the post. Anthropic's docs set the API's default effort to medium and say adaptive thinking is always on, while the showcased run identifies its setting as “Ultracode effort.”
The production limit is concrete: the model writes code and text, and the browser-rendering stack supplies the frames. The technical breakdown describes the approach as strongest for geometric shapes, kinetic typography, and interfaces, with complex photorealistic scenes still better suited to native video generators.