Skip to content
AI Primer

FLUX 3

One multimodal model for Image, Video, Audio, and Action-Prediction.

A multimodal foundation model that jointly handles images, video, audio, and action prediction in a unified architecture. Its available video capability generates or transforms clips with native audio, including text/image/video inputs, keyframes, multiple scenes, and continuation.

Pricing

Official site · Aug 23, 2026, 6:44 AM
Per second
$0.17

FLUX 3 full-render text-to-video and image-to-video: $0.17/s HD or $0.29/s FHD; draft: $0.06/s (HD). Video continuation full render: $0.43/s HD or $0.54/s FHD; draft: $0.12/s. Documentation states 1 credit = $0.01 USD and API and Playground use the same price.

Black Forest Labs’ official pricing documentation lists FLUX 3 video generation with synchronized audio on a per-second basis. The normalized price records the HD full-render text-to-video/image-to-video rate; other stated variants are captured in notes.

View source

Model Intelligence

Benchmarkable
Yes
Model level
release

Recent stories

1 linked story
AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.