FLUX 3
One multimodal model for Image, Video, Audio, and Action-Prediction.
A multimodal foundation model that jointly handles images, video, audio, and action prediction in a unified architecture. Its available video capability generates or transforms clips with native audio, including text/image/video inputs, keyframes, multiple scenes, and continuation.
Pricing
FLUX 3 full-render text-to-video and image-to-video: $0.17/s HD or $0.29/s FHD; draft: $0.06/s (HD). Video continuation full render: $0.43/s HD or $0.54/s FHD; draft: $0.12/s. Documentation states 1 credit = $0.01 USD and API and Playground use the same price.
Black Forest Labs’ official pricing documentation lists FLUX 3 video generation with synchronized audio on a per-second basis. The normalized price records the HD full-render text-to-video/image-to-video rate; other stated variants are captured in notes.