Skip to content
AI Primer

FLUX 3

One multimodal model for Image, Video, Audio, and Action-Prediction.

A multimodal foundation model that jointly handles images, video, audio, and action prediction in a unified architecture. Its available video capability generates or transforms clips with native audio, including text/image/video inputs, keyframes, multiple scenes, and continuation.

Pricing

Official site · Oct 5, 2026, 7:01 AM
Per second
$0.17

Pricing varies by operation and settings. Full render text-to-video/image-to-video: $0.17/s HD or $0.29/s FHD; draft: $0.06/s. Video continuation: $0.43/s HD or $0.54/s FHD; draft: $0.12/s. The normalized amount is the listed full-render HD text/image-to-video rate.

Black Forest Labs' own pricing article explicitly prices FLUX 3 video generation per second and lists separate rates by mode, resolution, and draft/full render.

View source

Model Intelligence

Benchmarkable
Yes
Model level
release

Recent stories

2 linked stories
AI PrimerAI Primer

Your daily guide to AI tools, workflows, and creative inspiration.

© 2026 AI Primer. All rights reserved.