FLUX 3
One multimodal model for Image, Video, Audio, and Action-Prediction.
A multimodal foundation model that jointly handles images, video, audio, and action prediction in a unified architecture. Its available video capability generates or transforms clips with native audio, including text/image/video inputs, keyframes, multiple scenes, and continuation.
Pricing
Pricing varies by operation and settings. Full render text-to-video/image-to-video: $0.17/s HD or $0.29/s FHD; draft: $0.06/s. Video continuation: $0.43/s HD or $0.54/s FHD; draft: $0.12/s. The normalized amount is the listed full-render HD text/image-to-video rate.
Black Forest Labs' own pricing article explicitly prices FLUX 3 video generation per second and lists separate rates by mode, resolution, and draft/full render.
Model Intelligence
Recent stories
FLUX 3 Image supports bounding-box layouts, targeted edits, native 4K output, and ten reference images. Commercial weights are available, open weights are planned, and API use is half-price through October 8.
Black Forest Labs released FLUX 3 Action, an open 7B model that jointly predicts future video and actions for robot policies. The company reports first place on RoboLab and released embodiment fine-tunes, training recipes, and Jetson deployment support.