FLUX 3
One multimodal model for Image, Video, Audio, and Action-Prediction.
A multimodal foundation model that jointly handles images, video, audio, and action prediction in a unified architecture. Its available video capability generates or transforms clips with native audio, including text/image/video inputs, keyframes, multiple scenes, and continuation.
Pricing
Pricing varies by operation and settings. Full render text-to-video/image-to-video: $0.17/s HD or $0.29/s FHD; draft: $0.06/s. Video continuation: $0.43/s HD or $0.54/s FHD; draft: $0.12/s. The normalized amount is the listed full-render HD text/image-to-video rate.
Black Forest Labs' own pricing article explicitly prices FLUX 3 video generation per second and lists separate rates by mode, resolution, and draft/full render.
Model Intelligence
Recent stories
Venturetwins reports making nine successive FLUX 3 edits to a Chloe ad without visible degradation. Magnific offers the model with multiple image references, element placement controls and targeted detail editing.
FLUX 3 Image adds bounding-box layouts, precise multi-turn edits, 4K output and support for up to 10 references. A creator comparison found it most accurate at following fine-detail editing instructions.
Black Forest Labs launched FLUX 3 Video Edit [fast] for targeted edits to objects, backgrounds, text, materials, events, and dialogue. The company says the mode better preserves unedited parts of a clip while making the specified change.
Flux 3 appeared in early preview on Nous Hermes Agent as an open-weights video model, with 720p sample tests circulating. Curious Refuge ranked its early results behind Seedance 2.0 and LTX 2.3, making output quality the main launch caveat.