OpenArt Arena ranks image and video models by production task with blind pairwise judgments
OpenArt Arena uses blind pairwise judgments from creative professionals to rank image and video models for ads, film, animation, editing, and lip sync. It also publishes overall leaderboards.

TL;DR
- OpenArt Arena separates image and video models into production-task boards, including ads, film, animation, editing, and lip sync, as hasantoxr's launch thread describes.
- Every result begins with an unlabeled, side-by-side choice against a specific creative criterion, according to hasantoxr's judging explainer.
- Seedance 2.5 led one captured overall-video table at 1,081, ahead of Wan 3.0 at 1,004 and Seedance 2.0 at 1,000 in MayorKingAI's leaderboard screenshot.
- Task boards already split the leaders: underwoodxie96's video-editing post reports Wan 3.0 at No. 1 for editing, while underwoodxie96's image-editing post puts GPT Image 2 first for image editing.
OpenArt's official announcement says its rankings are calculated from aggregated blind pairwise votes with a Bradley-Terry model. The public live Arena pools video cases into an Overall board, then breaks them back out by production lane.
The public board is linked in chrisfirst's launch post, with image and video coverage summarized by chrisfirst's overview.
Blind pairwise voting
A judgment starts with two unlabeled outputs matched to a creative brief. Judges pick the result that best satisfies the relevant criterion, while model identities stay hidden and left-right placement is randomized, according to hasantoxr's launch thread.
OpenArt's published explainer lists 28 expert judges and 1,003 Tastemakers. Its announcement names aesthetics, prompt adherence, realism, and motion quality among the criteria used across tasks.
Twelve production boards
Version 1 puts an Overall board beside specialized lanes. As hasantoxr's post puts it, cinematic film looks, short ads, lip sync, and cleanup are separate jobs.
- Video: Overall, Ads, Film, Animation, Motion Design, Video Editing, Lip Sync.
- Image: Overall, E-commerce, Film, Graphic Design, Image Editing.
The live Arena says each Overall board pools every case in that modality's bank. Film appears in both modalities; the remaining lanes separate around the medium and production task.
Video baseline
In MayorKingAI's leaderboard screenshot, the captured Overall video ranking reads:
- Seedance 2.5: 1,081
- Wan 3.0: 1,004
- Seedance 2.0: 1,000
- Seedance 2.0 Mini: 989
- Google Omni Flash: 981
- Flux 3 Video: 958
- MiniMax H3: 957
- Kling 3.0 Omni: 935
- HappyHorse 1.1: 931
- Grok Imagine 1.5: 916
The same post says Seedance 2.5 led prompt adherence, aesthetics, physics and motion, and consistency. Those figures are a captured state: the live v1 board currently lists the top three at 1,125, 1,047, and 1,044, in the same order.
Category leaders
The specialized boards publish different leaders from the pooled table. underwoodxie96's result places Wan 3.0 first in Video Editing while Seedance 2.5 remains the overall-video leader.
Image editing has its own three-model order:
- GPT Image 2
- Seedream 5.0 Pro
- Nano Banana Pro
FPV has no board
FPV is absent from the v1 category menu on the live Arena. In a Runway test outside the benchmark's listed lanes, CharaspowerAI said Seedance 2.0 delivered more speed and dynamic movement than 2.5 for FPV shots.