Skip to content
AI Primer
release

OpenArt Arena ranks image and video models by production task with blind pairwise judgments

OpenArt Arena uses blind pairwise judgments from creative professionals to rank image and video models for ads, film, animation, editing, and lip sync. It also publishes overall leaderboards.

3 min read
OpenArt Arena ranks image and video models by production task with blind pairwise judgments
OpenArt Arena ranks image and video models by production task with blind pairwise judgments

TL;DR

OpenArt's official announcement says its rankings are calculated from aggregated blind pairwise votes with a Bradley-Terry model. The public live Arena pools video cases into an Overall board, then breaks them back out by production lane.

The public board is linked in chrisfirst's launch post, with image and video coverage summarized by chrisfirst's overview.

Blind pairwise voting

A judgment starts with two unlabeled outputs matched to a creative brief. Judges pick the result that best satisfies the relevant criterion, while model identities stay hidden and left-right placement is randomized, according to hasantoxr's launch thread.

OpenArt's published explainer lists 28 expert judges and 1,003 Tastemakers. Its announcement names aesthetics, prompt adherence, realism, and motion quality among the criteria used across tasks.

Twelve production boards

Version 1 puts an Overall board beside specialized lanes. As hasantoxr's post puts it, cinematic film looks, short ads, lip sync, and cleanup are separate jobs.

  • Video: Overall, Ads, Film, Animation, Motion Design, Video Editing, Lip Sync.
  • Image: Overall, E-commerce, Film, Graphic Design, Image Editing.

The live Arena says each Overall board pools every case in that modality's bank. Film appears in both modalities; the remaining lanes separate around the medium and production task.

Video baseline

In MayorKingAI's leaderboard screenshot, the captured Overall video ranking reads:

  1. Seedance 2.5: 1,081
  2. Wan 3.0: 1,004
  3. Seedance 2.0: 1,000
  4. Seedance 2.0 Mini: 989
  5. Google Omni Flash: 981
  6. Flux 3 Video: 958
  7. MiniMax H3: 957
  8. Kling 3.0 Omni: 935
  9. HappyHorse 1.1: 931
  10. Grok Imagine 1.5: 916

The same post says Seedance 2.5 led prompt adherence, aesthetics, physics and motion, and consistency. Those figures are a captured state: the live v1 board currently lists the top three at 1,125, 1,047, and 1,044, in the same order.

Category leaders

The specialized boards publish different leaders from the pooled table. underwoodxie96's result places Wan 3.0 first in Video Editing while Seedance 2.5 remains the overall-video leader.

Image editing has its own three-model order:

  1. GPT Image 2
  2. Seedream 5.0 Pro
  3. Nano Banana Pro

FPV has no board

FPV is absent from the v1 category menu on the live Arena. In a Runway test outside the benchmark's listed lanes, CharaspowerAI said Seedance 2.0 delivered more speed and dynamic movement than 2.5 for FPV shots.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 5 threads
TL;DR4 posts
Blind pairwise voting1 post
Twelve production boards3 posts
Video baseline1 post
Category leaders2 posts
Share on X