Skip to content
AI Primer
update

ARC Prize reports lower matched-effort scores for GPT-6 Sol than GPT-5.6 Sol

ARC Prize reports lower ARC-AGI-1 and ARC-AGI-2 scores for GPT-6 Sol than GPT-5.6 Sol at matched effort. Lower prices and fewer reasoning tokens cut per-task costs by over 70%.

3 min read
ARC Prize reports lower matched-effort scores for GPT-6 Sol than GPT-5.6 Sol
ARC Prize reports lower matched-effort scores for GPT-6 Sol than GPT-5.6 Sol

TL;DR

  • GPT-6 Sol scored 78.1% on ARC-AGI-2 at xhigh effort, down from GPT-5.6 Sol’s 90.0%, according to ARC Prize’s comparison.
  • ARC-AGI-1 fell from 97.5% to 92.7% at xhigh, while cost per task dropped 73% across matched effort levels, per ARC Prize’s follow-up.
  • On interactive ARC-AGI-3, the harness changed Sol’s max-effort score from 4.6% to 23.0%, as ARC Prize’s results post reports.

The per-game results show Sol scoring 0.0% on TR87 with the standard harness and 80.3% with the provider adapter at max effort. OpenAI’s launch post measures its 50% API price cut against GPT-5.6 Sol’s promotional rates.

Matched-effort scores

ARC Prize’s GPT-5.6 Sol and GPT-6 Sol result tables show declines at both xhigh and max reasoning:

  • ARC-AGI-2, xhigh: 90.0% → 78.1%, −11.9 points, as ARC Prize’s comparison notes.
  • ARC-AGI-1, xhigh: 97.5% → 92.7%, −4.8 points, per ARC Prize’s follow-up.
  • ARC-AGI-2, max: 92.5% → 89.6%, −2.9 points, in the linked result tables.
  • ARC-AGI-1, max: 96.5% → 95.5%, −1.0 point, in the same tables.

Cost per task

OpenAI cut Sol’s standard input and output prices from $4/$20 to $2/$10 per million tokens relative to GPT-5.6’s promotional prices, according to its launch announcement. ARC Prize reports the resulting evaluation costs separately from accuracy:

ARC-AGI-3 harnesses

At max reasoning, ARC Prize’s results post gives Sol 4.6% under the standard harness for a $5.6K evaluation, versus 23.0% under the provider adapter for $8.7K. That is an 18.4-point score difference for the same model at the same effort level.

The testing policy describes what changed:

  • Standard: A provider-neutral interface that carries forward notes the model chooses to keep.
  • Provider adapter: Provider-designed context management that can preserve opaque reasoning state between requests and compact long conversations.

Public runs and verified scores

The verified testing policy says ARC Prize uses a single run per evaluation rather than averaging runs. Its ARC-AGI-1 and ARC-AGI-2 verified scores use a semi-private set; public-task outputs, costs, durations and task scores are published separately. ARC Prize links its open-source benchmarking runner for reproducing public results.

Share on X