Skip to content
AI Primer
breaking

Harvey LAB-AA v1.1 requires hallucination-free answers for benchmark credit

Harvey LAB-AA v1.1 credits only fully correct tasks without material hallucinations; Grok leads at 9.4%. Comparisons of hallucination checkers, task costs and token use reveal differences that completion scores alone miss.

5 min read
Harvey LAB-AA v1.1 requires hallucination-free answers for benchmark credit
Harvey LAB-AA v1.1 requires hallucination-free answers for benchmark credit

TL;DR

  • Legal tasks earn headline credit only when every rubric criterion passes and no material hallucination survives, under the v1.1 scoring rules. Minor errors do not trigger the gate.
  • Grok 4.7 (xhigh) leads at 9.4%, while more than 60% of otherwise passing results contain a material hallucination, according to the launch announcement.
  • Grounding reshuffles the rankings: Muse Spark 1.3 falls from 26.7% to 8.9% after gating in ArtificialAnlys's breakdown.
  • Checker choice changes the error count dramatically: GPT-6 Sol upheld 470 material hallucinations versus Claude Sonnet 5.5's 57 in the checker comparison.

LAB-AA rejects incorrectly named deliverables and strips out Harvey's custom document-generation tools. Its worked tax example catches an agent treating one loan's opening and closing balances as two separate debts.

Artificial Analysis evaluates agents on 120 private tasks across 24 practice areas. Harvey's broader LAB release contains more than 1,200 tasks and over 75,000 expert-written rubric criteria; each assignment combines instructions, client materials and a required work product.

The independent LAB-AA implementation makes three consequential harness choices:

  1. Stirrup: context compaction and simplified Artificial Analysis-authored agent and judge prompts.
  2. Code execution: a simple execution tool replaces Harvey's custom tools and document-generation skill scripts.
  3. Exact filenames: deliverables must use the specified filename rather than relying on fuzzy matching.

The v1.1 release notes also disclose an updated private dataset and changed judging. Its scores are not directly comparable with v1.0.

Hallucination gating

Harvey and Artificial Analysis added a welcome hard veto: one material hallucination zeroes a task's headline contribution, even when its rubric requirements pass. Minor hallucinations remain separately reported.

GPT-6 Sol (high) runs both stages of the grounding audit:

  1. Find candidate hallucinations by comparing deliverables with source documents. Flags fall into three categories:
  2. Recheck every flag against the sources, dismiss flags that do not hold up, and classify confirmed errors as material or minor.

General legal knowledge, including case law or statutes outside the supplied files, is outside this audit's scope. A task without a usable submission scores zero but is excluded from the hallucinations-per-task denominator.

Rubric grading uses a separate panel: GPT-6 Sol, Grok 4.7 and Claude Opus 5.5. The headline calculation averages each task's share of judges finding every criterion passed, then zeroes that contribution if the hallucination audit finds a material error.

The leaderboard reshuffle

Grok 4.7 (xhigh) takes first at 9.4%, ahead of Muse Spark 1.3 (max) at 8.9% and GPT-6 Astra (max) at 8.6% in the launch results. Astra moves from joint tenth before gating to third afterward.

The reductions below compare ungated and gated scores from the same evaluation, using the model breakdown and the launch results:

| Model and setting | All-pass | Hallucination-gated | Change |
| --- | ---: | ---: | ---: |
| Muse Spark 1.3 (max) | 26.7% | 8.9% | −17.8 points |
| GPT-6 Astra (max) | 8.9% | 8.6% | −0.3 points |
| GPT-6.1 Sol (max) | 7.5% | 6.9% | −0.6 points |
| GLM-5.3 (max) | 13.9% | 0.3% | −13.6 points |
| Kimi K3 (max) | 16.7% | 5.3% | −11.4 points |
| Claude Sonnet 5.5 (max with fallback) | 11.7% | 2.8% | −8.9 points |

Criterion Pass Rate pools individual rubric verdicts rather than requiring a complete task to pass. Kimi and Muse illustrate how high coverage can coexist with unsupported claims:

Six hallucination checkers

Artificial Analysis compared six checkers, all at high reasoning effort, on identical deliverables from eight models across 20 tasks:

  • GPT-6 Sol: upheld 470 material hallucinations; selected for production.
  • GPT-6 Luna: generally identified more material hallucinations alongside Sol.
  • Grok 4.7: upheld 219.
  • Claude Opus 5.5: upheld 99.
  • Claude Sonnet 5.5: upheld 57.
  • Gemini 3.8 Flash: generally identified far fewer material hallucinations.

All six found zero material hallucinations in Astra's outputs on this subset. Artificial Analysis cautioned that counts alone establish neither checker accuracy nor absence of self-preference.

Cost per task

Four models form the score-versus-cost Pareto frontier among models scoring above zero:

  • GPT-6 Luna (max): approximately $0.22 per task, scoring 3.3%.
  • GPT-6.1 Sol (max): scoring 6.9%.
  • Muse Spark 1.3 (max): approximately $4.20 per task, scoring 8.9%.
  • Grok 4.7 (xhigh): approximately $9.50 per task, scoring 9.4%.

The three Claude models cost approximately $18 to $22 per task.

Output tokens

Artificial Analysis counts both reasoning and answer tokens in its output-token totals:

  • GPT-6 Astra (max): approximately 81,000 output tokens per task, scoring 8.6%.
  • Grok 4.7 (xhigh): approximately 180,000, scoring 9.4%.
  • Three Claude models: approximately 202,000 to 562,000, scoring 2.8% to 6.4%.

Near-pass scores

Allowing one or two missed rubric criteria produces another ranking change. The near-pass analysis retains the material-hallucination veto in every band:

| Model | No misses | At most one miss | At most two misses |
| --- | ---: | ---: | ---: |
| GPT-6 Astra (max) | 8.6% | 20.3% | 31.7% |
| GPT-6.1 Sol (max) | 6.9% | 20.3% | 29.6% |
| Grok 4.7 (xhigh) | 9.4% | 15.3% | 23.1% |

GLM-5.3 reaches only 2.2% with two allowed misses; Gemini 3.8 Flash remains at zero.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 2 threads
TL;DR2 posts
The leaderboard reshuffle2 posts
Share on X