Skip to content
AI Primer
release

OpenAI releases MentalHealthBench, an open AI mental-health benchmark

OpenAI released MentalHealthBench, an open benchmark for AI responses to everyday support and crisis-related mental-health conversations, developed with mental-health experts. OpenAI reports GPT-6 Astra scored 57.3 versus 32.1 for GPT-4o.

3 min read
OpenAI releases MentalHealthBench, an open AI mental-health benchmark
OpenAI releases MentalHealthBench, an open AI mental-health benchmark

TL;DR

  • MentalHealthBench is an open benchmark for model replies to everyday support and crisis-related mental-health conversations, thekaransinghal announced.
  • Its 1,215 synthetic conversations receive custom expert rubrics, according to thekaransinghal.
  • GPT-6 Astra scored 57.3 against GPT-4o's 32.1, a 25.2-point gap, on OpenAI's results chart.
  • Similar aggregate scores can conceal sharply different context-seeking performance, thekaransinghal wrote.

Each criterion in the technical report carries a weight from -10 to +10. A separate ValsAI study found that 62% of its critical teen-safety failures emerged after the first assistant reply.

The 1,215 conversations

The benchmark spans adults, teens, caregivers, and clinicians, across multiple languages and regions, according to thekaransinghal. It covers the continuum from daily wellbeing and life advice to high-risk and acute situations.

MentalHealthBench uses synthetic dialogues built with privacy-preserving techniques. The official announcement says some scenarios supply personal context, such as a recent death in the family, so a response can be evaluated for whether it uses that context.

Weighted rubrics

Each model completion is evaluated against a rubric written for the conversation's final user message. The technical report says two clinicians independently authored weighted criteria and a third clinician adjudicated them.

  • Each criterion targets one response behavior, such as asking an appropriate question or giving suitable guidance.
  • Weights run from -10 for harmful behavior to +10 for beneficial behavior, with larger magnitudes reserved for more clinically significant behaviors.
  • OpenAI uses GPT-5.6 Sol to grade model answers against those expert-authored criteria.

More than 80 licensed psychologists and psychiatrists in 22 countries produced the rubrics, thekaransinghal said.

Model scores

GPT-6 Astra scored 57.3, while GPT-4o from March 2025 scored 32.1, on OpenAI's benchmark chart. GPT-6 Sol scored 53.9 and Claude Opus 5.5 scored 52.4.

The reported number measures satisfaction of weighted expert criteria for synthetic replies. GPT-5.6 Sol's role as automated evaluator remains part of the measurement stack described in the technical report.

Ten behaviors

MentalHealthBench also separates results across 10 behaviors. Context-seeking varies substantially across models, while empathy and support cluster more closely, thekaransinghal wrote.

OpenAI said the models still have room to improve at eliciting context and preserving a user's ability to make their own decisions, in thekaransinghal's results update. The benchmark materials are public for other researchers and model developers to inspect and extend, thekaransinghal said.

Teen conversations over turns

MentalHealthBench grades one completion after a conversation prefix. ValsAI's separate test ran 648 simulated, 10-turn teen conversations across nine model APIs; at least one critical safety check failed in 27.5% of conversations, and 62% of those failures appeared later in the exchange.

ValsAI also found that explicitly identifying the user as a teen in the system instruction reduced the critical-failure rate from 31.1% to 11.5% in a separate comparison. Its study measured simulated API conversations rather than what teens encounter in consumer apps, ValsAI noted.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 2 threads
The 1,215 conversations1 post
Ten behaviors2 posts
Share on X