blog

OpenAI's mental-health benchmark scores clinicians' own answers below most models, and its own decomposition says why

OpenAI released MentalHealthBench on 23 September 2026: 1,215 synthetic mental-health conversations, each graded against rubric criteria written by a cohort of more than 80 licensed psychiatrists and psychologists. In the paper's Figure 12, answers written by clinicians score 38.5%. That is below 12 of the 17 models OpenAI tested. GPT-6 Astra, OpenAI's own model, tops the chart at 57.3%.

That headline number is already out. Unite.AI's launch report the same day carried the 38.5% and the authors' explanation that clinicians write short replies. What the coverage left out is the paper's next sentence, and anyone about to quote the benchmark as proof that chatbots out-counsel professionals should read it first. From §4.5 of the PDF:

"The signed-score decomposition shows that clinician completions incur fewer penalties than ordinary model completions but also attain fewer positive points."

The clinicians took the smallest penalties on the chart

Each rubric item carries a weight from −10 to +10. Positive items reward a behaviour. In the paper's worked example (Figure 3), asking the user what may be helpful right now is worth +7. Negative items penalise one: telling the user they already know what to do about the friend costs 7. A grader model marks each item met or not met. Add the points for every item met, penalties included, and divide by the positive points on offer for that conversation: §2.3 calls that the signed score, and it can go below zero. The headline numbers, the clinicians' 38.5% among them, are a different measure. Before averaging, §2.3 clips each answer's signed score at zero, so an answer whose penalties outweigh its points counts as nothing rather than as a negative. The decomposition works on the unclipped signed score, split into positive points earned and penalty burden, each as a share of the positive-point budget. Its bars will not add up to the clipped headline figure.

Figure 5 prints both halves for every model. The smallest penalty any model takes is 17% of the positive-point budget (GPT-6 Sol). The largest is 44% (Gemini 3.1 Pro). GPT-6 Astra takes 19%. On the positive side, models earn between 45% (GPT-4o) and 73% (Claude Opus 5.5).

Figure 12 plots the clinicians' bars on the same axes without printing their values. Read against that axis, the clinicians' penalty bar comes to about 8% and their positive bar to about 43%. The penalty is less than half of any model's. The positive share is below every model's, and that is where their score goes.

A one-question reply cannot collect a long rubric

The paper's explanation is length and approach, not competence. The authors write that "clinicians tend to write extremely short responses, as if in an in-person conversation, by asking a single question or making a simple statement." Figure 13 backs it: on its log-scale length axis, the clinicians' answers average somewhere under 150 tokens, shorter than anything else plotted. The tersest model, GPT-6 Sol, sits near 190.

The clinicians weren't writing for a patient in the room. §3.2.3 says a fourth clinician, uninvolved in that conversation's rubric, was asked "to write the next response they would most want a safe and helpful AI system to provide", without seeing the rubric or any model's answer. They still wrote the way a clinician talks. One good question, then wait for the answer.

A rubric pays for every listed behaviour a reply hits in one turn. The paper never counts how many criteria a clinician's answer met, and the release holds no clinician answers to count, but a reply built around one question has little room to hit many. The paper's other reference point shows what full coverage looks like: GPT-6 Astra, handed the rubric and told to maximise it, scores 99.0% in Figure 12. Those rubric-aware answers are not long by the chart's standards either, plotted in Figure 13 at under 300 tokens against 600-plus for Claude Opus 5.5. The benchmark rewards covering what the rubric lists, and a reply written as the first move in a conversation covers little.

Who this lands on

Anyone quoting the benchmark to argue that chatbots handle mental-health conversations better than people. That covers product pages, investor decks and op-eds. The paper's own numbers support a narrower claim: on a single turn, graded against a checklist, current models hit more of the checklist and more of what it penalises. Nothing in the paper says they gave worse care, and the authors do not claim it.

The parties matter too. OpenAI wrote the benchmark and chose its reference points. Per §2.3, its GPT-5.6 Sol grades every answer. Its models take first, second and fourth place in Figure 5, with Claude Opus 5.5 third at 52.4%. The paper's conclusion (§7) asks for the benchmark to be read "not as a definitive leaderboard, but as an auditable diagnostic tool."

The audit has a gap exactly where this finding sits. The dataset is a zip on OpenAI's CDN, linked from §8 of the paper and MIT-licensed. Downloaded on 25 September 2026, it holds the 1,215 conversation prefixes and 5,262 rubric criteria, with no field for completions. The README asks users not to post examples online in plain text or images. The grader prompt is printed in Appendix B.2, so a model's score can be re-run by anyone with API access to the grader. The clinicians' answers are not in the release, so their 38.5% cannot be re-graded by anyone outside OpenAI.

Sources

Retrieved 25 September 2026 unless noted.

  • OpenAI, "MentalHealthBench: An Expert-Informed Benchmark of AI Capabilities in Realistic Mental Health Conversations" (PDF; §2.3, §3.2.3, §4.5, §7, §8, Appendix B.2, Figures 3, 5, 12 and 13) — cdn.openai.com
  • OpenAI, "Introducing MentalHealthBench", 23 September 2026, dated by OpenAI's news RSS feed. The page itself refused automated fetches. — openai.com/index/introducing-mentalhealthbench
  • MentalHealthBench dataset, README and MIT licence — cdn.openai.com/ctf-cdn/OAI_MentalHealthBench.zip
  • Jonas Reeve, Unite.AI, "OpenAI debuts MentalHealthBench for AI mental health conversations", 23 September 2026 — unite.ai

Primary evidence