Introducing MentalHealthBench
OpenAI has introduced MentalHealthBench, developed with over 80 mental health experts from 22 countries, to improve AI responses in sensitive conversations. This initiative builds on previous work like HealthBench and aims to ensure AI models prioritize user safety and well-being, especially as over a billion people use ChatGPT weekly. Enhancements include strengthened responses, expanded access to crisis resources, and the addition of Trusted Contact and ChatGPT for Teens with extra protections.
Time & source
Times shown in UTC
Display time zone: UTC
Local time zone unavailable; showing UTC.
PublishedOffset at this time: UTC+0Sep 23, 2026, 10:00 UTC
IngestedOffset at this time: UTC+0Sep 23, 2026, 21:01 UTC
- Published
- Sep 23, 2026, 10:00
- Ingested
- Sep 23, 2026, 21:01
- Source type
- Official
- Tier
- First-party
- Source status
- Healthy
Tier is a per-source editorial setting, not a per-item score.
Discussion trend
The percentage is based on collected discussion signal, not new comments or independent people. The curve only compares the same topic across time.
People turn to AI for many kinds of conversations: navigating a difficult relationship, working through everyday stress, supporting someone they care about, or deciding how to approach a challenging situation. These conversations require accuracy, practical judgment, and respect for people’s agency. With more than one billion people using ChatGPT each week, our research focuses on helping models respond with care across a wide range of needs and put people’s safety and well-being first.
Most evaluations of AI in this domain have focused primarily on emergency scenarios, given their importance to safety, and measure success using broad, predefined criteria. This has left a gap in understanding how models perform across the full range of mental health conversations, and how well their responses align with expert guidance for each situation, beyond whether they avoid disallowed responses. Assessing how models handle these different situations is essential for building towards AI that actively supports people’s long-term well-being and safety.
We’re introducing MentalHealthBench, a new open benchmark for measuring how AI systems respond in realistic mental health conversations. MentalHealthBench was co-created with a global cohort of more than 80 licensed mental health experts from 22 countries. It assesses model capabilities across key mental health behaviors like safety, seeking context, preserving user agency, and providing actionable guidance when appropriate. We’re releasing it openly so other researchers can examine the methods, run their own evaluations, and build on the work.
Results on MentalHealthBench show the steady improvement of AI systems in helping people navigate mental health situations. While ChatGPT is not a substitute for therapy or professional care, expert-informed insights help us measure the progress towards AI models that are able to respond with empathy, promote well-being, and guide people towards real-world support such as localized crisis hotlines (opens in a new window) or someone they trust .
Supporting mental health across all contexts
Conversations that involve well-being, life advice, or other mental health scenarios can vary widely in topic, urgency, and cultural context. MentalHealthBench is designed to capture the breadth of these realistic scenarios and user personas. Using privacy-preserving techniques, we created synthetic mental health conversations that accurately reflect real-world usage patterns of AI for mental health. Some scenarios also include relevant background information about the synthetic user—such as a recent loss in the family—so we can assess whether models use that context to tailor their responses appropriately.
MentalHealthBench includes scenarios involving adults, teens, caregivers, and clinicians, across multiple languages and regions. The conversations span multiple topical themes, and provide coverage across the full spectrum of acuity:
- Non-acute— Everyday conversations that may involve some emotional components.
- High-acuity —Conversations indicating more serious mental health concerns or significant distress, but not an immediate emergency.
- Emergencies —Conversations involving signs of a mental health emergency or immediate safety concerns that call for urgent real-world support.
Scenarios covered by MentalHealthBench. The mix of scenarios is designed to test model responses and does not represent how often these topics occur in ChatGPT.
Built in collaboration with experts
Building on our previous, clinician-informed work for HealthBench and HealthBench Professional (opens in a new window) , (opens in a new window) we developed MentalHealthBench in close collaboration with our cohort of mental health experts. This consisted of more than 80 licensed psychologists and psychiatrists across 22 countries, speaking 19 languages, and representing nearly 20 mental health subspecialties.
The experts were responsible for reading each synthetic conversation and producing a detailed list of rubric criteria to evaluate model responses to the last user message. Each criterion targets a single aspect of the model response such as asking the right question or providing the best possible advice. They each carry a weight ranging from -10 to +10: positive points reward beneficial behaviors, while negative points penalize harmful ones, and criteria with larger values indicate greater clinical importance in the context of a conversation.
Each conversation was reviewed by at least three experts, and we only retained criteria agreed upon by at least two experts and not contradicted by a third. The final rubrics in the benchmark therefore reflect a shared judgment about how models should respond in each context. For each conversation, we use an automated grader, GPT‑5.6 Sol, to assess model responses against the expert-written criteria. The paper describes the grading process and evaluation settings in detail.
Example conversation with associated rubric items. The rubric evaluates model responses to the last user turn.
Each criterion carries a weight reflecting expert judgment: positive points reward beneficial behaviors, while negative points penalize harmful ones. Criteria with greater clinical importance in the context of a conversation carry larger rewards or penalties—with 10 as the ceiling and 1 as the floor.
1 of 5
Performance of models
We evaluated a wide range of models on MentalHealthBench. The results below show the performance of these models on the entire dataset. The evaluation measures whether a model’s response demonstrates all the ideal behaviors experts identified for each scenario while avoiding less desirable behavior.
Recent frontier models on MentalHealthBench. Error bars show 95% confidence intervals. *Newest evaluated model from each provider (as of September 23, 2026).
The benchmark covers conversations at different levels of acuity. This lets us understand model performance on both everyday mental health scenarios and more urgent mental health emergencies requiring real-world support.
* Newest evaluated model from each provider (as of September 23, 2026).
Conversations in the benchmark represent four types of users: adults, teens aged 13-17, caregivers, and clinicians. For the teen persona, we explicitly stated that the user was between 13-17 through a system message. This approach is designed to work across model providers, though it may not capture all safeguards built into individual products. Conversations involving the teen persona were reviewed by clinicians with expertise in youth mental health to help assess whether responses are appropriate for teenagers’ unique needs.