Anthropic has announced a new research initiative funding the development of “wellbeing evaluations” — benchmarks designed to measure whether AI models can accurately assess and support human psychological wellbeing. The program represents a novel approach to AI safety that goes beyond preventing harm to actively promoting flourishing.
What Are Wellbeing Evaluations?
Wellbeing evaluations are benchmarks that test whether AI models can:
- Recognize signs of psychological distress in user interactions
- Provide appropriate responses that support rather than harm
- Avoid reinforcing negative patterns — for example, not encouraging rumination or avoidance
- Understand cultural context — wellbeing concepts vary across cultures and individuals
- Know their limits — recognize when a situation requires human professional support rather than AI assistance
The Research Program
Anthropic is funding research teams at several universities to develop these evaluations, with the goal of creating standardized benchmarks that can be applied to any AI model. The program includes:
- $12 million in research grants over three years
- Partnerships with psychology departments at Stanford, MIT, and University College London
- Open publication of all evaluation benchmarks and methodologies
- Cross-industry collaboration — inviting other AI companies to adopt and contribute to the benchmarks
Why It Matters
Traditional AI safety evaluations focus on preventing obvious harms — bias, toxicity, dangerous content. Wellbeing evaluations address a more subtle but equally important question: does interacting with an AI model make people’s lives better or worse?
This is particularly important as AI models become more conversational and emotionally engaging. Users increasingly turn to AI for advice on personal matters, and the quality of that advice has real consequences for mental health and life satisfaction.
The Challenge
Measuring wellbeing is inherently more complex than measuring bias or toxicity. There is no universal definition of “flourishing,” and what constitutes good psychological support varies across:
- Cultural contexts — different cultures have different norms around emotional expression and help-seeking
- Individual circumstances — what helps one person may harm another
- Situational factors — the same response might be appropriate in one context and inappropriate in another
Anthropic acknowledges these challenges and frames the research program as an iterative process — developing evaluations that improve over time as understanding of AI-human interaction deepens.
Industry Implications
By funding independent research on wellbeing, Anthropic is positioning itself as a leader in a new category of AI safety — one that measures positive impact rather than just negative outcomes. If the evaluations prove useful, they could become a standard part of AI model assessment, alongside existing benchmarks for accuracy, safety, and fairness.
The approach also signals a shift in how AI companies think about their products. Rather than simply asking “is this model safe?” the question becomes “is this model helpful?” — a harder but more important question.