Announcements

Funding better evaluations of AI’s impact on wellbeing

Funding better evaluations of AI’s impact on wellbeing

We’re launching a $5 million grant program to fund independent research into how AI impacts users’ wellbeing. The program will provide direct funding, access to our models, and technical support to grantees building open-source evaluations that help the AI industry measure how our models affect those who use them. Grantees will work fully independently, and will publish their work as open-source projects that any developer can make use of.

AI systems have become central to how many people work, learn, and solve problems. They’ve also become conversational partners and can be sources of emotional support during difficult times. But as an industry, we are still working towards developing clear standards for how models should behave in these conversations, for example, when a user begins to seek companionship from a model, or uses AI to navigate a mental health crisis.

Furthermore, wellbeing is a particularly difficult area to evaluate. For most model behaviors, we can look at a single answer and determine whether it is accurate and appropriate. But assessing wellbeing requires much more context. For example, a user in distress might not share thoughts of self-harm right away; the need for a more cautious response might only become clear over the course of a long conversation. And a response that might be reasonable in one context might be harmful in another. For example, Claude might give advice on balanced diets and workout routines to a user who asks about losing weight, but if the user has demonstrated a history of disordered eating, that response could be inappropriate, and potentially actively harmful.

We work to develop safeguards to identify such conversations and help ensure Claude responds appropriately, and we publish research into the types of conversations people have with Claude to better inform how we develop our safeguards, how we evaluate them, and other measures we can take to protect users’ wellbeing. But these are nuanced considerations, and the stakes are significant. The right approach will need to evolve alongside our models and their uses.

By funding the creation of independent evaluations and benchmarks of user wellbeing, we hope to invite more people to lend their expertise to this emerging and critical field, including clinicians, psychologists, methodologists, and others.

Towards more effective wellbeing evaluations and benchmarks

As part of this program, we’re sharing guidance from our Safeguards team on what we believe makes a wellbeing evaluation rigorous enough to build on, along with the common challenges that can limit an evaluation’s usefulness.

In brief, we’re seeking evaluations that:

  • State clearly what they are measuring (i.e., what counts as a pass or fail, and why it matters);
  • Involve clinical and subject-matter experts in the design and validation;
  • Test both precautions and harms (i.e., evaluate the risk of both overcompliance and overrefusal);
  • Reflect how users actually use AI (often, this means constructing scenarios that represent multi-turn conversations, where risk escalates and context shifts over the course of a long conversation);
  • Validate their graders against real subject-matter experts.

To learn more about the grant program and apply, see our application form. For more on building strong wellbeing evaluations and benchmarks, read our guidance. Applications are due by September 21; applicants who are selected to submit full proposals will be notified by October 5.

Related content

Expanding the Cyber Verification Program

We’re launching a new, expanded version of our Cyber Verification Program, which makes advanced cyber capabilities and reduced blocking classifiers available to qualifying security professionals.

Read more

Anthropic invests $100 million to train 10,000 engineers and tackle the enterprise AI talent gap

Anthropic is investing $100 million in Claude Frontier Academy to train 10,000 Frontier Deployed Engineers by the end of 2027, with cohorts from Accenture, Bain, CBA, Deloitte, McKinsey, Morgan Stanley and Novo Nordisk already underway.

Read more

Barclays scales Claude to upgrade operations and improve client experience

Barclays, the British universal bank, is expanding its strategic collaboration with Anthropic to integrate secure, enterprise-grade AI systems across its global operations.

Read more