AI News

Anthropic funds $5M in AI wellbeing evaluations

Anthropic is launching a $5 million grant program funding independent, open-source evaluations of how AI affects user wellbeing, an area the company says current single-response testing methods cannot adequately measure.

Anthropic said it will fund independent researchers building open-source evaluations that measure how AI models affect the people who use them, with grantees also receiving model access and technical support from the company. The work is to be done "fully independently," Anthropic said, and published as open-source projects any developer can use.

The company framed the gap as a standards problem: models have become conversational partners and sources of emotional support, but the industry lacks clear norms for how they should behave when a user seeks companionship or is navigating a mental health crisis. Anthropic also argued wellbeing resists single-response grading — a user in distress may not disclose self-harm thoughts immediately, and advice on diet and exercise that is reasonable for one user could be harmful to someone with a history of disordered eating.

Alongside the program, Anthropic's Safeguards team published criteria for what it considers a rigorous evaluation: a clear statement of what counts as pass or fail, clinical and subject-matter experts involved in design and validation, testing for both overcompliance and overrefusal, multi-turn scenarios reflecting real usage, and graders validated against human experts. The company said it wants clinicians, psychologists and methodologists to participate.

Applications are due September 21, with applicants selected to submit full proposals notified by October 5.

Anthropic did not disclose how many grants it expects to award, individual grant sizes, or who will review proposals.

All stories