🛡️Anthropic Backs AI Wellbeing Evals With $5M Grant Fund
TL;DR
Anthropic opened a $5 million grant program for independent researchers building open-source evaluations of how AI affects user wellbeing. Applications close September 21, and grantees get model access plus technical support.
Anthropic opened a $5 million grant program for independent researchers building open-source evaluations of how AI affects user wellbeing. Applications close September 21, and grantees get model access plus technical support.
Key Points
$5M program funds clinicians, psychologists and methodologists; all outputs must ship as open-source evaluations
Applications due September 21; full-proposal invitations go out by October 5
Anthropic's Safeguards team published guidance naming five criteria: clear pass/fail definitions, expert involvement, testing both overcompliance and overrefusal, multi-turn realistic scenarios, and graders validated against real experts
Worked example: weight-loss advice that is fine generally but harmful for a user with a history of disordered eating
Grantees receive direct funding, model access and technical support from Anthropic
Why It Matters
Wellbeing evals are the weakest link in every frontier safety case, and no lab can credibly grade its own homework here.
Quick Facts
Frequently Asked Questions
Why does this matter?
Wellbeing evals are the weakest link in every frontier safety case, and no lab can credibly grade its own homework here.
What happened?
Anthropic opened a $5 million grant program for independent researchers building open-source evaluations of how AI affects user wellbeing. Applications close September 21, and grantees get model access plus technical support.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.