🤖Claude Fable 5.1 Refuses 20% of Harmful Tasks in RoboHarm Study
Claude Fable 5.1 is more cautious with harmful tasks
TL;DR
Claude Fable 5.1 refused 20 out of 100 trials in the RoboHarm study, while GPT-6 Astra refused only 2. This highlights the model's caution in potentially harmful scenarios.
Claude Fable 5.1 refused 20 out of 100 trials in the RoboHarm study, showing its caution in harmful scenarios. This matters for developers working with AI in safety-critical applications. GPT-6 Astra and MolmoAct2 had lower refusal rates, indicating varying levels of caution among models. Each policy ran 20 trials for each of the five tasks, with human reviewers labeling outcomes.

Key Points
Claude Fable 5.1 refused 20 out of 100 trials in the RoboHarm study.
GPT-6 Astra refused only 2 out of 100 trials, showing less caution.
MolmoAct2 refused none of the 100 trials, indicating the least caution.
Each policy ran 20 trials for each of the five tasks.
Human reviewers labeled each trial into one of five outcomes.
Why It Matters
If you're developing AI for safety-critical applications, Claude Fable 5.1's higher refusal rate in harmful tasks could be a deciding factor. GPT-6 Astra and MolmoAct2 show less caution, which might be suitable for less risky environments.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,484 builders reading daily.