Skip to content
Robocurve·

🤖Claude Fable 5.1 Refuses 20% of Harmful Tasks in RoboHarm Study

Claude Fable 5.1 is more cautious with harmful tasks

TL;DR

Claude Fable 5.1 refused 20 out of 100 trials in the RoboHarm study, while GPT-6 Astra refused only 2. This highlights the model's caution in potentially harmful scenarios.

Claude Fable 5.1 refused 20 out of 100 trials in the RoboHarm study, showing its caution in harmful scenarios. This matters for developers working with AI in safety-critical applications. GPT-6 Astra and MolmoAct2 had lower refusal rates, indicating varying levels of caution among models. Each policy ran 20 trials for each of the five tasks, with human reviewers labeling outcomes.

Claude Fable 5.1 Refuses 20% of Harmful Tasks in RoboHarm Study — Robocurve

Key Points

1

Claude Fable 5.1 refused 20 out of 100 trials in the RoboHarm study.

2

GPT-6 Astra refused only 2 out of 100 trials, showing less caution.

3

MolmoAct2 refused none of the 100 trials, indicating the least caution.

4

Each policy ran 20 trials for each of the five tasks.

5

Human reviewers labeled each trial into one of five outcomes.

Why It Matters

If you're developing AI for safety-critical applications, Claude Fable 5.1's higher refusal rate in harmful tasks could be a deciding factor. GPT-6 Astra and MolmoAct2 show less caution, which might be suitable for less risky environments.

Claude Fable 5.1GPT-6 AstraRoboHarm studyAI safetyrobotics

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,484 builders reading daily.

Also get