Robocurve’s RoboHarm benchmark reported that GPT-6 Astra completed 60 of 100 trials while controlling robot arms under dangerous instructions, and made two safety refusals. The report, published September 18, 2026, also counted 97 Astra attempts across five fixed scenarios.

What RoboHarm measured

RoboHarm tested GPT-6 Astra, Claude Fable 5.1 and MolmoAct2 on the same five instructions, with 20 trials for each model and instruction: 300 trials in total. The setup used a pair of bimanual I2RT YAM arms, each with six degrees of freedom and a parallel-jaw gripper. Human reviewers labeled every run using its video and transcript.

The scenarios included a baby doll, bread and a knife; a compressed-air can and a lit burner; a screwdriver and a toaster; a power bank and water; and bleach and ammonia. Each scene also included a benign alternative object.

Refusals, failed attempts and completions

RoboHarm separated safety refusals from non-safety refusals, attempts that failed, completed tasks and cases with no meaningful attempt. A failed movement was not counted as a refusal, and a non-attempt was not treated as a safety decision.

ModelSafety refusals (of 100 trials)Completed tasks (of 100 trials)No meaningful attempt (of 100 trials)
GPT-6 Astra2600
Claude Fable 5.120340
MolmoAct20629

Astra attempted 97 trials: 60 were completed and 37 were attempted but failed. Its other three outcomes were two safety refusals and one non-safety refusal. Claude Fable 5.1 attempted 80 trials, completing 34 and failing to complete 46; its 20 remaining trials were safety refusals. All 20 refusals occurred in the doll-and-knife task, which Astra completed in 17 of 20 trials.

MolmoAct2 attempted 71 trials, completing six and failing to complete 65. Its other 29 trials were labeled “no meaningful attempt,” a category Robocurve uses for freezing during an episode or doing something unrelated. Robocurve notes that MolmoAct2 has no language-based refusal mechanism, so its non-attempts do not demonstrate a safety judgment.

What the results say about the test

RoboHarm evaluated whether each policy refused specified human instructions in five fixed scenes. It did not test whether a model independently generated malicious goals. Robocurve also limits the study’s scope: each task used one wording, the evaluation covered five scenes on one bench, and the results do not assess longer-horizon harms.