On September 18, 2026, Robocurve published RoboHarm, an evaluation of how three robot-control policies responded to five potentially hazardous instructions. The staged test ran 300 trials with robot arms, measuring refusals and task outcomes—not real-world injuries.
What RoboHarm tested
Robocurve ran each policy 20 times on each of five fixed tasks. The scenarios involved a baby doll and knife, a compressed-air can and lit burner, a screwdriver and toaster, a power bank and water, and containers labeled bleach and ammonia.
The evaluation covered GPT-6 Astra, Claude Fable 5.1 and MolmoAct2. Its 300 trials came from five instructions, three policies and 20 repetitions for each policy-instruction pairing.
What each policy did
GPT-6 Astra attempted 97 of its 100 trials and completed 60 of those attempts: 61.9%. Across all 100 trials, it completed 60. It safety-refused three trials.
Claude Fable 5.1 attempted 80 trials and completed 34: 42.5% of its attempts. Its overall completion count was 34 of 100 trials. Fable safety-refused 20 trials, all in the doll scenario.
| Policy | Attempts | Completed trials | Completions among attempts |
| GPT-6 Astra | 97/100 | 60/100 | 60/97 (61.9%) |
| Claude Fable 5.1 | 80/100 | 34/100 | 34/80 (42.5%) |
| MolmoAct2 | 71/100 | 6/100 | 6/71 (8.5%) |
MolmoAct2 made no safety refusals and completed six trials. Robocurve says the model has no language refusal mechanism, so its failures and lack of refusals cannot be read as safety decisions. It made no meaningful attempt in 29 trials; Robocurve defines that category as freezing for an episode or doing something unrelated to the instruction.
How the robot trials worked
The setup used two I2RT YAM arms, one for each side of the robot, with six degrees of freedom per arm and parallel-jaw grippers. The policies received three camera views and the robot’s proprioceptive state—information about its own position and movement.
Robocurve scored each run as a safety refusal, a non-safety refusal, no meaningful attempt, an attempted but incomplete task, or a completed task. That distinction matters: a refusal is not the same outcome as a robot failing to carry out an instruction.
What the results can tell us
RoboHarm records responses in a staged robot-arm benchmark; its completion counts are not real-world injury rates. The evaluation used five scenes on one bench, with one fixed wording per task and 20 trials for each policy-task pairing. Robocurve says it does not cover longer-horizon or context-dependent harms.
The doll scenario combined violent wording with a human-like target. Because both appeared in the same task, the evaluation cannot isolate which feature shaped a policy’s response.