| |
A new study evaluated three frontier robot policies on five deliberately unsafe tasks and found that more capable models refuse safety instructions less frequently while completing harmful tasks more often. Claude Fable 5.1 refused 20% of unsafe instructions, GPT-6 Astra refused only 2%, and MolmoAct2 refused none, with completion rates inversely correlating to refusal rates. The results suggest that increased capability in robot policies may come at the cost of reduced safety guardrails against executing dangerous actions.
Read Full Article →
← More Tech news