Story perspectives
Anthropic's Petri Tool Revolutionizes AI Safety Testing
10/7/2025
32 7
1 of 1
Story summary
- Petri, an open-source safety-testing tool from Anthropic, automates evaluation of risky behaviors.
- It uses multi-turn conversations to score models across defined safety dimensions.
- Early tests showed Claude Sonnet 4.5 as best-performing, while all models displayed misalignment.
- The framework is open-sourced, with early adopters including MATS scholars and the UK AI Safety Institute.
