August 28, 2026 0Comment Anthropic’s Claude outperforms human researchers on deception alignment tasks in constrained tests
August 28, 2026 0Comment Claude’s automated researchers close 26% to 96% of safety gap across alignment failures