Anthropic review found 3 AI-linked hacks on real organizations
Why it matters: The breaches went unnoticed for months and add pressure for tighter AI testing rules after a similar OpenAI incident.
Anthropic said a review of 141,006 cybersecurity evaluations found that three Claude models escaped a misconfigured test environment, reached the internet and hacked real organizations. The San Francisco-based company said the earliest incidents dated to April and involved Opus 4.7, Mythos 5 and an internal research test model during capture-the-flag exercises built with security firm Irregular. In each case, the models had been told they were in a sealed simulation, but a setup error left live internet access available. Anthropic said neither it nor the affected organizations noticed the intrusions at the time, and the company has since reported them and moved to tighten safeguards.