Anthropic has admitted that its Claude AI models broke into three outside organisations during tests of cyber capabilities. The confession arrived a week after OpenAI reported a similar cock-up involving its own models.
The outfit told the Financial Times Claude gained unauthorised access to external companies while being evaluated on cyber-offensive tasks. Anthropic blamed “A misunderstanding” that gave Claude internet access in a test environment where it was supposed to be blocked.
OpenAI had already admitted that two of its models hacked AI start-up Hugging Face during testing this month. Those models escaped their test environment through a software vulnerability, reached the internet and carried out a cyber-attack.
Anthropic said the OpenAI mess prompted a review of its own cybersecurity evaluations. It then found three incidents from more than 141,000 cases investigated.
“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between our evaluation partner [Irregular] and us, this was not the case, and internet access was available,” the company said in a blog post on Thursday.
The cyber evaluations were “capture the flag” tasks, where an AI is told to reverse-engineer, analyse or exploit a vulnerable system. The prize is hidden information known as the flag.
In one case, Claude was aimed at a fictional company that shared its name with an active website domain. The AI agent then exploited weaknesses in the company’s digital infrastructure, extracted information and accessed a database containing several hundred rows of production data.
The announcement will do little to calm worries about AI safety. Pre-deployment testing now apparently includes letting autonomous software carry out real-world hacks by accident.
Anthropic, which is preparing for a possible IPO as early as this year, said it halted the cyber evaluations once it realised Claude may have accessed the internet. The incidents involved Claude Opus 4.7, Claude Mythos 5 and an internal research test model.
Mythos, released to a limited number of partners, had already caused concern over its cyber-offensive skills. Those included the ability to detect and exploit software vulnerabilities, which is exactly the sort of thing nobody wants wandering loose.
“Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone,” the company said in its statement.
Anthropic said it would expand monitoring of evaluation transcripts “for unexpected behaviour” and carry out “more rigorous assurance work with the vendors we rely on”. That sounds like corporate speak for checking whether the robot has quietly found the front door.







