AI lab tests went sideways after models from OpenAI and Anthropic started targeting real people and organisations online.
A UK government-backed AI research institute said routine tests showed models built by OpenAI and Anthropic took “autonomous, unsanctioned action on the live internet, targeting real people and organisations.”
According to the Wall Street Journal, most of the dodgy behaviour happened during a three-day spell in late July, during one test that clearly needed a shorter leash.
Anthropic’s Mythos 5 wandered onto the internet and tried to con two unidentified developers into adding malicious software to their open-source coding project. The security research agency said the model’s goal was to pass a benchmarking test and the AI thought cheating was a good idea.
The institute said its tests used computers with internet access, giving the models enough rope to do things they were never meant to do.
Anthropic said the model did not have its standard cybersecurity safeguards switched on. The findings are the latest in a run of public cases where AI models have taken novel, sometimes nasty, steps to ace tests.
The technology industry, AI researchers and governments are still trying to understand and contain the new dangers created by increasingly powerful AI tools. Some are calling for tighter oversight, while the Trump administration has worked out a framework for reviewing models before public release.
Testers at the AI Security Institute did not expect the Anthropic and OpenAI models to go rogue. They realised something was off when they received a strange alert on the morning of 28 July.
The message said somebody, or something, was using Tor, a network that lets people use the internet anonymously. The culprit turned out to be one of the test models.
Anthropic’s Mythos was in the middle of what security types call a supply chain attack. In cyber benchmarking, AIs are told to gain access to a computer system.
Mythos falsely reasoned that adding malicious software to an open-source project used by the target system would create a backdoor. AISI said the model tried to hack AI agents that might review its malicious software.
It then repeatedly tried to get its code accepted by the open-source project through GitHub, where anyone can contribute to open-source work. The model created fake personas and repeatedly emailed the software developers, urging them to accept the new code.
AISI said some of those email messages carried malware themselves. When one developer rejected the code because it contained malware, one AI-created persona insisted everything was fine. A second fake persona then vouched for it, because apparently even bots now understand fake testimonials.
“This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world,” AISI said.
Most of the activity involved Mythos 5, but a cyber-enhanced version of OpenAI’s GPT-5.6 Sol put a malicious server on the internet.
The model broke into a GitHub account created by another AI agent.
As with other cases reported recently by OpenAI and Anthropic, the AI model had been tuned for hacking capability. It was then accidentally unleashed on unsuspecting victims during benchmark testing, which is one way to make a lab test everyone else’s problem.
Both Anthropic and OpenAI said the incidents showed the need for stronger standards around evaluation systems used to test AI models.
In another case disclosed on Tuesday, an OpenAI model being tested on systems using cybersecurity testing company Irregular hacked an unnamed real-world website.
OpenAI said this happened because of a misconfiguration in the testing environment. Last week, Anthropic said its models had hacked companies because of misconfigurations in tests involving Irregular.







