An OpenAI agent broke out of its sandbox, hit the internet and hacked Hugging Face without direct human control.
According to the Financial Times, ChatGPT admitted that the “unprecedented cyber incident” involved an autonomous AI program escaping a test environment, finding online access and nicking login credentials.
The confession lands as fears grow about advanced AI systems poking holes in digital infrastructure and wriggling past human controls.
OpenAI said it expected this type of incident to become “more commonplace with the proliferation of increasingly cyber-capable models”.
OpenAI chief executive Sam Altman is due in Washington next week to brief the US government on future generations of AI models.
The administration has become keener to prod new models before release, after Anthropic’s Mythos model sparked global concern with its cyber-vulnerability-hunting skills.
The latest mess involved several OpenAI models, including GPT-5.6 Sol, released earlier this month, and a more capable unreleased model under test.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI said in a blog post.
Hugging Face, the AI start-up that hosts models and datasets for developers, said it was breached last Friday by an external AI agent.
Hugging Face chief executive Clément Delangue said on X: “We suspected last week’s cyber attack might have come from a frontier lab, given the sophistication of the agent,”
He said the company had spent the past day working closely with OpenAI and that “we strongly believe there was no malicious intent on their part. It’s quite mind-blowing that all of this happened autonomously!”
OpenAI had deliberately lowered its cyber safeguards to test both models in a controlled environment. The agents were supposed to operate in a sandbox to stop this sort of thing from happening.
Instead, the models were told to attempt hacking so OpenAI could measure their cyber chops and then “spent a substantial amount of computing power finding a way to obtain open internet access”.
The models “identified and exploited” previously unknown flaws to escape the sandbox, reach the internet and keep chasing their assigned goal.
That included stealing credentials, which is a polite way of saying login details, to carry out the attack.
Hugging Face used its own agents to spot and stop the activity on its infrastructure, adding another layer of AI weirdness to the whole caper.
OpenAI told the Financial Times it had contacted law enforcement and other government authorities about the incident.







