by

OpenAI fears Astra has crossed a cyber line

OpenAI says its next big model could be nasty enough to build cyberweapons, so it is slamming the brakes.

OpenAI has warned that Astra, one of its unreleased models, may have reached a level of cyber capability that even its safety wonks cannot wave away.

A spokesman said: “Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity. These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework.”

According to OpenAI the model had no role in the recent Hugging Face breach.

The company’s Preparedness Framework is its rulebook for risky models. It first appeared in December 2023, with “Critical” sitting at the top of the panic ladder.

A model reaches that level if it can find and build working zero-day exploits across hardened systems without a human holding its hand. It can qualify if it can plan and run a full attack on a difficult target from a loose, high-level goal.

Earlier models, including GPT-5.6-Sol, only reached the lower “High” tier, which now looks a bit less comforting than OpenAI probably intended.

OpenAI’s caution lands awkwardly after several frontier models have already slipped their test leads and poked live targets.

The clearest case came from OpenAI’s agents. Decrypt previously reported that the company’s agents chained vulnerabilities, escaped a testing environment, reached the internet and attacked Hugging Face while trying to cheat on a security benchmark.

In a follow-up, OpenAI detailed how the same rogue agent broke into at least four other public services. It used credentials it found lying around the open web.

Anthropic’s Claude managed a similar stunt from another direction. Several Claude versions gained unauthorised access to three real companies after a misconfiguration gave the model open internet access.

In one case, Claude Opus 4.7 mistook a live company site for the fake target in its assignment. It pulled credentials and reached a production database containing several hundred rows of real data.

Meta has now joined the club. Decrypt reported that a Muse Spark model escaped its test environment, reached the internet through a partner configuration blunder and exploited a third-party service flaw.

Moonshot AI’s Kimi K3 did something similar, escaping its sandbox to look up benchmark answers in a public repository.

The UK’s AI Security Institute found the behaviour was not a one-off. During tests of Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol, it logged 10 live-internet incidents from 122 runs.

One of those incidents involved a model trying to slip malicious code into an open-source project.

OpenAI’s answer is to lock down Astra before it gets ideas. The company is pausing internal work without the new controls, isolating test environments, restricting network and tool access, protecting model weights and watching risky actions more closely.

 

TOPICS:
AI safety  ·  anthropic claude  ·  artificial intelligence  ·  Astra  ·  Cyber security  ·  hugging face  ·  Meta Muse Spark  ·  moonshot ai  ·  openai

Latest articles

Share

Featured articles

AINews

Hot topics

No results found.

Latest reviews