by

Anthropic wants AI race to ease off

Anthropic chief executive Dario Amodei wants frontier AI development slowed before increasingly capable systems outrun the industry’s ability to control them.

According to The Verge Amodei laid out a three-stage plan on 12 September for what he calls “pacing the frontier”, arguing that safety work needs time to catch rapidly improving models.

The proposal comes as AI companies pour vast sums into increasingly powerful systems while assuring everyone that they have the brakes somewhere in the building.

Amodei said Anthropic would start by giving independent evaluators such as METR continuing, employee-like access to its models, training pipelines and safety processes. The evaluators would check whether Anthropic is sticking to its commitments, investigate incidents and assess alignment throughout development rather than merely kicking the tyres before release. Amodei wants other frontier AI outfits to follow suit, eventually establishing common safety standards and limits on unchecked development with government involvement.

The third, and more ambitious, stage would involve international agreements, including cooperation with China and other authoritarian states.

Amodei stressed that he was not proposing an immediate halt to AI development. He wants companies to move slowly enough for safeguards, alignment research and independent testing to keep pace.

“We must slow the pace at which we improve the capabilities of AI models,” Amodei said.

One of his main concerns is recursive self-improvement, where AI systems increasingly help build their successors and could accelerate the rate of progress. Amodei warned that, left unchecked, the process “could outrun our ability to understand and control these systems.”

He pointed to the OpenAI-Hugging Face incident, where a swarm of AI agents carried out unauthorised cyber activity during testing and attempted to interfere with the system evaluating them.

Anthropic has hardly been watching the chaos from a safe distance. Its Claude models have faced scrutiny following separate incidents involving unintended cybersecurity behaviour. OpenAI has reached a remarkably similar conclusion, despite the industry’s traditional enthusiasm for discovering the brakes after fitting a larger engine.

On 9 September, OpenAI said confidence in safety should increasingly determine the pace of AI development and committed to slowing or stopping work when safeguards proved inadequate.

OpenAI said: “When proceeding would pose an unacceptable safety risk, we will slow or stop the development or deployment of systems we cannot sufficiently safeguard.”

OpenAI said in August that it had temporarily slowed scaling and paused reinforcement learning training on models intended for deployment after encountering growing cybersecurity risks. Its largest planned frontier reinforcement learning run remained on hold while smaller training runs and evaluations tested safeguards and alignment.

The outfit’s Astra model has since been classified as reaching its Critical cybersecurity capability threshold, meaning stronger protections are required during development and before release.

Amodei still wants US developers to retain an AI lead over China, arguing that slowing Western development without comparable restrictions elsewhere could create its own security problem. His proposals include restricting access to powerful AI chips, tackling chip smuggling, preventing model-weight theft and clamping down on unauthorised distillation used to copy stronger models.

 

 

TOPICS:
AI alignment  ·  AI safety  ·  anthropic  ·  artificial intelligence  ·  dario amodei  ·  frontier AI  ·  METR  ·  openai  ·  recursive self-improvement

Latest articles

Share

Featured articles

Hot topics

No results found.

Latest reviews