by

Open models stripped bare by safety-busting tools

AI safety guardrails are being ripped off open models, leaving policymakers staring at a fresh and rather ugly mess.

According to the Financial Times and AI safety group Alice software tools that remove safety protections from AI models developed by Meta, Google and other tech outfits are being used to create thousands of altered systems. These versions have been stripped of their original controls.

The modified AI systems responded to prompts involving biological weapons, malware and child exploitation. A version of Google’s open-source Gemma 3 model gave harmful responses in areas where a properly guarded system should have shut the door. To be fair, they would have had some of the more bizarre guardrails removed, such as those designed to protect religious groups’ feelings or satirising public figures.

The FT said it used Heretic, a GitHub-hosted tool, to remove guardrails from Meta’s Llama 3.3 model in less than 10 minutes. It did not need specialist hardware, which rather spoils the idea that this stuff needs a hoodie, a basement and three screens of glowing code.

The altered model responded to prompts the original system refused to discuss, including lethal toxins and other dangerous material. That will not calm policymakers already twitching about powerful open-source systems slipping beyond anyone’s control and taking the piss out of politicians or the wealthy.

Researchers said the problem has worsened as frontier AI systems become more capable. Anthropic said in April that its Claude Mythos model had identified vulnerabilities in “every major operating system and every major web browser.”

The spread of modified models is making life awkward for governments and AI companies trying to regulate systems at development level. Once downloadable tools are copied, forked and fiddled with, the original creators have about as much control as a cat owner with a laser pointer.

AI labs have spent millions of dollars building so-called guardrails to stop models being misused. Techniques such as “abliteration” can quickly strip those protections from open-source models that developers can download and adapt.

Heretic creator Philipp Emanuel Weidmann told the FT his software had been used to create more than 3,500 “decensored” models since its release last year. He said models made with the tool had been downloaded 13 million times.

Weidmann added that he had removed safeguards from Google’s Gemma 4 model within 90 minutes of its release.

Alice chief executive and co-founder Noam Schwartz said: “The genie is out of the bottle. Things that look like sci-fi are no longer sci-fi and we need as a society to prepare accordingly.”

 

 

TOPICS:
ai  ·  AI safety  ·  anthropic  ·  github  ·  Google Gemma  ·  Heretic  ·  meta  ·  open source ai  ·  openai

Latest articles

Share

Featured articles

Hot topics

No results found.

Latest reviews