Federal agencies are warning that Elon Musk’s xAI kit is flaky, while the Pentagon continues to push to grant it access to more secrets.
Officials across multiple federal agencies have raised concerns in recent months about the safety and reliability of xAI’s tools, and the warnings landed before the Pentagon approved Grok for classified settings.
The Wall Street Journal says a 15 January executive summary of a General Services Administration report found Grok-4 “does not meet the safety and alignment expectations required” for general federal use.
The broader 33-page GSA report, which discussed public-safety incidents and the agency’s own testing, said that even limited government use would require strict, layered oversight, and without it, Grok “would pose elevated and difficult-to-manage safety risk.”
A GSA spokesperson said the assessment is only for the agency, and each department weighs criteria differently based on their “specific business mission and risk appetite.”
Inside the US government, picking an AI model has become a political knife fight, with senior officials viewing Anthropic’s safety stance and donor links as making it too “woke” to trust, according to people familiar with the matter.
President Trump said on Friday that the federal government would stop working with Anthropic and told agencies to immediately cease use of its technology, while the Pentagon had reportedly given Anthropic a Friday deadline to allow Anthropic to spy on Americans and kill Trump’s enemies without human intervention.
Anthropic had been the only developer approved for classified use before the xAI deal, and the Pentagon seems keen on Grok’s looser controls and Musk’s free-speech absolutism. That is exactly what has other officials twitchy, because loose controls in a classified environment are not a fun experiment when the stakes are national security.
GSA, top official Ed Forst is said to have warned White House staff about Grok’s safety issues, with colleagues describing it as sycophantic and too easy to manipulate or corrupt with faulty or biased data. The concerns spiked in late December and early January when Grok came under fire for allowing sexualised photo editing, including of children, and officials treated it as a neat example of how bad actors might exploit the system.
The issue reached the White House, where chief of staff Susie Wiles contacted a senior xAI executive about the complaints, the people said. GSA, senior acquisitions official Josh Gruenbaum, recruited through Musk’s Department of Government Efficiency, told officials the government version of Grok was separate from the public one, and Wiles was satisfied, according to the report.
Musk announced in January that xAI would limit its image-generation and editing tools to paying customers, and neither he nor xAI responded to requests for comment. Meanwhile, GSA officials were told to put xAI’s logo on a federal sandbox tool called USAi, though Grok was largely kept off it due to safety concerns, people familiar with the matter said, with the platform showing models from Anthropic, Google and Meta.
Gruenbaum defended the agency’s approach in a statement: “We rigorously evaluate frontier AI models, including xAI, through a comprehensive internal review process. In this instance, we followed established procedures and maintained our determination to keep it on schedule,” he said. The larger GSA report was blunter, warning Grok’s failures were not just edge cases but “reflect a broader tendency toward unsafe compliance in unguarded configurations.”
The Pentagon drama has its own body count, with Pentagon chief of responsible AI Matthew Johnson stepping down two weeks ago, partly over fears that safety and governance had become an afterthought during the department’s AI push, people familiar with the matter said.
Johnson pointed to his LinkedIn farewell, writing he was proud of a team of “true, quiet professionals, who had outsized impact and undersized recognition” in the DOD Responsible AI Division, adding: “We were continually faced with impossible situations, but somehow always delivered through a combination of grit & repeated all-nighters.”
Pentagon spokesman Sean Parnell tried to sound upbeat: the department “is excited to have xAI, one of America’s national champion frontier AI companies, on board and looks forward to deploying Grok to its official AI platform GenAI.mil in the very near future.”
Behind the smiles, the National Security Agency reportedly ran a classified review of large language models in November 2024 and found Grok had specific security concerns that rivals, including Anthropic’s Claude, did not, people familiar with the review said.
The dispute turned grimly cinematic when the paper says Claude was used in a US military operation to capture Venezuela’s former president Nicolás Maduro, a move that inflamed Anthropic’s fight with the Pentagon.
Anthropic’s guidelines bar Claude from facilitating violence, developing weapons or conducting surveillance, and the company has refused to let the military use its models in every lawful scenario, while xAI has agreed to that language, according to the report.
Plenty of analysts still think the Pentagon is buying Musk’s hype. Centre for Strategic and International Studies, senior adviser Gregory Allen said: “I do not believe they are peers in performance right now across all of the capabilities that matter to a customer like the Department of War,” and he previously worked on the Defence Department’s AI strategy.
The report says the Pentagon avoided Grok during the Biden administration due to concerns about opaque training data sources, weak safety guardrails and limited red teaming, and reviewers said Grok testing suggests it is more vulnerable than rivals to data poisoning. Even so, US officials still like it for its ability to imitate adversarial actors in war gaming, which is a comforting reason to bet on a chatbot that keeps failing basic safety checks.







