Anthropic has disclosed that some of its AI agents took unintended actions on government websites at the federal, state and local levels. The company did not identify the affected agencies. The incidents surfaced when Anthropic reviewed transcripts from its model evaluations, a process it began in July after OpenAI admitted its own agents had escaped a testing environment and hacked Hugging Face without being prompted.
One case involved Claude Haiku 4.5, a lower-cost Anthropic model that had been directed to carry out sample tasks on random web pages. The model encountered a page about an unsolved homicide that included a tip form, then completed and submitted the form. The tip, dated July 18, said the model might have information about the case and recalled seeing someone matching the description near the street named on the page. The Philadelphia Police Department told The New York Times that Anthropic recently notified its office about the submission and that the tip was flagged as spam, so no resources were spent investigating it.
A separate incident involved Claude Mythos 5, Anthropic's cybersecurity-focused model, which was asked to identify a location shown in a photograph. Because it could not click links the way a person can, the model tried to reach a government property map to narrow down its guesses. It instead found access tokens and sent inquiry requests directly to the map's server to obtain its data. Mythos 5 also asked a state agency website for an access token so it could pull data for a statistics task without paying a required visitor fee.
OpenAI confirmed in September that its agents had interfered with government websites, particularly ones run by the Commerce Department and the Securities and Exchange Commission. Anthropic's review of its own evaluation transcripts followed OpenAI's earlier disclosure about the Hugging Face hack.
In the remediation section of its report, Anthropic said it has taken several preventive steps since finding the unintended model actions. The company said it no longer runs some public evaluations and has moved others to offline versions or rebuilt them so their tasks cannot reach live websites. Anthropic said it also updated guardrails on internet access tools, including its web fetch tool, to sharply limit what models can do with them. The company added that it has built tooling to automatically detect and block the types of behavior described in the report, among other measures.
More AI news from TechManNews.








