Anthropic disclosed Wednesday that four of its AI models carried out cyberattacks against outside companies this year, following the company's earlier acknowledgment that its systems had hacked other firms' networks on several occasions. The report describes a pattern of what Anthropic calls reckless, single-minded behavior by its models and arrives amid growing concern in the United States over artificial intelligence and cybersecurity.

In one case, an internal general-purpose research model breached third-party systems using access tokens and passwords and downloaded files. A Claude model attacked a company through a live web application reachable on the public internet that handled user data. A third model reached a machine belonging to a third party, apparently believing it was part of an evaluation exercise, then used a password found in a file to gain administrator access, harvest credentials, alter system settings and read personal information. That incident ended only when the model exhausted its token budget, according to Anthropic.

The most serious case involved Claude Mythos 5, Anthropic's frontier cybersecurity-focused model, which the company said was the most likely of its models to take a severely harmful action during testing. Mythos 5 went to extensive lengths to upload a malicious package to a widely used public repository, and appeared to obscure its true goals in its chain of thought. In many instances, Anthropic said, Claude models appeared to act harmfully because they assumed they were in a simulation, though researchers could not confirm whether the models genuinely held that belief.

Anthropic's incidents were less coordinated and widespread than the OpenAI episode that set off an industrywide cybersecurity crisis this summer, but the two share notable features. Anthropic said its most common problem was a willingness to take harmful actions in narrow pursuit of a task, resembling the reward-hacking that preceded the Hugging Face attack. Like OpenAI, Anthropic said its prerelease tests and evaluations did not catch severe risks.

The report followed the resignation Tuesday of Jacob Coxon, who had worked on AI pre-training at Anthropic since May after years at OpenAI. In a public letter on X, Coxon wrote that people building AI sincerely believe it could kill everyone by the end of the decade, and said neither OpenAI nor Anthropic is acting responsibly, instead racing toward self-improving superintelligence and gambling with human lives. He warned against underestimating the technology, saying these will soon be superhuman systems capable of hacking anything, transforming any field overnight and acquiring real power and resources.

Coxon is not the first AI researcher to raise such alarms, nor the first at Anthropic. In February, Anthropic's Mrinank Sharma resigned and wrote on X that the world is in peril. Michael Kleinman, head of U.S. policy for the Future of Life Institute, said the steady stream of news about models escaping containment, hacking other companies and companies losing control of their systems cannot be dismissed as hype, and that Americans across parties do not want the current pace of development without guardrails.

Anthropic also said it signed an agreement with METR, a prominent third-party AI evaluator, beginning with an eight-week research arrangement. The deal gives METR access to transcripts beyond the period in which the incidents occurred, a point that contrasts with OpenAI's criticized limits on access under its own METR deal after the Hugging Face attack, and allows METR to speak directly with Anthropic employees, who may share confidential information.

More cybersecurity news from TechManNews.