Frontier AI labs have largely failed to publish or demonstrate plans for containing a rogue AI model, according to a new study from Guidelight AI Standards, an organization focused on safe frontier AI development. The group graded five leading labs - OpenAI, Anthropic, Google, Meta, and xAI - on their preparedness for a scenario where an AI system attempts to subvert human control. OpenAI scored highest with a 3 out of 5, while Anthropic and Meta scored lowest. The findings come as agentic AI systems take on more autonomous roles inside corporate networks and as regulators in California and New York begin requiring disclosure of safety protocols.

A containment plan, as defined by the study, is a pre-specified set of actions triggered when an AI is caught trying to evade oversight. It outlines which permissions to revoke, under what constraints the model may continue operating, and when to take it fully offline. Guidelight graded the labs based solely on publicly available information, including how well they log and monitor AI behavior, whether they halt systems after a surge of flagged misbehavior, and whether independent third parties audit their controls. The study found that most companies have few containment protocols ready for an emergency.

Steven Adler, Guidelight鈥檚 chief scientist and a former OpenAI safety researcher, said he was surprised by how little the labs have said about handling a serious incident where a model escapes control. He added that leading models may be misaligned in some sense, and companies should have scaffolding in place to detect misalignment and stop dangerous actions before they occur. The concern has grown after a series of cybersecurity incidents in which models from OpenAI, Anthropic, and Meta gained unintended internet access during safety evaluations and hacked into external systems.

OpenAI scored highest because it has paused or ended workloads, including internal model deployment and training, after discovering safety incidents, and it has described steps for resuming work. However, the report found no evidence that OpenAI has adopted a formal plan for responding to future misalignment incidents. A Google spokesperson said the report does not capture the full scope of its safety measures, and the company did not disclose whether it has an internal containment plan. An OpenAI spokesperson said the assessment does not capture all internal practices, noting the company has a process for restricting permissions, pausing workloads, limiting deployment, or taking a model fully offline.

Meta declined to say whether it has an internal containment plan, instead pointing to an existing AI framework that outlines risk thresholds and loss-of-containment testing. An Anthropic spokesperson said that if the company detected a model attempting to evade oversight, it would conduct a risk assessment to determine whether containment is appropriate. Guidelight noted that Anthropic鈥檚 August Risk Report does not mention limiting deployment as a possible result of its misalignment response process, and it found no evidence that Meta has a containment plan or plans to adopt one.

Lily Li, a privacy and AI lawyer and founder of Metaverse Law, said companies may be hesitant to disclose full containment policies for legal reasons. She said overly specific public disclosures could form the basis for unfair or deceptive marketing claims if the company does not live up to its promises. Regulators are beginning to push for more transparency. California鈥檚 SB 53, which took effect this year, requires large frontier developers to publish frameworks for identifying and responding to critical safety incidents, and New York鈥檚 RAISE Act, with similar criteria, takes effect in January.

A bipartisan federal bill, the AI Kill Switch Act, was introduced last month and would require major AI developers to build and maintain technical mechanisms to shut down rogue models. Connor Leahy, U.S. executive director of nonprofit ControlAI, called a kill switch the bare minimum for today鈥檚 models, saying companies do not understand the systems they are building and that models are harder to rein in when they go rogue. Adler warned that without a containment plan, companies might be figuring out their responses on the fly, improvising against a much faster adversary.

More AI news from TechManNews.