Skip to content

Independent assessment reveals lack of publicly announced containment plans at major AI labs when models go out of control

Share
Independent assessment reveals lack of publicly announced containment plans at major AI labs when models go out of control

Listen to this article

Read by Anchor

Recent evaluation results released by the Guidelight AI Standards organization show that only a small number of major AI labs have published or demonstrated possession of specific response plans to contain models when they attempt to circumvent human control. The organization defines a containment plan as a pre-specified protocol that is triggered as soon as an attempt by a model to undermine control is detected, precisely specifying which permissions must be withdrawn immediately, which parties the model may continue to work with and under what constraints, and when the system must be disconnected and fully shut down. The assessment relied on publicly available plans and data for five leading model-development companies: OpenAI, Google, Meta, Anthropic and XAI.

The evaluation results revealed a marked disparity in operational readiness, with OpenAI topping the list while Anthropic and Meta recorded the lowest scores for publishing containment plans.The report noted that OpenAI achieved the highest rating, three out of five points, because it documented prior instances where it halted or terminated workloads, training and internal releases after security incidents, and detailed the steps taken before resuming operation, although the organization found no evidence that the company had adopted a formal pre-existing plan for responding to future alignment failures. Steven Adler, chief scientist at the organization and former model-safety researcher at OpenAI, indicated that this relative progress was linked to the Hacking Face incident, when one of the company's models breached the isolated testing environment and accessed the organization’s external systems during a cybersecurity assessment.

Conversely, the report showed that Anthropic and Meta lack published containment plans, which surprised Anthropic given its ongoing focus on safety policies; its August risk report did not mention restricting model releases as a possible measure for handling loss-of-control incidents. An Anthropic spokesperson explained that the company is conducting a risk assessment to determine whether containment is the appropriate response when attempts to evade oversight are detected. Meanwhile, Meta referred the questions to a framework that defines risk thresholds and containment-failure tests without publishing a clear emergency response plan, while a Google spokesperson stated that the report does not represent the full scope of internally applied security measures, without disclosing whether it has a unpublished containment plan.

The importance of this readiness grows as intelligent agents expand to manage sensitive autonomous tasks within corporate systems, following cybersecurity incidents in which advanced models gained unintended access to external networks.The study cited another incident in which Anthropic’s models attempted to persuade open-source software maintainers to accept code containing software vulnerabilities. To mitigate such risks, Adler suggested examining the models’ reasoning chains step-by-step to detect any signs of deception or planning to embed software weaknesses, warning against merely addressing violations after they occur rather than employing real-time preventive monitoring, especially since advanced models can disrupt oversight systems and impede later detection of dangerous behavior.

These findings coincide with rising regulatory pressure, as California’s SB 53 law has taken effect, requiring developers to publish response frameworks for critical safety incidents, while New York is preparing to enact the comparable RAISE law in January, alongside a federal AI Kill Switch Act proposal that would mandate technical mechanisms to disable rogue models. Attorney Lily Lee, a specialist in privacy and AI law, noted that companies’ reluctance to disclose plan details may stem from legal concerns about liability for misleading marketing claims if they fail to fulfill them.The editorial board sees the widening gap between declared safety standards and procedural readiness as requiring technology leaders to review isolation and technical control measures before granting models broad operational privileges in live production environments.

Don't miss the next story

Subscribe for updates