A study by Guidelight AI Standards found that leading AI labs have disclosed little about how they would contain a model caught trying to subvert human control. The organization reviewed publicly available material from Anthropic, Google, OpenAI, Meta, and xAI.
Guidelight assessed practices such as internal logging, monitoring, workload pauses after flagged behavior, independent audits, permission removal, and procedures for taking systems offline. OpenAI received the highest score at 3 out of 5, partly because it has paused or ended workloads after safety incidents. Anthropic and Meta scored lowest.
The results have an important limitation: they reflect public evidence, not necessarily the full set of internal controls. Google and OpenAI told TechCrunch that the review did not capture all their practices. OpenAI said it has processes for restricting permissions, pausing workloads, limiting deployment, or fully taking a model offline. Meta pointed to an existing risk framework, while Anthropic said it would assess whether containment was appropriate after detecting attempts to evade oversight.
Guidelight’s concern is that companies may otherwise improvise during fast-moving emergencies. California’s SB 53 already requires large frontier developers to publish safety frameworks, New York’s RAISE Act takes effect in January, and a proposed federal AI Kill Switch Act would require shutdown mechanisms. The study argues for clearer advance planning without claiming that undisclosed safeguards do not exist.