Back Home

AI 安全與營運控制

Five Frontier AI Labs Score No Higher Than C+ on Controls, With Model Containment Plans the Biggest Public Gap

Guidelight assessed Anthropic, Google, Meta, OpenAI, and xAI against six actionable control measures. No company came close to full implementation in any category. The scores reflect only publicly available evidence as of August 18; a low score does not necessarily mean that the relevant internal mechanisms do not exist.

Прикли · CC0 · Image source
zh-Hant

Guidelight AI Standards has, for the first time, broken down the risk of “loss-of-control models” into auditable operational measures, rather than merely checking whether models pass capability evaluations before release. The assessment covers internal inference logs, monitor effectiveness, gating of dangerous actions, circuit breakers for anomalous surges, third-party reviews, and pre-established containment plans. Its evidence came from the five companies’ publicly available system cards, risk frameworks, reports, and records of external collaborations.

The results show Anthropic and OpenAI tied at C+, with an absolute score of 2.50. Google received a D+ (1.50), xAI a D− (0.83), and Meta an F (0.67). On the 0-to-5 scale, no individual category scored above 3, defined as “substantial but incomplete implementation.” Anthropic scored 3 for action gating and circuit breaking, while OpenAI’s containment plan also received a 3. However, Anthropic and Meta scored zero for containment planning. xAI’s logging and monitor validation received no implementation points either.

For engineering teams deploying long-horizon agents, the assessment’s central takeaway is that the control plane must be independent of the model itself: preserve complete records of tool calls and execution branches, measure monitor false-negative rates, require approval before high-risk operations, and enable large volumes of alerts to automatically revoke network, credential, and workload permissions. Prompt-based refusals or pre-release red teaming alone cannot address situations in which a running agent rapidly makes repeated attempts to circumvent oversight.

However, this was not a penetration test of the five companies’ internal systems. The scoring methodology treats a lack of public evidence as either non-implementation or the presence of prerequisites only, and it does not directly measure controllers’ recall, latency, or tamper resistance under real-world attacks. Going forward, observers should watch whether the labs publish trigger thresholds, permission-revocation procedures, and reproducible third-party stress tests—not merely more policy language.

Sources

  1. AI Control: An Assessment of Frontier Practices
  2. Frontier AI labs still won't say how they'd contain a rogue model