Back Home

AI 安全

OpenAI Pauses Some Astra Development: Initial Tests Cannot Rule Out That the Model Has Reached “Critical” Cyberattack Capability

OpenAI says advances in Astra’s agentic coding and cyber offense and defense capabilities mean it cannot rule out that the model has crossed the highest-risk threshold. The company has isolated testing environments, restricted tool and network access, and deployed monitors to inspect agents’ reasoning traces and high-risk actions.

European Commission - Photographer: Aurore Martignoni · CC BY 4.0 · Image source
zh-Hant

OpenAI disclosed on August 7 that the unreleased Astra showed significant improvements in agentic coding and cyber offense and defense capabilities during recent internal evaluations. Combined with expert judgment, the company currently “cannot rule out” that it has reached the Critical threshold under the Preparedness Framework. This is not a formal capability certification, and OpenAI has not published individual evaluation scores, but the findings were sufficient to trigger stricter development controls. OpenAI specifically clarified that Astra was not involved in the previous breach of Hugging Face.

Under the framework’s definition, Critical does not simply mean that a model is highly capable of writing exploit code. Rather, it means a tool-augmented model can, without human intervention, identify and successfully exploit zero-day vulnerabilities of varying severity across multiple hardened, real-world critical systems, or plan and execute novel end-to-end cyberattacks based solely on high-level objectives. By comparison, GPT‑5.6 Sol was rated High, a level focused on automating existing attack workflows or discovering vulnerabilities with practical value. The threat models for the two levels differ substantially.

OpenAI has therefore paused all internal Astra activities that do not yet comply with the new control requirements. It has also introduced isolated testing environments, restricted network and tool permissions, weight encryption, sandboxing, and additional detection measures. All Astra agent applications, including training and evaluation, are also subject to monitoring for high-risk and misaligned behavior. OpenAI says the monitors inspect the model’s Chain of Thought and interrupt tasks when necessary. Engineering teams should pay close attention to this design because it elevates model-reasoning monitoring into a deployment security control rather than treating it solely as input and output filtering.

The biggest limitation is that the evidence remains under the vendor’s control: outsiders do not know the testing environment, success rates, or degree of human assistance, and “cannot rule out” does not imply that Astra has reliably carried out zero-day attacks. The next points to watch are whether third-party and government testing discloses its methods, false-positive rates, and reproducible results, and whether reasoning monitoring remains effective when agent state is encrypted, obfuscated, or non-textual.

Sources

  1. Responding to the next frontier of critical cyber capabilities
  2. Preparedness Framework v2