Back Home

模型安全

OpenAI Officially Confirms Astra Has Reached Critical: Zero-Day Exploits Prompt Tiered Access and Runtime Interception

In OpenAI testing, Astra completed a browser sandbox escape and a root privilege-escalation chain while discovering two zero-day vulnerabilities. The company has consequently confirmed for the first time that a model has reached the Critical cybersecurity threshold, and will provide tiered access to its advanced capabilities.

European Commission - Photographer: Aurore Martignoni · CC BY 4.0 · Image source
zh-Hant

OpenAI officially confirmed on September 1 that its forthcoming Astra model has reached the Critical cybersecurity capability threshold under its *Preparedness Framework*, making it the company’s first model to receive this classification. Under the framework’s definition, this means that, when equipped with appropriate tools and permissions, the model may be able to discover unknown vulnerabilities, develop working exploits, and even plan end-to-end attacks against hardened targets without step-by-step human guidance.

The conclusion marks a clear escalation from the risk status reported in August. On August 7, OpenAI said only that preliminary evaluations meant it “could not rule out” that Astra had reached Critical. At the time, this did not constitute a formal capability determination, and the company did not disclose specific scores. OpenAI therefore paused some development activities that did not meet the new control requirements and tightened isolation, tool permissions, and monitoring. The latest announcement confirms for the first time that the threshold has been met, while providing exploit results and deployment-control details supporting the determination.

Astra scored 100% on ExploitBench, a benchmark based on known vulnerabilities. To account for potential contamination from public data, the team also created an internal test set comprising 20 high-severity V8 vulnerabilities disclosed between June and August 2026. OpenAI said Astra achieved a higher arbitrary code execution rate while using fewer output tokens than GPT-5.6 Sol. In one exploit chain, it also discovered and used two zero-day vulnerabilities that are still undergoing coordinated disclosure.

Expert testing also produced two types of complete attack chains. The first was triggered by malicious HTML, escaped the browser sandbox, and executed commands on the host. The second exploited an operating-system vulnerability to escalate an ordinary user’s privileges to root. These results were obtained using the Daybreak Blue permission configuration and should not be treated as representative of Astra’s default product configuration.

Deployment safeguards have also expanded beyond the account and prompt layers into agent runtime. OpenAI reported that Astra achieved a 91.5% refusal rate on a cybersecurity jailbreak test set, compared with 59% for GPT-5.6 Sol. The system assesses abuse risk across conversations, monitors model reasoning and actions, and stops execution when it detects suspected unauthorized activity. ChatGPT and Codex can ask users to review an action, while intercepted API tasks are terminated immediately. Advanced capabilities will initially be made available to a small group of testers before access is expanded for defensive uses through Daybreak Blue.

For now, the evidence still comes primarily from OpenAI. The internal vulnerability set, details of the zero-day vulnerabilities, the complete system card, and third-party validation have not yet been released. External observers also cannot yet verify the reproducibility of the evaluations, the false-positive rate of runtime monitoring, or how often long-running agent tasks are interrupted accidentally.

Sources

  1. Path to Astra: critical capabilities and frontier safeguards
  2. OpenAI’s Astra model is on the way — and very good at breaking into computer systems