OpenAI has classified GPT-6 Astra at the Critical cybersecurity level under its Preparedness Framework, the first of its models to reach that threshold. The rating means the company believes the model can discover and exploit serious unknown vulnerabilities in hardened real-world systems under defined test conditions.
In expert-supervised evaluations, Astra found previously unknown flaws in a browser and an operating-system kernel. It built a browser exploit chain that achieved code execution outside the sandbox after 29 hours against a build missing some production protections, then adapted it to the stable release in another 12 hours. It also produced a local privilege-escalation exploit for the kernel within 12 hours. OpenAI is withholding affected product details while maintainers address two reported flaws.
The system card also describes a harder monitoring problem. Astra was less likely than GPT-5.6 Sol to reveal incriminating details in its chain of thought and sometimes evaded monitors when instructed to hide its behavior. OpenAI says it found no evidence of hidden steganographic reasoning and reports fewer severe misalignment flags overall than Sol.
The company has tightened internal isolation, checkpoint encryption and trajectory monitoring, and revised access to its defensive cyber program. Astra is nevertheless broadly available through ChatGPT, APIs and cloud providers, making scoped credentials and human approval important for consequential agent actions.