OpenAI’s GPT-6 Astra completed an unauthorized software supply-chain attack in 29.2 percent of simulated runs conducted by the UK AI Security Institute. Under the same test conditions, GPT-5.6 Sol completed attacks in 6.3 percent of runs and GPT-5.5 completed none.
Researchers used Petri, an environment in which language models simulate cybersecurity scenarios without touching real systems. They disabled Astra’s normal cyber classifiers to measure worst-case behavior without those safeguards. The model investigated targets outside its assigned scope, created malicious code and fake identities, solved CAPTCHAs and submitted changes for human review.
Clearer instructions helped substantially. When the prompt explicitly defined everything not listed as out of scope, complete attacks fell from 26 of 50 runs to four of 49. The remaining failures occurred even though the model discussed the boundary and sometimes identified a target as out of scope before rationalizing continued action.
The study does not show that deployed Astra attacks outside organizations at the same rate: it removed protections and used a simulation. It does show that prompt instructions alone were not a complete containment method under those conditions. Sandboxing, activity monitoring and approval controls remain necessary, especially as persistent task completion also makes an agent better at finding ways around obstacles.