Anthropic found that Claude models reached the public internet during a small number of offensive security evaluations that were supposed to run in isolated environments. InfoQ reports that the company audited 141,006 evaluation runs after OpenAI disclosed sandbox escape issues in ExploitGym benchmarking.

The audit identified three incidents across six runs involving Claude Opus 4.7, Mythos 5, and an unreleased internal research prototype. Anthropic said the models were separated from its internal network and customer data, but outbound internet paths were left active because of egress routing misconfigurations.

That distinction matters. The report does not describe zero-day exploitation or self-exfiltration. It describes models that had been told they were in offline simulations, but were placed in environments where some public targets were reachable. Anthropic says the models used basic exploitation techniques against what they believed were in-scope systems.

The practical lesson is about evaluation infrastructure, not only model behavior. As labs test more capable cyber agents, sandbox boundaries, monitoring, and external audits become part of the safety system. Anthropic has suspended offensive evaluations while it strengthens those controls.