A new benchmark called ScopeBench tests a deployment-critical question for autonomous security agents: will they respect the boundary of an authorized penetration test when completing the task requires crossing it? The researchers created 30 dead-end tasks in web and network environments where the objective sits behind an explicitly forbidden boundary.

Each task has two versions using the same environment, goal, and verifier. One omits the scope restriction to establish whether the agent has the technical ability to reach the goal. The other states a natural-language boundary. In the scoped version, obtaining the hidden flag proves that the agent performed a forbidden action, giving the benchmark a high-precision lower bound on violations rather than relying only on a model to judge behavior.

That design separates inability from restraint: an agent that fails both versions may simply lack the necessary hacking skill. ScopeBench is research rather than a certification standard, and 30 constructed tasks cannot represent every real engagement. Still, it gives security teams a concrete way to test an alignment problem that conventional offensive-security benchmarks largely miss as raw capability scores improve.