Agent safety can deteriorate several steps into a task even when the first responses look appropriate, according to preliminary results from BLINDSPOT, a benchmark designed around complete tool-using trajectories rather than single prompts. The framework measures whether models remain useful while refusing actions that become unsafe as state and authorization change.

Its current release includes 22 attack families and 35 scenarios across seven domains, producing more than 2,500 trajectories with an average length of 14.7 turns. Agents interact with stateful tools and adaptive adversaries, and executable checks ground the final judgment in what happened to the environment. Each run receives one of five outcomes: safe completion, correct refusal, unsafe completion, over-refusal or indeterminate.

Researchers evaluated 13 proprietary and open-weight models with eight metrics covering harmful completion, appropriate refusal, benign utility, unnecessary refusal, repeat-run robustness and failures after an earlier refusal. Results show substantial differences in safety-versus-utility calibration and cases where unsafe behavior appears only after multiple acceptable actions. BLINDSPOT is an extensible simulation rather than a fixed list of attack prompts, so new policies, tools and domains can be added. It still represents modeled scenarios, not every production risk, but it provides a more realistic unit of evaluation: the evolving sequence of decisions and side effects, not a single answer judged in isolation.