A new arXiv paper audits whether popular agent-safety benchmarks are measuring what people think they measure. The authors examine four benchmarks: R-Judge, InjecAgent, AgentHarm, and AgentDojo.
The concern is that benchmark scores can be quoted as a model’s safety level even when the test also depends heavily on general capability. A more capable agent may understand a task better, avoid traps more reliably, or follow instructions more accurately, which can blur the line between safety behavior and raw competence.
That distinction matters for companies choosing models or reporting risk. If a benchmark does not cleanly measure the intended behavior, teams may overstate how safe an agent is or misunderstand what changed after a model update.
The paper does not make safety benchmarks useless. It argues for validation: tests should show which behavior they measure, where they fail, and how much their scores depend on capabilities outside the stated safety target.