A new paper, Beyond Goodhart’s Law, focuses on compliance evaluation for multi-agent systems. The authors argue that as LLMs become execution-capable agents, static benchmarks are too easy to overfit or game.

The proposed dynamic benchmark is aimed at testing whether agents continue to follow constraints as environments and incentives change. That is closer to how deployed agent systems actually fail.

The work reflects a growing evaluation problem: agent safety cannot be reduced to whether a single model answers a prompt correctly in isolation.