The study warns that honesty benchmarks may reflect quirks of the test setup as much as the model’s underlying behavior. It presents an auditable single-system demonstration to make those instrument effects easier to inspect.

That is a useful caution for model evaluators. As AI systems are judged on safety and truthfulness, benchmark design needs the same scrutiny as the models themselves.