A new arXiv paper takes a closer look at how to evaluate lie detectors for language models.

The authors argue that useful testbeds must verify that a model internally believes something different from its outward answer. Without that condition, detector performance can be misleading.

The study introduces belief-verified reasoning model organisms and a broader prompted-lying testbed, making the work relevant for AI monitoring, audits, and post-hoc investigation.