A new arXiv paper examines when and why LLM self-reports predict behavior.

Psychometric tests are often applied to models by asking them questions about traits or preferences. The problem is that a self-report may not tell us much unless it predicts how the model behaves in relevant situations.

The study is useful for evaluation because it pushes model psychology claims toward behavioral validation rather than treating questionnaire answers as facts.