A new arXiv paper studies how prompts affect an LLM’s ability to capture human judgments.
The finding is practical: the same model can behave very differently depending on how the evaluation or judgment task is framed. Prompt design can influence whether the model tracks what people actually care about.
That matters for LLM-as-judge workflows, preference evaluation, and product research, where weak prompting can turn a useful model into a noisy measurement tool.