A new NLP preprint examines how citations affect preference data, the pairwise judgments often used to train and evaluate language models. Citations can help readers verify an answer and reduce unsupported claims, but their presence may also influence a reviewer before the underlying evidence has been checked.

The researchers focus on how humans and language models compare attributed outputs. That distinction matters for reward modeling and post-training systems, where a preference label can teach a model which style of answer to reproduce. If evaluators reward citation formatting rather than whether a source actually supports the claim, the training signal can favor answers that look well grounded without being more reliable.

The paper frames attribution as more than a presentation feature: it is part of the evaluation process itself. Its central practical implication is that preference-data pipelines should test citation correctness and support separately from general answer quality, rather than assuming that an answer with references is automatically better.

This is a newly announced preprint, and the available abstract does not establish that one evaluation protocol solves the problem across domains. The findings should be treated as research evidence about evaluator behavior, not a universal score for citation quality.