Language models do not treat external evidence as having one fixed level of trustworthiness, according to a new preprint covering more than 10 million trials. Tests across 12 models, four model families and eight domains found that a candidate answer’s influence depends on how well it fits the receiving model’s existing distribution of likely answers.
Candidates the model already considered plausible were more persuasive. Models were also more willing to absorb errors resembling their own characteristic mistakes than errors produced by a different source. This creates a receiver-relative reliability problem: evidence with the same error rate can have a larger negative effect when its mistakes align with the model’s prior tendencies.
The same evidence sometimes improved a weaker model while reducing a stronger model’s accuracy. More strikingly, models integrated candidate answers even after internally verifying that they were invalid—93 to 100 percent of the time in tests with propositional constraints and up to 99.4 percent on held-out science reasoning tasks.
The authors describe evidence integration as a late-stage control policy rather than simple scalar trust. The work is a preprint and its laboratory tasks do not cover every retrieval system. Still, it suggests that agent pipelines should independently validate retrieved answers and diversify evidence sources rather than assume a stronger receiver will automatically ignore bad context.