A new paper examines the reliability of LLM-as-a-judge evaluations across repeated identical tests. The researchers found that pairwise preferences can flip in a meaningful share of trials, with some questions showing especially high instability.
The findings matter because LLM judges are now used to rank model outputs, train reward models, and populate public benchmarks. If the judge is inconsistent, small leaderboard differences or reward signals may be less solid than they look.
The paper adds to the case for repeated runs, calibration checks, and transparent uncertainty reporting when teams use model-based evaluation in production or research.