An arXiv study questions whether clinician pairwise preferences are a reliable stand-in for clinical safety when evaluating medical language models. The work analyzes 26,804 judgments from more than 736 clinicians across 28 or more countries.
The authors used data from MOOVE, a clinician-led platform that collects blinded pairwise preferences alongside multi-criterion rubric ratings. They found that models favored in pairwise comparisons could still produce clinically meaningful failures in areas such as harmlessness and accuracy.
Those failures were also uneven across specialties, creating domain-specific no-go zones that aggregate leaderboards may hide. In other words, a model can look strong overall while being unsafe for a particular clinical context.
The lesson is that medical AI evaluation needs explicit safety rubrics, not only which answer a reviewer prefers. Preference data is useful, but it can reward surface quality while missing clinically important risk.