A new arXiv paper argues that clinical LLM evaluation should focus more directly on deployment behavior.

The study trains a classifier to predict whether a future interaction in an electronic health record system will lead a user to reject the LLM response. That turns sparse real-world feedback into a practical evaluation signal.

The approach highlights a gap in static benchmarks: aggregate accuracy can miss the queries where a model is most likely to fail in front of clinicians.