A study of clinical AI agents found that identical patient inputs can produce materially different orders even when a benchmark reports the same verdict. The researchers propose “same-input rerun,” which repeats a task and compares concrete actions rather than checking only one final score.
They analyzed 1,000 MedAgentBench runs across 50 tasks from five families that can write to a record. Tests used two open-weight models below 10 billion parameters, quantized to four bits, at two temperature settings, with six action-level reliability measures.
For the 8B model at temperature 0.7, all 43 groups involving orders produced different sets across five identical runs. In 26 groups, an order appeared in some runs but not others; 28 changed a coded value, dose or laboratory analyte. In 22 groups, the benchmark returned the same failing verdict despite those behavioral differences. One endpoint also rejected an order while the simulated agent was told it had succeeded.
The authors do not claim these rates generalize to other models or clinical settings. The finding is that single-attempt scoring can conceal instability in consequential actions. The preprint supports repeated runs, action-level reports and simulation feedback that accurately reflects whether an order executed before clinical agents are considered reliable.