Clinical language models can make an earlier decision look more obvious when their evaluation record includes what happened later, according to a new benchmark of temporal reasoning. That setup differs from real care, where a clinician acts without knowing the final diagnosis or treatment response.
The dataset contains 171 case reports from the PubMed Central Open Access collection: 40 involving sepsis and 131 involving GLP-1 treatment or diabetes. Each question is tied to a meaningful decision cutoff and includes both a prospective reference answer and an outcome-consistent “hindsight trap.”
Models answered using either a timeline truncated at the cutoff or the complete record. Across GPT 5.6 Sol, Gemma 4, GLM 5.2 and Opus 5, access to the full timeline produced consistent shifts toward answers influenced by later outcomes. Temporal masking reduced measured hindsight bias without reducing accuracy.
The benchmark separates accuracy from trap rate, answer instability and an explicit hindsight-bias rate. That matters because a model can appear accurate while relying on information unavailable at the moment being judged. The study is a new arXiv preprint and does not establish clinical safety. It does show why medical evaluations should enforce time boundaries rather than hand a model a retrospective chart and treat its answer as prospective reasoning.