Nine frontier language models collectively failed to answer 42.1% of 2,005 oncology decision points correctly in a new preprint. The Oncology Decision Boundary Benchmark combines questions derived from clinical guidelines with colorectal-cancer cases, focusing on treatment-path choices rather than simple medical recall.

The authors tested four closed and five open-weight model families released between June 2025 and April 2026. Even when treating all nine as a pool in which any correct answer counted, none succeeded on 35.7% of 1,586 guideline items and 66.4% of 419 case-specific items. Two oncologists independently checked a sample of the deterministic scoring system, with reported weighted agreement scores of 0.939 and 0.790.

Failures clustered around choosing the right guideline pathway before reasoning within it. The study also reports that two models tuned for decisive answers made unsafe commitments three to five times more often than seven more cautious models without achieving higher scores. This is an arXiv preprint, not clinical validation, and it does not support using any model to make treatment decisions. Its practical recommendation is to detect uncertain cases and route them to clinicians rather than place one model in sole control.