A new arXiv paper introduces OncoTriad-QA, a benchmark for testing AI systems on patient-level cancer reasoning across multiple evidence streams. The dataset contains 86,100 semantic questions from 9,281 TCGA patient cases across 32 cancer cohorts.

The benchmark combines radiology, whole-slide pathology, somatic mutations, copy-number alterations, DNA methylation, RNA sequencing, and clinical metadata. That breadth matters because real oncology decisions rarely depend on a single image or text note.

The authors also introduce OncoVLM, a reference multimodal model that connects evidence from different medical modalities to a language-model interface through learned projectors. The goal is to test whether systems can reason across a patient’s combined record.

The work is research, not a clinical product. Its importance is evaluation: medical AI systems need benchmarks that reflect integrated patient evidence before their limits can be understood safely.