AI chatbots used to interpret X-rays can be dangerously confident even when their conclusions are wrong. That combination of fluent explanations and unreliable accuracy is a particular risk in medical settings, where users may overtrust a plausible answer.
The report adds to mounting evidence that general-purpose language and vision systems need careful limits in healthcare. Accuracy, uncertainty calibration, audit trails, and escalation to clinicians are all part of safe deployment.
For hospitals and health-tech vendors, the lesson is clear: impressive model interfaces cannot substitute for validated clinical performance.