Large multimodal models diagnose intracranial and spinal conditions far more reliably when they can pick from answer options than when asked to reason from scratch, a study from Turkish neurosurgery departments finds.
Open-ended generative diagnosis ranged from 51.6% to 62.7% accuracy across models. When the same models received answer choices, accuracy jumped to 75.4% to 81.0%, a gap of 15 to 24 percentage points. Between 43% and 59% of completely wrong open-ended answers were self-corrected once options appeared.
A neurosurgical panel graded 17.8% to 26.9% of all cases as clinically significant errors, enough to matter for treatment decisions. The pattern persisted in a 100-case subcohort with repeated runs and in fresh-session evaluations, so the gap was not an artifact of the testing design.
The authors conclude the models are not ready for autonomous diagnostic use. Whether their stronger multiple-choice performance can support an assistive role, they write, requires prospective human-AI evaluation, and future assessments should move beyond multiple-choice accuracy.
