Medical AI’s real argument is not about whether the models work. It is about the human in the loop. Clinicians, one essay argues, can degrade a capable model’s output rather than improve it. The essay claims an autonomous system will soon outscore a doctor who is using AI.
The authors are Ezekiel Emanuel and Abe Baker-Butler, joined by Vinod Khosla and Neal Khosla. An August JAMA paper by the same group is the origin of the claim. It named five cognitive areas ripe for automation, spanning patient history-taking, differential diagnosis, test selection, prescribing and chronic-illness management.
Their evidence comes from studies published since January 2024. Nine of 13 comparisons that pit autonomous AI against AI-assisted clinicians favored the autonomous system. Diagnoses from AI alone beat AI-assisted clinicians by 21.3 percentage points.
Individual results are striking. Google’s AMIE elicited fuller patient histories than physicians. ChatGPT beat doctors on differential diagnosis by 18 points, 92% to 74%. Microsoft’s AI Diagnostic Orchestrator reached correct final diagnoses 4.02 times more often at 19.1% lower testing cost. A Stanford diabetes study reached a stable insulin dose in 15 days where clinicians did not after eight weeks.
Why would adding a doctor hurt? The authors say that when AI is highly capable, clinician corrections introduce more errors than insight. They also point to a 461-visit Annals of Internal Medicine study where physicians, given the AI’s recommendation, still produced worse treatment plans.
AMA chief executive John Whyte counters that licensure and liability frameworks are inadequate, that empathy and trust remain human work, and that most studies are simulations. Emanuel and Baker-Butler accept the regulatory gap but say the fix is to build oversight, not to block testing.
