Google’s SymptomAI outperformed independent clinicians in a national-scale study of conversational AI for differential diagnosis, according to research published July 22 by Google Research. The system, deployed through the Fitbit app with 13,917 participants, demonstrated diagnostic accuracy more than 2.5 times higher than human doctors given the same patient conversations.
The study tested multiple conversational AI agents built on Gemini models. SymptomAI’s differential diagnoses were significantly more accurate than those from independent clinicians in a blinded randomized comparison. The odds ratio of 2.56 with a p-value under 0.001 indicates the result is highly statistically significant. The system used an agentic interviewing approach that actively gathered symptom information before suggesting diagnoses, rather than simply responding to whatever the user chose to share.
Researchers found that the structured symptom interview strategy substantially outperformed baseline user-guided conversations. An auxiliary analysis on 1,509 participants from a general US population panel confirmed the results generalized beyond wearable device users. The study also linked SymptomAI conversations that led to infectious disease diagnoses with physiological trends from participants’ Fitbit data, including elevated heart rate and temperature changes consistent with immune response.
The research addresses a critical gap in AI diagnostics: most existing evaluations rely on curated medical vignettes rather than real-world patient communication, which includes incomplete information, varying health literacy, and conversational complexity. Google said the findings demonstrate the benefits of dedicated symptom interviewing compared to the user-guided approach that most consumer AI chatbots default to today.
