Most people have Googled their symptoms at some point. But typing symptoms into a search bar is a far cry from the structured, back-and-forth interview a doctor conducts. Google Research just published a study asking a pointed question: what if an AI could run that interview itself, at national scale, on real patients with real symptoms -- and what if it was actually better at it than human clinicians?

The answer, according to their new paper "SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment", is yes -- at least by the metrics they measured. The results are striking enough to warrant a close look at both what the system does and where its limits lie.

The gap that motivated this

Language models have been benchmarked on medical diagnosis before, and they've done well. But those benchmarks almost always use curated patient vignettes -- carefully written case studies with complete, structured information. Existing studies focus on complex scenarios with rich context, making it difficult to draw conclusions about how these systems perform for patients reporting symptoms in everyday life. Real patients ramble. They forget details. They use lay terms. They report symptoms at odd times during an illness. None of that shows up in a vignette.

There's also a systemic access problem. A large proportion of clinical diagnoses can be derived from language-based interviews alone, but these interactions can suffer from financial, geographic, and systemic barriers that limit their accessibility. Close to 20% of health-related AI chat conversations already involve symptom assessment or condition discussion. The demand is there -- the question is whether AI can meet it safely and accurately.

Five agents, one massive study

Google deployed SymptomAI, a set of conversational AI agents for end-to-end patient interviewing and differential diagnosis (DDx), via the Fitbit app in a study that randomized participants (N=13,917) to interact with five AI agents. All five agents were built on Gemini Flash 2.0, but they differed in how they ran the symptom interview. The five strategies were:

Alpha Signal

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves