Google's AMIE AI matches doctors in real-time video consultations

Google Research presented AMIE (Video), a real-time audio-visual configuration of its AMIE (Articulate Medical Intelligence Explorer) research system, built on Gemini and Project Astra. AMIE previously showed expert-level performance in text-based diagnostic dialogue and worked as a differential diagnosis aid for clinicians; Google has since extended it toward treating and managing disease over time, toward specialist-level evaluations in oncology, cardiology and ophthalmology, and toward multimodal reasoning over images and clinical documents, alongside a framework for physician oversight and early real-world studies, including a clinical feasibility study with Beth Israel Deaconess Medical Center and an ongoing nationwide randomized study with Included Health.
The motivation for AMIE (Video) is that text-based interfaces discard the visual and auditory dimensions of a real consultation. Patients must translate physical symptoms into written descriptions, a process that loses diagnostic information and can disadvantage patients with limited digital or health literacy; a text-only system also cannot observe non-verbal cues or guide a patient through a physical examination.
AMIE (Video) conducts synchronous clinical video consultations: it perceives non-verbal cues, guides patient actors through virtual physical-examination maneuvers, and reasons diagnostically, all in real time. To reconcile natural conversational pace with the time deep clinical reasoning requires, it runs an asynchronous multi-agent architecture that splits the work across three specialized agents operating continuously in parallel, a decoupled design intended to keep conversational latency natural while reasoning and audio-visual perception happen alongside it.
To characterize the system's perceptual and reasoning abilities before human evaluation, Google built a taxonomy of clinical audio-visual competencies drawn from the medical literature, covering non-verbal visual cues, auditory signals and physical-exam maneuvers, and used it to structure an automated evaluation suite combining single-turn audio-visual tests (for example, correctly identifying anatomical laterality or signs of respiratory distress) with multi-turn simulated audio consultations.
The headline result comes from a randomized Objective Structured Clinical Examination (OSCE) study using a synchronous video interface. It covered 100 clinical scenarios across five body systems: cardiopulmonary, abdominal, head/eyes/ears/nose/throat, neurological/psychiatric and musculoskeletal. Fifteen trained patient actors carried out 300 standardized consultations across three study arms, which compared AMIE (Video), the text-based AMIE (Text), and a group of 30 board-certified primary care physicians. A separate, independent panel of 20 experienced primary care physicians then scored every consultation using established clinical competency scales and case-specific rubrics.
On core clinical competencies, history-taking thoroughness, diagnostic accuracy, management appropriateness and communication quality, the evaluator panel rated AMIE (Video) on par with the human primary care physicians, and AMIE (Video) also matched or exceeded AMIE (Text) on these same dimensions. AMIE (Video) was rated significantly higher, on average, than both the physicians and AMIE (Text) at eliciting physical signs and proactively guiding patient actors through virtual examination maneuvers, an advantage that also showed up in the case-specific perception and examination scores. The patient actors themselves strongly preferred the video interface over text-based chat, rating it significantly easier to use and more effective for communicating health concerns, and they rated AMIE (Video) favorably on empathy, rapport and confidence in care compared with both the physicians and AMIE (Text).
Google frames the result with substantial caveats. The study used professional patient actors in simulated settings rather than real patients with their own conditions, and the scenarios were limited to presentations that could be authentically acted, excluding cases where audio-visual perception would matter most. Separate automated evaluations found occasional perceptual and reasoning errors despite overall high conversation and diagnostic quality, and the system still shows intermittent technical issues, tied to the prototype nature of Project Astra, that can disrupt the naturalness of the conversation. Google says validating the findings with real patients and real clinical conditions, and building out safety frameworks, is an essential next step before any conclusion about real-world utility, and points to the ongoing Beth Israel Deaconess and Included Health studies as the path toward that evidence.
Key facts
- AMIE (Video), Google's real-time audio-visual medical AI built on Gemini and Project Astra, was tested in a randomized OSCE study of 100 clinical scenarios across five body systems, with 15 trained patient actors carrying out 300 standardized consultations across three study arms.
- An independent panel of 20 experienced primary care physicians rated AMIE (Video) on par with a separate group of 30 board-certified primary care physicians on history-taking, diagnostic accuracy, management appropriateness and communication quality, and found it matched or exceeded the text-only AMIE (Text) on the same measures.
- AMIE (Video) scored significantly higher on average than both the physicians and AMIE (Text) at eliciting physical signs and guiding patient actors through virtual physical-examination maneuvers.
- Patient actors preferred the video interface over text chat, rating it easier to use and more effective for communicating health concerns, and rated AMIE (Video) higher on empathy, rapport and confidence in care than both the physicians and AMIE (Text).
- The study used professional patient actors in simulated settings rather than real patients, and Google says validating the results with real patients is an essential next step before drawing conclusions about real-world utility.
Why it matters
Text-based medical AI throws away most of what a real consultation runs on: gait, breathing, visible discomfort, the physical exam itself. AMIE (Video) is Google's attempt to close that gap, and the OSCE study is the first published demonstration of an AI system performing at expert level in a real-time, synchronous video consultation rather than a text exchange, a step toward systems that could extend access to clinical expertise through telehealth.
Who it affects
Patients, particularly those with limited digital or health literacy who struggle to translate symptoms into text, stand to gain the most from a system that can watch and listen rather than only read. The result also matters to primary care physicians and telehealth operators, and to Google's existing clinical partners, Beth Israel Deaconess Medical Center and Included Health, whose real-world studies are the next step Google points to.
How to use it
AMIE (Video) is a research system, not a public product; the post gives no access path, pricing or release plan. Technically, it runs on Gemini and Project Astra with an asynchronous multi-agent architecture that splits the work across three specialized agents running in parallel, letting it hold a natural conversational pace while reasoning and audio-visual perception happen alongside it.
How solid is it
The core claim rests on a randomized OSCE study: 100 clinical scenarios across five body systems, 15 trained patient actors, 300 standardized consultations split across three arms (AMIE Video, AMIE Text, and 30 board-certified physicians), scored by an independent panel of 20 primary care physicians using established clinical rubrics. That is a reasonably sized, structured evaluation, but every consultation in it involved a patient actor, not a real patient.
Risks and caveats
Google is explicit that professional actors cannot fully replicate the unpredictability of real encounters, and that scenarios were limited to conditions that could be authentically acted, leaving out presentations where visual or auditory perception would matter most. Automated evaluations separately turned up occasional perceptual and reasoning errors, and the system still has intermittent technical issues tied to Project Astra's prototype status. Google states that testing with real patients and real conditions is still required before any real-world utility conclusion can be drawn.
“This work demonstrates that the transition from text-based to audio-visual clinical AI is achievable at expert-level quality.”
— Google Research blog post, by Anil Palepu and Mike Schaekermann