Google’s latest medical AI scored higher than primary-care doctors on several measures in a simulated video exam. It did not treat a real patient.
That distinction does not make the result trivial. It tells us what the evidence is actually about.
What changed. Google’s new AMIE preprint describes a system that can talk with a patient over live video, watch and listen for clinical signs, and guide parts of a physical examination. Three agents work at once: a Talker keeps the conversation moving, a Planner tracks symptoms and possible diagnoses, and a Perception agent reviews the audio and video for cues the fast conversational agent might miss.
This is a genuine step beyond Google’s earlier 20-scenario proof of concept, where the medical agent approached doctors on some measures but physicians remained stronger overall.
The new study used 100 scripted cases and three versions of each consultation: AMIE over video, AMIE through text chat, and a doctor over video. Fifteen professional actors played the patients. Ten board-certified primary-care physicians conducted the human consultations, while a separate panel of 20 physicians evaluated the recordings and transcripts.
On the case-specific score, AMIE Video received 83% versus 68% for the doctors. Its first-choice diagnosis matched the scripted answer in 91% of cases versus 77% for doctors. When the comparison expanded to the top three possible diagnoses, the gap narrowed to 98% versus 90% and was not statistically significant.
Video helped most where text is weakest. AMIE Video scored 74% versus 47% for doctors on physical observation and guided examination. Patient actors also rated video as easier and more effective than text chat. But diagnosis and management were comparable between AMIE’s video and text versions. The strongest video-specific result is therefore not “a camera made the AI smarter.” It is that the system could turn live perception into useful questions and examination instructions.
The system design matters beyond medicine. A single model must either answer quickly or stop to think carefully. AMIE splits those jobs: one agent speaks, another plans, and another watches. That is a practical pattern for any real-time assistant that cannot make the user wait while it reasons.
What the headline does not establish. Every patient was an actor following a case pack with a known answer. Conditions that could not be portrayed reliably were left out. Real clinics add unclear histories, multiple illnesses, poor connections, interrupted conversations and signs that do not arrive on cue.
The comparison also changed the normal clinical setting. Doctors could see patients, but the researchers asked doctors to turn off their own cameras to match AMIE’s lack of an avatar. AMIE also used extra inference after the visit to produce its diagnostic questionnaire. The paper is an Alphabet-funded preprint; the recordings, model code and weights are not public.
Targeted tests show why average scores are not enough. AMIE performed poorly on subtle or fast-moving signs, including nystagmus, tremors and emotional affect. The authors also report missed visual cues, occasional failures to correct a wrongly performed exam and cases where red-flag advice was absent.
Google calls AMIE a research system and says real-patient validation must come before clinical use. That is the next test that matters: not whether AMIE can pass a scripted exam, but whether it can fail safely amid the disorder of actual care.
Source graph: Semble source collection