Video consultations with AMIE: Google's AI system performed at the level of doctors

19 August 20262 views

Google's experimental medical AI system conducted synchronous video consultations with actor-patients and received evaluations from clinical experts comparable to those of live physicians. The researchers emphasize that final conclusions about practical application are only possible after trials with real patients.

Video consultations with AMIE: Google's AI system performed at the level of doctors

In recent years, artificial intelligence has been making increasingly confident inroads into medicine, but most systems work with text data or individual tasks. New research from Google takes the conversation to a different level: the company tested its experimental AMIE system on live video consultations, and the results were comparable to the work of practicing physicians. This doesn't mean AI is ready to replace doctors, but it can already conduct examinations and collect medical histories at the level of a primary care specialist.

What the study showed

The trial involved 15 professional actors who played patients with complaints across five areas: cardiopulmonary issues, abdominal pain, ENT and eye symptoms, neurological and psychiatric conditions, as well as musculoskeletal pathologies. Each consultation took place in a synchronous video format: AMIE conversed with the actor in real time, asked questions, and attempted to perform a virtual examination.

An independent panel of 20 experienced primary care physicians evaluated these dialogues using standardized clinical rubrics. The surprise was that the AI's scores were on par with those of live doctors working under the same scheme. This applied not only to gathering complaints but also to diagnostic accuracy, appropriateness of recommendations, and quality of communication. The video format also gained an advantage over text chat: physicians rated the system working with a video stream higher for identifying physical signs and guiding actors through virtual examinations. The actors themselves also preferred video—they found it more convenient and effective for describing problems, and noted a high level of empathy and confidence in treatment.

How AMIE works

The secret to its success lies in an unusual architecture. AMIE uses an asynchronous multi-agent scheme where three independent agents are responsible for the process. The first, talker, directly conducts the conversation with the patient: its task is to maintain dialogue, ask clarifying questions, and respond to statements. The second agent, planner, works quietly in the background: it continuously updates the differential diagnosis and treatment plan, tracks what data is missing, and adjusts the questioning strategy. The third—perception—handles sensing: it analyzes video and audio streams, picks up nonverbal cues, records visible physical signs, and even sounds such as coughing or breathing.

This division of labor allows the talker agent to respond almost instantly without waiting for slow background computations to finish. As a result, the conversation feels natural, without pauses, while all the intelligent work happens behind the scenes. This is a key difference from conventional chatbots, which process a request in full and only then return a response.

Methodology and results

The study was designed as a randomized multi-arm trial. Synchronous video AMIE was compared with a text version of the same system and with a group of ten certified primary care physicians using the same video interface. All consultations were then blindly assessed by an expert panel against two sets of criteria: overall clinical competence and scenario-specific skills.

Additionally, Google developed an automated evaluation suite based on a taxonomy of telemedicine competencies. It was used to test specific perception tasks—for example, the ability to distinguish anatomical sides or notice signs of respiratory distress. In some simulations, visual information was provided as text descriptions to test dialogue skills without an end-to-end video stream. This made it possible to quickly identify weaknesses in the system and refine its design before the trial with actors.

The final assessment showed that each of the three agents contributes to improving clinical metrics: history taking, clinical reasoning, treatment recommendations, communication, and response latency. Video AMIE either matched or surpassed physicians across all key metrics and outperformed the text version in most aspects. The only area where video lagged was, perhaps, speed—but even that indicator proved acceptable.

Limitations and prospects

Despite the impressive results, the study authors themselves warn of serious limitations. The main one is that actors cannot fully convey the variability of real patients: their symptoms, behavior, connection quality, environment, and medical history always extend beyond a prepared script. Moreover, cases where audiovisual perception could provide particularly important diagnostic information were deliberately excluded from the scenarios—simply because actors could not convincingly portray them.

Neither automated assessments nor simulated consultations prove that the system will perform equally well in a real clinic. Google currently has almost no production data—all real-world constraints remain text-based. Therefore, before any talk of deploying AMIE in practice, full-scale trials with real patients are needed, where people's symptoms and behavior are far less predictable.

Nevertheless, the very fact that AI can conduct a differential diagnosis, gather a medical history, and build a virtual examination at the level of a general practitioner looks promising. In the future, such systems could become a useful assistant for doctors, especially in telemedicine—but for now, this is only a research prototype, not a ready-made clinical tool.

Frequently asked questions

Video consultations with AMIE: Google's AI system performed at the level of doctors