English

Towards Conversational Medical AI with Eyes, Ears and a Voice

Artificial Intelligence 2026-05-12 v1 Computation and Language Computer Vision and Pattern Recognition

Abstract

The practice of medicine relies not only upon skillful dialogue but also on the nuanced exchange and interpretation of rich auditory and visual cues between doctors and patients. Building on the low-latency voice and video processing capabilities of Gemini, we introduce AI co-clinician, a first-of-its-kind conversational AI system utilizing continuous streams of audio-visual data from live patient conversations to inform real-time clinical decisions. Its dual-agent architecture balances deep clinical reasoning with the low latency required for natural dialogue. To assess this system, we implemented a video-based interface emulating telemedicine consultations. We crafted 20 standardized outpatient scenarios requiring proactive real-time auditory and visual reasoning and designed "TelePACES" evaluation criteria alongside case-specific rubrics. In a randomized, interface-blinded, crossover simulation study (n = 120 encounters) with 10 internal medicine residents as patient actors, we compared AI co-clinician with primary care physicians (PCPs), GPT-Realtime, and a baseline agent. AI co-clinician approached PCPs in key TelePACES dimensions, including management plans and differential diagnosis, while significantly outperforming GPT-Realtime across all general criteria. While our agent demonstrated parity with PCPs in case-specific triage measures, physicians maintained superior overall performance in case-specific assessments. Although AI co-clinician marks a significant advance in real-time telemedical AI, gaps remain in physical examination and disease-specific reasoning. Our work shows that text-only approaches fail to capture the true challenges of medical consultation and suggests that high-stakes real-time diagnostic AI is most safely advanced in collaborative, triadic models where AI can be a supportive co-clinician for doctors and patients.

Keywords

Cite

@article{arxiv.2605.09272,
  title  = {Towards Conversational Medical AI with Eyes, Ears and a Voice},
  author = {Meet Shah and Jason Gusdorf and Anil Palepu and Chunjong Park and Jack W. O'Sullivan and Vishnu Ravi and Tim Strother and Pavel Dubov and Aliya Rysbek and Toshiyuki Fukuzawa and Yana Lunts and Jan Freyberg and Michael B. Chang and Aniruddh Raghu and David Stutz and Devora Berlowitz and Eliseo Papa and Taylan Cemgil and JD Velasquez and Jack Chen and Arthur Chen and Doug Fritz and Charlie Taylor and Katya Tregubova and Jing Rong Lim and Richard Green and Sara Mahdavi and Mahvish Nagda and Jihyeon Lee and Craig Schiff and Liviu Panait and Sukhdeep Singh and Valentin Liévin and David G. T. Barrett and Hannah Gladman and Anna Cupani and Francesca Pietra and Uchechi Okereke and Katherine Tong and Clemens Meyer and Erwan Rolland and Mili Sanwalka and Michael D. Howell and Shixiang Shane Gu and Bibo Xu and Euan A. Ashley and S. M. Ali Eslami and Gregory Wayne and Pushmeet Kohli and Vivek Natarajan and Adam Rodman and Alan Karthikesalingam and Ryutaro Tanno},
  journal= {arXiv preprint arXiv:2605.09272},
  year   = {2026}
}

Comments

Video examples are available on Youtube: https://youtu.be/y5Vaa_SN1t0, https://youtu.be/dC4icb75vLQ, and https://youtu.be/E7iEvWo-E6c