推动AMIE迈向专家级视听临床会诊

qimuai 发布于 阅读:3 一手编译

推动AMIE迈向专家级视听临床会诊

内容来源:https://research.google/blog/advancing-amie-towards-expert-level-audio-visual-clinical-consultations/

内容总结:

谷歌旗下研究团队8月11日发布一项最新研究成果,其研发的医疗人工智能系统AMIE在实时视频问诊中展现出与人类医生相当的专业水平。这是AI系统首次在模拟问诊的随机对照研究中达到专家级表现,标志着医疗AI从文本交互向视听交互的重要跨越。

据介绍,AMIE系统基于谷歌Gemini和Project Astra技术构建,采用异步多智能体架构,由三个专业化智能体并行协作,分别负责对话响应、临床推理和视听感知,从而在保持自然对话节奏的同时完成复杂的诊断推理。该系统能够实时观察患者步态、呼吸等非语言线索,并引导患者完成虚拟体格检查。

研究团队设计了一项大规模随机对照研究,涵盖100个临床场景、300次实时问诊,并由30名持证初级保健医生参与对比评估。独立评审专家组按照既定临床评分标准对所有问诊进行评估,结果显示AMIE(视频版)在病史采集、诊断准确性、处置合理性和沟通质量等核心临床能力上与人类医生持平,并在体征观察和体格检查引导方面显著优于人类医生和文本版系统。

参与模拟问诊的患者演员对视频问诊体验给予高度评价,认为该界面比文本聊天更易用、沟通健康问题更有效,并认为AMIE(视频版)在共情、融洽度和就医信心方面表现良好。

研究人员强调,该研究仍存在局限:所有问诊均在模拟环境中进行,使用的是专业患者演员而非真实患者,且研究范围限于可通过表演真实呈现的病症。此外,系统仍存在偶发的感知和推理错误及技术问题。团队表示,下一步将在真实患者和真实临床条件下验证这些发现,目前已在贝斯以色列女执事医疗中心开展真实世界可行性研究,并与Included Health合作推进全国性随机研究。

研究团队表示,这项工作证明了从文本到视听交互的临床AI转型能够在专家级质量水平上实现,是朝着未来AI系统增强临床诊疗能力迈出的重要里程碑。

中文翻译:

2026年8月11日
Anil Palepu,高级研究科学家,Mike Schaekermann,研究负责人,Google

我们推进了研究型医疗AI系统AMIE,使其能够进行实时视频问诊,在一项采用模拟问诊的随机对照研究中,首次展示了专家级表现的验证。

当医生接诊患者时,问诊远不止于言语交流。医生观察患者的步态,留意可见的不适迹象,注意其呼吸状况,并引导患者完成体格检查动作。这一连续的视觉和听觉信息流与口述的病史无缝融合。这些非语言的视觉和听觉线索对于有效诊断、患者信任和临床沟通至关重要。

具备临床推理和对话能力的AI系统有潜力大幅提升医疗专业知识和护理的可及性,从而开创一个医生可以将时间集中在医患互动中最有意义方面的未来。在早期工作中,Articulate Medical Intelligence Explorer(AMIE)——我们用于临床推理和对话的研究型AI系统——在基于文本的诊断对话中展现了专家级表现,并被证明是临床医生的有效鉴别诊断辅助工具。最近,我们将AMIE的能力从诊断进一步拓展到长期疾病治疗和管理。

我们还拓展了AMIE的能力,使其能够在模拟环境中对患者演员进行肿瘤学、心脏病学和眼科学领域的专科级评估,以及基于图像和临床文档的多模态诊断推理。与此同时,我们已开始将这些研究进展向临床实践转化,通过一个以医生为中心的监督框架,以及我们的首批真实世界临床研究,包括与Beth Israel Deaconess医学中心合作开展的临床可行性研究,以及与Included Health合作开展的正在进行的全国性随机研究。

尽管取得了这些进展,我们研究中仍存在一个根本性限制,即基于文本的界面丢弃了临床实践中的视觉和听觉维度。患者必须将复杂的身体症状转化为书面描述,这一过程会丢失诊断信息,并可能对数字素养或健康素养有限的患者产生不利影响。仅基于文本的系统无法独立观察那些为临床推理提供依据的视觉和听觉线索,也无法引导患者完成有助于鉴别诊断的体格检查动作。

今天,在“迈向专家级实时视频问诊医疗AI”一文中,我们以实时视频配置呈现了AMIE,即AMIE(Video),以解决这些限制。AMIE(Video)基于Gemini和Project Astra构建,能够进行同步临床视频问诊,感知非语言临床线索,引导患者演员完成虚拟体格检查,并进行诊断推理,且全部实时完成。在一项包含100个临床场景、300次实时问诊和30名经过认证的家庭医生(PCP)的多组随机研究中,我们首次展示了一个AI系统在实时临床视频问诊中表现出专家级水平。

通过视频进行有效的临床对话需要平衡相互竞争的需求:系统必须以自然的对话速度回应患者,同时进行细致的临床推理并持续处理视觉和听觉信息流。目前,单一智能体无法满足所有这些要求。深度推理需要时间,但对话中的停顿会削弱患者的信任和融洽感。

为应对这一挑战,AMIE(Video)采用了一种异步多智能体架构,将任务分配给三个专业化智能体,它们持续并行工作:

这种解耦设计使AMIE(Video)能够在进行诊断推理和视听感知的同时保持自然的对话延迟,否则这些推理和感知将带来不可接受的延迟。自动评估证实,这种三智能体架构中的每个智能体都对临床指标的改善做出了重要贡献,例如病史采集能力、临床推理和治疗建议,以及对对话质量相关指标的改善,包括以患者为中心的沟通技能和响应延迟。

构建视听医疗AI的一个关键挑战是如何大规模地表征系统的感知和推理能力。为了指导开发,我们从医学文献中提炼出一个与远程医疗相关的临床视听能力分类体系,涵盖非语言视觉线索、听觉信号和体格检查动作。然后,我们围绕这一分类体系构建了一套自动化评估套件。

该评估框架将针对性的单轮视听评估与多轮模拟音频问诊相结合。单轮视听评估测试临床感知和推理的具体实例(例如,正确识别解剖学左右方位或识别呼吸窘迫的体征)。多轮模拟音频问诊则评估端到端的对话表现,同时将视觉线索作为文本描述注入模拟中(例如,一个针对帕金森场景的AI患者模拟器在被提示展示其笔迹时,可能会注入一段口头描述为“[将纸举到镜头前,展示 cramped、微小的字迹]”)。这些互补的评估共同实现了系统设计的快速迭代,并在人体评估之前充分表征了AMIE(Video)的能力和失败模式。

为了在更具挑战性和更真实的端到端视听临床问诊环境中评估临床能力,我们开展了一项大规模的随机客观结构化临床考试(OSCE)研究,使用同步视频问诊界面。

为覆盖评估中广泛的医学状况,该研究涵盖了100个临床场景,涉及五个身体系统,包括心肺、腹部、头/眼/耳/鼻/喉(HEENT)、神经/精神和肌肉骨骼系统。十五名训练有素的患者演员在三组中完成了300次标准化问诊:

一个由20名经验丰富的家庭医生组成的独立评审小组使用既定的临床评分标准对所有问诊进行了评估,包括通用临床能力量表和针对每个场景定制的详细病例特定评分标准。

专家级临床表现:在核心临床能力方面,包括病史采集全面性、诊断准确性、管理适当性和沟通质量,临床评审员对AMIE(Video)的评分与PCP相当。AMIE(Video)在这些维度上也达到或超过了AMIE(Text)。

体格观察和检查方面的优势:AMIE(Video)在引导患者演员主动完成虚拟检查动作和获取体格体征方面的平均评分显著高于PCP和AMIE(Text)。这一优势也反映在病例特定感知和检查评分标准中。

患者演员更偏好视频体验:患者演员强烈偏好同步视频界面而非文本聊天,认为其在传达健康问题方面显著更易于使用且更有效。与PCP和AMIE(Text)相比,他们还在同理心、融洽感和对护理的信心方面对AMIE(Video)给予了积极评价。

这项研究存在重要局限性,在解释这些结果时必须将其置于这些局限性的背景中。该研究完全由专业患者演员在模拟临床环境中进行,而非真实的患有自身健康问题的患者。患者演员无论多么熟练,都无法完全复制真实临床诊疗的复杂性和不可预测性,而且场景仅限于可以通过表演真实呈现的病症,省略了视听感知会产生诊断性影响的重要临床表现。在研究范围之外,有针对性的自动化评估揭示了偶发的感知和推理错误,尽管整体对话质量和诊断准确性很高,但该系统仍存在间歇性的技术问题,可能破坏对话的自然性。鉴于Project Astra的原型性质,这包括未来开发可能在系统层面解决的技术考量,这些考量超出了本工作中探索的特定医疗应用。在得出关于现实世界实用性的任何结论之前,在真实患者和真实临床条件下的研究中评估这些发现是必不可少的下一个步骤。

这项工作表明,从基于文本的临床AI向视听临床AI的转变可以以专家级质量实现。AMIE(Video)参与到临床实践的感知丰富性中,观察非语言线索,引导体格检查,并通过语音对话自然交流——这些能力更接近于远程医疗视频问诊的体验。

在通往负责任的现实世界证据的道路上,仍存在重要问题。我们的发现需要在真实患者中得到验证,扩展到无法演绎的临床表现,并得到稳健的安全框架的支持。我们已经在朝这个方向迈出了初步步伐:与Beth Israel Deaconess医学中心合作开展的一项现实世界可行性研究为基于文本的AMIE在临床实践中的安全性和实用性提供了初步证据,我们与Included Health合作开展的正在进行的全国性随机研究正在进一步评估AI在现实世界虚拟护理中的应用。这些研究经验共同将有助于为如何负责任地将视听能力整合到临床实践中提供信息。虽然还有大量工作要做,但这些结果标志着朝着有朝一日能够通过与临床实践的感官复杂性互动来增强护理的AI系统迈出了重要里程碑。

本文所述研究是Google Research和Google DeepMind多个团队的联合工作。我们感谢所有合著者——Mahvish Nagda、Jihyeon Lee、Matthew Thompson、CJ Park、Tim Strother、Valentin Liévin、Roma Ruparel、Akshay Goel、Teya Bergamaschi、Suhana Bedi、Meet Shah、Pavel Dubov、Liviu Panait、Toshiyuki Fukuzawa、Sam Schmidgall、Craig Schiff、Joseph Xu、Aliya Rysbek、Yana Lunts、Jan Freyberg、Rebecca Hemenway、Sunny Virmani、David Racz、Carey Radebaugh、Joëlle Barral、Kavi Goel、Dale R. Webster、Katherine Chou、Avinatan Hassidim、Yossi Matias、James Manyika、Gregory Wayne、Tao Tu、Yun Liu、Ethan Goh、Christina Chen、Ryutaro Tanno和Cameron Chen。

英文来源:

August 11, 2026
Anil Palepu, Senior Research Scientist, and Mike Schaekermann, Research Lead, Google
We advance AMIE, our research medical AI system, to conduct real-time video consultations, with a first-of-its-kind demonstration of expert-level performance in a randomized controlled study with simulated consultations.
When a physician meets a patient, the consultation extends far beyond the words exchanged. The physician observes the patient's gait, registers visible signs of discomfort, notes their breathing, and guides the patient through physical examination maneuvers. This continuous stream of visual and auditory information is seamlessly integrated with the spoken clinical history. These non-verbal visual and auditory cues are central to effective diagnosis, patient trust, and clinical communication.
AI systems capable of clinical reasoning and dialogue have the potential to dramatically increase access to medical expertise and care, fostering a future where physicians can focus their time on the most meaningful aspects of patient interactions. In early work, the Articulate Medical Intelligence Explorer (AMIE), our research AI system for clinical reasoning and dialogue, demonstrated expert-level performance in text-based diagnostic dialogue and proved effective as a differential diagnosis aid for clinicians. Recently, we advanced AMIE’s capabilities beyond diagnosis towards treating and managing disease over time.
We have also extended AMIE's capabilities towards specialist-level evaluations in oncology, cardiology and ophthalmology, and multimodal diagnostic reasoning over images and clinical documents, in simulated settings with patient actors. In parallel, we have begun translating these research advances towards clinical practice, through a framework for physician-centered oversight, as well as our first real-world clinical studies including a clinical feasibility study with Beth Israel Deaconess Medical Center, and an ongoing nationwide randomized study in partnership with Included Health.
Despite these advances, a fundamental constraint in our research remained that text-based interfaces discard the visual and auditory dimensions of clinical practice. Patients must translate complex physical symptoms into written descriptions, a process that discards diagnostic information and can negatively affect patients with limited digital or health literacy. Text-only systems cannot independently observe the visual and auditory cues that inform clinical reasoning, nor can they guide patients through the physical examination maneuvers that shape differential diagnosis.
Today, in “Towards expert-level medical AI for real-time video consultations”, we present AMIE in a real-time video configuration, AMIE (Video), that addresses these limitations. Built on Gemini and Project Astra, AMIE (Video) conducts synchronous clinical video consultations, perceiving non-verbal clinical cues, guiding patient actors through virtual physical examinations, and reasoning diagnostically, all in real time. In a multi-arm randomized study with 100 scenarios, 300 live consultations, and a group of 30 board-certified primary care physicians (PCPs), we present the first demonstration of an AI system exhibiting expert-level performance in real-time clinical video consultations.
Conducting an effective clinical conversation over video requires balancing competing demands: the system must respond to patients at natural conversational speed while simultaneously performing careful clinical reasoning and continuously processing visual and auditory streams. Currently, a single agent cannot satisfy all these requirements. Deep reasoning takes time, but conversational pauses erode patient trust and rapport.
To address this challenge, AMIE (Video) uses an asynchronous multi-agent architecture that divides labor across three specialized agents working continuously in parallel:
This decoupled design allows AMIE (Video) to maintain natural conversational latency while performing diagnostic reasoning and audio-visual perception that would otherwise introduce unacceptable delays. Automated evaluations confirm that each agent in this three-agent architecture makes important contributions towards improvements on clinical metrics, such as competency in history-taking, clinical reasoning and treatment recommendations, as well as on metrics related to dialogue quality, including patient-centered communication skills and response latency.
A key challenge in building audio-visual medical AI is characterizing a system's perceptual and reasoning capabilities at scale. To guide development, we derived a taxonomy of clinical audio-visual competencies relevant to telehealth from the medical literature, covering non-verbal visual cues, auditory signals, and physical examination maneuvers. We then built an automated evaluation suite structured around this taxonomy.
This evaluation framework combines targeted single-turn audio-visual assessments with multi-turn simulated audio consultations. The single-turn audio-visual assessments test specific instances of clinical perception and reasoning (e.g., correctly identifying anatomical laterality or recognizing signs of respiratory distress). And the multi-turn simulated audio consultations assess end-to-end conversational performance while injecting visual cues as textual descriptions into the simulation (for example, an AI patient simulator for a Parkinson’s scenario prompted to show their handwriting may inject a verbal description of “[holding up paper to camera showing cramped, tiny script]”). Together, these complementary evaluations enabled rapid iteration on system design and richly characterized capabilities and failure modes of AMIE (Video) prior to human evaluation.
To evaluate clinical competence in the more challenging and realistic setting of an end-to-end audio-visual clinical consultation, we conducted a large-scale, randomized Objective Structured Clinical Examination (OSCE) study with a synchronous video consultation interface.
To cover a breadth of medical conditions in our evaluation, the study spanned 100 clinical scenarios covering five body systems, including cardiopulmonary, abdominal, head/eyes/ears/nose/throat (HEENT), neurological/psychiatric, and musculoskeletal conditions. Fifteen trained patient actors carried out 300 standardized consultations across three study arms:
An independent panel of 20 experienced primary care physicians evaluated all consultations using established clinical rubrics, including both general clinical competency scales and detailed case-specific scoring criteria tailored to each scenario.
Expert-level clinical performance: Across core clinical competencies, history-taking thoroughness, diagnostic accuracy, management appropriateness, and communication quality, clinical evaluators rated AMIE (Video) on par with PCPs. AMIE (Video) also matched or exceeded AMIE (Text) on these dimensions.
Strength in physical observation and examination: AMIE (Video) was rated significantly higher, on average, than both PCPs and AMIE (Text) eliciting physical signs and proactively guiding patient actors through virtual examination maneuvers. This advantage was also reflected in case-specific perception and examination rubric scores.
Patient actors preferred the video experience: Patient actors strongly preferred the synchronous video interface over text-based chat, rating it as significantly easier to use and more effective for communicating health concerns. They also rated AMIE (Video) favorably on empathy, rapport, and confidence in care compared to both PCPs and AMIE (Text).
This research has important limitations and it is critical to interpret these results within the context of these limitations. This study was conducted entirely with professional patient actors in simulated clinical settings, not with real patients presenting with their own health conditions. Patient actors, however skilled, cannot fully replicate the complexity and unpredictability of real clinical encounters, and the scenarios were limited to conditions that can be authentically portrayed through acting, omitting important clinical presentations where audio-visual perception would be diagnostically consequential. Beyond the scope of the study, targeted automated evaluations revealed occasional perceptual and reasoning errors, despite overall high-quality conversation and diagnostic accuracy, and the system still exhibits intermittent technical issues that can disrupt conversational naturalness. Given the prototype nature of Project Astra, this includes technical considerations that future development may address at a system level that go beyond the specific medical application explored in this work. Assessing these findings in studies with real patients and real clinical conditions is an essential next step before any conclusions about real-world utility can be drawn.
This work demonstrates that the transition from text-based to audio-visual clinical AI is achievable at expert-level quality. AMIE (Video) engages with the perceptual richness of clinical practice, observing non-verbal cues, guiding physical examination, and conversing naturally through spoken dialogue — capabilities that more closely approximate the experience of a telehealth video encounter.
Important questions remain on the path towards responsible real-world evidence. Our findings need to be validated with real patients, expanded to encompass clinical presentations that cannot be enacted, and supported by robust safety frameworks. We have already taken early steps in this direction: a real-world feasibility study with Beth Israel Deaconess Medical Center provided initial evidence for the safety and utility of text-based AMIE in clinical practice, and our ongoing nationwide randomized study with Included Health is further evaluating AI in real-world virtual care. Together, these research experiences will help inform how audio-visual capabilities might be responsibly integrated into clinical practice. While much remains to be done, these results mark an important milestone towards AI systems that could one day augment care by engaging with the sensory complexity of clinical practice.
The research described here is joint work across many teams at Google Research and Google Deepmind. We are grateful to all our co-authors - Mahvish Nagda, Jihyeon Lee, Matthew Thompson, CJ Park, Tim Strother, Valentin Liévin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg, Rebecca Hemenway, Sunny Virmani, David Racz, Carey Radebaugh, Joëlle Barral, Kavi Goel, Dale R. Webster, Katherine Chou, Avinatan Hassidim, Yossi Matias, James Manyika, Gregory Wayne, Tao Tu, Yun Liu, Ethan Goh, Christina Chen, Ryutaro Tanno, and Cameron Chen.

谷歌研究进展

文章目录


    扫描二维码,在手机上阅读