SymptomAI:面向日常症状评估的对话型人工智能代理

内容总结:
谷歌发布全国性AI诊断研究:对话式AI在症状评估中展现临床级准确率
2026年7月22日,谷歌研究团队公布了一项里程碑式研究成果——通过全国性大规模随机对照试验,首次系统评估了对话式人工智能在真实场景中进行症状鉴别诊断的能力。这项名为“SymptomAI”的研究以13,917名知情同意的参与者为样本,对比了五种不同交互策略的AI代理(基于Gemini Flash 2.0模型)与临床医生的诊断表现。
核心发现:AI诊断准确率优于临床医生基准
研究显示,在由三名认证临床医生组成的盲评小组中,SymptomAI生成的鉴别诊断(DDx)在超过50%的案例中被认为优于其他临床医生提出的诊断方案。同时,在“前5位准确率”指标(即AI判断是否包含最终确诊结果)上,SymptomAI的表现也持续超过临床医生队列。值得注意的是,AI提升最为显著的案例恰恰是临床医生自身诊断信心最低的情境。
交互设计决定诊断质量
研究设置了五种实验条件:两种允许AI主动追问的“动态”模式、两种基于标准化问诊清单的“固定”模式,以及不做主动引导的“基线”模式。结果显示,所有具备主动问诊能力的代理方案均显著优于基线模式,证实了通过交互式信息采集提升AI诊断准确性的核心价值。
跨维度验证:可穿戴设备数据佐证AI诊断
研究进一步引入Fitbit设备采集的生理信号数据作为验证依据。分析表明,当SymptomAI将参与者症状归类为感染性疾病时,其诊断结论与用户报告症状日期前后出现的生理指标波动(如心率、呼吸、皮肤温度、睡眠质量变化)高度吻合。这种“生理证据呼应主观症状”的模式,为AI诊断提供了客观验证路径。
研究局限性声明
研究团队强调,所有AI生成的诊断标签仅用于研究分析,不构成临床确诊或医学评估。鉴于鉴别诊断本身存在时间动态性,且AI无法获取患者体征语言、医疗历史等临床背景信息,当前成果仍需在特定病程阶段(如代谢综合征早期、呼吸道感染初期)进行针对性验证。
未来应用前景
研究认为,SymptomAI类系统有潜力突破传统医疗在时间、地理和经济上的可及性壁垒,通过即时症状评估辅助临床决策,并为大规模人群健康分析提供自动化标签——例如基于可穿戴生理数据的疾病表型研究。相关工作由谷歌研究与谷歌DeepMind团队共同完成。
中文翻译:
2026年7月22日
约瑟夫·布雷达(Joseph Breda),学生研究员;杰克·桑夏恩(Jake Sunshine),研究科学家,谷歌研究院
我们通过一项全国规模的研究,首次展示了人工智能在鉴别诊断和症状检查领域的开创性成果。
相当一部分临床诊断仅凭基于语言的问诊即可得出。这类诊断性问诊通常由临床医生在面诊或远程就诊中通过医患互动完成。尽管这种互动是症状评估的黄金标准,但往往受到经济、地域和系统性障碍的限制,影响其可及性。当前的语言模型在基于精选医学病例研究的评估中已展现出强大的鉴别诊断能力,凸显了其辅助诊断流程的潜力。然而,现有评估大多依赖经过精心筛选、高度详细甚至有时是人工合成的患者案例,这些案例可能无法反映真实世界的体验和临床表现的差异性。这些评估未能捕捉到普通患者日常报告自身症状的方式,例如患者医学素养参差不齐、信息不完整,以及自然对话中出现的其他复杂性。这构成了一个关键缺口,导致我们对语言模型在真实场景中的表现缺乏把握。
为填补这一缺口,我们开展了一项原位比较研究,针对一组实验性的对话式原型人工智能代理,旨在探索对话式人工智能如何为研究基准测试目的,执行端到端的症状问诊与鉴别诊断评估。在我们最新发表的研究论文《SymptomAI:面向日常症状评估的对话式人工智能代理》中,我们分享了一项随机化全国规模研究(样本量 n=13,917)的结果。在该研究中,经过知情同意的研究参与者与五个可能的 Gemini Flash 2.0 SymptomAI 代理之一进行互动。研究期间生成的所有诊断、标签和疾病关联仅用于研究分析,不构成经确认的临床诊断或官方医学评估。在与人工智能代理互动两周后,我们请研究参与者报告他们从医疗服务提供者处获得的任何诊断结果。利用这些数据,我们开展了一项临床专家标注研究,将 SymptomAI 的诊断表现与真实临床医生的医学评估进行了比较。
在评估了 SymptomAI 鉴别诊断的准确性后,我们进一步将 SymptomAI 的诊断结果与参与者在与 SymptomAI 对话前一段时间内 Fitbit 可穿戴设备的生物信号进行了比较。我们发现,最终诊断为传染性疾病的 SymptomAI 对话,与可能指示免疫反应的生理变化趋势相吻合,这为 SymptomAI 的性能提供了进一步证据。
我们招募了 13,917 名经知情同意的研究参与者,每位参与者向五个随机分配的 SymptomAI 代理之一描述自己的症状。这些代理在症状问诊方式上具有不同程度的灵活性。在这些对话中,参与者描述症状,SymptomAI 提出后续问题,对话最终给出一个鉴别诊断列表(即一系列可能的诊断)和下一步行动建议。随后,参与者可以自行就诊,并在两周后通过调查问卷反馈就诊结果。为了评估 SymptomAI 的评估结果并建立基线,我们进行了一项临床专家标注研究。由三位持有委员会认证的临床医生组成的小组审查对话记录,并提供他们自己的评估(即鉴别诊断)。然后,每位临床医生以盲法方式,对 SymptomAI 提供的鉴别诊断和其余临床医生提供的鉴别诊断进行排序。
我们发现,在超过 50% 的病例中,临床医生更倾向于选择 SymptomAI 生成的鉴别诊断,而非其他临床医生提供的诊断。这表明,SymptomAI 的鉴别诊断与临床医生医学评估的一致性,至少与其他临床医生之间的互评结果相当,甚至更高。
同样,我们通过前五名准确率(即参与者个人医疗服务提供者给出的真实诊断,是否出现在鉴别诊断的五个可能诊断之一中)来比较 SymptomAI 生成和真实临床医生提供的鉴别诊断的准确性。我们请每位临床医生判断所提供的诊断是否出现在每个鉴别诊断列表中,这既包括 SymptomAI 生成的鉴别诊断,也包括临床医生提供的诊断。结果发现,临床医生认为 SymptomAI 生成的鉴别诊断在许多情况下比临床医生提供的鉴别诊断更准确。
作为本研究的一部分,我们评估了不同的病史采集问诊方法。参与者被随机分配到五个研究组,每组采用不同的提示策略。其中两组(动态实时组和动态终末组)被授予完全自主权,可提出不受限制的后续问题;另外两组(固定规范组和灵活规范组)则从医学院教授的一套标准病史采集问题中提问;最后是一组无提示的基础语言模型,代表了当前查询语言模型聊天机器人时完全由用户驱动的状态。我们发现,所有由代理驱动的提示策略(即 SymptomAI 主动提问)的表现都显著优于基础条件,证明了从参与者处主动获取信息对于提高鉴别诊断准确性的价值。
我们还发现,在临床医生自身对其鉴别诊断最不自信的病例中,SymptomAI 相对于临床医生基线的表现最为突出。
鉴于 SymptomAI 相对于临床基线的准确性,我们还可以探索其在规模化应用中的潜力。目前,临床标注的成本阻碍了对人群规模数据集的真实世界分析。像 SymptomAI 这样准确度高的症状检查系统,有望实现临床质量诊断的自动化参考标注,从而开启对生理数据的大规模分析——这项任务目前在大规模上是无法实现的。
其中一个例子是将可穿戴设备的生物信号与不同疾病类别相关联。可穿戴设备生物信号最显著的变化出现在急性呼吸道感染中。为在人群规模上进行研究,我们收集了参与者在与 SymptomAI 互动前最多 30 天的每日生物特征数据。我们发现了清晰的生物信号变化,表明在用户报告症状前的几天内出现了症状发作。重要的是,组间区分是通过对 SymptomAI 排名第一的候选诊断进行分类,并将归类为呼吸道感染的诊断分组而得到的。该队列排除了过敏性鼻炎或慢性阻塞性肺疾病等非传染性呼吸道疾病。对于这些参与者,可穿戴设备生物信号变化的峰值与症状报告日期相吻合,这作为观察性生理学证据,与参与者报告的症状相符。
基于人工智能的症状评估为新的研究打开了大门。通过使用 SymptomAI 分析大量的症状报告,并将其与实时的 Fitbit 数据配对,我们可以探索跨越多种疾病的数字生物信号表型。我们的分析揭示了用户与 SymptomAI 对话前几天内生理指标的显著变化——包括心血管功能、呼吸、皮肤温度和睡眠质量。这些客观变化与症状对话的时间点紧密吻合,为验证患者报告的症状提供了潜在方法,或提供了被动数据,有助于结合症状对话来辅助鉴别诊断。此外,这种实时可及性突显了人工智能症状检查器的一个核心优势。与可能因排期延误的传统临床预约不同,参与者可以在症状刚刚出现时就同步参与 SymptomAI 研究。这可能有助于提高患者报告症状发作时间线的准确性——这是人群规模健康分析中的关键细节。
SymptomAI 是一项探索性研究工作,可能代表了基于人工智能的症状评估领域的一项重大研究进展,并展示了其为寻求了解自身症状的公众提供帮助的潜力。虽然人群部署评估揭示了通过远程患者问诊进行症状评估的准确性,但在与临床医生的评估进行比较时,仍存在一些微妙的局限性。
首先,鉴别诊断本身就是一个模棱两可的任务,即使是报告出来的诊断也可能随着时间的推移而改变和发展。一次症状评估只是某个时间点的快照,捕捉的是该时刻呈现的症状。由于我们部署的规模,我们无法控制症状报告的频率和时间点。因此,一些参与者可能在更具代表性的指标出现之前就报告了症状,而另一些参与者则可能在多年慢性病经历后,在了解背景的情况下报告了明显的指标。未来的研究可能会聚焦于特定症状发展阶段的特定疾病,例如早期发作的代谢综合征或呼吸道感染初期的相关症状。研究期间生成的所有诊断、标签和疾病关联均为人工智能生成,仅用于研究分析,不构成经确认的临床诊断或官方医学评估。
其次,在我们的评估中,临床医生审查的是静态的对话记录,没有被赋予提问后续问题的自主权。如果他们主导症状问诊,可能会凭直觉获取不同的信息。此外,尽管近期研究表明,对话式人工智能系统能够以临床医生级别的细节和准确性收集临床数据,但这类系统可能会遗漏其他信号,如肢体语言、视觉评估、医疗记录,或者在初级诊疗环境中与患者已有的信赖关系。
总之,我们推出了 SymptomAI,这是一个用于进行真实世界患者问诊和症状评估的实验性对话式人工智能代理。我们通过在人群样本上验证其鉴别诊断的准确性,展示了 SymptomAI 端到端的真实世界性能,并说明了 SymptomAI 的诊断如何能够实现对人群规模信号(如可穿戴设备生物信号)的分析,以识别生理信号与报告疾病之间的关联。
本研究由乔·布雷达(Joe Breda)、杰克·桑夏恩(Jake Sunshine)和丹尼尔·麦克达夫(Daniel McDuff)共同完成。我们要感谢来自谷歌研究院和谷歌 DeepMind 的合作者对本研究的贡献。
英文来源:
July 22, 2026
Joseph Breda, Student Researcher, and Jake Sunshine, Research Scientist, Google Research
We present a first-of-its-kind research of AI for differential diagnosis and symptom checking through a national-scale study.
A large proportion of clinical diagnoses can be derived from language-based interviews alone. These diagnostic interviews are typically conducted by clinicians through doctor-patient interactions during in-person or remote visits. While these interactions are the gold standard for symptom assessment, they can often suffer from financial, geographic, and systemic barriers that limit their accessibility. Current language models (LMs) have demonstrated strong differential diagnosis assessment capabilities when evaluated on curated medical case-studies, highlighting their potential to support the diagnostic process. However, existing evaluations have largely relied on curated, highly detailed and sometimes synthetic patient vignettes, which may not reflect real world experience and clinical presentation variability. These evaluations do not capture how everyday patients report their health symptoms, for example with varying levels of medical literacy, incomplete information, and other complexities that arise through natural conversation. This represents a key gap, leading to uncertainty of how LMs might perform in real-world contexts.
To address this gap, we conduct an in-situ comparative research study of a set of experimental conversational prototype AI agents designed to explore how conversational AI might conduct end-to-end symptom interviews and differential diagnostic assessment for research benchmarking purposes. In our recent research paper, “SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessment”, we share results from a randomized national scale study (n=13,917) in which consented research participants interact with one of five possible Gemini Flash 2.0 SymptomAI agents. All diagnoses, labels, and disease associations generated during the study were for research analysis only and did not constitute confirmed clinical diagnoses or official medical assessments. Two weeks after their interaction with the AI agents, we asked research participants to report any diagnoses they received from a visit with a healthcare provider. Using this data, we conducted a clinical expert annotation study comparing SymptomAI’s diagnostic performance relative to real clinicians' medical assessments.
After assessing the accuracy of SymptomAI’s differential diagnoses (DDx), we further compare SymptomAI’s diagnoses against biosignals from participants’ Fitbit wearable devices in the time leading up to their conversation with SymptomAI. We show that SymptomAI conversations that led to diagnosis with an infectious disease etiology coincide with physiological trends that may indicate an immune response, suggesting further evidence of SymptomAI’s performance.
We enrolled 13,917 consenting research study participants who each describe their symptoms to one of five randomized SymptomAI agents, each with varying degrees of flexibility in how they conducted the symptom interview. During these conversations, participants described their symptoms and SymptomAI asked follow-up questions, with conversations culminating in a final differential diagnosis (DDx, a list of plausible diagnoses) and recommendations for next steps. Participants could then go on to see a healthcare provider and were asked to share the outcome of that visit via a survey two-weeks later. To evaluate and baseline SymptomAI’s assessment, we conducted a clinical-expert annotation study in which a panel of three board-certified clinicians reviewed the conversation transcripts and provided their own assessment (i.e., differential diagnosis). Then each clinician, in a blinded fashion, ranked the DDx provided by SymptomAI and those provided by the remaining clinicians.
We found that the clinicians preferred the DDx generated by SymptomAI over those provided by the other clinicians in over 50% of the cases. This indicates that SymptomAI DDx aligned with our clinicians’ medical assessments just as often or more often than that of other clinicians.
Similarly, we compare the accuracy of the DDx generated by SymptomAI and provided by real clinicians via top-5 Accuracy (i.e., whether the true diagnosis provided by our participants' personal healthcare provider appears as one of the five possible diagnoses in the DDx). We had our clinicians each identify whether the provided diagnosis was in each DDx, including both the DDx generated by SymptomAI and those provided by clinicians. We found that the clinicians ranked the DDx generated by SymptomAI to be accurate more often than the DDx provided by other clinicians.
As part of this research, we assessed different approaches for conducting history taking interviews. Participants were randomly assigned to five study arms, each employing different prompting strategies. Two (Dynamic Live and Dynamic Final) were given total agency to ask unrestricted follow up questions, two more (Fixed Canonical and Flexible Canonical) each asked questions from a set of standard history taking questions taught in medical school, and finally a Base unprompted LM, representing the fully user-driven experience that is the current status quo when querying LM chatbots. We found that all agent-driven prompting strategies (i.e., where SymptomAI actively asked follow up questions) significantly outperformed the Base condition, demonstrating the value of eliciting information from participants for improving differential diagnostic accuracy.
We found that SymptomAI’s performance above clinical baselines was greatest for cases where the clinician’s felt least confident in their own DDx.
Given SymptomAI's accuracy against clinical baselines, we can also explore its potential at scale. Currently, the cost of clinical labels prohibits real-world analyses of population-scale datasets. Accurate symptom checking systems like SymptomAI have the potential to enable automated reference labeling of clinical quality diagnosis, which can open up large-scale analyses of physiological data — a task that is currently impossible at scale.
One such example is correlating wearable biosignals with different categories of illness. The most notable changes in wearable biosignals are observed for acute respiratory infections. To study this at population scale, we collected daily biometric data from our consenting participants for up to 30 days prior to their interaction with SymptomAI. We find clear biosignal shifts indicating symptom onset in the days approaching the user's symptom reporting. Importantly, the separation between cohorts was derived through categorizing SymptomAI's top-1 candidate diagnosis and grouping diagnoses that were classified as respiratory infections. This cohort excludes non-infectious respiratory illnesses like allergic rhinitis or chronic obstructive pulmonary disease. The correlation of wearable biosignals shift peaks aligning with the date of symptom reporting for these participants serves as observational physiological evidence that align with their reported symptoms.
AI-based assessment of symptom presentations opens the door to new research. By using SymptomAI to analyze a large volume of symptom reports and pairing those with real-time Fitbit data, we can explore digital biosignal phenotypes across a wide range of diseases. Our analysis revealed distinct shifts in physiological metrics — including cardiovascular function, respiration, skin temperature, and sleep quality — in the days leading up to a user's SymptomAI conversation. These objective changes align closely with the timing of the symptom conversation, offering a potential way to validate patient-reported symptoms or provide passive data to help inform a differential diagnosis alongside their symptom conversation. Additionally, this real-time accessibility highlights a core benefit of AI symptom checkers. Unlike traditional clinical appointments that can suffer from scheduling delays, participants could take part on the SymptomAI research study contemporaneously while symptoms are fresh. This potentially could improve the accuracy of patient-reported onset timelines — a crucial detail for population-scale health analysis.
SymptomAI is an exploratory research effort that could represent a significant research advancement in AI-based symptom assessment and demonstrates the potential it could provide for the general public seeking understanding of their symptoms. While a population deployment evaluation reveals the accuracy of symptom assessment through remote patient interviews, there are nuanced limitations when comparing against clinician’s assessments.
Firstly, differential diagnosis itself is an ambiguous task and even reported diagnoses may change and develop longitudinally. A symptom assessment is a snapshot in time and captures the symptoms as they present in that moment. Due to the scale of our deployment, we were unable to control for frequency and timing of symptom reporting. As a result, some participants may have reported their symptoms well before more representative indicators developed, while others may have reported obvious indicators from an informed context after years of experience with chronic illness. Future work may focus on specific illnesses at specific points during symptom development such as early-onset metabolic syndrome or symptoms discussed at the start of respiratory infections. All diagnoses, labels, and disease associations generated during the study are AI-derived for research analysis only and do not constitute confirmed clinical diagnoses or official medical assessments.
Secondly, in our evaluation the clinicians reviewed static chat transcripts and were not given agency to ask their own follow-up questions. Clinicians may have intuitively sourced different information had they directed the symptom interview. Moreover, while recent research has shown that conversational AI systems can source clinical data with a clinician-level of detail and accuracy, such systems may miss alternative signals like body language, visual assessment, medical records, or in the context of primary care, existing rapport with the patient.
In conclusion, we introduce SymptomAI, an investigational conversational AI agent for conducting real-world patient interviews and symptom assessments. We demonstrate SymptomAI’s end-to-end real-world performance through DDx accuracy on a population sample, and show how SymptomAI diagnoses can enable analysis of population-scale signals like wearable biosignals for identifying associations in physiological signals with reported illness.
This work is the result of equal contributions from Joe Breda, Jake Sunshine and Daniel McDuff. We would like to thank our co-authors and collaborators from Google Research and Google DeepMind for their contributions to this work.
文章标题:SymptomAI:面向日常症状评估的对话型人工智能代理
文章链接:https://news.qimuai.cn/?post=4629
本站文章均为原创,未经授权请勿用于任何商业用途