一款用于从可穿戴传感器数据中优先筛选候选生物标志物的人工智能工具

qimuai 发布于 阅读:55 一手编译

一款用于从可穿戴传感器数据中优先筛选候选生物标志物的人工智能工具

内容来源:https://research.google/blog/an-ai-tool-for-prioritizing-candidate-biomarkers-from-wearable-sensor-data/

内容总结:

谷歌研究团队于2026年8月21日发布了一项新成果,推出名为“生物标志物发现框架”(Biomarker Discovery Framework)的多智能体系统。该系统旨在从可穿戴设备传感器数据中,通过迭代假设生成、统计分析和基于文献的推理,自动挖掘候选生物标志物,以加速临床研究中的假设验证周期。

研究指出,当前可穿戴设备虽能大规模采集心率、睡眠等连续生理信号,但将这些数据转化为可靠、有临床意义的生物标志物仍是瓶颈。现有基于语言模型的智能体系统在处理生理时间序列数据时,常因优化预测性能而忽视统计严谨性,导致虚假关联或特征脆弱。新框架则通过“编排器”智能体将自然语言研究指令分解为执行计划,并引导多个专业智能体完成六阶段流程,包括假设生成、并行统计分析、模型训练、对抗验证与文献推理,全程保留人类监督和可追溯性。

在验证中,研究团队将该框架应用于三个大规模队列(共9279人次观测),覆盖心理健康(DWB和GLOBEM)和代谢疾病(WEAR-ME)领域,自动识别出41个心理健康和25个代谢结果的候选数字生物标志物。例如,在心理健康领域,框架发现睡眠时长变异性和入睡时间变异性与抑郁严重程度显著相关(如睡眠时长变异性与PHQ-8评分的斯皮尔曼相关系数为0.252);在代谢领域,它构建了“步数除以静息心率”的心血管适能指数,作为胰岛素抵抗的非侵入性关联指标。框架还生成了新的复合特征,而非简单选用已有变量。

在预测性能方面,将框架生成的数字生物标志物与人口统计学特征结合,可显著提升模型表现(抑郁预测ΔR²=0.040,胰岛素抵抗预测ΔR²=0.021)。此外,在由15名医学、生物信息学等领域专家进行的盲审评估中,该框架在全部七个质量维度上均获最高平均分。在模拟编辑评审中,它获得2项“接受”、8项“小修”、8项“大修”和3项“拒稿”,是唯一获得“接受”或“小修”建议的系统;评审专家估计平均会保留其生成手稿内容的56.9%,远超其他基线系统的18.8%至30.4%。

研究团队强调,随着可穿戴健康数据规模不断扩大,数字医学的瓶颈已从数据收集转向严谨的假设生成。单纯扩展模型能力并不能解决科学严谨性问题。这一框架尝试通过将确定性计算与生成式推理分离,并让智能体进行防御性辩论,在人类监督下实现结构化的假设生成、验证和优先级排序,从而安全地加速临床研究从假设到验证的循环。该研究由谷歌研究团队主导,主要作者包括麻省理工学院博士生Yubin Kim(在谷歌实习期间完成核心工作)、Hamid Palangi和Daniel McDuff。

中文翻译:

2026年8月21日
Yubin Kim,学生研究员
我们引入了生物标志物发现框架(Biomarker Discovery Framework),这是一个多智能体系统,通过迭代假设生成、统计分析和基于文献的推理,支持从可穿戴传感器数据中发现候选生物标志物。
可穿戴设备以人群规模捕获连续的生理信号。这些信号流,从心率动态到睡眠模式,可以在症状出现之前揭示早期的生理变化。瓶颈不再是数据收集,而是将这些信号转化为可靠的、具有临床意义的生物标志物。
现有的基于语言模型的智能体系统自动化了科学工作流程的某些部分,但在处理生理时间序列数据时往往容易失效。这些系统优化预测性能的同时忽视了统计有效性,导致虚假相关性、数据泄漏和脆弱特征。
为此,我们引入了生物标志物发现框架,这是一个多智能体系统,将候选生物标志物的优先排序构建为在人类监督下的迭代研究循环。通过结合假设生成、并行统计分析、模型训练、对抗性验证和基于文献的推理,生物标志物发现框架在保持严格统计严谨性和人类监督的同时加速了发现过程。在三个队列(N = 9,279个参与者观测)中,生物标志物发现框架恢复了已知的临床信号,识别了跨独立数据集的收敛生物标志物,并在与人口统计学特征结合时改善了下游预测。
生物标志物发现框架将用于数值分析的确定性计算与用于假设形成和解释的生成性推理相结合。一个编排者智能体将自然语言研究指令分解为执行计划,并引导专门的智能体通过六个阶段的流程。同时,共享内存、结构化事实表和通用工具在整个工作流程中保持可追溯性:
例如,给定一个优先排序与抑郁严重程度相关的可穿戴候选标志物的请求,生物标志物发现框架对DWB数据集进行了画像分析,提出了睡眠时间变异性特征,并估计了睡眠时长变异性与PHQ-8严重程度之间的关联(ρ = 0.252)。工作流程随后检查了稳定性、数据泄漏、亚组一致性和替代解释,然后将结果构建为基于文献的昼夜节律不稳定假设,供人工审查。
为了评估生物标志物发现框架从噪声数据中提取合理生理洞察的能力,我们将其独立应用于三个大规模队列,总计9,279个参与者观测,涵盖心理健康(DWB和GLOBEM)和代谢疾病(WEAR-ME)领域。该流程自主识别了41个心理健康候选数字生物标志物和25个代谢结局候选数字生物标志物。
下表显示了精选的候选关联样本。Spearman的ρ总结了关联的方向和强度。95%置信区间量化了不确定性,调整后的p值考虑了多重比较。机制部分提出了基于文献的假设而非因果结论。重要的是,最后一列描述了先前证据的强度——而非本研究中的临床验证——星号表示证据层级标记而非统计显著性代码。
生物标志物发现框架不仅仅是选择现有变量;它构建了新颖的复合特征。例如,在心理健康领域,它识别出睡眠时长变异性和入睡时间变异性作为抑郁严重程度的首要相关因素。在代谢领域,它推导出一个心血管健康指数(步数除以静息心率)作为胰岛素抵抗的非侵入性相关指标,并将其与此前关于葡萄糖调节和心肺代谢健康的研究联系起来。
我们在三个不同的大规模队列中部署了生物标志物发现框架,总计9,279个参与者观测,涵盖心理健康和代谢疾病领域。
在两个抑郁领域中,生物标志物发现框架对同一昼夜节律不稳定构念的不同操作化方式进行了优先排序。在DWB中,睡眠时长变异性与PHQ-8严重程度相关(ρ = 0.252, p < 0.001)。在GLOBEM中,入睡时间变异性作为探索性的低信号关联出现在PHQ-4中(ρ = 0.126, p < 0.001;CV AUC = 0.535)。由于队列、终点和特征定义不同——且没有相同的候选标志物被重复验证——这种模式应被解释为构念层面的暗示性收敛,而非直接复制。
总体而言,生物标志物发现框架识别了41个心理健康候选数字生物标志物和25个代谢结局候选数字生物标志物。虽然效应量反映了被动感知数字表型分析中典型的适度量级,但将这些生物标志物发现框架衍生的特征与人口统计学变量相结合,改善了预测性能(抑郁ΔR² = 0.040,胰岛素抵抗ΔR² = 0.021)。
为了评估稿件质量,15位来自医学、生物医学数据科学、机器学习、生物信息学和数字健康领域的专家审查了来自生物标志物发现框架和三个当代AI研究系统(Google DeepMind的AI合作科学家、Biomni和Google ADK的数据科学智能体)的盲审报告。生物标志物发现框架、Biomni和数据科学智能体在21个评审场次中一起评分,而生物标志物发现框架在另外13个场次中使用相同评估工具进行单独评分。
在盲审评估中,生物标志物发现框架在所有七个质量维度上获得了最高的平均分。在该研究模拟的编辑评审标准下,它是唯一获得“接受”或“小修”建议的系统:2项接受、8项小修、8项大修和3项拒稿。评审者估计他们平均会保留56.9%的生物标志物发现框架生成的稿件内容,而基线系统的保留率为18.8%–30.4%,并且在13个四系统排名场次中,评审者在9个场次将生物标志物发现框架列为第一。
随着可穿戴健康数据在人群中持续扩展,数字医学的瓶颈不再是数据收集,而是有原则的、严谨的假设生成。仅扩展模型能力并不能解决科学严谨性的问题。然而,当部署在精心构建的架构中,将确定性计算与生成性推理分离,并促使智能体防御性地辩论其发现时,AI可以在人类监督下支持结构化的假设生成、验证和优先级排序。通过从黑箱自动化转向透明的人机协同工作流程,我们可以构建能够安全加速临床研究中假设到验证循环的AI系统。
本文由Google Research的Yubin Kim、Hamid Palangi和Daniel McDuff撰写。这项工作由麻省理工学院博士生Yubin Kim在Google实习期间牵头,由Daniel McDuff和Hamid Palangi指导。我们感谢来自Google Research、Google DeepMind和学术界的合著者和合作者对本工作的贡献。

英文来源:

August 21, 2026
Yubin Kim, Student Researcher
We introduce the Biomarker Discovery Framework, a multi-agent system that supports the discovery of biomarker candidates from wearable sensor data through iterative hypothesis generation, statistical analysis, and literature-grounded reasoning.
Wearable devices capture continuous physiological signals at population scale. These streams, ranging from heart rate dynamics to sleep patterns, can reveal early physiological changes before symptoms appear. The bottleneck is no longer data collection, but turning these signals into reliable, clinically meaningful biomarkers.
Existing language model-based agent systems automate parts of the scientific workflow, but can often break down on physiological time-series data. These systems optimize for predictive performance while overlooking statistical validity, leading to spurious correlations, leakage, and brittle features.
To this end, we introduce the Biomarker Discovery Framework, a multi-agent system that structures candidate biomarker prioritization as an iterative research loop under human supervision. By combining hypothesis generation, parallel statistical analysis, model training, adversarial validation, and literature-grounded reasoning, Biomarker Discovery Framework accelerates the discovery process while maintaining strict statistical rigor and preserving human oversight. Across three cohorts (N = 9,279 participant-observations), Biomarker Discovery Framework recovered known clinical signals, identified convergent biomarkers across independent datasets, and improved downstream prediction when combined with demographic features.
Biomarker Discovery Framework combines deterministic computation for numerical analysis with generative reasoning for hypothesis formation and interpretation. An Orchestrator agent decomposes natural-language research directives into execution plans and guides specialized agents through a six-phase process. Meanwhile, shared memory, a structured fact sheet, and common tools preserve traceability across the workflow:
For example, given a request to prioritize wearable candidates associated with depression severity, Biomarker Discovery Framework profiled the DWB dataset, proposed sleep-timing variability features, and estimated an association between sleep-duration variability and PHQ-8 severity (ρ = 0.252). The workflow then checked stability, leakage, subgroup consistency, and alternative explanations before framing the result as a literature-grounded circadian-instability hypothesis for human review.
To assess the Biomarker Discovery Framework's capability to extract plausible physiological insights from noisy data, we applied it independently across three large-scale cohorts totaling 9,279 participant-observations, spanning both mental health (DWB and GLOBEM) and metabolic disease (WEAR-ME) domains. The pipeline autonomously identified 41 candidate digital biomarkers for mental health and 25 for metabolic outcomes.
The table below shows a curated sample of candidate associations. Spearman’s ρ summarizes the direction and strength of an association. The 95% confidence interval quantifies uncertainty, and the adjusted p-value accounts for multiple comparisons. The mechanism presents a literature-grounded hypothesis rather than a causal conclusion. Importantly, the final column describes the strength of prior evidence — not clinical validation in this study — and the stars denote evidence-tier markers rather than statistical-significance codes.
The Biomarker Discovery Framework did not simply select existing variables; it constructed novel composite features. For instance, in the mental health domain, it identified sleep duration variability and sleep onset variability as top correlates of depression severity. In the metabolic domain, it derived a cardiovascular fitness index (steps divided by resting heart rate) as a non-invasive correlate of insulin resistance, linking it to prior work on glucose regulation and cardiometabolic fitness.
We deployed the Biomarker Discovery Framework across three distinct large-scale cohorts totaling 9,279 participant-observations, spanning both mental health and metabolic disease domains.
Across the two depression domains, Biomarker Discovery Framework prioritized different operationalizations of a related circadian-instability construct. In DWB, sleep-duration variability was associated with PHQ-8 severity (ρ = 0.252, p < 0.001). In GLOBEM, sleep-onset variability emerged as an exploratory, low-signal association with PHQ-4 (ρ = 0.126, p < 0.001; CV AUC = 0.535). Because the cohorts, endpoints, and feature definitions differ — and no identical candidate was replicated — this pattern should be interpreted as suggestive construct-level convergence, not direct replication.
In total, the Biomarker Discovery Framework identified 41 candidate digital biomarkers for mental health and 25 for metabolic outcomes. While the effect sizes reflect the modest magnitudes typical of passive-sensing digital phenotyping, the integration of these Biomarker Discovery Framework-derived features alongside demographic variables improved predictive performance when combined with demographic features (ΔR² = 0.040 for depression, 0.021 for insulin resistance).
To assess manuscript quality, 15 experts in medicine, biomedical data science, machine learning, bioinformatics, and digital health reviewed blinded reports from the Biomarker Discovery Framework and three contemporary AI research systems (Google DeepMind’s AI co-scientist, Biomni, and Google ADK’s Data Science Agent). Biomarker Discovery Framework, Biomni, and the Data Science Agent were scored together in 21 sessions, and Biomarker Discovery Framework was scored in a separate 13-session set using the same evaluation instrument.
In the blinded evaluation, the Biomarker Discovery Framework received the highest mean scores across all seven quality dimensions. Under the study’s simulated editorial rubric, it was the only system to receive any “Accept” or “Minor Revision” recommendations: 2 Accept, 8 Minor Revision, 8 Major Revision, and 3 Reject. Reviewers estimated that they would retain 56.9% of Biomarker Discovery Framework-generated manuscript content on average, compared with 18.8%–30.4% for the baselines, and ranked the Biomarker Discovery Framework first in 9 of 13 four-system ranking sessions.
As wearable health data continues to scale across populations, the bottleneck in digital medicine is no longer data collection, but rather principled, rigorous hypothesis generation. Scaling model capability alone does not address the problem of scientific rigor. However, when deployed within a meticulously structured architecture that separates deterministic computation from generative reasoning, and forces agents to defensively debate their findings, AI can support structured hypothesis generation, validation, and prioritization under human supervision. By shifting from black-box automation to transparent, human-in-the-loop workflows, we can build AI systems capable of safely accelerating the hypothesis-to-validation cycle in clinical research.
This blog post was written by Yubin Kim, Hamid Palangi, and Daniel McDuff from Google Research. This work was spearheaded by MIT PhD student Yubin Kim during a Google internship advised by Daniel McDuff and Hamid Palangi. We are grateful to our co-authors and collaborators from Google Research, Google DeepMind, and academia for their contributions to this work.

谷歌研究进展

文章目录


    扫描二维码,在手机上阅读