GlucoFM:用于连续血糖监测的基础模型

内容来源:https://research.google/blog/glucofm-foundation-model-for-continuous-glucose-monitoring/
内容总结:
谷歌研究院日前发布新一代连续血糖监测(CGM)基础模型“GlucoFM”,该模型采用双流架构,将血糖监测数据中的缓慢趋势与短期波动分离处理,在多项临床代谢预测任务中刷新性能纪录,为糖尿病风险评估、胰岛素抵抗及β细胞功能障碍预测等提供了更高效的工具。
据谷歌研究院介绍,当前消费级可穿戴设备虽能通过运动与生理传感器估算活动量和睡眠状况,但对血糖调节只能提供间接观察。连续血糖监测仪虽可每隔数分钟通过皮下传感器捕捉空腹、夜间及餐后血糖模式,但解读这些数据仍具挑战,尤其当高质量临床标注稀缺且获取成本高昂时。
针对现有CGM基础模型(如CGMformer、GluFormer、CGM-JEPA)多采用单一表征流处理血糖数据的局限,GlucoFM创新性地将血糖动态分解为较缓的基线趋势与短期偏差两个独立流,同时保留每日时刻信息与数据缺失状态。模型通过潜在预测目标学习日常情境与时间演化,并在预训练阶段引入CGM感知的数据增强,模拟真实记录中的基线漂移、压缩样下降、稀疏采样及短暂断连等情况。
研究团队基于10.9万小时未标注CGM数据对GlucoFM进行预训练,并在四个不同队列(CGMacros、Stanford、Hall、ShanghaiT2DM)上针对七项临床预测任务(糖尿病风险、胰岛素抵抗、β细胞功能障碍、高脂血症、低血糖、肥胖及血糖类型)展开14项队列-任务组合评估。结果显示,GlucoFM的平均PR-AUC较最优GluFormer变体提高5.8个百分点,从基准的54.7升至58.8,绝对提升4.1点,相对提升约7.5%。在所有糖尿病风险与β细胞功能障碍评估中,GlucoFM均取得最高PR-AUC,并在四项胰岛素抵抗评估中三项领先。
在餐后血糖反应(PPGR)预测测试中,GlucoFM利用餐前信息预测餐后两小时的完整血糖变化轨迹,在Dexcom与Libre两款CGM设备分别建模的条件下,以21.88 mg/dL的平均绝对误差(MAE)优于最佳基线的22.90 mg/dL及训练集均值基线的27.69 mg/dL。此外,在多日表征聚合测试中,将最多七天的逐日表征取平均能显著提升受试者级预测性能,例如Stanford队列的β细胞功能障碍任务提升9.6点,Hall队列的糖尿病预测提升14.0点。
跨数据集迁移测试进一步验证了模型的泛化能力:在12项跨队列评估中,GlucoFM在11项中领先第二名0.5至8.6个PR-AUC点,仅在1项中落后0.6点,绝对PR-AUC值介于61.6%至90.0%之间。在少样本学习场景下,GlucoFM在每类仅一名标注受试者或仅1%观测数据的极端条件下仍保持最高任务平均PR-AUC,展现出高效的数据利用能力。
研究团队还通过消融实验证实双流设计的必要性:仅含事件(短期波动)的版本性能最弱,仅含状态(缓慢趋势)或直接处理原始数据的版本虽具竞争力,但完整双流模型持续最优。这表明将快慢血糖动态作为互补流显式建模后再融合,优于依赖单一数据流。
谷歌研究院表示,GlucoFM的当前预训练人群规模仍属中等,下一步计划在更大、更多样化的人群上训练模型,并将模型从独立处理的24小时窗口扩展至原生多日建模,以捕捉数周至数月尺度的趋势,同时探索实时动态变化的处理能力。该研究已获得相关伦理委员会批准,参与者均签署知情同意书,用于去标识化二次研究与算法开发。
中文翻译:
2026年8月26日
艾哈迈德·A.梅特沃利,高级研究科学家,与李泽辰,学生研究员,谷歌研究院
GlucoFM是一个轻量级、自监督的连续血糖监测(CGM)基础模型,它将较缓慢的血糖趋势和短期波动分别建模为独立的数据流,生成可迁移的表征,并在多种代谢预测任务中树立了新的性能标杆,这些任务包括糖尿病风险评估、胰岛素抵抗、β细胞功能障碍以及餐后血糖反应。
消费级可穿戴设备利用运动和生理传感器来估算活动和睡眠,但这些信号只能间接反映血糖调节状况。连续血糖监测仪(CGM)通过在皮下植入小型传感器,每隔几分钟追踪一次组织间液葡萄糖,捕捉空腹、夜间和餐后模式,从而对这些测量形成补充。然而,理解这些血糖轨迹仍然充满挑战,尤其是当用于辅助解读的高质量临床标签数据稀缺且获取成本高昂时。
许多现有的CGM基础模型——包括CGMformer、GluFormer和CGM-JEPA——都通过单一表征流来处理血糖数据,而非明确区分缓慢的基线动态和瞬态事件动态。但CGM数据并非无差别的数据流:它包含相对缓慢的基线模式,其间穿插着可能反映进餐、活动或传感器伪影的短期波动。如果我们能够利用日常CGM数据,在仅有有限标注数据的情况下,估算糖尿病风险、胰岛素抵抗和β细胞功能障碍等指标,会怎样呢?
这正是我们构建GlucoFM的原因——一个具有双流设计的自监督基础模型,它将较缓慢的血糖趋势与短期波动分开处理,同时保留一天中的时间信息和数据缺失模式。随后,潜在预测目标学习它们的日常背景和时间演化规律。我们在四个不同的队列上、针对七项临床预测任务——糖尿病风险、胰岛素抵抗、β细胞功能障碍、高脂血症、低血糖、肥胖和血糖分型——对GlucoFM进行了评估,共计14项队列-任务组合评估。在这些评估中,GlucoFM的PR-AUC(精确率-召回率曲线下面积)平均比表现最佳的GluFormer变体高出5.8个百分点,且两者均在同一语料库上进行预训练。在PR-AUC指标上,GlucoFM在所有糖尿病风险和β细胞功能障碍评估以及四项胰岛素抵抗评估中的三项中均处于领先地位。我们还在餐后血糖反应(PPGR)预测上对GlucoFM进行了评估。在匹配输入和评估协议的情况下,GlucoFM在两种CGM设备(Dexcom和Libre)上取得了最低的平均绝对误差(MAE)。此外,GlucoFM实现了最佳的跨数据集迁移性能,并展现出强大的少样本适应能力,即使在来自新队列或标注受试者数据极其有限的情况下也是如此。
我们在来自Wear-CGM的109,066小时无标注CGM数据上对GlucoFM进行了预训练。
CGM记录可能包含数据缺口、不同的采样间隔和传感器伪影。GlucoFM将每次记录对齐到一个24小时、五分钟间隔的网格上,并保留一个观测掩码,以区分测量位置和未观测位置。其双流编码器将较低频率的状态分量(代表较缓慢的血糖趋势)与残差事件分量(捕捉可能源于生理、行为或传感伪影的短期波动)分离开来。
GlucoFM并不重建可能受测量噪声和传感器伪影影响的精确原始血糖读数,而是使用带有两个互补任务的潜在预测预训练方法:
最后,CGM感知的数据增强引入了基线漂移、类似压缩的骤降、更稀疏的采样和短暂断连,使模型接触到真实CGM记录中会遇到的变化和数据缺失情况。
我们在四个队列(CGMacros、Stanford、Hall和ShanghaiT2DM)和七项临床预测任务上评估了GlucoFM,并单独评估了两小时餐后血糖反应预测。具体来说,我们探究了以下问题:其冻结表征是否能从未见参与者的单个24小时窗口提供有用信息;这些表征是否为预测餐后血糖轨迹提供了有用的历史背景;组合多天数据是否能改善受试者级别的预测;这些表征对新队列的迁移效果如何;以及当标注数据有限时它们的适应效果如何。
首先,我们使用了受试者不相交的窗口级线性探针。我们冻结每个模型的编码器,在单个24小时表征上训练线性分类器,并确保没有任何参与者同时出现在训练集和测试集中。这检验了单日表征能否为未见参与者提供表型信息,同时保留逐日变异性。
在评估的方法中,GlucoFM取得了最高的任务平均PR-AUC。在14项队列-任务评估中,它将平均PR-AUC从最强CGM特定基线(在同一数据上重新训练)的54.7提升至58.8——绝对增益为4.1个百分点,相对该基线提升约7.5%。GlucoFM在所有糖尿病风险和β细胞功能障碍评估以及四项胰岛素抵抗评估中的三项中取得了最高PR-AUC。
为了在动态预测任务上测试GlucoFM,我们利用每次记录的进餐前可用信息,预测相对于进餐开始值的完整两小时血糖变化轨迹。我们使用受试者不相交的交叉验证评估了来自34名参与者的874次配对进餐事件,Dexcom和Libre(两种CGM设备)在完全相同的划分下分别建模。
我们逐步将每个冻结模型表征与餐前一小时CGM数据、膳食营养信息——包括能量、碳水化合物、脂肪、蛋白质和膳食纤维——以及参与者级别的信息(如空腹血糖、BMI和糖尿病状态)相结合。在完整上下文下,GlucoFM在评估模型中取得了最低的平均MAE:21.88 mg/dL,而最佳基线为22.90 mg/dL,训练集均值基线为27.69 mg/dL。这些结果表明,GlucoFM为预测餐后血糖变化提供了互补的历史背景。
单个24小时轨迹可能无法完全反映一个人的血糖模式,因此我们测试了组合多天数据是否能改善受试者级别的预测。GlucoFM分别对每天进行编码,并将最多七天的表征取平均,每位参与者贡献权重相等。
如下表所示,在大多数数据集的大多数设置下,增加天数改善了PR-AUC,包括Stanford β细胞功能障碍提高了9.6个百分点,Hall糖尿病预测提高了14.0个百分点。CGMacros在Dexcom、Libre和融合传感器数据上也大多呈现正向增益。ShanghaiT2DM的胰岛素抵抗是简单平均下的主要例外,表明最佳聚合策略可能因任务而异。总体而言,GlucoFM的冻结每日表征可以在无需重新训练编码器的情况下进行组合,以增强受试者级别的预测。
接下来,我们想知道我们的模型学习到的生理模式是否能够泛化。如果我们使用一个临床队列的数据训练一个下游分类器来识别糖尿病风险,它能否在来自完全不同研究的患者身上同样有效?跨数据集迁移条形图凸显了GlucoFM如何应对这一挑战,具体展示了其相对于第二佳模型在糖尿病风险和胰岛素抵抗上的直接提升。
在下图中,正条形表示GlucoFM优于最强的竞争方法:它在12项评估中的11项中以0.5至8.6个PR-AUC百分点的优势领先,仅在1项中落后0.6个百分点。其绝对PR-AUC范围从Stanford到Hall两项任务的61.6%到Hall到CGMacros胰岛素抵抗的90.0%,表明对潜在生理学的关注帮助冻结表征超越队列特定的噪声,找到普遍的代谢模式。
标注临床数据的获取成本很高,因此我们还在两种少样本设置下测试了GlucoFM:左图改变每类的标注参与者数量,右图改变每位参与者可用的观测比例。向右移动添加更多的标注数据,更高的点表示更好的任务平均PR-AUC。
橙色的GlucoFM标记在每个评估的数据预算下都是最高的,包括每类仅1名参与者和仅1%观测值的最有限设置。当标注受试者稀缺时,这一优势尤为明显,表明该模型即使只有少量样本也能高效地捕捉到正确的信号。
我们还仔细研究了将信号分成两个流是否真的产生了差异。我们将完整的双流设计与更简单的替代方案进行了比较:一种直接处理原始血糖,一种设计用于强调较慢趋势,另一种设计用于强调较快、短期的波动。
正如我们的编码器设计分析所示,“仅事件”版本是最弱的,证明了仅凭瞬态波动不足以形成稳定的代谢图景。虽然原始输入和“仅状态”版本具有竞争力,但完整的双流模型始终表现最佳。这些结果支持将较慢和较快的血糖动态组织为互补的数据流后再进行组合,而不是仅依赖其中任何一个流。
我们的结果表明,CGM模型可以受益于明确考虑血糖动力学的多尺度结构,包括较慢趋势、短期波动、每日时间和传感器数据缺失。通过从无标注CGM数据中学习可复用的模式,GlucoFM生成的表征在评估的预测、迁移和少样本设置中均表现强劲,为更好地利用有限的标注临床数据提供了一种途径。
代谢反应因人群、队列和传感器设备而异,而我们当前的预训练人群规模仍然有限。我们的下一步是在更大、更多样化的人群上进行训练,并将GlucoFM从独立处理的24小时窗口扩展为原生多日建模,以捕捉跨越数周或数月的趋势,同时探索这些表征如何处理实时变化。关于代谢健康,我们还有很多要学习的地方,我们期待看到这些工具将引领我们走向何方。
以下研究人员为这项工作做出了贡献:李泽辰、基尔塔娜·纳塔拉詹、张伟志、西蒙·A.李、张玉伟、马克斯韦尔·A.徐、周梦莲、泽纳布·埃斯梅尔普尔、弗洛拉·D.萨利姆(来自新南威尔士大学)、马克·马尔霍特拉、林赛·桑登、什韦塔克·帕特尔、杨宇哲和艾哈迈德·A.梅特沃利。
我们衷心感谢鲍巴克·J.莫尔塔扎维和里卡多·古铁雷斯-奥苏纳(德克萨斯A&M大学)提供了本研究中使用的CGMacros数据集。
两项Wear-CGM研究已获得Advarra的批准(IRB编号:Pro00059582和Pro00069880),参与者签署了书面知情同意书,同意将去标识化数据用于二次研究和算法开发。已发表的数据集均在其各自的伦理审批和同意程序下收集。
英文来源:
August 26, 2026
Ahmed A. Metwally, Staff Research Scientist, and Zechen Li, Student Researcher, Google Research
GlucoFM is a lightweight, self-supervised CGM foundation model that models slower glucose trends and short-term deviations in separate streams, producing transferable representations and setting new performance standards across diverse metabolic prediction tasks, including diabetes risk assessment, insulin resistance, beta-cell dysfunction, and post-prandial glycemic response.
Consumer wearables use motion and physiological sensors to estimate activity and sleep, but these signals provide only an indirect view of glucose regulation. Continuous glucose monitors (CGM) complement these measurements by tracking interstitial glucose every few minutes through a small sensor inserted under the skin, capturing fasting, overnight, and post-meal patterns. Yet making sense of these traces remains challenging, especially when high-quality clinical labels that help interpret them are sparse and costly to obtain.
Many existing CGM foundation models — including CGMformer, GluFormer, and CGM-JEPA — process glucose through a single representation stream rather than explicitly separating slow baseline and transient event dynamics. But CGM is not an undifferentiated data stream: it contains relatively slow baseline patterns punctuated by short-term deviations that may reflect meals, activity, or sensor artifacts. What if we could leverage daily CGM data to estimate things like diabetes risk, insulin resistance, and beta-cell dysfunction using limited labeled data?
That's why we built GlucoFM, a self-supervised foundation model with a dual-stream design that separates slower glycemic trends from short-term deviations while preserving time-of-day and missingness. Latent-prediction objectives then learn their daily context and temporal evolution. We evaluated GlucoFM across four diverse cohorts on seven clinical prediction tasks — diabetes risk, insulin resistance, beta-cell dysfunction, hyperlipidemia, hypoglycemia, obesity, and glucotype — comprising 14 cohort–task evaluations. Across these evaluations, GlucoFM’s PR-AUC was 5.8 percentage points higher on average than that of the best-performing GluFormer variant evaluated, with both pre-trained on the same corpus. On PR-AUC, GlucoFM led all diabetes-risk and beta-cell-dysfunction evaluations and three of four insulin-resistance evaluations. We also evaluated GlucoFM on postprandial glycemic response (PPGR) forecasting. Under matched inputs and evaluation protocols, GlucoFM achieved the lowest mean absolute error (MAE), averaged across two CGM devices (Dexcom and Libre). Moreover, GlucoFM achieved the best overall cross-dataset transfer performance and demonstrated strong few-shot adaptation, even when data from a new cohort or labeled subjects are extremely limited.
We pre-trained GlucoFM on 109,066 hours of unlabeled CGM data from Wear-CGM
CGM recordings can contain gaps, different sampling intervals, and sensor artifacts. GlucoFM aligns each recording to a 24-hour, five-minute grid and retains an observation mask, keeping measured and unobserved positions distinct. Its dual-stream encoder separates a lower-frequency state component, representing slower glycemic trends, from a residual event component capturing short-term deviations that may arise from physiology, behavior, or sensing artifacts.
Rather than reconstructing exact raw glucose readings, which can be affected by measurement noise and sensor artifacts, GlucoFM uses latent predictive pre-training with two complementary tasks:
Finally, CGM-aware augmentations introduce baseline drift, compression-like drops, sparser sampling, and short disconnections, exposing the model to variation and missingness encountered in real CGM recordings.
We evaluated GlucoFM across four cohorts (CGMacros, Stanford, Hall and ShanghaiT2DM) and seven clinical prediction tasks, alongside a separate assessment of two-hour postprandial glycemic response prediction. Specifically, we asked whether its frozen representations are informative for individual 24-hour windows from unseen participants; whether they provide useful historical context for predicting postprandial glucose trajectories; whether combining multiple days improves subject-level prediction; how well the representations transfer to new cohorts; and how effectively they adapt when labeled data are limited.
First, we used subject-disjoint window-level linear probing. We froze each model’s encoder, trained a linear classifier on individual 24-hour representations, and ensured that no participant appeared in both the training and test folds. This tests whether a single-day representation is phenotype-informative for unseen participants while retaining day-to-day variability.
GlucoFM achieved the strongest task-averaged PR-AUC among the evaluated methods. Across 14 cohort–task evaluations, it increased average PR-AUC from 54.7 for the strongest CGM-specific baseline retrained on the same data to 58.8 — an absolute gain of 4.1 points, or approximately 7.5% relative to that baseline. GlucoFM achieved the highest PR-AUC in all diabetes-risk and beta-cell-dysfunction evaluations and in three of four insulin-resistance evaluations.
To test GlucoFM on a dynamic prediction task, we used information available before each logged meal to predict the complete two-hour glucose-change trajectory relative to the meal-start value. We evaluated 874 paired meal events from 34 participants using subject-disjoint cross-validation, with Dexcom and Libre (two CGM devices) modeled separately under identical splits.
We progressively combined each frozen model representation with one hour of pre-meal CGM, meal nutrition — including energy, carbohydrate, fat, protein, and dietary fiber — and participant-level information such as fasting glucose, BMI, and diabetes status. With the full context, GlucoFM achieved the lowest mean MAE among the evaluated models: 21.88 mg/dL, compared with 22.90 mg/dL for the best baseline and 27.69 mg/dL for the train-fold mean baseline. These results suggest that GlucoFM provides complementary historical context for predicting postprandial glucose changes.
A single 24-hour trace may not fully capture a person’s glucose patterns, so we tested whether combining multiple days improves subject-level prediction. GlucoFM encoded each day separately and averaged representations across up to seven days, with each participant contributing equally.
As the chart below shows, additional days improved PR-AUC in most settings across most datasets, including gains of 9.6 points for Stanford beta-cell dysfunction and 14.0 points for Hall diabetes prediction. CGMacros also showed mostly positive gains across Dexcom, Libre, and fused sensor data. ShanghaiT2DM insulin resistance was the main exception under simple averaging, indicating that the best aggregation strategy can vary by task. Overall, GlucoFM’s frozen daily representations can be combined to strengthen subject-level prediction without retraining the encoder.
Next, we wanted to know if the physiological patterns our model learns can generalize. If we train a downstream classifier to spot diabetes risk using data from one clinical cohort, will it still work on patients from an entirely different study? The cross-dataset transfer bar chart highlights how GlucoFM handles this challenge, specifically plotting its direct improvement over the second-best model for diabetes risk and insulin resistance.
In the chart below, positive bars indicate that GlucoFM outperformed the strongest competing method: it led in 11 of 12 evaluations by 0.5 – 8.6 PR-AUC points and trailed once by 0.6 points. Its absolute PR-AUC ranged from 61.6% for both Stanford-to-Hall tasks to 90.0% for Hall-to-CGMacros insulin resistance, showing that focus on underlying physiology helps the frozen representations look past cohort-specific noise to find universal metabolic patterns.
Labeled clinical data are expensive to obtain, so we also tested GlucoFM under two few-shot settings: the left plot varies the number of labeled participants per class, while the right varies the fraction of observations available from each participant. Moving right adds labeled data, and higher points indicate better task-averaged PR-AUC.
The orange GlucoFM markers are highest at every evaluated data budget, including the most limited settings of one per class and 1% of observations. The advantage is especially clear when labeled subjects are scarce, showing the model is highly efficient at picking up the right signals even with just a handful of examples.
We also looked closely at whether splitting the signal into two streams actually made a difference. We compared the full dual-stream design with simpler alternatives: one that processes raw glucose directly, one designed to emphasize slower trends, and one designed to emphasize faster, short-term deviations.
As our encoder design analysis showed, the "event-only" version was the weakest, proving that transient fluctuations alone are not enough for a stable metabolic picture. While the raw-input and "state-only" versions were competitive, the full dual-stream model consistently came out on top. These results support organizing slower and faster glucose dynamics as complementary streams before combining them, rather than relying on either stream alone.
Our results suggest that CGM models can benefit from explicitly accounting for the multiscale structure of glucose dynamics, including slower trends, short-term deviations, daily timing, and sensor missingness. By learning reusable patterns from unlabeled CGM, GlucoFM produced representations that performed strongly across the evaluated prediction, transfer, and few-shot settings, offering a way to make better use of limited labeled clinical data.
Metabolic responses vary across people, cohorts, and sensor devices, while our current pre-training population remains modest. Our next steps are to train on larger and more diverse populations and extend GlucoFM beyond independently processed 24-hour windows toward native multi-day modeling to capture trends that unfold over weeks or months, and explore how these representations handle real-time changes. There is still so much to learn about metabolic health, and we are excited to see where these tools take us next.
The following researchers contributed to this work: Zechen Li, Keerthana Natarajan, Weizhi Zhang, Simon A. Lee, Yuwei Zhang, Maxwell A Xu, Menglian Zhou, Zeinab Esmaeilpour, Flora D. Salim (from the University of New South Wales), Mark Malhotra, Lindsey Sunden, Shwetak Patel, Yuzhe Yang, and Ahmed A. Metwally.
We gratefully acknowledge Bobak J. Mortazavi and Ricardo Gutierrez-Osuna (Texas A&M University) for providing the CGMacros dataset used in this study.
The two Wear-CGM studies were approved by Advarra (IRB nos. Pro00059582 and Pro00069880), and participants provided written informed consent for de-identified secondary research and algorithm development. The published datasets were collected under their respective ethics approvals and consent procedures.