超越BMI:利用智能手机图像评估心血管代谢风险

内容总结:
谷歌研发智能手机拍照评估胰岛素抵抗新方法,准确率接近临床金标准
谷歌研究院的科研团队近日发布了一项最新研究成果,展示了一种名为“PhotoScan”的深度学习技术,该技术能够仅通过智能手机照片估算人体体成分,并在预测胰岛素抵抗方面达到了与临床DXA扫描相当的准确度。相关研究由谷歌研究院研究科学家Cassie Zhou和Staff Research Scientist Ahmed Metwally主导。
胰岛素抵抗是当代代谢疾病最隐蔽的驱动因素之一,往往在2型糖尿病临床确诊前数年就已存在,并悄悄损害血管健康、肝脏功能和能量代谢。目前,临床上常用HOMA-IR指数评估胰岛素抵抗,该指数高于2.9即被视为胰岛素抵抗。近年研究表明,结合可穿戴设备与常规化验的多模态机器学习可较准确预测HOMA-IR,而体成分的客观测量则提供了对肥胖程度的直接结构评估,与可穿戴设备形成互补。
研究指出,除了总体脂率外,Android-Gynoid脂肪比(A/G比)和内脏脂肪与皮下脂肪面积比(V/S比)等指标具有更深的临床意义。目前测体成分的金标准是双能X射线吸收测定法(DXA),但其价格昂贵、需要专业设备且伴有微量辐射,不适合日常筛查。
为此,谷歌团队开发了PhotoScan框架,利用在UK Biobank超过3.5万名参与者数据上预训练、并在677名成年人新队列中微调的深度神经网络,直接通过二维智能手机照片估算三维体成分指标,包括体脂率、A/G比和V/S比。在临床队列验证中,PhotoScan在体脂率预测上的准确度优于基于智能手表的生物电阻抗分析(BIA)技术,并能提供BIA无法测得的A/G和V/S比值,具备接近DXA的准确度,同时实现无创、可扩展。
研究分三阶段展开。在PhotoBIA队列(经5折交叉验证)中,PhotoScan模型对体脂率的平均绝对误差为2.15,而BIA模型为2.91;A/G和V/S的平均绝对误差分别为0.107和0.094。在独立验证的MetabolicMosaic队列中,体脂率、A/G和V/S的误差分别达到2.13、0.085和0.085,与PhotoBIA队列高度一致,显示出良好稳定性。
为评估临床价值,团队使用梯度提升分类器,在MetabolicMosaic队列上测试了五种不同输入特征集预测胰岛素抵抗的能力,包括基础人口学特征、传统卷尺测量、智能手表BIA、PhotoScan指标以及DXA扫描。结果显示,基线人口学模型的AUROC为0.692;加入PhotoScan体成分数据后,AUROC提升至0.760,净重分类指数(NRI)达0.593,与使用临床DXA数据的效果(AUROC 0.773,NRI 0.748)几乎相当。而仅添加BIA数据(限于体脂率)对提高胰岛素抵抗分类能力几乎没有帮助。
研究团队总结指出,PhotoScan提供了一种介于临床DXA(高精度但难以普及)和可穿戴BIA(便捷但信息有限)之间的可行方案,从普通智能手机图像中即可获得接近DXA水平的细致体成分数据。这项研究证明了光学体成分估算在技术上可行且具有临床信息价值,尽管目前仍属研究原型,但为便捷、无创的胰岛素抵抗风险筛查开辟了新路径。未来团队计划探索多模态数据融合,将体成分估算与连续穿戴数据、葡萄糖动态及临床血液标志物结合,以支持更全面、更可及的个人代谢健康管理。
中文翻译:
2026年8月17日
Cassie Zhou,研究科学家,Ahmed Metwally,高级研究科学家,谷歌研究
我们证明了PhotoScan——一种通过智能手机照片估算身体成分的深度学习方法——在临床研究环境中预测胰岛素抵抗的准确度可与DXA扫描相媲美。
胰岛素抵抗是现代代谢性疾病中最关键但最容易被漏诊的驱动因素之一。在2型糖尿病临床发病前数年,胰岛素敏感性受损就已悄然损害血管健康、肝脏功能和能量代谢,远在空腹血糖升高至诊断范围之前。胰岛素抵抗稳态模型评估(HOMA-IR)模拟稳态空腹条件下肝脏葡萄糖生成与胰岛素分泌之间的反馈回路,根据流行病学综述,HOMA-IR评分大于2.9即被视为胰岛素抵抗。近期研究表明,将可穿戴传感器数据与常规实验室检测相结合的多模态机器学习框架能够准确预测HOMA-IR,从而及早标记代谢风险。将身体成分的客观测量纳入其中,为可穿戴技术提供了重要补充;可穿戴设备追踪日常生理行为,而身体成分则提供对肥胖程度的独立结构性评估,共同构成代谢风险的完整图景。
虽然了解全身脂肪百分比是衡量肥胖程度与瘦体重的良好基线,但额外的身体成分生物标志物能提供更深入的临床洞察。例如,安卓型与梨形脂肪比(A/G比)比较的是躯干(“苹果”形)与臀部和大腿(“梨”形)储存脂肪的比例;内脏脂肪与皮下脂肪面积比(V/S比)则区分包裹器官的高代谢性内脏脂肪与紧贴皮肤下方的皮下脂肪。A/G比升高和内脏脂肪量增加与胰岛素抵抗患病率密切相关。目前,测量真实身体成分的金标准是双能X射线吸收法(DXA)扫描。这类扫描极其精确,但不适合日常筛查,因为其成本高昂、需要专门的临床基础设施,且会使患者暴露于低剂量辐射。
基于智能手机在日常使用中被动监测用户健康能力的不断增强,例如连续心率监测,我们推出PhotoScan:一种研究性深度学习框架,可直接从标准2D智能手机照片中估算三维身体成分指标,包括体脂百分比(BF%)、A/G比和V/S比。为此,我们在英国生物银行超过35,000份参与者记录上预训练了一个深度神经网络,并使用一个由677名成年人组成的多样化新队列对其进行了微调。经过临床队列验证,PhotoScan在体脂百分比准确度上优于基于智能手表的生物电阻抗分析(BIA)传感器,同时解锁了BIA无法提供的A/G比和V/S比,提供了一种可扩展、无创的框架,以接近DXA的准确度预测胰岛素抵抗。
PhotoScan绕过了临床测量,直接从智能手机图像中提取几何体形信息。我们通过三个关键阶段构建并评估了这一框架:
对于PhotoBIA队列,通过5折交叉验证评估,微调后的PhotoScan模型在PhotoBIA队列上的BF%预测平均绝对误差(MAE)为2.15,而基于BIA的模型MAE为2.91。A/G的平均MAE为0.107,V/S为0.094。对于MetabolicMosaic队列(独立验证),BF%、A/G和V/S的MAE均与PhotoBIA队列相当(BF%为2.13,A/G为0.085,V/S为0.085)。总体而言,微调后的PhotoScan模型在微调数据集和验证数据集之间表现出高度一致性。MetabolicMosaic队列中A/G和V/S的轻微下降源于其女性记录比例(67%)高于PhotoBIA队列(57%),因为女性通常由于以梨形、皮下脂肪储存为主而表现出较低的绝对A/G和V/S比值,这缩小了区域比值的方差和预测误差。
接下来,我们比较了不同数据组合预测胰岛素抵抗的效果,将基线人口统计学数据分别与标准卷尺测量、智能手表BIA传感器、PhotoScan和金标准DXA扫描相结合进行堆叠。
我们在MetabolicMosaic队列上使用梯度提升分类器测试了模型,以识别胰岛素抵抗受试者。为确保结果完全无偏且无信息泄露,我们实施了严格的测试流程,反复在未见数据上评估模型。我们还确保每个测试组在BMI和胰岛素抵抗状态上均均衡分布,以保证公平且真实的性能测试。在这一稳健框架下,我们系统地向分类器输入五组不同的特征集以比较其预测能力:基线人口统计学数据(如年龄、性别和体质指数)、标准卷尺人体测量数据、智能手表生物电阻抗数据、我们的智能手机PhotoScan指标,以及临床金标准DXA扫描数据。通过比较模型在每组独立输入下的表现,我们确立了智能手机光学表型分析的临床价值。
为评估模型,我们聚焦于两个关键指标:受试者工作特征曲线下面积(AUROC)和净重分类指数(NRI)。简单来说,AUROC衡量模型区分胰岛素抵抗患者与非患者的能力(越高越好)。而NRI则量化我们的新数字指标相对于旧基线模型在正确分类人群能力上的提升幅度。如下图所示,我们的基线人口统计学模型达到了0.692的AUROC。当加入基于PhotoScan的身体成分特征(人口统计学+PhotoScan)后,分类准确度提升至AUROC 0.760,NRI提升至0.593,几乎与使用临床DXA数据本身相当——后者的AUROC最高为0.773,NRI为0.748。相比之下,在人口统计学基础上加入BIA(下图中的人口统计学+BIA)在胰岛素抵抗分类方面未带来AUROC或NRI的提升,因为BIA仅提供BF%估算,而在人口统计学+PhotoScan模型中,其特征重要性显著低于A/G比和V/S比。
总体而言,我们的研究结果表明,基于智能手机的身体成分估算作为一种可扩展的心血管代谢研究工具是可行的。临床DXA成像提供最准确的身体成分数据但缺乏可扩展性,而可穿戴BIA传感器虽便捷但仅限于基础体脂百分比。我们的PhotoScan方法提供了一条有前景的中间路径,可从标准智能手机图像中以接近DXA的准确度估算精细身体成分。
归根结底,这项研究凸显了数字表型分析如何解决BMI等传统人体测量指标的局限性——这些指标常常遗漏身体成分中具有临床意义的差异。我们的研究表明,从智能手机图像进行光学身体成分估算在技术上是可行的,在临床上也具有参考价值。尽管仍是研究原型,这种方法为可及的、无创的胰岛素抵抗风险筛查指明了一条道路。
虽然这些结果令人鼓舞,但身体成分只是心血管代谢健康的一个组成部分。展望未来,我们的研究旨在探索多模态数据整合,将身体成分估算与连续可穿戴数据、血糖动态和临床血液生物标志物相结合。通过汇聚这些多样化信号,我们希望能为个人代谢健康的整体化、可及化管理提供支持。
英文来源:
August 17, 2026
Cassie Zhou, Research Scientist, and Ahmed Metwally, Staff Research Scientist, Google Research
We demonstrate the feasibility of PhotoScan, a deep learning approach estimating body composition from smartphone photos, to predict insulin resistance with accuracy comparable to DXA scans in a clinical research setting.
Insulin resistance is one of the most critical yet underdiagnosed drivers of modern metabolic disease. Predating the clinical onset of type 2 diabetes by years, impaired insulin sensitivity stealthily impairs vascular health, liver function, and energy metabolism long before fasting blood sugar rises into diagnostic ranges. Homeostasis Model Assessment for Insulin Resistance (HOMA-IR) models the feedback loop between liver glucose production and insulin secretion under steady-state fasting conditions, and a HOMA-IR score greater than 2.9 is considered insulin resistant based on epidemiological reviews. Recent studies demonstrate that multimodal machine learning frameworks integrating wearable sensor data with routine lab tests can accurately predict HOMA-IR to flag early metabolic risk. Integrating objective measures of body composition offers a vital complement to wearable technology; while wearables track daily physiological behaviors, body composition provides a distinct structural assessment of adiposity to form a complete picture of metabolic risk.
While knowing your total body fat percentage is a good baseline to measure adiposity versus lean mass, additional body composition biomarkers provide much deeper clinical insights. For instance, the Android-to-Gynoid fat ratio (A/G ratio) compares the fat stored in your trunk (an "apple" shape) versus your hips and thighs (a "pear" shape); the Visceral-to-Subcutaneous fat area ratio (V/S ratio) distinguishes between the highly metabolic internal fat surrounding your organs and the subcutaneous fat stored just beneath your skin. Elevated A/G ratios and higher visceral fat mass strongly correlate with insulin resistance prevalence. Currently, the gold standard for measuring true body composition is Dual-Energy X-Ray Absorptiometry (DXA) scans. These scans are incredibly precise, but aren't built for everyday screening because they are expensive, require specialized clinical infrastructure, and expose patients to low doses of radiation.
Building on the growing capability of smartphones to passively monitoring user health during daily use, such as continuous heart-rate monitoring, we introduce PhotoScan: an investigational deep learning framework that estimates three-dimensional body composition metrics including body fat percentage (BF%), A/G ratio and V/S ratio, directly from standard 2D smartphone photos. To build this, we pre-trained a deep neural network on over 35,000 participant records from the UK Biobank and fine-tuned it with a diverse new cohort of 677 adults. Validated across clinical cohorts, PhotoScan demonstrates higher body fat percentage accuracy than smartwatch-based bioelectrical impedance analysis (BIA) sensors while unlocking A/G and V/S ratios beyond BIA's capabilities, offering a scalable, non-invasive framework to predict insulin resistance with near-DXA accuracy.
PhotoScan bypasses the clinical measurements by extracting geometric body information directly from smartphone images. We built and evaluated this framework in three key phases:
For the PhotoBIA Cohort, evaluated via the 5 fold cross-validation, the fine-tuned PhotoScan model demonstrated an average mean absolute error (MAE) of 2.15 for BF% prediction across the PhotoBIA cohort, while the BIA-based model achieved an MAE of 2.91. The averaged MAE is 0.107 for A/G and 0.094 for V/S. For the MetabolicMosaic cohort (independent validation), the MAE of BF%, A/G and V/S are all comparable with the PhotoBIA Cohort (2.13 for BF%, 0.085 for A/G and 0.085 for V/S). Overall, the fine-tuned PhotoScan models demonstrated strong consistency between the fine-tuning and validation datasets. The minor reduction in A/G and V/S observed in the MetabolicMosaic cohort is driven by its higher proportion of female records (67%) than the PhotoBIA cohort (57%), as females generally exhibit lower absolute A/G and V/S ratios due to predominantly gynoid, subcutaneous fat storage, which reduces regional ratio variance and the prediction error.
Next, we compared how well different combinations of data predicted insulin resistance, stacking our baseline demographics against combining it with standard tape measurements, smartwatch BIA sensors, PhotoScan, and gold-standard DXA scans.
We tested our models on the MetabolicMosaic cohort using a gradient boosting classifier to identify subjects with insulin resistance. To ensure our results were completely unbiased and leak-free, we implemented a rigorous testing process that repeatedly evaluated the model on unseen data. We also made sure each test group was evenly balanced by both BMI and insulin resistance status, ensuring a fair and realistic performance test. With this robust framework in place, we systematically fed the classifier five distinct feature sets to compare their predictive power, baseline demographics like age, sex, and body mass index, standard tape measure anthropometrics, smartwatch bioelectrical impedance, our smartphone PhotoScan metrics, and the clinical gold-standard DXA scans. By comparing how the model performed with each of these isolated inputs, we established the clinical value of our smartphone optical phenotyping.
To evaluate our models, we focused on two key metrics: the Area Under the Receiver Operating Characteristic curve (AUROC) and the Net Reclassification Index (NRI). Simply put, AUROC measures how accurately a model can distinguish between someone who has insulin resistance and someone who does not (higher is better). NRI, on the other hand, quantifies exactly how much our new digital metrics improve our ability to correctly categorize people compared to our old baseline model. As the figure below indicates, our baseline demographic model achieved an AUROC of 0.692. When we added the photoscan-based body composition features (demo + photoscan), the classification accuracy improved to an AUROC to 0.760 and NRI improved to 0.593, nearly as effective as using clinical DXA data itself, which topped out at an AUROC of 0.773 and an NRI of 0.748. In contrast, adding BIA with demographics (demo + bia below) yielded no improvement in AUROC or NRI for insulin resistance classification for IR classification, as BIA only provides BF% estimation, whose feature importance is significantly lower than A/G ratio and V/S ratio in the demo + photoscan model.
Overall, our findings demonstrate the feasibility of smartphone-based body composition estimation as a scalable tool for cardiometabolic research. Clinical DXA imaging delivers the most accurate body composition but lacks scalability, whereas wearable BIA sensors offer convenience but are limited to basic body fat percentage. Our PhotoScan approach offers a promising middle ground, estimating granular body composition from standard smartphone imagery with near-DXA accuracy.
Ultimately, this research highlights how digital phenotyping can address key limitations of traditional anthropometrics like BMI, which often miss clinically significant variations in body composition. Our study demonstrates that optical body composition estimation from smartphone imagery is both technically feasible and clinically informative. While still a research prototype, this approach suggests a path toward accessible, non-invasive screening for insulin resistance risk.
While these results are encouraging, body composition is just one component of cardiometabolic health. Looking ahead, our research aims to explore multi-modal data integration, combining body composition estimation with continuous wearable data, glucose dynamics, and clinical blood biomarkers. By bringing these diverse signals together, we hope to support more holistic, accessible approaches to managing personal metabolic health.
文章标题:超越BMI:利用智能手机图像评估心血管代谢风险
文章链接:https://news.qimuai.cn/?post=4831
本站文章均为原创,未经授权请勿用于任何商业用途