教导人工智能说病理学的语言

内容来源:https://news.microsoft.com/signal/articles/teaching-ai-to-speak-the-language-of-pathology/
内容总结:
AI病理学新突破:微软与Paige联合发布多模态基础模型PRISM2
一项发表于《自然·医学》的最新研究显示,微软研究院与Paige(现为Tempus旗下公司)联合开发了病理学基础模型PRISM2。该模型创新性地将组织切片图像与真实病理报告中的语言信息相结合,旨在让AI从海量文本数据中学习,从而支持未来病理学的研究与开发。
在测试中,PRISM2在前列腺癌、乳腺癌及乳腺淋巴结转移检测等多个基准任务上,其表现达到甚至超越了专门的癌症检测系统。值得注意的是,研究团队并未为每项任务单独建模,仅凭一个统一模型即取得了上述成果。
意义何在?
病理学是癌症诊断的核心环节,病理医生通过观察组织样本撰写报告,指导后续治疗方案。随着医疗数据规模激增,学界正积极探索AI在辅助临床决策和提供新见解方面的潜力。当前多数病理AI系统均为单一任务设计,如仅识别某类癌症,而PRISM2作为一项研究性成果,可通过单一模型响应用户提示和问题,覆盖更广泛的任务范围。研究人员指出,这有望大幅降低未来病理工具开发的门槛,避免每次新应用都需从零开始。
技术原理:让AI“看懂”病理,也“听懂”病理语言
PRISM2的设计理念在于,病理学不仅是视觉科学,更是语言驱动的学科。研究团队将病理图像与报告中的诊断信息配对,生成了数百万个问答示例,帮助模型将视觉特征与临床诊断术语紧密关联。该模型既可单独处理图像,也能结合图文信息,用户通过自然语言提示即可交互,无需依赖特定任务软件。
这一路径与早期多数病理基础模型不同——后者主要侧重于从图像中学习视觉表征,而PRISM2从设计之初就致力于打通组织中的视觉模式与临床诊断语言之间的通道。
目前,PRISM2的完整模型权重已在Hugging Face平台向研究界公开提供。
中文翻译:
预计阅读时间:3分钟。
教会AI说病理学的语言
作者
人工智能已经在帮助临床医生识别诊断图像中的模式并发现疾病迹象。然而,当今许多病理学AI系统都是为单一任务而构建的,每开发一个新应用都必须重新搭建。
最近发表在《自然医学》上的一项研究描述了一种旨在推动该领域研究的不同方法。来自微软研究院和Paige(现为Tempus一部分)的研究人员开发了PRISM2,这是一种基于组织图像和真实病理报告语言训练的病理学基础模型。其目标是帮助AI从大量现有文本数据中学习,并创建一个能够支持未来病理学研究与开发的系统。
在测试中,该模型在多项基准任务上达到或超过了专门的癌症检测系统的性能,包括前列腺癌、乳腺癌和乳腺淋巴结转移检测。重要的是,研究人员报告称,这些结果无需为每项任务单独创建模型即可实现。
PRISM2的完整模型权重已在Hugging Face上公开供研究使用,链接见此处和此处。
为何重要
病理学处于许多癌症诊断的核心位置。病理学家检查组织样本并撰写报告,这些报告指导治疗决策。随着医疗保健领域产生的数据量越来越大,研究人员一直在探索AI如何支持临床医生做出这些决策,甚至可能带来新的见解。
当今大多数病理学AI系统都是为单一目的而设计的,例如检测某一种特定类型的癌症。PRISM2则是一项研究工作,旨在通过一个可以对提示和问题做出响应的单一模型来支持更广泛的任务。研究人员表示,这可以使未来病理学工具的构建和适配变得更加容易,而无需为每个新应用从头开始。
工作原理
PRISM2背后的核心思想是,病理学不仅仅是一门视觉学科,它也是一门由语言驱动的学科。
为了训练该模型,研究人员将病理图像与从病理报告中提取的信息配对,创建了数百万个问答示例,帮助将视觉发现与诊断语言联系起来。由此产生的模型既可以仅使用图像,也可以同时使用图像和文本,使用户能够通过提示与其交互,而不仅仅依赖特定任务的软件。
这种方法与许多早期的病理学基础模型不同,那些模型主要侧重于从图像中学习视觉表征。PRISM2从一开始就被设计为将组织中的视觉模式与临床医生在做出诊断时使用的语言相连接。
封面图片由Microsoft Copilot生成。
英文来源:
– The estimated reading time is 3 min.
Teaching AI to speak the language of pathology
Author
Artificial intelligence is already helping clinicians spot patterns in diagnostic images and detect signs of disease. But many of today’s pathology AI systems are built for a single task and must be rebuilt for each new application.
A study recently published in Nature Medicine describes a different approach aimed at advancing research in this field. Researchers from Microsoft Research and Paige, now part of Tempus, developed PRISM2, a pathology foundation model trained on both tissue images and language based on real pathology reports. The goal was to help AI learn from the large amount of available text data and create a system that could support future pathology research and development.
In testing, the model matched or exceeded the performance of specialized cancer-detection systems on several benchmark tasks, including prostate cancer, breast cancer and breast lymph node metastasis detection. Importantly, researchers reported these results without creating a separate model for each task.
The full PRISM2 model weights are publicly available for research use on Hugging Face here and here.
Why it matters
Pathology sits at the center of many cancer diagnoses. Pathologists examine tissue samples and create reports that guide treatment decisions. As healthcare generates larger volumes of data, researchers have been exploring how AI can support clinicians in making these decisions and possibly even deliver new insights.
Most pathology AI systems today are designed for a single purpose, such as detecting a particular type of cancer. PRISM2 was developed as a research effort to support a broader range of tasks through a single model that can respond to prompts and questions. Researchers say that could make it easier to build and adapt future pathology tools without starting from scratch for each new application.
How it works
The central idea behind PRISM2 is that pathology is not only a visual discipline. It is also a language-driven one.
To train the model, researchers paired pathology images with information derived from pathology reports, creating millions of question-and-answer examples that helped connect visual findings with diagnostic language. The resulting model can work with either images only or images and text, enabling users to interact with it through prompts rather than relying solely on task-specific software.
This approach differs from many earlier pathology foundation models, which focused primarily on learning visual representations from images. PRISM2 was designed from the outset to link visual patterns in tissue with the language clinicians use when making diagnoses.
Lead image created with Microsoft Copilot.