人工智能如何帮助科学家设计下一代药物

内容总结:
AI赋能药物研发:阿斯利康如何用机器学习重塑新药发现流程
随着生成式AI进入公众视野,另一种人工智能正在悄然改变药物研发领域。机器学习模型正在帮助将长达十年的研发周期压缩,并攻克以往无法解决的难题。据估计,将一款新药推向市场平均需要10年以上的时间和数十亿美元的投资,而大多数候选药物最终都无法惠及患者。对于由工程化蛋白质制成的生物药物而言,其复杂性更是远超传统化学合成药物。
AI加速“设计-制造-测试-分析”闭环
科学家们以往需要在海量分子中寻找那些既能精准结合靶点、在人体内保持稳定,又具备规模化生产潜力的“幸运儿”。如今,AI正以核心基础设施的角色深度融入制药研发全流程。阿斯利康高级副总裁兼生物制剂工程与肿瘤靶向发现负责人普贾·萨普拉表示:“我们做的每一件事,无论是设计、制造、测试还是分析,现在都已通过计算增强。周期时间正在缩短,而生产力和创新力却在提升。”
阿斯利康采用“构建-测量-学习”的循环模式:AI首先生成或优先排序候选分子,预测哪些设计最有可能成功;科学家随后仅对排名靠前的候选分子投入实验室资源。这种紧密的反馈循环大幅减少了研究死胡同,实现了更快的迭代,并使过去被认为无药可治的疾病靶点成为可能。
攻克“不可成药”靶点,设计新一代多特异性药物
传统生物制剂通常只针对单一疾病通路,而新一代药物能够同时攻击多个靶点,或精准地将治疗载荷递送到特定细胞。实现这一目标需要同时优化大量变量。萨普拉指出,AI驱动模型有望帮助设计这些日益复杂的多特异性生物制剂:“这类模型可以基于底层生物学识别优先靶向哪些靶点,然后跨多个参数优化,平衡分子的效力、稳定性、可制造性和安全性。”她表示,“靶向‘不可成药’靶点正成为现实,这些技术最终将使我们对过去认为不可能触及的靶点开发药物。”
数据护城河:专有多模态数据集是核心竞争力
麦肯锡估计,结合其他计算工具,生成式AI可将药物发现时间缩短多达50%。但AI模型的质量完全取决于其训练数据。在药物发现中,这意味着需要充足的高质量生物数据。阿斯利康的数据集具有专有性和多模态特征,涵盖分子结构、结合测量值、安全性概况和生产结果。萨普拉说:“数据是我们的差异化优势。我们在多个疾病领域和药物类型上构建了有意多样化的组合,这些数据使我们能够用更丰富、更具代表性的训练集来微调前沿AI模型。”
打造“未来实验室”:AI与机器人形成闭环发现系统
阿斯利康正在马萨诸塞州剑桥市肯德尔广场建设名为“未来实验室”的设施。在这里,AI和机器人自动化将形成一个连续、闭环的发现系统。“就像自动驾驶汽车用传感器和模型导航一样,这个系统用AI进行预测,用机器人系统执行实验,用仪器生成数据。”萨普拉解释。这些数据直接反馈回模型,加速每个后续循环。科学家将始终处于核心位置,提供监督、判断和战略方向。最终,自动化高通量系统每周可制造并评估数千种分子相互作用。
终极目标:“从头设计”全新药物
萨普拉表示,AI在生物药物发现中的终极愿景是“从头设计”——让AI生成全新的蛋白质序列,精确匹配所需的药物特性,包括设计结构、预测安全性、体内行为及可制造性。“该领域正朝着完全由AI生成的、从零开始设计直至临床候选药物的方向取得巨大进展。”她强调。
但这需要几个关键要素:更丰富、更标准化的行业训练数据;针对AI生成候选药物的稳健评估基准;以及懂得在机器学习和生物学交叉领域工作的团队。其中,安全性预测可能是最重要也最容易被忽视的问题。阿斯利康正通过“虚拟临床试验”——先进的细胞系统和微尺度器官模型作为物理测试平台,配合AI学习其输出——来应对这一挑战。
人机协同:科学家与工程师是AI潜力释放的关键
当前正在发生的转变是向“智能体AI系统”迈进,这类系统能同时生成候选分子并预测其功效和安全性。“生物学的复杂性与分子设计密不可分。”萨普拉总结道。
对科学家而言,与AI合作是协作过程。“科学家将与这些模型系统携手工作,模型设计分子,科学家测试这些分子并整合所有数据。”对于工程师而言,设计高效、透明的人机协作系统至关重要。阿斯利康的工程团队正开发充当“思维伙伴”而非“黑箱”的系统。“工程师们设计了能够快速生成、验证和学习的系统。问题非常硬核:多模态数据融合、闭环优化、不确定性量化以及临床决策点的可解释性。”
萨普拉表示,迎接这些技术挑战的工程师和科学家有机会为许多疾病的潜在治疗突破做出贡献:“我们今天能开发的生物药物,以及明天将设计的药物,都取决于将世界级的AI和工程人才与深厚的科学专业知识相结合。”
中文翻译:
赞助内容
人工智能如何帮助科学家设计下一代药物
随着生成式人工智能吸引公众目光,另一种类型的人工智能正在重塑药物发现领域。机器学习模型正帮助压缩长达十年的研发时间线,并攻克此前无法解决的难题。
与阿斯利康合作
设计和开发一种新药是一项代价高昂、极易失败的科学挑战。一种新药可能需要多年时间才能开发出来,并需要投入巨额资金。即便如此,大多数候选药物也从未真正惠及患者。对于生物药物——即由工程化蛋白质而非合成化学方法制成的疗法(常用于治疗大多数主要急性和慢性疾病)——其复杂性甚至更高。
科学家们探索数量庞大的可能分子,寻找极少数能够与正确靶点结合、在人体内保持稳定且可大规模生产的分子。如今,人工智能正在加速这些过程,并迅速成为医药研发基础设施的核心组成部分。
人工智能辅助设计正日益成为生物候选药物开发的一部分,像阿斯利康这样的公司正在积极建设其工程团队以推动这一进程。“我们所做的一切,无论是设计、制造、测试还是分析,现在都得到了计算技术的增强,”阿斯利康高级副总裁、研发生物制剂工程与肿瘤靶向发现负责人普贾·萨普拉表示。“周期时间正在缩短,而生产力和创新力却在提高。”
萨普拉解释说,阿斯利康的方法遵循“构建-测量-学习”的循环。人工智能通过计算生成或优化候选分子,预测哪些设计最有可能成功。然后,科学家们将实验室资源只集中在排名最高的候选分子上。这带来了更紧密的反馈循环,减少了死胡同,加快了迭代速度,并使得能够攻克以前被认为无药可治的疾病靶点。由于可能的分子组合数量远超任何人类团队所能系统探索的范围,利用人工智能来缩小和优化待测试的选择范围,已成为生物制剂药物设计的一个主要焦点。
应对复杂的药物设计难题
除了加快时间线,人工智能还被应用于发现全新类别的药物。传统生物制剂通常靶向单一疾病通路。而下一代药物能够同时作用于多个靶点,或将治疗有效载荷精确递送到特定细胞。要实现这一点,需要同时对多个变量进行优化。展望未来,由人工智能驱动的模型可以帮助设计这些日益复杂的、多特异性的生物制剂,普贾·萨普拉解释道。“例如,”她继续说道,“这类模型可以根据潜在的生物学机制帮助确定优先考虑哪两个或三个靶点,然后跨多个参数进行优化,以平衡分子的效力、稳定性、可制造性和安全性。”“攻克不可成药靶点正成为现实,”萨普拉说。“这些技术最终将使我们能够针对那些曾经被认为无法触及的靶点开发药物。为患者带来的潜在益处是巨大的。”
数据护城河
麦肯锡估计,生成式人工智能与其他计算工具相结合,可将药物发现的时间线缩短多达50%。但每个人工智能模型的性能都取决于其训练数据的质量。在药物发现领域,这意味着需要大量高质量的生物学数据。实验可以提供此类数据的丰富来源。无论成功还是失败,每个实验都会生成一个关于什么有效、什么无效的信号。
“数据是我们的差异化因素,”萨普拉说,她解释了公司的数据集如何具有专有性和多模态性,并包含分子结构、结合测量数据、安全性概况和生产结果。“我们在多个疾病领域和药物类型上构建了有意多样化的组合。所有这些数据使我们能够用更丰富、更具代表性的训练集来微调前沿的人工智能模型。”她继续说道,“此外,我们投资了深度筛选技术,以生成批量所需的数据集,用于不断优化和验证我们的模型。”
构建自主发现引擎
为了将所有数据集中到一个地方,阿斯利康正在马萨诸塞州剑桥市肯德尔广场建设其所谓的“未来实验室”设施。在那里,人工智能和机器人自动化将能够形成一个连续的、闭环的发现系统。“无人驾驶汽车利用传感器和模型来导航其环境,而这个系统则利用人工智能进行预测,利用机器人系统执行实验,利用仪器生成数据,”萨普拉解释道。这些数据会直接反馈到模型中,加速后续的每一个循环。
“在整个过程中,科学家仍将是核心,提供监督、判断和战略指导,确保输出是可解释的、可耐受的,并旨在为患者带来潜在益处,”她补充道。
最终,自动化高通量系统将能够每周制造和评估数千种分子相互作用。“这将产生传统工作流程无法比拟的大量、人工智能就绪的数据,”萨普拉说。“机器人样品处理、自动化质量检查和集成数据管道也有望显著加速早期药物开发的时间线。”
下一个前沿:从头生成药物
最终,萨普拉表示,人工智能在生物制剂药物发现领域的最终愿景是该领域所谓的“从头设计”。为此,目标是让人工智能生成全新的蛋白质序列,使其精确符合所需的药物特性。这包括设计结构、预测安全性、预测其在体内的行为方式以及如何实现可制造性。
“该领域在实现完全由人工智能生成的、从零开始设计直至成为临床候选药物的生物制剂方面正取得巨大进展,”萨普拉说。“随着我们继续利用前沿模型,并用正确的数据集对其进行微调,我们正越来越接近这一现实。我相信它会到来。这只是时间问题。”
然而,要达到这一点,需要几个关键要素。首先,整个行业需要更丰富、更标准化的训练数据。其次,需要针对人工智能生成的候选药物建立稳健的评估基准。第三,需要懂得如何在机器学习和生物学交叉领域工作的团队。然而,在所有这些先决条件中,安全性预测可能是最重要的,也许也是最少被讨论的,萨普拉说。
“从头设计中最困难的问题之一是预测计算生成的分子在人体内是否安全,”萨普拉解释道。阿斯利康正通过所谓的虚拟临床试验来应对这一挑战。这些是先进的细胞系统和微型器官模型,它们充当物理测试平台,并与从它们的输出中学习的人工智能配对。
“这些系统有潜力在不经历传统测试瓶颈的情况下生成增强的生物信号,并且它们是连接人工智能生成设计与临床候选药物之间闭环的关键缺失环节,”萨普拉补充道。
目前正在进行的一个转变是向能够同时生成候选分子并预测其有效性和安全性的“智能体”AI系统迈进。这些自主工作流程可以将疾病层面的洞察直接与分子设计联系起来,弥合了以前孤立的数据孤岛。“生物学的复杂性与分子的设计是相辅相成的,”萨普拉总结道。
人类人才释放人工智能潜力
生物制剂领域正在进行的变革不仅仅是关于技术。“随着系统更加自主,人类监督仍然是这种方法的核——确保可解释和合乎道德的人工智能惠及患者,”萨普拉说。
对科学家而言,与人工智能合作是一个协作过程。“科学家将与这些模型系统携手合作,”她说。“未来将有一个世界,模型设计分子,然后科学家与系统合作测试这些分子,并将所有数据整合在一起。”通过这种人类检查、制衡和判断的过程,模型将不断进化并持续改进,最终有可能造福患者。
对于工程师而言,设计和构建准备好进行人机协作的有效系统,将意味着确保模型的高度透明度和可解释性。据萨普拉称,阿斯利康的工程团队包括数据科学家、自动化专家和人工智能工程师,他们正在开发充当“思考伙伴”而非“黑箱”的系统。“工程师正在设计能够快速生成、验证和学习的系统。而这些问题确实很棘手:多模态数据融合、闭环优化、不确定性量化以及在临床决策点的可解释性,”她补充道。
萨普拉说,在应对这些技术要求极高的挑战时,工程师和科学家有机会为许多疾病的研究和开发可能改变生命的治疗方法做出贡献。“我们今天能够开发的生物药物,以及我们明天将设计的药物,都取决于将世界级的人工智能和工程人才与深厚的科学专业知识相结合。”
本文由阿斯利康发起并资助。Z4-85058,2026年7月。
本内容由《麻省理工科技评论》定制内容部门 Insights 制作。非《麻省理工科技评论》编辑人员撰写。由人类作者、编辑、分析师和插画师进行研究、设计和撰写。这包括撰写调查问卷和收集调查数据。可能使用的人工智能工具仅限于经过全面人工审查的二级制作流程。
深度探索
人工智能
一家初创公司声称突破了阻碍大语言模型发展的瓶颈
Subquadratic 现已分享了其新模型的更多细节。但一些人仍持怀疑态度。
对人工智能就业恐慌的现实审视
关于人工智能对劳动力市场的影响,数据到底说明了什么?答案可能会让你大吃一惊。
Anthropic 发现了一个隐藏空间,克劳德在其中思考概念
一项新技术让该公司比以往任何时候都更深入地探究了大语言模型的奇妙工作机制。
Claude Science 是 Anthropic 最新的旗舰产品
该公司正加倍押注人工智能在科学领域的应用。
保持联系
获取来自
MIT Technology Review
的最新更新
发现特别优惠、重要报道、即将举行的活动等更多信息。
英文来源:
Sponsored
How AI helps scientists design the next generation of medicines
As generative AI captures public attention, a different kind of AI is reshaping drug discovery. Machine learning models are helping to compress decade-long timelines and cracking problems that were previously unsolvable.
In partnership withAstraZeneca
Designing and developing a new medicine is an expensive, failure-prone scientific challenge. A new drug can take many years to develop, at the cost of a significant investment. And even then, most possible candidates never reach the patient. For biologic medicines, therapies made from engineered proteins rather than synthetic chemistry (which are often used to treat conditions across most major acute and chronic diseases), the complexity is even greater.
Scientists explore vast quantities of possible molecules, looking for the rare few that will bind to the right target, remain stable in the human body, and be manufacturable at scale. Today, AI is speeding up these processes and has quickly become a core part of the infrastructure in pharmaceutical R&D.
AI-assisted design is a growing part of how biologic drug candidates are developed, and companies like AstraZeneca are actively building its engineering teams to push this further. “Everything we do, whether it’s design, make, test, or analyze, is now computationally enhanced,” says Puja Sapra, senior vice president and head of R&D biologics engineering and oncology targeted discovery at AstraZeneca. “The cycle times are getting shorter while productivity and innovation increase.”
Sapra explains that AstraZeneca’s approach follows a build-measure-learn loop. AI generates or prioritizes candidate molecules computationally, predicting which designs are most likely to succeed. Scientists then focus lab resources only on the top-ranked candidates. This leads to a tighter feedback cycle with fewer dead ends, faster iteration, and the ability to go after disease targets that were previously considered untreatable by medicine. Because the number of possible molecular combinations far exceeds what any human team can systematically explore, using AI to narrow and refine the options for testing has become a major focus in biologics drug design.
Navigating complex drug design problems
Beyond accelerating timelines, AI is also being applied to the discovery of entirely new classes of medicines. Traditional biologics typically target one disease pathway. The next generation of drugs can hit multiple targets simultaneously or precisely deliver therapeutic payloads to specific cells. Achieving this requires optimization across many variables at once. Looking ahead AI-driven models could help design these increasingly complex, multi-specific biologics, explains Puja Sapra. “For example,” she continues, “such models could help identify which two or three targets to prioritize based on the underlying biology, then optimize across multiple parameters to balance a molecule’s potency, stability, manufacturability, and safety.” “Drugging the undruggable is becoming a reality,” Sapra says. “These technologies will eventually enable us to develop medicines against targets once thought impossible to reach. The potential for benefit to patients is remarkable.”
The data moat
McKinsey estimates that generative AI, combined with other computational tools, could cut drug discovery timelines by as much as 50%. But every AI model is only as good as its training data. In drug discovery, that means ample quantities of high-quality biological data. Experiments can provide a rich source of such data. Whether they succeed or fail, each experiment generates a signal about what does and does not work.
“Data is our differentiator,” says Sapra, explaining how the company’s datasets are proprietary and multimodal and include molecular structures, binding measurements, safety profiles, and manufacturing outcomes. “We’ve built an intentionally diverse portfolio across multiple disease areas and drug types. All of that data empowers us to fine-tune frontier AI models with richer, more representative training sets.” She continues, “Further, we have invested in deep screening technologies to generate additional datasets required in volume to constantly refine and validate our models.”
Building an autonomous discovery engine
To bring all of that data together in one place, AstraZeneca is building what it calls a “lab of the future” facility in Kendall Square, Cambridge, Massachusetts where AI and robotic automation will be able to form a continuous, closed-loop discovery system. “Where a self-driving car uses sensors and models to navigate its environment, this system uses AI to make predictions, robotic systems to execute experiments, and instruments to generate data,” explains Sapra. That data feeds directly back into the models, accelerating each subsequent cycle.
“Throughout, scientists will remain central to the process, providing the oversight, judgement, and strategic direction that ensure outputs are explainable, tolerable, and directed toward potential patient benefit,” she adds.
Eventually, automated high-throughput systems will be able to make and evaluate thousands of molecular interactions on a weekly basis. “This will generate AI-ready data at a scale that traditional workflows cannot match,” Sapra says. “Robotic sample handling, automated quality checks, and integrated data pipelines also have the potential to help accelerate early drug development timelines significantly.”
The next frontier: Generating medicines from scratch
Ultimately, Sapra says, the end-state vision for AI in biologic drug discovery is what the field calls “de novo” design. For this, the goal is for AI to generate entirely new protein sequences that precisely fit the desired drug properties. This includes designing the structure, predicting safety, how it will behave in the body and how to make it manufacturable.
“The field is making great progress toward a completely AI-generated biologic, designed from scratch all the way to a clinical candidate,” Sapra says. “As we continue to leverage frontier models and fine-tune them with the right datasets, we bring ourselves closer to this reality. I believe it will come. It’s a matter of time.”
Several key elements are needed to reach this point, however. First is richer and more standardized training data across the industry. Second, robust evaluation benchmarks for AI-generated candidates. And third, teams that know how to work at the intersection of machine learning and biology. Of all the prerequisites, however, safety prediction may be the most consequential, and perhaps the least discussed, Sapra says.
“One of the hardest problems in de novo design is predicting whether a computationally generated molecule will be safe in the human body,” Sapra explains. AstraZeneca is tackling this with what amounts to virtual clinical trials. These are advanced cell systems and micro-scale organ models that function as physical testbeds, paired with AI that learns from their outputs.
“These systems have the potential to generate enhanced biological signals without traditional testing bottlenecks, and they're a critical missing piece in closing the loop between AI-generated designs and clinical-ready candidates,” Sapra adds.
A shift currently underway is the move toward agentic AI systems that can simultaneously generate molecule candidates and predict how efficacious and safe they are likely to be. These autonomous workflows can connect disease-level insights directly to molecule design, bridging what were previously separate data silos. “The complexity of the biology goes hand-in-hand with the design of the molecule,” summarizes Sapra.
Human talent unlocks AI potential
The transformation underway in biologics is not just about technology. “With more autonomous systems, human oversight remains at the heart of this approach—ensuring explainable and ethical AI for the benefit of patients,” says Sapra.
For scientists, working with AI is a collaborative process. “Scientists will work hand-in-hand with these model systems,” she says. “There will be a world where models will design molecules, then scientists will work with the systems to test those molecules and put all that data together.” Through this process of human checks, balances, and judgement calls, the models will evolve and constantly improve, ultimately with potential to benefit patients.
For engineers, designing and building effective systems ready for human-AI collaboration will mean ensuring high levels of model transparency and explainability. According to Sapra, AstraZeneca’s engineering teams include data scientists, automation specialists, and AI engineers, who are developing systems that act as “thinking partners” rather than black boxes. “Engineers are designing systems that generate, validate, and learn at speed. And the problems are genuinely hard: Multimodal data fusion, closed-loop optimization, uncertainty quantification, and interpretability at the point of clinical decision-making,” she adds.
In taking on such technically demanding challenges, engineers and scientists have the opportunity to contribute to the research and development of potentially life-changing treatments for many diseases, says Sapra. “The biologic medicines we can develop today, and those we’ll design tomorrow, depend on combining world-class AI and engineering talent with deep scientific expertise.”
This article has been initiated and funded by AstraZeneca. Z4-85058, July 2026.
This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by MIT Technology Review’s editorial staff. It was researched, designed, and written by human writers, editors, analysts, and illustrators. This includes the writing of surveys and collection of data for surveys. AI tools that may have been used were limited to secondary production processes that passed thorough human review.
Deep Dive
Artificial intelligence
A startup claims it broke through a bottleneck that’s holding back LLMs
Subquadratic has now shared more details about its new model. But some are still skeptical.
A reality check on the AI jobs hysteria
What do the numbers really say about the impact of artificial intelligence on the labor market? The answer might surprise you.
Anthropic found a hidden space where Claude puzzles over concepts
A new technique has let the company probe deeper than ever into the weird workings of an LLM.
Claude Science is Anthropic’s newest flagship product
The company is doubling down on AI for science.
Stay connected
Get the latest updates from
MIT Technology Review
Discover special offers, top stories, upcoming events, and more.