为什么传奇的埃尔德什问题正在被人工智能攻克

qimuai 发布于 阅读:15 一手编译

为什么传奇的埃尔德什问题正在被人工智能攻克

内容来源:https://www.quantamagazine.org/why-the-legendary-erdos-problems-are-falling-to-ai-20260803/

内容总结:

传奇数学家埃尔德什的难题,正被人工智能接连攻克

2026年5月20日,OpenAI的一项公告震动了数学界。其内部研发(未公开)的一款AI模型,为匈牙利数学家保罗·埃尔德什于1946年提出的“单位距离”猜想找到了一个反例。这个问题既表述简单又极具数学深度,这标志着AI模型首次在历史上取得意义重大的数学证明。尽管该结果并非定论——人类数学家随后在数周内大幅改进了它,但它极具开创性,引入了来自数学遥远分支、此前无人成功应用的新思想,并迅速催生了解决其他重要问题的相关技术。同年8月1日,OpenAI再度宣布,其未发布的模型“Astra”取得了10项新进展,其中包括解决埃尔德什提出的另外三个问题。

许多数学家将此类进展视为AI数学能力的“相变”时刻。普林斯顿大学教授诺加·阿隆表示,这些模型正“戏剧性地改变数学研究的方式”。

埃尔德什本人及其猜想长期以来令数学家着迷。他居无定所、终身漂泊,将问题发表于论文或写给世界各地同行的信中,并自掏腰包为最先解出问题的人提供奖金(金额从10美元到数千美元不等)。他于1996年在华沙参加数学会议时因心脏病去世,但爱荷华州的一家非营利基金会承诺继续兑现其奖金。他性格古怪且受人爱戴,终生未婚、只穿丝绸衣物、避免身体接触,将大部分收入赠与他人。讽刺的是,他提出的问题如今却成为全球最大科技公司展示AI实力的核心“试验场”和“公关胜利”。

这一切的源头,很大程度上归功于一位名叫托马斯·布鲁姆的英国数学家。他对算术组合学等交叉领域深感兴趣。自2023年初起,布鲁姆在ChatGPT的帮助下搭建了网站“erdosproblems.com”,系统性地收集并整理了埃尔德什提出的近千个问题,并明确了每个问题最合理的表述。他原以为网站可能无人问津,但这个平台逐渐聚集了志同道合的研究者,后来还开放了评论区,促成了包括业余爱好者在内的全球数学社区的协作。

在这个社区里,有尚未完成学业的年轻爱好者凯文·巴雷托和利亚姆·普莱斯,也有著名数学家陶哲轩。他们发现,用特定方式“引导”AI模型(如GPT-5.2),能够解决一些此前悬而未决的埃尔德什问题。尽管过程中不乏乌龙(例如AI“解出”了埃尔德什早已解决的题目),但他们随后找到了确属首次被人类解决的问题,并学会了用AI进行自我验证。这些成果大多出自使用公开模型的业余爱好者或本科生之手。

情况在2026年5月20日发生巨变。OpenAI宣布其内部模型攻克了最著名的“单位距离问题”。数学家们原本普遍相信猜想的正确性,但AI模型却从代数数论领域找到了一个巧妙的反例。多位世界级数学家,包括蒂姆·高尔斯,在随附的评论文章中对这一成果给予高度评价,认为如果这是人类投稿,他们会毫不犹豫地建议顶级期刊接受发表。

此后,受该技术启发,包括布鲁姆在内的四位数学家又推翻了埃尔德什另一个长期未解的“和积猜想”(针对实数集版本)。AI对数学界的影响深远:有的数学家(如阿隆)认为既然AI已能解决此类问题,便失去了继续尝试的动力;而像范·多恩这样的业余爱好者则热情拥抱AI,将其视为强大的思考伙伴,尽管最终他们会将AI的成果简化和推广为更易懂的版本。与此同时,AI能力的提升使部分人产生担忧——大量无法验证结果真伪的“百页论文”开始泛滥。尽管如此,诺加·阿隆仍预测,许多乃至大多数优秀数学家将会使用AI,而顶尖人才向AI公司的流动也已成为一种趋势。曾在7月获颁菲尔兹奖的雅各布·齐默尔曼,同一天宣布将离开学术界加入OpenAI。

中文翻译:

为何传奇的埃尔多什问题正被人工智能攻克

2026年5月20日,OpenAI发布了一项震动数学界的公告。一个内部AI模型——未向公众开放——对“单位距离”问题提出了反例,这是多产、四处漂泊的匈牙利数学家保罗·埃尔多什于1946年提出的猜想。

埃尔多什提出了数千个问题,但这个问题尤为特殊:它既易于阐述,又在数学上深奥难解。这是历史上第一个具有重大意义的、由AI模型产出的证明。尽管该模型的结果并非定论——人类数学家在数周内就对其进行了大幅改进——但它富有创新性,引入了来自数学一个遥远分支的思想,此前从未有人成功将该分支的方法应用于此问题。而且它影响深远:几天之内,相关技术就被用于解决其他重要问题。

随后,8月1日,OpenAI宣布一个名为Astra的未发布模型取得了另外10项数学进展,包括为埃尔多什提出的另外三个问题找到了解答。

许多数学家将此类进展誉为AI模型数学能力的一次相变。普林斯顿大学的诺加·阿隆表示,这些模型“正在彻底改变数学研究的方式”。阿隆在数十年的职业生涯中解决了数十个埃尔多什问题。

埃尔多什及其猜想长期令数学家着迷。他常年旅行——数年如一日地过着行李箱生活,寄宿在朋友家中,几乎一无所有。他在发表的论文和致世界各地数学家的信中源源不断地抛出问题,常常附带奖金,由他自掏腰包奖励第一个给出解答的人。奖金可能只是象征性的10美元或25美元,而对于他认为重要或困难的问题,奖金可能高达数千美元。1996年,埃尔多什在华沙参加数学会议时因心脏病发作去世,但一个总部位于艾奥瓦州的非营利基金会承诺兑现他的悬赏。

他是一个受人爱戴的人物,但也着实古怪。他只穿丝绸衣物,且避免与他人有身体接触。他对权威极度怀疑,将大部分收入捐出,并依靠一位朋友管理财务和其他实务。他称上帝为“最高法西斯”,并以持续服用安非他命来支撑他不间断的数学思想产出。历史的奇特讽刺在于,他所提出的问题如今已成为世界上最大、最强大的科技公司的一个核心试验场——实际上也是一系列公关胜利。

但极有可能的是,若非一位名叫托马斯·布鲁姆的英国数学家,这一切都不会发生。

多次相遇

和埃尔多什一样,布鲁姆对数论和组合学都有兴趣。他的研究重点是算术组合学这一领域,正处于两者的交叉地带。2014年获得博士学位后,布鲁姆成为该领域冉冉升起的新星,获得了英国皇家学会的著名奖学金,使他几乎可以在任何他想去的大学工作。(他目前在曼彻斯特大学。)

布鲁姆从记事起就喜欢埃尔多什的风格。但他一直觉得很难追踪哪些问题已被解决,哪些已被完全遗忘。因此,在2023年初,他决定尽可能多地将问题收集成一个列表。

他最初是为了自己使用。但他说:“我想如果无论身在何处都能访问会更容易”;他觉得“不妨建个网站,不过预期大概没人会用。”他收集了两百多个问题,上线了erdosproblems.com。布鲁姆使用ChatGPT编写了运行该网站的Python代码,这在当时对大型语言模型而言已是了不起的能力。用AI协作做数学本身似乎仍是一个遥远可能。

他的目标不仅仅是勾掉列表中的项目。他在一篇博客文章中写道,他好奇“现代数学——常常使用埃尔多什所不知的技术——能否澄清这些较为冷门问题中的许多。”他说:“到那时,我们将会留下一个有趣而困难的核心问题集,用于展示我们知识的边界。”

布鲁姆在整理列表方面做了关键工作:有时埃尔多什对问题的表述含糊不清,布鲁姆则厘清了每个问题最合理的版本。他不断向网站添加问题,受众也逐渐增长。在2024年全年及2025年前八个月中,列表上有111个问题的状态从“未解决”变为“已解决”(尽管其中一些问题早在多年前就已解决,状态变更反映的是证明被重新发现或验证)。

随后,在2025年8月,一些同事建议布鲁姆添加评论功能,让人们可以讨论他们感兴趣的问题。他很快用ChatGPT编写了代码并实现了这一功能。此时他已收录了近1000个问题。

布鲁姆的时机把握得很好。他让志同道合的人得以互相交流,这“真正催生了一个社区的建立”,他说。大多数时候,评论是零星的——一个问题可能只有一条评论,指出一个例子或说明该问题看起来有多难。但活动在稳步增长,有些问题引发了陌生人之间细致入微的数学讨论。

“汤姆可能从未真正意识到这一点,但对我来说这确实改变了我的生活,”沃特·范·多恩说。他是布鲁姆网站上评论数量第四多的用户。和许多在2025年秋季活跃于该网站的人一样,范·多恩并非严格意义上的职业数学家。用他自己的话说,他在“一家被其他公司聘请做客户服务支持的公司”工作。但他也不完全是业余爱好者——十年前,他差点在比利时鲁汶大学完成数学硕士学位。2024年,部分受到大型语言模型能力显著提升的触动,他休假六个月专注于数学。当时,虽然他并不特别想使用AI,但他记得自己在想:“目前我在数学上仍然比AI强,但谁知道一两年、五年后会怎样?如果我想完成这些项目,而且想以我的名义完成,现在就是时候。”

于是,2025年10月,已回到日常工作的范·多恩在问题1102的页面上留下了第一条评论。这个问题由埃尔多什于1981年提出,询问关于“无平方因子”整数集合的性质——即没有重复质因数的整数。(例如,30是无平方因子数,因为它等于2 × 3 × 5;但18不是,因为它等于2 × 3 × 3,其中3重复了。)

11月初,范·多恩以评论的形式在问题页面上分享了向解答推进的进展——这是他在不依赖AI的情况下想出来的。

当天稍晚,网站上的另一位评论者回复,声称发现范·多恩论证中的一个缺陷。两人快速交替回复,范·多恩说服了对方自己的论证是正确的。“我现在明白你的论证是如何运作的了。漂亮!”那位数学家回复道。那位数学家是陶哲轩,加州大学洛杉矶分校教授,可以说是当今在世最著名的数学家,且毫无疑问是最有影响力的数学家之一。(值得一提的是,陶哲轩年仅10岁时就与埃尔多什有过交集。)

布鲁姆的网站带着一种更早时代的观感,正在成为互联网民主最佳状态的例证。“没有汤姆的网站和那里的评论区,这整次合作是不可能实现的,”范·多恩说。你是否拥有终身教职并不重要,你是年轻还是年长并不重要,你是在名校还是甚至根本不在大学也不重要。只要你想做数学并有好的想法,就能找到合作者。

但随着冬天来临——大约在范·多恩发现自己在与特里·陶合作的同时——事情开始发生变化。

通往十字路口之旅

凯文·巴雷托和利亚姆·普莱斯都二十出头,于2025年夏天在一个专注于AI的Discord服务器上成为朋友。巴雷托目前是剑桥大学的本科生;普莱斯在大学学过一些数学但未完成学业就离开了。12月,两人确信最新的AI模型可能成功解决一些埃尔多什问题,于是开始向它们批量投喂问题。他们很早就意识到,如果告诉GPT-5.2某个问题的答案未知,它就不会有多大进展,所以正如巴雷托所说,他们学会了“以一种非常特殊的方式提示它,让它产生错觉以为问题比实际更容易。”

他们认为自己在埃尔多什问题333上取得了首次胜利。圣诞节清晨,巴雷托在布鲁姆的网站上发布了一个证明,写道:“我们相信,据我们所知,这是首个由大型语言模型完全自主解决此前未被人类解决的埃尔多什问题的案例。”尽管问题333——涉及整数集合的和——并非特别重要的问题,但用AI解决它仍然令人觉得意义重大。

但几小时后,另一位用户指出埃尔多什本人在1977年发表的一篇论文中已给出了问题333的解答。巴雷托承认了错误。“我向网站所有成员正式请求,对目前标记为未解决的问题加强文献检索,”他写道,“作为一个已经两次栽在这个问题上的人,这确实令人痛心疾首。”

他没有气馁,与普莱斯继续努力,到2026年1月4日,他们用GPT-5.2 Pro找到了埃尔多什问题728的解答,该问题涉及某些数何时可被其他数整除。这次没有人能找到先前已证明该结果的文献。巴雷托使用了另一个名为Aristotle的AI工具(由一家名为Harmonic的初创公司开发)来验证证明在逻辑上成立。Nat Sothanaphan——一位软件工程师,也是该论坛上比布鲁姆、陶和范·多恩更活跃的唯一参与者——让ChatGPT将形式化结果写出来并发布到网上。

普莱斯开发了一种方法论,用于如何向大型语言模型提问以解决开放问题。首先,他会让聊天机器人给出一个解答。然后他将该解答输入一个新的聊天机器人实例,要求其检查前一个聊天机器人的成果。他反复执行此过程,直到得到一个看起来可行的解答。(这呼应了各公司内部一直在做的一些工作,即创建所谓的外壳或脚手架,将普莱斯手动进行的那种迭代自动化。)

巴雷托和普莱斯的论文只是过去几个月来许多至少部分由AI解决的埃尔多什问题中的一小部分。这些特定问题成为大型语言模型如此肥沃的试验场,有多重原因。最主要的原因是,埃尔多什问题大体上属于数论、组合学和图论,而这些数学领域已被证明比其他领域对大型语言模型更为友好。这些问题在难度和数学意义上也差异很大。这种差异使它们适合一种能力同样差异很大的新兴技术。

图片由乔治·奇克塞里拍摄,出自纪录片《N是一个数字:保罗·埃尔多什画像》©1993。版权所有。

“我最近很多论文的功劳主要应归于AI,”范·多恩说,“其中的想法并非我自己想出来的。”和许多活跃在埃尔多什网站上的人一样,范·多恩对大型语言模型让他能以更快的速度做更多事情感到兴奋。“如果我读到大型语言模型的一个想法,我会消化它、试着理解它、简化它并推广它,”他说。他用AI来更好地理解数学。

并非所有人都对自己有同样的要求。“一个大问题是,很多并非数学家、没有深厚数学背景、也没有能力验证输出的人大量使用AI,”布鲁姆说,“他们喜欢快速推进,让他们的AI检查结果,它不断膨胀。我们看到越来越多这种100到200页的论文被贴出来:‘我解决了这个定理;我让AI生成证明并检查证明、撰写论文。’但没有人读过它,也不会有人去读。这是一个巨大的挑战。”

按普莱斯自己的评估,他没有足够的数学理解力来验证他最终从大型语言模型中诱导出的解答。但借助巴雷托的帮助,他找到了足够有知识且愿意验证结果的人。普莱斯和巴雷托都是与陶哲轩、斯坦福大学的贾里德·杜克·利希特曼以及其他有成就的数学家共同署名一篇2026年5月论文的作者,该论文解决了埃尔多什问题1196,这是他们较重要的成果之一。(问题1196询问所谓原始集合的可能大小——即整数集合,如{2, 5, 9, 21},其中没有一个数能整除其他任何数。)

布鲁姆感到惊讶的是,尽管OpenAI、谷歌DeepMind和多家初创公司给予了大量关注,但大多数新成果来自使用公开大型语言模型的爱好者和本科生,而非使用更先进内部模型的企业实验室。

但几周后的2026年5月20日,情况发生了变化,OpenAI宣布他们解决了所有埃尔多什问题中最著名的问题之一——单位距离问题。

多次告别

2026年头几个月,各大科技公司开始看到erdosproblems.com上的机会。正如利希特曼所解释的:“埃尔多什有超过1000篇论文。它们分散各处。”匈牙利的一家研究所收集了许多论文的扫描件,但没有人收集过所有问题。“这种任何人都能访问的单一资料库——各实验室意识到这实际上可以成为一个基准。”

1月,由谷歌DeepMind领导的24人研究团队分享了一篇论文,解决了四个问题并为另外九个问题找到了旧有的被遗忘的解答,他们“使用Gemini系统性地评估了布鲁姆埃尔多什问题数据库中标记为‘未解决’的700个猜想”。5月,DeepMind另一个由21名研究人员组成的团队宣布,“我们最有能力的智能体以每个问题数百美元的成本,自主解决了353个开放埃尔多什问题中的9个。”(截至本文发表时,布鲁姆的数据库包含565个已解决问题和652个未解决问题,但DeepMind团队将搜索范围限定在用形式逻辑编写的问题。)

5月20日,OpenAI分享了单位距离问题的解答,同时发布了一篇解释该工作的博客文章和一篇配套论文,其中九位世界级数学家对证明的正确性及其所做工作的重要性发表了评论(并给出了该结果的一个简化人类版本)。数学家们普遍认为埃尔多什的猜想——关于平面上可以放置多少个等距点——是正确的。出乎所有人的意料,OpenAI的内部模型找到了一个反例。为了做到这一点,它找到了一种巧妙的方法,使用了代数数论这一数学领域的工具。多伦多大学的雅各布·齐默尔曼在配套文章中写道:“这确实是一项令人印象深刻的工作……这绝对是一个令人望而生畏的构造。”

在同一篇文章中,剑桥大学和法兰西公学院的蒂姆·高尔斯写道:“如果是一个人写了这篇论文并投稿到《数学年鉴》,而有人征求我的快速意见,我会毫不犹豫地建议接受。此前没有任何AI生成的证明接近过这个水准。”

那篇提出原始解答的论文上,作者署名仅为“OpenAI”。

比利·格雷斯·陶

后来,使用与AI模型应用于单位距离问题的技术相关的技术,包括布鲁姆在内的四位数学家否证了另一个长期存在的埃尔多什猜想的一个版本。“和积猜想”提出,如果你有若干数字集合,它们的和或积中必有一个快速增长。这些数学家找到了一个实数集合,其和与积都比预期增长得更慢。对整数的该猜想仍然开放。

弄清AI将对数学和数学家产生怎样的影响,不仅要看其最重要的成果,还要审视它如何改变解决日常问题的日常实践。普林斯顿大学的数学家诺加·阿隆估计,他在职业生涯中解决了大约几十个埃尔多什问题。他现在已不再尝试。“一旦AI开始解决它们,就没有意义了,”他说。陶哲轩已退出埃尔多什问题社区,专注于完成自己的工作。

范·多恩目前仍在客户服务公司做他的日常工作,他说大型语言模型“在思考数学和做数学方面显然比我强。我无法与当前的AI系统相提并论。”然而他补充道:“我最终写出的证明比ChatGPT生成的要更简洁、更普遍、对其他人更易读。”AI无疑提高了他的产出效率,而且他仍然乐在其中。“如果你想弹钢琴,你不会雇一台弹得比你好的钢琴机器。你会弹钢琴,是因为你喜欢弹钢琴。我喜欢思考数字、做数学、写论文。我不会雇一台论文制造机器替我做。”

对范·多恩来说,消化他从大型语言模型得到的回复中也有乐趣。“由于埃尔多什问题社区,我最近做了很多数学。以前我总是独自做一切,独自在房间里挣扎。我不知道是怎么发生的,但如今人们联系我说:‘我有这个想法。你想和我一起思考吗?’”

AI不断增强的能力使普莱斯或范·多恩这样数学训练较少的人更容易解决谜题,同时也使这些谜题对阿隆这样毕生致力于理解它们的人变得不再那么有趣。

尽管如此,“许多——也许大多数——优秀的数学家将使用AI,”阿隆说。他指出,一些一流数学家已离开学术界去AI公司工作,不仅因为报酬优厚,而且因为“也许现在的前沿就在那里。”2026年7月,在齐默尔曼获得数学界最高荣誉菲尔兹奖的同一天,他宣布离开学术界,加入OpenAI工作。

英文来源:

Why the Legendary Erdős Problems Are Falling to AI
On May 20, 2026, OpenAI made an announcement that shook the mathematical world. An internal AI model — one not available to the public — had come up with a counterexample to the “unit distance” problem, a conjecture made in 1946 by Paul Erdős, the prolific, itinerant Hungarian mathematician.
Erdős posed thousands of questions, but this one was special: It was both simple to explain and mathematically deep. It was the first historically significant proof to come from an AI model. Though the model’s result wasn’t definitive — human mathematicians would substantially improve on it within weeks — it was innovative, bringing in ideas from a distant branch of math that no one had successfully applied to this problem before. And it was influential: Within a few days, related techniques were used to solve other important problems.
Then on August 1, OpenAI announced that an unreleased model named Astra made 10 additional mathematical advances, including finding solutions to three more problems posed by Erdős.
Many mathematicians have hailed developments such as these as a phase transition in the mathematical capability of AI models. These models are “changing dramatically the way mathematical research is being done,” said Noga Alon of Princeton University, who has solved dozens of Erdős problems over his decades-long career.
Erdős and his conjectures have long fascinated mathematicians. He traveled constantly — living out of a suitcase for years at a time, staying with friends, owning almost nothing. He rattled off problems in published papers and letters to mathematicians around the world, often attaching prize money that he would pay out of pocket to the first person to come up with a solution. The reward might be a token $10 or $25, or, for problems he considered important or difficult, it could range into the thousands. Erdős died of a heart attack in 1996 while attending a math conference in Warsaw, but a nonprofit foundation based in Iowa has promised to make good on his bounties.
He was a beloved figure, but also a downright weird one. He only wore silk, and he avoided the touch of other people. Deeply cynical about authority, he gave away most of the money he earned and relied on a friend to manage his finances and other practical affairs. He referred to God as the “Supreme Fascist” and fueled his incessant output of mathematical ideas with a steady diet of amphetamines. It is a strange irony of history that the problems he suggested have now become a central proving ground — and, in effect, a series of PR coups — for the world’s biggest and most powerful technology companies.
But in all likelihood none of this would have happened had it not been for an English mathematician named Thomas Bloom.
Many Meetings
Like Erdős, Bloom was interested in both number theory and combinatorics. His focus has been an area called arithmetic combinatorics, which lies at the intersection of the two. After getting his doctorate in 2014, Bloom established himself as a rising star in the field, landing a prestigious fellowship from Britain’s Royal Society, which let him work at almost any university he wanted to. (He’s now at the University of Manchester.)
Bloom has liked Erdős’ style for as long as he can remember. But he always found it hard to keep track of which problems had been solved and which had been forgotten entirely. So in early 2023, he decided to gather as many problems as he could into a list.
He intended it for his own use. But “I thought it would be easier if I could access it wherever I was,” he said; he figured he “might as well make a website, kind of with the expectation that maybe nobody would use it.” He gathered a couple hundred problems and launched erdosproblems.com. Bloom used ChatGPT to write the Python code that ran the website, which was, at the time, a remarkable thing for a large language model to be able to do. Using one to collaborate on the math itself still seemed like only a distant possibility.
His goal was not just to cross items off a list. He wondered if “modern day mathematics, often using techniques unknown by Erdős, could clear up many of these more obscure problems,” he wrote in a blog post. “We will then be left with a core of interesting, difficult problems, which can serve to demonstrate the limits of our knowledge.”
Bloom did crucial work in curating the list: Sometimes Erdős stated problems in ambiguous or unclear ways, and Bloom figured out what the most sensible version of each problem should be. He kept adding problems to the site, and gradually its audience grew. Over the course of 2024 and the first eight months of 2025, the statuses of 111 problems on the list were changed from “open” to “solved” (although some of these had been solved years earlier, and their status change reflected the rediscovery or verification of a proof).
Then, in August 2025, some colleagues suggested that Bloom add a commenting function, so that people could talk about problems they were interested in. He was able to do so quickly, using ChatGPT to write the code. By now he’d cataloged nearly 1,000 problems.
Bloom’s timing was good. He made it possible for like-minded people to talk to one another, and that “really let a community build up,” he said. For the most part, comments were sporadic — a problem might attract a single comment pointing out an example or noting how hard the problem looked. But activity steadily grew, and some problems catalyzed nuanced mathematical discussions between strangers.
“Tom probably never really realized this, but for me it’s honestly changed my life,” said Wouter van Doorn, the fourth-most-prolific commenter on Bloom’s website. Like many people who became active on the site in the autumn of 2025, van Doorn isn’t exactly a professional mathematician. He works “for a company that gets hired by other companies to do customer service support,” as he put it. But he isn’t exactly an amateur either — a decade prior, he almost completed a master’s degree in math at KU Leuven in Belgium. In 2024, spurred in part by how capable he saw LLMs getting, he took a six-month leave of absence from work to focus on math. At the time, while he didn’t particularly want to use AI, he remembers thinking, “Right now I’m still better at mathematics than an AI is, but who knows what it’ll be in a year, two years, five years? If I want to finish these projects, and I want them to be mine, now is the time.”
And so, in October 2025, van Doorn, now back at his day job, left the first comment on the page for Problem 1102. The problem, which Erdős posed in 1981, asks about properties of sets of “square-free” integers — that is, integers that have no repeated prime factors. (For instance, 30 is square-free because it is equal to 2 × 3 × 5, but 18 is not, because it is equal to 2 × 3 × 3; the 3 repeats.)
In early November, van Doorn shared progress toward an answer — which he’d figured out without relying on AI — as a comment on the problem page.
Later that day, another commenter on the site replied, claiming he had found a flaw in van Doorn’s argument. The two traded remarks in rapid succession, and van Doorn convinced his interlocutor that his argument was correct. “I see how your argument works now. Nice!” the other mathematician replied. That other mathematician was Terence Tao, a professor at the University of California, Los Angeles who is arguably the best-known mathematician alive today, and inarguably one of the most influential. (Not incidentally, when Tao was just 10 years old, he crossed paths with Erdős.)
Bloom’s website, which has the look and feel of an earlier time, was becoming an example of the internet at its democratic best. “This entire collaboration would not have been possible without Tom’s website and the comments section there,” van Doorn said. It didn’t matter if you had tenure or not, if you were young or old, if you were at a fancy university or even at a university at all. If you wanted to work on math and had good ideas, you could find people to collaborate with.
But as the winter set in — around the same time that van Doorn found himself collaborating with Terry Tao — things started to change.
Journey to the Cross-Roads
Kevin Barreto and Liam Price, both in their early 20s, became friends in the summer of 2025 on a Discord server dedicated to AI. Barreto is currently an undergraduate at the University of Cambridge; Price studied some math in college but left before finishing. In December, convinced that the newest AI models might succeed in resolving some Erdős problems, the pair started throwing batches of problems at them. They realized early on that if they told GPT-5.2 that a problem’s answer wasn’t known, it wouldn’t make much headway, so as Barreto put it, they learned how to “prompt it in a very particular way, gaslighting it into thinking the problem is easier than it actually is.”
They had what they thought was their first triumph on Erdős Problem 333. Early on Christmas morning, Barreto posted a proof to Bloom’s website, writing, “We believe, to the best of our knowledge, this is the first case of an LLM fully autonomously resolving an Erdős problem, not previously resolved by humans.” Even though 333, which dealt with the sums of sets of integers, was not a particularly important problem, solving it with AI still felt important.
But a few hours later, another user pointed out that Erdős himself had provided a resolution to 333 in a paper published in 1977. Barreto owned up to the mistake. “My formal request to all members of the website is to put greater focus on literature search on the problems currently marked as open,” he wrote. “As someone who has fallen for this twice now, it’s quite gut-wrenching.”
Undeterred, he and Price kept at it, and by January 4, 2026, they’d used GPT-5.2 Pro to find a solution to Erdős 728, a problem about when certain numbers are divisible by other numbers. This time nobody could find prior work already proving it. Barreto used another AI tool called Aristotle (developed by a startup called Harmonic) to certify that the proof held together logically. Nat Sothanaphan, a software engineer and the only forum participant more prolific than Bloom, Tao, and van Doorn, had ChatGPT write up the formalized result and posted it online.
Price developed a methodology for how to ask LLMs to solve open questions. First, he would ask a chatbot for a solution. Then he would feed that solution into a fresh instance of the chatbot, asking it to check the previous chatbot’s work. He’d repeat this process until he had what looked like a workable solution. (This echoes some of the work that companies have been doing internally to create what they call harnesses or scaffolds, which automate the sort of iteration that Price does by hand.)
Barreto and Price’s papers represent just a fraction of the many Erdős problems solved at least in part by AI over the past few months. There are multiple reasons why these problems in particular have become such a fertile test bed for LLMs. The primary one is that, by and large, Erdős problems are in number theory, combinatorics, and graph theory, all areas of math that have proved more accessible than others to large language models. The problems also vary widely in difficulty and mathematical significance. This variation makes them appropriate for a nascent technology whose abilities also vary widely.
Photo by George Csicsery from the documentary N is a Number: A Portrait of Paul Erdős ©1993. All Rights Reserved.
“A lot of my recent papers should be mostly credited to AI,” van Doorn said. “The ideas involved were ideas I did not come up with myself.” Like many people active on the Erdős site, van Doorn is excited about the way LLMs are allowing him to do more things more quickly. “If I read an idea by an LLM, I digest it, try to understand it, simplify it, and generalize it,” he said. He uses AI to better understand the math.
Not everyone holds themselves to this standard. “A big problem is AI is being used a lot by people who aren’t mathematicians, who don’t have a huge mathematical background and are not capable of verifying the output,” Bloom said. “They like to move fast, ask their AI to check it, it grows and grows. We’re seeing a lot more of these 100- to 200-page papers that people are posting. ‘I solved this theorem; I got AI to generate the proof and check the proof and write the paper.’ But no human has read it, and no human is going to read it. It’s a huge challenge now.”
By Price’s own assessment, he doesn’t have enough mathematical understanding to verify the solutions he ultimately coaxes from the LLMs. But with Barreto’s help, he’s been able to find mathematicians knowledgeable and willing enough to check the results. Both Price and Barreto are co-authors with Tao, Jared Duker Lichtman of Stanford University, and other accomplished mathematicians on a May 2026 paper resolving Erdős Problem 1196, one of their more significant results. (1196 asks about the possible size of so-called primitive sets — collections of integers, such as {2, 5, 9, 21}, in which no number divides any other.)
Bloom was surprised that despite lots of attention from OpenAI, Google DeepMind, and several startups, most of the new results had come from hobbyists and undergraduates using publicly available LLMs, not from corporate labs using more advanced internal models.
But that would change a few weeks later, on May 20, 2026, when OpenAI announced that they had solved one of the most well known Erdős problems of all, the unit distance problem.
Many Partings
In the first months of 2026, the major tech companies began to see opportunity in erdosproblems.com. As Lichtman explained, “Erdős had over 1,000 papers. They were scattered.” An institute in Hungary had collected scanned images of many of the papers, but nobody had collected all the problems. “This kind of single repository that anyone can access — labs realized that this could effectively be a benchmark.”
In January, a team of 24 researchers led by Google DeepMind shared a paper solving four problems and finding old, forgotten solutions to nine more, after “using Gemini to systematically evaluate 700 conjectures labeled ‘Open’ in Bloom’s Erdős Problems database.” In May, a separate DeepMind team of 21 researchers announced that “our most capable agent autonomously resolved 9 of 353 open Erdős problems at the per-problem cost of a few hundred dollars.” (As of this article’s publication, Bloom’s database contains 565 solved problems and 652 open ones, but the DeepMind team narrowed their search to problems that have been written in formal logic.)
And, on May 20, OpenAI shared a solution to the unit distance problem, along with a blog post explaining the work and a companion paper that featured nine world-class mathematicians commenting on the correctness of the proof and the importance of what had been done (as well as presenting a streamlined human version of the result). Mathematicians had generally believed that Erdős’ conjecture — about how many evenly spaced points can be placed on a plane — was correct. To general surprise, OpenAI’s internal model found a counterexample. To do so, it had found a sophisticated way to use tools from an area of math called algebraic number theory. As Jacob Tsimerman of the University of Toronto wrote in the companion article, “This is a really impressive piece of work. … It is definitely an intimidating construction.”
In the same article, Tim Gowers of Cambridge and the Collège de France wrote that “if a human had written the paper and submitted it to the Annals of Mathematics and I had been asked for a quick opinion, I would have recommended acceptance without any hesitation. No previous AI-generated proof has come close to that.”
The author on the paper that presented the original solution was given simply as “OpenAI.”
Billy Grace Tao
Later, using techniques related to the ones the AI model had applied to the unit distance problem, a group of four mathematicians, including Bloom, disproved a version of another long-standing Erdős conjecture. The “sum-product” conjecture proposed that if you have sets of numbers, either their sum or their product must grow quickly. The mathematicians found a set of real numbers for which both the sum and the product grow more slowly than expected. The conjecture for integers remains open.
Figuring out what impact AI will have on math and mathematicians means not only looking to its most important results, but also examining how it changes the everyday practice of solving quotidian problems. Noga Alon, the Princeton mathematician, estimates that he has solved a few dozen Erdős problems over his career. He has now stopped trying. “Once AI started to solve them, there is no point anymore,” he said. Terry Tao has stepped away from the Erdős problem community to focus on getting work done.
Van Doorn, who for now still has his day job at a customer service company, said that LLMs “are clearly better at thinking and doing math than I am. I don’t hold a candle to current AI systems.” However, he added, “the eventual proofs that I write are simpler, more general, and easier to read for other people than the thing that ChatGPT came up with.” AI has indisputably boosted his productivity, and he’s still having fun. “If you want to play piano, you aren’t going to hire a piano-playing machine that does it better than you. You will play the piano because you like playing the piano. I enjoy thinking about numbers, doing math, writing papers. I’m not going to hire a paper-making machine that does it for me.”
For van Doorn, there is joy to be found in digesting the responses he gets from LLMs. “I’ve been doing a lot of math recently thanks to the Erdős-problems community. It used to be the case I just did everything all by myself, struggled alone in my room. I don’t know how it happened, but nowadays people contact me saying, ‘I have this idea. Do you want to join me thinking about this?’”
The increasing capability of AI has made it easier for people like Price or van Doorn, with less mathematical training, to solve puzzles, while making those puzzles less interesting to people like Alon who have devoted a lifetime to understanding them.
Nonetheless, “many and maybe most good mathematicians will use AI,” Alon said. He notes that a number of first-rate mathematicians have left academia to work at AI companies, not only because they are well paid to do so but because “maybe this is where the action now is.” In July 2026, on the same day that Tsimerman was awarded the Fields Medal, the highest honor in math, he announced that he was leaving academia for a job at OpenAI.

quanta

文章目录


    扫描二维码,在手机上阅读