人类刷新了一项数学纪录,随后AI便来挑战它

内容来源:https://www.sciencenews.org/article/human-math-record-twin-primes-openai
内容总结:
数学界近日上演了一场人类与人工智能的激烈竞速。8月的最后一天,伊利诺伊大学厄巴纳-香槟分校的数学家朱莉娅·斯塔德尔曼在孪生素数猜想研究上取得十余年来首个突破,将“小素数间隙上界”的世界纪录从246降至240。然而这一纪录仅保持三天,便被AI初创公司Axiom Math以212超越;两小时后,OpenAI宣布将纪录进一步推至186。五天后,OpenAI又声称解决了另一数学难题。
斯塔德尔曼的导师形容这是“大卫战胜歌利亚、人类摘得金牌”的故事,但金牌在手仅三天。这一连串事件在数学界引发强烈不安。菲尔兹奖得主、加州大学洛杉矶分校教授陶哲轩批评部分AI公司“放弃了对人类理解的追求,仅将AI工具用于刷榜”。多位数学家指出,OpenAI在斯塔德尔曼已投入该问题的情况下未直接告知她,违背了学术界避让或合作的不成文规范。OpenAI回应称已与斯塔德尔曼的博士导师梅纳德取得联系,但未直接联系她本人。
数学家们担忧的不仅是优先权归属。他们强调,这类难题的真正价值不在于最终数字,而在于求解过程中发展的技术、新的思维方式以及对数学整体的更深理解——这些无法被一个数字所传达。福特教授担心,当AI可以轻易抢先给出答案,有才华的年轻人将不愿进入该领域,“这不仅对数学,对整个科学都将是巨大损害”。目前数学界仍在消化OpenAI的论文及其策略。斯塔德尔曼本人表示,她更关心其中的数学思想,而非具体数字。
中文翻译:
8月的最后一天,一位人类数学家在一项数论中最著名的问题上掀起了一连串进展——几乎立刻就被AI超越了。
伊利诺伊大学厄巴纳-香槟分校的朱莉娅·斯塔德尔曼在攻克孪生素数猜想的征程中报告了一项新纪录。该猜想认为,存在无穷多对相差仅为2的素数(例如3和5是孪生素数,17和19也是)。这是十多年来该领域的首次此类进展。
斯坦福大学数学家坎南·桑达拉拉詹称这是一个“极其困难的问题”,并表示他钦佩斯塔德尔曼的毅力和勇气。“这确实非常令人印象深刻。”
不到三天,一家名为Axiom Math的AI初创公司便在斯塔德尔曼的方法基础上实现了超越。Axiom发布消息后两小时内,两项纪录又被OpenAI刷新。仅仅五天后,OpenAI宣称解决了数学领域另一个重大难题。
“这是一个大卫对战歌利亚的故事,而这次人类拿到了金牌,”斯塔德尔曼在伊利诺伊大学的博士后导师凯文·福特说。至少,维持了三天。
这一连串迅速发生的事件让数学家们对AI的焦虑急剧升温。在许多研究者看来,一些AI公司不仅颠覆了职业规范,还破坏了研究者对知识的追求。“我们在这里看到的,是一些人刻意选择放弃任何获取人类理解的伪装……并将AI工具仅用于达成基准测试目的所带来的后果,”加州大学洛杉矶分校的陶哲轩在mathstodon.xyz上写道。他是菲尔兹奖得主,也可能是数学界最有影响力的博主。
追逐素数对
素数——即只能被1和自身整除的数——对数学家来说似乎有着无穷无尽的魅力,他们长期以来一直困惑于素数在数轴上的分布规律及其原因。孪生素数猜想至少从19世纪起就一直是数学家们努力攻克的难题,人们普遍相信它是正确的,但数论学家至今未能证明。该问题的一个较简单的版本不问是否存在无穷多对间隔为2的素数,而是问是否存在无穷多对间隔为任意相同大小的素数。
2013年,现就职于中国广州中山大学的数学家张益唐证明了答案是肯定的:某些间隔必然无限次重复出现。数学家们不知道间隔为2是否如此,但张益唐证明了某个小于7000万的间隔确实如此,从而创下了“小素数间隔上界”的首个纪录。
数学家们迅速轮番降低这个数字。他们能把它压得越低,就越接近解决孪生素数猜想。到2014年中期,大规模合作努力已将上界降至246,意味着间隔大小等于或小于246的素数对必然无限重复。这个纪录保持了十二年。
斯塔德尔曼大约两年前开始集中研究这个问题,当时她刚在英国牛津大学詹姆斯·梅纳德指导下完成博士学位。(梅纳德曾帮助创下此前的纪录,并于2022年获得菲尔兹奖,部分原因正是他在素数间隔方面的工作。)斯塔德尔曼承担了一项艰巨的任务:试图将较新的技术与张益唐早期的成果更好地整合起来。
两种方法都使用了“筛法”,即用不同的权重部分过滤掉合数,例如对2的倍数比对3的倍数进行更强的过滤。悖论式地是,这种部分过滤的思路反而比完全过滤能更好地揭示素数的分布。斯塔德尔曼通过拼接高维空间中不同方法占优的区域来确定最优权重。她必须在仅对这些区域有粗略了解的情况下计算它们的体积。
“这比大海捞针还要难,”蒙特利尔大学数论学家安德鲁·格兰维尔说。
到2026年初夏,斯塔德尔曼成功将世界纪录从246降到了240。数论学家们一致认为,假以时日,她的方法本可以做得更好。不幸的是,她没有时间。
8月中旬,斯塔德尔曼开始听到传言说OpenAI计划宣布一项关于素数间隔的重大新成果。福特告诉她:“放下你手头的一切,把这个结果发到数学预印本库上。赶紧发出去。即使第二天就被超越,你至少也拥有过一天的纪录。”
结果,这个纪录维持了三天。这可能是人类最后一次保持该纪录。
斯塔德尔曼将结果发布到网上后,Axiom Math的研究人员——他们一直在核查2013年和2014年发表的证明——转移了精力,将她的新思路整合进一个定理证明AI中。在接下来的三天里,Axiom的研究人员通宵达旦,将上界降到了212,这大概是对斯塔德尔曼若有更多时间和计算资源可能取得的成果的一个合理估计。
与此同时,OpenAI一直在研究一个密切相关的问题——不是求素数之间的最小间隔,而是最大间隔。团队一直将其作为基准测试,用来展示其最新大语言模型GPT-6 Astra的解题能力。据OpenAI计算机科学家塞巴斯蒂安·布贝克称,团队也在探索小间隔问题,这算是附带项目,因为公司内部的数学家对此感兴趣。OpenAI在9月3日随模型发布一起发表的论文中报告,Astra将纪录降到了186。
布贝克表示,OpenAI不打算在素数间隔问题上继续推进。“我们的目标不是先发制人地抢占所有可能的成果,”布贝克说。“我们的策略是赋能数学家。”
但数学界的争议随之迅速升温。
为理解而数学
没有人质疑AI可以成为一个极其有用的数学工具。“上周,我在两小时内完成了一个以前要花一个月才能做出的证明,”格兰维尔说。AI可以比人类更快地梳理大量文献,更快地识别死胡同,并尝试更多想法组合。
但AI也带来了许多问题:谁将有机会使用这些工具?成本会有多高?如何最好地理解机器生成的证明——其中可能充斥着人类无法立即理解的“AI垃圾”?如果计算机能突然介入并使某人的工作立即过时,这个领域的文化将如何改变?
尽管斯塔德尔曼说她并不觉得受到了不公正对待,但许多数学家对OpenAI在未告知斯塔德尔曼的情况下研究同一问题感到不满。在数学界,通常如果其他人——尤其是资历较浅的研究者——已经在研究某个问题,学者们会退让。或者,他们可能会主动联系合作。“我看不出OpenAI认真考虑过他们是如何对待她的,”格兰维尔说。
OpenAI表示他们与梅纳德有过联系,但没有直接联系斯塔德尔曼。
数学家们目前正在消化OpenAI的论文以及使新纪录成为可能的策略。“我对具体数字不太感兴趣,但我对数学思想非常感兴趣,”斯塔德尔曼说。
正是这种对思想的追求,让数学家们担心当AI被放开来攻克未解难题时会消失。陶哲轩——他曾参与此前的小素数间隔研究——指出,对于这类问题,价值不在于答案本身,而在于探索过程中产生的东西。在这个过程中,数学家们发展出技术、新的思考问题的方式,以及对数学整体的更深理解,这些都不是一个像186这样的数字所能传达的。
“我热爱解决问题的过程,”福特说。“用头撞墙、感到沮丧,然后突然间看到那个奏效的想法。”他担心随着AI夺走这些,有天赋的年轻人会决定不进入这个领域。未来的研究者可能会被劝退,不再去追求那些机器能比他们更快找到答案的问题。“那将是极其有害的,”他说,“不仅对数学,而且对整个科学都是如此。”
英文来源:
On the last day of August, a human mathematician kicked off a flurry of advances on one of the most famous problems in number theory — and was almost instantly superseded by AI.
Julia Stadlmann of the University of Illinois Urbana–Champaign reported a new record in the quest to solve the twin prime conjecture, which posits that there are infinitely many pairs of prime numbers that are just two numbers apart (3 and 5 are twin primes, for example, as are 17 and 19). It was the first such advance in more than a decade.
It’s a “fiendishly difficult problem,” says mathematician Kannan Soundararajan of Stanford University, who says he admires Stadlmann’s persistence and courage. “It was really quite impressive.”
Within three days, an AI startup called Axiom Math had built on Stadlmann’s approaches to best her effort. Within two hours of Axiom’s announcement, both records were surpassed by OpenAI, which just five days later claimed to have solved another of math’s biggest puzzles.
“It’s a David-versus-Goliath story where the human gets the gold,” says Kevin Ford, Stadlmann’s postdoctoral mentor at Illinois. For three days, at least.
The rapid-fire series of events has sent mathematicians’ angst over AI skyrocketing. In the view of many researchers, some AI companies are not only subverting professional norms but also undermining researchers’ pursuit of knowledge. “What we are seeing here is a consequence of some very deliberate choices to abandon any pretense of gaining human understanding … and using AI tools for the sole purpose of achieving a benchmark,” Terence Tao of UCLA, a Fields Medalist and probably the most influential blogger in mathematics, wrote on mathstodon.xyz.
Chasing pairs
Prime numbers, those that can be divided only by one and themselves, are of seemingly endless fascination to mathematicians, who have long puzzled over where along the number line they appear and why. The twin primes conjecture, which has been wrestled with since at least the 19th century, is widely believed to be true but still unproven by number theorists. An easier version of the problem asks not whether there are infinite pairs of primes with a gap of two but whether there are infinite pairs of primes that all have the same gap of any size.
In 2013, mathematician Yitang Zhang, now at Sun Yat-sen University in Guangzhou, China, proved that answer is yes: Some gaps must repeat infinitely many times. Mathematicians don’t know that a gap of two does, but Zhang showed that a gap of some number less than 70 million does, thus setting the first record for the “small prime-gap bound.”
Mathematicians quickly took turns lowering that number. The lower they can push it, the closer they are to solving the twin prime conjecture. By the middle of 2014, massive collaborative efforts had brought the bound down to 246, meaning pairs with some gap size equal to or less than 246 must repeat infinitely. That is where the record stood for a dozen years.
Stadlmann began working intensively on the problem about two years ago, after finishing her doctorate under James Maynard at the University of Oxford in England. (Maynard had helped set the previous record, and won the Fields Medal in 2022, in part for his work on prime gaps.) Stadlmann took on the formidable task of trying to better integrate the more recent techniques with Zhang’s earlier efforts.
Both approaches use a “sieve” method, which partially filters out composite numbers using various weights, for example filtering out multiples of 2 more strongly than multiples of 3. Paradoxically, this idea of partial filtering offers a better idea of the distribution of primes than a complete filter would. Stadlmann determined the optimal weights by piecing together regions of high-dimensional space where the different approaches prevail. She had to calculate the volumes of those regions with only a rough idea of what the regions look like.
“It goes beyond finding a needle in a haystack,” says Andrew Granville, a number theorist at the University of Montreal.
By early summer 2026, Stadlmann had managed to move the world record down from 246 to 240. With time, number theorists agree, her methods could have done better. Unfortunately, she did not have time.
In mid-August, Stadlmann began hearing rumors that OpenAI was planning to announce a big new result on prime gaps. Ford told her: “Drop everything you’re doing and put this result on the math preprint archive. Get it out there. Even if it’s improved the next day, you’ve got the record at least for a day.”
As it turned out, the record lasted three days. It may be the last time the record will ever be held by a human.
After Stadlmann posted her result online, researchers at Axiom Math, who had been checking the proofs published in 2013 and 2014, shifted their energies and incorporated her new ideas into a theorem-proving AI. Over the next three days, pulling all-nighters, researchers at Axiom got the bound down to 212, a good estimate of what Stadlmann might have achieved with more time and computing resources.
Meanwhile, OpenAI had been working on a closely related problem that asks for the largest gaps between primes instead of the smallest ones. The team had been using it as a benchmark to showcase the problem-solving abilities of its newest large language model, GPT-6 Astra. According to OpenAI computer scientist Sébastien Bubeck, the team was also exploring the small gaps problem, a sort of add-on because mathematicians at the company were interested in it. Astra got the record down to 186, OpenAI reported in a paper that accompanied the model’s release on September 3.
Bubeck says that OpenAI does not intend to push their work on the prime-gaps problem further. “Our goal is not to preemptively strike and capture all the results possible,” says Bubeck. “Our strategy is to empower the mathematician.”
But controversy in the math community swiftly ensued.
Math for understanding
No one disputes the idea that AI can be a stunningly useful mathematical tool. “Last week, I did a proof in two hours that would have taken me a month before,” Granville says. AI can sift through the vast literature faster than a human, identify dead ends more quickly, and try many more combinations of ideas.
But AI also brings many questions: Who will have access to these tools? How much will they cost? What’s the best way to make sense of machine-generated proofs, which can be filled with “AI slop” that humans can’t immediately understand? How will the culture of the field change if a computer can swoop in and make someone’s work immediately obsolete?
Though Stadlmann says she doesn’t feel mistreated, many mathematicians were upset that OpenAI was working on the same problem as Stadlmann without informing her. Typically in mathematics, academics will stand back if others, particularly less senior researchers, are already at work on a problem. Alternatively, they might reach out to collaborate. “I do not see that OpenAI carefully considered how they treated her,” Granville says.
OpenAI says that they were in touch with Maynard, but that they did not contact Stadlmann directly.
Mathematicians are now in the process of digesting OpenAI’s paper and the strategy that made the new record possible. “I am not so much interested in the exact number, but I am very interested in the mathematical ideas,” Stadlmann says.
It’s this pursuit of ideas that mathematicians worry will be lost when AI is let loose on unsolved puzzles. For problems like these, notes Tao, who collaborated on previous small prime-gap efforts, the value is not in the answer, but in what comes out of the struggle. Along the way, mathematicians develop techniques, new ways of thinking about problems and a deeper understanding of mathematics as a whole, which cannot be conveyed by a single number like 186.
“I love the process of problem solving,” Ford says. “Banging our heads against the wall and getting frustrated, and all of a sudden seeing the idea that works.” He worries that as AI takes that away, talented young people will decide not to enter the field. Future researchers may be discouraged from pursuing problems for which a machine can beat them to the answer. “That would be massively detrimental,” he says, “not just to math but to science in general.”