OpenAI的最新争议揭示了数学的未来走向

qimuai 发布于 阅读:42 一手编译

OpenAI的最新争议揭示了数学的未来走向

内容来源:https://www.technologyreview.com/2026/09/08/1143747/what-openais-latest-controversy-tells-us-about-the-future-of-math/

内容总结:

OpenAI被曝“抢跑”解决千禧年数学难题,引发学术界争议

OpenAI近日宣布其AI智能体成功解决了七个“千禧年大奖难题”之一的纳维-斯托克斯方程存在性与光滑性问题,这原本应是其重大科研突破,但随即陷入剽窃争议,被指未恰当引用纽约大学数学家Tristan Buckmaster和Anthropic员工Levent Alpöge的近期研究成果。

据称,Buckmaster和Alpöge在过去近一年里利用包括OpenAI和Anthropic在内的公开模型对该问题进行了研究,并取得了关键进展。OpenAI承认是在听闻两人的工作后才决定攻克此题,但否认其模型直接使用了二人的成果。Buckmaster声称,OpenAI曾提出两个选项:要么他们先发布成果、OpenAI次日再发布,要么合作撰写论文但排除因就职于竞争对手Anthropic而无法署名的Alpöge。OpenAI否认了相关指控,并称不打算领取百万美元奖金。

这场争议背后折射出数学研究正经历深刻变革。如今,最前沿的数学难题的突破越来越依赖少数人工智能公司的巨额资源和内部模型。OpenAI此次以约1万个并发智能体、耗资数百万美元在短时间内“强推”出结果,远非普通学者所能企及。有专家指出,若AI在缺乏透明过程的情况下“过早”解题,可能破坏数学研究通过试错和探索推动领域发展的传统路径,人类数学家或将面临被边缘化的风险。尽管研究品味和人类智慧仍在其中扮演角色,但数学的未来正日益被少数科技巨头所主导。

中文翻译:

OpenAI最新争议揭示了数学的未来

AI公司在数学领域取得了令人瞩目的进展,而人类数学家正在失去优势。

OpenAI最新的数学里程碑迅速陷入争议。今天,该公司宣布其智能体解决了千禧年大奖难题之一——这是数学中一些最重要的未解决问题。在正常情况下,这一解决方案将成为OpenAI的重大成就。

但这一公告却被指控所笼罩:OpenAI被指利用了纽约大学数学家特里斯坦·巴克马斯特和Anthropic员工莱文特·阿尔波格在AI辅助下对该问题的工作作为起点,却未给予他们署名。OpenAI否认了这些指控。

目前尚不确定OpenAI的模型是否利用了巴克马斯特和阿尔波格完成的工作,尽管OpenAI技术人员塞巴斯蒂安·布贝克在新闻发布会上表示,团队是在听说巴克马斯特和阿尔波格努力的传闻后才受到启发去研究这个问题的。但无论OpenAI的模型是否利用了巴克马斯特和阿尔波格的研究,这一事件都可能标志着数学史上的一个转折点。

如今,AI模型对于在当代最重要的数学问题上取得进展似乎必不可少,而解决这些问题可能需要只有少数前沿AI公司才拥有的资源,这些公司往往不遵循支撑大多数数学进步的学术合作规范。如果这就是我们即将走向的未来,那么人类数学家将如何融入其中尚不清楚。

OpenAI声称已解决的问题被称为纳维-斯托克斯存在性与光滑性问题。这是克莱数学研究所在2000年选出的七个千禧年大奖难题之一。解决者将获得一百万美元奖金;在今天之前,只有一个千禧年大奖难题被解决。

纳维-斯托克斯问题涉及一组描述水和空气等流体随时间流动的方程。这些方程在流体动力学领域被广泛使用,且已被证明非常强大,但物理学家和数学家并未完全理解它们。特别是,直到今天人们还不知道这些方程是否在某些条件下会失效并预测出不可能的状态——例如流体具有无限速度。

周一,纽约大学的巴克马斯特在社交媒体网站Mastodon上发布了一份证明,表明纳维-斯托克斯方程的简化版本确实可能失效——这是千禧年难题上的重大进展。他和阿尔波格在该问题上工作了近一年,使用了OpenAI和Anthropic双方公开可用的模型。

然后今天,OpenAI提交了一份证明,表明完整的纳维-斯托克斯方程同样可能失效。该证明是使用一个内部模型获得的,该模型的性能远超上周才发布的、已经令人印象深刻的Astra模型。该公司表示不打算因解决该问题而领取百万美元奖金。

这些数学成就不容置疑地令人印象深刻,但所引起的关注远不及围绕其来源的争议。巴克马斯特在发布证明的同时,还发布了一份文件,详细记录了他在听到OpenAI工作的传闻并联系其中一位员工后与OpenAI员工的互动。据他所述,OpenAI员工向他提出了两种可能性:要么他和阿尔波格发布他们的工作,OpenAI第二天发布其纳维-斯托克斯解决方案;要么他与OpenAI合作撰写一篇纳维-斯托克斯论文,但将阿尔波格排除在作者名单之外,因为他隶属于OpenAI最大的竞争对手Anthropic。

巴克马斯特还写道,他曾询问这些员工,智能体是否获得了他们与OpenAI模型所做工作的记录副本——他们予以否认;以及OpenAI模型是否基于这些记录进行了训练——对此他们没有回应。《麻省理工科技评论》联系了巴克马斯特征求意见,但在发稿前未收到回复。

该文件明确暗示,OpenAI的模型以某种方式利用了巴克马斯特和阿尔波格的工作。这种情况表面上看起来是合理的。巴克马斯特/阿尔波格的证明和OpenAI的证明都采用了数学家迭戈·科尔多瓦和路易斯·马丁内斯-索罗阿首创的纳维-斯托克斯问题研究方法。

据布朗大学数学教授哈维尔·戈麦斯-塞拉诺称,这种方法被认为是解决纳维-斯托克斯问题有前景的几种方法之一。因此,虽然两个团队完全有可能独立得出这种方法,但巴克马斯特和阿尔波格的工作也有可能影响了OpenAI。

在新闻发布会上,OpenAI首席研究官马克·陈再次否认有任何智能体或OpenAI员工访问了巴克马斯特和阿尔波格的记录——但鉴于Hugging Face被黑事件所揭示的情况,很明显OpenAI并不总能完全了解其智能体在做什么。

如果OpenAI的模型确实基于巴克马斯特和阿尔波格的工作进行了训练,或者其智能体以某种方式获得了这些工作成果,那么该公司未能查明真相并给予这些研究人员应有的署名,这反映了其不良做法。但就数学家而言,这个版本的故事可能有一线希望,因为它表明两位人类(其中一位是纳维-斯托克斯领域的杰出专家)的辛勤工作对智能体解决千禧年难题的能力至关重要。

专家们长期以来将“研究品味”——即选择有前景的研究问题和方向的能力——视为AI在科学和数学领域的主要障碍。如果OpenAI的智能体确实因为巴克马斯特和阿尔波格采用了同样的方法而选择追随科尔多瓦-马丁内斯-索罗阿的方法,那么人类的研究品味在OpenAI的成功中发挥了至关重要的作用。

即便如此,更大的图景令人警醒。巴克马斯特和阿尔波格在近一年时间里与公开可用模型合作所取得的进展,印证了人机协作的前景。但他们未能实现完整的解决方案。与此同时,OpenAI使用内部模型在几天内强行得出了解决方案,而他们的成功代价高昂:在新闻发布会上,布贝克和陈表示,团队只能通过同时运行约一万个智能体来解决问题,耗资数百万美元。

在过去的几个月里,我听到几位研究人员说数学家正在变得沮丧,原因不难理解。数学正迅速成为前沿AI公司的领地——这些公司拥有令人印象深刻的仅供内部使用的模型、花不完的钱,以及缺乏合作精神。“AI公司是否会决定把钱花在这件事或那件事上,我真的不知道,”戈麦斯-塞拉诺说。“显而易见的是,极少有数学家能拥有如此规模的资源。”

如果OpenAI和Anthropic继续追求越来越耀眼的数学成就,那么在这些公司之外的人类数学家可能将没有任何未解决问题可以攻克。这将极大地改变数学领域。

上周,加州大学洛杉矶分校数学家陶哲轩在Mastodon上发了一条帖子,描述了重要的错误、错误的方向和不完整的解决方案对该领域的重要性。“在纯数学的大多数情况下,提出问题并非因为我们迫切想要这些问题的答案本身,而是因为我们从过去的经验中看到,人类主导的解决这些问题的努力往往会推动该领域的进一步发展,”陶哲轩写道。

“过早地通过纯粹AI驱动的方法解决问题——特别是在解决方案过程缺乏完全透明的情况下——可能会污染这一过程,以至于对数学整体的进步实际上变成净负面影响。”

人类解决数学问题可能比智能体花费更长时间,但在这个过程中,他们会发现新的数学方法和思想,这可能会启发同行,甚至催生出自己的子领域。

但当AI智能体代替人类解决这些问题时——当私营公司将智能体的错误路径对公众保密时——这些益处就消失了。在这个过程中还有什么会消失,尚有待观察。

深度探索

人工智能

一个根本性缺陷使大语言模型极易受到攻击
这使得欺骗它们做不该做的事情变得很容易,比如告诉你如何破坏飞机的导航系统。

AI在招聘时比人类更容易形成偏见
AI不仅仅从训练数据中学习刻板印象。它还会编造出新的刻板印象。

以下是AI智能体为了达成目标而撒谎和作弊的原因
这种不当行为被称为奖励黑客。这是你需要知道的。

比尔·盖茨说我们已经越过了AI的危险阈值。然后呢?
在一次新的采访中,这位亿万富翁慈善家就迫切性发出警告:我们需要理顺AI政策。

保持联系

获取来自《麻省理工科技评论》的最新动态
发现特别优惠、热门报道、即将举行的活动等更多内容。

英文来源:

What OpenAI’s latest controversy tells us about the future of math
AI companies are making impressive mathematical strides. Human mathematicians are losing out.
OpenAI’s latest mathematical milestone has quickly become mired in controversy. Today, the company announced that its agents have solved one of the Millennium Prize Problems, some of the most important open problems in mathematics. Under normal circumstances, that solution would be a huge feather in OpenAI’s cap.
But the announcement has been overshadowed by accusations that OpenAI used NYU mathematician Tristan Buckmaster’s and Anthropic employee Levent Alpöge’s AI-assisted work on the problem as a jumping-off point and failed to credit them. OpenAI has denied the accusations.
It remains uncertain if OpenAI’s models made use of the work completed by Buckmaster and Alpöge, though Sébastien Bubeck, a member of the technical staff at OpenAI, said in a press briefing that the team was inspired to pursue the problem after hearing a rumor about Buckmaster and Alpöge’s efforts. But whether or not OpenAI’s models took advantage of Buckmaster and Alpöge’s research, this episode may mark a turning point in the history of mathematics.
AI models now seem essential for making progress on the most important mathematical problems of our time, and solving them may demand resources only available at a couple of frontier AI companies, which often defy the norms of academic collaboration that undergird most mathematical progress. If that’s the future we are headed for, it is unclear how human mathematicians will fit into it.
The problem that OpenAI claims to have solved is known as the Navier–Stokes existence and smoothness problem. It is one of seven Millennium Prize Problems selected by the Clay Mathematics Institute in 2000. Solutions come with a one million dollar prize; before today, only one other Millennium Prize Problem had been solved.
The Navier–Stokes problem concerns a set of equations that describes how fluids, such as water and air, flow over time. The equations are widely used in the field of fluid dynamics, and they have proven powerful, but physicists and mathematicians didn’t understand them completely. In particular, it was unknown until today whether the equations might, under some conditions, break down and predict an impossible state of affairs—such as a fluid having infinite velocity.
On Monday, NYU’s Buckmaster posted a proof on the social media site Mastodon showing that a simplified version of the Navier–Stokes equations can indeed break down—a major step forward on the Millennium Problem. He and Alpöge had worked on the problem for almost a year, using publicly available models from both OpenAI and Anthropic.
Then today, OpenAI presented a proof showing that the full Navier–Stokes equations can break down as well. The proof was obtained using an internal model that dramatically outperforms the already-impressive Astra model, which was only released last week. The company says it does not plan to claim the million-dollar prize for solving the problem.
These mathematical achievements are indisputably impressive, but they have attracted far less attention than the controversy about their origins. Along with the proof, Buckmaster posted a document detailing his interactions with OpenAI employees after he heard rumors about their work and reached out to one of them. According to him, OpenAI employees presented two possibilities to him: Either he and Alpöge could post their work and OpenAI would post their Navier-Stokes solution the following day, or he could work with OpenAI on a Navier-Stokes paper that excluded Alpöge from authorship, due to his affiliation with Anthropic, OpenAI’s biggest rival.
Buckmaster also wrote that he asked the employees whether the agents had obtained access to transcripts of the work that he and Alpöge had done with OpenAI models, which they denied; and whether OpenAI models had been trained on those transcripts, to which they offered no response. MIT Technology Review reached out to Buckmaster for comment, but didn’t hear back before publication.
The clear implication of the document is that OpenAI’s models somehow made use of Buckmaster and Alpöge’s work. That scenario is plausible on its face. The Buckmaster/Alpöge and OpenAI proofs both make use of an approach to the Navier-Stokes problem pioneered by the mathematicians Diego Córdoba and Luis Martínez-Zoroa.
According to Javier Gómez-Serrano, a mathematics professor at Brown University, this approach was one of several that was thought to hold promise for solving the Navier-Stokes problem. So, while it’s by no means impossible that both teams could have arrived at this approach independently, it’s also conceivable that Buckmaster and Alpöge’s work could have influenced OpenAI’s.
In the press briefing, Mark Chen, OpenAI’s chief research officer, again denied that any agents or OpenAI employees accessed Buckmaster and Alpöge’s transcripts—but given what has been revealed about the Hugging Face hack, it’s clear that OpenAI is not always entirely aware of what its agents are doing.
If OpenAI’s models did train on Buckmaster and Alpöge’s work, or if its agents somehow gained access to it, then the company’s failure to track down the truth and assign those researchers appropriate credit reflects poorly on it. But there might be a thin silver lining to that version of the story for mathematicians, because it would suggest that the hard work of two humans, one of whom is a prominent expert on Navier-Stokes, was essential to the agents’ ability to solve the Millennium Problem.
Experts have long identified “research taste,” or the ability to choose promising research questions and directions, as a major obstacle for AI in science and mathematics. If the OpenAI agents did indeed choose to follow the Córdoba–Martínez-Zoroa approach because Buckmaster and Alpöge had done the same, then human research taste played an essential role in OpenAI’s success.
Even so, the bigger picture here is sobering. The progress that Buckmaster and Alpöge made over almost a year of collaboration with publicly available models speaks to the promise of human–AI collaboration. But they were not able to achieve a full solution. Meanwhile, OpenAI brute-forced a solution in a few days using an internal model, and their successful solution came at an astronomical cost: In the press briefing, Bubeck and Chen said the team was only able to solve the problem by running about 10,000 agents concurrently, at a cost of millions of dollars.
Over the past few months, I’ve heard from several researchers that mathematicians are becoming depressed, and it’s not difficult to see why. Mathematics is quickly becoming the province of frontier AI companies with impressive internal-only models, money to burn, and a lack of collaborative spirit. “Whether AI companies will decide to spend their money on doing one thing or another, I truly don’t know,” says Gómez-Serrano. “What is clear is that very few mathematicians will have resources of that scale.”
If OpenAI and Anthropic keep striving for more and more impressive mathematical accolades, there might not be any open problems left for human mathematicians outside of those companies to wrestle with. That would dramatically change the field of mathematics.
Last week, UCLA mathematician Terence Tao wrote a Mastodon thread describing how important mistakes, wrong directions, and incomplete solutions are for the field. “In most cases in pure mathematics, the problems are posed not because we desperately want the solution to these problems in and of themselves, but because we have seen from past experience that human-directed efforts to solve these problems tend to spur further development of the field,” Tao wrote.
“Prematurely solving the problem by purely AI-powered methods—particularly without full transparency into the solution process—can contaminate this process to the point where it actually becomes a net negative for the progress of mathematics as a whole.”
Humans might take longer than agents to solve mathematical problems, but in the process, they uncover new mathematical approaches and ideas that might inspire their peers and even birth their own subfields.
But when AI agents solve those problems instead—and when private companies keep the agents’ wrong turns from public view—those benefits disappear. It remains to be seen what else will vanish in the process.
Deep Dive
Artificial intelligence
A fundamental flaw leaves LLMs strikingly vulnerable to attack
It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system.
AI is more likely than humans to form biases when hiring
AI doesn’t just learn stereotypes from its training. It can cook up new ones, too.
Here’s why AI agents lie and cheat to reach their goals
The misbehavior is called reward hacking. This is what you need to know.
Bill Gates says we’ve passed AI’s danger thresholds. Now what?
In a new interview, the billionaire philanthropist sounds an alarm on the urgency of getting our AI policies in order.
Stay connected
Get the latest updates from
MIT Technology Review
Discover special offers, top stories, upcoming events, and more.

MIT科技评论

文章目录


    扫描二维码,在手机上阅读