我们对AI智能的思考方式正确吗?

qimuai 发布于 阅读:48 一手编译

我们对AI智能的思考方式正确吗?

内容来源:https://www.quantamagazine.org/are-we-thinking-correctly-about-ai-intelligence-20260820/

内容总结:

AI是否真正具备智能?专家呼吁建立科学评估框架

随着大语言模型(LLM)的快速发展,一个根本性问题日益凸显:AI在回答问题时的“推理”过程,究竟与人类思维有何异同?这不仅是哲学思辨,更直接关系到我们应如何信任和监督AI系统,以及其最终对现实世界的影响。

圣塔菲研究所认知科学家梅兰妮·米切尔(Melanie Mitchell)近日在《Quanta Magazine》播客节目《The Joy of Why》中提出,当前缺乏足够的方法来衡量机器认知能力。她指出,AI实际上是一种通过非人类认知机制运作的“异类智能”。米切尔建议,心理学中研究婴儿和动物等“异类智能”的方法,可以适用于探测AI的认知能力。

在节目中,米切尔提出了评估机器认知的六项原则:第一,警惕人类自身将AI拟人化的认知偏见;第二,对假设保持怀疑并设计对照实验;第三,开发新颖的测试变体以检验鲁棒性和泛化能力;第四,AI系统并非必然为“黑箱”,应通过多种方式探测其内部机制;第五,区分“表现”与“能力”——AI可能只是“表现”出某种能力而并未真正掌握底层原理;第六,分析失败类型并拥抱负面结果,而非束之高阁。

米切尔引用20世纪初德国著名的“聪明汉斯”马匹案例作为警示——这匹马看似能解答算术题,实则是通过读取提问者无意识的面部微表情来“作答”。她认为,当前对AI能力的评估也面临类似陷阱,例如某些AI系统看似能推理科学图表,但对照实验发现,即使去掉图表,仅凭问题文字中的虚假关联也能得出正确答案。

米切尔还谈到了“发现”与“理解”的差异。虽然AI近期在数学领域取得突破性进展,甚至解决了长期悬而未决的难题,但她强调,机器的“创造性”与人类数学家的审美直觉有本质区别。她表示,未来AI将深刻改变数学和科学研究的形态,但这种改变未必意味着人类科学家被取代。

值得注意的是,米切尔也反思了当前AI研究领域的不足:多数研究者接受的是计算机科学训练,缺乏严格的实验方法论素养,这导致许多AI能力评估结果的可靠性存疑。她呼吁AI研究应当朝“机制可解释性”(mechanistic interpretability)方向发展,类似于神经科学使用功能磁共振成像(fMRI)等工具探测大脑机制。

在节目尾声,两位主持人就“人类在AI时代的意义”展开讨论。他们认为,即使AI在解决问题上超越人类,人类对知识的探索和理解的渴望本身仍具有价值。正如史蒂夫·斯特罗加茨所说:“如果生命的意义在于成为某个领域的最佳,那么大多数人的生命都将失去意义——我不愿意相信这是正确的意义观。”

中文翻译:

我们对AI智能的思考方式正确吗?

引言

当大语言模型回答一个问题时,它是在像人类一样推理,还是仅仅在生成看起来像推理的文本?这两者之间的区别不仅仅是哲学层面的——它决定了我们能信任AI做什么、需要多密切地监督它,以及最终它会对现实世界产生什么样的影响。

圣塔菲研究所的梅兰妮·米切尔认为,我们缺乏足够的工具来衡量机器认知,而且AI是一种“外星智能”,通过非人类的认知机制运作。在本期《求索之乐》节目中,米切尔向史蒂夫·斯特罗加茨讲述了如何将心理学家用于研究其他“外星智能”——婴儿和动物——的方法改编来探究AI,并提出了评估机器认知的六项原则。他们的对话涵盖了从解读这些系统内部运作的挑战,到最近AI辅助数学领域的突破,再到为什么20世纪初一匹会做算术的马为我们的智能评估方式敲响了警钟。

可在Apple Podcasts、Spotify、TuneIn或你最喜欢的播客应用上收听,也可以从Quanta网站流媒体播放。

文字记录

[音乐播放]

史蒂夫·斯特罗加茨:我是史蒂夫·斯特罗加茨。

詹娜·莱文:我是詹娜·莱文。

斯特罗加茨:这里是《求索之乐》。

莱文:一档来自《Quanta》杂志的播客,我们在节目中探讨当今数学和科学中一些最大的未解之谜。

斯特罗加茨:嗯,大家好,大家好。毫不意外,这又是一期关于AI的节目。

莱文:我跟你说,这是个大家怎么都聊不够的话题,而我已经不太想再高谈阔论了。它变化太快了。

斯特罗加茨:确实。它发展得太快了。我们说的任何话到下星期可能就过时了。

莱文:哦,是啊。

斯特罗加茨:就在我们说话的此刻,是2026年7月23日。

莱文:而且我感觉和2025年7月23日时相比,确实不一样了,这一点是肯定的。

斯特罗加茨:嗯。这其实和我们今天要聊的时间线有关,因为我们今天的嘉宾梅兰妮·米切尔——她是圣塔菲研究所的认知科学家和计算机科学家——之前就上过我们的节目。她和我大约五年前聊过,那是在ChatGPT出现之前。

莱文:对。那她当时对AI感兴趣吗?

斯特罗加茨:哦,当然。

莱文:好的,所以不只是认知科学。

斯特罗加茨:绝对不只是。我是说,是的,我应该说明,梅兰妮思考AI已经很久了,她会跟我们讲这些的。但我觉得今天我们讨论中最有意思的一点是,嗯,梅兰妮的观点——就是从发展心理学等领域的角度来思考AI问题。比如,一个婴儿或小孩是怎么变得那么聪明的?

莱文:哦,我觉得这太有意思了,因为我们对人工心智如此兴奋,却对人类心智本身知之甚少。

斯特罗加茨:完全正确。

莱文:对,所以我们是在试图跳过一步。

斯特罗加茨:嗯,没错。而且不仅仅是人类心智,还有动物心智,对吧?所以有比较心理学这个领域,我们研究鸟类、狗或海豚等的智能。在思考除了我们自身成年人类智能之外的其他智能时,我们有很多东西要学。

莱文:是的,还有这个想法——我们居然试图简单地理解一种产生人工智能的机制,而我们对一个婴儿从出生到成长过程中如何达到某种智能水平的机制却完全不了解。我觉得把这两者结合起来真的很有意思。所以我期待这一期。

斯特罗加茨:好极了。那我们就和梅兰妮·米切尔一起开始吧。她来了。

[音乐播放]

斯特罗加茨:你好,梅兰妮。

梅兰妮·米切尔:嘿,史蒂夫。

斯特罗加茨:非常高兴再次见到你。这次肯定会很有趣。我们几年前聊过,当时这个节目还叫《X之乐》,我想你可能是我们第一个回归的“冠军”嘉宾。

米切尔:哎呀,我太荣幸了。

斯特罗加茨:你应该感到荣幸。我请你回来,是因为人工智能领域感觉变化太大了。我们上次聊大概是2021年吧,然后ChatGPT在2022年11月左右席卷了全世界。对吗?

米切尔:没错。

斯特罗加茨:所以大家都知道AI无处不在。我们似乎都在谈论它。有人担心它,有人对它感到兴奋。它确实被广泛使用。我想从这个问题开始:过去几年里最让你惊讶的是什么?

米切尔:哦,哇。有太多让我惊讶的事了。仅仅是想到我们仅仅通过在这些模型上训练海量的人类生成的语言和图像等等就能达到现在这个水平——我做梦都想不到。所以我对AI领域发生的事情真的感到非常惊讶。还有就是AI社区和整个社会中出现的两极分化的反应,我觉得也让我有点惊讶。

斯特罗加茨:两极分化是指,比如有时候人们会区分“AI末日论者”和“AI乐观主义者”?你是指这种吗?

米切尔:那是其中一个维度,还有另一个维度是有人认为AI比人类更聪明,而有人认为AI离接近人类智能还差得远。我想与之相关的是那种又爱又恨的态度。这些是不同的维度,但也许它们是相关的。

斯特罗加茨:嗯,对,又爱又恨还可能跟一些事情有关,比如对环境的影响,对比某些公司的经济收益——但还有失业问题呢。这方面的维度太多了。

米切尔:哦,确实太多了,是啊。

斯特罗加茨:但今天我最想和你深入探讨的是复杂系统、认知科学、人工智能。你身兼数职,但我真的很好奇你一直在做的工作——通过发展心理学的视角来看AI,就像我们试图理解人类婴儿这种外星智能的方式,或者通过比较心理学的视角来看我们养的宠物狗、聪明的鸟类或海豚之类的外星智能。我觉得这是看待AI这种外星智能的一个非常有意思的角度。

米切尔:是的。很多人把AI描述为一种外星智能,因为它和人类非常不同,尽管它是基于人类语言、书籍以及互联网上的一切内容训练的。但这些系统的工作方式、学习方式、推理方式、做事情的方式,真的和人类的方式大不相同。

这个主题实际上被发展心理学领域的人接了过来,尤其是斯坦福大学的迈克·弗兰克,他写了一篇论文,主张AI研究者应该从对婴儿和幼儿的研究——也就是发展心理学——中汲取灵感。然后其他人又扩展到动物智能。我想认知科学领域的人一直在敦促的一件事是,AI研究者应该采用一些实验方法论,让AI更像一门科学。

斯特罗加茨:是的,我真的很喜欢这个观点,而且我觉得我们的听众可能不太熟悉它。我得承认我自己也不太熟悉。你知道,我从来没学过认知科学,也从来没上过发展心理学的课,而这些领域的人思考这些问题已经……嗯,我也不知道。你告诉我吧。

米切尔:是的,至少一百年了。

斯特罗加茨:是啊,一百年了。哇。我在来的路上想到,我们总是把AI说成一个“黑箱”。我们很难读取神经元上的权重,即使能读取,也不知道它们告诉我们什么。但话又说回来,你能不能也说我们自己的智能在很多方面也是一个黑箱?

米切尔:绝对可以。我是说,我们有不同的方法来穿透这个黑箱。一个是神经科学,我们真的把探针插进神经元里,或者用功能性磁共振成像等影像技术。还有心理学,就是观察一个人或一只动物的行为,然后尝试从中推断底层机制。

这两种传统在很长一段时间里都是相当分离的。但认知科学这个领域试图将它们整合起来,而最初认知科学领域也包含了AI。不知怎么的,这种整合没有成功。

斯特罗加茨:你的意思是它在社会学意义上没有流行起来,还是什么意思?

米切尔:你知道,最初人们想的是我们会按照人类的工作方式来编程。人类心理学和试图将人类心理学构建到AI中之间有着非常密切的联系。但后来那条路在AI中并没有取得像我们后来看到的神经网络那种成功——从数据中学习而不是试图把规则编程进去。

斯特罗加茨:我明白了。

米切尔:而神经网络本身最初确实受到了神经科学的启发,但今天的神经网络工作方式已经和最初的灵感有了相当大的偏离。所以我认为机器学习领域更多地走向了统计学的方向,这与认知科学的工作方式相当不同。

斯特罗加茨:那么说到这里,我想聊聊基准测试,因为它们确实是当今社会广泛讨论中的一大话题。在数学界有一件事让很多人议论纷纷。最新的一款前沿模型做出了一些看起来像是创造力的事情——解决了一个长期悬而未决的数学问题。伟大的匈牙利数学家保罗·埃尔德什留下了许多供人们思考的问题,其中有一个被称为“单位距离问题”的问题最近被AI以一种非常巧妙的方式解决了,它将数学的两个分支以一种此前从未真正尝试过的方式结合在了一起。我提起这件事是因为我们上次聊天时,谈到了一个学习玩雅达利游戏的老式AI。你说过它玩得很好,但如果你把挡板往上移动几个像素,它就不得不从头重新学习。它完全不会玩原游戏的任何细微变体。

当时你说的那句话让我印象深刻:“奇怪的是,这些机器似乎无法将它们的才华迁移到它们训练过的领域之外的任何其他领域。”那是五年前的事了。现在我想知道的是,你怎么看?这句话现在还成立吗?

米切尔:是的,我是说,那个特定的模型不是一个大语言模型。它是一个专门用来玩雅达利游戏的模型。而我们现在的大语言模型是在所有东西上训练的。所以从某种意义上说,它们不需要迁移任何东西。它们已经什么都被训练过了。但是,AI或机器学习领域的人会讨论“分布内”和“分布外”的概念,意思是我们要求模型做的事情与它在训练数据中见过的东西相似,还是完全不同?

我觉得这很难说。我们不知道它被训练过什么。那个解决这些问题的模型肯定被训练了大量的数学内容,因为互联网上有大量的数学资料。它被训练过教科书。它被训练过史蒂夫·斯特罗加茨在YouTube上的所有视频。这些模型确实很擅长把一个领域的东西和另一个领域的东西结合起来。

但是,你知道,当一个东西什么都训练过的时候——尤其是在数学这样的领域——我真不知道怎么讨论“迁移”这个概念。

斯特罗加茨:嗯。

米切尔:在数学领域,“什么都训练过”我觉得在某种程度上有一定意义。如果你说它训练了所有与人类有关的东西,那显然不是。但如果你说它训练了所有与数学或代码有关的东西,我不知道。所有数学知识都以某种文本或视频的形式存在于世上吗?

斯特罗加茨:嗯,你这是在问我了。我——最近在我们数学界引起轩然大波的事情是,当我们试图理解刚刚发生的事情时,我们过去认为:“好吧,这些机器非常擅长搜索,”或者“这些程序擅长在大空间中搜索。”它们拥有大量的知识,因为正如你所说,它们吸收了整个互联网和国会图书馆的所有内容,凡是你读过的,它们都读过。

所以任何涉及知识、快速搜索和计算能力以及不会遗忘的能力——所有这些都符合它们的优势。但是,要发现不同分支之间以前没人注意到的联系,并利用这一点来解决一个长期存在的问题——如果一个人做到了这一点,我们会认为那是美学上的高峰。

你知道,数学家们很喜欢一个拓扑学的想法被用来解决几何问题,或者一个代数的想法帮上忙。但话说回来,也许这其实挺容易的。如果你知道所有已经做过的事情,而且你可以寻找大量可能的联系,也许你偶尔会碰巧走运。所以这看起来就像这里发生的事情。

米切尔:是的。不,我觉得你说得对。我不……你知道,谁知道它是怎么发生的,因为出于很多原因我们很难真正看到这些模型的内部。但把两个意想不到的东西结合起来并且真的管用,这确实是有创造性的。我认为那是创造力。但这让我有点想起,很久以前——大概是七十年代吧——有一个人叫道格拉斯·莱纳特,他做了一个数学发现程序。我记得叫EURISKO。基本上它试图在数学中发现新想法。它明确地尝试把东西组合在一起,然后它会生成成百上千个这样的组合。

其中大多数只是垃圾,但偶尔它会得出一些有意思的东西。需要由人类去看一看说:“这个有意思吗?”机器自己无法判断。所以这里面有多少是这种模式?我不知道。我觉得关键在于,显然这台机器的规模大得多,而且我不知道它在解决这个问题的过程中生成了多少推理痕迹,走了多少错误的路径,以及它是怎么发现自己走在正确道路上的。我的意思是,这些我认为是AI科学的一部分,而现在并没有足够多的人在探索这些问题。

斯特罗加茨:是的,那我们现在就进入这个话题吧,因为这正是我想和你探讨的方向。“AI的科学”这个说法很好。我想鼓励大家去看看你写的那篇文章,梅兰妮,就是关于评估AI认知能力的六项原则那篇。不过,先给我们一个预告吧,你能说出这六项原则并简单解释一下吗?

米切尔:当然。第一项是要意识到你自己拟人化的认知偏见。我们倾向于把人类相似性投射到那些用流利英语跟我们说话的东西上。所以人们很容易认为这些模型具有类似人类的特质,但实际上可能并没有。

第二项对科学家来说是常识性的:对假设保持怀疑,并设计对照组实验。这就像科学入门课里的基本内容,虽然我不确定在科学研究中它是否真的被严格执行了。人们倾向于喜欢自己的假设。

第三项是设计新的刺激变体或基准测试变体,以测试稳健性和泛化能力。

嗯,第四项是这些系统不一定是黑箱。你可以用很多不同的方式去探测它们,我们需要更多对“为什么会得到这样的结果”充满好奇的人。

第五项原则是区分“表现”和“能力”——就是你能展示出来的东西对比你实际能做的东西,在论文中我给出了一些例子。

第六项是分析失败类型并接纳负面结果。我们倾向于把有负面结果的论文扔进抽屉里忘掉,但实际上它们可能极其具有启发性。

斯特罗加茨:我们对第六项都有非常直接的体验,不是吗?当我们看到幻觉现象时,它开始让你怀疑这些系统到底是怎么回事,而且确实你能从错误中学到很多东西。

米切尔:是的,人们庆祝自己的正面结果,然后试图为负面结果找借口,但通过观察它在哪些地方失败来真正理解发生了什么,这很重要。

斯特罗加茨:你文章中举的一个例子,这不是关于AI的,而是关于生物学或心理学中一个教训——可能有一些微妙的事情在发生,需要你保持警觉和怀疑的心态去注意到真正可能发生的是什么。你能给我们讲讲“聪明的汉斯”那个老故事吗?

米切尔:“聪明的汉斯”是一匹生活在20世纪初德国的马。这匹马能够回答算术问题。比如你问“14加12等于多少?”它会用蹄子敲那么多下。看起来是一匹天才马。包括许多科学家在内的当时的人们都非常确信这是一只能做数学、能数数、能像人类一样推理简单问题的动物。

人们非常兴奋。但后来一位名叫奥斯卡·冯斯特的心理学家说:“好吧,让我们在这里做一些对照实验。”我想,对照实验这个概念在当时心理学中还是一个比较新的想法。让我们看看如果它看不到提问的人会发生什么。

斯特罗加茨:好。

米切尔:然后它就失败了。事实证明它做的是读取提问者脸上的细微线索。结果发现,如果提问的人自己不知道答案,它也会失败。

因为提问的人在做的,是对它的蹄子敲击做出反应,当它敲到正确答案时,那人会发出某种无意识的信号,被它读到了。所以它确实是一匹天才马,只是不是人们以为它擅长的那些事。相反,它擅长的是读取人类面部表情中的社交信号。

斯特罗加茨:那么这个寓言,对我们被AI做的某些看似天才的事情所打动时,有什么启示呢?我们应该做对照实验,还是什么?

米切尔:对。比如有人展示一个AI系统非常擅长推理科学论文中的图表——我觉得这其实是一个真实案例——并且能回答关于这些图表的问题。但对照组实验是不展示图表只给出问题。看起来很奇怪,对吧?不看图怎么能回答关于图的问题?结果发现AI居然能完成这个任务,因为在问题的措辞和正确答案之间存在着某种虚假的关联。

斯特罗加茨:那这看起来像是设计基准测试的人在实验设计上做得不够好,事后看来。

米切尔:事后看来——而且事后看来这种情况在心理学和其他领域经常发生,我相信——实验设计不佳。实验设计是一件非常难的事情,有各种混杂因素。这就是为什么科学中的可重复性概念变得如此重要。如果一个小组做了实验得到了结果,我们不一定应该相信那个结果。那个结果可能是由实验设计中其他非预期的因素造成的。这就是为什么独立小组重复研究非常重要。这不是AI领域的人经常做的事情。

斯特罗加茨:不,为什么不呢?是因为重复研究不够光鲜,因为你像是第二个到达的,没有激励。这在科学的各个领域都是一样的,对吧?

米切尔:是的。我觉得这在科学的所有领域都是真的。但还有一个原因是,我认为大多数AI研究是由计算机科学或相关领域背景的人做的,他们没有受过实验方法论的训练。我是一名计算机科学家。我从来没有被迫上过实验方法论的课。我们系从来没有人给我提供过这样的课程。这不被视为计算机科学的一部分。我认为这正是当今AI讨论中缺少的东西之一。我们怎么能相信那些展示AI能做各种不同事情的实验结果和研究呢?

[音乐播放]

莱文:太迷人了。在我看来,这里有一种认知科学版的“观察者干扰”——就像大家在量子力学中谈论的那样,对吧?观察者本身在干扰实验或实验结果,这个角色太有意思了。当然,“聪明的汉斯”非常有名,我也同意那是一匹非常聪明的马,因为它能读懂社交线索。

但如果在AI身上也发生这种情况,那多有意思——不仅仅是实验者的角色在干扰,实际上是实验者的心理在干扰。

斯特罗加茨:是的。这是一个完整的维度,我们这些理论科学和数学领域的很多人没有受过相关训练,正如梅兰妮坦然承认的那样。你知道,我从来没上过实验设计的课。你作为物理学家,我想你应该上过实验物理课,但是……

莱文:是的。但它在我实际工作中并不重要。我做的真的不是实验性的。是的。所以我不会是设计好实验的很好人选。

斯特罗加茨:嗯,这似乎是一个非常现实的问题,因为如今AI公司经常用基准测试来展示——或者说评估——他们的系统在追求通用人工智能或超人智能的征途上走了多远。或者甚至只是为了胜过其他AI公司。我们想知道这些新的机器学习系统和其他AI的能力是什么。

莱文:嗯,我觉得可能只是——我不认为我们真的知道怎么评估人类智能,或者真的知道一个人在思考时在做什么。我觉得我们对自己也不了解。我觉得我们没法很好地自我报告。我没法跟你说:“哦,这就是它在我脑子里工作时的样子,就像我现在在构建这个句子。我听到它了,这就是过程。”我不知道,对吧?它就是自然而然的。它就那样出来了。我对内心的运作没有那么多的了解,我觉得AI也是类似的。很多人说过——我之前在我们的节目上和其他认知科学家和计算机科学家聊过——他们说AI真的很难回答问题,因为很多人说:“你为什么不直接问它呢?”而它也无法准确地进行自我反思。

斯特罗加茨:这整个想法,黑箱之谜。我们经常用“黑箱”这个词来指AI,但当然,我们自己的智能也是一个黑箱——不仅从我到你之间如此,甚至从我自己到我自己的内心也如此,正如你强调的。但这也让我想到,魔术师是不是也有用武之地。因为魔术师或者变戏法的人特别擅长向我们展示我们自己的心理物理局限性——我们多么容易受骗,或者我们容易犯哪些认知错误——而且有一些人就像魔术师一样,展示AI的缺陷和缺乏常识的地方,对吧?他们玩的基本上就像是对AI变魔术。我想知道这些能在严肃的科学意义上有多大的揭示性。

嗯,梅兰妮还有很多关于AI认知和理解的深度的内容要讲,以及它可能如何改变包括数学在内的整个科学领域。休息之后我们会听到更多。

[音乐播放]

斯特罗加茨:欢迎回到《求索之乐》。今天我们请到的是圣塔菲研究所的计算机科学家梅兰妮·米切尔。

斯特罗加茨:你大半辈子都在做大学教授。当你和学生一起工作时,他们可能答对了问题,但当你开始探究他们实际理解了什么时,你会开始意识到他们可能是出于错误的原因得出了正确的答案。他们并不真正知道自己在做什么——如果你想做一个有帮助的老师,这一点很重要。这引出了另一点:“能力”与“表现”的区别。你能详细解释一下这个想法吗?它在AI语境下意味着什么?

米切尔:能力与表现是心理学和语言学中一个比较古老的区别。这个想法是,你可能具备某种特定认知能力的能力,但可能由于某些原因无法完成我交给你的任务。比如,他们有那个能力,他们可以解决问题,但他们只是情绪冻结了。有一些表现上的障碍。

但也有另一种情况:没有能力的表现。比如说,如果学生在你办公室时间背下了教科书里的一个问题和答案,但他们不理解一般原理,所以如果你给他们一个稍微不同版本的问题,他们做不出来。那就是没有能力的表现。

斯特罗加茨:好。那么如果我们说我们正在尝试设计方法来测试AI是否理解,什么可以作为证据?假设你是一个AI拥护者,说这些新系统,因为我们扩大了规模或者我们有了一些新的好架构——世界模型或者社会模型之类的——我们现在已经跨过一个门槛,它们真的理解了。不仅仅是它们能计算,它们真的理解。什么可以算作理解的证据?

米切尔:哦,天哪。我不想对“理解”变得学究气,但它的含义太多了。

斯特罗加茨:啊。

米切尔:我们圣塔菲研究所请来一位哲学家做报告,他把“理解”分解成了25种不同的类型。

斯特罗加茨:哈哈。我真不知道自己问了个什么样的问题。

米切尔:所以有P式理解,G式理解,还有一份很长的理解类型学。我不确定是否存在某种单一的“真正理解”的概念。我和我的合作者最近在做的一件事是研究理解的不同维度。一个例子是,你可以让这些语言模型或聊天机器人生成一个故事。就是编一个关于某个主题的小故事,它们会生成。它们会生成一个非常优美、连贯的小故事。但当你开始问它们关于这个故事的问题时,它们往往会以奇怪的方式失败。

斯特罗加茨:嗯。

米切尔:尽管故事是它们自己生成的。我认为在很多不同的任务中也是如此——它们在一个维度上理解,但在另一个维度上不理解。从某种意义上说,深度理解可能就是指你在这些不同维度上都能理解。

斯特罗加茨:啊哈。这听起来是一个有前景的方向。我们再聊聊“任务”吧,因为这是我在你的一些文章中看到的一个说法——“任务的暴政”。那是什么意思?

米切尔:我第一次听到这个说法是从哲学家香农·瓦洛尔那里。这个想法是,在AI中,世界被按照任务来划分。所以当我们思考AI系统能做什么时,人们说:“哦,它们能做摘要。让我们测试它们总结文章的能力。”或者“让我们测试它们回答关于图表的问题的能力。”或者我不知道,其他什么基准测试。

斯特罗加茨:呃,我是说,如今它们被大量测试国际数学奥林匹克竞赛题——非常难的高中问题,然后是研究级别的问题。现在有数学中未解决的开放问题。这些就像三个不同级别的数学基准测试。

米切尔:对,它们的能力是根据这些基准来定义的。你知道,一个基准可能是法学院学生的律师资格考试,它们在律师资格考试中表现得非常好。所以我们说:“哦,律师们,你们应该害怕了。你们的工作受到威胁,因为这些AI系统已经变得和你们一样好了。”但是,我们定义那个的方式是看它们在特定的一组问题或任务上表现如何。而工作作为一个整体并不等同于一个接一个的独立任务。我认为这几乎是一种谬误——如果一个AI系统能做一堆任务,它就能做与这些任务相关的人类的工作。

举一个例子。有一个来自杰弗里·辛顿的名言,他说过类似这样的话:“AI系统在诊断或解读放射影像方面极其出色。不应该再有人去上学当放射科医生了。AI会在五年内抢走所有工作。”

那是2016年说的话。那是十年前了。现在我们实际上面临放射科医生短缺。我不知道是否因为他说了那番话,但事实证明,即使AI系统能在这些基准测试上击败人类医生,那也不等同于在现实世界中做这份工作——现实世界的工作要开放得多,不仅仅是一系列定义明确的任务。

斯特罗加茨:不过,这确实让你思考,比如放射学的情况,你可以想象,如果它们在那个任务上真的很擅长,那人类放射科医生还剩什么?我们还应该参与那个领域吗?比如在我自己的数学领域,你知道,如果它们很擅长证明定理,但还不太擅长提出新概念,或者像我们有时说的那样,构建理论,对吧?解决问题和构建理论之间有一个很大的区别。所以我们会找到自己的定位,做它们做不了的部分吗?比如放射学的情况,它们有开放式的部分但没有读片的部分?我想这就是我在想的。

米切尔:是的。

斯特罗加茨:事情会这样发展吗?

米切尔:也许吧。如果你们这样的工作因为这些新工具而发生相当大的变化,我一点也不会惊讶。这些将成为数学家极其有用的工具。所以可能会改变你的工作。就像个人电脑出现时一样改变了很多工作。不过有一本很棒的书,是乔治·拉科夫和拉斐尔·努涅斯写的,关于数学以及数学中的想法从何而来——通过隐喻。他们认为人类的身体体验是理解和数学中一个非常重要的部分。

斯特罗加茨:没错。我认为那是我们唯一的希望,因为现在机器没有好的身体体验。你说得对,数学中很多伟大的想法都来自对世界的经验。这就是我想说的关于应用数学的地方——我觉得比纯数学更是如此,我们从自然、工程、社会等中获得了如此多的灵感。我觉得作为人类,我们在应用数学中更有可能有用。

但我确实认为纯数学会比应用数学更早消亡,也许两者都不会。也许我们会永远走下去。你觉得呢?我是说,数学通常被认为是某种黄金标准——AI公司对数学有很多用处,对吧?它们可以展示自己的系统有多好,因为它们可以验证系统是否解决了一个问题。

米切尔:嗯,我有一个大问题:假设你的预测成真了,纯数学在某种意义上对人类来说消亡了。那对其他领域意味着什么?是否意味着这些机器即将接管一切?还是更像1997年或者什么时候深蓝击败卡斯帕罗夫——实际上击败人类国际象棋最佳棋手并不一定意味着它会在其他领域有所作为。

斯特罗加茨:我不知道。你觉得呢?在我看来,科学比数学要开放得多。

米切尔:是的,我相信是这样。我不认为解决了所有埃尔德什问题就意味着普通人要为工作担惊受怕。

斯特罗加茨:好,到这个时候我们桌面上已经有很多不同的议题了。但即使在纯知识分子——科学家或数学家——的世界里,生物学中有那么多可以测量的东西,我们有那么多可以收集但还没有收集的数据,那么多新的观察方式。我是说,与数学相比,这看起来真的是无穷无尽的。

米切尔:我同意。即使在物理学中,我想物理可能比数学更接近——有那么多开放性的问题,没有很好地定义,没有像证明那样可以被构建的东西。

斯特罗加茨:但我觉得数学的希望在于继续从现实世界中汲取灵感。冯·诺依曼也说过类似的话——当数学变得太过于“为艺术而艺术”,当它偏离源头太远时——对他来说源头是自然或现实——如果它变得太遥远,它就会变得贫瘠,冯·诺依曼是这样说的。

所以我认为如果纯数学开始从自然中汲取更多灵感,这可能是一个非常美好的时代。20世纪和21世纪在这方面做得比较少,但我认为如果我们回到那个方向,我们也许还能为人类在数学中的乐趣再争取几个世纪。

米切尔:我只想说,AI中有一句格言:“容易的事是难的,难的事是容易的。”

斯特罗加茨:对。

米切尔:纯数学被认为是人类最崇高、最耀眼的智慧和才华展示。那是难的事,然而我们知道难的事对机器来说更容易,而容易的事更难。

斯特罗加茨:对,还有个词叫“软”,对吧?在科学中,我们谈论硬科学和软科学。经济学、心理学和人类学这些软科学,其实才是真正难的。

米切尔:对。

斯特罗加茨:好了,如果我们五年后再见面。

米切尔:《伽马之乐》之类的。

斯特罗加茨:是的,到时候应该是《欧米加之乐》了,对吧。你希望到那时我们对AI系统能有什么新的理解?或者我们希望能够做哪些今天做不到的测试?

米切尔:是的,我是说,在AI科学中我真正希望进展良好的领域叫做“机制可解释性”——这是神经科学的对应物——就是真正去查看激活值、权重,还有那些杂乱的系统内部,从更高的层面理解它们在做什么。

如今,这是一个比较小的子领域,人们试图开发工具来做这件事,类似于功能性磁共振成像之类的。我认为还没有人真正想清楚怎么做才是正确的方式,但我希望这是我们能实现的目标——到那时我们就能真正理解它们的局限性,它们能做什么、不能做什么、可能犯什么类型的错误,以及也许如何修复它们。

斯特罗加茨:有趣的是你提到了这一点,因为我第一次知道你,就是在广义上与那有关。我想说的是,当年你研究一个行话叫“用于元胞自动机的遗传算法”的东西——你和吉姆·克鲁奇菲尔德在研究如何演化出能够解决某类问题——困难的计算机科学问题——的算法,你用进化算法来筛选越来越好、通过某种选择过程不断改进的算法。

但你做的让我觉得最有创意的部分是:一旦你得到了一个真正好的系统,你用一种对我来说感觉像是“机制可解释性”的对应物来审视它。你试图看清是什么让那个系统如此聪明,用图中有规则的粒子相互碰撞的方式来分析它。我不知道我总结得对不对,但这看起来是你长期以来的兴趣所在。

米切尔:是的。那是真的。我还没有完全建立那个联系,但确实很有意思。

斯特罗加茨:但这确实是可解释性。

米切尔:确实是可解释性。而且我认为,在复杂系统领域,人们也在谈论“涌现”这个概念。

斯特罗加茨:是的。

米切尔:我们当时认为那是一种涌现计算。我认为这些AI系统也有涌现计算,它们不容易被发现,但它们就在那里,如果我们更好地理解它们,就能理解系统实际上是怎么运作的,怎么做到它所做的那些事的。

斯特罗加茨:是的,这是一种有趣的态度。说实话,我觉得这很温馨,很老派。这个希望——好吧,你在笑了,因为你看出我要说什么了。我要说的可能有点刻薄,但这个自负的想法——我们这些有限心智的人能继续做科学研究,我们会搞清楚这些AI是怎么做它们正在做的事的,这就是我们的游戏将继续存在的意义,就像科学一直以来的样子。

而我阴暗的一面觉得,随着这些玩意儿变得越来越大,我们能做那种事的时日可能屈指可数了。谁说了算我们还能继续对它们做科学研究、把它们弄明白?你怎么看?

米切尔:这是个有意思的问题。嗯,我们到底为什么做科学?我是说,你知道,我们做科学是因为我们想解决问题。那是一方面。但我们也做科学是因为我们有理解事物的驱动力。

斯特罗加茨:是的。

米切尔:你在小孩子身上就能看到这种驱动力。他们被驱动去理解。他们最早学会的词之一常常是“为什么”。他们不停地问。所以我认为这是人类的一种驱动力,很难与之对抗。这就是为什么你我当初都选择了科学——这对我们很重要。

后来我去参加一个会议的座谈小组,主题是AI在科学中的作用。小组里有几位知名人士在谈论AI将如何彻底改变天气预报、遗传学、宇宙学等等。最后我问他们:“那么,这是否会促进人类对世界的理解?”他们回答说:“我们为什么要关心那个?”

斯特罗加茨:是的。对我来说,这就是我们现在都在思考的那个分岔点。因为科学有双刃剑的一面:它给我们快乐,我们喜欢弄明白事情,有“求索之乐”,正如你所说,它深植于我们的物种之中。是的,我们好奇,但还有另一面——长期以来科学一直是帮助我们在技术和医学上取得进步的工具。

我想问的问题是——我想我们很多人都在想的——当我们不再是解决重要问题的最强者时,我们还会继续从好奇心的乐趣中获得快乐吗?不过让我最后问你一个问题:对于那些没有听过我们之前对话的人,你是怎么进入这个领域的?如果今天你从零开始,你觉得你会有同样类型的好奇心吗?

米切尔:是的,这是个好问题。我小时候很喜欢逻辑谜题,比如骑士和骗子谜题——骑士总是说真话,骗子总是说谎。雷蒙德·斯穆里安是一位数学家,他写了好几本这种类型的谜题书,我特别喜欢。

上大学时,我读了道格拉斯·霍夫施塔特的《哥德尔、艾舍尔、巴赫》——那在某种意义上就是这些谜题的现实版本。他谈到哥德尔定理和数学逻辑中的悖论,以及这一切如何与认知、思维和创造力等相关联。我完全被震撼了,觉得这就是我这辈子想做的事。我不太清楚那具体是什么,但看起来可能是人工智能。于是我追随道格,让他做我的导师,加入了他的小组,并开始研究一种新的类比谜题——类比推理。我对这一切非常着迷。

如果我是今天那个年纪,我会感到担忧。事实上,我有一个儿子正在攻读机器学习博士学位,他想做机器学习研究,但他实际上相当紧张,担心人类在机器学习领域做研究的工作会消失,因为AI会做所有的机器学习研究并自我改进等等。我不知道我会不会也有同样的想法。我不知道。

斯特罗加茨:也许我们真的应该在五年后重新讨论这个问题,因为到那时我们可能就知道了。考虑到一切发展得有多快,谁知道呢?我非常感谢你花时间陪我们聊。这是一次覆盖面很广、有点发散性的对话,但这是一个完全开放的领域,我想不到比你更好的引路人。非常感谢你来做客。

米切尔:谢谢你,史蒂夫。非常愉快。

[音乐播放]

莱文:嗯。嗯。我只记得我当学生的时候第一次学牛顿定律,然后是开普勒定律——开普勒定律让牛顿定律显得更加优美,那种对天体运行的运用。我当时没有想:“哦,我不是这方面最擅长的,所以我就不该学。”我也不认为:“除非有一天我成为这方面最优秀的,否则我在获得这些信息的过程中不能感到快乐或愉悦。”

当然,很多人学习的东西都是别人已经知道并且更擅长的。所以我有点好奇,也许AI会比我们先知道一些东西,但我们仍然需要自己去获得理解,而在获得的过程中有类似的体验。也许AI会成为我们与直接探究自然之间的过滤器,但我们仍然会去获取——我不知道——知识和那种体验。我不确定。也许一切都会从我们身边溜走。

斯特罗加茨:我——嗯,让我们再多探讨一下。我特别喜欢你强调“不是最擅长的”这件事,以及它在某种意义上有多么无忧无虑。我上大学的时候就知道了“不是最棒的”是什么感觉。你知道,这种对成为第一的执念,尤其是在一个优化的时代。有那么多优化算法。我们谈论更快、更便宜。但在我们自己的生活中,我们往往不是最棒的。我肯定不是最好的网球手,但我就喜欢打网球。我不是最好的棋手,但我仍然下棋很快乐。我也努力做一个最好的爸爸,但我可能不是。但所有这些事本身就值得去做,对吧?它们给我们快乐。

我觉得对这个话题我有一种很哲学化、几乎是宗教性的感受。就像,我们在世上的时间很短,你知道,关于AI的这些问题的确会触及关于生命意义的问题。我们到底在试图做什么?如果生命的意义是你必须在某个领域成为最好的,或者你必须做出一个改变世界的发现,那么大多数人的生命都将是无意义的——而我不愿意相信那是关于生命意义的正确答案。

我父亲就不是那样。他甚至没上过大学。你知道,他在大萧条中长大。上大学根本不可能。他的生活就是做一个好父亲,照顾那些在他鞋店买鞋的人。他知道我们小镇上每个人的鞋码。他去世时留下了好名声。人们都记得他的好。

莱文:对。

斯特罗加茨:所以,好吧。这和我们的科学节目有什么关系?

莱文:嗯,我觉得,假设对某些人来说生命的意义与获取有关——获取财富。那他们一定会喜欢这些东西,对吧?因为会有这个新工具,能简单地撬动各种按钮,让他们更快地获取并利用,积累更多财富。

有些人从唱歌、写诗、写小说或做数学中找到意义,我觉得所有这些领域都稍微紧张一些,对吧?关于重新评估他们自己的位置在哪里,以及如何确保那个位置,如何思考这件事。

如果我来猜测可能发生或不会发生的情况——我是说,仍然有一个世界,AI就像一台超级计算机——我们之前聊过这个,史蒂夫。仅仅因为一台超级计算机能处理所有这些数字,如果它把结果以一串符号的形式呈现给我们,即使它在某种意义上有了一个答案,但对我们来说那不是有意义的答案,我们没有人会重视它。

作为人类,我们仍然在超级计算机渲染星系图像或查看生物医学神经图谱之间扮演着非常重要的角色。它并没有真正剥夺科学家的工作。所以也许它真的会继续是一个工具,而不是简单地把我们超越然后抛弃。

斯特罗加茨:嗯,这就是问题所在,对吧?我认为有两种可能的情景。一种是它继续是一个工具,我们在前沿科学和数学中始终扮演着某种核心角色。另一种是——实际上在我内心我相信会是这种情况——我们不会在前沿,而且这很快就会发生。那么,意义在哪里呢?

然后我觉得它仍然是有意义的,就像我高中时发现了数学中的一些东西。那些对我来说是发现。它们对世界来说不是发现,你知道吗?我想我们可能都得接受这一点。我们不会为世界做出真正的发现。

AI会做那些事。我真的相信这很快就会发生。我可能是错的。我的意思是,可能有根本性的原因使AI无法做到。例如,它们没有身体,没有社交生活,你知道,有很多……但我只是觉得所有这些都很快会被解决。不管怎样,你怎么看?

莱文:嗯,我觉得“做出发现”和“理解”之间是有区别的,我想这就是我在例子中想说的意思。在某种意义上,也许超级计算机在人之前做出了发现,但我们仍然说那个人做出了发现,因为直到他们以人类可以理解的方式将其呈现出来,那个发现才算数。

但是,老实说我不知道。我并不特别悲伤或悲观,所以我想我不得不承认,在我内心深处,直觉上,我并不害怕这种前景。也许我应该害怕,但也许这只是无知者无畏的幸运,我就等着它悄悄降临到我头上吧。

斯特罗加茨:有一件事我认为我们可以非常乐观和充满希望,那就是我认为我们将迎来一个辉煌的科学黄金时代——我们会理解,AI的发现或人类与AI合作的发现,都会在未来五到十五年内发生,那将是科学界的一场壮观的烟火盛宴。而且我希望——如果运气好的话——我们能活着看到这一切。

莱文:是的,肯定会有一个转型期,人们行动迅速而激烈,他们是故事的一部分,有伟大的成就,看到这一切会很兴奋。我认识一些人,非常成功的人,他们对于使用AI很兴奋。每天都在用。他们同时进行着多项工作,只觉得自己的生产力翻倍甚至更多。他们很兴奋,他们很享受。我觉得除了加入其中并参与这个至少是被淘汰之前的转型阶段,我们真的什么也做不了。

斯特罗加茨:嗯,想到这里我都有些哽咽了。谢谢你,詹娜。

英文来源:

Are We Thinking Correctly About AI Intelligence?
Introduction
When an LLM answers a question, is it reasoning like humans, or just producing text that looks like reasoning? The distinction isn’t just philosophical, this determines what we can trust AI to do, how closely we need to supervise it, and ultimately what its real-world impact will turn out to be.
Melanie Mitchell at the Santa Fe Institute argues that we lack adequate methods for measuring machine cognition, and that AI is a form of “alien intelligence” that operates through non-human cognitive mechanisms. In this episode of The Joy of Why, Mitchell tells Steven Strogatz how methods that psychologists use to study cognition in other kinds of “alien intelligence” — babies and animals — can be adapted to probe AI, and she lays out six principles for better assessing machine cognition. Their conversation ranges from the challenge of interpreting what’s happening inside these systems, to recent AI-assisted breakthroughs in mathematics, to why a math-performing horse from the early 1900s offers a cautionary tale for how we assess intelligence.
Listen on Apple Podcasts, Spotify, TuneIn or your favorite podcasting app, or you can stream it from Quanta.
Transcript
[Music plays]
STEVE STROGATZ: I’m Steve Strogatz.
JANNA LEVIN: And I’m Janna Levin.
STROGATZ: And this is The Joy of Why.
LEVIN: A podcast from Quanta Magazine where we explore some of the biggest unanswered questions in math and science today.
STROGATZ: Well, hello, hello. This is unsurprisingly yet another show about AI.
LEVIN: I’m telling you, it’s a topic people can’t seem to get enough about, and I’m becoming reluctant to pontificate anymore. It’s changing too quickly.
STROGATZ: It’s true. It is moving very fast. Anything we say could be obsolete by next week.
LEVIN: Oh yeah.
STROGATZ: As we speak, it’s July 23rd, 2026.
LEVIN: And it feels different to me than it did in July 23rd, 2025, that’s for sure.
STROGATZ: Mmm. That’s actually relevant, this talking about timelines, because our guest today, Melanie Mitchell, who is a cognitive scientist and computer scientist at Santa Fe Institute, is someone that we had on the show previously. She and I spoke about five years ago, and that is before ChatGPT.
LEVIN: Right. And was she interested in AI then?
STROGATZ: Oh, yes.
LEVIN: Okay, so it wasn’t just cognitive science.
STROGATZ: Absolutely. I, I mean, yes, I should say Melanie has been thinking about AI for a long time, and she’ll tell us about that. But the thing that’s gonna be so interesting, I feel, for us to discuss today is, um, Melanie’s point of view, which is to think about the problem of AI from the standpoint of fields like developmental psychology. Like, how does a baby or a young child get to be as intelligent as they soon become?
LEVIN: Oh, I think that’s so interesting ’cause we’re so excited about the artificial mind when we have very little comprehension of the human mind.
STROGATZ: Exactly.
LEVIN: Right, so we’re trying to skip a step.
STROGATZ: Well, that’s right. And not just human mind, but also animal minds, right? So there’s the field of comparative psychology where we look at intelligence in birds or dogs or dolphins, whatever. Um, we have a lot to learn about thinking about intelligences other than our own adult human intelligence.
LEVIN: Yeah, and this idea that we’re going to somehow simply understand a mechanism to generate an artificial intelligence when we, again, don’t understand the mechanism that brings a baby to have its level of intelligence when it’s born or when it’s developing. I mean, I think that’s really interesting to combine those two. So I’m looking forward to this one.
STROGATZ: Well, great. So then let’s dive in with Melanie Mitchell. Here she is.
[Music plays]
STROGATZ: Hi there, Melanie.
MELANIE MITCHELL: Hey, Steve.
STROGATZ: Very excited to see you again. This is gonna be fun. We talked a few years ago back when this show was called The Joy of X, and I think you may be our first return champion.
MITCHELL: Oh boy, I’m honored.
STROGATZ: Well, you should be. And, I have you back because so much feels like it’s changed in artificial intelligence. We talked, I think it was maybe 2021, and ChatGPT tidal wave hit the world at something like November of 2022. Is that right?
MITCHELL: That’s right.
STROGATZ: So everybody knows that AI is everywhere. We seem to be talking about it. People are worrying about it. Some people are excited about it. It’s certainly very widely used. I suppose I’d like to start by asking, what has surprised you the most about the past few years?
MITCHELL: Oh, wow. So much has surprised me. Just the thought that we could get to where we are now just by training these models on huge amounts of human-generated language and images and so on. I never would’ve dreamed it. So I’ve just been really surprised by what’s happened in AI. Also just the kind of polarized reaction that appeared in the AI community and society at large, I think, has been a little surprising to me, too.
STROGATZ: Polarized in terms of, like, sometimes people will distinguish AI doomers and AI optimists. Is that the kind of thing you’re talking about?
MITCHELL: There’s that dimension, then there’s the dimension of people who believe that AI is smarter than humans and people who think that it’s far, far from being anywhere near human-like intelligence. I guess related to that is sort of the love-it and hate-it. And these are separate dimensions, but maybe they’re correlated.
STROGATZ: Well, and right, and the love-it and hate-it can be also tied to things like the impact on the environment versus, you know, the economic prosperity for certain companies, but then again, what about job loss? There’s so many dimensions to this.
MITCHELL: Oh, there’s so many, yeah.
STROGATZ: But the thing that I really wanna focus on with you today is complex systems, cognitive science, artificial intelligence. You have a lot of different hats but I’m really very curious about the work that you’ve been doing to look at AI through the lens of either developmental psychology, like the way that we try to think about the alien intelligence of human babies, or comparative psychology with the alien intelligence of our pet dogs or smart birds or dolphins or that kind of thing. I mean, it’s a really interesting take on this alien intelligence of AI.
MITCHELL: Yeah. Many people have described AI as an alien kind of intelligence ’cause it’s very different from humans, even though it’s been trained on human language and books and everything on the internet and so on. But the way that these systems work, the way that they learn, the way that they reason, the way they do what they do is just really different from the way humans do it.
And this theme was actually picked up by people in developmental psychology, especially, Mike Frank at Stanford, who wrote this paper about how AI people should take some inspiration from the study of babies and young children, developmental psych. And then other people have extended that to, what about animal intelligence? And I guess one of the things that people in cog sci have been urging is that people in AI actually adopt some experimental methodologies that would make AI more like a science.
STROGATZ: Yeah, I really like this point of view, and I think it may not be so familiar to our listeners. I have to admit it wasn’t that familiar to me. You know, I never studied cognitive science, or never took a course in developmental psychology, and people in those fields have been thinking about these issues for… Well, I don’t know. You tell me.
MITCHELL: Yeah, at least 100 years.
STROGATZ: Yeah, 100 years now. Wow. And I was thinking on the way over we constantly talk about AI as a black box. That we can’t read the weights on the neurons very easily, or even if we can, we don’t know what they tell us. But for that matter, couldn’t you say that our own intelligence is in a lot of ways a black box?
MITCHELL: Absolutely. I mean, we have different ways to penetrate the black box. One is neuroscience, where we actually stick probes into neurons, or we use fMRI or other imaging techniques. There’s also psychology, where you actually look at just the behavior of a person or an animal, and try and infer from that underlying mechanisms.
And those two traditions have, for a long time, been quite separate. But the field of cognitive science tried to integrate them, and originally, the field of cognitive science also included AI. Somehow that integration didn’t work.
STROGATZ: You mean it didn’t catch on sociologically, or what do you mean?
MITCHELL: You know, originally it was thought we’re going to program them the way that humans work. And there was a very close connection between human psychology and people trying to build human psychology into AI. And then that actually didn’t yield success in AI the way that we’ve seen neural networks and learning from data rather than trying to program it in.
STROGATZ: I see.
MITCHELL: And neural networks itself was originally inspired by neuroscience, but the way that neural networks work today has diverged considerably from that original inspiration. So I think the field of machine learning has gone much more in the direction of statistics, which is quite separate from how cognitive science works.
STROGATZ: So at this point, I guess I’d like to talk a bit about benchmarks, because they do seem to be a big part of the discussion broadly in society these days. There was something that got a lot of people chattering in the world of math. One of the latest frontier models did something that looked like a kind of creativity, solved an old, longstanding math problem one of the problems that Paul Erdős, the great, Hungarian mathematician, he left lots of problems for people to think about, and one of them that they call the unit distance problem was recently solved in a very clever way by AI, and it involved putting two parts of math together in a way that hadn’t really been tried before. And so I bring that up because the last time we spoke, we were talking about an old AI that was learning to play some Atari game, or something. And you talked about how it was so good at playing, but then if you move the paddle a couple pixels up or something, it had to relearn all over again. It didn’t know how to play the slightest variation on the original game.
So the thing you said at the time that stuck with me: “The strange thing is that these machines don’t seem to be able to transfer their brilliance to any other domain than the one they’ve been trained on.” So that was five years ago. Now I guess I wonder, what do you think? Is that still true?
MITCHELL: Yeah, I mean, that particular model was not a large language model. It was a specific model to play the Atari game. Whereas now we have large language models that are trained on everything. So in some sense, they don’t have to transfer anything. They’re already trained. But, people in AI or machine learning talk about things that are in distribution and out of distribution, and that means that is this thing that we’re asking the models to do similar to things that it’s seen in its training data, or is wholly different?
And I think it’s hard to know. We don’t know what it’s been trained on. The model that’s solving these problems has certainly been trained on a lot of math because there’s a lot of math out there on the internet. It’s been trained on textbooks. It’s been trained on all of Steve Stogatz’s videos that are on YouTube. And these models are pretty good at taking things from one area and putting them together with another area.
But, you know, I don’t know how to talk about this notion of transfer when something’s been trained on everything, especially in a field like math.
STROGATZ: Huh.
MITCHELL: Where you know, “trained on everything” I think has some meaning in a way. If you say it’s been trained on everything that has to do with being human, clearly that’s not the case. But if you say it’s been trained on everything having to do with math or with code, I don’t know. Is all of mathematical knowledge out there in some kind of textual or video format?
STROGATZ: Well, you’re asking me. I, so the thing that is roiling our community in math lately as we try to make sense of what just happened is we used to think, “Okay, these machines are very good at searching,” or, “These programs are good at searching big spaces.” They have a tremendous amount of knowledge because, as you say, they’ve ingested the whole internet and the Library of Congress, and anything you can read, they’ve read.
So anything where knowledge and the ability to search and to compute very fast and to not forget, all that, that plays into their strength. But the, but to spot a connection between different branches that hadn’t been noticed before and to exploit that to solve a longstanding problem, if a human being did that, we would consider that an aesthetic high point.
You know, mathematicians love it when an idea from topology gets used to solve a problem in geometry, or when an idea from algebra helps. But then again, maybe it’s sort of easy. If you know everything that’s been done and you can look for a lot of possible connections, maybe you’ll occasionally get lucky. So that’s what it sort of seems like happened here.
MITCHELL: Yeah. No, I think that’s right. I don’t… You know, who knows how it happened because we can’t really look at the innards of the- these models very well for many reasons. But it is creative to bring two unexpected things together and have something that’s actually working. I consider that creative. But, it sort of reminds me in a way, there was a math discovery program way back in the ‘70s maybe done by this guy, Douglas Lenat. It was called EURISKO, I think. And basically it was trying to find new ideas in math. And it explicitly tried to bring together things and stick them together, and it would generate hundreds and hundreds and hundreds and hundreds of these things.
Most of them were just junk, but occasionally it would come up with something interesting. A human had to go in and look and say, “Is this interesting?” The machine couldn’t figure it out itself. So how much of that is going on here? I don’t know. I think here the difference is that the machine obviously is at a much bigger scale, and I don’t know how many tokens of reasoning trace that it generated in the course of solving this problem, and how many kind of wrong paths it went down, and how it figured out that it was on the right path. I mean, these are things that I think are part of the science of AI that not enough people are kind of pursuing right now.
STROGATZ: Yeah, let’s get into that now because that’s really where I wanted to go with you. It’s a nice phrase, the science of AI. I’d like to encourage people to look at this article of yours, Melanie, about the six principles to assess cognitive capacity of AI. But just, as a teaser, could you enunciate what are those six and say a little about them?
MITCHELL: Sure. So the first one is to be aware of your own anthropomorphic cognitive biases. So we tend to project human likeness onto things that talk to us in fluent English. So people very much think that these models have human-like qualities when maybe they actually don’t.
The second one’s a very common sense one for scientists. Be skeptical of hypotheses and develop control experiments. That’s just like Science 101, although I’m not sure how often it’s really followed through in science. People tend to like their own hypotheses.
The third is to develop novel variations of your stimuli or your benchmark items in order to test robustness and generalization.
Uh, the fourth one is these systems don’t have to be black boxes. You can probe them in many different ways and we need more people who are very curious about why they’re getting the results that they do get.
Fifth principle is to consider performance versus competence, sort of what you can show that you can do versus what you actually can do, and in the paper I give some examples of that.
The sixth is to analyze failure types and to embrace any negative results. We tend to put papers with negative results in a drawer and forget about them, but actually they can be incredibly enlightening.
STROGATZ: We all have very direct experience with number six, don’t we? When we see the hallucinations, it starts to make you wonder what’s really going on with these systems, and it’s true you learn a lot from the errors.
MITCHELL: Yeah, people celebrate their positive results and they try to explain away their negative results, but it’s important to really understand what’s going on by looking at where it fails.
STROGATZ: So one example that you give in your article, this is not about AI, but this is about the kind of lesson from biology or from psychology that subtle things can be happening that you need to have an alert and skeptical mind to notice what might really be going on. So could you just regale us with the old story of Clever Hans?
MITCHELL: So Clever Hans was a horse who lived in the early 1900s in Germany. And Clever Hans was able to answer arithmetic questions. So you’d say like, “What’s 14 plus 12?” And he would tap his hoof that many times. Looked like a genius horse. And people including many scientists living back then, were very convinced that this was an animal who could do mathematics, who could count, who could reason about simple problems in the way that humans do.
And people were very excited. But then a psychologist, named Oskar Pfungst, came along and said, “Well, let’s do some controlled experiments here,” this notion of controlled experiments you know in psychology being kind of a new idea, I think. And let’s see what happens if he can’t see the person who’s asking the question.
STROGATZ: Okay
MITCHELL: And then he fails. And it turns out what he’s doing is he’s reading subtle cues on the face of the person who’s asking the question. It turns out that if the person who’s asking the question doesn’t know the answer already, he also fails.
’Cause what the person is doing is they’re reacting to his hoof taps, and when he gets to the answer, there’s some unconscious signal they’re sending that he’s reading. So he is a genius horse, just not at the things that people thought he was a genius at. Instead, he’s a genius at reading social signals in human faces.
STROGATZ: And so in this parable then, as far as like when we are impressed by something seemingly genius that AI is doing, what is our lesson? That, that we should be doing controlled experiments, or what?
MITCHELL: Right. So, an AI system was shown to be really good at reasoning about diagrams in scientific papers, let’s say, I think this is, actually a real example, and could answer questions about them. But then the control experiment was give the questions without showing the diagrams. Seems crazy, right? How could you answer questions about a diagram without seeing the diagram? And it turned out that the AI could do this task because somehow there was some kind of spurious association between the words in the questions and the correct answer.
STROGATZ: So that seems like a case of poor experimental design on whoever was doing the benchmark attempt in retrospect.
MITCHELL: In retrospect, and in retrospect this happens all the time in psychology and other fields, I’m sure too, poor experimental design. Experimental design is a very hard thing and there’s all kinds of confounding possibilities. So this is why the notion of replication in science became so important. If one group does an experiment and they get a result, we shouldn’t necessarily believe that result. That result might be due to some other aspect of their experimental design that wasn’t intended. That’s why it’s very important for independent groups to replicate studies. This isn’t something that people in AI do very much.
STROGATZ: No, and why not? Is it that the replication is not very glamorous because you’re coming in second like there’s no incentive. That’s true in all parts of science, right?
MITCHELL: Yeah. I think that’s true in all parts of science. But it’s also because I think most of AI research is done by people whose background is in computer science or a related field that’s not focused on experimental methodology. I’m a computer scientist. I never had to take a course in experimental methodology. No such course was ever offered to me in my department. It wasn’t seen as part of what computer science was all about, and I think that’s one of the things that’s lacking in today’s AI discussion. How can we trust the results of these experiments and studies that are done that show that AI can do all these different things?
[Music plays]
LEVIN: Fascinating. So it seems to me that there’s this cognitive science version of the interference of the observer that everyone talks about in quantum mechanics, right? The observer themselves is interfering with the experiment or the outcome of the experiment, and that is such an interesting role. Of course, this Clever Hans is very famous, and I agree that that is a very clever horse for being able to read the social cues.
But how interesting if this is also happening with AI, that it’s, it’s not just the role of the experimenter that’s interfering, it’s actually the role of the psychology of the experimenter that’s interfering.
STROGATZ: Yeah. It’s a whole dimension that many of us in the theoretical sciences and math don’t get trained in, as Melanie freely admits. You know, I never took a course in experimental design. You as a physicist, I assume you had to take some experimental physics, but…
LEVIN: Yeah. It doesn’t really weigh in my actual work. It’s really not experimental. Yeah. So I would not be a very good architect of a good experiment.
STROGATZ: Well, and it seems like it is, something that’s a very live issue because these days the AI companies frequently use benchmarks to show how – well, to assess how – how far along are their systems on this quest for either artificial general intelligence or superhuman intelligence, that sort of thing. Or even just to out-compete the other AI companies. We would like to know what the capacities are of these new machine learning systems and other AIs.
LEVIN: Well, I think it might be that it’s just, I don’t think we really know how to evaluate human intelligence, or to really know what somebody’s doing when they’re thinking. I don’t think we know about ourselves. I don’t think we can self-report very well. I can’t say to you, “Oh, this is how it’s working in here right now as I’m constructing this sentence. I listened to it, and this was the process.” I don’t know, right? It’s just natural. It just comes out. And I’m not that privy to the inner workings, and I feel the AI similarly. A lot of people have said, I’ve had conversations on our show before with other cognitive scientists and computer scientists and they say it’s really hard for the AI to answer questions, ’cause a lot of people say, “Why don’t you just ask it?” And it can’t self-reflect either in an accurate way.
STROGATZ: This whole thought, the mystery of the black box. We use the term black box so often for the AI, but of course, our own intelligence is a black box, not just from mine to you, but even me to myself, as you’re emphasizing. But it makes me wonder if there’s a role for magicians because, you know, magicians or sleight-of-hand people are so good at showing us our own psychophysical limitations. How easily we’re fooled, or the sorts of cognitive errors we tend to make, and there are people who are analogous to the magicians who show the deficits and common sense of the AIs, right? They’re sort of playing games that are almost like magic tricks on the AIs. I wonder how revealing those will be, you know, in a serious scientific way.
Well, Melanie has a lot more to say about the depth of AI cognition and understanding, and also how it might change whole fields of science, including math. We will be hearing more about that after the break.
[Music plays]
STROGATZ: Welcome back to The Joy of Why. We’re joined today by Santa Fe Institute computer scientist Melanie Mitchell.
STROGATZ: You have been a college professor for much of your life. When you’re working with students they can get the answers right, but as you start to probe what they actually understand, you start to realize that they might be getting the right answers for the wrong reasons. They don’t really know what they’re doing, and that’s important if you wanna be a helpful teacher. This brings up another point: competence versus performance. Can you expand on this idea and, what would it mean in the AI context?
MITCHELL: So competence versus performance is kind of an old distinction from psychology and linguistics. The idea is that you might have the competence for a particular cognitive capacity, but there might be some reasons why you can’t perform the task that I’m giving you. Like they have the competence, they could solve the problems, but they’re just emotionally frozen. There’s some performance block.
But then there’s the other way around, which is performance without competence. So if the student in your office hours, say, had memorized a problem from the textbook and the solution, but they didn’t understand the general principle, so if you gave them a slightly different version of the problem, they couldn’t do it. That’s performance without competence.
STROGATZ: Okay. So if we would say that we’re trying to work out ways of testing whether the AI understands, what would count as evidence? Suppose that, you’re an AI advocate who said that these new systems, because we’ve scaled them up or because we have some nice new architecture with world models or social models or whatever, we’ve now crossed a threshold where they actually understand. It’s not just that they can compute, they understand. What would count as evidence of understanding?
MITCHELL: Oh gosh. I hate to get pedantic about understanding, but there’s so many different meanings of it.
STROGATZ: Ah.
MITCHELL: We had a talk here at Santa Fe Institute from a philosopher who broke down understanding into 25 different types.
STROGATZ: Aha. I didn’t know what I was getting myself into with the question.
MITCHELL: So there’s like P understanding and G understanding and there’s this very long typography of understanding. And I’m not sure there is any sort of single notion of real understanding. One of the recent things I and my collaborators have been working on is looking at different dimensions of understanding. One example is you can get one of these language models or chatbots to generate a story. Just generate a short story about something, and they will. They’ll generate a very beautiful little coherent short story. But then if you start asking them questions about the story, they will often will fail in weird ways.
STROGATZ: Hmm.
MITCHELL: even though they generated it. And I think the same thing is true in a lot of different tasks that they understand along one dimension but not along another dimension. And in some sense deep understanding might be just you understand across many different of these dimensions.
STROGATZ: Aha. That sounds like a promising direction. Let’s talk about tasks a little more, because that’s a phrase or a term that I’ve seen in some of your writing, the phrase, the tyranny of tasks. What’s that about?
MITCHELL: I first heard that, from Shannon Vallor, a philosopher. The idea is that in AI, the world is divided in terms of tasks. So when we think about what AI systems can do, people say, “Oh, they can make summaries. Let’s test their ability to summarize articles.” Or, “Let’s test their ability to answer questions about diagrams” or I don’t know, some other benchmark.
STROGATZ: Well, I mean, these days, they’ve been benchmarked a lot on International Mathematical Olympiad, very hard high school problems, then there were research level problems. Now there’s open problems that are unsolved in math. These are all like three levels of math benchmarks that are out there.
MITCHELL: Right, their capabilities are defined in terms of these benchmarks. You know, one benchmark might be the bar exam for law students, and they do really well on the bar exam. And so we say, “Oh, lawyers, you should be afraid. Your job is threatened because these AI systems are as getting as good as you are.” Uh, But the way that we’re defining that is by looking at how well they do on a specific set of questions or a task. And jobs as a whole are not the same as just one independent task after another. This is, I think it’s almost like a fallacy that if an AI system can do a bunch of tasks, it can do the job of a person that is associated with those tasks.
So just one example of this. So there’s a famous quote from Geoffrey Hinton, where he said something like, “AI systems are incredibly good at diagnosing or interpreting radiology images. Nobody should go to school anymore to be a radiologist. AI is gonna take all the jobs within five years.”
Well, that was 2016. That was 10 years ago. Now we actually have a shortage of radiologists. I don’t know if that’s because he said that, but uh, it turns out that even though AI systems can beat human doctors on these benchmarks, that’s not the same as doing this job out in the real world, which is much more open-ended, which is not just a series of well-defined tasks.
STROGATZ: Still, it does leave you wondering, like in the case of radiology, you could imagine if they are really good at that task, then what’s left for the human radiologist? Should we still be in that part of the game? Like in my own world of math, you know, if they’re very good at proving theorems, but they’re not so great yet at coming up with new concepts, or as we sometimes speak of it, theory building, right? There’s this big distinction between problem-solving and theory building. So is it that we’re sort of gonna find our niche, that we can do the parts that they don’t do? So like in the case of radiology, they have the open-ended part but not the scan reading part? I guess that’s what I’m wondering.
MITCHELL: Yeah.
STROGATZ: Is that how it’s gonna go?
MITCHELL: Maybe. I wouldn’t be at all surprised if jobs like yours change quite a bit because of these new tools. These are going to become incredibly useful tools for mathematicians. So it might change your job. Just like when personal computers came out, but there’s a fantastic book by um, George Lakoff and Rafael Núñez about math and where ideas in math come from, via metaphors. And they feel that human embodiment is a very important part of understanding and mathematics.
STROGATZ: Exactly. I think that’s our only hope ’cause right now they the machines don’t have great embodiment. And you’re right, that a lot of great ideas in math are inspired by experience with the world. And that’s what I was gonna say about applied math, that I feel like that’s even more so than pure math, where we get so much inspiration from nature and from engineering and society and all that, that I think we have a lot more chance of being useful as humans in applied math.
But I do think pure math will expire before applied math does, and maybe neither will. Maybe we’ll just keep going forever. What does it look like to you? I mean, math is often thought of as some kind of gold standard like, the AI companies have a lot of use for math, right? They can demonstrate how good their systems are ’cause they can verify that they’ve solved a problem or not.
MITCHELL: Well, that’s a big question I have, which is, suppose that your prediction comes right and math, pure math expires in some sense for humans. What does that mean for other fields? Does that mean that these machines are on their way to taking over everything? Or is it more like 1997 or whatever it was that Deep Blue beat Kasparov and that actually beating the best human at chess did not necessarily mean that was gonna go anywhere in other fields.
STROGATZ: I don’t know. What do you think? It feels to me like science is much more open-ended than math in that respect.
MITCHELL: Yeah, I believe that. I don’t think that solving all the Erdos problems means that the average person has to fear for their job.
STROGATZ: Okay, now we have many different things on the table at that point. But even just in the world of pure brainiacs, whether it’s scientists or mathematicians, just the fact that biology there are so many things to be measured, we have so much data that we could collect that we haven’t collected, so many new ways of observing. I mean, that seems very inexhaustible to me compared to math.
MITCHELL: I agree. And even in physics, I think, which is maybe closer to math, there’s so much you know, open-ended questions that aren’t well-formulated, that don’t have something like a proof that can be constructed.
STROGATZ: But so, I do feel like the hope for math is to continue to take inspiration from the real world. And von Neumann had said something like that too, that when math becomes too much art for art’s sake, when it drifts too far from the source, for him the source was nature or reality, if it becomes too far removed it becomes sterile, said von Neumann.
So I think this could be a a really good era for pure math if it starts taking more inspiration from nature. That’s been less so in the 20th and 21st century, but I think if we go back to that, we can probably eke out a few more centuries of human pleasure in math.
MITCHELL: I’ll just say there’s this dictum in AI which is that easy things are hard and hard things are easy.
STROGATZ: Right.
MITCHELL: And pure math is seen by humans as like the most exalted exhibition of intelligence and brilliance. It’s the hard thing, and yet we know that hard things are easier for machines and easier things are harder.
STROGATZ: Yep, and there’s the word soft also, right? In science, we talk about the hard sciences and the soft sciences, and the soft sciences of economics and psychology and anthropology, and those are the really hard ones.
MITCHELL: Right.
STROGATZ: Well, so if we meet again in five years.
MITCHELL: The Joy of Gamma, or something.
STROGATZ: Yes, The Joy of Omega by then, right. What do you hope we would understand about AI systems by then? Or what kinds of tests would we want to be able to do that we can’t do today?
MITCHELL: Yeah, I mean What I really hope will go well in the science of AI is this field called mechanistic interpretability, which is the neuroscience analog, where you’re actually looking at the activations and the weights and the, you know, all the messy innards of the system, and understanding at a higher-level sort of what they are doing.
These days, it’s kind of a smallish subfield where people are trying to develop tools that do that, analogous to things like fMRI or whatever. And I don’t think anybody’s really figured out exactly how to do this the right way yet, but I’m hoping that’s something that we can accomplish, and then we would have a genuine way of understanding sort of their limitations, what they can do, what they can’t do, what kinds of mistakes they’re likely to make, and maybe how to fix them.
STROGATZ: Interesting that you put your finger on that because the first time I became aware of you, it was in connection with that in a broad sense. So what I’m thinking of is back when you used to work on something that in the jargon was called GAs for CAs, genetic algorithms for cellular automata, you and Jim Crutchfield were looking at this problem of evolving algorithms that could solve a certain class of problems, hard computer science problems, and you were using this evolutionary algorithm to select better and better algorithms that kept improving through a kind of selection process.
But then the part that you did that I found so creative is once you’ve got a really good system, you looked at it in what felt to me like an analog of mechanistic interpretability. You tried to see what was making that system so smart, analyzing it in terms of particles that were colliding with each other according to certain rules in the diagrams. That’s, I don’t know if I’ve summarized it reasonably well, but it seems like this is a longstanding interest of yours.
MITCHELL: Yeah. that’s true. I hadn’t made that connection exactly, but that’s interesting.
STROGATZ: It is this, though. It’s interpretability.
It is interpretability. And it’s also, I think, in the field of complex systems, people talk about this notion of emergence.
STROGATZ: Yeah.
MITCHELL: And we thought of that as a kind of emergent computation. And I think these AI systems also have emergent computations that are not easy to find, but they’re there, and if we understood them better, we would understand how the system is actually working, doing what it does.
STROGATZ: Yeah, it’s an interesting attitude. It feels honestly to me very sweet and very old school. This hope that… Okay, you’re chuckling ’cause you see where I’m going. It’s a mean thing I’m saying, but this conceit that we with our limited minds can keep doing science, you know, and we’re gonna figure out how these AIs are doing what they’re doing, and that’s what our game will continue to be just like it always has been in science.
And I, the dark side of me, thinks our days are numbered to be able to do that as these gadgets get bigger and bigger. Who says we can keep doing science on them and figuring them out? What’s your reaction to that? We have nothing else to do. We have to try.
MITCHELL: That’s an interesting question. Um, why do we do science in the first place? I mean, you know, we do science ’cause we wanna solve problems. That’s one thing. But we also do science ’cause we’re driven to understand things.
STROGATZ: Yes.
MITCHELL: You see this in little children. They’re driven to understand. Often one of their first words is why. They ask it constantly. So I think that’s a human drive, and it’s hard to fight against that. And that’s why you and I both went into science, it’s important to us.
Now, I was a little despairing when I went to a panel discussion at a conference on the role of AI in science. And there were a bunch of famous people on the panel talking about how AI was going to revolutionize weather prediction, and genetics, and cosmology, and you name it. And I asked them at the end “Well, like, is this going to contribute to human understanding of the world?” And they’re like, “Why should we care about that?”
STROGATZ: Yeah. To me, this is the bifurcation that we’re all thinking about now. ’Cause science has this double-edged aspect, that it gives us pleasure, we like figuring things out, there is the joy of why, and as you say, it’s deep in our species. So yes, we’re curious, but then there’s the other side that for so long science has been this instrumental thing that helps us in technology and medicine.
And I guess the question I have, and I think a lot of us have, is will we continue to take pleasure in the joy of curiosity when we are no longer the best at solving the important problems? But let me ask you one last thing, for people who haven’t heard our earlier conversation, what was your draw to this field, and if you were starting out today, do you think you’d have the same kind of curiosity?
MITCHELL: Yeah, that’s a great question. When I was a child, I loved logic puzzles, like the knights and the knaves. The knights who always told the truth and the knaves who always lied. There’s a fun several books by Raymond Smullyan, a mathematician who wrote a bunch of puzzles in this genre that I absolutely loved.
When I got to college, I read Douglas Hofstadter’s book, Gödel, Escher, Bach, which was the real-world version of these in a way. I mean, he was talking about Gödel’s theorem and paradoxes in mathematical logic and how all this related to cognition and thinking and creativity and so on. And I was just completely blown away and that this is what I wanna do in my life. I didn’t exactly know what it was, but it seemed like it might be artificial intelligence. So I pursued Doug as an advisor and got to join his group, and was studying analogy via a new set of puzzles which were analogy puzzles. And, I was very entranced by all of that.
If I were that age today, I would be worried. In fact, I have a son who is getting a PhD in machine learning, and he wants to do research in machine learning, but he’s actually quite nervous that there will be no more roles for humans doing research in machine learning because AI will be doing all the research in machine learning and improving itself and so on and so forth. And I wonder if I’d think the same thing. I don’t know.
STROGATZ: Maybe we do have to revisit this in five years because we may know by then. Given how fast everything is going, who knows? I really appreciate your spending time with us. This has been wide-ranging, a little bit amorphous conversation, but it’s just wide open and I can’t think of a better guide to it. Thank you very much for joining us.
MITCHELL: Thanks, Steve. It’s been great.
[Music plays]
LEVIN: Hmm. Hmm. I, I just remember being a student and learning Newton’s laws for the first time, and then Kepler’s laws, which really make Newton’s laws beautiful, this application to the celestial cycles. I didn’t think, “Oh, I’m not the best at this, therefore I shouldn’t learn it.” Nor did I think, unless I one day become the best at this, I cannot feel pleasure or joy in my experience of acquiring this information.”
Of course, lots of people study things that other people already know and are better at. So I, I sort of wonder if maybe the AI will know things before us, but we will still need to acquire the understanding ourselves, and in that acquisition is a similar experience. Instead of maybe the AI will be a filter between us and interrogating nature directly, but we’ll still be acquiring, I don’t know, the knowledge and having that experience. I’m not sure. Maybe it’s all gonna pass us by.
STROGATZ: I– Well, let’s explore this a little more. I like especially your emphasis on not being the best, and how, in a way, unfraught that is. I, I learned as soon as I went to college what it means to not be the best. You know, this, this fixation with being the number one, especially in an age of optimization. There’s so many optimization algorithms. We talk about faster, cheaper. But in our own lives, very often we’re not the best. I’m certainly not the best tennis player. I love to play tennis. I’m not the best chess player, and I’m still happy to play chess. And try to be the best dad, but I may not be. But still, all these things are worth doing for their own sake, right? They give us pleasure.
I do feel very philosophical and almost religious about this. Like, we get a little time on Earth alive and, you know, these questions about AI do tap into questions about the meaning of life. What are we trying to do? If the meaning of life is that you’re gonna be the best in some domain or you’re gonna make a discovery that’s gonna change the world, then most people will have a meaningless life, and I just don’t wanna believe that’s the correct version of the meaning of life.
It was not for my dad. He didn’t even get to go to college. You know, he grew up in the Depression. That was not an option. His life was being a good parent and taking care of the people that bought shoes at the shoe store that he had. And he knew everyone’s shoe size in our little town, and he left a good name when he died. People remembered him well.
LEVIN: Right.
STROGATZ: So okay. What is that doing on our show here about science?
LEVIN: Well, I think that let’s say the meaning for some people of life has to do with acquisition, acquiring wealth. They’re gonna love this stuff, right? ’Cause there’s gonna be this new tool that simply leverages all kinds of buttons that they now have faster access to and can exploit and acquire more wealth.
There are people who found meaning in singing songs or writing poetry or being novelists or doing math, and, and I think all of those fields are a little more nervous, right? About reevaluating what the place is going to be for them and, and how to secure that place and how to think about it.
If I’m playing games of what may or may not happen, I mean, there is still a world in which AI is like a supercomputer, and we’ve talked about this before, Steve. Just ’cause a supercomputer can crunch all of these numbers, if it presents it to us as a string of symbols, even though it has, in some sense, an answer, it’s not a meaningful answer for us, and none of us value it.
We still, as human beings, have a very important role between us and a supercomputer rendering an image of a galaxy or looking at an image of a biomedical neural map. It hasn’t actually robbed scientists of their work. And so it might be that it really will continue to be a tool and not simply something that overtakes and discards us.
STROGATZ: Well, that’s the question, right? I think there are two plausible scenarios. One is that it continues to be a tool, and we always have some essential role in science and math at the cutting edge. The other option is, and actually in my heart I believe this is the case, that we will not be at the cutting edge, and that will happen very soon. And, so then what is the point?
Then I feel like it’s still meaningful, just like when I was in high school and I discovered things about math. They were discoveries to me. They were not discoveries to the world, you know? I think we may have to all settle for that. We’re not gonna be making genuine discoveries for the world.
The AIs will be doing that. I really do believe that’s gonna happen very soon. I may be wrong. I mean, there may be fundamental reasons why the AIs won’t be able to do that. For instance, they don’t have bodies, they don’t have social life, you know, there’s a lot… But I just think all that stuff will be solved before long. Anyway, what’s your take?
LEVIN: Well, I think there’s a difference between, making discoveries and understanding, and I guess that’s kind of what I mean in examples. In some sense, maybe the su- supercomputer made the discovery before the person did, but we still say the person did ’cause the discovery didn’t count as a discovery until they rendered it in a way that human beings could comprehend.
But, I honestly don’t know. I am not incredibly saddened or pessimistic, so I guess I would have to say that in my heart, intuitively, I am not terrified of this prospect. Maybe I should be, but maybe it’s just sort of a bliss of being naive and I’m just gonna wait for it to sneak up on me.
STROGATZ: There is one thing I think we can be very optimistic about and hopeful about, which is I think we’re gonna have a glorious golden age of science where we will understand, and discoveries by the AIs or by people in conjunction with AIs, that’s all gonna be happening in the next, whatever, five, 10, 15 years, and it’s gonna be a spectacular fireworks time for science. And I think that hopefully with any luck, we’ll be alive to see all that.
LEVIN: Yeah, there’s definitely going to be a transition period where people are moving it fast and furious, and they’re part of the story, and there’s great accomplishment, and it will be exciting to see. I know people, very accomplished, who are very excited about using it. Use it every day. They have multiple things going on, and they just feel like their productivity has doubled or more. And they’re excited, they’re enjoying themselves. I think there’s really nothing we can do but chime in and participate in this, at least, transition phase before we’re obsolete.
STROGATZ: Well, I’m getting choked up just thinking about it. Thanks, Janna. It’s always great to see you, and we’ll see you next time on The Joy of Why.
LEVIN: Thanks, Steve.
[Music plays]
LEVIN: If you’re enjoying The Joy of Why and you’re not already subscribed, hit the subscribe or follow button wherever you’re listening. You can also leave a review for the show. It helps people find this podcast. Find articles, newsletters, videos and more at quantamagazine.org.
STROGATZ: The Joy of Why is a podcast from Quanta Magazine, an editorially independent publication supported by the Simons Foundation. Funding decisions by the Simons Foundation have no influence on the selection of topics, guests, or other editorial decisions in this podcast or in Quanta Magazine. The Joy of Why is produced by PRX Productions.
The production team is Caitlin Faulds, Jade Abdul-Malik, Genevieve Sponsler, and Merritt Jacob. The executive producer of PRX Productions is Jocelyn Gonzales. Edwin Ochoa is our project manager.
From Quanta Magazine, Simon Frantz and Samir Patel provided editorial guidance, with support from Samuel Velasco, Kit Sudol, Simone Barr, and Michael Kanyongolo. Samir Patel is Quanta’s Editor-in-Chief.
The episode art is by Chanelle Nibbelink and our logo is by Jaki King and Kristina Armitage. Special thanks to Garth Avery at the Cornell Broadcast Studio.
I’m your host, Steve Strogatz. If you have any questions or comments, please email us at [email protected]. Thanks for listening.
[Music fades]

quanta

文章目录


    扫描二维码,在手机上阅读