这位人工智能创业者正在开发能够提前规划、应对突发情况的智能体。

内容来源:https://www.technologyreview.com/2026/09/08/1142088/danijar-hafner-developing-plan-ahead-agents/
内容总结:
这位企业家正在开发能提前规划的智能体
丹尼哈·哈夫纳的初创公司处于隐身模式,但他在教AI智能体理解我们这个世界方面有着长期而丰富的履历。
丹尼哈·哈夫纳位于旧金山SoMa区的办公室几乎空无一物。他的全新初创公司仍处于隐身模式,甚至连门上都还没有挂上公司名称。我去拜访那天,办公室里只有另外一个人,家具也寥寥无几。但它在装饰上的欠缺,却由机器人弥补了回来。各种大小和形态各异的人形机器人像提线木偶一样悬挂在贯穿这片开阔空间中央的架子上。
31岁的哈夫纳目前还不愿过多透露他的新项目,但他形容这是自己长期工作的延续——让AI能够在其训练中未曾遇到的环境中自如行动。他从中国进口的这些人形机器人,正是这项工作的下一步演进,也是其物理形态的体现。它们在未曾测试过的场景中做出反应的能力,将是让机器人进入人类生活空间的关键。因为如果你想派一个机器人进入某人的家中,它就需要能够应对从未见过的户型布局和家具陈设。
为了实现这一目标,哈夫纳依赖于一种名为“基于模型的强化学习”的技术。他开发世界模型——旨在模拟物理现实的AI模型——并在其中训练智能体。智能体本质上将模型视为真实世界的模拟环境,并学习如何在其中行动。随后,它利用这些经验对未来结果做出预测(哈夫纳可能会称之为“做梦”或“想象”)。这使得智能体——或者说搭载它们的机器人——能够在现实世界中应对陌生情境。
与其他努力不同,哈夫纳的技术让智能体及其控制的机器人能够执行极其复杂的任务,而无需依赖机器人领域传统上使用的真实世界试错训练。
哈夫纳在德国东北部的一个乡村小镇长大,父母都是古典音乐家。他从一位邻居那里学会了编程,高中时开始选修AI在线课程,这很快发展为他的热情所在。“我一直对思维如何运作深感着迷,”他说。AI为他提供了一种在计算机上模拟思维的方式。
2015年,作为波茨坦哈索·普拉特纳工程学院的一名大二本科生,他获得了在谷歌大脑担任学生研究员的机会。此后,他在该公司历任十余个实习和其他职位,包括在谷歌大脑和谷歌深度思维(两者现已合并为深度思维)工作,足迹遍及英国、加拿大和美国。他曾与业界传奇人物共事,包括常被称为“AI教父”之一的杰弗里·辛顿,以及开创性论文《注意力就是一切》的合著者阿希什·瓦斯瓦尼——该论文描述了当今大语言模型所使用的Transformer技术。
哈夫纳在谷歌的前经理兼合著者蒂莫西·利利克拉普形容他是“佼佼者中的佼佼者”。“在谷歌研究部门,我有机会与许多非常聪明的人打交道,而他轻松跻身顶尖1%之列,”利利克拉普说。“在很多情况下,他能单枪匹马完成通常需要整个工程师团队才能构建的东西。”
多年来,哈夫纳通过让在其世界模型中训练的智能体挑战热门电子游戏,不断打磨并验证了自己的方法。他的第一个突破是PlaNet,一个允许智能体通过提前规划来执行动作的模型。他的Dreamer 2是首个仅凭世界模型就在Atari 2600游戏中达到人类水平表现的智能体。Dreamer 3是第一个独自完成《我的世界》钻石挑战——成功在游戏中开采钻石——的智能体。而Dreamer 4则更进一步,通过从录制的游戏视频离线数据集中学习开采钻石,完全不与游戏直接互动。
最近,他开始将智能体从虚拟世界迁移到物理现实中。他的DayDreamer项目使用Dreamer算法,让机器人在全新环境中自主操作,并对新的体验(例如被推倒)做出反应,而无需任何特定训练。
如今,哈夫纳正致力于他的新初创公司——他于2025年秋天离开谷歌深度思维创立了这家公司。尽管他对下一步计划讳莫如深,但显然他志存高远:“我感兴趣的是解决一个问题,”他暗示道,“一个会改变世界的问题。”
深度专题
人工智能
一个根本性缺陷使大型语言模型极易受到攻击
这让人们可以轻易诱骗它们做不该做的事,比如告诉你如何破坏飞机的导航系统。
Anthropic发现了一个隐藏空间,Claude在那里思考概念
一项新技术让该公司比以往任何时候都更深入地探索大语言模型的奇异运作机制。
AI在招聘时比人类更容易形成偏见
AI不仅从训练数据中学习刻板印象——它还能自己“发明”新的偏见。
AI智能体为什么会为了达成目标而撒谎和作弊
这种不当行为被称为“奖励黑客”。这是你需要了解的内容。
保持联系
获取麻省理工科技评论的最新资讯
发现特别优惠、热门报道、即将举办的活动等更多内容。
中文翻译:
这位创业者正在开发能提前规划的智能体
丹尼贾尔·哈夫纳的初创公司尚处于隐身模式,但他教授AI智能体认识我们这个世界已有很长一段历史。
丹尼贾尔·哈夫纳在旧金山SoMa区的办公室基本空着。他全新的初创公司仍处于隐身模式,门上连公司名字都没有。我去拜访那天,那里只有另外一个人,家具也寥寥无几。但装饰上的欠缺,机器人全都补回来了。各种形状和大小的仿人机器人像牵线木偶一样悬挂在贯穿这个开阔空间中央的架子上。
31岁的哈夫纳对自己的新事业还不愿多说,但他将其描述为自己多年工作的延续——让AI能够应对训练中未遇到过的环境。他从中国进口的这些仿人机器人,是这项工作的下一步演进——也是它的物理实体形态。它们在未经测试的场景中做出反应的能力,将是让机器人进入人类生活空间的关键。因为如果你想派一个机器人进入某人的家中,举例来说,它需要能够应对从未见过的户型和家具。
为了实现这一目标,哈夫纳依赖于一种叫做基于模型的强化学习的技术。他开发世界模型——旨在模拟物理现实的AI模型——并在其中训练智能体。智能体本质上将模型视为真实世界模拟器,并学习如何在其中行动。然后它利用这些经验对未来的结果做出预测(哈夫纳可能会称之为做梦或想象)。这使得智能体——或者说它们所嵌入的机器人——能够在现实世界中应对不熟悉的情境。
“我在Google能接触到很多非常聪明的研究人员,而他轻松跻身最顶尖的1%之列。”
——蒂莫西·利利克拉普,Google DeepMind
与其他方法不同,哈夫纳的技术使智能体及其控制的机器人能够执行极其复杂的任务,而无需传统机器人学中常用的真实世界试错训练。
哈夫纳在德国东北部的一个乡村小镇长大,父母都是古典音乐家。他从一位邻居那里学会了编程,高中时开始上关于AI的在线课程,这很快发展成了一项热情。“我一直对思维是如何运作的感到着迷,”他说。AI为他在计算机上模拟思维提供了一种方式。
2015年,作为波茨坦哈索·普拉特纳学院工程专业的二年级本科生,他获得了Google Brain学生研究员的职位。从那时起,他在该公司积累了十几次实习和其他职位经历,包括在Google Brain和Google DeepMind(两者后来已合并为DeepMind)位于英国、加拿大和美国的工作。他与业界传奇人物合作过,包括经常被称为AI教父之一的杰弗里·辛顿,以及开创性研究论文《注意力就是一切》的合著者阿希什·瓦斯瓦尼——该论文描述了当今大语言模型所使用的Transformer技术。
哈夫纳在Google的一位前经理兼合著者蒂莫西·利利克拉普将他描述为精英中的精英。“我在Google研究部门能接触到很多非常聪明的人,而他轻松跻身最顶尖的1%之列,”利利克拉普说。“在很多情况下,他单枪匹马就能做出通常需要整个工程师团队才能做出的东西。”
多年来,哈夫纳通过让他的世界模型中训练的智能体挑战热门电子游戏,不断完善并证明了他的方法。他的第一个突破是PlaNet,一个让智能体能够通过提前规划来执行动作的模型。他的Dreamer 2是第一个使用世界模型在玩雅达利2600游戏时达到人类水平的智能体。Dreamer 3是第一个解决《我的世界》钻石挑战的智能体——独立成功挖掘游戏中的宝石。而Dreamer 4则更进一步,通过从录制的游戏视频的离线数据集中学习挖钻石,而从未直接与游戏互动。
最近,他开始将智能体从虚拟世界迁移到物理现实中。他的DayDreamer项目使用Dreamer算法让机器人在新环境中自主运行,并对新的体验(比如被推倒)做出反应,而无需任何特定训练。
如今,哈夫纳正致力于他的新初创公司——他于2025年秋天离开Google DeepMind创办了这家公司。虽然他对下一步计划守口如瓶,但很明显他有着远大的梦想:“我感兴趣的是一直在解决的问题,”他暗示道,“那将改变世界。”
深度探索
人工智能
一个根本性缺陷使大语言模型极容易受到攻击
这使得人们可以轻易诱骗它们做不该做的事,比如告诉你如何破坏飞机的导航系统。
Anthropic发现了一个隐藏空间,Claude在其中思考概念
一项新技术让该公司比以往任何时候都更深入地探究大语言模型的奇异运作方式。
AI在招聘时比人类更容易形成偏见
AI不仅从训练数据中学习刻板印象——它还能创造新的刻板印象。
这就是AI智能体撒谎和欺骗以达成目标的原因
这种不当行为被称为奖励黑客。这是你需要知道的。
保持联系
获取来自
《麻省理工科技评论》的最新资讯
发现特别优惠、头条新闻、即将举办的活动等更多内容。
英文来源:
This entrepreneur is developing agents that can plan ahead
Danijar Hafner has a startup in stealth and a long track record of teaching AI agents about our world.
Danijar Hafner’s office in San Francisco’s SoMa district sits mostly empty. His brand-new startup is still in stealth mode and doesn’t even have its name on the door. On the day I visit, there’s only one other person there, and little in the way of furniture. But what it lacks in decor, it makes up for in robots. Humanoids of various shapes and sizes hang like marionettes from racks that run down the center of the wide-open space.
While Hafner, 31, won’t say too much about his new venture just yet, he describes it as a continuation of his longtime work to enable AI to navigate environments it has not encountered in training. The humanoids, which he imports from China, are the next evolution of this work—and its physical embodiment. Their ability to react in previously untested scenarios will be key to getting robots into human spaces. Because if you want to send a robot into a person’s home, for example, it needs to be able to handle a floor plan and furniture it’s never seen before.
To achieve this, Hafner relies on something called model-based reinforcement learning. He develops world models—AI models designed to emulate physical reality—and trains agents within them. The agent essentially treats the model as a real-world simulation and learns how to act there. It then uses those experiences to make predictions (to dream or imagine, Hafner might say) about future outcomes. That allows agents—or the robots they’re embedded in—to navigate unfamiliar situations IRL.
“I get to interact with a lot of really smart people in research at Google, and he easily sits in the top half of 1%.”
Timothy Lillicrap, Google DeepMind
Unlike other efforts, Hafner’s technique enables agents and the robots they control to execute massively complicated tasks without the real-world trial-and-error training that’s traditionally been used in robotics.
Hafner grew up in a rural town in northeastern Germany, where his parents were both classical musicians. He learned programming from a neighbor, and in high school he began taking online courses about AI, which quickly developed into a passion. “I was always fascinated with how thinking works,” he says. AI offered him a way to emulate it on a computer.
In 2015, as a second-year undergraduate studying engineering at Hasso Plattner Institute in Potsdam, he won a role as a student researcher at Google Brain. From there, he went on to a dozen internships and other positions at the company, including stints with Google Brain and Google DeepMind (the two have since merged under DeepMind) in the UK, Canada, and the US. He worked with industry legends including Geoffrey Hinton, who is often referred to as one of the godfathers of AI, and Ashish Vaswani, coauthor of the groundbreaking research paper “Attention Is All You Need,” which described the transformer technology used by today’s large language models.
One of Hafner’s former managers and coauthors at Google, Timothy Lillicrap, describes him as a standout among standouts. “I get to interact with a lot of really smart people in research at Google, and he easily sits in the top half of 1%,” Lillicrap says. “In many cases he would build, single-handedly, things it would take entire teams of engineers to build.”
Over the years, Hafner has honed and proved his approach by pitting agents trained within his world models against popular video games. His first breakthrough was PlaNet, a model that allowed agents to execute actions by planning ahead. His Dreamer 2 was the first agent to hit human-level performance playing Atari 2600 games using a world model. Dreamer 3 was the first one to solve the Minecraft Diamond challenge—successfully mining in-game gems on its own. And Dreamer 4 went a step beyond that by learning to mine diamonds from an offline data set of recorded game-play videos, without ever interacting with the game directly.
More recently, he’s begun to migrate his agents out of the virtual world and into physical reality. His DayDreamer project used the Dreamer algorithm to let robots operate themselves in novel environments and react to new experiences (such as being pushed over) without any specific training.
Today, Hafner is working on his new startup, which he left Google DeepMind to form in the fall of 2025. Though he’s coy about his next steps, it’s clear he’s dreaming big: “I was interested in solving a problem,” he hints, “that would change the world.”
Deep Dive
Artificial intelligence
A fundamental flaw leaves LLMs strikingly vulnerable to attack
It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system.
Anthropic found a hidden space where Claude puzzles over concepts
A new technique has let the company probe deeper than ever into the weird workings of an LLM.
AI is more likely than humans to form biases when hiring
AI doesn’t just learn stereotypes from its training. It can cook up new ones, too.
Here’s why AI agents lie and cheat to reach their goals
The misbehavior is called reward hacking. This is what you need to know.
Stay connected
Get the latest updates from
MIT Technology Review
Discover special offers, top stories, upcoming events, and more.
文章标题:这位人工智能创业者正在开发能够提前规划、应对突发情况的智能体。
文章链接:https://news.qimuai.cn/?post=5009
本站文章均为原创,未经授权请勿用于任何商业用途