我在一台能够即时学习的机器人身上,看到了人工智能的未来。

内容来源:https://www.wired.com/story/generalist-ai-robots-learn-like-clever-toddlers/
内容总结:
上周,我驱车仅15分钟,便目睹了一场令人瞠目结舌的机器人“智能秀”。在位于马萨诸塞州剑桥市的一家名为“通用人形智能”(Generalist AI)的初创公司办公室里,我看到机器臂们熟练地堆叠杯子、将积木块放入碗中,其反应之迅捷、决策之灵活,令人恍惚间以为面对的是血肉之躯的人类。
这些机器臂仅通过观看一段简短的演示视频,便掌握了多项任务,且无需针对特定任务进行专门训练,堪称惊艳。最令人称奇的一幕是,当机器人被要求用簸箕和刷子将积木块扫入碗中时,刷子却不见了踪影。只见它随机应变,直接用簸箕当刷子,轻轻一拨,将积木块准确送入碗中。
另一项演示中,一台双臂机器人观看了一段人类拉开钱包拉链并取出纸币的视频。随后,它面对一个截然不同的钱包,竟也熟练地拉开拉链,小心翼翼地取出纸币。更令人惊叹的是,当第一次抓取失败后,它果断将右夹爪换成左夹爪,以更佳角度再次尝试。一旁围观的技术人员不禁笑出声:“这个动作,它以前可没做过。”
“这正是当年GPT-3问世时,人们为之兴奋的那种能力。”该公司联合创始人兼CEO皮特·弗洛伦斯(Pete Florence)告诉笔者。他提到的GPT-3,是OpenAI于2020年发布的突破性大语言模型。“你可以直接给模型一个指令去完成新任务,它真有可能会做到。”
“通用人形智能”公司似乎专注于让机器人理解物理世界的运行规律,这一思路借鉴了人类幼年时期表现出的对物理法则的直觉认知。这种“物理直觉”可能正是模型能够将所学技能举一反三、灵活迁移的关键。事实上,公司内部的多次演示也让研究人员感到意外——比如,当一根香蕉出现在机器人面前时,它会顺势用它当“扫帚”来清扫物品。这看似小儿科,但“物理智能”恰恰是当前机器人的短板。人类婴幼儿如何高效地认知世界,或许能为人工智能研究提供重要启示。
笔者在公司会议室见到了弗洛伦斯和联合创始人兼CTO安德鲁·巴里(Andrew Barry)。窗外,工作人员正佩戴特制的夹爪手套进行机器人训练。该公司的另一位联合创始人、首席科学家是安迪·曾(Andy Zeng)。三人履历亮眼,均曾任职于Google DeepMind和波士顿动力等顶尖机构,深耕前沿机器人与AI模型领域。
传统上,训练一个AI机器人掌握新技能需要向其输入成千上万个样本,但这种方式弊端明显:哪怕只是改变一下现场光线,机器人就可能“犯迷糊”。如今,“通用人形智能”等多家初创公司正押注另一条路径——建立由人类训练出的“通用机器人模型”。该公司特别设计了一种类似机器人钳爪、带有摄像头的手套,供人类穿戴后执行各种日常操作。笔者看到,一个大木箱里堆满了数百只这样的手套,它们即将被运往墨西哥等地的数据采集人员手中。
关于具体训练方法,弗洛伦斯团队略显神秘,仅透露公司已积累海量高质量训练数据。与部分同业不同,他们选择完全从零构建AI模型,而非基于开源语言模型进行适配。
佐治亚理工学院的机器人学专家丹费·徐(Danfei Xu)对该公司颇为熟悉,他认为在众多追逐“通用机器人”的创业公司中,“通用人形智能”表现突出。“他们把这条路推到了极致,并且执行得极为出色。”徐评价道,该公司不仅数据质量高、获取量大,且团队本身科研功底扎实,“做出了真正扎实的科学。”
徐还指出,从公司目前的演示成果看,他们已经明确瞄准了真实的商业应用场景。“他们是目前最接近可落地部署的团队。”斯坦福大学机器人学教授凯伦·刘(Karen Liu)也对该公司的数据策略表示认可:“他们大规模采集物理交互数据,并不局限于某一种特定的机器人本体。从目前最亮眼的结果来看,这一押注似乎正在奏效。”
当然,该公司也坦言,其模型的学习能力尚未达到完全可靠的程度。目前,机器人对演示过任务的完成率平均仅为约59%,而理想的成功率应达到99%以上。此外,这些技能究竟能在多大程度上泛化至所有场景和任务,仍是未知数。
即便如此,机器人快速学习新技能的潜力,尤其在制造业等领域的应用前景,已相当可观。最近一个深夜,该公司一名工程师就亲身体验了这份惊喜。当时他正随手将小杯子码放在一台双臂机器人面前的桌上,想看它作何反应,不料机器人竟主动加入,伸出双爪,学着一起码放,最终将所有杯子整齐堆叠成一摞。视频记录了这一过程,工程师抑制不住兴奋,对着空气大声喝彩。
本文为威尔·奈特(Will Knight)AI实验室新闻简报中的一篇。
中文翻译:
上周,我斗胆从家出发,走了足足15分钟,去看机器人表演一些令人瞠目结舌、叹为观止的绝活。
我去了位于马萨诸塞州剑桥市的一家名为Generalist AI的初创公司办公室,在那里,我目睹了机械臂执行一些简单的家务活,比如叠杯子、把积木块放进碗里等等。它们掌握新事物的速度之快,令我震惊——让人联想到有血有肉的人。
这些机械臂在摄取了一小段教学视频后,就掌握了一系列任务,而且最令人印象深刻的是,它们没有针对特定任务进行过专门训练。其中最引人注目的例子之一,是一个机器人被指示用簸箕和刷子将积木块扫入碗中。当刷子从现场被拿走后,这个机器人随机应变,把簸箕当刷子用,将积木块弹进了碗里。
另一个案例中,一个双臂机器人观看了一段某人拉开手包拉链并取出纸币的视频剪辑。我看着机器人拉开一个不同款式手包的拉链,小心翼翼地取出纸币,不禁有些目瞪口呆。最神奇的是,当它抓不住钱币时,它会从使用右夹爪切换到使用左夹爪,以获得更好的攻击角度。“哈,”旁边一位工程师说道,“它以前从没这么干过。”
“这正是人们对GPT-3感到兴奋的那类事情,”Generalist的联合创始人兼首席执行官皮特·弗洛伦斯告诉我,他指的是OpenAI于2020年发布的突破性大语言模型。“你可以拿那个模型,直接提示它去执行一个新任务,它就有很大可能做到。”
Generalist似乎专注于教会其机器人理解世界的物理规律,这似乎受到了人类幼年时就展现出的直觉物理感的启发。这很可能有助于该模型将其在一个场景中学到的东西迁移到另一个场景的能力。事实上,公司的一些演示让我想到,当向孩子们展示一项任务时,他们是如何即兴发挥和尝试的。研究人员时常对机器人决定采取的行动感到惊讶——例如,当一根香蕉被放在它面前时,它选择用香蕉把物品扫到一起。这看起来可能微不足道,但物理智能仍然是机器在很大程度上所缺乏的,而婴儿高效地了解世界的方式,可能会为人工智能研究人员提供重要的见解。
我在一间会议室里见到了弗洛伦斯和联合创始人兼首席技术官安德鲁·巴里,会议室俯瞰着成队的工作人员,他们手上戴着特制的夹爪进行机器人训练。公司的另一位联合创始人兼首席科学家是安迪·曾。这三人背景出众:他们曾在Google DeepMind和波士顿动力公司,参与研发一些最先进的硬件和机器人模型。
传统上,训练一个由人工智能驱动的机器人执行不同任务,意味着要向模型输入数千个示例。不过,这是一种公认的不完美的学习方式,如果你改变一些简单的东西,比如光照条件,机器人就会在任务中遇到困难。
Generalist和其他一些机器人初创公司正在大力投资于一种由人类训练的通用机器人模型。该公司制造了类似机器人钳爪的特制手套,上面装有摄像头,人们戴上它们来执行各种家务。我看到一个板条箱里堆放着几百个这样的夹爪,它们将被送往墨西哥等地的工作人员手中。
弗洛伦斯和团队对他们训练机器人的具体方法守口如瓶,但他们表示,公司已经收集了大量高质量的训练数据。与其他一些追逐更智能机器人的公司相比,他们还完全从零开始构建了自家的人工智能模型,而不是依赖开源语言模型。
佐治亚理工学院的机器人专家徐丹飞熟悉Generalist的工作,他表示,这家初创公司在追逐更通用机器人模型的公司中脱颖而出。“他们把这推向了极致,并且执行得非常好,”徐说。除了收集大量高质量数据外,他说,“他们是出色的机器人专家,做了非常扎实的科学工作。”
徐还说,Generalist迄今展示的内容表明,他们着眼于在真实的商业环境中部署机器人。“他们是最接近可部署状态的,”他说。
“Generalist的数据方法,是大规模收集物理交互数据,而不将其与某一个特定的机器人绑定得太紧,”斯坦福大学的机器人专家凯伦·刘说,她也认识这家公司。“他们最有力的结果表明,这个赌注可能正在奏效。”
话虽如此,Generalist也表示,其模型的学习技能还不那么可靠。一个机器人平均只有大约59%的时候能完成向它展示过的任务;理想情况下,其成功率应该在99%以上。此外,这些技能究竟能在多大程度上泛化到所有能想到的任务或场景,目前也似乎还不清楚。
即便如此,机器人在制造业等领域快速学习技能的潜力似乎巨大。最近一个深夜,Generalist的一位工程师似乎就发现了这一点。一段记录了这一事件的视频显示,这位工程师在双臂机器人面前的桌子上叠放小杯子,只是想看看机器可能会做什么。机器人突然加入进来,用它的两个夹爪抓起并堆叠其他杯子。当机器人把所有杯子整齐地堆成一摞时,这位工程师开始对着空气欢呼,对这一操作感到欣喜不已。
这是威尔·奈特《AI实验室》新闻通讯的一期。点击此处阅读往期通讯。
评论
返回顶部
英文来源:
Last week, I ventured a whopping 15 minutes from my house to see robots do some mind-boggling, jaw-dropping stuff.
I visited the Cambridge, Massachusetts, offices of a startup called Generalist AI, where I watched robot arms perform simple chores like stacking cups, putting blocks into bowls, and the like. I was astonished by how quickly they figured things out—it was reminiscent of a flesh-and-blood person.
The arms mastered a range of tasks after ingesting a short, instructional video and, most impressively, no specific training for a given task. One of the most striking examples involved a robot that was instructed to sweep a block into a bowl using a dustpan and brush. When the brush was removed from the scene, the robot improvised by using the dustpan like a brush and flicking the block into the bowl.
In another case, a two-armed robot watched a videoclip of someone unzipping a purse before removing some banknotes. I watched—somewhat slack-jawed—as the robot unzipped a different kind of purse and carefully removed the notes. Most amazingly, when it couldn’t grab the money, it switched from using its right gripper to its left to get a better angle of attack. “Ha,” said one engineer standing nearby. “It never did that before.”
“This is exactly the kind of thing people were really excited about with GPT-3,” Generalist cofounder and CEO Pete Florence told me, in reference to OpenAI’s breakthrough large language model, released in 2020. “You could take that model and just prompt it to do a new task and it would have a real shot at doing it.”
Generalist appears to be focused on teaching its robots about the physics of the world, which seems inspired by the intuitive sense of physics humans exhibit from an early age. That may well contribute to the model’s ability to transfer what it has learned in one scenario to another. In fact, some of the company’s demos made me think of how children improvise and experiment when shown a task. The researchers have often been surprised by what the robot decides to do—one chose to sweep up items with a banana when it was placed in front of it, for example. This might seem trivial, but physical intelligence is something still largely lacking in machines, and the way babies learn so efficiently about their world may offer important insights for AI researchers.
I met Florence and Andrew Barry, cofounder and CTO, in a conference room overlooking teams of people doing robot training with special grippers on their hands. The company’s other cofounder and chief scientist is Andy Zeng. The trio have impressive backgrounds: They previously worked at Google DeepMind and Boston Dynamics on some of the most advanced hardware and robotic models around.
Traditionally, training an AI-powered robot to do different tasks has meant feeding thousands of examples into the model. This is a notoriously imperfect kind of learning, though, and a robot will struggle with the task if you change something as simple as the lighting.
Generalist and some other robotics startups are investing heavily in a general robotic model trained by humans. The company builds special gloves resembling robot pincers that have cameras attached to them, which people then use to perform different chores. I saw a crate piled high with several hundred of these grippers destined for workers in Mexico and elsewhere.
Florence and team are cagey about exactly what recipe they’re using to train the robots, but they say the company has already gathered a huge amount of high-quality training data. In contrast to some other companies chasing smarter robots, they have also built their AI models entirely from scratch rather than relying on an open-source language model.
Danfei Xu, a roboticist at Georgia Tech who is familiar with Generalist’s work, says that the startup stands out among companies chasing more general robot models. “They have pushed this to the extreme, and they’ve done a really good job executing,” Xu says. Besides gathering a huge amount of high-quality data, he says, “they are excellent roboticists, and they have done really good science.”
Xu also says that the stuff Generalist has demo’d so far suggests that they have an eye on deploying robots in real commercial settings. “They are the closest to something that's deployable,” he says.
“Generalist's data approach is collecting physical interaction data at large scale without tying it too closely to one particular robot,” says Karen Liu, a roboticist at Stanford University who also knows the company. “Their strongest results suggest that this bet may be working.”
That said, Generalist says the learning skills of its models are not yet all that reliable. A robot is only able to complete a task it has been shown about 59 percent of the time, on average; ideally, its success rate would be somewhere upwards of 99 percent. It also seems unclear how well these skills will generalize to every imaginable task or setting.
Even so, the potential for robots to quickly learn skills in, say, manufacturing seems huge. One of Generalist’s engineers seemed to discover this late one recent evening. A video that captured the episode shows the engineer stacking small cups on the table in front of a two-armed robot, just to see what the machine might do. The robot suddenly joined in, grabbing and stacking other cups with its two grippers. As the robot finished stacking the cups into one neat pile, the engineer began yelling to no one in particular, delighted by the maneuver.
This is an edition of Will Knight’s AI Lab newsletter. Read previous newsletters here.
Comments
Back to top
文章标题:我在一台能够即时学习的机器人身上,看到了人工智能的未来。
文章链接:https://news.qimuai.cn/?post=4848
本站文章均为原创,未经授权请勿用于任何商业用途