AgentHands:在XR中为空间具身智能体对话生成交互式手部手势

内容总结:
谷歌发布AgentHands原型系统,推动XR空间交互实现具身化AI对话
2026年8月25日,谷歌X研究团队推出了一项名为AgentHands的创新研究原型,该成果发表于2026年计算机人机交互大会(CHI 2026)。AgentHands是一套由大语言模型驱动的扩展现实(XR)交互系统,通过为AI对话代理赋予同步、表达自然的手势动作,在物理空间中提供具象化的操作指引,有效弥合用户心理认知差距,显著提升实体任务中的用户参与度。
随着AI助手从纯文本界面迈向多模态交互,行业正经历从被动应答到主动情境感知的转型。项目Astra及Gemini 3.1 Flash Live等最新技术已允许用户实时讨论周围环境,通常借助视觉边界框叠加在摄像头画面中识别物体。然而,此类二维屏幕方案在过渡至Android XR等沉浸式平台时面临独特挑战:如何突破平面界面限制,构建真正具备空间感知的具身对话体验?
为破解这一难题,AgentHands将共语言手势(co-speech gestures)引入三维世界。人类交流中,手势不仅用于指代,更能描述形状、模拟动作、强调重点,并与语音保持同步。AgentHands依托XR空间理解能力,复现这一自然协同效应。该研究延续了团队在Human I/O与Sensible Agent方面的前期积累,进一步为AI代理配备富有表现力的同步手势,将抽象的口头指令转化为直觉化的实体演示,使用户与周边环境的对话更加自然生动。
在研究初期,团队与谷歌XR及人机交互领域的专家开展了形成性研究,明确虚拟手部在三维环境中的“可读性”标准,并据此提炼出多维度分类体系,定义代理如何在用户物理空间中通过手势实现对话锚定。
AgentHands的核心创新在于将大语言模型的高层推理映射为精确、实时的物理动作,使手势与代理“语音”及用户XR环境高度匹配。其工作流程包含以下关键环节:
系统启动轻量级物体注册模块,借助眼动追踪与场景重建,用户可快速“标记”物品(如兰花或笔记本电脑),建立带有三维边界框的空间注册表供代理引用。
团队构建了覆盖三类语义的手势行为库:指向型(deictic)用于指示参照,象似型(iconic)用于描绘动作或形态,表达型(expression)用于传递社交信号与情感。
用户提问时,后端大语言模型生成响应并嵌入内联手势事件(GestureEvents)。每个事件绑定特定触发词,按分类维度编码手势行为原语。
XR头显上的本地解析器协调语音合成与动画引擎,利用词级时间戳确保代理手部动作与语音输出完美同步,实现清晰、富有表现力的空间参照。
通过模块整合,AgentHands在语言意图与物理动作间构建无缝桥梁,将标准大语言模型输出转化为丰富的多模态表现,使代理的响应同时通过语音与空间精确动作呈现,复杂指令得以在用户环境中的准确位置进行演示。
研究团队展示了具身手势结合XR空间感知如何增强对物理环境的理解,典型应用场景包括:
互动教学:在兰花养护场景中,代理不仅说“检查根系”,还会将手移至植物基部,勾勒气根轮廓并同步讲解其功能。
技术指导:针对3D打印机操作,代理可演示精确的“旋转并点击”序列,指导用户操作控制旋钮与文件选择,使复杂物理界面步骤直观易懂。
生活陪伴:代理可作为健康教练,与用户互动。例如,做出“警示”手势并辅以视觉效果,提示用户避免不健康行为。
为评估手势效果,团队开展了一项受试者内设计研究(N=12),将AgentHands与纯语音基线进行对比。两种条件均使用相同的研究员脚本化语言内容,唯一变量为具身手势及同步动作。参与者完成涵盖日常照料与技术操作的两项程序性任务。结果表明,XR与共语言手势的结合在空间锚定交互中极为有效,多项沟通效率关键指标均获验证。
AgentHands标志着AI系统正从单纯分析世界迈向动态融入世界。通过共语言手势与XR空间能力将对话锚定于物理运动,系统可显著降低复杂任务的认知负荷,推动空间计算向更易用、更人性化的方向发展。
在持续拓展Android XR生态的过程中,团队正探索手势个性化功能,如适配用户惯用手或学习特定空间习惯,以构建更为无缝的人机协作体验。
该研究主要由Ziyi Liu在谷歌学生研究员任职期间主导,系多团队联合合作成果。团队感谢David Li、Zhongyi Zhou与David Kim的关键贡献,以及Adarsh Kowdle、Guru Somadder和Shahram Izadi的战略指导与细致评审。
中文翻译:
2026年8月25日
荀千,研究科学家,杜若飞,Google XR 交互感知与图形负责人
AgentHands 是一个由大语言模型驱动的 XR 原型,它为对话式智能体赋予同步且富有表现力的手势,以提供基于空间定位的引导,弥合心理映射的鸿沟,并增强用户在实体任务中的参与感。
随着人工智能助手从简单的文本界面演变为多模态伴侣,我们看到它正在转向更具主动性、更具情境感知的辅助方式。像 Project Astra 和 Gemini 3.1 Flash Live 这样的最新创新已经让用户能够实时讨论周围环境,通常利用视觉边界框叠加层来识别摄像头画面中的物体。虽然这些叠加层在二维屏幕上非常有效,但向 Android XR 等沉浸式平台的过渡带来了一个独特的挑战:我们如何超越平面界面,创造出一种真正具身化、具备空间感知的对话?
为了弥合这一差距,我们推出了 AgentHands(发表于 CHI 2026),这是一个研究原型,将伴随言语手势的力量带入三维世界。在人类交流中,我们的双手所做的远不止指指点点;它们描述形状、模仿动作、强调要点,并且全部与我们的声音同步。通过利用扩展现实(XR)的空间理解能力,AgentHands 复现了这种自然的协同效应。继我们之前在 Human I/O 和 Sensible Agent 方面的研究之后,AgentHands 进一步为 AI 智能体配备了富有表现力且同步的手势,将抽象的言语指令转化为直观的实体演示,让你与周围环境的对话变得更加自然和引人入胜。
首先,我们与 Google 的 XR 和人机交互(HCI)专家进行了一项形成性研究,以确定是什么让虚拟手在三维环境中“清晰可辨”。我们将这些见解提炼为一个多维分类体系,该体系定义了智能体应如何利用双手在用户物理空间中将对话锚定于具体情境。
AgentHands 的核心创新在于它能够将大语言模型的高层推理映射为精确的、实时的物理动作,与智能体的“声音”和用户的 XR 环境相匹配。我们引入了以下关键步骤来构建 AgentHands 的工作流程。
该系统首先从一个轻量级的物体注册模块开始。利用视线追踪和场景重建,用户可以快速“标记”物品——比如一株兰花或一台笔记本电脑——创建一个附带三维边界框的空间注册表,供智能体参考。
我们构建了一个手势行为库,涵盖三个语义类别:a)指示性手势,用于指代;b)象似性手势,用于描绘动作或形态;c)表达性手势,用于传达社交线索和情感。
当用户提问时,后端大语言模型会生成一个响应,其中包含内联的 GestureEvents(手势事件)。每个事件都附着在特定的触发词上,并根据分类体系的维度编码手势行为的基元。
XR 头显上的本地解析器将文本转语音(TTS)播放与动画引擎协调起来。通过利用词级时间戳,智能体的双手能够与说出的话语完美同步地执行伴随言语手势,提供清晰、富有表现力的空间参考。
通过整合这些模块,AgentHands 在语言意图与物理动作之间创建了一座无缝的桥梁。该系统将标准的大语言模型输出转化为丰富的多模态表演,智能体生成的响应通过语音和空间精确的动作同时呈现,使复杂的指令能够在用户环境中准确对应的位置得到演示。
我们展示了这些具身手势与 XR 的空间感知相结合,如何增强我们对物理环境的理解。
互动式教学:在兰花养护场景中,智能体不只是说“检查根部”;它会将手移到植物基部,在勾勒气生根轮廓的同时解释其功能。
技术操作讲解:对于 3D 打印机操作,智能体可以演示操作控制旋钮和选择文件所需的精确“旋转并点击”顺序,让复杂的物理界面步骤变得直观易懂。
生活方式陪伴:智能体可以充当健康教练,与你的身体选择进行互动。例如,智能体可以做出一个互动式的“警告”手势,握住用户的手并配合视觉效果,提醒用户注意不健康的行为。
为了评估这些手势的影响,我们进行了一项被试内研究(N = 12),将 AgentHands 与仅语音的对照组进行比较。两种条件使用相同的由研究人员预先编写的口头内容,确保唯一的差异在于是否具备具身化的双手及其同步手势。参与者完成了两项程序性任务,兼顾了日常养护与技术操作。
结果证实,XR 与伴随言语手势的结合对于空间锚定的交互极为有效。我们从多个沟通效能的关键指标维度对数据进行了分析。
AgentHands 代表了迈向未来的一步,在那个未来中,AI 系统不仅分析我们的世界,还能在其中动态地运作。通过利用伴随言语手势和 XR 的空间能力,将对话锚定于身体动作之中,我们可以降低复杂任务的认知负荷,让空间计算变得更易用、更以人为本。
随着我们继续为 Android XR 生态系统进行开发,我们正在探索如何让这些手势更加个性化,适应用户的惯用手或学习他们的特定空间习惯,以创造出更加无缝的人机协作体验。
这项研究主要由刘子毅在 Google 学生研究员任职期间主导完成,是多团队联合协作的一部分。我们衷心感谢主要贡献者 David Li、周中一和 David Kim 的支持,以及 Adarsh Kowdle、Guru Somadder 和 Shahram Izadi 的战略指导和细致审阅。
英文来源:
August 25, 2026
Xun Qian, Research Scientist, and Ruofei Du, Interactive Perception & Graphics Lead, Google XR
AgentHands is an LLM-powered XR prototype that augments conversational agents with synchronized, expressive hand gestures to provide spatially grounded guidance, bridging the mental mapping gap and enhancing user engagement in physical tasks.
As AI assistants evolve from simple text interfaces to multimodal companions, we are seeing a shift toward more proactive, situated assistance. Recent innovations like Project Astra and Gemini 3.1 Flash Live already allow users to discuss their physical surroundings in real time, often utilizing visual bounding box overlays to identify objects in a camera feed. While these overlays are highly effective for 2D screens, the transition to immersive platforms like Android XR presents a unique challenge: how do we move beyond flat UI to create a truly embodied, spatially aware dialogue?
To bridge this gap, we introduce AgentHands, published at CHI 2026, a research prototype that brings the power of co-speech gestures to the 3D world. In human communication, our hands do more than just point; they describe shapes, mimic actions, and emphasize points, all synchronized with our voice. By leveraging the spatial understanding capabilities of Extended Reality (XR), AgentHands replicates this natural synergy. Following up our prior research in Human I/O and Sensible Agent, AgentHands further equips AI agents with expressive, synchronized hand gestures that transform abstract verbal instructions into intuitive, physical demonstrations, making conversations about your surroundings more natural and engaging.
To start, we conducted a formative study with XR and human–computer interaction (HCI) experts at Google to determine what makes a virtual hand “legible” in a 3D environment. We distilled these insights into a multi-dimensional taxonomy that defines how an agent should use its hands to ground a conversation within a user's physical space.
The core innovation of AgentHands is its ability to map the high-level reasoning of LLMs into precise, real-time physical motions that match the agent's “voice” and the user's XR environment. We introduce the following key steps to compose the AgentHands workflow.
The system begins with a lightweight object registration module. Using eye gaze and scene reconstruction, users can quickly “tag” items — like an orchid or a laptop — creating a spatial registry with 3D bounding boxes that the agent can reference.
We constructed a library of hand gesture behaviors across three semantic categories: a) deictic for referencing, b) iconic for depicting actions or forms, and c) expression for conveying social cues and emotion.
When a user asks a question, the backend LLM generates a response that includes inline GestureEvents. Each event is attached to specific trigger words and encodes the primitives for a hand behavior following the taxonomy dimensions.
A local parser on the XR headset coordinates the text-to-speech (TTS) playback with the animation engine. By using word-level timestamps, the agent’s hands perform co-speech gestures in perfect sync with the spoken words, providing clear, expressive spatial references.
By integrating these modules, AgentHands creates a seamless bridge between linguistic intent and physical action. The system transforms a standard LLM output into a rich, multimodal performance where the agent's generated responses are manifested through both speech and spatially accurate movement, allowing for complex instructions to be demonstrated exactly where they occur in the user's environment.
We demonstrated how these embodied gestures, paired with the spatial awareness of XR, enhance our understanding of our physical surroundings.
Interactive tutoring: In an orchid-care scenario, the agent doesn’t just say “check the roots”; it moves its hands to the base of the plant and outlines the air roots while explaining their function.
Technical walkthroughs: For 3D printer operations, the agent can demonstrate the exact ''turn and click'' sequence needed to navigate control knobs and select files, making complex physical interface steps intuitive.
Lifestyle companionship: The agent can serve as a wellness coach that interacts with your physical choices. For instance, the agent can perform an interactive “warning” gesture by holding the user’s hand and a visual effect to caution the user against unhealthy behavior.
To evaluate the impact of these gestures, we conducted a within-subjects study (N = 12) comparing AgentHands to a speech-only baseline. Both conditions used the same researcher-scripted verbal content, ensuring the only difference was the presence of the embodied hands and their synchronized gestures. Participants completed two procedural tasks that balanced everyday care with technical operation.
The results confirmed that the combination of XR and co-speech gestures is highly effective for spatially grounded interactions. We analyzed the data across several key metrics of communication effectiveness.
AgentHands represents a step toward a future where AI systems aren’t just analyzing our world, but dynamically operating within it. By leveraging co-speech gestures and the spatial power of XR to ground conversation in physical movement, we can reduce the cognitive load of complex tasks and make spatial computing more accessible and human-centric.
As we continue to develop for the Android XR ecosystem, we are exploring ways to make these gestures even more personalized, adapting to a user’s dominant hand or learning their specific spatial routines, to create an even more seamless human-AI collaboration.
This research was primarily conducted by Ziyi Liu during his Student Researcher tenure at Google, as part of a joint collaboration across multiple teams. We extend our sincere gratitude to key contributors David Li, Zhongyi Zhou, and David Kim for their support, and to Adarsh Kowdle, Guru Somadder, and Shahram Izadi for their strategic guidance and thoughtful reviews.
文章标题:AgentHands:在XR中为空间具身智能体对话生成交互式手部手势
文章链接:https://news.qimuai.cn/?post=4891
本站文章均为原创,未经授权请勿用于任何商业用途