Gemini 3.8 Live 变革对话式AI

内容来源:https://aibusiness.com/generative-ai/gemini-3-8-live-transforms-conversational-ai
内容总结:
谷歌云近日发布两款全新音频模型Gemini 3.8 Live与Live Extended Thinking,允许企业部署具备深度推理能力、更强交互性和对话速度的智能AI代理。该模型于9月15日正式推出,开发者可借此构建在保持对话流畅的同时进行推理和执行任务的语音代理。其核心能力包括:在持续向用户流式传输音频响应的同时执行API和工具调用、提供实时视觉输入以帮助代理理解用户所言所见,以及支持超过97种语言。其中,3.8 Live Extended Thinking版本支持可配置思考功能,用户可自行开启或关闭AI模型的内部推理过程。
此次音频模型的发布距谷歌首次推出Gemini 3.8 Flash和3.8 Cyber基础模型仅过去两周。Futurum Group分析师布拉德利·希明表示,凭借3.8 Live的能力,谷歌似乎正在追赶苹果等厂商——苹果已在Notes和Voice Memos等生产力应用中提供实时转录功能。但希明指出,谷歌的技术通过改变用户与AI的交互方式,将这一能力提升到了新层次。
希明表示:“从其架构方式来看……你可以对其进行定制,使其以多种不同方式运行。”他指出,企业虽可将这些模型用于传统翻译场景,但谷歌的设计意图是让用户不再以传统的提示-响应方式与之交互,而是可以进行真正的对话。“这就是一种自然对话,并且能够实时打断并注入内容,”希明说,“你可以在不中断其当前操作的情况下,向对话中注入上下文信息。”
Tekonyx创始人西德·纳格表示,借助谷歌先进的语音技术能力,新模型不会仅停留在回答问题的表面层次。“该模型被设计用于进行后台推理,”纳格说,“它在对话过程中同步进行推理,因此可以在推理的同时保持对话持续进行。”
这项技术的进步在于将交互延迟与推理延迟分离。交互延迟是指用户执行操作与系统响应之间的延迟;推理延迟则是AI模型处理信息所需的总时间。“它是实时完成的,而不是先暂停再返回另一个答案,”纳格说。
纳格继续表示,这些语音模型还改变了AI代理的响应和交互方式,因为“AI代理不必再在快速响应和深思熟虑之间二选一”,“它实际上可以维持实时交互。”
纳格指出,新模型的企业应用场景包括客户支持聊天机器人、销售流程,以及结合文本、语音和视频的多模态代理应用。
中文翻译:
由谷歌云赞助
选择你的首批生成式AI用例
要开始使用生成式AI,首先应关注那些能够改善人类信息体验的领域。
这些模型支持实时对话,使用户无需遵循传统的提示-响应结构即可进行交互。
谷歌推出了两款新的音频模型,使企业能够部署具有推理深度、更强交互性和对话速度的智能AI代理。
谷歌表示,9月15日发布的Gemini 3.8 Live和Live Extended Thinking让开发者能够构建语音代理,这些代理可以在保持对话流畅的同时进行推理和执行任务。这些模型的核心能力包括:在执行API和工具调用的同时持续向用户流式传输音频响应;提供实时视觉输入,帮助代理理解用户所说和所见的内容;以及支持超过97种语言。3.8 Live Extended Thinking版本支持可配置思考功能,该功能允许用户开启或关闭AI模型的内部推理。
这两款音频模型的发布距离谷歌最初推出Gemini 3.8 Flash和3.8 Cyber基础模型仅两周。Futurum Group分析师布拉德利·希明表示,凭借3.8 Live的能力,谷歌似乎在追赶苹果等厂商,后者已在Notes和Voice Memos等生产力应用中提供实时转录功能。不过,他说,谷歌的技术通过改变用户与AI的交互方式,将其提升到了一个新的水平。
“它的架构方式……你可以对其进行定制,使其以多种不同方式运行,”希明说。他指出,虽然企业可以将这些模型用于传统的翻译用例,但谷歌在设计这些模型时也考虑到了让用户不以提示-响应方式进行交互,而是可以进行真正的对话。
“这就是一种自然对话,并且能够实时打断以注入内容,”希明说。“你可以在不打断它正在做的事情的情况下,将上下文注入对话中。”
Tekonyx创始人西德·纳格表示,借助谷歌先进的语音技术能力,这些新模型不仅仅是表面化地回答一个问题。
“该模型被设计用于进行后台推理,”纳格说。“它在对话过程中进行推理,因此你可以在它推理的同时保持对话持续进行。”
这项技术的进步在于将交互延迟与推理延迟分离。交互延迟是用户执行操作与系统响应之间的延迟。推理延迟是AI模型处理信息所需的总时间。
“它是实时进行的,而不是停下来然后再给出另一个答案,”纳格说。
纳格继续表示,这些语音模型还改变了AI代理的响应和交互方式,因为“AI代理不必在快速和深思熟虑之间做出选择”。“它实际上可以维持实时交互。”
纳格说,这些新模型的企业应用包括客户支持聊天机器人、销售流程以及结合文本、语音和视频的多模态代理应用。
英文来源:
Sponsored by Google Cloud
Choosing Your First Generative AI Use Cases
To get started with generative AI, first focus on areas that can improve human experiences with information.
The models support real-time conversation by enabling users to interact without the traditional prompt-response structure.
Google introduced two new audio models that enable enterprises to deploy intelligent AI agents with reasoning depth, more interactivity and conversational speed.
Launched on September 15, Gemini 3.8 Live and Live Extended thinking let developers build voice agents that can reason and execute tasks while maintaining the flow of the conversation, Google said. The models’ key capabilities include performing API and tool calls while continuing to stream audio responses to users, providing live visual inputs that help agents understand what users say and see, and supporting more than 97 languages. The 3.8 Live Extended Thinking version supports configurable thinking, a feature that lets users turn an AI model’s internal reasoning on or off.
The release of the audio models comes two weeks after Google initially introduced the Gemini 3.8 Flash and 3.8 Cyber foundation models. With 3.8 Live’s capabilities, Google appears to be catching up to vendors such as Apple, which already offers real-time transcription in productivity apps such as Notes and Voice Memos, said Bradley Shimmin, an analyst at Futurum Group. Still, Google tech takes it to the next level by changing how users interact with AI, he said.
“The way it's architected … you can customize this to behave in a lot of different ways,” Shimmin said. He noted that while enterprises can use the models for traditional translation use cases, Google has also designed the models so users don’t interact with it in a prompt-response way; rather, they can have a real conversation.
“It's just a natural conversation with the ability to interrupt it in real time to inject,” Shimmin said. “You can inject context into the conversation without interrupting what it's doing.”
With Google’s advanced speech technology capability, the new models don’t just answer a question at face value, said Sid Nag, founder of Tekonyx.
“The model is designed to do background reasoning,” Nag said. “It’s doing it during the conversation so you can keep the conversation alive while it’s reasoning.”
This technology is an advancement in which interactive latency is separated from reasoning latency. Interactive latency is the delay between a user performing an action and the system responding. Reasoning latency is the total time it takes for an AI model to process information.
“It’s doing it in real time, rather than halting and then coming back with another answer,” Nag said.
The speech models also change how an AI agent responds and interacts because “an AI agent doesn’t necessarily have to choose between being fast and being thoughtful,” Nag continued. “It can actually maintain a real-time interaction.”
Enterprise applications for the new models include customer support chatbots, sales processes and multimodal agentic applications that combine text, voice and video, Nag said.
文章标题:Gemini 3.8 Live 变革对话式AI
文章链接:https://news.qimuai.cn/?post=5102
本站文章均为原创,未经授权请勿用于任何商业用途