Gemini 3.8 文本转语音功能进一步完善了语音 AI 的能力

qimuai 发布于 阅读:0 一手编译

Gemini 3.8 文本转语音功能进一步完善了语音 AI 的能力

内容来源:https://aibusiness.com/generative-ai/gemini-3-8-text-speech-refines-voice-ai-capabilities

内容总结:

谷歌云赞助

选择你的首批生成式人工智能应用场景

要开始使用生成式人工智能,首先应关注能够改善人类信息体验的领域。

谷歌此次推出的功能建立在现有技术基础之上,为企业提供了一个汇聚其所有AI模型能力的中心平台。

谷歌为Gemini 3.8新增了一款文本转语音模型,虽然并非开创性突破,但对该领域已有的AI音频技术进行了优化和提升。

谷歌于9月23日推出了Gemini 3.8 TTS和Gemini 3.8 Flash-Lite TTS,号称是该公司迄今为止能力最强的音频生成模型。

谷歌表示,这些模型能够帮助创作者、开发者和企业生成更高质量、更具表现力的音频内容。Gemini 3.8 Flash TTS面向创意指导和角色设计,可通过自然语言提示从零开始创建全新语音,适用于游戏、有声书、播客和互动媒体等多种媒介。该模型还支持在100多种语言和方言中自定义角色、口音和声音特征。Gemini 3.8 Flash-Lite TTS则面向音频配音、内容创作和语音代理等场景。

Gemini 3.8 Flash TTS和Flash-Lite TTS的许多功能在音频或AI语音市场并非首创。ElevenLabs、Baseten等厂商已提供同类音频技术。但通过此次发布,谷歌对自身的AI音频技术进行了改进,为娱乐、营销等领域的企业用户提供了更多选择和更逼真的效果。

Futurum Group分析师Bradley Shimmin表示:“我们在这里看到的是文本转语音技术的精细化,并针对特定用例进行了重点优化。”他说,该模型对制作有声书以及长篇和短篇叙事类炉边谈话等格式的用户非常有用。

Shimmin表示:“这为游戏和娱乐等市场带来了许多有趣的机会,因为你突然有机会创建高度表现力、高度逼真的环境,在其中进行互动……消费极具个性化的音频内容。”

Gemini 3.8 TTS的另一个卖点是它属于Gemini模型家族。

AI语音厂商Modulate首席执行官Carter Huffman表示:“行业和技术前沿正从ElevenLabs等提供的单一独立端点……转向一个 cohesive 模型,能够完成你在语音方面所需的一切。拥有一个统一的端点是一种优势,前提是这个统一模型能提供出色的性能,最好是一流的性能。”

Shimmin表示,对企业而言,语音模型属于同一模型家族这一事实意味着更紧密的集成。

他说:“数据专业人士指出,他们在使用AI时面临的最大挑战是集成,不是数据集成,而是技术集成。”

Huffman表示,尽管Gemini基础模型支持集成,但语音AI能力仍处于早期阶段。因此,谷歌要想真正在市场上脱颖而出,需要提供别人没有的能力或前所未有的应用。

中文翻译:

由谷歌云赞助

选择你的首批生成式AI用例

要开始使用生成式AI,首先应关注那些能够改善人类信息体验的领域。

谷歌推出的这些功能建立在现有技术基础之上,为企业提供了一个汇聚其所有AI模型能力的中心。

谷歌为Gemini 3.8新增了一款文本转语音模型,虽然算不上突破性创新,但优化并提升了已有的AI音频技术。

谷歌于9月23日推出Gemini 3.8 TTS和Gemini 3.8 Flash-Lite TTS,这是该科技巨头迄今能力最强的音频生成模型。

谷歌表示,这些模型使创作者、开发者和企业能够生成更高质量、更具表现力的音频。Gemini 3.8 Flash TTS面向创意指导和角色设计,可通过自然语言提示从零开始创建全新语音,适用于游戏、有声书、播客和互动媒体等媒介。该模型还支持在100多种语言和方言中自定义角色、口音和声音特征。Gemini 3.8 Flash-Lite TTS则面向音频配音、内容创作和语音代理。

Gemini 3.8 Flash TTS和Flash-Lite TTS的许多功能在音频或AI语音市场上并不新鲜。ElevenLabs、Baseten等厂商也提供同类音频技术。不过,通过此次发布,谷歌正在改进自身的AI音频技术,为娱乐和营销等领域的企业用户提供更多选择和更逼真的效果——如果他们选择谷歌的模型的话。

“我们在这里看到的是文本转语音技术的一次精进,并且针对特定用例做了大量定向优化,”Futurum Group分析师布拉德利·希明表示。他说,该模型对制作有声书以及长篇和短篇 narrated fireside chats 等格式的用户很有用。

“它为游戏和娱乐等市场开辟了许多有趣的机遇,因为你突然有机会创造高度表现力、高度逼真的环境,在其中你可以互动……并消费极具个性化的音频内容,”希明说。

Gemini 3.8 TTS的另一个卖点是它属于Gemini模型家族。

“行业和技术前沿正在从ElevenLabs等提供的单个独立端点……转向一个能够完成你所有语音需求的统一连贯模型,”AI语音厂商Modulate的CEO卡特·哈夫曼表示。“拥有一个统一连贯的端点是一种优势,前提是这个统一模型能提供卓越的性能,最好是一流的性能。”

希明表示,对企业而言,语音模型属于同一模型家族这一事实带来了更紧密的集成。

“数据专业人士指出,他们在使用AI时面临的最大挑战是集成,”他说。“不是数据集成,而是技术集成。”

哈夫曼表示,尽管Gemini基础模型支持集成,但语音AI能力仍处于早期阶段。因此,他说,谷歌要想在市场上真正脱颖而出,需要提供别人没有的能力或前所未有的应用。

英文来源:

Sponsored by Google Cloud
Choosing Your First Generative AI Use Cases
To get started with generative AI, first focus on areas that can improve human experiences with information.
The features Google introduced build on existing technology and provide enterprises with a hub for all their AI model capabilities.
Google has added a new text-to-speech model to Gemini 3.8, which, while not groundbreaking, refines and improves AI audio technology that already exists.
Google introduced Gemini 3.8 TTS and Gemini 3.8 Flash-Lite TTS on September 23 as the tech giant’s most capable audio generation models yet.
The models enable creators, developers and enterprises to create higher quality, more expressive audio, the vendor said. Gemini 3.8 Flash TTS is for creative direction and character design. It creates new voices from scratch using natural language prompts across mediums such as gaming, audiobooks, podcasts and interactive media. The model can also customize role, accent and voice characteristics across more than 100 languages and dialects. Gemini 3.8 Flash-Lite TTS is for audio dubbing, content creation and voice agents.
Many of the features in Gemini 3.8 Flash TTS and Flash-Lite TTS are not new to the audio or AI speech market. Vendors such as ElevenLabs, Baseten and others provide the same type of audio technology. However, with this release, Google is improving its own AI audio technology, giving enterprise users in entertainment and marketing, among others, more choice and realism if they choose its model.
“What we’re seeing here is a refinement of your text-to-speech with some heavy targeting to specific use cases,” said Bradley Shimmin, an analyst at Futurum Group. He said the model is useful for users creating audiobooks and formats such as long-form and short-form narrated fireside chats.
“It opens up a lot of interesting opportunities for markets like gaming and entertainment because you suddenly have the opportunity to create highly expressive, highly realistic environments where you can interact with ... and consume audio content that is extremely personable,” Shimmin said.
Another selling point for Gemini 3.8 TTS is that it is part of the Gemini family of models.
“The industry and the state of the art are going from individual single endpoints like what ElevenLabs offers … to one cohesive model that can do everything you need to with voice,” said Carter Huffman, CEO of Modulate, an AI voice vendor. “Having one cohesive endpoint is a win so long as that one cohesive model gives you excellent performance, ideally best-in-class performance.”
For enterprises, the fact that the voice model is in the same model family enables tighter integration, Shimmin said.
“Data professionals cite that the biggest challenge they have in using AI is integration,” he said. “Not data integration but integration of technologies”
Although the Gemini foundation model allows for integration, the voice AI capabilities are still nascent, Huffman said. So, for Google to really stand out in the market, it needs to provide a capability no one has or an application that hasn’t been seen before, he said.

商业视角看AI

文章目录


    扫描二维码,在手机上阅读