新的Gemini 3.5 Flash模型速度更快、成本更低,但并未变得更智能。

qimuai 发布于 阅读:59 一手编译

新的Gemini 3.5 Flash模型速度更快、成本更低,但并未变得更智能。

内容来源:https://aibusiness.com/generative-ai/new-gemini-3-5-flash-models-are-faster-cheaper-not-smarter

内容总结:

谷歌发布新一代Gemini模型,聚焦成本优化与网络安全

谷歌云赞助消息:在选择首个生成式AI应用场景时,应优先关注能提升人类信息体验的领域。

当地时间周二,谷歌发布三款全新Gemini模型,在AI模型市场竞争日益激烈的背景下,既回应了企业对成本的关注,又通过提升速度和专业化程度,与OpenAI及Anthropic展开较量。

主打高性价比:Flash系列加速降本

此次发布的Gemini 3.6 Flash模型专注于编码与知识密集型任务。谷歌重点优化了成本效率,在部分基准测试中,其每输出Token的成本低于前代Gemini 3.5 Flash。该系列中最具性价比的是Flash-Lite模型,每秒可输出约350个Token。谷歌表示,当前全球AI市场正转向更实惠的模型,中国AI厂商阿里和月之暗面本周也推出了高性能、低成本模型。

AI分析机构Tekonyx总裁兼首席研究官Sid Nag指出:“AI模型的成本效益已成为行业热点,如何在优化成本的同时保证速度,将是业内持续讨论的话题。”他进一步分析,对于企业而言,模型训练和推理是最大开支。如果谷歌能通过3.6 Flash降低延迟和Token成本,“对于大规模部署,尤其是企业级部署,将意味着巨大的成本节约。”

从定价看,Gemini 3.5 Flash-Lite最为亲民:每百万输入Token仅0.3美元,每百万输出Token为2.5美元。相比之下,中国模型Kimi K3的定价为每百万输入Token 3美元、输出15美元;OpenAI的GPT-5.6系列则更高,输入价格在1至5美元之间,输出在6至30美元之间。

不过,康奈尔大学信息科学讲师、AI创新主管Ayham Boucher指出,并非所有新模型都完全达标。例如,Gemini 3.6 Flash每百万输出Token为7.5美元,高于SpaceXAI的Grok 4.5(输出6美元)。他提醒,在智能体系统中,输入Token通常较短,而智能体在响应前会生成大量输出,“输出Token比输入Token更重要。”

速度领先但智能尚有差距

Boucher肯定谷歌在速度上的优化:“它们确实在速度上遥遥领先于所有封闭前沿模型,只有开源模型能与之竞争。”但在智能水平上,他认为谷歌未能达到Anthropic和OpenAI封闭模型的水平。不过,这并非意味着谷歌落后,而是其战略正聚焦于打造更快、更便宜的基础设施。“推理成本和高吞吐管理对AI系统成功部署至关重要。谷歌押注的方向没错——市场存在巨大需求,企业并不总是需要最聪明的模型,而是需要既快又聪明的廉价模型。”

他展望,如果明年发布的Gemini 3.6版本能在智能水平上超越Anthropic的Fable,同时保持速度和低价,必将引发强烈关注。“明年的Fable会更智能,但许多企业任务并不需要最高智能。”

网络安全模型:AI专业化新方向

谷歌还推出了专注网络安全的Gemini 3.5 Cyber模型,并将其与AI智能体CodeMender结合,用于检测、修复和预防安全漏洞。Nag认为,这体现了AI正走向更专业化的细分领域。“AI正融入软件开发生命周期,成为安全团队的效率倍增器,低成本的安防智能体也具备了商业可行性。”

不过,Guardrail Technologies首席执行官TJ Marlin提醒,对这类自主识别并修复漏洞的系统仍需谨慎。“当系统自主发现并修复漏洞时,谁来监督制衡?已知的局限性、可靠性问题和风险并不会因为有了编排模型就消失。”他指出,谷歌只展示了Cyber模型的正面效果,其可靠性尚未可知。“所有AI被有意或无意操控、导致不良后果的风险,同样存在于谷歌这套漏洞检测系统中。”

中文翻译:

由谷歌云赞助
选择您的首个生成式人工智能应用场景
要开始使用生成式AI,首先应聚焦于能够改善人类与信息交互体验的领域。

更新后的模型旨在降低企业使用成本。全新网络安全模型则专为自动化协调而设计。

谷歌于周二发布三款全新Gemini模型,表明该供应商在AI模型市场关注成本问题的同时,仍通过提升模型速度与领域专用性,与OpenAI和Anthropic展开竞争。

Gemini 3.6 Flash模型专注于编码与知识领域。谷歌在此模型上重点优化成本效率,部分基准测试中每输出令牌成本低于Gemini 3.5 Flash。据谷歌介绍,Flash-Lite是该系列性价比最高的模型,每秒可输出约350个令牌。而3.6 Flash Cyber是一款全新的网络安全模型,将与谷歌AI代理CodeMender配合使用,用于检测、修复及预防安全漏洞。

新模型在Gemini 3.5 Pro之前发布——该模型目前正在测试中,即将面世。更新版Flash系列推出之际,全球AI市场正转向低成本模型,中国AI供应商阿里巴巴与Moonshot本周也发布了高性能低成本模型。

“AI模型在行业中的成本效益已成为焦点话题,”AI分析公司Tekonyx总裁兼首席研究官Sid Nag表示。“在成本与速度之间寻求优化,将是我们行业参与者持续讨论的议题。”

企业最大成本之一在于模型训练与推理。Nag进一步指出,若谷歌能通过3.6 Flash降低延迟与令牌成本,“便可为大规模部署(尤其是企业部署)带来显著成本节约”。

与近期模型相比,Gemini 3.5 Flash-Lite似乎是该系列最实惠选择:每百万输入令牌0.3美元,每百万输出令牌2.5美元。作为对比,中国模型Kimi K3每百万输入令牌3美元,每百万输出令牌15美元。Gemini 3.5 Flash-Lite同样低于OpenAI的GPT-5.6系列(每百万输入令牌1至5美元,每百万输出令牌6至30美元)。

但康奈尔大学信息科学讲师兼AI创新主管Ayham Boucher表示,谷歌并未在所有新模型上完全实现成本最优。例如,3.6 Flash的单位输出成本仍高于SpaceXAI的模型。Gemini 3.6 Flash每百万输入令牌1.5美元,每百万输出令牌7.5美元;而SpaceXAI的Grok 4.5每百万输入令牌2美元,每百万输出令牌6美元。

“众所周知,在代理系统中,输入提示可能非常简短,但代理在响应前会生成大量输出,”Boucher说。“输出令牌比输入令牌更为关键。”

Boucher指出,除价格外,谷歌还优化了模型速度。
“他们确实实现了速度领先,远超所有前沿闭源模型,只有开源模型能与之竞争。”

但在智能水平方面,他表示谷歌未能达到Anthropic和OpenAI闭源模型的高度。这并非意味着谷歌落后,而是其正瞄准构建更快、更便宜模型的基础设施。

“推理成本及高吞吐推理管理,将随AI系统成功部署而变得至关重要,”Boucher说。“这也是谷歌押注的方向——存在巨大市场无需最智能模型,只需快速、智能且廉价的模型,这一判断完全正确。”

他补充道,若明年发布的Gemini 3.6版本比Anthropic当前Fable版本更智能,同时兼具快速与低成本优势,必将引发广泛关注。因为明年的Fable可能更智能,而许多企业无需最高智能模型完成无需顶尖能力的任务。

但Boucher强调,谷歌聚焦平价模型,并不意味着可以放弃在智能层面的前沿竞争。
“若止步于此,便会在前沿智能竞赛中落后。提升模型速度的同时,不应放弃在智能前沿的领先地位。”

谷歌发布Gemini 3.5 Cyber亦明确显示其专注网络安全模型。在Nag看来,该模型印证了AI正走向更专业化。
“这带来三大企业级影响:AI成为软件开发生命周期的一环;AI成为安全团队的效率倍增器;低成本安全代理变得切实可行。”

Guardrail Technologies CEO TJ Marlin表示,尽管Gemini 3.5-Cyber已创建可自主识别修复安全问题的代理系统,但该模型仍需深入认知。
“若由自主系统查找并修复漏洞,制衡机制何在?已知的局限性、可靠性问题及风险,不会因创建了协调模型而消失。”

他补充道,谷歌虽展示了Cyber的积极面,但模型可靠性尚未明朗。
“我们未能了解全貌。AI被有意或无意操纵导致不良后果的风险,同样存在于谷歌为漏洞检测创建的代理系统中。”

英文来源:

Sponsored by Google Cloud
Choosing Your First Generative AI Use Cases
To get started with generative AI, first focus on areas that can improve human experiences with information.
The updated models are intended to be more affordable for enterprises. The new cyber model is designed to orchestrate.
Google introduced three new Gemini models on Tuesday, showing the vendor is paying attention to cost concerns in the AI model market while still competing against OpenAI and Anthropic by making its models faster and more domain-specific.
The Gemini 3.6 Flash model is coding and knowledge focused. With this model, Google concentrated on cost efficiency, with a lower cost per output tokens on some benchmarks compared to Gemini 3.5 Flash. Flash-Lite is the most cost-effective model in the series, delivering about 350 output tokens per second, Google said. And 3.6 Flash Cyber is a new cybersecurity model that will be paired with CodeMender, Google’s AI agent for detecting, patching and preventing security vulnerabilities.
The new models arrive before the Gemini 3.5 Pro model, which Google is currently testing and will be made available soon. However, as updated Flash series arrives as momentum in the worldwide AI market shifts toward less expensive models, with Chinese AI vendors Alibaba and Moonshot this week launching new high-performing, low-cost models.
“Cost utilization of AI models in general in the industry has become a topic of interest,” said Sid Nag, president and chief research officer of AI analyst firm Tekonyx. “Cost optimized with speed is going to be a continuing dialogue that we as industry participants are going to be talking about.”
One of the biggest enterprise costs is model training and inference. So, if Google can reduce latency and token cost with 3.6 Flash, it can “translate into substantial savings for large deployments, especially enterprise deployments,” Nag continued.
Compared to recent models, Gemini 3.5 Flash-Lite appears to be the cheapest option in the line, costing $0.3 per million input tokens and $2.5 per million output tokens. By comparison, the Chinese model Kimi K3 costs $3 per million input tokens and $15 per million output tokens. Gemini 3.5 Flash-Lite is also cheaper than OpenAI’s GPT-5.6 models, which range from $1 to $5 per million input tokens and $6 to $30 per million output tokens.
However, Google did not completely hit the mark on cost with all the new models it released, said Ayham Boucher, the Head of AI Innovations and a lecturer in information science at Cornell University. The cost of 3.6 Flash is still higher per unit of output than that of models from SpaceXAI, for example. Gemini 3.6 Flash runs at $1.50 per million input tokens and $7.50 per million output tokens. Comparatively, SpaceXAI’s Grok 4.5 costs $2 per million input tokens and $6 per million output tokens.
“As we all know, with the input tokens with agentic systems, your prompt could be very short, and then the agent is going to go and generate a lot of output before its response,” Boucher said. “Output tokens do matter more than input tokens.”
Boucher noted that, in addition to price, Google optimized the speed of its models.
“They definitely achieved the speed, which is definitely way ahead of all the frontier closed models, and you could only compete with it with open source models,” Boucher said.
However, regarding intelligence, he said that Google itself failed to reach the level of other closed models from Anthropic and OpenAI. This does not mean that Google is falling behind, but rather that it is targeting infrastructure to enable faster, cheaper models, he said.
“Cost of inference, and managing high throughput of inference, is going to be very pivotal and essential as we continue to deploy AI systems successfully,” Boucher said. “This is what Google is betting on, and they are not wrong that there is a huge market where you do not necessarily need the smartest model. You need a fast, smart model that is cheap.”
He added that if the Gemini 3.6 version expected next year is more intelligent than the current Fable version from Anthropic, while also being fast and less expensive, it could garner a lot of interest. This is because next year’s Fable will likely be more intelligent and many enterprises do not always need the most intelligent model because many tasks that don’t require the highest intelligence.
However, Google’s focus on more affordable models does not mean it should not also provide a frontier model that also competes on intelligence, Boucher said.
“If it doesn't, then it has fallen behind in the race of frontier intelligence,” he said. “Just because you're making your models faster, you shouldn’t also give up your spot in the leading edge of intelligence.””
It is also clear that Google is focusing on cybersecurity models with its release of Gemini 3.5 Cyber. For Nag, the cyber model illustrates another way AI is becoming more specialized.
“It has three major enterprise implications,” he said. “AI becomes part of the software development lifecycle. AI becomes a force multiplier for security teams, and of course, lower-cost security agents become practical.”
While with Gemini 3.5-Cyber, Google has created an agentic system that autonomously identifies and fixes security problems, there is still much to learn about this model, said TJ Marlin, CEO of Guardrail Technologies.
“If you're going to have an autonomous system that's finding the vulnerabilities and fixing them, where are the checks and balances there?” Marlin said. “There’s a number of known limitations, reliability issues and risks that don’t go away just because they created an orchestration model.”
In addition, while Google has showcased the positives of Cyber, it is unclear how reliable the model is, he said.
“We're not learning the full picture,” Marlins said. “All the risks that exist for AI being manipulated intentionally or unintentionally, resulting in bad outcomes, they also exist for this agentic system that Google has created for vulnerability detection.”

商业视角看AI

文章目录


    扫描二维码,在手机上阅读