Z.AI在新模型中使用中国芯片,关键在于优化。

内容来源:https://aibusiness.com/generative-ai/z-ai-s-use-chinese-chips-new-model-about-optimization
内容总结:
中国AI企业Z.ai实现重大突破:十万国产芯片驱动大模型,加速摆脱对英伟达依赖
近日,中国人工智能企业Z.ai宣布其最新AI模型GLM-5.3 Flash已全部采用10万颗国产芯片处理在线查询请求。这一举措标志着中国科技企业在降低对美国芯片巨头英伟达依赖方面迈出实质性步伐,同时也折射出AI模型优化水平不断提升的新趋势。
国产芯片扛起推理重任
Z.ai此次发布的GLM-5.3 Flash为开源权重多模态模型,拥有3200亿参数,主打低成本优势,专注于长文本处理、视觉驱动智能体任务及代码生成。该模型于8月20日以代号“Ox Alpha”首次亮相,8月26日正式以MIT开源许可协议发布。定价方面,9月9日前每百万输入tokens收费0.075美元、每百万输出tokens收费0.25美元,此后价格将调整为输入0.15美元、输出0.5美元。相比之下,OpenAI的GPT-5.6 Luna输入定价为0.2美元/百万tokens、输出定价为2美元/百万tokens,而Anthropic的Claude Opus 5则高达输入输出各5美元/百万tokens。
政策驱动下的自主可控之路
Z.ai坚持采用纯国产芯片供应商的决策,与近年来中国着力降低对英伟达依赖的大背景密切相关。美国持续收紧出口管制,阻止中国科技企业获取最先进的英伟达芯片,尽管部分限制有所放宽,但2025年9月中国政府已要求阿里巴巴、字节跳动等科技巨头停止采购英伟达AI芯片,11月更禁止国有数据中心使用所有外国AI芯片。不过,此前多数观察人士仍认为中国在基础设施竞争中处于落后地位。Z.ai的实践表明,随着智能体AI兴起带来的推理需求激增,中美在AI芯片领域的差距正在收窄。
推理市场或成弯道超车关键
RPA2AI Research创始人兼CEO Kashyap Kompella指出:“在训练顶级前沿模型方面,英伟达在芯片性能、网络连接和软件生态上仍有显著优势。但推理是另一回事。”他强调,Z.ai能用国产芯片规模化支撑GLM-5.3 Flash服务,证明中国硬件已足以应对大量高负载推理场景。“这一区别至关重要,因为推理很可能成为未来更大的AI芯片市场。如果中国企业能通过高效模型架构、软件优化和大规模国产芯片集群来弥补,或许无需在芯片性能上与英伟达逐一对标。”
Futurum Group分析师Bradley Shimmin则认为,中国厂商迈向自主化的最早信号始于DeepSeek重写英伟达CUDA驱动程序以提升推理速度。他指出,除响应美中两国政府先后施加的限制外,企业还需“更高效利用有限的电力资源”。谷歌TPU系列和亚马逊Trainium芯片等垂直整合路线,正是AI价值驱动因素从模型能力向“全栈经济性”转变的体现。
开发者用脚投票:盲测表现亮眼
值得关注的是,GLM-5.3 Flash在盲测中获得了开发者的高度认可。大量开发者本月早些时候在不知情的情况下涌向该模型,使其迅速成为OpenRouter平台上下载量最大的模型。Kompella认为,这比“又一个声称中国模型接近美国前沿模型的基准测试”更有说服力,表明中国开源权重模型“继续充当制衡行业高定价的力量”。他建议企业应设计灵活的AI系统,使工作负载能够根据能力、质量、成本和战略考量在多个模型间路由,而非依赖单一供应商。
前路仍存挑战
尽管GLM-5.3 Flash迅速吸引大量开发者对Z.ai而言是积极信号,但该企业仍面临诸多考验。Kompella表示,一是需证明所用芯片能持续提供可靠性、经济性和规模化能力;二是在模型训练的绝对前沿,算力获取仍是重大制约,英伟达在此领域的优势难以复制;三是需要赢得对数据安全、监管合规及采购风险持谨慎态度的欧美企业客户信任。
中文翻译:
由谷歌云赞助
选择你的首个生成式AI应用场景
要开始使用生成式AI,首先要聚焦于能够改善人类信息体验的领域。
使用国产芯片表明中国厂商正变得更加自力更生,同时也在提升AI芯片上的推理性能。
中国AI厂商Z.ai表示,已使用10万颗国产芯片来处理其最新AI模型GLM-5.3 Flash的在线查询。
此举表明中国厂商正在减少对美国芯片制造商英伟达的依赖,同时也凸显了AI模型进一步优化的趋势。
这款全新的开放权重多模态模型成本低廉,拥有3200亿参数。Z.ai最初于8月20日以代号Ox Alpha预览了该模型,并于8月26日在MIT开源许可下正式发布。该模型专长于长上下文处理、视觉驱动的智能体任务和代码合成。
GLM-5.3 Flash目前至9月9日期间,每百万输入tokens收费0.075美元,每百万输出tokens收费0.25美元。此后,价格将调整为每百万输入tokens 0.15美元,每百万输出tokens 0.50美元。相比之下,OpenAI的GPT-5.6 Luna每百万输入tokens收费0.20美元,每百万输出tokens收费2美元。Anthropic的Claude Opus 5每百万输入tokens收费5美元,每百万输出tokens收费5美元。
这家中国厂商决定使用纯国产芯片供应商,正值中国力求减少对英伟达依赖之际。在多年美国收紧出口管制、阻止中国科技厂商获取最强大英伟达芯片之后,中国转向了自力更生。尽管其中一些管制有所放松,但在2025年9月,中国政府下令阿里巴巴和字节跳动等科技巨头停止购买英伟达AI芯片,11月,北京又禁止所有外国AI芯片进入政府资助的数据中心。
尽管有这些举措,许多观察人士仍认为中国在基础设施竞赛中处于落后地位。然而,Z.ai的策略表明,中美在AI芯片上的差距可能正在缩小,尤其是在过去一年智能体AI迅速崛起、业界更加关注推理的情况下。
RPA2AI Research首席执行官兼创始人Kashyap Kompella表示:“在训练最前沿的大模型方面,英伟达在芯片性能、网络和更广泛的软件生态上仍具有显著优势。但推理是另一回事。”
他补充说,Z.ai能够使用国产芯片大规模服务GLM-5.3-Flash,表明中国硬件已经足以应对许多高吞吐量的推理工作负载。
Kompella表示:“这一区别很重要,因为推理很可能随着时间的推移成为更大的AI芯片市场。如果中国公司能够通过高效的模型架构、软件优化和大规模国产AI芯片集群来弥补差距,它们可能并不需要与英伟达在芯片层面完全对等。”
然而,中国厂商走向独立的最初信号始于DeepSeek重写英伟达CUDA驱动程序以加速推理,Futurum Group分析师Bradley Shimmin如是说。
他表示,虽然中国厂商是在回应美国政府首先施加、后来中国政府也加入的限制,但另一个因素是各公司“需要用可用的瓦时做更多的事情”。
谷歌(TPU系列)和亚马逊(Trainium芯片)等厂商正专注于垂直整合,即AI堆栈经过优化,将AI模型、基础设施和软件结合起来,以实现最高效的AI处理。
Shimmin表示:“我们在这里见证的是一个真正的认识——AI的价值不仅取决于模型的能力,同样取决于堆栈的经济性。这些公司所做的只是展示在优化硬件上投入一点点的价值。”
至于GLM-5.3-Flash本身,它在盲测中获得开发者高度认可是意义重大的,Kompella说。本月早些时候,许多开发者涌入该模型,在不知道它来自中国厂商的情况下使用它;它很快成为OpenRouter上下载量最高的模型。
Kompella表示:“这比又一个声称中国模型接近美国前沿模型的基准测试更有意思。”他补充说,这表明中国的开放权重模型“继续充当制衡行业高定价的力量”。
他继续说道,对于企业而言,这意味着它们需要设计自己的AI系统,“使工作负载能够根据能力、质量、成本和战略考量在不同模型之间路由,而不是依赖单一模型提供商”。
虽然该模型迅速吸引了大量开发者的关注对Z.ai来说是好兆头,但该厂商仍面临前路挑战。其一,它必须证明所使用的芯片能够持续提供可靠性、经济性和可扩展性,Kompella说。
他表示:“在模型训练的最前沿,计算资源的获取仍然是一个重大制约,因为英伟达在那里的优势更难复制。”
Kompella还表示,该厂商还需要说服那些出于数据安全、监管和采购方面的顾虑而对使用中国托管AI服务持谨慎态度的欧美企业。
英文来源:
Sponsored by Google Cloud
Choosing Your First Generative AI Use Cases
To get started with generative AI, first focus on areas that can improve human experiences with information.
Using domestic chips shows that Chinese vendors are becoming more self-reliant and improving inference performance on AI chips.
Chinese AI vendor Z.ai said it used 100,000 China-made chips to handle online queries to its latest AI model, GLM-5.3 Flash.
The move shows how Chinese vendors are reducing their dependence on U.S. chipmaker Nvidia, while also highlighting the trend toward greater optimization of AI models.
The new open-weight, multimodal model is low-cost, with 320 billion parameters. Z.ai originally previewed the model under the code name Ox Alpha on August 20 and officially released it under the MIT open source license on August 26. The model is specialized for long-context processing, vision-driven agentic tasks and code synthesis.
GLM-5.3 Flash costs $0.075 per million input tokens and $0.25 per million output tokens from now until Sept. 9. After that, the price will be $0.15 per million input tokens and $0.50 per million output tokens. Comparatively, GPT-5.6 Luna from OpenAI costs $0.20 per million input tokens and $2 per million output tokens. Anthropic Claude Opus 5 is $5 per million input tokens and $5 per million output tokens.
The Chinese vendor’s decision to use purely Chinese chip providers comes as China aims to rely less on Nvidia. China’s pivot toward self-reliance comes after years in which the U.S. tightened export controls to prevent Chinese tech vendors from obtaining the most powerful Nvidia chips. Although some of these controls have eased, in September 2025, China’s government ordered tech giants such as Alibaba and ByteDance to stop buying Nvidia AI chips, and in November, Beijing banned all foreign AI chips from state-funded data centers.
Despite these moves, many observers still saw China as behind in the infrastructure contest. However, Z.ai’s strategy suggests that the gap between China and the U.S. in AI chips may be narrowing, especially with the increased focus on inference over the past year amid the sharp rise of agentic AI.
“For training the largest frontier models, Nvidia still has a significant advantage in chip performance, networking and the broader software ecosystem,” said Kashyap Kompella, CEO and founder of RPA2AI Research. “But inference is a different story.”
He added that Z.ai’s ability to serve GLM-5.3-Flash at scale using Chinese chips indicates that Chinese hardware is already good enough for many high-volume inference workloads.
“This distinction is important because inference is likely to become the larger AI chip market over time,” Kompella said. “Chinese companies may not need exact chip-for-chip parity with Nvidia if they can compensate through efficient model architectures, software optimization and large clusters of domestically made AI chips.”
However, the first signals of Chinese vendors’ move toward independence began with DeepSeek rewriting Nvidia CUDA drivers to make inference faster, said Bradley Shimmin, an analyst at Futurum Group.
He said that while Chinese vendors are responding to constraints first imposed by U.S. government and later by the Chinese government, another factor is that companies “need to do more with the watt hours they have available to them.”
Vendors such as Google, with its TPU series and Amazon, with its Trainium chips, are focusing on vertical integration, in which the AI stack is optimized to use AI models, infrastructure and software together for the most efficient AI processing.
“What we’re witnessing here is this real realization that AI value is driven by economics of the stack, as much as it is about the capabilities of the models,” Shimmin said. “What these guys are doing is simply showcasing the value of a little bit of investment in optimizing that hardware.”
As for GLM-5.3-Flash itself, the fact that it received strong developer mindshare in a blind test is significant, Kompella said. Many developers flocked to the model early this month and used it without knowing it came from a Chinese vendor; it quickly became the most downloaded model on OpenRouter.
“That is more interesting than another benchmark claiming that a Chinese model is close to a U.S. frontier model,” Kompella said, adding that it shows that Chinese open-weight models “continue to function as a counterweight to higher industry pricing.”
For enterprises, this means they need to design their AI system “so workloads can be routed across models based on capability, quality, cost and strategic considerations rather than becoming dependent on a single model provider,” he continued.
Although the model attracted a large number of developers right away is a good sign for Z.ai, the vendor still faces challenges ahead. For one, it must prove that the chips it is using can continue to deliver reliability, economy and scale, Kompella said.
“At the absolute frontier of model training, access to compute also remains a meaningful constraint because Nvidia’s advantage is much harder to replicate there,” he said.
The vendor will also need to convert U.S. and European enterprises that remain cautious about using Chinese-hosted AI services due to data security, regulatory and procurement concerns, Kompella continued.
文章标题:Z.AI在新模型中使用中国芯片,关键在于优化。
文章链接:https://news.qimuai.cn/?post=4913
本站文章均为原创,未经授权请勿用于任何商业用途