作者介绍了新的人工智能模型,并升级了相关配套机制以控制代币成本。

内容总结:
人工智能行业正面临部署成本攀升的压力,用户对降低开支的需求日益迫切。尽管开源模型在单位令牌成本上更具优势,但如何精准匹配具体任务仍是难题。本周四,面向营销人员提供AI工具与智能体的公司Writer正式发布旗舰模型Palmyra X6,旨在解决这一痛点。该模型基于Z.ai的开源模型GLM-5.2进行后训练优化,官方称其能在更低价格下提供可直接部署的能力。据估算,新模型结合公司智能体框架的调整,可将用户执行基础任务的成本降低至多50%。
与新款模型同步推出的还有Writer标准智能体框架的重大升级。上述功能自本周四起向所有客户开放。公司首席执行官May Habib向TechCrunch表示:“企业早已厌倦追逐下一个基准分数。他们渴望成本持续下降,而目前似乎无人能做到这一点。”
此次发布的核心在于优化复杂多步任务的执行效率,以更少令牌实现更快响应。Writer视智能体框架优化为实现降本的关键杠杆。其研究人员近期发布的论文验证了这一思路:通过在多款模型中微调框架效率,结果显示多数情况下框架调整比更换模型更能稳定降低成本,测试中平均降幅达40%。研究团队在论文中指出:“框架是唯一能对组织运行的所有模型(包括现有及未来模型)产生倍增效率的组件。”
对于Writer客户而言,使用体验依然保持模型无关性:Palmyra X6将与Writer其他模型或通过Azure、Amazon Bedrock导入的外部模型并列运行。Habib同时认为,降本压力正加剧企业对大型AI实验室的不信任——后者存在驱动令牌消耗量增长的经济动机。她补充道:“成本暴涨对客户而言史无前例,首席信息官们对实验室的失望程度同样罕见。这些实验室并未真正理解如何帮助企业从AI中获益。”
中文翻译:
在整个AI行业,用户越来越清楚地意识到自己的部署可能有多昂贵——并且感受到削减成本的新的紧迫性。然而,尽管开源模型提供了显著更低的单token成本,但为特定任务找到合适的模型可能并不容易。
周四,为营销人员提供AI工具和智能体的Writer公司推出了一款名为Palmyra X6的新旗舰模型,旨在为其用户解决这一问题。作为基于Z.ai开源模型GLM-5.2的后训练变体,Writer表示,新系统应以低得多的价格提供可直接部署的能力。该公司估计,这款新模型加上对其智能体基础设施的改动,将为其客户在基础任务上削减多达50%的成本。
除了新模型,该公司还发布了其标准智能体框架的重大升级。这两项功能将于周四起对Writer客户开放。
“我认为企业界已经厌倦了追逐下一个基准,”CEO May Habib告诉TechCrunch。“他们想要成本持平,而似乎没人能做到这一点。”
新方法特别强调复杂、多步骤的任务,以更快的速度和更少的token完成。而Writer认为,框架优化是实现这一目标的关键杠杆。
Writer的研究人员最近发表的一篇论文为这一方法提供了支持,他们在多种不同模型上测试了框架效率的微小改动。研究发现,在许多情况下,框架的改动比模型选择更能可靠地降低成本,在测试中成本平均下降了40%。
“框架是这样一个组件,其效率会在组织运行的每一个模型上成倍放大——无论是现在的还是未来的,”研究人员写道。
对于Writer的客户而言,体验仍然是模型无关的:Palmyra X6将与其他Writer模型或通过Azure或Amazon Bedrock导入的外部模型并列运行。但Habib也认为,削减成本的推动力正在引发对大型AI实验室的更广泛不信任,因为这些实验室有经济利益驱动token使用量的增长。
“对客户来说,这里的成本爆炸是前所未有的,CIO们放弃这些实验室的程度也是前所未有的,”Habib告诉TechCrunch,并补充说,AI实验室“并不深入理解如何帮助企业从AI中获得收益。”
英文来源:
Across the AI industry, users are becoming more conscious of just how expensive their deployments can be —and feeling a new urgency to cut costs. But while open source models offer significantly lower per-token costs, it can be difficult to find the right model for a given job.
On Thursday, Writer, which offers AI tools and agents for marketers, launched a new flagship model called Palmyra X6, aimed at solving that problem for its users. Built as a post-training variation on Z.ai’s open source model GLM-5.2, Writer says the new system should provide deployment-ready capabilities at a much lower price. The company estimates the new model, combined with changes to the companies harness infrastructure, will cut costs for its customers by as much as 50% for basic tasks.
Together with the new model, the company also released significant upgrades to its standard agentic harness. Both features will be available to Writer clients starting Thursday.
“I think the enterprise is absolutely sick of chasing the next benchmark,” CEO May Habib told TechCrunch. “They want flattening cost, and it seems like nobody can deliver that.”
The new approach puts particular emphasis on complex, multi-step tasks, executed faster and with fewer tokens. And Writer sees harness optimization as a crucial lever toward making that happen.
A recent paper from Writer researchers lends credence to this approach, testing small changes in harness efficiency across multiple different models. The research found that, in many cases, changes in the harness were a more reliable way to reduce costs than model choice, with costs falling an average of 40% across their testing.
“The harness is the one component whose efficiency multiplies across every model an organization runs—present and future,” the researchers wrote.
For Writer’s clients, the experience is still model-agnostic: Palmyra X6 will sit alongside other Writer models or outside models imported through Azure or Amazon Bedrock. But Habib also sees the push to cut costs as driving a broader distrust toward major AI labs, which have a financial incentive to drive up token use.
“The cost explosion here is just unprecedented for customers, and so is the degree to which CIOs are giving up on the labs,” Habib told TechCrunch, adding that the AI labs “don’t deeply understand right how to help an enterprise get benefit from AI.”
文章标题:作者介绍了新的人工智能模型,并升级了相关配套机制以控制代币成本。
文章链接:https://news.qimuai.cn/?post=4806
本站文章均为原创,未经授权请勿用于任何商业用途