微软框架旨在降低AI代理训练成本

内容来源:https://aibusiness.com/agentic-ai/microsoft-framework-cut-ai-agent-training-costs
内容总结:
谷歌云赞助:生成式AI首批应用场景的选择之道
在生成式人工智能的起步阶段,企业应优先聚焦于能够改善人类信息获取体验的领域。这是当前行业给出的核心建议,旨在通过技术落地切实提升效率与价值。
与此同时,AI服务市场正呈现出对低成本解决方案的强劲需求。为顺应这一趋势,科技巨头微软研究院于本周一正式发布了名为“Orchard”的开源框架。该框架旨在降低训练自主AI代理的难度与成本,使得这一过程更加便捷高效。此举是微软自今年三月首次就该主题发表研究以来的持续成果。
Orchard框架能够跨多种任务训练和评估AI代理,涵盖编程、网页浏览及工具使用等。其最大亮点在于通过消除研究人员为单个模型和用例搭建沙盒基础设施、数据管道及评估系统的繁琐需求,极大地简化了开发流程。
微软在官方博客中指出,尽管业界对代理式AI的能力充满期待,但研究社区始终面临一个持续性的瓶颈:构建最先进的代理系统往往依赖专有的基础设施,而这对于大多数研究人员和从业者而言,既难以获取也无法复现。
为破解这一难题,微软推出了Orchard。该框架的核心组件是“Orchard Env”,微软将其描述为一个“轻量级的Kubernetes环境,提供可复用的隔离组件”。这些组件可用于大规模运行和构建代理,涵盖从收集训练数据到强化学习 rollouts 及评估的各个环节,而无需重建底层架构。
为展示其实用性,微软同步发布了三个特定领域的训练配方——分别是面向软件工程代理的Orchard-SWE、专注于浏览器导航的Orchard-GUI,以及覆盖日常生产力任务的Orchard-Claw,并附带了构建这些配方所用的训练数据和评估方法。微软表示,通过这些配方训练出的模型在既有基准测试中均表现出较强的竞争力。
微软强调,这一成果意义重大。其在博客中称:“通过使底层基础设施开放、轻量且可复用,Orchard显著降低了代理式AI研究的门槛。团队无需再从头构建自定义隔离环境,或依赖专有云服务。” 展望未来,微软认为将训练经验复用是实现累积式代理学习的可行方向。他们设想,不再在每次训练结束后丢弃轨迹数据,而是将其视为持久资产,例如提炼为可复用的价值模型。
中文翻译:
由谷歌云赞助
选择你的首个生成式AI应用场景
要开始使用生成式AI,首先要聚焦于那些能够改善人类获取信息体验的领域。
该供应商正致力于抓住市场对低成本AI服务日益增长的需求。
微软研究院发布了Orchard,这是一个开源框架,旨在让训练自主AI智能体变得更简单、更具成本效益。
这家科技巨头的研究部门于周一发布了这一成果,此前该供应商自今年3月首次发表相关主题论文以来一直在持续推进这项工作。
Orchard可以在多项任务上训练和评估智能体,包括编码、网页浏览和使用工具,同时通过消除研究人员为单个模型和应用场景构建沙盒基础设施、数据管道和评估系统的需求来降低复杂性。
“尽管人们对智能体AI的能力感到兴奋,但研究界仍面临着一个持续的瓶颈。构建最先进的智能体系统通常需要专有基础设施……而大多数研究人员和从业者无法访问或复现这些设施,”该供应商在一篇博文中表示。
微软创建Orchard就是为了解决这个问题。这一新框架的核心是Orchard Env,微软将其描述为“一个轻量级的Kubernetes环境,提供可复用的隔离组件”。
这些组件可用于大规模运行和构建智能体,从收集训练数据到强化学习部署和评估,而无需重建底层架构。
为了展示其方法,微软发布了三个特定领域的训练配方——Orchard-SWE、Orchard-GUI和Orchard-Claw——以及用于构建它们的训练数据和评估方法。
Orchard-SWE用于训练软件工程智能体,Orchard-GUI专注于浏览器导航,Orchard-Claw则涵盖日常生产力任务。微软表示,通过每种配方训练的模型在既有基准测试上均表现出竞争力。
据微软称,这些结果可能具有深远意义。该供应商在博文中表示:“通过使底层基础设施开放、轻量且可复用,Orchard降低了智能体AI研究的成本。团队不再需要从零开始构建自定义隔离环境,也不必依赖专有云服务。”
“展望未来,我们认为复用训练经验是实现智能体累积学习的一个有前景的方向。我们不再在训练运行结束后丢弃轨迹,而是将其视为持久资产——例如,将它们提炼为可复用的价值模型。”
英文来源:
Sponsored by Google Cloud
Choosing Your First Generative AI Use Cases
To get started with generative AI, first focus on areas that can improve human experiences with information.
The vendor is aiming to capitalize on the growing demand for lower-cost AI services.
Microsoft Research released Orchard, an open source framework that aims to make it easier and more cost-effective to train autonomous AI agents.
The release by the tech giant’s research division on Monday follows ongoing work since the vendor’s initial publication on this topic in March.
Orchard can train and evaluate agents across several tasks, including coding, web browsing and using tools, while reducing complexity by eliminating the need for researchers to build sandbox infrastructure, data pipelines and evaluation systems for individual models and use cases.
“While there is excitement around agentic AI’s capabilities, the research community faces a persistent bottleneck. Building state-of-the-art agentic systems often requires proprietary infrastructure … that most researchers and practitioners cannot access or reproduce,” the vendor said in a blog post.
Microsoft created Orchard to address this problem. At the core of the new framework is Orchard Env, which Microsoft describes as a “lightweight, Kubernetes environment that provides reusable isolated components".
These can be used to run and build agents at scale, from collecting training data to reinforcement learning rollouts and evaluation, without rebuilding the underlying architecture.
To demonstrate its approach, Microsoft released three domain-specific training recipes -- Orchard-SWE, Orchard-GUI, and Orchard-Claw -- along with the training data and evaluation methods used to build them.
Orchard-SWE trains software-engineering agents, Orchard-GUI focuses on browser navigation, and Orchard-Claw covers everyday productivity tasks. Results published for models trained through each showed they were competitive on established benchmarks, Microsoft said.
According to Microsoft, the results could have significant implications. It said: “By making the underlying infrastructure open, lightweight, and reusable, Orchard lowers the cost of agentic AI research,” the vendor said in the blog. “Teams no longer need to build custom isolated environments from scratch or depend on proprietary cloud services.”
“Looking ahead, we see reusing training experience as a promising direction toward cumulative agent learning. Instead of discarding trajectories once a training run finishes, we treat them as persistent assets -- for example, distilling them into reusable value models.”