微软通过AMD扩展Azure AI与高性能计算基础设施

内容总结:
微软推出全新Azure虚拟机,强化AI基础设施布局
随着AI工作负载的扩展速度超越单一基础设施架构的承载能力,微软正通过深化与AMD的合作,进一步优化其云端AI算力布局。近日,微软宣布将在Azure平台引入AMD最新的Helios AI平台及下一代EPYC数据中心处理器,并推出三款全新虚拟机产品:HDv2(数据处理)、HXv2(电子设计自动化)及ND MI455X v7(AI推理),旨在满足从模型训练、推理到芯片设计等多样化场景对高性能、高能效计算的需求。
HDv2虚拟机:专为AI数据系统打造
AI系统的高效运行离不开高性能CPU基础设施。新推出的Azure HDv2虚拟机由微软与AMD联合设计,搭载近500颗第六代AMD EPYC CPU核心、4TB内存、32TB本地NVMe存储及400Gb Azure Boost网络,可显著消除数据处理瓶颈,支撑大规模数据准备、搜索、强化学习及智能体协同等AI工作负载。
HXv2虚拟机:优化芯片设计与技术计算
基于2023年发布的HX系列,HXv2虚拟机进一步提升了单线程性能与内存容量。其配备176颗主频超5GHz的第六代AMD EPYC核心,每核心可寻址缓存增加50%,并提供接近2TB或4TB内存选项。该机型支持RTL仿真、科学模拟及工程分析等分布式内存应用,同时集成800Gb InfiniBand互联,适用于大规模MPI并行计算。AMD自身也在使用HX系列进行EPYC CPU及Instinct GPU设计,其首席技术官Mark Papermaster表示:“HXv2旨在为全球最严苛的工程与科学计算提供更强性能与可扩展性。”此外,HXv2还优化了Synopsys的AI驱动EDA解决方案,帮助客户缩短芯片设计周期。
ND MI455X v7虚拟机:规模化AI推理
专为现代AI服务中的推理、搜索及智能体工作负载设计,ND MI455X v7基于AMD Helios机架级方案,可提供高效的大规模推理能力,进一步丰富Azure在AI推理场景中的基础设施选择。
微软表示,上述新产品的推出体现了“客户选择”的核心设计原则,通过提供开放、异构的云端平台,帮助企业在性能、成本与能效之间取得最佳平衡。更多详情可访问Azure官网。
中文翻译:
人工智能工作负载的扩展速度已超出任何单一基础设施架构的承载能力——模型数量增多、新型智能体驱动的工作负载涌现以及计算需求激增,都促使全栈各层面需要更高程度的专业化。为满足这一需求,微软持续演进 Azure 的基础设施,包括引入 AMD 最先进的 AI 和高性能计算(HPC)解决方案来扩充其 AI 计算集群。
我们设计 AI 基础设施的方法旨在全面支撑 AI 系统的构建与运行方式。我们不仅与 AMD 等行业创新者紧密合作,还结合自身定制的芯片与系统,为客户提供一个全面、开放且异构的平台,以实现最佳的性能、成本与能效产出。
基于与 AMD 的深度合作,微软即将在 Azure 上部署 AMD 最新的 Helios AI 平台和下一代 EPYC 数据中心处理器。这些技术将为三项即将推出的 Azure 服务提供动力:用于数据处理的 HDv2 虚拟机、用于电子设计自动化(EDA)的 HXv2 虚拟机,以及用于 AI 推理工作负载的 ND MI455X v7 虚拟机。
扩展面向推理、AI 数据系统与芯片设计的基础设施
专为 AI 数据系统打造——Azure HDv2
CPU 基础设施对于现代 AI 系统的性能与效率至关重要。AI 加速器需要高密度、高能效的 CPU 计算能力来处理数据、协调工作负载并保持流水线大规模运行。否则,训练任务将缺乏足够的数据进行学习,智能体也将没有足够的能力代表客户执行任务。Azure HDv2 虚拟机是我们最新推出的产品之一,旨在从零开始消除这些瓶颈,并支持大规模智能体工作负载的采用。
HDv2 虚拟机由微软与 AMD 联合设计,扩展了 Azure 针对 AI 客户最苛刻 CPU 工作负载(包括数据准备、搜索、强化学习和大规模智能体协调)的专用解决方案组合。HDv2 虚拟机配备近 500 个第六代 AMD EPYC CPU 物理核心、4TB 内存、32TB 本地 NVMe 存储以及 400Gb Azure Boost 网络,专为满足我们要求最严苛的 AI 客户的工作负载需求而构建。
针对芯片设计与技术计算优化——Azure HXv2
人工智能时代为开发支撑该基础设施的芯片产品的公司创造了巨大的需求和机遇。因此,于 2023 年与 AMD 合作推出、并采用 AMD 独特的 3D V-cache 技术的 Azure HX 虚拟机,在致力于将功能更强、效率更高的 AI 芯片推向市场的芯片设计公司中获得了广泛采用。今天,我们宣布为这些客户推出工作负载优化之旅的下一步——HXv2。
HXv2 虚拟机延续并扩展了 HX 的优势。两者都通过再次采用 3D V-cache 技术,继续为 RTL 仿真工作负载提供 Azure 的差异化优势,同时在单线程性能和内存方面实现了显著提升。HXv2 虚拟机将配备 176 个 AMD 第六代 EPYC CPU 核心,时钟频率超过 5GHz,每核心可寻址缓存增加 50%,并提供近 2TB 或 4TB 内存的虚拟机规格,帮助客户针对内存需求优化其工作负载。
Azure HXv2 还旨在支持更广泛的技术计算工作负载,包括科学模拟、工程分析和其他分布式内存应用。虚拟机级别和核心级别的性能显著提升,再加上 800Gb InfiniBand 的引入,使得基于 MPI 的大规模模拟成为可能,让 HXv2 成为各类 HPC 客户的理想选择。
作为 HX 系列的主要客户,AMD 直接强调了这一影响:
“工程团队正在不断突破模拟、芯片设计和科学计算的极限。在 AMD,我们在设计未来 AMD EPYC CPU 和 AMD Instinct GPU 时,亲身感受到了这些需求。Azure HX 是扩展复杂 EDA 工作负载的重要平台,我们对旨在提供更强性能和可扩展性的 Azure HXv2 感到兴奋。我们期待继续与微软合作,共同为世界上最苛刻的工程和科学工作负载推进基础设施建设。”
—— Mark Papermaster,AMD 执行副总裁兼首席技术官
HXv2 还利用了微软与 Synopsys 的长期合作关系,在 Azure 上优化其 AI 驱动的 EDA 解决方案:
“随着 AI 计算不断突破半导体设计的极限,我们与微软在 Azure HX 系列上的合作展现了一个共同的愿景:使客户能够在加速的设计周期中,以精准和大规模的方式交付下一代 AI 系统。这些系统使 Synopsys 客户能够可靠且高效地利用基于云的计算,将 EDA 工作负载延伸至传统基础设施的限制之外,从而在满足雄心勃勃的开发计划的同时,最大化设计质量并带来显著的性能提升。”
—— Shankar Krishnamoorthy,Synopsys 首席产品开发官
生产级 AI 推理——ND MI455X v7
ND MI455X v7 专为现代 AI 服务背后的推理、搜索和智能体工作负载而设计。它由 AMD Helios 机架级解决方案提供动力,扩展了 Azure 用于大规模推理的基础设施选项,旨在为苛刻的 AI 工作负载提供强大的性能和效率。
这些新能力共同扩展了 Azure 的功能,同时为客户提供了更大的灵活性,以便为每个独特的 AI 工作流(从推理、数据系统到芯片设计)选择正确的计算资源。客户选择权是内置于 Microsoft Azure 的核心设计原则,我们很高兴能将 AMD 最先进的创新以生产规模推向市场。
要了解有关 Azure 高性能计算和 AI 基础设施功能的更多信息,请访问 Azure.com。
Scott Guthrie 负责一系列超大规模云计算解决方案和服务,包括微软云计算平台 Azure、生成式 AI 解决方案、数据平台以及信息和网络安全。这些平台和服务帮助全球各地的组织解决紧迫的挑战,并为未来进行转型。
英文来源:
AI workloads are scaling faster than any single infrastructure approach can support — with more models, new agent-driven workloads and surging compute demand driving the need for greater specialization across the stack. To meet this need, Microsoft continues to evolve Azure’s infrastructure, including expanding its AI fleet with AMD’s most advanced AI and high-performance computing (HPC) solutions.
Our approach to AI infrastructure is designed to support the breadth of how AI systems are built and run. We closely work with industry innovators like AMD as well as our own purpose-built silicon and systems to provide customers with a comprehensive, open and heterogenous platform to achieve the best performance, cost and energy efficiency outcomes.
Building on our close collaboration with AMD, Microsoft is bringing AMD’s latest Helios AI platform and next-generation EPYC datacenter processors to Azure. These technologies will power three upcoming Azure offerings: HDv2 VMs for data processing, HXv2 VMs for electronic design automation (EDA) and ND MI455X v7 VMs for AI inference workloads.
Expanded infrastructure for inference, AI data systems and chip design
Built for AI data systems — Azure HDv2
CPU infrastructure is essential to the performance and efficiency of modern AI systems. AI accelerators depend on high-density, power-efficient CPU compute to process data, coordinate workloads and keep pipelines running at scale. Without this, training jobs don’t have enough data to learn from, and agents don’t have enough capacity to perform tasks on behalf of customers. Azure HDv2 virtual machines are one of our latest offerings designed from the ground up to eliminate these bottlenecks and empower massive agentic workload adoption.
Co-designed with AMD, HDv2 VMs expand Azure’s portfolio of purpose-built solutions for the most demanding CPU workloads from AI customers, including data preparation, search, reinforcement learning and agent coordination at scale. Featuring nearly 500 physical 6th Gen AMD EPYC CPU cores, 4 terabytes of RAM, 32 terabytes of local NVMe storage and 400 Gb Azure Boost networking, HDv2 VMs are built for the workload needs of our most demanding AI customers.
Optimized for silicon design and technical computing — Azure HXv2
The AI era has created tremendous need and opportunity for firms developing the silicon products that power this infrastructure. For this reason, Azure HX virtual machines, launched in partnership with AMD in 2023 and featuring AMD’s unique 3D V-cache technology, have seen significant adoption among silicon design firms working to bring more capable and efficient AI silicon to market. Today, we are announcing the next step in our workload optimized journey for these customers, HXv2.
HXv2 virtual machines build on and extend the strengths of HX. They both continue the differentiation Azure offers for RTL simulation workloads by again employing 3D V-cache technology, while offering significant improvements to single threaded performance and memory. HXv2 VMs will feature 176 AMD 6th Gen EPYC CPU cores with a clock frequency of more than 5 GHz, 50% more addressable cache per core and VM sizes with nearly 2 or 4 terabytes of RAM, helping customers optimize their workloads to memory needs.
Azure HXv2 is also designed to support a broader range of technical computing workloads including scientific simulation, engineering analysis and other distributed memory applications. The significantly increased per VM and per core performance, and the inclusion of 800 Gb InfiniBand, enable large-scale MPI-based simulations and make HXv2 an ideal fit for a wide variety of HPC customers.
AMD, a leading HX-series customer, highlights this impact directly:
“Engineering teams are pushing the limits of simulation, chip design and scientific computing. At AMD, we experience those demands firsthand as we design future AMD EPYC CPUs and AMD Instinct GPUs. Azure HX is an important platform for scaling complex EDA workloads, and we’re excited about Azure HXv2, which is designed to deliver even greater performance and scalability. We look forward to continuing our collaboration with Microsoft as we help advance infrastructure for the world’s most demanding engineering and scientific workloads.”
— Mark Papermaster, Executive Vice President and CTO, AMD
The HXv2 also leverages Microsoft’s long-standing collaboration to optimize Synopsys AI-powered EDA solutions on Azure:
“As AI compute continues to push the limits of semiconductor design, our collaboration with Microsoft on the Azure HX-series demonstrates a shared vision for enabling customers to deliver next-generation AI systems with precision and scale in accelerated design cycles. These systems have enabled Synopsys customers to reliably and efficiently leverage cloud-based compute, extending EDA workloads beyond traditional infrastructure constraints so they can meet ambitious development schedules while maximizing design quality and delivering dramatic performance gains.”
— Shankar Krishnamoorthy, Chief Product Development Officer, Synopsys
Production-scale AI inference — ND MI455X v7
ND MI455X v7 is designed for the reasoning, search and agentic workloads behind modern AI services. Powered by the AMD Helios rackscale solution, it expands Azure’s infrastructure options for large-scale inference and is designed to deliver strong performance and efficiency for demanding AI workloads.
Together, these new capabilities expand Azure capabilities while giving customers more flexibility to choose the right compute for each unique AI workflow: from inference, to data systems, to chip design. Customer choice is a core design principle built directly into Microsoft Azure, and we’re excited to bring AMD’s most advanced innovations at production scale.
To learn more about Azure’s high-performance computing and AI infrastructure capabilities, visit Azure.com.
Scott Guthrie is responsible for a set of hyperscale cloud computing solutions and services including Azure, Microsoft’s cloud computing platform, generative AI solutions, data platforms and information and cybersecurity. These platforms and services help organizations across the globe solve urgent challenges — and transform for the future.
文章标题:微软通过AMD扩展Azure AI与高性能计算基础设施
文章链接:https://news.qimuai.cn/?post=4616
本站文章均为原创,未经授权请勿用于任何商业用途