在人工智能时代构建内存与存储架构

内容来源:https://www.technologyreview.com/2026/09/04/1140872/architecting-memory-and-storage-in-the-ai-era/
内容总结:
AI推理时代来临:企业需重构内存与存储架构以释放AI潜力
随着AI推理(AI inference)日益成为企业级工作负载的核心驱动力,传统的数据中心基础设施正面临严峻挑战。行业专家指出,仅靠堆砌高性能计算硬件已无法满足实时AI服务的需求,企业必须从系统层面出发,重新思考内存、存储与网络的协同架构,以平衡速度、效率、可扩展性及每瓦性能。
从“单一算力”到“系统协同”
Tirias Research创始人兼首席分析师Jim McGregor强调,AI并非单一工作负载,而是成千上万乃至数十亿种不同任务的集合。AI推理改变了优化问题的本质——从单纯追求计算峰值转向内存、存储与网络的整体协同。他表示:“当前最大的挑战在于如何高效地将数据从一个位置移动到另一个位置,并确保我们能有效利用它。”传统IT架构基于相对稳定的负载假设,而推理与智能体AI(agentic AI)对延迟、数据流动和系统利用率提出了全新要求,使得架构选择比以往任何时候都更具决定性。
数据流动成为新瓶颈
随着企业部署高级推理系统,实时查询的数据量激增,使得数据搬运成为最紧迫的制约因素。现代AI技术如检索增强生成(RAG)需要系统不断扫描海量数据库以生成准确响应,这不仅要求强大的计算力,更要求数据的即时可及性。内存带宽、缓存效率、存储 proximity 以及信息检索的一致性,都直接影响AI服务的质量。在机器人、金融服务、医疗健康和客户交互等领域,延迟已不仅是技术瑕疵,更关乎安全、响应速度与用户信任。
构建面向未来的AI基础设施
专家建议,企业规划AI基础设施时,应从以下几方面构建采购与设计框架:
- 明确优化目标:基础设施选型必须匹配具体业务需求,而非追求空泛的“AI就绪”,避免资金浪费与瓶颈并存。
- 采用模块化架构:在计算、内存、存储、供电与散热方面保持灵活性,避免过早锁定僵化设计。
- 深化生态合作:与供应商和集成商紧密协作,降低供应链风险,获取关键组件。
- 持续评估采购策略:AI需求、硬件和商业模式变化迅速,固定的长期规划难以适应。
- 兼顾效率与投资回报:追求极致性能可能导致成本过高。同时,提升资源利用率和能效(如功耗、用水)也是企业应对公众审视的重要指标。
内存与存储:从后台仓库到战略资产
美光科技(Micron)作为本内容的合作伙伴,其观点与行业分析一致:在AI推理时代,内存和存储不再是被动的数据仓库,而是AI系统的“活性血液”。企业若想从AI中获得最大价值,未必需要最大的计算集群,而在于能否将每一层基础设施精准对齐业务成果、消除数据瓶颈,并构建能随工作负载演进而灵活调整的体系。
Jim McGregor总结道:“采购如今已成为战略问题,系统设计则是领导层议题。每个高管都必须自问:AI将如何改变我的商业模式?”未来的竞争优势,将属于那些把计算、内存、存储和网络视为一个集成系统来设计的企业——以高效、可规模化且能带来可衡量回报的方式交付AI能力。
本文内容由MIT Technology Review Insights(定制内容部门)撰写,非编辑部出品,由人工调研与写作完成。
中文翻译:
赞助内容
AI时代的存储与内存架构设计
随着AI推理如今驱动企业工作负载,各组织必须重新思考基础设施,以在速度、效率、可扩展性和每瓦性能方面实现突破,从而释放AI在现实世界中的潜力。
与美光科技联合呈现
AI推理时代已经到来。想象一个医疗系统实时分析数百万个数据点以加速挽救生命的医学研究,或者一个智能助手同时即时解决数千个复杂的客户需求。这些现实世界的突破依赖于先进的基础设施,它充当持续智能的引擎,为实时服务提供动力,同时支持日益智能化的物联网和消费设备边缘。然而,在这个由推理驱动的格局中,每一次延迟、每一个瓶颈或每一瓦电力的浪费都直接影响到人类的成果和运营成本。
这种转变改变了基础设施必须交付的价值。性能、延迟、内存带宽、存储吞吐量和网络不能各自为政地优化。推理工作负载是持续性的、地理分布式的,并且对响应时间高度敏感,要求系统从一开始就为规模、韧性和效率而设计。
“我们倾向于将AI视为单一工作负载,但它不是。它是成千上万、数百万、数十亿种不同的工作负载,”Tirias Research创始人兼首席分析师吉姆·麦格雷戈表示。AI推理将优化问题从单纯的计算能力转变为协调的基础设施——内存、存储和网络。
对于企业领导者来说,优先级是明确的:AI基础设施决策必须在成本、灵活性和未来准备度之间取得平衡。赢家将是那些提升每瓦性能、减少环境足迹,并在内存和存储瓶颈限制增长之前将其消除的组织。
AI推理需要新的架构方法
AI系统需要重新架构,因为将现代AI系统硬塞进传统基础设施会限制AI的变革潜力。专用架构对于实现AI的真正价值至关重要,无论是加速科学发现还是创造真正自主的数字代理。
传统企业IT一直能够依赖相对稳定的基础设施假设,但推理和代理式AI带来了围绕延迟、数据移动、可扩展性和利用率的新需求,使得架构选择的影响远比以往更加深远。
“数据中心现在必须支持持续的、分布式的、且日益实时的AI服务——这些都不是单一工作负载,”麦格雷戈说。“从系统层面来看,它们都有不同的需求。”
为了支持实时AI,企业不能再将内存和存储仅仅视为辅助硬件,而应将其视为系统的核心。各组织需要构建一条能够快速摄取、清洗、转换、存储、移动和交付数据的数据管道。推理工作负载以与早期以训练为中心的部署截然不同的方式对基础设施施加持续压力,要求持续的数据检索和缓存,这是传统应用从未需要的。
相应地,性能本身不再是最重要的唯一基准。企业越来越需要在性能与效率、成本和可扩展性之间取得平衡,尤其是在试图支持不同AI服务而不为峰值条件过度建设基础设施的情况下。
“你必须围绕计划运行的工作负载类型来优化整个网络,包括内存和存储,”麦格雷戈说。“你必须真正详细了解这些工作负载将是什么样子的。”
任何AI基础设施战略都必须从工作负载认知开始。推理、代理式AI和其他新兴AI用例要求各组织将数据中心视为一个集成系统。
数据移动是新的瓶颈,也是竞争优势的机遇
随着企业部署先进的推理和代理系统,实时查询的数据量之大已使数据移动成为最紧迫的约束。诸如检索增强生成(RAG)等现代AI技术要求系统不断扫描大规模数据库以生成准确的响应。这需要巨大的计算能力,但更重要的是,它需要即时访问数据。
麦格雷戈表示,关注点转向如何在更广泛的架构中高效地移动、缓存和交付数据,这使内存和存储从后台基础设施提升为战略资产。“我们现在做的最重要的事情就是将数据从一个地方移动到另一个地方,并确保我们能够有效地使用它。”
由于AI不是一个单一的工作负载类别,仅仅购买最快的处理器是不够的。推理高度依赖内存带宽、缓存、存储邻近性,以及快速且一致地检索相关信息的能力。理解每种资源在技术栈中的位置以及这些层在实际运行条件下如何交互,已成为一项业务要务。
麦格雷戈表示,最有效的AI基础设施看起来不像是一堆顶级部件的集合,而更像是一个由计算、内存、存储和网络组成的均衡系统,因为瓶颈往往会从一层迁移到另一层。“你必须将这四者一起架构才能实现高效,这就是挑战所在。”
数据平面设计与网络带宽的相互依存意味着AI基础设施规划已成为一项业务决策,与工程决策同等重要:延迟如今与价值密不可分。在机器人技术、金融服务、医疗保健和面向客户的AI系统中,延迟不仅仅是技术缺陷;它们可能损害安全性、响应性或信任。AI基础设施性能成为声誉管理的问题。
从AI中获得最多收益的组织可能不是拥有最大集群的那些,而是那些最清楚如何协调每个基础设施元素以有效执行AI工作负载的组织。
构建AI基础设施采购框架
规划AI基础设施不仅仅是选择最快的硬件。它关乎如何在不让组织被可能迅速过时的假设所束缚的情况下进行扩展。“你需要保持灵活,因为需求会迅速变化,技术也在迅速变化,”麦格雷戈说。
面向未来的AI基础设施需要在工作负载、经济性和架构持续变化时保持选择的开放性:
- 明确正在优化的AI工作负载。基础设施选择必须匹配业务需求,而非麦格雷戈所称的泛泛的“AI准备度”,后者可能导致在某些领域超支,同时在其他领域留下未解决的瓶颈。
- 为计算、内存、存储、电力和冷却构建模块化架构,以便容量能随需求变化而调整,而不是过早承诺僵化的架构。
- 与完整的供应商和集成商生态系统合作,以降低供应风险并改善获取正确组件的渠道。麦格雷戈表示,买家不能再假设仅靠其OEM或云服务提供商就能使他们免受供应限制或架构复杂性的影响。
- 持续重新评估采购策略。AI需求、硬件和商业模式变化太快,无法采用固定的长期设计。
- 优化效率和投资回报率,而不仅仅是峰值性能。最强大的配置可能过于昂贵而难以维持。效率也是一个面向公众的指标——更好的利用率和更具工作负载感知的系统设计可以帮助企业应对围绕电力消耗和水资源使用日益增长的审视。
更智能的AI数据中心设计的战略目标不是不计代价地追求最大性能,而是一种能够交付价值、吸收变化并证明其存在合理性的适应性架构。
AI基础设施如今是一项业务战略
AI数据中心已迅速从后端技术问题演变为战略性业务系统,帮助决定一个组织能否有效地将AI转化为收入、改善人类成果并创造竞争优势。
在推理时代,内存和存储不再是被动存储库,麦格雷戈解释说,它们是AI活跃的生命线。从AI中获得最多收益的组织未必是计算足迹最大的那些,而是那些将基础设施投资与业务成果对齐、减少数据瓶颈并建立适应工作负载演变的灵活性的组织。他预测,竞争优势将越来越多地属于那些将计算、内存、存储和网络视为一个集成系统来高效、规模化地交付AI并实现可衡量投资回报率的企业。
采购如今已成为战略,系统设计则是一个领导力问题,麦格雷戈总结道。“每位高管必须提出的最大问题之一是:AI将如何改变我的商业模式?”
本内容由Insights制作,即《麻省理工科技评论》的定制内容部门,而非其编辑团队。内容由人工研究和撰写,任何可能使用的AI工具仅限于在人工监督下的生产流程。
深度探索
人工智能
一个根本性缺陷使LLM极易受到攻击
它使得诱骗模型做出不该做的事情变得轻而易举,例如告诉您如何破坏飞机的导航系统。
Anthropic发现了一个隐藏空间,Claude在其中思考概念
一项新技术使该公司能够比以往更深入地探究LLM的奇特运作机制。
AI在招聘时比人类更容易形成偏见
AI不仅从训练数据中学习刻板印象。它还能制造出新的刻板印象。
以下是AI代理为何会为了实现目标而撒谎和作弊
这种不当行为被称为奖励黑客。这是您需要了解的内容。
保持联系
获取《麻省理工科技评论》的最新资讯
发现特别优惠、热门故事、即将举办的活动等更多内容。
英文来源:
Sponsored
Architecting memory and storage in the AI era
With AI inference now driving enterprise workloads, organizations must rethink infrastructure for speed, efficiency, scalability, and performance per watt to unlock AI’s real-world potential.
In partnership withMicron
The era of AI inference has arrived. Imagine a healthcare system analyzing millions of data points in real time to accelerate life-saving medical research, or an intelligent assistant instantly resolving thousands of complex customer needs at once. These real-world breakthroughs rely on advanced infrastructure acting as the engine of continuous intelligence, powering real-time services while also supporting an increasingly intelligent edge of IoT and consumer devices. However, in this inference-driven landscape, every delay, bottleneck, or wasted watt directly affects human outcomes and operating costs.
This shift changes what infrastructure must deliver. Performance, latency, memory bandwidth, storage throughput, and networking cannot be optimized in silos. Inference workloads are continuous, geographically distributed, and highly sensitive to response time, requiring systems designed for scale, resilience, and efficiency from the start.
“We tend to think of AI as a single workload, and it’s not. It’s thousands, it’s millions, it’s billions of different workloads,” says Jim McGregor, founder and principal analyst, Tirias Research. AI inference changes the optimization problem from one of raw compute to coordinated infrastructure—memory, storage, and networking.
For business leaders, the priority is clear: AI infrastructure decisions must balance cost, flexibility, and future readiness. The winners will be organizations that improve performance per watt, reduce environmental footprint, and remove memory and storage bottlenecks before they limit growth.
AI inference requires a new architectural approach
Systems for AI need to be rearchitected because shoehorning modern AI systems into legacy infrastructure limits AI’s transformative potential. Purpose-built architectures are essential to realize the true value of AI, from accelerating scientific discovery to creating truly autonomous digital agents.
Traditional enterprise IT has been able to rely on relatively stable infrastructure assumptions, but inference and agentic AI introduce new demands around latency, data movement, scalability, and utilization that make architecture choices far more consequential.
“Data centers must now support continuous, distributed, and increasingly real-time AI services—none of which are a single workload,” says McGregor. “They all require different requirements from a system-level perspective.”
To support real-time AI, enterprises can no longer view memory and storage merely as supporting hardware, but at the heart of the system. Organizations need to architect a data pipeline that can rapidly ingest, clean, transform, store, move, and deliver data. Inference workloads place sustained pressure on infrastructure in ways that look very different from earlier training-centric deployments, demanding continuous data retrieval and caching that traditional applications never required.
Accordingly, performance by itself is no longer the sole benchmark that matters. Enterprises increasingly must balance performance with efficiency, cost, and scalability, especially as they try to support different AI services without overbuilding infrastructure for peak conditions.
“You have to optimize the entire network, and that includes memory and storage, around the types of workloads you plan on running,” says McGregor. “You have to really have a detailed understanding of what those workloads are going to be.”
Any AI infrastructure strategy must start with workload awareness. Inference, agentic AI, and other emerging AI use cases require organizations to treat the data center as an integrated system.
Data movement is the new bottleneck and an opportunity for competitive advantage
As enterprises deploy advanced inference and agentic systems, the sheer volume of data being queried in real time has made data movement the most pressing constraint. Modern AI techniques like retrieval-augmented generation (RAG) require systems to constantly scan massive databases to generate accurate responses. This requires immense computing power, but more importantly, it requires immediate access to data.
McGregor says the focus shift to how efficiently data can be moved, cached, and delivered across the broader architecture elevates memory and storage from background infrastructure to strategic assets. “The biggest thing we’re doing right now is moving data from one place to another and making sure that we can use it effectively.”
Because AI is not a single workload category, simply buying the fastest processors is insufficient. Inference depends heavily on memory bandwidth, caching, storage proximity, and the ability to retrieve relevant information quickly and consistently. Understanding where each resource belongs in the stack and how those layers interact under real operating conditions has become a business imperative.
The most effective AI infrastructure looks less like a collection of best-in-class parts and more like a balanced system of compute, memory, storage, and networking, McGregor says, because bottlenecks tend to migrate from one layer to the next. “You have to architect all four together to be efficient, and that’s the challenge.”
The interdependence of data-plane design and network bandwidth means AI infrastructure planning has become a business decision just as much as an engineering one: latency is now inseparable from value. In robotics, financial services, healthcare, and customer-facing AI systems, delays are not merely technical imperfections; they can undermine safety, responsiveness, or trust. AI infrastructure performance becomes a matter of reputation management.
The organizations that gain the most from AI may not be those with the largest clusters, but those with the clearest understanding of how to align every infrastructure element to effectively execute AI workloads.
Building an AI infrastructure procurement framework
Planning AI infrastructure is not simply about choosing the fastest hardware. It is about how to scale without locking the organization into assumptions that may quickly become obsolete. “You need to be flexible because the demands are going to change rapidly and the technology is changing rapidly,” McGregor says.
Future-proofing AI infrastructure requires keeping your options open as workloads, economics, and architectures keep shifting:
- Define the AI workloads that are being optimized. Infrastructure choices must match business needs rather than what McGregor calls generic “AI readiness,” which risks overspending in some areas while leaving bottlenecks unresolved in others.
- Build a modular architecture for compute, memory, storage, power, and cooling so capacity can change as demand shifts rather than committing too early to a rigid architecture.
- Work with the full ecosystem of suppliers and integrators to reduce supply risk and improve access to the right components. McGregor says buyers can no longer assume their OEM or cloud provider alone will insulate them from supply constraints or architectural complexity.
- Reassess your procurement strategy continuously. AI requirements, hardware, and business models are changing too quickly for a fixed long-term design.
- Optimize for efficiency and ROI, not just peak performance. The most powerful setup may be too costly to sustain. Efficiency is also a public-facing metric—better utilization and more workload-aware system design can help companies respond to growing scrutiny around power consumption and water use.
The strategic goal of smarter AI data center design is not maximum performance at any cost, but an adaptable architecture that can deliver value, absorb change, and justify its footprint.
AI infrastructure is now a business strategy
AI data centers have quickly evolved from a back-end technical concern to becoming strategic business systems that help determine how effectively an organization can turn AI into revenue, improve human outcomes, and create a competitive advantage.
In the inference era, memory and storage are no longer passive repositories, explains McGregor, they are the active lifeblood of AI. The organizations that gain the most from AI will not necessarily be those with the largest computing footprint, but those that align infrastructure investments to business outcomes, reduce data bottlenecks, and build the flexibility to adapt as workloads evolve. He predicts that competitive advantage will increasingly belong to enterprises that treat compute, memory, storage, and networking as an integrated system designed to deliver AI efficiently, at scale, and with measurable ROI.
Procurement is now strategy and system design is a leadership issue, McGregor concludes. “One of the biggest questions every executive has to ask is how is AI going to change my business model?”
This content was produced by Insights, MIT Technology Review’s custom content arm, not its editorial staff. It was researched and written by humans, with any AI tools that may have been used limited to production processes under human oversight.
Deep Dive
Artificial intelligence
A fundamental flaw leaves LLMs strikingly vulnerable to attack
It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system.
Anthropic found a hidden space where Claude puzzles over concepts
A new technique has let the company probe deeper than ever into the weird workings of an LLM.
AI is more likely than humans to form biases when hiring
AI doesn’t just learn stereotypes from its training. It can cook up new ones, too.
Here’s why AI agents lie and cheat to reach their goals
The misbehavior is called reward hacking. This is what you need to know.
Stay connected
Get the latest updates from
MIT Technology Review
Discover special offers, top stories, upcoming events, and more.