为AI提供动力是一个架构问题

内容来源:https://www.technologyreview.com/2026/09/10/1141649/powering-ai-is-an-architecture-problem/
内容总结:
一场发生在弗吉尼亚州阿什伯恩的电网故障,将AI数据中心集群的电力架构问题推到了台前。2026年7月22日,一条输电线路故障在几秒内从电网中切除了超过3吉瓦的负荷。这并非首次——两年前,一个失效的避雷器曾同时导致约60处弗吉尼亚设施、共计1500兆瓦负荷脱网。没有人预料到如此大规模的均质负荷会以相同方式、在同一时刻对电网故障做出响应。
关于AI电力的讨论大多集中在发电侧:更多涡轮机、更多太阳能、更多输电线路。电网确实需要更多电子。但弗吉尼亚的停电事件并非供电失败,而是架构失败。而一大波互联接入请求正涌向同一套架构,将电网可靠性置于风险之中。这是一个没人愿意接手的问题。
电网承受之重
电网的设计围绕可预测负荷展开:钢厂、炼油厂、晚餐时分的住宅。负荷规模不同,但过程相同——平稳取电,偶尔异常,优雅恢复。但AI数据中心不是这样运行的。一个AI园区在训练运行期间可以在毫秒内摆动70%的负荷,随后在上游出现故障迹象时同样迅速地脱网,以保护数十亿美元的算力。每个数据中心单独来看都是理性的。但在吉瓦级别上合在一起,它们构成了电网从未解决过的问题——而下一波数据中心园区正是按照这一规模规划的。
旧架构在哪里断裂
标准的数据中心电力架构几十年来没有变化:中压电力送入,变压器降压,低压不间断电源(UPS)调节,电力到达机架。将这一设计推到AI规模,它会在三个地方开裂。
第一,UPS深居建筑内部,紧邻机架。但其电池只是一个尺寸不足的备胎,设计用于应对几分钟的停电,而非全天候吸收如此快速且剧烈的负荷波动。
第二,UPS大部分时间处于旁路状态。传统变换器浪费的电力足以让运营商运行在eco模式:静态开关直接从电网向机架供电,两个方向都没有任何滤波。算力的波动原样输出,而电网瞬态——可能损坏或击穿设备的亚毫秒级事件——来得太快,任何开关都来不及捕捉。
第三,保护逻辑的编写年代,“大负荷”意味着50兆瓦。这套保护逻辑看不见它现在所连接的电网,因此当上游出现故障时,它做了完全错误的事:脱网。在2024年弗吉尼亚事件中,大部分损失负荷可追溯到保护方案——计算电压跌落次数,在第三次时断开——正如设计所要求的那样,在最糟糕的时刻。
这不是粗劣的工程。这是精心的工程,只是负荷已经超越了它。
移入电力路径
解决方案是三个同步动作:向上移——从480伏升至中压(13.8千伏及以上),即大型站点从电网取电的电压等级。向外移——从数据大厅移至变电站附近的模块化外壳,使建筑内仅保留算力和维持其运行的冷却设备。移入路径——不再是旁观反应的电池,而是每个电子始终流经的系统。无需检测,无需切换,因为没有任何东西被绕过去。
纸面上,三个直截了当的升级。实践中,它们改写了下游每一个项目。
实施变革
当数千块GPU同时启动时,系统吸收波动,向电网输出平稳的负荷曲线。当扰动来袭时,其后的设备毫无察觉。一个难缠的邻居变成了可预测的邻居。当公用事业公司需要帮助时,它变成了有用的邻居。
互联接入也随之改变。公用事业公司只需认证一个中压箱体,而非梳理其后的每一台变压器、UPS、冷水机组、水泵和开关柜排列。工程师可以在不重新进行互联研究的情况下更换芯片代际。许可时间线缩短数月。
在围栏内部,UPS房间变成算力或冷却空间。每建设一美元的密度攀升。经济性也随之翻转。在中压下运行、置于室外、自储能量的设备可以获得税收抵免,并在削峰和需求响应等电网项目中赚取收入。备用电源不再是保险,而是开始自己养活自己。
架构测试
2026年初,我们在国家落基山实验室——美国能源部设施,也是西半球唯一能在同一回路中同时复制真实电网故障和AI规模负荷波动的地方——测试了一套全尺寸系统。我们从两个方向同时冲击:真实AI负荷曲线以全中压冲击算力侧。电网故障冲击公用事业侧,包括一次完全零电压事件。算力侧纹丝不动。电网侧也没有。它通过了得克萨斯电力可靠性委员会(ERCOT)的大负荷电压穿越要求,且余量充足。
这些规则的存在是因为运营商不再凭空信任这种规模的设施,而更多设施正在到来。大多数行业将这些规则视为障碍。一套中压、在线式系统可以开箱即用地通过。合规不是附加功能。它就是这套架构本身的属性。
新的一层
在AI建设中,许多看似电网问题的东西其实就在围栏内部,在于那些为已不复存在的负荷而 sizing 的设备。把正确的部件向上移、向外移、移入路径,电网的负担就变成电网的资产。密度上升。许可时间下降。备用电源自食其力。工程可行——下一波AI工厂正建立在其之上。行业还没有给这一层命名。我们称之为中压AI UPS。名字不重要,重要的是选择:这些工厂可以作为电网的负担到来,也可以作为电网的力量到来。我们已经知道如何建造第二种。
中文翻译:
赞助内容
为AI供能是一个架构问题
将电力保护上移至更高电压等级、移至建筑外部、并纳入供电通路,不只是解决断电问题;它改变了密度、审批时间线和备用电源的经济性。
由ON.energy提供
2026年7月22日,弗吉尼亚州阿什本——全球最大数据中心集群的核心地带——一条输电线路故障在数秒内从电网切除了超过3吉瓦的负荷。而这并非首次。两年前,一个浪涌保护器失效,一次性导致弗吉尼亚约60处设施、1500兆瓦负荷脱网。没有人能预料到如此庞大的均质负荷会以同样的方式、在同一时刻对电网故障做出响应。
关于AI电力的争论大多集中在发电环节:更多涡轮机、更多太阳能、更多输电线路。电网需要更多电子。但弗吉尼亚的断电并非供给失败,而是架构失败。而一大波互联接入请求正涌向同一套架构,将电网可靠性置于风险之中。这是一个没人愿意接手的问题。
向电网索取更多
电网是围绕可预测的负荷建造的:钢厂、炼油厂,以及晚餐时分的千家万户。负荷规模各异,但过程相同——平稳取电,偶尔异常,优雅恢复。
但AI数据中心的行为方式并非如此。
一个AI园区在一次训练运行中可以在毫秒内摆动70%的负荷,然后在发现上游出现麻烦的第一时间以同样快的速度脱网,以保护数十亿美元的计算设备。单独来看,每一个都是理性的。但合在一起,在吉瓦级别上,它们构成了电网从未解决过的问题——而下一波数据中心园区正是按照这一规模规划的。
旧架构在哪里断裂
标准的数据中心供电架构几十年来没有改变。中压电力送入,变压器降压,低压不间断电源(UPS)调节电能,然后送达机架。将这一设计推向AI规模,它会在三个地方开裂。
第一,UPS深置于建筑内部,紧邻机架。但它的电池是一个尺寸不足的备胎,设计用于应对几分钟的断电,而非全天候吸收如此快速且剧烈波动的负荷摆动。
第二,UPS一生中大部分时间处于旁路状态。传统变换器浪费的电力足以让运营商运行在生态模式:静态开关直接从电网向机架供电,两个方向都没有任何滤波。计算的波动原样输出,而电网瞬变——可能损坏或击倒设备的亚毫秒级事件——来得太快,任何开关都来不及捕捉。
第三,保护逻辑是在“大负荷”意味着50兆瓦的时代编写的。这一保护逻辑看不到它现在所连接的电网,因此当上游出现问题时有,它做的恰恰是错误的事:脱扣退出。在2024年弗吉尼亚事件中,大部分损失负荷可追溯到保护方案——它们计算电压跌落次数,在第三次时断开——正如设计所愿,在最糟糕的时刻。
这不是粗劣的工程。这是精心的工程,只是负荷已经超出了它的设计能力。
移入通路
解决方案是三个动作,同步实施。
上移——从480伏提升到中压(13.8千伏及以上),即大型场地从电网取电的电压等级。
外移——从数据大厅移至变电站附近的模块化机柜,使建筑内只保留计算设备和维持其运转的冷却系统。
移入通路——不再是旁观并做出反应的电池,而是一个每个电子始终流经的系统。没有什么需要检测,也没有什么需要切换,因为从来没有东西被绕过去。
纸面上,三个直截了当的升级。实践中,它们改写了每一个下游条目。
实现变革
当数千块GPU同时启动时,系统吸收波动,向电网输出平坦的负荷曲线。当扰动来袭时,其背后的设备毫无察觉。一个难缠的邻居变成了可预测的邻居。而当公用事业需要帮助时,它变成了有用的一个。
互联接入也随之改变。公用事业只需认证一个中压箱体,而不必理清其背后每一台变压器、UPS、冷水机组、水泵和开关柜的组合。工程师无需重新进行互联研究即可更换芯片代际。审批时间线缩短数月。
在围栏内部,UPS机房变成计算或冷却空间。每单位建设资金的密度上升。
经济性也随之翻转。运行在中压、置于室外、自带储能的设备可以获得税收抵免,并在削峰和需求响应等电网项目中赚取收入。备用电源不再是保险,而是开始自我回本。
架构测试
2026年初,我们在落基山国家实验室——美国能源部设施,也是西半球唯一能在同一回路中同时复现真实电网故障和AI级负荷摆动的地方——测试了一套全尺寸系统。
我们从两个方向同时施压:真实AI负荷曲线以全中压冲击计算侧。电网故障冲击公用事业侧,包括一次完整的零电压事件。计算侧纹丝不动。电网侧也是。它以充裕的余量通过了得克萨斯电力可靠性委员会(ERCOT)——电网运营商——的大负荷电压穿越要求。
这些规则之所以存在,是因为运营商不再凭信任接纳这种规模的设施,而更多设施正在到来。业内大多数企业将其视为障碍。一套中压在线系统开箱即通过。合规不是附加功能。它就是这套架构本身的属性。
新的一层
在AI建设中,许多看似电网问题的东西其实就在围栏内部,在于那些按照已不复存在的负荷来 sizing 的设备。将正确的部件上移、外移、移入通路,电网的负担就变成了电网的资产。密度上升。审批时间下降。备用电源物有所值。
工程可行——下一波AI工厂正建立在其之上。业界还没有给这一层命名。我们称之为中压AI UPS。名字不如选择重要:那些工厂可以作为电网的负担到来,也可以作为电网的力量到来。我们已经知道如何建造第二种。
本内容由ON.energy制作。非MIT Technology Review编辑团队撰写。
深度洞察
人工智能
一个根本性缺陷使大语言模型极易受到攻击
它让人们很容易诱骗它们做不该做的事,比如告诉你如何破坏飞机的导航系统。
AI在招聘时比人类更容易形成偏见
AI不仅从训练中学习刻板印象。它还能编造新的。
以下是AI智能体为何为实现目标而撒谎和作弊
这种行为被称为奖励黑客。以下是你需要了解的。
比尔·盖茨说我们已经越过了AI的危险阈值。接下来呢?
在一次新采访中,这位亿万富翁慈善家对我们整顿AI政策的紧迫性发出警报。
保持联系
获取MIT Technology Review的最新动态
发现特别优惠、热门故事、即将举办的活动等。
英文来源:
Sponsored
Powering AI is an architecture problem
Moving power protection up the voltage stack, outside the building, and into the power path doesn't just solve outages; it changes density, permitting timelines and backup power economics.
Provided byON.energy
On July 22, 2026, a transmission line fault in Ashburn, Virginia—the heart of the world's largest data center cluster—knocked more than 3 gigawatts of load off the grid in seconds. And it wasn't the first time. Two years earlier, a single failed surge arrester dropped roughly 60 Virginia facilities and 1,500 megawatts at once. No one could anticipate so much uniform load responding to grid faults the same way, at the same time.
The AI power debate is mostly about generation: more turbines, more solar, more transmission. The grid needs more electrons. But the outages in Virginia weren't supply failures; they were architecture failures. And a giant wave of interconnections is arriving on that same architecture, putting grid reliability at risk. It's a problem nobody wants to own.
Asking more from the grid
The grid was built around predictable loads: steel mills, refineries, and houses at dinnertime. Different load sizes, same process—drawing power smoothly, misbehaving occasionally, and recovering gracefully.
But AI data centers don't behave that way.
An AI campus can swing 70% of its load in milliseconds during a training run, then trip offline just as fast at the first sign of trouble upstream to protect billions in compute. Each is rational alone. Together, at gigawatt scale, they're a problem the grid has never solved—and the next wave of data center campuses is planned at exactly that scale.
Where the old stack breaks
The standard data center power stack hasn't changed in decades. Medium-voltage power arrives, transformers step it down, low-voltage uninterruptible power supply (UPS) units condition it, and it reaches the racks. Push that design to AI scale, and it cracks in three places.
First, the UPS sits deep inside the building, close to the racks. But its batteries are an undersized spare tire, designed to handle an outage for a few minutes, not to absorb load swings this fast and volatile around the clock.
Second, the UPS spends most of its life in bypass. Legacy converters waste enough power that operators run in eco-mode: A static switch feeds the racks directly from the grid and nothing filters in either direction. The compute's swings go out raw, and grid transients—sub-millisecond events that can damage or take down equipment—come in too fast for any switch to catch.
Third, the protection logic was written when "large load" meant 50 megawatts. This protection logic can't see the grid it is now a part of, so when trouble hits upstream, it does exactly the wrong thing: it drops out. In the 2024 Virginia event, most of the lost load traced to protection schemes that count voltage dips and disconnect on the third one—as designed, at the worst moment.
This isn't sloppy engineering. It's careful engineering the load has outgrown.
Moving into the path
The fix is three moves, made together.
Move it up—from 480 volts to medium voltage (13.8 kilovolts and higher), the voltage large sites draw from the grid.
Move it out—from the data hall to modular enclosures near the substation so the building holds only compute and the cooling that keeps it alive.
Move it into the path—instead of a battery that watches and reacts, a system every electron runs through, all the time. There's nothing to detect and nothing to switch because nothing was ever routed around it.
On paper, three straightforward upgrades. In practice, they rewrite every line item downstream.
Making the change
When thousands of GPUs spin up together, the system absorbs the swing and hands the grid a flat load profile. When a disturbance hits, the equipment behind it never notices. A difficult neighbor becomes a predictable one. And when the utility needs help, it becomes a useful one.
Interconnection changes, too. The utility certifies one medium-voltage box instead of untangling every transformer, UPS, chiller, pump, and switchgear lineup behind it. Engineers swap chip generations without a fresh interconnection study. Months come off the permitting timeline.
Inside the fence, UPS rooms become compute or cooling space. Density per construction dollar climbs.
And the economics flip. Equipment that runs at medium voltage, sits outside, and stores its own energy can qualify for tax credits, and earn revenue in grid programs like peak shaving and demand response. Backup power stops being insurance and starts paying for itself.
The architecture test
In early 2026, we tested a full-scale system at the National Laboratory of the Rockies, a U.S. Department of Energy facility and the only place in the Western Hemisphere that can replicate real grid faults and AI-scale load swings concurrently in the same loop.
We hit it from both directions: real AI load profiles hit the compute side at full medium voltage. Grid faults hit the utility side, including a full zero-voltage event. The compute side didn't flinch. Neither did the grid side. It cleared the large-load voltage ride-through requirements from the Electric Reliability Council of Texas (ERCOT), the grid operator, with room to spare.
Those rules exist because operators no longer take facilities this size on faith, and more are coming. Most of the industry treats them as hurdles. A medium-voltage, inline system clears them out of the box. Compliance isn't an added feature. It's what the architecture does.
The new layer
Much of what looks like a grid problem in the AI buildout sits inside the fence, in equipment sized for a load that no longer exists. Move the right pieces up, out, and into the path, and a grid liability becomes a grid asset. Density goes up. Permitting time comes down. Backup power earns its keep.
The engineering works—and the next wave of AI factories is being built on it. The industry hasn't named this layer yet. We call it the medium-voltage AI UPS. The name matters less than the choice: those factories can arrive as a strain on the grid or as strength for it. We already know how to build the second kind.
This content was produced by ON.energy. It was not written by MIT Technology Review’s editorial staff.
Deep Dive
Artificial intelligence
A fundamental flaw leaves LLMs strikingly vulnerable to attack
It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system.
AI is more likely than humans to form biases when hiring
AI doesn’t just learn stereotypes from its training. It can cook up new ones, too.
Here’s why AI agents lie and cheat to reach their goals
The misbehavior is called reward hacking. This is what you need to know.
Bill Gates says we’ve passed AI’s danger thresholds. Now what?
In a new interview, the billionaire philanthropist sounds an alarm on the urgency of getting our AI policies in order.
Stay connected
Get the latest updates from
MIT Technology Review
Discover special offers, top stories, upcoming events, and more.