
面向下一代AI基础设施的800V直流架构
800 VDC Architecture for Next-Generation AI Infrastructure
Jared Huntington & Mike Tu
800 VDC架构与集成储能之势在必行
The Architectural Imperative of 800 VDC and Integrated Energy Storage
就在几年前,数据中心的核心设计仍以计算空间为中心——庞大的数据大厅布满服务器,而供电和冷却系统仅占据较小空间。然而,GPU的革命性发展将数据中心转变为“AI工厂”。如图1所示,GPU机架的功率密度已达到网络服务器的近100倍,且正以近乎指数级的速度增长,彻底改变了这种平衡。曾经处于次要地位的电力基础设施,如今其所需空间已能与计算空间相匹敌,甚至更胜一筹。
Only a few years ago, data centers were built around the compute space, vast data halls of servers, with power and cooling systems taking up a smaller share of the footprint. Then came the GPU revolution, transforming the data center into “AI Factories”. As shown in Figure 1, GPU racks have reached nearly 100 times the power density of web servers and their power density is growing at a near-exponential pace, completely reversing the balance. Power infrastructure, once secondary, now rivals or even exceeds the space dedicated to compute.

图1:跨代增长的发电量
Figure 1: Increasing Power Generation over Generation
随着CPU与GPU的迭代升级,GPU的热设计功耗通常会出现代际递增约20%的阶梯式增长。这导致单台服务器所需的功耗随时间推移持续攀升。英伟达的NVLink技术允许多个GPU通过网络互联,协同运作如同一颗大型同步GPU,相比基于以太网的连接方式可显著提升性能。从功耗与成本角度考量,通过铜缆实现GPU互联能获得最佳效益,但其代价是因信号完整性导致的传输距离受限。由于在有限的铜互连域内集成更多GPU可实现极致性能,最大性能实际上与最大功率密度直接挂钩。这意味着功耗增长不再局限于每代20%的幅度,随着NVLink互联域规模的扩大,功耗水平可轻松实现2倍、4倍甚至8倍的跃升。
With each generation, CPU and GPU improvements typically bring an incremental ~20% increase in GPU Thermal Design Power. This leads to an increase in power required per server over time. Nvidia’s NVLink allows multiple GPUs to be networked together to effectively act as one large synchronous GPU and delivering significant performance gains when compared to operating over Ethernet connections. This networking of GPUs is most effective when implemented over copper cabling from both a power and cost standpoint, but at the expense of a limited reach due to signal integrity. Since the highest performance can be achieved when more GPUs are on the same copper domain with a limited reach, the maximum performance is tied to the maximum power density. This means power increases are no longer 20% generation over generation, but can easily be 2x, 4x or 8x with the increase in NVLink networking domain size.

图2:TDP功耗增加75%,但性能提升50倍
Figure 2: A 75% increase in TDP power, but a 50x increase in performance
图2示例展示了从Hopper架构到GB300架构的性能跃升。虽然热设计功耗仅增加75%,但性能却实现了50倍的提升。这一变革同时使得机柜功率密度增长3.4倍——NVLink互联域从4x8 GPU配置(机柜内共32个GPU)扩展为72个GPU的互联单元。随着GPU集成技术与封装工艺的持续进步,以及网络拓扑向更大规模互联域演进,功率密度的提升态势仍将持续。
An example in Figure 2 above is the performance increase from Hopper to GB300. The TDPincreased by 75%, but the performance increased by 50x with these changes. This also led to a 3.4x increase in rack power density going from an 4x8 GPU NVLink domains (32 in the rack) to a 72 GPU NVLink domain. As GPU packing and packaging improves and networking topologies move to larger domain sizes, this power density can continue to increase.
GPU每代性能提升与NVLink互联域扩增共同推动了功率需求的飙升,其增速远超前代GPU的发展轨迹。另一个关键目标在于尽可能将电源组件移出NVLink域的辐射范围——因为该区域是机柜中支撑算力性能的核心地带。功率等级的持续攀升与电源组件外移这两大趋势相互叠加,正催生对新型机柜电源架构的迫切需求。
This increase in power generation over generation and with NVLink domain drives a much more rapid increase in power than has been seen in the past with GPUs. A secondary goal is to remove as many power components from the NVLink domain radius as possible since this is the most critical area in the rack for performance. This combination of driving to higher power level and pushing the power components away from the GPUs drives requirements for a different rack power architecture.
为满足这些前所未有的需求,800 VDC已成为下一代配电系统的最优架构。该架构能最大程度减少计算空间的电能转换环节与线路路由体积,同时有效降低数据中心配电损耗及端到端整体转换级数。与机架内采用的54 VDC或设施级480 VAC系统相比,800 VDC在确保安全性与可扩展性的同时,显著降低了电流强度、铜材用量及线缆体积。这一架构得益于碳化硅与氮化镓功率转换器件的日益成熟,以及800 VDC系统在电动汽车行业的广泛应用,从而实现了从电网到机柜的无缝端到端整合,使功率密度突破1兆瓦成为可能。返回搜狐,查看更多