智算中心论文观察｜2026-06-25

Current Issue

Volume 2026 · Issue 06-25

按期刊卷期页方式整理本期论文。每条仅使用日报已列出的可追溯公开来源，不新增未经核验事实。

Research Article算电协同

Peer-to-Peer Cloud Service Market for Data Centers Oriented to Computation-Electricity Coordination

Yugui Liu、Yibo Ding、Xudong Li、Jing Qu、Wenyi Zhang、Tong Qian、Wuyou Xiao、Zhengyang Hu

Published 2026-06-03 · arXiv · Credibility S

Abstract, interpretation and reference

Abstract

Energy-intensive data centers (DCs) have emerged as substantial and flexible loads in modern power systems, underscoring the critical need for computation-electricity coordination. Harnessing the spatio-temporal flexibility of DC workloads is a promising approach to facilitate this coordination. However, existing studies overlook the collaborative potential of computational resource sharing among geo-distributed DCs, thereby failing to fully unlock this flexibility. In this paper, a bi-level computation-electricity coordination framework is proposed to explicitly capture the bidirectional interactions between DCs and power grid. Firstly, a peer-to-peer cloud service market (P2P-CSM) for geo-distributed DCs is proposed, which enables bilateral cloud service transactions to leverage regional heterogeneities (e.g., electricity prices, cooling efficiency). Secondly, locational marginal prices are embedded into the framework to reflect network congestion and nodal price disparities. Thirdly, a dual consensus alternating direction method of multipliers (ADMM)-based decentralized algorithm is developed as the P2P market clearing algorithm, and a bisection-assisted iterative algorithm is proposed to ensure rigorous convergence of the framework. Case studies conducted on modified IEEE 30-bus system validate that the P2P-CSM achieves a win-win computation-electricity coordination: it not only increases total DC operational profit by 22.8\%, but also effectively alleviates grid congestion and yields a 3.2\% reduction in total energy consumption.

中文解读

背景：AI 数据中心负载、功率密度和能源约束同步上升，算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题：论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法：摘要显示作者采用框架构建和频域/系统级分析，把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果：研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义：对日报读者而言，它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Yugui Liu, Yibo Ding, Xudong Li, 等. Peer-to-Peer Cloud Service Market for Data Centers Oriented to Computation-Electricity Coordination[J/OL]. (2026-06-03)[2026-06-25]. http://arxiv.org/abs/2606.04981v1.

Full text 中文海报

Research ArticleAI 运维优化

Hosting Capacity Assessment and Enhancement for Edge Data Centers in Active Distribution Networks

Linhan Fang、Xingpeng Li

Published 2026-06-01 · arXiv · Credibility S

Abstract, interpretation and reference

Abstract

With the increasing demand for edge computing and AI-driven workloads, integrating small and medium-sized edge data centers into distribution networks has become increasingly important. This paper investigates the hosting capacity of distribution networks for data center integration and identifies the key physical mechanisms that limit the maximum allowable data center load. The baseline analysis shows that data center hosting capacity varies significantly across candidate buses due to network topology and electrical distance. Three dominant limiting mechanisms are identified: current-constrained locations, voltage-constrained locations, and mixed-constrained locations where both current loading and voltage deviation jointly affect hosting capacity. To increase the hosting capacity, this study evaluates multiple flexible resources, including battery energy storage systems (BESS), dispatchable distributed generators (DDG), and static synchronous compensators (STATCOM). Numerical results demonstrate that these resources provide complementary benefits through active power support, sustained local generation, and reactive power compensation, effectively expanding data center hosting capacity in distribution systems.

中文解读

背景：AI 数据中心负载、功率密度和能源约束同步上升，AI 运维、负载预测和设施调优正在成为智算中心设计的关键变量。问题：论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法：摘要显示作者采用仿真建模和情景分析，把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果：研究重点指向跨地域数据中心负载与电力资源之间的调度关系。意义：对日报读者而言，它可用于判断AI 工具是否能降低运维复杂度并提升可用性。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Linhan Fang, Xingpeng Li. Hosting Capacity Assessment and Enhancement for Edge Data Centers in Active Distribution Networks[J/OL]. (2026-06-01)[2026-06-25]. http://arxiv.org/abs/2606.01407v1.

Full text 中文海报

Research Article热管理与液冷

Maximizing Compute Capacity in AI Data Centers through Cooling, Energy Storage, and Computing Adaptation

Shaolei Ren、Mohammad A. Islam、Adam Wierman

Published 2026-05-30 · arXiv · Credibility S

Abstract, interpretation and reference

Abstract

The deployment of artificial intelligence is increasingly constrained by limited site-level power capacity, which must support both compute systems and non-compute systems (primarily cooling) at all times. Cooling power demand, especially in non-evaporative cooling systems, can increase substantially with ambient temperature in the summer, producing recurring periods of elevated cooling power that often lasts for multiple hours per day. Therefore, maximizing compute capacity under a limited site-level power budget is an important planning and operational challenge. Sizing the compute system conservatively based on peak cooling power can leave part of the site-level power capacity underutilized when the cooling power is below its peak, particularly in cooler months. On the other hand, sizing the compute system aggressively based on low cooling power can cause the total site-level power demand to exceed the site-level power capacity during hot days in the summer. This paper proposes ComputeAmp (Compute Amplifier), a framework that maximizes the compute capacity by jointly and dynamically leveraging cooling, battery energy storage, and computing-based adaptation. We discuss the opportunities and limitations of ComputeAmp and illustrate its potential to significantly expand usable compute capacity within local power and water resource limits. We also present a problem formulation for ComputeAmp and highlight a few algorithmic and operational challenges.

中文解读

背景：AI 数据中心负载、功率密度和能源约束同步上升，液冷、热管理和数据中心能效正在成为智算中心设计的关键变量。问题：论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法：摘要显示作者采用框架构建和频域/系统级分析，把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果：研究重点指向冷却效率、能源利用或运维策略的改进方向。意义：对日报读者而言，它可用于判断液冷方案、热管理路线和高密度部署节奏。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Shaolei Ren, Mohammad A. Islam, Adam Wierman. Maximizing Compute Capacity in AI Data Centers through Cooling, Energy Storage, and Computing Adaptation[J/OL]. (2026-05-30)[2026-06-25]. http://arxiv.org/abs/2606.00457v1.

Full text 中文海报

Research Article芯片与算力

Space-CIM: Enabling Compute-In-Memory Accelerators for Thermally-Constrained Space Platforms

Sohan Salahuddin Mugdho、Md. Shahedul Hasan、Cheng Wang

Published 2026-06-04 · arXiv · Credibility S

Abstract, interpretation and reference

Abstract

The rapid growth in compute demand from artificial intelligence (AI) has driven a massive surge in data center construction, precipitating an energy and sustainability crisis. Motivated by the abundant solar energy in outer space and the recent sharp reduction in space launch costs, orbital data centers are emerging as a potential pathway for the future scaling of AI compute infrastructure. While the cold background in vacuum seems appealing for cooling, computing systems operating in space without convection ultimately rely on radiative cooling, requiring large-area radiators. Such limitations in thermal management pose a significant challenge for deploying the standard liquid/air-cooled computers in space. In this work, we investigate the impact of the thermal constraints in space on both graphics processing units (GPUs) with high-bandwidth memory (HBM) and the emerging compute-in-memory (CIM) accelerators. We develop a radiator-in-the-loop co-design methodology that directly links the permitted system TOPS (terra-operations per second) with the practical radiator cooling capacity in space. Our thermal simulations reveal that the separately located GPU die and HBMs create severe thermal hotspots under limited radiator capacity, necessitating GPU thermal throttling. In contrast, CIM accelerators exhibit a much more uniform heat distribution and consistently outperform GPUs in TOPS/W across a wide range of radiator budgets. We systematically evaluated the performance of CIM and GPU across various AI workloads and demonstrated that CIM has a magnified advantage for deployment in space under realistic thermal constraints.

中文解读

背景：AI 数据中心负载、功率密度和能源约束同步上升，芯片、服务器和高密度算力部署正在成为智算中心设计的关键变量。问题：论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法：摘要显示作者采用综述归纳和指标比较，把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果：研究重点指向跨地域数据中心负载与电力资源之间的调度关系。意义：对日报读者而言，它可用于判断芯片路线和服务器密度变化如何传导到机房设计。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Sohan Salahuddin Mugdho, Md. Shahedul Hasan, Cheng Wang. Space-CIM: Enabling Compute-In-Memory Accelerators for Thermally-Constrained Space Platforms[J/OL]. (2026-06-04)[2026-06-25]. http://arxiv.org/abs/2606.05741v1.

Full text 中文海报

Research Article算电协同

Power Grid Infrastructure for AI Data Centers

Amir Sajadi、Muhy Eddin Za'ter、Maria Vabson、Kyri Baker、Bri-Mathias Hodge

Published 2026-05-31 · arXiv · Credibility S

This article addresses recent advances in artificial intelligence, which have set off an astounding race among technology frontiers to build large data centers. It provides insights into impacts of large data centers on the planning and operation of the power grid.

Abstract, interpretation and reference

Abstract

中文解读

背景：AI 数据中心负载、功率密度和能源约束同步上升，算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题：论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法：摘要显示作者采用文献摘要中的模型、实验或案例分析，把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果：研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义：对日报读者而言，它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Amir Sajadi, Muhy Eddin Za'ter, Maria Vabson, 等. Power Grid Infrastructure for AI Data Centers[J/OL]. (2026-05-31)[2026-06-25]. http://arxiv.org/abs/2606.00941v1.

Full text 中文海报

Research Article算电协同

GridPilot: Real-Time Grid-Responsive Control for AI Supercomputers

Denisa-Andreea Constantinescu、David Atienza

Published 2026-05-26 · arXiv · Credibility S

Abstract, interpretation and reference

Abstract

At global scale, data-center electricity demand is growing faster than the grids that supply it, while system operators increasingly require large flexible loads that can adjust power within seconds to absorb variable wind and solar generation. For multi-megawatt AI/HPC facilities, the key unresolved question is practical and measurable: how quickly can the software stack translate a grid request into a real change in GPU power at the facility meter, where commitments are settled? We answer this on real hardware with GridPilot, a three-tier predictive controller operating across milliseconds, seconds, and hours, augmented by a deterministic safety-island bypass for fast response. On a three-GPU NVIDIA V100 testbed, GridPilot achieves a measured end-to-end trigger-to-target response of 97.2 ms, which is 6.9x faster than the 700 ms requirement of Nordic Fast Frequency Reserve. We further incorporate an instantaneous Power Usage Effectiveness (PUE) correction so dispatched commitments remain robust at meter level rather than only at IT load level. In replay experiments across six representative European grids (from Sweden to Poland), the PUE-aware controller closes 2.5-5.8 percentage points of cooling-overhead drag. GridPilot is released as open source and serves as a proof of concept that MW-scale AI/HPC demand can be engineered as controllable, grid-responsive flexibility by design.

中文解读

背景：AI 数据中心负载、功率密度和能源约束同步上升，算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题：论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法：摘要显示作者采用实验验证、原型测试或测量对比，把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果：研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义：对日报读者而言，它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Denisa-Andreea Constantinescu, David Atienza. GridPilot: Real-Time Grid-Responsive Control for AI Supercomputers[J/OL]. (2026-05-26)[2026-06-25]. http://arxiv.org/abs/2605.26384v1.

Full text 中文海报

Research Article算电协同

Grid Capacity Expansion under Data Centers and Electrified Manufacturing Large Loads

Jiyong Lee、Melody Agustin、Joanne Langsdorf、Erhan Kutanoglu、Michael Baldea、Ilias Mitrai

Published 2026-05-28 · arXiv · Credibility S

Abstract, interpretation and reference

Abstract

In this paper, we consider the expansion of power grids under emerging large loads from data centers and electrified manufacturing. We develop a multi-period grid capacity expansion model to determine optimal investment profiles for power generation, storage, and transmission capacity while accounting for hourly power dispatch, such that electricity demand is satisfied and the total planning and operation cost is minimized. We also propose a new modeling approach regarding the spatial distribution of demand from large loads. The model is used to analyze the expansion of a synthetic grid that follows key characteristics of the ERCOT system over a seven-year planning horizon, under loads from data centers and electrified oil refining, which account for 17.5% and 4.7% of total annual electricity demand by the end of the planning horizon. The optimal investment policy leads to an 83.6% increase in generation capacity and exploits the short construction times of solar and storage as well as the operational flexibility of thermal generators. Finally, sensitivity analysis reveals that the construction time of grid assets substantially impacts investment timing, generation technology mix, and transmission capacity expansion. The proposed modeling framework is general and can be extended to other grid systems, enabling the exploration of diverse demand scenarios, policy assumptions, and regional characteristics.

中文解读

参考文献

Jiyong Lee, Melody Agustin, Joanne Langsdorf, 等. Grid Capacity Expansion under Data Centers and Electrified Manufacturing Large Loads[J/OL]. (2026-05-28)[2026-06-25]. http://arxiv.org/abs/2605.29053v2.

Full text 中文海报

Research Article算电协同

Revisiting "Cooler is Better": ITD-Aware Per-CPU Thermal Optimization for Sustainable Data Center Operation

Jason Crop、Hayden Moore、Sudeep Pasricha

Published 2026-06-10 · arXiv · Credibility S

Abstract, interpretation and reference

Abstract

As data center energy demand approaches grid-level constraints, optimizing conventional server infrastructure is essential for sustainable growth. The long-standing assumption that "cooler is better", i.e., lower CPU temperatures reduce power, does not fully hold for modern low-voltage CPUs, where inverse temperature dependence (ITD) drives higher supply voltages at lower temperatures. This creates a non-monotonic performance-per-watt curve where efficiency peaks at an intermediate thermal point. In this paper, for the first time, we empirically characterize ITD on production Intel Xeon CPUs and demonstrate that efficiency-optimal temperatures are CPU part-specific, and frequently higher than typical data center operating conditions. Measurements from commercial cloud data center platforms (Amazon, Equinix) reveal that approximately half of modern high-power CPUs operate about 10°C below their efficiency-optimal thermal point. By implementing ITD-aware thermal grouping of CPUs and inlet temperature adjustments, data center operators can optimize facility-level cooling and overall sustainability. Our case study shows that this approach can reduce total data center energy by 4-13% without sacrificing performance or reliability.

中文解读

背景：AI 数据中心负载、功率密度和能源约束同步上升，算力负载与电网侧资源的协同调度正在成为智算中心设计的关键变量。问题：论文聚焦现有方案在效率、可靠性或工程协同上的瓶颈。方法：摘要显示作者采用建模优化、调度分析或算法评估，把运行负载、冷却/能源系统和基础设施约束放在同一分析框架中。结果：研究重点指向AI 负载波动对电网设备寿命和调频边界的影响。意义：对日报读者而言，它可用于判断智算中心建设是否受电网容量、负载波动和调度机制约束。仍需结合全文实验条件、样本范围和成本假设核验。

参考文献

Jason Crop, Hayden Moore, Sudeep Pasricha. Revisiting "Cooler is Better": ITD-Aware Per-CPU Thermal Optimization for Sustainable Data Center Operation[J/OL]. (2026-06-10)[2026-06-25]. http://arxiv.org/abs/2606.11163v1.

Full text 中文海报

智算中心论文专站

Abstract

中文解读

参考文献

Abstract

中文解读

参考文献

Abstract

中文解读

参考文献

Abstract

中文解读

参考文献

Abstract

中文解读

参考文献

Abstract

中文解读

参考文献

Abstract

中文解读

参考文献

Abstract

中文解读

参考文献