Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Data center thermal management”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Improving Thermal Management Strategies for Data Centers: A Physical Testbed Incorporating Small Modular Reactor and Microreactor Technology

This study aims to accelerate the demonstration of various thermal management systems for data centers using nuclear-generated heat to enhance energy and grid reliability. Utilizing mobile containerized and stationary test beds at INL's High Performance Computing (HPC) facility, this project integrates with various nuclear-related energy systems testing facilities. Key components include immersion cooling apparatus, absorption chillers, and adjustable thermal management simulators. Tasks involve acquiring necessary hardware, sensors, and cooling apparatus, engaging with data center industry stakeholders, and providing a testing platform for algorithms, models, tools, and software. The objective is to expedite the deployment of nuclear-powered data centers, thereby improving energy reliability and affordability.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN

Holistic energy analysis method for thermal management architectures of data centers

Modern high-performance computing (HPC) data centers (DCs), particularly those supporting energy-intensive artificial intelligence (AI) workloads, face escalating thermal management challenges that degrade performance through thermal throttling and drive up cooling power consumption and operational costs. To address this challenge, many have developed a wide variety of thermal management solutions (single-phase, two-phase, direct, indirect, hybrid, and more) which attempt to cool HPC DCs effectively while attempting to minimize overall system power consumption. However, the analysis of these solutions and methods to effectively compare one with another is lacking. Overall power usage effectiveness (PUE) and total-power usage effectiveness (TUE) provide a metric to quantify power consumption but fail to identify components in the system which require further optimization. To address this, we propose a holistic analytical framework – the waterfall diagram (WFD) – which leverages a waterfall chart methodology, offering a comprehensive visualization of both the thermal management system loop and heat flow pathways from individual server components to the outdoor ambient. Use of the WFD enables graphical estimations of power efficiency and cooling performance across each component of a DC cooling system and complements Sankey-style energy flow visualizations by additionally resolving stage-wise temperature changes and incremental TUE contributions. The framework is used in conjunction with simulation-based approaches, to conduct a detailed pressure drop and flow distribution analysis aimed at identifying the optimal coolant distribution architecture for a single-phase direct-to-chip water-cooled DC, which serves as the baseline for subsequent WFD analysis. Among the evaluated architectures, the 3 U modular coolant distribution architecture is found to demonstrate the best performance, considering minimal pressure drop and uniform flow distribution. In addition, TUE is calculated for each cooling loop component based on its associated pressure drop and corresponding pumping power, which are integrated into the WFD. This correlation between TUE and local temperature offers immediate insight into the power efficiency and thermal performance contributions of individual components, facilitating further development and optimization. Examples of WFD applications are presented under varying thermal loads and ambient conditions, demonstrating reasonable cooling strategies. Notably, the 3 U modular architecture maintains a consistent chip case temperature of 85°C, achieving a TUE of 1.016 at ambient temperature of 47°C, and a TUE of 1.026 at ambient temperature of 52°C. The WFD methodology provides an efficient, holistic, and streamlined framework for DC thermal management architecture assessment and enables design optimization which is important for addressing the thermal-fluidic energy challenges of current and next-generation DCs.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Accelerating nuclear-integrated data center pursuits in the USA: SWOT analysis, power-thermal management strategies and demonstration plan

Here, this study explores the increasing interest in leveraging nuclear power to meet the escalating energy demands of data centers in the United States (U.S.) by focusing on key factors that contribute to accelerated deployment. The study highlights the importance of N+1/N+2 power supplies (where N is the required number of units), outlines research and innovations in nuclear-integrated data center thermal management and demonstration plan. It also provides updates about status and costing of various reactor system designs. A summarized strengths, weaknesses, opportunities, and threats (SWOT) analysis shows the potential options for grid connectivity, reactors, and site selection. Suitable site discussions consider land and water availability, grid access, and optical fiber connectivity, and the study presents graded prospects for Department of Energy (DOE) sites with a specific example. Community engagement and partnerships are emphasized, particularly the roles of local government, federal agencies, utilities, and data center industry partners, which are crucial for accelerating deployment, business outreach, and approvals. The study provides actionable insights for stakeholders to accelerate the deployment of nuclear-powered data centers.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

Naug, Avisek [Hewlett Packard Enterprise]

A Novel and Scalable Method for Microencapsulating Salt Hydrate Phase Change Materials in Core–Shell Fibers

Phase change materials (PCMs) are in high demand for applications such as thermal energy storage in buildings, electronics cooling, and thermal management of electric vehicle batteries and data centers. Among these materials, salt hydrate PCMs are particularly attractive due to their high thermal energy storage capacity and low cost. However, they suffer from two major issues: leakage in the melted phase and phase segregation during phase transitions. Microencapsulation is the primary process capable of addressing both of these challenges. However, there is no reliable or scalable method available for microencapsulating salt hydrate PCMs. As a result, the full potential of salt hydrates for building and data center applications has yet to be realized. In this work, we present an innovative method for the microencapsulation of salt hydrate PCMs using a co‐axial pushing technique. This process creates core–shell fibers, with the salt hydrate as the core and a polymer as the shell. Our approach demonstrates strong potential for scalable microencapsulation of salt hydrate PCMs. In conclusion, achieving scalability could enable their widespread use in applications such as data center cooling, battery thermal management, and building climate control.

Sharma, Jaswinder [Oak Ridge National Laboratory (

Revolutionizing thermal Management in Next-Generation AI data centers: Challenges and breakthrough innovations

Data centers (DCs) serve as critical infrastructure for powering the growth and evolution of AI. Next-generation AI DCs present unique challenges in thermal management driven by unprecedented computational demands. This paper provides a comprehensive summary of key stakeholder perspectives on technology gaps, infrastructure requirements, test bed needs, emerging opportunities, and preliminary solutions related to thermal management for AI DCs. It establishes six strategic pillars of thermal management for next generation AI DC: reliability, deployability, efficiency, resilience, measurability, and valorization. The discussion spans a range of critical topics, including advanced cooling technologies, thermal strategies for emerging modular and edge DCs, system-level optimization and control frameworks, infrastructure planning and grid integration designs, benchmarking approaches, and pathways for waste heat recovery and reuse. The proposed research, development, and demonstration efforts are aimed at accelerating the deployment of AI DCs while ensuring energy efficiency, reliability, safety, and regulatory compliance.

Wang, Pengtao [ORNL] (ORCID:0000000214713429)

Ultra-low thermal resistance and pressure drop copper and copper-tungsten diamond-shaped pin fin cold plates for liquid cooling of electronics

Modern and future data centers face increasing cooling challenges due to increasing chip thermal design power and die size, along with the need to reduce energy consumption used for cooling. High performance cooling solutions that maintain a low chip junction temperature are needed to ensure electronics reliability. This work develops an ultra-low thermal resistance and low pressure drop 75 mm × 75 mm cold plate, intended for next-generation electronics cooling. The cold plate features an array of diamond-shaped pin fins and integrated copper tungsten heat spreader, selected for its low coefficient of thermal expansion which reduces thermomechanical deformation and allows for closer integration of the cold plate with silicon dies. Starting with 300 candidate designs, three-dimensional computational fluid dynamics simulations predict the thermal-hydraulic performance of cold plate subsections. The highest performing geometries are evaluated with high fidelity simulations. Four cold plates are manufactured for experiments: three with diamond-shaped pin fins and one with straights fins for comparison purposes. The cold plates are fabricated from copper-tungsten (CuW), copper (Cu), or aluminum-silicon-magnesium alloy (AlSi10Mg). The diamond-shaped pin fins achieve a roughly 15 % lower thermal resistance compared to the conventional straight fin microchannel. The highest performing design achieves a chip-to-coolant (including thermal interface material) thermal resistance of 9.0 K/kW in CuW and 6.9 K/kW in Cu under a 1 kW heat load with an inlet-to-outlet pressure drop of 9.0 kPa and water as the working fluid. This work demonstrates ultra-low thermal resistance and pressure drop cold plates for large die, high heat load applications, and shows that CuW is an attractive cold plate material for improved reliability in next generation data center cooling.

Coefficient of thermal expansion

Accelerating Nuclear-Integrated Data Centers in the USA: SWOT Analysis, Power-Thermal Management Strategies, and Industrial-Scale Demonstration and Potential Deployment

Driven by the growth in digital services, cloud computing, AI, and manufacturing, data centers face rising energy demands that challenge traditional power sources and cooling efficiency. This study explores using nuclear power to meet these demands, focusing on accelerated reactor technology deployment and highlighting needs such as N+1/N+2 power supplies and integrated power-thermal management. A SWOT analysis addresses grid connectivity, reactors, and site selection, particularly DOE sites. Reactor technology demonstration and deployment could be accelerated by leveraging test facilities such as MARVEL, MAGNET, TED, FAS, DOME, LOTUS, ATR, Energy System Proving Grounds, and upcoming Energy Launch Pads, along with modeling and simulation tools such as RELAP5, MOOSE, VERA, RAVEN, and FORCE. The potential power and thermal management options, including various cooling technologies, waste-heat utilization, and an industrial-scale demonstration plan, aim to accelerate the integration of nuclear power and data centers in the USA, while emphasizing community and stakeholder engagement and synergistic efforts.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN

Potential of Data Center Controls in Grid Services

The rapid proliferation of large data centers brings both challenges and opportunities for grid reliability. The data center resources and their potential flexibility have the potential to contribute resources to grid operations. Through capabilities like energy shifting and resource coordination, data centers can help reduce their net demand on the transmission network, as well as provide additional grid services to support reliable operation on the grid. While transient and long-term grid planning and operations are the scenarios that draw most attention, the quasi-steady state timeseries (QSTS) operation of data centers and grid bring interesting scenarios that can help evaluate the data center controls to aid grid services. This work is focused on modeling data centers for QSTS applications – incorporating the AI data center load profiles and building on the PNNL digital twin model for the thermal management loads to enable simulation studies to reveal the impact of data center controls on grid performance. This includes the integration of a QSTS battery and natural gas generator model to incorporate local resource impacts to the system. The simulation study is performed with a modified IEEE 24-Bus transmission system. Scenarios are focused on evaluating the data center load impacts on the transmission system and leveraging both data center and local generation controls to mitigate those impacts and provide additional grid services. The data center controls revealed the ability to contribute to two main kinds of grid services: preventing congestion on a weak grid by coordinating the data center resources with the collocated BESS and onsite generation; and the ability to help the grid operations during stressed times of operation like during a contingency. Leveraging these and other capabilities has the potential to help data centers become grid responsive assets, aiding in both their integration into the power system and grid reliability.

power grid simulation

Metal additively manufactured wavy fin cold-plate architecture for improved thermal-hydraulic performance

Rapid growth in artificial intelligence and data center workloads demands high-performance liquid cooling to manage increasing chip power. This study presents two metal-additive-manufactured cold plates with sinusoidal fins, constant-amplitude wavy fins and linearly variable-amplitude wavy fins and compares them against metal-additive-manufactured straight fins using experiments conducted at 1 kW heat dissipation as well as high-fidelity 3D conjugate computational fluid dynamic simulations. The cold plates were printed in AlSi10Mg material and underwent design using a Python-automated workflow prior to manufacture and testing. The experiments show that wavy fins reduce the normalized thermal resistance by 35 to 45 % at water flow rates from 1 to 4 LPM. At a fixed 20 kPa pressure drop, the variable-waviness design lowered peak surface temperature by 9 °C and thermal resistance by 51 %, while edge-channel maldistribution in the constant wavy fin design limited gains. A thermal resistance breakdown revealed that 55–63 % of the total thermal resistance in wavy designs comes from base heat conduction, 27–33 % from fin heat conduction, and 9–13 % from fin heat convection, indicating the need to address conduction bottlenecks. Parametric sweeps identify a 3 mm fin pitch as optimal, and that horizontal inlet/outlet manifolds further reduce pressure drop by 30–60 % and thermal resistance by 9–16 % relative to vertical inlet-outlet manifolds. The results yield comprehensive guidelines for fin geometry, manifold alignment, material selection and additive-manufacturing constraints to realize high-performance liquid-cooled cold plates for power-dense electronics.

3d printing

A Physics-Informed Reinforcement Learning Framework for Economic-Thermal Co-Optimization of Crypto Mining Data Centers: Preprint

The rapid expansion of cryptocurrency mining has created a new class of high-density data centers characterized by extreme thermal flux and high sensitivity to volatile economic markets. Traditional thermal management strategies, typically reliant on rule-based control, maintain static setpoints that fail to account for fluctuating electricity prices and cryptocurrency values - factors critical to mining profitability. To address this, we present a physics-informed reinforcement learning (PIRL) framework for economic-thermal co-optimization in crypto mining data centers. This framework consists of a proximal policy optimization (PPO) agent, a virtual testbed powered by high-fidelity physics-based models, and an interactive frontend dashboard. The PPO agent is trained using the virtual testbed and strict hardware safety limits. This physics-informed approach allows the agent to learn a stochastic policy that dynamically balances mining revenue against operational costs by co-optimizing HVAC cooling setpoints and IT computational hashrate. The simulation results demonstrate that the integrated framework achieved an 8.62% increase in net operational profit compared to traditional baseline strategies while strictly adhering to safety-critical temperature constraints (coolant supply temperature < 32 degrees C). This work provides a scalable template for the deployment of reinforcement learning in mission critical facilities where economic volatility and physical safety must be managed simultaneously.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Reliability of Copper Inverse Opal Surfaces for Extreme-Heat-Flux Micro-Coolers in Low-Global-Warming-Potential Refrigerant R-1233zd Pool Boiling Experiments

This presentation provides a brief snapshot of the InterPACK paper InterPACK2023-113781. The paper explores copper inverse opal (CIO) surface reliability in pool boiling experiments in water and a new, low-global-warming-potential (GWP = 1) hydrofluoroolefin (HFO) refrigerant R-1233zd. The CIO-based structure is intended to develop enhanced two-phase heat transfer surfaces for extreme-heat-flux (approximately 1 kW/cm2) micro-coolers. In this study, a limited number of pool boiling experiments were performed using water and HFO-1233zd fluid, and the reliability of the CIO-based surfaces was evaluated. Critical heat flux (CHF) values in HFO-1233zd at 40 Degrees Celsius to 45 Degrees Celsius saturation temperatures and the corresponding saturation pressures were also measured. The CHF values with the refrigerant are significantly lower compared to those with water, but the refrigerant allows for a wider usable temperature range in the end application of the micro-coolers and is not limited to data centers with controlled ambient conditions. Reliability experiments with CIO surface samples - involving pool boiling with water on the CIO surfaces for approximately 48 hours and with HFO-1233zd for 144 hours - showed no structural degradation of the enhanced surface or any significant performance drop in heat transfer coefficients. The CIO surface samples in water were oxidized, most likely due to the presence of air in water and in the experimental vessel.

critical heat flux

Reducing Data Center Peak Cooling Demand and Energy Costs with Underground Thermal Energy Storage (UTES)

By recent estimates, data center energy demands are projected to consume between 6.7% and 12% of U.S. annual electricity generation by the year 2028, driven primarily by expanded demands from cloud services, big data analytics, and Artificial Intelligence (AI) (Shehabi et al., 2024). As much as 40% of data center total energy consumption are loads associated with the site infrastructure cooling systems, and these are often highly water consumptive (Aljbour et al., 2024). For energy system planners, this presents significant challenges to meeting and managing the anticipated loads, and especially the peak loads of projected data center deployments. Geothermal technologies offer two unique solutions to these challenges: 1) by serving loads through the deployment of new conventional and/or next-generation geothermal power technologies such as EGS and 2) through an often-overlooked opportunity to reduce data center peak cooling loads. The latter is the focus of this paper which explores Cold Underground Thermal Energy Storage ("Cold UTES") as an emerging industrial-scale geothermal cooling solution. This cooling solution is energy efficient, non-water-consumptive, and utilizes long duration energy storage (LDES) on both diurnal and seasonal time scales. Cold UTES has the potential to also function as a virtual power plant (VPP). The US Department of Energy's Geothermal Technologies Office is supporting R&D to understand the grid and system-wide value, costs, and impacts of deploying this emergent cooling solution at scale.

AI

Foreword: Special Section on Multiphysics Aspects of Power Electronics Packaging—Power Die, Power Module, and Converter Level: Part 2

Power electronics are increasingly being used to condition electricity for a wide array of applications, such as transportation (on land, air, and water), data centers, radio frequency, directed energy, wind, solar, and grid-tied applications. Here, to increase power density, performance, efficiency, and reliability-as well as to reduce cost-innovations and developments are needed in the multiphysics packaging of power electronics at a die, module, and converter level. This includes fundamental R&D related to emerging high-voltage, high-temperature, and high-switching-frequency power electronics, packaging materials, thermal materials and interfaces, fluid-based thermal management technologies, reliability, condition monitoring, and prognostics. Latest developments in this area are published as a Special Section on Multiphysics Aspects of Power Electronics Packaging. The first part was published in the May 2024 issue of the IEEE Transactions on Components, Packaging and Manufacturing Technology (Volume 14, Issue 5). The second part of that Special Section is being published in this issue. A brief summary of the papers included in the second part are given below.

24 POWER TRANSMISSION AND DISTRIBUTION

A Survey on the Expanding Scope and Interdisciplinary Opportunities for Processing-in-Memory Techniques

Processing-in-Memory (PIM) is emerging as a practical path to overcome the limitations of traditional von Neumann architectures. At its core, PIM systems implement computing primitives such as logic operations and multiply-accumulate acceleration through compute-in-memory, near-memory processing, or hybrid designs. The role of memory cells varies widely across technologies, acting as inputs, outputs, or analog accumulators through bit-lines and sense amplifiers. This diversity creates trade-offs in precision, bandwidth, latency, and programmability, making it difficult to build a unified understanding on the progress of the field. In this survey, we organize recent advances of PIM into three areas. First, we discuss the progress on the architectural optimizations of PIM and its integration with both DRAM and emerging non-volatile memories. Second, we examine how PIM is being used to accelerate key computing domains, including generative AI workloads and high-performance kernels, along with new approaches. Third, we highlight the growing adoption of PIM in computational sciences, where it is being applied to solve interdisciplinary problems such as genome analysis, mRNA quantification, mass spectrometry, quantum circuit simulation, wave modeling, and secure computation. Finally, we synthesize the major challenges that continue to slow PIM adoption, including manufacturing constraints, power delivery, thermal reliability, data consistency, runtime and memory-management coordination, and the difficulty of building portable software abstractions without sacrificing commercial viability. This work provides an updated, structured perspective on PIM’s potential across computing and computational sciences and the barriers that must be solved for it to reach its full impact.

Asifuzzaman, Kazi [Oak Ridge National Laboratory (

Hybrid Power Purchase Agreements for Flexible 24/7 Energy Delivery – A Comprehensive Review of Current Practices and Research Pathways

Power Purchase Agreements (PPAs) are becoming increasingly preferred among large energy consumers, such as data centers, to secure cost-effective energy and meet accelerating demand growth. Traditionally, variable renewable energy (VRE)-based PPAs operate on a pay-as-produced basis, balancing supply and demand for a relatively longer duration (e.g., annually). However, the focus is shifting toward matching supply and demand on an hourly basis to fully meet energy needs. This shift requires the integration of flexible energy resources, such as hydropower, thermal generation, and energy storage, to complement VRE sources like wind and solar, forming the foundation for 24/7 PPA. This work contributes by: (i) reviewing emerging market trends and current practices in PPA procurement, supported by data on PPA prices and technology portfolios; (ii) synthesizing the existing literature on modeling approaches for contract pricing, quantities, hybrid resource procurement, and risk management in 24/7 PPA design, while identifying key research gaps; and (iii) proposing an integrated 24/7 PPA design framework along with two contracting mechanisms from the perspectives of both PPA providers and consumers. The proposed framework highlights critical modeling challenges, risk-allocation issues, and future research opportunities for 24/7 PPA design.

24/7