Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Center”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Unlocking Synergistic Benefits of Colocation of Data Centers at Airports

This white paper evaluates the potential of a new energy system configuration: Strategic colocation of data centers - on airport property or adjacent to airports - in order to realize significant operational, economic, and environmental benefits to airports, data centers, and the communities they serve. While increasing demand from both airports and data centers requires both parties to make significant infrastructure investments, colocation can mitigate some of the energy, land use, permitting, and connectivity challenges of siting and powering these industries.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Addressing Data Center Cooling Needs through the Use of Subsurface Thermal Energy Storage Systems

This study aims to evaluate the feasibility of addressing the cooling needs for information technology (IT) equipment in data centers by using reservoir thermal energy storage (RTES) to provide reliable and sustainable low-temperature fluid. This project focuses on the technical viability of operating such a system in Houston, Texas, which is a representative location for crypto mining data centers. An analysis has been performed to investigate the technical feasibility with climate data for Houston. Results show that data centers on the scale of 30 MW, and operating at a temperature of 27 degrees C, can be reliably cooled by a combined RTES and dry coolers setup. A techno-economic analysis will be performed and energy/water saving benefits will be quantified in the future. In addition, the study will be extended to other representative data center locations with different climates and geographical locations.

cooling↗

The water use of data center workloads: A review and assessment of key determinants

The global importance of data center water use is increasing with the rapid growth of digitalization and artificial intelligence. This study analyzes the factors influencing workload-level water use, measured in liters consumed per workload, to guide water-saving strategies in data centers. Our findings reveal workload-level water use variations exceeding 10,000-fold, driven by over 1000-fold differences in water consumption per kilowatt hour of server electricity consumed and approximately 10-fold differences in server workload efficiency. Key determinants are ranked as server efficiency, electrical grid water consumption factors, server utilization, cooling system type, infrastructure efficiency, climate zone, inactive server percentage, and server refresh cycle. Notably, there is no single recipe for minimizing water use; instead, optimal outcomes depend on tailored combinations of these factors. This analysis addresses critical knowledge gaps by identifying the determinants of data center water use and exploring their achievable minima under diverse site-specific constraints.

Data centers↗

Navigating Economies of Scale and Multiples for Nuclear-Powered Data Centers and Other Applications with High Service Availability Needs

Nuclear energy is increasingly being considered for such targeted energy applications as data centers in light of their high capacity factors and low carbon emissions. This paper focuses on assessing the tradeoffs between economies of scale versus mass production to identify promising reactor sizes to meet data center demands. A framework is then built using the best cost estimates from the literature to identify ideal reactor power sizes for the needs of the given data center. Results should not be taken to be deterministic but highlight the variability of ideal reactor power output against the required demand. While certain advocates claim that with the gigawatts of clean, firm energy needed, large plants are ideal, others advocate for SMRs that can be deployed in large quantities and reap the benefits from learning effects. The findings of this study showcase that identifying the optimal size for a reactor is likely more nuanced and dependent on the application and its requirements. Overall, the study does show potential economic promise for coupling nuclear reactors to data centers and industrial heat applications under certain key conditions and assumptions.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Integrated SOAs enable energy-efficient intra-data center coherent links

Coherent optical links are becoming increasingly attractive for intra-data center applications as data rates scale. Realizing the era of high-volume short-reach coherent links will require substantial improvements in transceiver cost and power efficiency, necessitating a reassessment of conventional architectures best-suited for longer-reach links and a review of assumptions for shorter-reach implementations. In this work, we analyze the impact of integrated semiconductor optical amplifiers (SOAs) on link performance and power consumption, and describe the optimal design spaces for low-cost and energy-efficient coherent links. Placing SOAs after the modulator provide the most energy-efficient link budget improvement, up to 6 pJ/bit for large link budgets, despite any penalties from nonlinear impairments. Increased robustness to SOA nonlinearities makes QPSK-based coherent links especially attractive, and larger supported link budgets enable the inclusion of optical switches, which could revolutionize data center networks and improve overall energy efficiency.

Maharry, Aaron (ORCID:0000000273732697)↗

Deployment of inference as a service at the US CMS Tier-2 data centers

Coprocessors, especially GPUs, will be a vital ingredient of data production workflows at the HL-LHC. At CMS, the GPU-as-a-service approach for production workflows is implemented by the SONIC project (Services for Optimized Network Inference on Coprocessors). SONIC provides a mechanism for outsourcing computationally demanding algorithms, such as neural network inference, to remote servers, where requests from multiple clients are intelligently distributed across multiple GPUs by a load-balancing service. This talk highlights the recent progress in deploying SONIC at selected U.S. CMS Tier-2 data centers. Using realistic CMS Run3 data processing workflows, such as those containing transformer-based algorithms, we demonstrate how SONIC is integrated into the production-like environment to enable accelerated inference offloading. We will present developments from both the client and server sides, including production job and data center configurations for NVIDIA and AMD GPUs. We will also present performance scaling benchmarks and discuss the challenges of operating SONIC in CMS production, such as server discovery, GPU saturation, fallback server logic, etc.

Holzman, Burt↗

Data Center Cybersecurity, Supply Chain Risk Management, and Emerging Regulation Cohort Summary: Takeaways and Action Plans

This report summarizes the outcomes of the Data Center Cohort under the Department of Energy’s Technical Assistance for Digital Assurance (TADA) initiative, aimed at enhancing grid resilience through cybersecurity, supply chain risk management (SCRM), and Cyber-Informed Engineering (CIE). The cohort engaged 17 organizations across utilities, data center operators, vendors, and technology providers in three sessions combining presentations, discussions, and exercises. Key topics included AI-driven load behavior, cybersecurity vulnerabilities in UPS/BESS and cooling systems, governance gaps at utility–data center boundaries, and supply chain integrity. Five cross-cutting themes emerged: interconnection architecture vulnerabilities, fragmented governance, AI-driven stability risks, lack of regulatory frameworks, and long-term supply chain concerns. Actionable recommendations were developed, including implementing DMZ segmentation, formalizing vendor access agreements, designing AI workload limits, and advancing standards through NERC and state-level programs. These strategies aim to strengthen resilience, clarify responsibilities, and ensure secure integration of data centers into the grid.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Data Center Facility Monitoring with Physics Aware Approach

U.S. Department of Energy's National Renewable Energy Laboratory (NREL) hosts one of the world's most energy-efficient HPC data centers; this system uses component-level warm-water liquid cooling to efficiently remove heat from the data center and capture it for reuse in the building or rejection to the atmosphere. Given the complexity of this system, building data-driven tools for holistically monitoring and operating the entire data center is a priority for ensuring maximal efficiency and resiliency. In this advanced smart facility, over one million metrics are recorded per minute using state-of-the-art streaming data architecture and software to capture and process the state of the system in real time. Here we detail two efforts to effectively analyze, visualize, and interpret this large volume streaming data. We have developed a novel, flexible system for identifying and visualizing individual metric anomalies and component performance across the data center through automatic metadata extraction and physically-motivated visualization for quick interpretation. Additionally, to directly connect system maintenance to data stream processing we explore a physics informed multi-metric drift and anomaly detection application to detect scale-build up in heat exchangers.

anomaly detection↗

Organizational and psychological measures for data center energy efficiency: barriers and mitigation strategies

Abstract It was last estimated that in 2020, data centers comprised approximately 2% of total US electricity consumption, with an estimated annual growth rate of 4%. As our country increasingly relies on information technology (IT), our data centers (DCs) will need to increase their energy efficiency (EE) to stabilize their energy consumption. The task of studying EE in DCs is complicated by the interconnected nature of humans and mission-critical technical systems. Moreover, the literature tends to focus on technology solutions such as improvements to IT equipment, cooling infrastructure, and software, without addressing organizational and psychological drivers. Our research demystifies the complex interactions between humans and DCs, by asking What non-technical barriers impede EE investment decision-making and/or implementing energy management strategies? To begin to answer this question, we perform a literature review of 86 resources, ranging from peer-reviewed journal publications to handbooks. We also consider related fields such as organizational behavioral management and energy intensive buildings. We develop a public Zotero library, perform content coding, and complete a rudimentary network analysis. Our findings from the literature review suggest that (1) technological solutions are abundant in the literature but fall short of providing practical guidance on the pitfalls of implementation, (2) making energy efficiency a priority at the executive level of organizations will be largely ineffective if the IT and facilities staff are not directly incentivized to increase EE, and (3) there is minimal current understanding of how the individual psychologies of IT and facilities staff affect EE implementation in DCs. In the next phase of our research, we plan to interview data center operators/experts to ground-truth our literature findings and collaboratively design decarbonization policy solutions that target organizational structure, empower individual staff, and foster a supportive external market.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

Naug, Avisek [Hewlett Packard Enterprise]↗

Server rack for improved data center management

Methods and systems for data center management include collecting sensor data from one or more sensors in a rack; determining a location and identifying information for each asset in the rack using a set of asset tags associated with respective assets; communicating the sensor and asset location to a communication module; receiving an instruction from the communication module; and executing the received instruction to change a property of the rack.

Bermudez Rodriguez, Sergio A.↗

Server rack for improved data center management

Methods and systems for data center management include collecting sensor data from one or more sensors in a rack; determining a location and identifying information for each asset in the rack using a set of asset tags associated with respective assets; communicating the sensor and asset location to a communication module; receiving an instruction from the communication module; and executing the received instruction to change a property of the rack.

Bermudez Rodriguez, Sergio↗

Integrating AI Data Centers with the Power Grid

The rapid expansion of artificial intelligence (AI) has triggered an unprecedented surge in electricity demand, with US data center energy use projected to double or triple 2023 levels by 2028. This exponential growth places strain on grid infrastructure, which can hinder timely construction of desired computing capacity. To bridge this supply-demand gap, utilities and AI developers are increasingly turning to demand flexibility, a strategy that incentivizes shifting or reducing power use during peak periods of grid stress. Data centers are uniquely equipped for flexible operations due to their digital workloads, built-in redundancy, and onsite energy assets. This article outlines four primary mechanisms to enable data center flexibility: computational load flexibility (shifting tasks temporally or geographically), flexible use of core facility infrastructure adjustments, energy storage utilization, and onsite electricity generation. To encourage adoption, utilities are deploying new tariff designs, including voluntary interruptible service riders, mandated flexibility requirements, and streamlined interconnection processes for flexible loads. For the highly capitalized and rapidly growing AI industry, the primary motivators for embracing these strategies are expediting facility interconnection, satisfying emerging regulatory mandates, and mitigating community resistance. While demand flexibility cannot substitute the long-term need for new bulk power generation, it serves as an essential, immediate solution for enabling near-term deployment. By transforming data centers from grid stressors into stabilizing assets, flexible operations can ensure reliable grid integration, ease market pressures, and support a resilient power system.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Revolutionizing thermal Management in Next-Generation AI data centers: Challenges and breakthrough innovations

Data centers (DCs) serve as critical infrastructure for powering the growth and evolution of AI. Next-generation AI DCs present unique challenges in thermal management driven by unprecedented computational demands. This paper provides a comprehensive summary of key stakeholder perspectives on technology gaps, infrastructure requirements, test bed needs, emerging opportunities, and preliminary solutions related to thermal management for AI DCs. It establishes six strategic pillars of thermal management for next generation AI DC: reliability, deployability, efficiency, resilience, measurability, and valorization. The discussion spans a range of critical topics, including advanced cooling technologies, thermal strategies for emerging modular and edge DCs, system-level optimization and control frameworks, infrastructure planning and grid integration designs, benchmarking approaches, and pathways for waste heat recovery and reuse. The proposed research, development, and demonstration efforts are aimed at accelerating the deployment of AI DCs while ensuring energy efficiency, reliability, safety, and regulatory compliance.

Wang, Pengtao [ORNL] (ORCID:0000000214713429)↗

A Physics-Informed Reinforcement Learning Framework for Economic-Thermal Co-Optimization of Crypto Mining Data Centers: Preprint

The rapid expansion of cryptocurrency mining has created a new class of high-density data centers characterized by extreme thermal flux and high sensitivity to volatile economic markets. Traditional thermal management strategies, typically reliant on rule-based control, maintain static setpoints that fail to account for fluctuating electricity prices and cryptocurrency values - factors critical to mining profitability. To address this, we present a physics-informed reinforcement learning (PIRL) framework for economic-thermal co-optimization in crypto mining data centers. This framework consists of a proximal policy optimization (PPO) agent, a virtual testbed powered by high-fidelity physics-based models, and an interactive frontend dashboard. The PPO agent is trained using the virtual testbed and strict hardware safety limits. This physics-informed approach allows the agent to learn a stochastic policy that dynamically balances mining revenue against operational costs by co-optimizing HVAC cooling setpoints and IT computational hashrate. The simulation results demonstrate that the integrated framework achieved an 8.62% increase in net operational profit compared to traditional baseline strategies while strictly adhering to safety-critical temperature constraints (coolant supply temperature < 32 degrees C). This work provides a scalable template for the deployment of reinforcement learning in mission critical facilities where economic volatility and physical safety must be managed simultaneously.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Accelerating nuclear-integrated data center pursuits in the USA: SWOT analysis, power-thermal management strategies and demonstration plan

Here, this study explores the increasing interest in leveraging nuclear power to meet the escalating energy demands of data centers in the United States (U.S.) by focusing on key factors that contribute to accelerated deployment. The study highlights the importance of N+1/N+2 power supplies (where N is the required number of units), outlines research and innovations in nuclear-integrated data center thermal management and demonstration plan. It also provides updates about status and costing of various reactor system designs. A summarized strengths, weaknesses, opportunities, and threats (SWOT) analysis shows the potential options for grid connectivity, reactors, and site selection. Suitable site discussions consider land and water availability, grid access, and optical fiber connectivity, and the study presents graded prospects for Department of Energy (DOE) sites with a specific example. Community engagement and partnerships are emphasized, particularly the roles of local government, federal agencies, utilities, and data center industry partners, which are crucial for accelerating deployment, business outreach, and approvals. The study provides actionable insights for stakeholders to accelerate the deployment of nuclear-powered data centers.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗