Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data center design”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Implementation and Demonstration of the Digital Twin Certification System Remote Operations Framework

Microreactors are one promising advanced-reactor concept being pursued by the nuclear industry. They are distinguished by a relatively low power output of 20 MWth or less. These microreactors are intended for deployment in applications where conventional small-capacity power solutions, such as diesel generators, are either economically unfeasible or logistically challenging. Such applications include providing electric power and/or heat for remote communities, mining sites, defense installations, and humanitarian and disaster-relief missions. An important feature for the successful deployment of microreactors is their capability to be operated remotely. This capability can significantly reduce staffing costs by eliminating the need for licensed operators to be physically present at each reactor site. Instead, operators can be centralized in a single remote operations center placed in an economically advantageous location, thereby optimizing resources by consolidating expertise and enhancing operational efficiency. However, the implementation of a remote operation system for nuclear reactors raises new concerns regarding the security, reliability, and resilience of such a system. One way in which remote operations can be supported in a manner that maintains system security, reliability, and resilience is through the use of digital twins in a novel framework designed to verify and validate sensor data and commands communicated between the remote operations center and reactor. This framework, known as the Digital Twin Certification System (DTCS), has previously been proposed as an operations architecture that can bring security and resiliency levels of remote nuclear-reactor operations to a level acceptable for commercial deployment. This paper moves the proposed DTCS architecture from concept to reality by presenting the implementation and testing of the system. The rationale and implementation of the DTCS using tools such as DeepLynx and Apache Airflow, is covered in-depth. This is followed by a demonstration of the DTCS by applying the implemented system architecture to the Single Primary Heat Extraction and Removal Emulator, a small-scale non-nuclear test bed that emulates thermal behavior of a microreactor. The demonstration includes both normal and abnormal operating scenarios to highlight how the DTCS can increase the security, reliability, and resilience of a remote operations system.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Implementation and Demonstration of the Digital Twin Certification System Remote Operations Framework

Microreactors are one promising advanced-reactor concept being pursued by the nuclear industry. They are distinguished by a relatively low power output of 20 MWth or less. These microreactors are intended for deployment in applications where conventional small-capacity power solutions, such as diesel generators, are either economically unfeasible or logistically challenging. Such applications include providing electric power and/or heat for remote communities, mining sites, defense installations, and humanitarian and disaster-relief missions. An important feature for the successful deployment of microreactors is their capability to be operated remotely. This capability can significantly reduce staffing costs by eliminating the need for licensed operators to be physically present at each reactor site. Instead, operators can be centralized in a single remote operations center placed in an economically advantageous location, thereby optimizing resources by consolidating expertise and enhancing operational efficiency. However, the implementation of a remote operation system for nuclear reactors raises new concerns regarding the security, reliability, and resilience of such a system. One way in which remote operations can be supported in a manner that maintains system security, reliability, and resilience is through the use of digital twins in a novel framework designed to verify and validate sensor data and commands communicated between the remote operations center and reactor. This framework, known as the Digital Twin Certification System (DTCS), has previously been proposed as an operations architecture that can bring security and resiliency levels of remote nuclear-reactor operations to a level acceptable for commercial deployment. This paper moves the proposed DTCS architecture from concept to reality by presenting the implementation and testing of the system. The rationale and implementation of the DTCS using tools such as DeepLynx and Apache Airflow, is covered in-depth. This is followed by a demonstration of the DTCS by applying the implemented system architecture to the Single Primary Heat Extraction and Removal Emulator, a small-scale non-nuclear test bed that emulates thermal behavior of a microreactor. The demonstration includes both normal and abnormal operating scenarios to highlight how the DTCS can increase the security, reliability, and resilience of a remote operations system.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Electricity Rate Designs for Large Loads: Evolving Practices and Opportunities

Electricity demand from large load customers such as data centers is projected to grow significantly in the near term. While data centers play an important role in advancing technology innovation and economic growth in the United States, data center energy needs present challenges and opportunities for electricity supply and infrastructure. This technical brief serves as a foundation for the discussion of issues and sharing of perspectives among utilities, regulators, large load customers, and other stakeholders. As utilities and regulators explore rate structures to address growing data center electricity demand, several issues have emerged: -Fair allocation of electricity system costs to large-load customers without unfair shifting of costs to other customers -Appropriate mitigation of the financial risks associated with stranded assets from underutilized utility system investments -Mitigation of operational and resource adequacy risks if electricity demand exceeds supply -Appropriate risk-sharing in commercializing newer electricity technologies such as advanced geothermal, small modular reactors, and long duration energy storage -Accommodating the diverse needs of large-load customers, such as having the option to match electricity consumption with output from carbon-free resources or using onsite generation to provide system capacity The technical brief also identifies key design elements that aim to address these issues and uses leading examples from pending and approved rate structures, agreements, and special contracts to ground the elements in practice.

24 POWER TRANSMISSION AND DISTRIBUTION↗

AEOLUS: Advances in Experimental Design, Optimal Control, and Learning for Uncertain Complex Systems

The AEOLUS Center is dedicated to developing a unified optimization-under-uncertainty framework for (1) learning predictive models from data and (2) optimizing experiments, processes, and designs governed by these models, all driven by complex, uncertain energy systems. AEOLUS addressed the critical need for principled, rigorous, scalable, and structure-exploiting capabilities for exploring parameter and decision spaces of complex forward simulation models---the so-called outer loop. This report summarizes the work done under DE-SC0021077 on (1) nonlocal models for solidification problems, (2) a multifidelity method for a nonlocal diffusion model, and (3) multifidelity Monte Carlo methods.

97 MATHEMATICS AND COMPUTING↗

Performance and Reliability Assessment of the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) Data Advisor (ADA)

The Atmospheric Radiation Measurement (ARM) User Facility provides one of the world's largest openly accessible repositories of atmospheric observations through the ARM Data Discovery platform. Although the repository contains more than three decades of measurements collected from permanent observatories, mobile facilities, aircraft campaigns, and field experiments, identifying appropriate datasets can be challenging, particularly for new users unfamiliar with ARM instrumentation and datastream organization. To improve data accessibility, the ARM Data Center developed the ARM Data Advisor (ADA), an artificial intelligence-powered assistant designed to facilitate scientific data discovery, dataset interpretation, and user guidance. This report evaluates ADA's performance as a domain-specific scientific assistant using realistic atmospheric science workflows. The evaluation examines five key capabilities: data retrieval and curation efficiency, hallucination resistance, scientific reasoning, response to ambiguous queries, and content retention and session continuity. Representative prompts were developed to simulate typical interactions between researchers and the ARM Data Discovery platform, and ADA's responses were assessed for retrieval completeness, scientific accuracy, consistency, and practical usefulness. In these representative tests, ADA reduced the complexity of discovering and accessing ARM datasets by recommending appropriate datastreams, explaining instrumentation, interpreting metadata, and assisting with data processing workflows. ADA also exhibits strong domain knowledge of atmospheric science terminology and generally resists hallucination by acknowledging unavailable datasets and requesting clarification when appropriate. Overall, the results indicate that ADA represents a promising advancement in scientific data discovery within the ARM User Facility and has considerable potential to improve researcher productivity, particularly for new users and interdisciplinary scientists seeking efficient access to ARM observations.

Salvador, Christian [ORNL] (ORCID:0000000283287777↗

Accelerating nuclear-integrated data center pursuits in the USA: SWOT analysis, power-thermal management strategies and demonstration plan

Here, this study explores the increasing interest in leveraging nuclear power to meet the escalating energy demands of data centers in the United States (U.S.) by focusing on key factors that contribute to accelerated deployment. The study highlights the importance of N+1/N+2 power supplies (where N is the required number of units), outlines research and innovations in nuclear-integrated data center thermal management and demonstration plan. It also provides updates about status and costing of various reactor system designs. A summarized strengths, weaknesses, opportunities, and threats (SWOT) analysis shows the potential options for grid connectivity, reactors, and site selection. Suitable site discussions consider land and water availability, grid access, and optical fiber connectivity, and the study presents graded prospects for Department of Energy (DOE) sites with a specific example. Community engagement and partnerships are emphasized, particularly the roles of local government, federal agencies, utilities, and data center industry partners, which are crucial for accelerating deployment, business outreach, and approvals. The study provides actionable insights for stakeholders to accelerate the deployment of nuclear-powered data centers.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

Advances in Experimental Design, Optimal Control, and Learning for Uncertain Complex Systems (Final Report for AEOLUS)

The AEOLUS Center is dedicated to developing a unified optimization-under-uncertainty framework for: (1) learning predictive models from data; and (2) optimizing experiments, processes, and designs governed by these models, all driven by complex, uncertain energy systems. AEOLUS addresses the critical need for principled, rigorous, scalable, and structure-exploiting capabilities for exploring parameter and decision spaces of complex forward simulation models. This report summarizes the key highlights of our research during the period of performance.

97 MATHEMATICS AND COMPUTING↗

Improving Cost and Efficiency of the Scalable Solid Oxide Fuel Cells Power System

The objective of this project was to design and develop a 20kW range small-scale solid oxide fuel cells (SOFC) power system for applications such as data centers and commercial buildings. The original plan included a 5,000 hours demonstration and a Techno-Economic Analysis (TEA) which were dropped as part of project termination. The original project plan was to use a stack with a cross-flow cell design which had previously been tested for 500 hours at a community college in Malta, NY. However, it was decided to move to the advanced R-SOFC co-flow cell developed under Department of Energy Award DE-FE0031971. The advanced cell design has the advantage of a larger active area for the same manufacturing footprint which results in fewer required cells for the same stack power, hence a higher volumetric power density (kW/L) and lower cost per kW than the original cross-flow cell design. A full SOFC system Simulink model was developed and calibrated with testing data from a fuel cell stack and BOP (balance of plant) components. The simulation results from the calibrated model showed an acceptable match with the experimental data. A structural analysis conducted for various load scenarios indicated no high stress areas for all spatial directions. Major electrical system components were acquired, built and successfully tested. System sensors were verified and validated against controls. Safety checks, a diagnostic check, PID tuning, and control software commissioning tasks were also conducted. The power electronics prototype was delivered and trial testing completed. Balance of Plant component testing and simulation work was conducted to characterize Reformer-Heat Exchanger heat transfer and backpressure and reformer catalyst methane conversion and product selectivity. Simulations were conducted to design the Anode and Cathode fluid passages and size the air-air and fuel-fuel heat exchangers. A Burner operation map was created from test data and the Anode Gas Recirculation blower was tested to evaluate its durability. The SOFC system used a horizontal style design where components sit directly on a casting with a direct connection to the skid. This design has efficient packaging and a small footprint with approximate dimensions of 750 mm x 700 mm x 1700 mm. An SOFC system was built and successfully tested at the Malta, NY facility The system for over 500 hours under load of which over 300 hours was at full load of 20 kW.

30 DIRECT ENERGY CONVERSION↗

Data Center Cybersecurity, Supply Chain Risk Management, and Emerging Regulation Cohort Summary: Takeaways and Action Plans

This report summarizes the outcomes of the Data Center Cohort under the Department of Energy’s Technical Assistance for Digital Assurance (TADA) initiative, aimed at enhancing grid resilience through cybersecurity, supply chain risk management (SCRM), and Cyber-Informed Engineering (CIE). The cohort engaged 17 organizations across utilities, data center operators, vendors, and technology providers in three sessions combining presentations, discussions, and exercises. Key topics included AI-driven load behavior, cybersecurity vulnerabilities in UPS/BESS and cooling systems, governance gaps at utility–data center boundaries, and supply chain integrity. Five cross-cutting themes emerged: interconnection architecture vulnerabilities, fragmented governance, AI-driven stability risks, lack of regulatory frameworks, and long-term supply chain concerns. Actionable recommendations were developed, including implementing DMZ segmentation, formalizing vendor access agreements, designing AI workload limits, and advancing standards through NERC and state-level programs. These strategies aim to strengthen resilience, clarify responsibilities, and ensure secure integration of data centers into the grid.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

AEOLUS: Advances in Experimental Design, Optimal Control, and Learning for Uncertain Complex Systems

Sustained advances in the mathematics of modeling and simulation have resulted in the capability today for routine simulation of a number of large scale complex DOE-relevant systems. As remarkable as this capability for solving the so-called forward problem is, it is typically only the first step-an inner loop within an outer loop that explores the simulation model's parameter space and decision space to characterize uncertainty in the model's predictions, learn unknown model parameters from data, design the most informative experiments, determine optimal control strategies, and create optimal designs. Broadly, what unifies all of these outer loop problems is that they are, in one form or another, optimization problems over parameter/control/design space that are constrained by complex uncertain models. To fully realize the power of scientific simulation as a basis for scientific discovery, technological innovation, and rational decision-making, it is imperative to move beyond simulation to tackle the outer loop of optimization for learning from data, experimental design, and control with complex uncertain models. When the models under consideration are large-scale and complex, and when the optimization variable and uncertain parameter spaces are high (or infinite) dimensional, this constitutes a grand challenge of the highest order, and is intractable with conventional methods. To overcome these challenges, the AEOLUS Center was established to develop a unified mathematical, computational, and statistical framework for (1) Learning predictive models from complex data via Bayesian inference and optimization, and (2) Optimizing experiments, processes, and designs using the resulting uncertain models. These problems are intractable with conventional methods, for several reasons: (1) The simulation problems that govern the inner loops of the optimization problems are expensive to execute (due to severe nonlinearity, heterogeneity, multiphysics/multiscale coupling); (2) The optimization variable and uncertain parameter spaces are high dimensional, often stemming from discretizations of infinite dimensional fields such as initial conditions, sources, or material properties. We argue that the key to overcoming these challenges is to develop new mathematical, computational, and statistical methods that exploit the structure of the Bayesian inference and optimization problems mediated by their underlying complex uncertain models. This structure includes the regularity, sparsity, geometry, low intrinsic dimensionality, and multifidelity nature of the maps from uncertain parameter/optimization variable spaces to the specific objectives targeted: Bayesian inference, optimal experimental design, and optimal control design. Black box methods developed as generic tools are incapable of exploiting this structure. To be successful, we must create, integrate, and cross-fertilize ideas across multiple areas of applied math--including approximation theory, Bayesian inference, data science, experimental design, information theory, machine learning, model reduction, optimal control theory, parallel algorithms, PDE-constrained optimization, randomized algorithms, stochastic optimization, and uncertainty quantification--all while exploiting the structure of the problems at hand. With this goal in mind, we have marshaled a team of leading authorities in these areas. While the methods we develop will be broadly applicable across a wide spectrum of DOE problems in which experiments inform models and the systems those models describe must be optimized under uncertainty, we have chosen a specific area, advanced manufacturing and materials, to drive our work. AMM is characterized by complex models across multiple scales, and is a rich source of challenging problems in inference, experimental design, and optimal control, requiring multifaceted and integrated advances in applied mathematics. As such, AMM serves as an excellent vehicle to motivate and demonstrate the advances in applied mathematics developed by our center.

97 MATHEMATICS AND COMPUTING↗

Energy efficient integrated photonic systems based on inverse design

The energy footprint of modern information processing and communications systems is immense and utilizes a significant fraction of total global energy usage. Data centers alone consume over 70 billion kilowatt-hours per year. Much of this energy usage is intrinsic to the use of electronic wiring, making optical-based technologies a necessary and promising route to mitigating energy consumption in short- to medium-distance communication links. In this project, we developed and implemented a framework for the design of optical components, based on machine learning, which enables optical components relevant to optical information processing to be realized at their physical performance limits.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Solar Resource Measurements in Eugene, OR: Cooperative Research and Development Final Report, CRADA Number CRD-07-00252

Site-specific, long-term, continuous, and high-resolution measurements of solar irradiance are important for developing renewable resource data. These data are used for several research and development activities consistent with the NLR mission: establish a national 3-year climatological database of measured solar irradiances; provide high quality ground-truth data for satellite remote sensing validation; support development of radiative transfer models for estimating solar irradiance from available meteorological observations; provide solar resource information needed for technology deployment and operations. Data acquired under this agreement will be available to the public through NLR's Measurement & Instrumentation Data Center – MIDC (http://www.nlr.gov/midc) Or the Renewable Resource Data Center - RReDC (http://rredc.nlr.gov). The MIDC offers a variety of standard data display, access, and analysis tools designed to address the needs of a wide user audience (e.g., industry, academia, and government interests).

14 SOLAR ENERGY↗

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.

crossanalysis↗

Electricity Rate Designs for Large Loads: Evolving Practices and Opportunities 2026 Update

Electricity demand from large-load customers such as data centers is projected to grow significantly in the near term. While these large loads play an important role in advancing technology innovation and economic growth in the United States, meeting their energy needs requires utilities and regulators to consider important operational and financial risks, such as insufficient energy supply or underutilized investments, that can impact all customers. This paper builds on similar research published in January 2025, providing an overview of how utilities and regulators are managing these risks through different tariffs, including rate structures and electric service agreements. Regulators, utilities, customers, and other stakeholders can use this paper as a foundation when discussing issues and sharing perspectives on developing or reviewing large-load tariffs.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Integrating AI Data Centers with the Power Grid

The rapid expansion of artificial intelligence (AI) has triggered an unprecedented surge in electricity demand, with US data center energy use projected to double or triple 2023 levels by 2028. This exponential growth places strain on grid infrastructure, which can hinder timely construction of desired computing capacity. To bridge this supply-demand gap, utilities and AI developers are increasingly turning to demand flexibility, a strategy that incentivizes shifting or reducing power use during peak periods of grid stress. Data centers are uniquely equipped for flexible operations due to their digital workloads, built-in redundancy, and onsite energy assets. This article outlines four primary mechanisms to enable data center flexibility: computational load flexibility (shifting tasks temporally or geographically), flexible use of core facility infrastructure adjustments, energy storage utilization, and onsite electricity generation. To encourage adoption, utilities are deploying new tariff designs, including voluntary interruptible service riders, mandated flexibility requirements, and streamlined interconnection processes for flexible loads. For the highly capitalized and rapidly growing AI industry, the primary motivators for embracing these strategies are expediting facility interconnection, satisfying emerging regulatory mandates, and mitigating community resistance. While demand flexibility cannot substitute the long-term need for new bulk power generation, it serves as an essential, immediate solution for enabling near-term deployment. By transforming data centers from grid stressors into stabilizing assets, flexible operations can ensure reliable grid integration, ease market pressures, and support a resilient power system.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Design of Experiments for Dynamic Test Runs in Solvent-Based CO 2 Capture Pilot Plants

Test runs in the pilot plants consume significant resources, and therefore, the learning from test runs should be maximized. Test runs conducted in the pilot plants are often steady state. It takes several hours for reaching steady-state in the pilot plants, and thus, the duration of the test runs needs to be long even for collecting few steady-state data points. On the other hand, a large number of measurements can be collected through dynamic test runs in a short span of time. This paper presents a systematic design of dynamic experiments (DoDEs) for identifiability of model parameters, which is achieved by persistently exciting the inputs signals. A pseudorandom binary sequence (PRBS) is designed as the input signal for DoDE due to its efficiency in obtaining sufficient spectral content. However, due to the long sequence size of the PRBS signal, a Schroeder-phase input signal, which is a multisine signal, is also designed. Tests for both types of signals are run in the Pilot Solvent Test Unit (PSTU) at the National Carbon Capture Center in Wilsonville, Alabama. The transient data are used to solve dynamic data reconciliation and parameter estimation problem. The estimated parameters are found to be not only superior to those estimated from using data collected from hundreds of steady-state test runs in a nonreactive (air–water) system, but the parameters could be estimated by using the dynamic data collected for about 24 h from the pilot plant for the MEA-H 2 O–CO 2 system.

CO2 capture↗

Hands-On, Heads-Up: Blending Cyber T&E with Data Science-Driven Training in Jupyter Notebooks

In an era of increasingly sophisticated threats to critical infrastructure, cybersecurity professionals must be more than just aware; they must be immersed, agile, and equipped to operate in environments where failure is not an option. Nowhere is this truer than in the nuclear sector, where cyber-physical systems, regulatory scrutiny, and insider threat potential demand a new generation of hands-on, technically fluent defenders. This paper presents a unified training approach that integrates Cybersecurity Test and Evaluation (T&E) with data science techniques using Jupyter Notebooks as the interactive lab environment. The program centers on a modular, scenario-driven curriculum designed to build not just knowledge but practical capability in the assessment and defense of radiation detection systems, firmware interfaces, and operational security postures.

98 - NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL↗

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

Naug, Avisek [Hewlett Packard Enterprise]↗