Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “HPC system”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Liquid Cooling for HPC - National Renewable Energy Lab

NREL has nearly a decade of experience in liquid cooling or HPC. The technical approach will be presented along with success stories on energy, cost and emission savings over two generations of HPC systems. Lessons learned will also be shared.

energy efficient computing↗

Coatings for CSP Lifetime

The feasibility and performance of tower-technology-based concentrated solar power (CSP) is highly dependent on the efficiency of the energy transformation from sun to heat at the receiver. The higher the solar absorptivity of the receiver coating, the higher the efficiency of the plant as a whole. BrightSource Energy (BSE) has developed a series of High-Performance Coating (HPC) systems in order to achieve high absorptivity over the plant lifetime (25-35 years). This requires stable coatings that are easily applicable on the large receiver surface and will maintain their optical properties under intense solar flux and thousands of heating and cooling cycles in desert conditions. Different coating formulations are required according to the differing plant conditions: receiver materials, operating conditions (temperatures, daily cycles, etc.), and environmental conditions. BSE also developed a coating for the next generation of CSP receivers, such as those being under DOE’s CSP Gen3 program, which will be operated with high temperature heat transfer fluids at temperatures of up to 800°C, which is significantly hotter than the operating temperature of current systems. Once the coating is formulated, the next challenge is evaluating its lifetime properties. BSE has found several independent failure modes that impact HPC absorptivity degradation: • Decrease in HPC optical properties due to oxidation in the receiver tubes surface below the HPC; • HPC film deterioration due to cycling of temperatures and humidity due to daily operation startup and shutdown as well as changing ambient conditions; • Mechanical degradation due to erosion by sand and wind. Existing test methods examine various aspects independently, but do not provide a combined accelerated lifetime result. Creating such a combined test suite, with a way to interpret the results to predict the coating’s projected lifetime, was the ultimate goal of this project. The project was divided into three major workstreams: lab testing (individual failure mode tests and combined failure mode tests); developing a theoretical model for aging; and validation of the test apparatus via on-sun testing in near real-world conditions at CIEMAT-PSA. Developing a test apparatus that accurately controlled the temperature while also introducing the desired solar flux proved more challenging than expected. While in the end we did succeed in creating a test apparatus that can control temperature, solar flux, and humidity, the results did not appear to accelerate the lifetime of the samples as desired. We suspect that to properly accelerate the samples we must also subject the samples to increased amounts of oxygen. Similarly, while we successfully created a combined model that is publicly available, we were unable to validate it sufficiently to feel comfortable recommending it as a general guideline.

14 SOLAR ENERGY↗

CACTI Radar b1 Processing: Corrections, Calibrations, and Processing Report

The U.S. Department of Energy’s (DOE) Atmospheric Radiation Measurement (ARM) user facility deployed a large number of instruments to a region nearby the Sierra de Córdobas mountains in Argentina as part of the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) field campaign (1). During this campaign, four radars were installed at a site outside of Villa Yacanto as shown in Figure 1. As part of a post-campaign effort to improve the usability of these data, a significant activity was undertaken towards the calibration, correction, and improvement of the data quality of these radar datastreams. This process in ARM nomenclature is referred to as generating a “b1” datastream. While these “b1” standards may imply different corrections or standards for various ARM instruments, for radars it refers to a datastream that has been calibrated (and cross-calibrated), including a serious effort to deliver the highest-quality (well-characterized) data possible. This report details (i) the status/quality of the original “a1” (raw) data, (ii) the corrections and calibrations that are applied to generate these b1 datastreams, (iii) the details of the applied algorithms and how radar offset/calibration numbers were determined for the eventual corrections, and (iv) the new and flexible plug-in-based processing system designed during CACTI for radar b1 activities (current, future) that interfaces with ARM’s Data Integrator (ADI) and high-performance computing (HPC) system.

47 OTHER INSTRUMENTATION↗

Softwarized Federations of Science Instruments with Edge-Continuum Containers

Significant expansion of capabilities of DOE science complex of supercomputers, instruments and networks is expected as powerful experimental facilities, exascale computers and terabit networks are added. Combined with the advances in edge and cloud computing technologies, DOE science users now have the promise of unprecedented execution of complex, continuum workflows, namely, small and latency-sensitive computations at the edge and site, massive computations at remote HPC systems, and everything in between on the cloud. But, bringing this capability to the science user requires overcoming the overwhelming complexity of forming the federations of systems, and efficiently and effectively orchestrating the workflows while ensuring high utilization of the expensive facilities. Current manual configuration of the federated systems simply will not scale, since the coordination across sites may take weeks to months, often leading to under-utilized and hard-to-diagnose compositions. A powerful, composable software stack will be developed to (i) wrap the systems so that federations can be composed fast in software, and (ii) containerize computations to be orchestrated across the edge, site, cloud and HPC resources.

Rao, Nageswara S.↗

Execute BEE workflows on private cloud infrastructure-2.3.6.01 - LANL ATDM ST / STNS01-22 Milestone Completion Documentation (BEE-FY21 P6-2) [Slides]

This work involves the creation of the Cloud Launcher, a new subcomponent of BEE, and the extension of the BEETaskManager to run on Cloud systems. BEE will be able to interact with the Google Compute Engine and OpenStack cloud APIs to set up simple Cloud clusters for launching HPC job scripts. BEE will use existing functionality to launch jobs that previously could only be launched on HPC systems. The BEETaskManager will handle launching tasks on the Cloud cluster.

97 MATHEMATICS AND COMPUTING↗

Bridging paradigms: Designing for HPC-Quantum convergence

Here, this paper presents a comprehensive software stack architecture for integrating quantum computing (QC) capabilities with High-Performance Computing (HPC) environments. While quantum computers show promise as specialized accelerators for scientific computing, their effective integration with classical HPC systems presents significant technical challenges. We propose a hardware-agnostic software framework that supports both current noisy intermediate-scale quantum devices and future fault-tolerant quantum computers, while maintaining compatibility with existing HPC workflows. The architecture includes a quantum gateway interface, standardized APIs for resource management, and robust scheduling mechanisms to handle both simultaneous and interleaved quantum–classical workloads. Key innovations include: (1) a unified resource management system that efficiently coordinates quantum and classical resources, (2) a flexible quantum programming interface that abstracts hardware-specific details, (3) A Quantum Platform Manager API that simplifies the integration of various quantum hardware systems, and (4) a comprehensive tool chain for quantum circuit optimization and execution. We demonstrate our architecture through implementation of quantum–classical algorithms, including the variational quantum linear solver, showcasing the framework’s ability to handle complex hybrid workflows while maximizing resource utilization. This work provides a foundational blueprint for integrating QC capabilities into existing HPC infrastructures, addressing critical challenges in resource management, job scheduling, and efficient data movement between classical and quantum resources.

97 MATHEMATICS AND COMPUTING↗

HPC Info Sheet

A flyer with current HPC systems.

97 - MATHEMATICS AND COMPUTING↗

Understanding GPU Memory Corruption at Extreme Scale: The Summit Case Study

GPU memory corruption and in particular double-bit errors (DBEs) remain one of the least understood aspects of HPC system reliability. Albeit rare, their occurrences always lead to job termination and can potentially cost thousands of node-hours, either from wasted computations or as the overhead from regular checkpointing needed to minimize the losses. As supercomputers and their components simultaneously grow in scale, density, failure rates, and environmental footprint, the efficiency of HPC operations becomes both an imperative and a challenge. We examine DBEs using system telemetry data and logs collected from the Summit supercomputer, equipped with 27,648 Tesla V100 GPUs with 2nd-generation high-bandwidth memory (HBM2). Using exploratory data analysis and statistical learning, we extract several insights about memory reliability in such GPUs. We find that GPUs with prior DBE occurrences are prone to experience them again due to otherwise harmless factors, correlate this phenomenon with GPU placement, and suggest manufacturing variability as a factor. On the general population of GPUs, we link DBEs to short- and long-term high power consumption modes while finding no significant correlation with higher temperatures. We also show that the workload type can be a factor in memory’s propensity to corruption.

Oles, Vlad↗

Benchmark Tracking System for Performance Monitoring

Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.

97 MATHEMATICS AND COMPUTING↗

OptiBench: An Optimization Benchmark Tool for Renewable Energy Problems

We propose a benchmark framework and visualization tool, OptiBench, for analyzing the performance of state-of-the-art optimization solvers across a variety of optimization problems in renewable energy research. Our framework is designed from the ground up in the Julia programming language and enables analysis at scale on high performance computing (HPC) systems. Our visualization tool allows effortless evaluation of optimization solver performance, robustness, and accuracy through intuitive plots, e.g., performance profiles, heat maps, and distribution plots. We have tested three benchmark suites relevant to the modeling of renewable energy systems, viz., CUTEst, PGLib-OPF, and WaterTAP water treatment optimization problems. We illustrate benchmarking of CUTEst using OptiBench on the National Renewable Energy Laboratory's (NREL) HPC Kestrel. Our findings indicate that MA57 HSL linear solver demonstrated the best overall performance for an experimental IPOPT implementation. Our work is ongoing and we intend to add support for more optimization solvers and benchmark test suites in the future.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

OptiBench: An Optimization Benchmark Tool for Renewable Energy Problems

We propose a benchmark framework and visualization tool, OptiBench, for analyzing the performance of state-of-the-art optimization solvers across a variety of optimization problems in renewable energy research. Our framework is designed from the ground up in the Julia programming language and enables analysis at scale on high performance computing (HPC) systems. Our visualization tool allows effortless evaluation of optimization solver performance, robustness, and accuracy through intuitive plots, e.g., performance profiles, heat maps, and distribution plots. We have tested three benchmark suites relevant to the modeling of renewable energy systems, viz., CUTEst, PGLib-OPF, and WaterTAP water treatment optimization problems. We illustrate benchmarking of CUTEst using OptiBench on the National Laboratory of the Rockies's (NLR) HPC Kestrel. Our findings indicate that MA57 HSL linear solver demonstrated the best overall performance for an experimental IPOPT implementation. Our work is ongoing and we intend to add support for more optimization solvers and benchmark test suites in the future.

97 MATHEMATICS AND COMPUTING↗

Automated Integration of Continental-Scale Observations in Near-Real Time for Simulation and Analysis of Biosphere–Atmosphere Interactions

The National Ecological Observatory Network (NEON) is a continental-scale observatory with sites across the US collecting standardized ecological observations that will operate for multiple decades. To maximize the utility of NEON data, we envision edge computing systems that gather, calibrate, aggregate, and ingest measurements in an integrated fashion. Edge systems will employ machine learning methods to cross-calibrate, gap-fill and provision data in near-real time to the NEON Data Portal and to High Performance Computing (HPC) systems, running ensembles of Earth system models (ESMs) that assimilate the data. For the first time gridded EC data products and response functions promise to offset pervasive observational biases through evaluating, benchmarking, optimizing parameters, and training new machine learning parameterizations within ESMs all at the same model-grid scale. Leveraging open-source software for EC data analysis, we are already building software infrastructure for integration of near-real time data streams into the International Land Model Benchmarking (ILAMB) package for use by the wider research community. We will present a perspective on the design and integration of end-to-end infrastructure for data acquisition, edge computing, HPC simulation, analysis, and validation, where Artificial Intelligence (AI) approaches are used throughout the distributed workflow to improve accuracy and computational performance.

Durden, David J.↗

Clippy

Clippy (CLI + PYthon) is a Python language interface to HPC resources. Precompiled binaries that execute on HPC systems are exposed as methods to a dynamically-created Clippy Python object, where they present a familiar interface to researchers, data scientists, and others. Clippy allows these users to interact with HPC resources in an easy, straightforward environment - at the REPL, for example, or within a notebook - without the need to learn complex HPC behavior and arcane job submission commands.

Bromberger, SethA.↗

Advanced Computing Annual Report 2024

In fiscal year (FY) 2024, the National Renewable Energy Laboratory (NREL) took a major leap forward with the completed full buildout of Kestrel, the Office of Energy Efficiency and Renewable Energy's newest high-performance computing (HPC) system. Kestrel is already supporting science across the portfolio, bringing roughly 44 petaflops of computing power, which is more than five times the capacity of our previous supercomputer, Eagle. By delivering greater GPU capacity, Kestrel enables faster progress in artificial intelligence (AI) and opens new avenues in energy research - from defining long-term planning scenarios to accommodate a growing power system to material discovery to improving energy efficiency in photovoltaics (PV). Across the portfolio, research is being accelerated by Kestrel's impressive power. During FY24, 427 projects and more than 700 researchers used NREL's HPC, supporting the U.S. Department of Energy's Office of Energy Efficiency and Renewable Energy across 13 funding areas. Through these collaborations, researchers produced more than 450 technical outputs, including 195 articles in peer-reviewed publications, pushing the boundaries of science and engineering. This year's report features new sections spotlighting the expanding roles of Artificial Intelligence and Accelerated Computing. We also introduce an early career section to celebrate the accomplishments of our up-and-coming researchers, whose pioneering work is shaping the future of energy. We hope you enjoy the new insights and discoveries highlighted in these pages.

97 MATHEMATICS AND COMPUTING↗

Holistic Measurement Driven Resilience: Combining Operational Fault and Failure Measurements and Fault Injection for Quantifying Fault Detection, Propagation and Impact. Final report

For HPC systems to date, application resilience to faults and failures has been accomplished by the brute- force method of checkpoint/restart, which allows an application to make forward progress in the face of system and application faults, errors, and failures independent of root cause or end result. It has remained the primary resilience mechanism because we lack a way to identify faults and anticipate consequences early enough to take meaningful mitigating action. However, checkpoint/restart implementations put a tremendous burden on system resources and on the applications themselves and is becoming less feasible at scale. Because we have not yet operated at scales at which checkpoint/restart fails to provide forward progress, despite increasing costs, vendors have had little motivation to provide the instrumentation necessary for early identification of faults and failures. However, as we move from petascale to exascale, component mean time to failure (MTTF) will render the existing techniques ineffectual and/or too expensive. Furthermore, fault recovery mechanisms such as failover and/or error correction introduce performance inconsistency. Instrumentation allowing early indication of problems and tools to enable use of such information by systems, operating systems, and applications offer an alternative, more scalable and less costly solution. In the HMDR project, we built on our experience and expertise developed and accumulated over years of research on design, monitoring, measurement, and assessment of resilient computing systems. Analysis of field data on the current and past generations of extreme-scale systems revealed several challenges that, if not addressed in increasingly larger and more complex systems, may hinder the effectiveness of future exascale computing systems. Specifically, i) file systems and interconnects in current-generation large-scale systems already operate at the margins of resiliency, including consistent performance, and may not scale to larger deployments; ii) automated, software-based failover mechanisms are frequently inadequate and can introduce wider failures, such that failures during recovery may lead to system/application failures, including system-wide outages; and iii) silent data corruption represents a critical fault mode and will require efficient detection mechanisms if next-generation applications are to take full advantage of exascale hardware. To address the above challenges, we assembled a team of world-renowned experts in resilient extreme- scale computing from the University of Illinois (Electrical and Computer Engineering, Computer Science, and NCSA), SNL, LANL, NERSC, and Cray. Our team includes representatives from centers that house many of the largest HPC resources in the world, both today and over the coming years. The team has a unique track record of research in i) system and application failure characterization based on the analysis of field data, ii) data-driven design of fault/error detection mechanisms, and iii) experimental characterization of system/application resiliency. The team includes system owners/operators who provide continuous data collection and access and ensure installation of appropriate analysis tools.

97 MATHEMATICS AND COMPUTING↗

Artificial Intelligence for Multiphysics Nuclear Design Optimization with Additive Manufacturing

The geometric flexibility of additively manufactured metals and ceramics generates a very large and open design space that requires advanced modeling and simulation tools for physics simulations and the rigorous definition of design problems. This effort deploys artificial intelligence (AI) and machine learning (ML) algorithms to understand the design space, evaluate potential designs, and more efficiently generate optimized results. The Transformational Challenge Reactor (TCR) program is leveraging advances in several scientific areas—including materials, manufacturing, sensors and control systems, data analytics, and high-fidelity modeling and simulation—to accelerate the design, manufacturing, qualification, and deployment of advanced nuclear energy systems. Through a manufacturing-informed design approach, the TCR program seeks to integrate digital data for rapid nuclear innovation; accelerate the adoption of advances in manufacturing, materials, and computational sciences for nuclear applications; and dramatically reduce deployment costs and timelines for new nuclear reactor technologies. This report documents efforts under the TCR program to leverage advanced modeling and simulation techniques driven by AI/ML algorithms on high-performance computing (HPC) systems to yield more optimized TCR core designs. A multiphysics ML surrogate model was developed to run on the HPC architectures. The surrogate model is trained on high-fidelity simulation data of coupled neutronics and thermofluidics and is used to quickly evaluate thousands of candidate core designs in parallel, which drives the evolution of the cooling channel shapes to minimize temperature peaking and material stress. Outcomes from these activities provide design information and feedback into the core design efforts.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Multiphase Simulations of the SLS Launch Environment

NASA’s Space Launch System (SLS), which will send astronauts back to the Moon in the next few years, is powered by four RS-25 engines and two RSRMV solid rocket boosters (SRBs). During launch the SLS propulsion system generates intense acoustics and other powerful waves, such as ignition overpressure (IOP) which, if unmitigated, have the potential to damage the vehicle and possibly cause loss of mission or crew. To protect the vehicle from these powerful waves, the SLS launch pad design includes an ignition overpressure/sound suppression (IOP/SS) system which sprays 270,000 gallons per minute of water very close to the SRB and RS-25 nozzles. The SRB and RS-25 engine plumes, and the proximity of the IOP/SS water, create a complex multiphase (gas and liquid) environment during the SLS ignition sequence. The interplay among these systems creates challenges related to water spray into/onto engine nozzles, potential debris transport, and additional transient loads due to strong plume-water interactions - all of which the SLS vehicle must be able to withstand. Prior to the Artemis I launch, the SLS multiphase liftoff environment was largely unknown due to differences from the Space Shuttle and other programs. Some data was available from tests of individual systems, but no integrated testing or analysis was available. Even post-launch analysis of Artemis I cannot provide a full understanding of the complex physics involved due to limited (or obstructed) camera views and instrumentation. Computational fluid dynamics (CFD) is being used to investigate the details of the multiphase environment which could not be measured, help comprehend the data gathered from the launch, and ultimately identify phenomena that are a concern for future flights. Project Details Engineers at NASA’s Marshall Space Flight Center (MSFC) have executed simulations using the Loci/STREAM-Volume of Fluid (VoF) multiphase CFD solver to understand this environment. Initial efforts successfully validated the CFD solver on various tests, giving confidence to simulate the SLS multiphase liftoff environment prior to the Artemis I launch. The CFD simulation of the SLS ignition sequence was conducted in three phases. First the IOP/SS water system was simulated for approximately 6 seconds to reach a quasi-steady state. Next, the RS-25 engine plumes were activated and held at full power for 1 second. Lastly, the SRB booster was activated and the simulation was carried out until just prior to vehicle motion. This simulation process mimics the conditions that exist at launch. Results and Impact The SLS ignition sequence simulation results provide deep understanding of the underlying physics occuring during launch. Observations from the simulation include reduction of water splashing into/onto the engine nozzles, change in angling of the dense water sheets, and the origin of the powerful ignition overpressure (IOP) wave. These observations directly inform the SLS program on subjects including plume-water induced side loads, debris transport, and the acoustic launch environment. Additionally, with post launch comparison of CFD observations to flight data, these tools can be applied to launch vehicles and environments other than SLS with confidence. Why HPC Matters The SLS ignition sequence CFD simulations are conducted on meshes up to hundreds of millions of cells on thousands of processors for weeks at a time. These simulations generate terabytes of data that must also be stored and archived for future use on HPC systems. Simply put, the CFD simulations would not be possible without NASA HPC resources. What’s Next Comparisons between the Artemis I flight data and the CFD simulations will be continued to both improve confidence in the CFD results and provide deeper understanding into the SLS multiphase launch environment. This will be used to provide insight for decision making for the first manned SLS flight, Artemis II. Future simulations will target new configurations of the SLS IOP/SS water required to support the more powerful variants of the SLS vehicle, such as Block 1B. Additionally, this capability provides NASA the ability to investigate launch environments for vehicles other than SLS to support other missions.

Travis Rivord↗

Quantum Computing Strategy 2026

Quantum computing (QC) is a rapidly maturing technology with the potential for revolutionary impacts on stockpile stewardship science and national security. Recent developments in fault-tolerant architectures have compressed vendor roadmaps, and predictions of a production-ready quantum computer by the mid-2030s are becoming increasingly credible. This strategy provides a roadmap for integrating QC into the Advanced Simulation and Computing (ASC) program by investing in four strategic focus areas: 1. Develop Capabilities in Mission-Relevant Quantum Applications: ASC will prioritize developing quantum-ready applications in mission areas that have shown significant promise for quantum advantage, including simulations of materials in extreme environments, nuclear dynamics, solving linear and nonlinear partial differential equations, and uncertainty quantification. These applications directly support stockpile stewardship science and modernization objectives. 2. Conduct R&D in Algorithms, Software, and Hardware: Sustained research into quantum algorithms, robust software tools, and quantum hardware is essential. ASC will develop efficient quantum algorithms; invest in quantum compilers, debuggers, and performance tools; and explore specialized quantum hardware tailored to NNSA’s unique requirements. 3. Engage with Vendors and Partners: Early and active collaboration with commercial quantum hardware vendors and academic partners is critical. Through testbeds, co-design agreements, and quantum demonstration facilities, ASC will influence hardware design, gain early access to emerging technologies, and ensure that quantum platforms evolve to meet mission needs. 4. Build Knowledge, Experience, and Workforce: Expanding and upskilling the quantum-trained workforce is essential to long-term success. This includes hiring, internal training, university outreach, and postdoctoral support to ensure ASC maintains the expertise required to operate, program, and integrate quantum systems as they become available. While quantum computing will never replace classical computing, it has the potential to solve certain problems with speed and accuracy that would be unachievable using any conceivable classical high-performance computing (HPC) system. By investing strategically in QC, ASC will help propel the emergent QC industry, maintain U.S. technological leadership, ensure mission readiness, and position itself to rapidly adopt quantum technologies as they mature.

97 MATHEMATICS AND COMPUTING↗