Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “HPC Utilization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Grid-based minimization at scale: Feldman-Cousins corrections for light sterile neutrino search

High Energy Physics (HEP) experiments generally employ sophisticated statistical methods to present results in searches of new physics. In the problem of searching for sterile neutrinos, likelihood ratio tests are applied to short-baseline neutrino oscillation experiments to construct confidence intervals for the parameters of interest. The test statistics of the form Δχ2 is often used to form the confidence intervals, however, this approach can lead to statistical inaccuracies due to the small signal rate in the region-of-interest. In this paper, we present a computational model for the computationally expensive Feldman-Cousins corrections to construct a statistically accurate confidence interval for neutrino oscillation analysis. The program performs a grid-based minimization over oscillation parameters and is written in C++. Our algorithms make use of vectorization through Eigen3, yielding a single-core speed-up of 350 compared to the original implementation, and achieve MPI data parallelism by employing DIY. We demonstrate the strong scaling of the application at High-Performance Computing (HPC) sites. We utilize HDF5 along with HighFive to write the results of the calculation to file.

Wospakrik, Marianette↗

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.

crossanalysis↗

Opportunities for enhancing MLCommons efforts while leveraging insights from educational MLCommons earthquake benchmarks efforts

MLCommons is an effort to develop and improve the artificial intelligence (AI) ecosystem through benchmarks, public data sets, and research. It consists of members from start-ups, leading companies, academics, and non-profits from around the world. The goal is to make machine learning better for everyone. In order to increase participation by others, educational institutions provide valuable opportunities for engagement. In this article, we identify numerous insights obtained from different viewpoints as part of efforts to utilize high-performance computing (HPC) big data systems in existing education while developing and conducting science benchmarks for earthquake prediction. As this activity was conducted across multiple educational efforts, we project if and how it is possible to make such efforts available on a wider scale. This includes the integration of sophisticated benchmarks into courses and research activities at universities, exposing the students and researchers to topics that are otherwise typically not sufficiently covered in current course curricula as we witnessed from our practical experience across multiple organizations. As such, we have outlined the many lessons we learned throughout these efforts, culminating in the need for benchmark carpentry for scientists using advanced computational resources. The article also presents the analysis of an earthquake prediction code benchmark while focusing on the accuracy of the results and not only on the runtime; notedly, this benchmark was created as a result of our lessons learned. Energy traces were produced throughout these benchmarks, which are vital to analyzing the power expenditure within HPC environments. Additionally, one of the insights is that in the short time of the project with limited student availability, the activity was only possible by utilizing a benchmark runtime pipeline while developing and using software to generate jobs from the permutation of hyperparameters automatically. It integrates a templated job management framework for executing tasks and experiments based on hyperparameters while leveraging hybrid compute resources available at different institutions. The software is part of a collection called cloudmesh with its newly developed components, cloudmesh-ee (experiment executor) and cloudmesh-cc (compute coordinator).

58 GEOSCIENCES↗

How Cloud is Accelerating Research at NREL

This presentation coincides with AWS's announcement of their new Parallel Computing Service (PCS) which allows for easy creation of HPC-style clusters in their AWS cloud computing platform. I helped them beta test this service before it was made generally available in August. AWS asked if we would be interested in discussing our experience with the PCS service, and our experience with HPC workloads in the cloud in general, so this slideshow discusses a brief history of scientific computing at NREL and shares a bit of our experiences and approach to utilizing cloud services for HPC-style workloads.

97 MATHEMATICS AND COMPUTING↗

High-Performance Transmission and Distribution Co-simulation with 10,000+ Inverter-Based Resources

The inverter-based resource (IBR) has become avery important component in the distribution system. The impacts on system transient stability introduced by high IBR penetration are not fully addressed because of the lack of high-fidelity models. The aggregate IBR model at the transmission level cannot precisely reproduce the dynamics of distributed IBR at the distribution system because of the oversimplification. In this paper, we will develop a high-penetration fully-connected transmission and distribution (T&D) co-simulation platform that supports the simulation of 10,000+ dispersed IBR models. The interfacing and iterative initialization techniques for the co-simulation have been implemented to maintain stable operation and simulation of large-multitude of IBR models. The phasor-domain IBR models with grid-forming (GFM) and grid-following (GFL) control are implemented in the distribution systems simulators. The developed platform is tested on high-performance computing (HPC) resources and can be utilized to explore the hierarchical control strategies of IBRs for the large-scale T&D hybrid system.

Liu, Yuan↗

Scalable training of graph convolutional neural networks for fast and accurate predictions of HOMO-LUMO gap in molecules

Abstract Graph Convolutional Neural Network (GCNN) is a popular class of deep learning (DL) models in material science to predict material properties from the graph representation of molecular structures. Training an accurate and comprehensive GCNN surrogate for molecular design requires large-scale graph datasets and is usually a time-consuming process. Recent advances in GPUs and distributed computing open a path to reduce the computational cost for GCNN training effectively. However, efficient utilization of high performance computing (HPC) resources for training requires simultaneously optimizing large-scale data management and scalable stochastic batched optimization techniques. In this work, we focus on building GCNN models on HPC systems to predict material properties of millions of molecules. We use HydraGNN, our in-house library for large-scale GCNN training, leveraging distributed data parallelism in PyTorch. We use ADIOS, a high-performance data management framework for efficient storage and reading of large molecular graph data. We perform parallel training on two open-source large-scale graph datasets to build a GCNN predictor for an important quantum property known as the HOMO-LUMO gap. We measure the scalability, accuracy, and convergence of our approach on two DOE supercomputers: the Summit supercomputer at the Oak Ridge Leadership Computing Facility (OLCF) and the Perlmutter system at the National Energy Research Scientific Computing Center (NERSC). We present our experimental results with HydraGNN showing (i) reduction of data loading time up to 4.2 times compared with a conventional method and (ii) linear scaling performance for training up to 1024 GPUs on both Summit and Perlmutter.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Developing a Vorticity-Velocity-Based Off-Body Solver to Perform Multifidelity Simulations of Wind Farms

Wind power has become a key player in satisfying the global energy needs. With increased market penetration, unanticipated unsteady loading induced failures, installation related reductions in power generation, and significant maintenance costs have underscored the need to predict the unsteady fluid-structure interactions related to turbine layout and off-design wind conditions. Contemporary turbine design tools are incapable of accounting for such loadings. As a result, researchers have started utilizing high-Performance-Computing (HPC) based Computational Fluid Dynamics (CFD) solvers, such as the U.S. Department of Energy sponsored ExaWind software package, to investigate these phenomena. Unfortunately, such HPC tools are computationally expensive for routine industrial use, often because of the sheer number of cells required to resolve the wake flowfield. This paper describes a preliminary effort to address this issue by developing a vorticity-velocity based CFD off-body solver, VorTran-M2-AMReX, that integrates directly with DOE's ExaWind wind turbine analysis system to perform accurate and reliable simulations of wind turbine/farm at a lower computational cost than ExaWind alone. This article summarizes work undertaken to date concerning the assembly of the proposed analysis tool, and provides preliminary validation and verification of the VorTran-M2-AMReX off-body solver.

adaptive mesh refinement↗

Three practical workflow schedulers for easy maximum parallelism

Runtime scheduling and workflow systems are an increasingly popular algorithmic component in HPC because they allow full system utilization with relaxed synchronization requirements. There are so many special-purpose tools for task scheduling, one might wonder why more are needed. Use cases seen on the Summit supercomputer needed better integration with MPI and greater flexibility in job launch configurations. Preparation, execution, and analysis of computational chemistry simulations at the scale of tens of thousands of processors revealed three distinct workflow patterns. A separate job scheduler was implemented for each one using extremely simple and robust designs: file-based, task-list based, and bulk-synchronous. Comparing to existing methods shows unique benefits of this work, including simplicity of design, suitability for HPC centers, short startup time, and well-understood per-task overhead. All three new tools have been shown to scale to full utilization of Summit, and have been made publicly available with tests and documentation. This work presents a complete characterization of the minimum effective task granularity for efficient scheduler usage scenarios. Here, these schedulers have the same bottlenecks, and hence similar task granularities as those reported for existing tools following comparable paradigms.

97 MATHEMATICS AND COMPUTING↗

Evaluating the potential of disaggregated memory systems for HPC applications

Summary Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such systems to improve overall system memory utilization, but performance can vary across workloads. High‐performance computing (HPC) is crucial in scientific and engineering applications, where HPC machines also face the issue of underutilized memory. As a result, improving system memory utilization while understanding workload performance is essential for HPC operators. Therefore, learning the potential of a disaggregated memory system before deployment is a critical step. This paper proposes a methodology for exploring the design space of a disaggregated memory system. It incorporates key metrics that affect performance on disaggregated memory systems: memory capacity, local and remote memory access ratio, injection bandwidth, and bisection bandwidth, providing an intuitive approach to guide machine configurations based on technology trends and workload characteristics. We apply our methodology to analyze thirteen diverse workloads, including AI training, data analysis, genomics, protein, fusion, atomic nuclei, and traditional HPC bookends. Our methodology demonstrates the ability to comprehend the potential and pitfalls of a disaggregated memory system and provides motivation for machine configurations. Our results show that eleven of our thirteen applications can leverage injection bandwidth disaggregated memory without affecting performance, while one pays a rack bisection bandwidth penalty and two pay the system‐wide bisection bandwidth penalty. In addition, we also show that intra‐rack memory disaggregation would meet the application's memory requirement and provide enough remote memory bandwidth.

Ding, Nan↗

EVs@Scale Next-Gen Profiles - Fleet Utilization 2023

As U.S. fleet operators begin transitioning to electric vehicles (EVs), critical questions arise regarding how to manage this shift without disrupting fleet operations or placing undue stress on the electric grid. A major challenge for fleets is maintaining effective operational schedules while accommodating charging requirements, particularly with high-power charging (HPC) infrastructure, which presents grid stability concerns for utilities. Proposed solutions such as charging substations, megawatt charging systems (MCS), and smart charge management systems (SCMS) offer potential pathways forward, but their effectiveness depends on alignment with real-world fleet behavior and operational constraints. This report investigates the charging and utilization behavior of EV and EVSE fleets actively employing HPC technologies by conducting detailed case study analyses based on telematics data. A suite of predefined metrics—covering charging, routing, and other operational behaviors—is developed to evaluate the impact of fleet activities on grid infrastructure and identify opportunities for optimization. Results highlight variations in charging behavior across fleets, such as weekday versus weekend usage, diurnal charging trends, and the role of operational predictability in enabling SCMS effectiveness. While SCMS can help lower costs and improve energy efficiency for fleets with stable schedules, they may be insufficient for fleets with highly variable or long-haul operations, which may require more robust solutions like MCS. Visualization of aggregated hourly energy metrics reveals that while fleet behaviors are diverse, there are common temporal patterns that could inform infrastructure planning and energy management. These insights emphasize the need for fleet-specific charging strategies that minimize grid impact while supporting reliable fleet operations. Additionally, the report underscores the broader economic stakes of electrification, particularly in high-value markets such as freight, where misaligned transitions could stall EV adoption. By examining current EV and EVSE fleet deployments using predetermined standardized metrics, this study offers a foundation for developing technologies and operational frameworks that support scalable, grid-compatible electrification across a variety of fleet types while establishing a baseline understanding of operational behaviors. In doing so, we aim to ensure that future charging solutions reflect actual fleet needs and grid constraints—an essential step toward maintaining operational continuity and achieving a successful transition to electric fleet operations.

Charging↗

A parallel and performance portable implementation of a full-field crystal plasticity model

We have developed a parallel implementation of an Elasto-Viscoplastic Fast Fourier Transform-based (EVPFFT) micromechanical solver to enable computationally efficient crystal plasticity modeling for polycrystalline materials. Our primary focus lies in achieving performance portability, allowing a single EVPFFT implementation to run optimally on various homogeneous architectures, including multi-core Central Processing Units (CPUs), as well as on heterogeneous computer architectures comprising multi-core CPUs and Graphics Processing Units (GPUs) from different vendors. To accomplish this goal, we have leveraged MATAR, a C++ software library that simplifies the creation and utilization of multidimensional dense or sparse matrix and array data structures. These data structures are designed to be portable across diverse architectures through the use of Kokkos, a performance-portable library. Additionally, we have employed the Message Passing Interface (MPI) to efficiently distribute the computational workload among processors. The heFFTe (Highly Efficient FFT for Exascale) library is used to facilitate the performance portability of the fast Fourier transforms (FFTs) computation. The computational performance of EVPFFT is evaluated and presented in terms of parallel scalability and simulation runtime on different high-performance computing (HPC) architectures. As a result, the utility of the developed framework to efficiently simulate the micro-mechanical fields in polycrystalline microstructures in engineering applications is discussed.

36 MATERIALS SCIENCE↗

Development of a DC Distribution Testbed for High-Power EV Charging

This paper explores the design and implementation of a power hardware-in-the-loop (P-HIL) setup for DC distribution infrastructure integrated with high-power charging (HPC) of electric vehicles (EVs), DC loads, and sources. The utilization of DC distribution holds significant potential for enhancing the operation of a HPC station architecture. However, there are challenges establishing a DC charging hub including interoperability, commoditization, distributed energy resource integration, stability, DC protection, and lack of common system level controllers. To address these challenges, a testing setup is required that accommodates commercial off-the-shelf (COTS) products to evaluate different use cases at rated power and voltage levels. The developed P-HIL setup features a dedicated DC charging hub, DC-coupled chargers, DC loads/sources, DC protection, and a communication architecture. The integrated P-HIL system provides a versatile testing environment to address technology and interoperability gaps and implements a smart energy management system (SEMS). This platform enables comprehensive and robust testing of COTS devices, charger prototypes, SEMS controllers and protection schemes, which together will accelerate transition to EVs at scale. The setup is tested for various use-cases at full-scale, integrating 950 V DC bus voltage, 660 kW grid-tied inverter, 150 kW COTS charger, and 100 kW energy storage system within an open-source SEMS platform.

ADVANCED PROPULSION SYSTEMS,POWER TRANSMISSION AND↗

Enabling power measurement and control on Astra: The first petascale Arm supercomputer

Astra, deployed in 2018, was the first petascale supercomputer to utilize processors based on the ARM instruction set. The system was also the first under Sandia's Vanguard program which seeks to provide an evaluation vehicle for novel technologies that with refinement could be utilized in demanding, large-scale HPC environments. In addition to ARM, several other important first-of-a-kind developments were used in the machine, including new approaches to cooling the datacenter and machine. Here we document our experiences building a power measurement and control infrastructure for Astra. While this is often beyond the control of users today, the accurate measurement, cataloging, and evaluation of power, as our experiences show, is critical to the successful deployment of a large-scale platform. While such systems exist in part for other architectures, Astra required new development to support the novel Marvell ThunderX2 processor used in compute nodes. In addition to documenting the measurement of power during system bring up and for subsequent on-going routine use, we present results associated with controlling the power usage of the processor, an area which is becoming of progressively greater interest as data centers and supercomputing sites look to improve compute/energy efficiency and find additional sources for full system optimization.

97 MATHEMATICS AND COMPUTING↗

Techno-Economic Assessment of Data Center Load Demand Powered by Small Modular Reactors and Distributed Energy Resources

The rapid increase in data center energy demand, driven by AI and large-scale data processing, poses significant challenges to global energy infrastructure. Data centers require substantial and reliable energy for continuous operations and high-performance computing. Current electrical grids face issues such as transmission bottlenecks and aging infrastructure, making it difficult to meet these demands. Integrating inverter-based-resources (IBRs) like solar and wind presents both opportunities and challenges due to their intermittent nature. Small Modular Reactors (SMRs) offer a promising solution with their enhanced safety, modularity, reliability, and scalability, providing consistent base load power ideal for data center operations. This study presents a comprehensive techno-economic assessment of powering data center load demand using a combination of SMRs and IBRs with grid-connected and islanded mode. This study utilized Idaho National Laboratory’s (INL) HPC data center hourly load profiles and Xendee microgrid optimization platform to conduct the analysis. In this configuration, SMRs serves as the primary base load power source, consistently providing a steady supply of electricity necessary to meet the minimum load demand of the data center with support from the IBRs. Key performance indicators such as Levelized Cost of Electricity (LCOE), Net Present Value (NPV) has been calculated to assess the economic feasibility. The findings from this research will underscore the strategic benefits of integrating SMR plant with DERs – particularly for critical infrastructure load such as data centers.

14 - SOLAR ENERGY↗

Quantifying the Impact of Advanced Web Platforms on High Performance Computing Usage

The deployment of Science Gateways for High Performance Computing (HPC) systems can alter long-accepted usage patterns on supercomputing systems in positive ways as an ever-increasing number of users migrate their workflows to HPC systems. Idaho National Laboratory (INL) has deployed two separate advanced web platforms, Open OnDemand and NICE DCV, for integration with HPC resources to improve web accessibility for HPC users. Researchers conducted a multi-year study on how HPC usage pat- terns changed in the presence of these platforms. This work reports the results of that study and quantifies the observed impacts, including adoption by visualization and Jupyter Notebook/Lab users, decreased job submission friction, rapid uptake of HPC by Windows users, and increased overall system utilization. The most significant impacts were observed from the deployment of Open OnDemand, and this work also identifies some best practices for Open OnDemand deployment for HPC datacenters.

97 MATHEMATICS AND COMPUTING↗

Scalable Deep-Learning-Accelerated Topology Optimization for Additively Manufactured Materials

Topology optimization (TO) is a popular and powerful computational approach for designing novel structures, materials, and devices. Two computational challenges have limited the applicability of TO to a variety of industrial applications. First, a TO problem often involves a large number of design variables to guarantee sufficient expressive power. Second, many TO problems require a large number of expensive physical model simulations, and those simulations cannot be parallelized. To address these issues, we propose a general scalable deep-learning (DL) based TO framework, referred to as SDL-TO, which utilizes parallel schemes in high performance computing (HPC) to accelerate the TO process for designing additively manufactured (AM) materials. Unlike the existing studies of DL for TO, our framework accelerates TO by learning the iterative history data and simultaneously training on the mapping between the given design and its gradient. The surrogate gradient is learned by utilizing parallel computing on multiple CPUs incorporated with a distributed DL training on multiple GPUs. The learned TO gradient enables a fast online update scheme instead of an expensive update based on the physical simulator or solver. Using a local sampling strategy, we achieve to reduce the intrinsic high dimensionality of the design space and improve the training accuracy and the scalability of the SDL-TO framework. The method is demonstrated by benchmark examples and AM materials design for heat conduction. The proposed SDL-TO framework shows competitive performance compared to the baseline methods but significantly reduces the computational cost by a speed up of around 8.6x over the standard TO implementation.

Bi, Sirui↗

An Efficient Storage-Driven Machine Learning Model for Performance in the Era of Multimodal Scientific Data

Scientific workflows are increasingly relying on machine learning (ML), simulation, and hybrid techniques to predict, understand, and optimize the behavior of complex experiments. High-performance computing has greatly improved researchers’ ability to acquire diverse data modalities in these workflows. Recent studies suggest that the performance of machine learning models can be improved by integrating data from various sources. Unfortunately, these workloads pose unprecedent pressure on the network storage to meet the demands associated with accessing these multimodal data. To mitigate the impact of intensive IO, we propose a solution that utilizes a multi-tier High-Performance Computing (HPC) distributed storage and data processing framework, placing computation where the data resides for better performance. By adopting this project, the scientific community will gain new opportunities to explore multimodal storage-driven possibilities, integrating multiple scientific data sources with advanced streaming frameworks. Additionally, our framework effectively utilizes computing resources and bridges the gaps identified by HPC experts. Our proposed approach tackles scalability and persistence challenges by leveraging native persistency, which has posed difficulties in traditional approaches. Furthermore, we seek to enhance fault-tolerance and load-balance of computations by leveraging real-time streaming in diverse scientific computing environments, thereby propelling advanced scientific computing research into the next generation.

97 MATHEMATICS AND COMPUTING↗