Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “GPU Failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

SMC 2021 Data Challenge: Analyzing Resource Utilization and User Behavior on Titan Supercomputer

Resource utilization statistics of submitted jobs on a supercomputer can help us understand how users from various scientific domains use HPC platforms and better design a job scheduler. We explore to generate insight regarding workload distribution and usage pattern domains from job scheduler trace, GPU failure information, and project-specific information collected from Titan supercomputer. Furthermore, we want to know how the scheduler performance varies over time and how the users' scheduling behavior changes following a system failure. These observations have the potential to provide valuable insight, which is helpful to prepare for system failures. These practices will help us develop and apply novel machine learning algorithms in understanding system behavior, requirement, and better scheduling of HPC systems. There are two datasets, RUR and GPU: RUR dataset is the job scheduler traces collected from the Titan supercomputer from 01/01/2015 to 07/31/2019 (2015.csv - 2019.csv). These were collected using resource Utilization Report (RUR), a Cray-developed resource-usage data collection and reporting system. It contains the usage information of its critical resources (CPU, Memory, GPU, and I/O) of each running job on Titan during that period (https://ieeexplore.ieee.org/abstract/document/8891001). It includes ProjectAreas as additional information, every job is associated with a project ID. TheProjectAreas.csv dataset provides a mapping of the project ID to its domain science. GPU dataset has information regarding GPU failure on Titan. There have been some hardware-related issues in the GPUs in Titan that caused some GPUs to fail, sometimes irrecoverably during some job runs. This dataset provides information regarding these failures during the execution of the submitted jobs. GPUs on Titan are uniquely identified by a serial number (SN), and they are installed in a location. A GPU can be installed in a location, then removed from that location following a failure, and then re-installed in a different location after fixing the problem. If the failure can't be recovered, the GPU might be removed entirely from Titan. There are two prominent types of failures that resulted in the removal of GPUs from Titan: Double Bit Error (DBE) and Out of the Bus (OTB). The dataset (gc_full.csv) has seven attributes, we provided a short description of these attributes in the ReadMe file. To learn more about this dataset, please refer to the git repository https://github.com/olcf/TitanGPULife and the related publication (https://ieeexplore.ieee.org/abstract/document/9355319).

42 ENGINEERING↗

SMC 2021 : Analyzing Resource Utilization and User Behavior on Titan Supercomputer

Resource utilization statistics of submitted jobs on a supercomputer can help us understand how users from various scientific domains use HPC platforms and better design a job scheduler. We explore to generate insight regarding workload distribution and usage pattern domains from job scheduler trace, GPU failure information, and project-specific information collected from Titan supercomputer. Furthermore, we want to know how the scheduler performance varies over time and how the users' scheduling behavior changes following a system failure. These observations have the potential to provide valuable insight, which is helpful to prepare for system failures. These practices will help us develop and apply novel machine learning algorithms in understanding system behavior, requirement, and better scheduling of HPC systems. There are two datasets, RUR and GPU. RUR: This dataset is the job scheduler traces collected from the Titan supercomputerfrom 01/01/2015 to 07/31/2019 (2015.csv - 2019.csv). These were collected usingResource Utilization Report (RUR), a Cray-developed resource-usage data collectionand reporting system. It contains the usage information of its critical resources (CPU,Memory, GPU, and I/O) of each running job on Titan during that period [2]. ProjectAreas: Every job is associated with a project ID. TheProjectAreas.csvdatasetprovides a mapping of the project ID to its domain science. GPU: There have been some hardware-related issues in the GPUs in Titan that caused some GPUs to fail, sometimes irrecoverably during some job runs. This dataset provides information regarding these failures during the execution of the submitted jobs. GPUs on Titan are uniquely identified by a serial number (SN), and they are installed in a location. A GPU can be installed in a location, then removed from that location following a failure, and then re-installed in a different location after fixing the problem. If the failure can't be recovered, the GPU might be removed entirely from Titan. There are two prominent types of failures that resulted in the removal of GPUs from Titan: Double Bit Error (DBE) and Out of the Bus (OTB). The dataset (gc_full.csv) has the following fields: 1. SN : Serial number of a GPU 2. location : The location where it is installed 3. insert : The time when it was inserted into that location 4. remove : The time when it was removed from that location 5. duration : Amount of time the GPU spent in this location 6. out : If the device was taken out entirely w/o a re-installment into a new location. 7. event : If the GPU was taken out entirely, the reason for its removal. To learn more about this dataset, please refer to the git repositoryhttps://github.com/olcf/TitanGPULifeand the related publication [1]. References [1] George Ostrouchov, Don Maxwell, Rizwan A Ashraf, Christian Engelmann, MallikarjunShankar, and James H Rogers. Gpu lifetimes on titan supercomputer: Survival analysisand reliability. InSC20: International Conference for High Performance Computing,Networking, Storage and Analysis, pages 1-14. IEEE, 2020. [2] Feiyi Wang, Sarp Oral, Satyabrata Sen, and Neena Imam. Learning from five-yearresource-utilization data of titan system. In2019 IEEE International Conference onCluster Computing (CLUSTER), pages 1-6. IEEE, 2019.

42 ENGINEERING↗

Revealing power, energy and thermal dynamics of a 200PF pre-exascale supercomputer

As we approach the exascale computing era, the focused understanding of power consumption and its overall constraint on HPC architectures and applications are becoming increasingly paramount. Summit, located at the Oak Ridge Leadership Computing Facility (OLCF), is one of the fastest and largest pre-exascale platforms in operation today. This paper provides a first-order examination and analysis of power consumption at the component-level, node-level, and system-level, from all 4,626 Summit compute nodes, each with over 100 metrics at 1Hz frequency over the entire year of 2020. We also investigate the power characteristics and energy efficiency of over 840k Summit jobs and 250k GPU failure logs for further operational insights. To the best of our knowledge, this is the first systematic analysis of power data of HPC system at this scale.

Shin, Woong↗

Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis

HPC datacenters offer a backbone to the modern digital society. Increasingly, they run Machine Learning (ML) jobs next to generic, compute-intensive workloads, supporting science, business, and other decision-making processes. However, understanding how ML jobs impact the operation of HPC datacenters, relative to generic jobs, remains desirable but understudied. In this work, we leverage long-term operational data, collected from a national-scale production HPC datacenter, and statistically compare how ML and generic jobs can impact the performance, failures, resource utilization, and energy consumption of HPC datacenters. Our study provides key insights, e.g., ML-related power usage causes GPU nodes to run into temperature limitations, median/mean runtime and failure rates are higher for ML jobs than for generic jobs, both ML and generic jobs exhibit highly variable arrival processes and resource demands, significant amounts of energy are spent on unsuccessfully terminating jobs, and concurrent jobs tend to terminate in the same state. We open-source our cleaned-up data traces on Zenodo (https://doi. org/10.5281/zenodo.13685426), and provide our analysis toolkit as software hosted on GitHub (https://github.com/atlarge-research/2024-icpads-hpc-workload-characterization). This study offers multiple benefits for data center administrators, who can improve operational efficiency, and for researchers, who can further improve system designs, scheduling techniques, etc.

crossanalysis↗

OLCF Summit Supercomputer GPU Snapshots During Double-Bit Errors and Normal Operations

As we move into the exascale era, the power and energy footprints of high-performance computing (HPC) systems have grown significantly larger. Due to the harsh power and thermal conditions the system, components are exposed to extreme operating conditions. Operation of such modern HPC systems requires deep insights into long term system behavior to maintain its efficiency as well as its longevity. To help the HPC community to gain such insights, we provide double-bit errors using system telemetry data and logs collected from the Summit supercomputer, equipped with 27,648 Tesla V100 GPUs with 2nd-generation high-bandwidth memory (HBM2). The dataset relies on Nvidia XID records internally collected by GPU firmware at the time of failure occurrence, on the reboot-time logs of each Summit node, on node-level job scheduler records collected after each job termination, and on a 1Hz data rate from the baseboard management controllers (BMCs) of each Summit compute node using the OpenBMC event subscription protocol. Technical details can be found in the paper Oles et. al “Understanding GPU Memory Corruption at Extreme Scale: The Summit Case Study” ICS’24 (https://doi.org/10.1145/3650200.3656615).

97 MATHEMATICS AND COMPUTING↗

GPU Lifetimes on Titan Supercomputer: Survival Analysis and Reliability

The Cray XK7 Titan was the top supercomputer system in the world for a long time and remained critically important throughout its nearly seven year life. It was an interesting machine from a reliability viewpoint as most of its power came from 18,688 GPUs whose operation was forced to execute three rework cycles, two on the GPU mechanical assembly and one on the GPU circuitboards. We write about the last rework cycle and a reliability analysis of over 100,000 years of GPU lifetimes during Titan’s 6-year-long productive period. Using time between failures analysis and statistical survival analysis techniques, we find that GPU reliability is dependent on heat dissipation to an extent that strongly correlates with detailed nuances of the cooling architecture and job scheduling. We describe the history, data collection, cleaning, and analysis and give recommendations for future supercomputing systems. We make the data and our analysis codes publicly available.

Ostrouchov, George↗

GPU Lifetimes on Titan Supercomputer: Survival Analysis and Reliability

George Ostrouchov, Don Maxwell, Rizwan Ashraf, Mallikarjun Shankar, and James Rogers. 2020. GPU Lifetimes on Titan Supercomputer: Survival Analysis and Reliability. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC '20). Association for Computing Machinery, New York, NY, USA. Data and code for SC20 paper about Titan GPU reliability analysis: https://github.com/olcf/TitanGPULife. Includes R code to generate graphics for paper and additional analyses. See code/README for instructions. Includes original Titan GPU reliability data on over 100,000 collective hours of operation: data/titan.gpu.history.txt - history data, data/titan.service.txt - service nodes for exclusion. Includes output data files produced by code/TitanGPUmodel.Rmd: data/gc_full.csv - cleaned up data (see paper and R code); data/gc_summary_loc.csv - one record per GPU (variables: SN, time, nlife, nloc, last, col, row, cage, slot, node, max_loc_events, time_max_loc, dbe, dbe_loc, otb, otb_loc, out, batch, days, years, dead, dead_otb, dead_dbe) (see paper and R code). Includes .Rmd analysis document as TitanGPUmode.html. Includes Python code to process data/gc_full.csv into graphics from time-between-failure analyses: See code/tbf-analyses/README for instructions.

42 ENGINEERING↗

Standardizing Microprocessor and GPU Radiation Test Approaches

Microprocessor, Graphics Processing Units (GPUs) and DDRx memory devices have emerged as promising next-generation technologies that enables both high performance processing and acceleration of complex algorithms for the latest challenges in human spaceflight, autonomous vehicles and artificial intelligence (AI). The feature sets of these devices offer exponential increases to throughput, calculation capability and system autonomy when compared to legacy flight systems. NASA's Electronic Part and Packaging (NEPP) Program has conducted an investigation into the radiation susceptibility of leading edge devices and process technologies by establishing standardized test approaches. Unlike most discrete devices, these require state of the art test systems to induce specific hardware activity similar to application software, thus allowing the characterization of failure modes within the system. To best characterize the tested part, NEPP eliminates variables that may impact device performance under radiation. Simplification of remaining system-level variables leads to an improved understanding of complex computational devices and their intended applications. The failure modes and error signatures that are recorded during testing are used to determine radiation sensitivity of the semiconductor process and the microcode architecture of the design. This presentation will discuss the test methodology that NASA Electronic Parts and Packaging (NEPP) is working to establish for its microprocessor, GPU and DDRx memory test programs to provide guidance on these devices and their underlying technology, in regards to their potential usage in future space flight systems.

GPU↗

Proton Testing of nVidia GTX 1050 GPU, Part 2

Single-Event Effects (SEE) testing was conducted on the nVidia GTX 1050 Graphics Processor Unit (GPU); herein referred to as device under test (DUT). Testing was conducted at Massachusetts General Hospital's (MGH) Francis H. Burr Proton Therapy Center on April 28th, 2018 using 200-MeV protons. This testing trip was purposed to provide additional radiation susceptibility data from payloads compiled in Q3FY18. While not all radiation-induced errors are critical, the effects on the application need to be considered. More so, failure of the device and an inability to reset itself should be considered detrimental to the application. Radiation effects on electronic components are a significant reliability issue for systems intended for space.

Single-Event Effects (SEE)↗

Understanding GPU Memory Corruption at Extreme Scale: The Summit Case Study

GPU memory corruption and in particular double-bit errors (DBEs) remain one of the least understood aspects of HPC system reliability. Albeit rare, their occurrences always lead to job termination and can potentially cost thousands of node-hours, either from wasted computations or as the overhead from regular checkpointing needed to minimize the losses. As supercomputers and their components simultaneously grow in scale, density, failure rates, and environmental footprint, the efficiency of HPC operations becomes both an imperative and a challenge. We examine DBEs using system telemetry data and logs collected from the Summit supercomputer, equipped with 27,648 Tesla V100 GPUs with 2nd-generation high-bandwidth memory (HBM2). Using exploratory data analysis and statistical learning, we extract several insights about memory reliability in such GPUs. We find that GPUs with prior DBE occurrences are prone to experience them again due to otherwise harmless factors, correlate this phenomenon with GPU placement, and suggest manufacturing variability as a factor. On the general population of GPUs, we link DBEs to short- and long-term high power consumption modes while finding no significant correlation with higher temperatures. We also show that the workload type can be a factor in memory’s propensity to corruption.

Oles, Vlad↗

LaueMatching: an approach for rapid and robust indexing of Laue diffraction patterns

Traditional Laue diffraction pattern indexing often struggles with noisy data, weak signals, peak overlap and missing reflections, particularly from complex or deformed microstructures. Here, we introduce LaueMatching, a high-throughput indexing algorithm designed to overcome these limitations. LaueMatching utilizes a fundamentally different approach based on direct pattern correlation: experimentally pre-processed images are compared against a comprehensive pre-computed library of simulated diffraction patterns corresponding to a dense grid of possible orientations. This approach bypasses the need for explicit peak identification and fitting, steps that are often a failure point for traditional methods. The algorithm rapidly and robustly indexes multiple crystallographic orientations and crystal systems simultaneously, even from challenging patterns. LaueMatching's effectiveness and accuracy have been rigorously tested and validated on diverse experimental (Ni, Al, EuAl 2 O 4 ) and simulated diffraction patterns, demonstrating high-fidelity orientation refinement. Code to implement this approach on both CPU and GPU resources can be downloaded from https://github.com/AdvancedPhotonSource/LaueMatching.

36 MATERIALS SCIENCE↗

Genomic Language model for Annotation of Repetitive Elements (GLARE) v1.0

GLARE (Genomic Language model for Annotation of Repetitive Elements) is a tool that classifies transposable elements (TEs)—the mobile, repetitive DNA sequences that make up large fractions of eukaryotic genomes. GLARE fine-tunes the NTv3-650M genomic language model on a harmonized collection of curated TE sequences from the PanTEon and Repbase reference databases, assigning each input sequence to one of 11 orders and 32 superfamilies in a Wicker-compatible taxonomy. Features. From nucleotide FASTA input, GLARE outputs per-sequence predictions, class summaries, composition figures, and an annotated FASTA. It provides calibrated confidence scores with optional abstention and runs on CPU or GPU. Uses. GLARE serves as a classification component in genome-annotation pipelines, downstream of TE discovery, supporting genome annotation and comparative and evolutionary genomics. Advantages. GLARE is the first repeat-element classifier to leverage a pretrained genomic language model. Combined with multi-database training, this approach outperformed all nine classifiers in the PanTEon benchmark, generalized better to unseen taxonomic clades, and remained robust to sequence orientation—a common failure mode of existing tools.

Bruna, Tomas [Lawrence Berkeley National Laborator↗

Distributed Quantum-Enhanced Optimization: A Topographical Preconditioning Approach for High-Dimensional Search

Optimization problems become fundamentally challenging as the number of variables increases. Because the volume of the search space grows exponentially, classical algorithms frequently fail to locate the global minimum of non-convex functions. While quantum optimization offers a potential alternative, mapping continuous problems onto near-term quantum hardware introduces severe scaling limits and barren plateaus. To bridge this gap, we propose the Distributed Quantum-Enhanced Optimization (D-QEO) framework. Instead of forcing the quantum processor to find the exact minimum, we use it simply as a topographical preconditioner. The QPU maps the landscape to locate the most promising basin of attraction, generating high-quality seed points for a classical GPU-accelerated solver to refine. To make this approach viable for utility-scale problems, we exploit the mathematical structure of separable functions. This allows us to cut a 50-qubit (i.e., $2^{50}$) global search space into independent and manageable sub-spaces using 5-qubit subcircuits. By executing these fragments concurrently with CUDA-Q, we completely bypass the overhead of cross-register entanglement and classical tensor knitting for separable functions. Benchmarks on the 10-dimensional Rastrigin and Ackley functions show that D-QEO prevents the exponential failure rates observed in purely classical algorithms. Furthermore, this quantum warm-start significantly reduces the number of classical BFGS iterations required to converge, providing a highly practical blueprint for utilizing near-term quantum resources in complex global search.

Soos, Dominik [Old Dominion U.]↗

Effect of non-uniform void distributions on the yielding of metals

High-throughput (several thousand) calculations have been carried out to investigate the yield behavior of porous materials with randomly distributed pores, porosity levels over four orders of magnitude and up to a hundred pores per simulation box. To this end, a Galerkin based fast Fourier transform (FFT) formulation was enhanced to deal with high phase contrast materials. In addition, GPU parallelization was employed in solving the governing equation for strain fluctuations using a Krylov iterative solver. Emphasis is laid on the conditions under which percolation of plastically non-deforming zones through the porous network emerge, a regime termed unhomogeneous yielding. By way of contrast, the regime where the plastic strain fluctuations (associated with the heterogeneous void-matrix aggregate) fall below the percolation threshold is defined as homogeneous yielding. Here, we find that nonuniform pore distributions only affect unhomogeneous yielding and have a universal softening effect. The extent of this distribution softening is analyzed as a function of porosity, cell size and number of realizations. Whether the uncovered universal distribution softening has direct implications on failure resistance of porous materials is discussed.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Simulation and Analysis on Reactor Pressure Vessel (RPV) subjected to Pressurized Thermal Shock (PTS) under SBLOCA scenario by using Cardinal to support the fracture mechanics analyses

The structural components that comprise nuclear reactors and their supporting structures are subjected to harsh operating environments that can challenge their integrity, especially after exposure for extended durations or under accident condition. As one of the most significant components of a Reactor, the Reactor Pressure Vessel (RPV) is exposed to an aggressive environment during the operation time (e.g. more than 40 years). Ageing degradation mechanisms (e.g. thermos-fatigue) could grow initial defects up to a critical size, increasing the susceptibility to failure in the RPV. The conventional methods are mostly based on simple crack and structure geometries. Very limited studies consider the real conditions of the RPV subjected to a thermal shock due to a Loss of Coolant Accident (LOCA). During a LOCA event, the most severe conditions take place when the emergency core cooling (ECC) water is injected inside the cold legs filled initially with hotter water and/or steam. The rapid cooling of the down-comer and the internal RPV surface followed probably by re-pressurization of the RPV causes large temperature gradients and variation of pressure which induces thermal-mechanical stresses. In order to develop the model for integrity assessment of a reactor pressure vessel (RPV) subjected to pressurized thermal shock (PTS), a multi-physics simulation, which includes the thermo-hydraulic, thermo-mechanical and fracture mechanics analyses is necessary. The multi-physics simulations are performed using Cardinal, a wrapping of the GPU-oriented spectral element Computational Fluid Dynamics (CFD) code NekRS and other multi-physics sub-modules within the MOOSE framework. Cardinal now fully supports MOOSE stochastic perturbations of NekRS models with varying boundary conditions, initial conditions, material properties, and any other quantity which is defined by a kernel (such as coefficients in a momentum source model). The implementation is designed in a flexible manner so that scalar values are sent from MOOSE into a user scratch space in NekRS, which can then be applied for any purpose within the NekRS case files (both on the host and device). When modeling PTS, several factors can impact the results significantly. In this report, the impacts of the geometry of the model, Reynolds number and buoyancy effect are investigated. Two geometry, i.e., a simplified model and a realistic RPV model, with both laminar and turbulent flow condition are adopted for the PTS simulation with and without buoyancy effect. The purpose of the investigation is to understand the impact of these factors on the prediction of temperature history of RPV. The accurate prediction on the temperature evolution, which will be exported to Grizzly code for further analyses on the progression of aging mechanisms and their effects on the integrity of RPV structures, is very crucial. Based on the understanding of these factors, a more sophisticated model is built to analysis the PTS under SBLOCA scenario. A literature survey is conducted to pick the SBLOCA scenario for the multi-physics simulation. The analysis helps to explain the form and the transformation of the cold plum when the ECC is activated under SBLOCA. This model can be can be applied to study the PTS effect for different RPV configurations. The results can help to assess structural component degradation for advanced reactors.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

PipeSight: A High-Performance Computing Platform for Pipeline Integrity Management

The Phase I feasibility study completed as part of this project has led to a number of innovative technologies being developed and has laid the foundation for a successful Phase II effort to commercialize a platform for managing the integrity of pipelines for the damage mechanisms of the new, hybrid-energy based economy. To ground the development efforts and direction of the project, an extensive market research and customer discovery effort was undertaken early in Phase I. Through this effort, a number of pipeline owners and operators were interviewed, and the following key findings were discovered about the pipeline industry: • Small pipeline operators do not have the central engineering groups necessary to perform their own independent analysis of inspection data, but instead rely on summarized tally sheets provided to them by inspection service providers. • The time it takes to go from an inspection to a completed engineering assessment, even for small segments of pipeline, can take anywhere from 30-120 days. During this delay, critical threats can (and have been known to) cause failures. • Uncertainty is often not accounted for in the assessment of pipeline integrity. The tally sheets provided by third-party service providers are almost always deterministic in nature, identifying threats that present a concern only to the current (not the future) integrity of the pipeline. • It is uncommon to apply the latest technologies to perform advanced assessments of damaged pipelines. There is a desire to use more advanced analysis capabilities to assess threats. Many pipeline operators indicated that they would often excavate a pipeline to perform an inspection and find that the damage was not as bad as they anticipated, thus using limited resources unnecessarily. Companies are not consistent in their use of inspection data to determine corrosion rates, and those that do only calculate deterministic corrosion rates. • The industry has prominently relied on time-based inspections but has recently started to transition to risk-based inspections. However, there appears to be no uniform guidance on how to do so while properly accounting for all sources of uncertainty. • Companies are not storing inspection data in a manner that allows for the ready determination of temporal trends. • Predictive maintenance principles and practices are beginning to be used by early adopters • Some pipelines are being re-purposed to transport different process fluids than they were designed for, e.g., H 2 and CO 2 rich process streams to serve the new hybrid-energy based economy, which are presenting new integrity concerns for the existing pipeline network that crisscrosses the United States. As a result of these discoveries, we were able to target the development efforts in Phase I to best serve the needs of the industry. In Phase I, we developed a way to correlate multiple large-scale scans of the pipeline to determine a probabilistic corrosion rate that accounts for all sources of error and uncertainty in the inspection process. This probabilistic corrosion rate can be used to predict the future thickness distribution of the pipe wall. We demonstrate how this analysis may be performed in an analytical fashion and has been implemented in such a manner that it can be readily distributed using GPU computing through integration of the Kokkos programming model. We also make a very novel extension of the analytical corrosion rate model to Bayesian Networks (an explainable AI technique) that can account for non-parametric distributions of corrosion rates. With the predictions made above for the probabilistic corrosion rate and corresponding future distribution of the pipe wall thickness, we can assess the integrity of the pipeline through the use of a probabilistic engineering assessment. We developed a novel screening data analysis approach that can rapidly identify ‘hotspots’ (local thin areas) where the integrity of the pipeline is a concern. Once more, we implemented this screening approach in C++ to leverage GPU computing via the Kokkos programming model. After the critical hotspots are identified, we developed a program that can automatically generate an advanced finite element model of the damaged regions. Since the number of damaged regions that require advanced analysis can number in the thousands, we integrated an open-source container-native workflow engine for orchestrating parallel jobs on the cloud. Initially, these advanced numerical models were only designed to account for loading due to internal pressure. However, in a slight pivot from the initial Phase I proposal, we developed a complete pipe stress analysis program (called Simflex) which can simulate the complete pipeline and its response to thermal expansion, pressure, thermal bowing, weight, wind, earthquake, support displacement, support friction and external forces. This pipe stress analysis program was written generically, to handle any piping system, but contains the features needed to model long pipelines (i.e., it incorporates a model for soil mechanics and can account for the nonlinear boundary conditions necessary to simulate long underground pipelines). This pipe stress analysis program can simulate any segment of the pipeline (simple or complex) under any set of conditions and loads, to determine the supplemental loads (axial forces and bending moments) at the location of damage. This enables the most accurate state of stress to be accounted for in the pipeline, which can prove critical when evaluating the integrity of a damaged region. In the process of developing the technologies to perform the integrity assessment of the pipeline, we also extended one of the industry standard approaches for performing the assessment of local thin areas that extend more in the circumferential direction than the longitudinal direction of the pipeline. This approach was presented to the API 579-1/AS ME FFS-1 steering committee in November 2021 for consideration in the next edition of the industry standard for Fitness-For-Service (expected to be released in 2023). To help pipeline operators make decisions with the results on any integrity assessment, we developed a new approach to the life-cycle management of pipelines which uses a Bayesian Decision Network. The network is designed to help pipeline operators plan and prioritize inspection activities and ultimately make smarter, more cost-effective decisions. The Bayesian approach accounts for all sources of uncertainty and carries them through to the final optimal decisions, providing a probabilistic framework for optimizing inspection intervals. The proof-of-concept networks developed in the feasibility study are complete, verified, and are focused on a subset of the pipeline. To expand this novel approach to the scale necessary for an entire network of pipelines in Phase II, we will leverage the DOE-funded Bengi solver for industrial-scale decision making with Bayesian Networks [22]. Once implemented, we will be able to provide the pipeline industry with a much-needed tool for optimal inspection planning using truly explainable artificial intelligence (XAI). To handle all of these advanced capabilities into a cloud-based platform, the architecture of the Equity Engineering Cloud (EEC) was extended to include Argo Workflows, a framework capable of distributing and managing a massive number of jobs that consume their own resources, such that thousands of serial finite element simulations can be run in parallel. As part of this substantial undertaking, we also integrated Argo Continuous Delivery (CD) into the EEC, to aid with the rapid prototyping and iterations that will be imperative to the success of the PipeSight platform’s Agile development process in Phase II. As part of the pipe stress analysis program, we also developed a custom visualizer that leverages the DOE-funded VTK visualization library. We added custom contouring capabilities and a means for interacting visually with both the inputs and outputs of the pipe stress analysis program. We also developed routines for automating the post-processing of the finite element simulations to determine if any failure criteria are met and to visualize the deformations, stresses and strains in ParaView using the exodus II file format (a subset of netCDF).

24 POWER TRANSMISSION AND DISTRIBUTION↗

Experiences with SYCL on AMD GPUs with Kokkos

With the recent diversification of the hardware landscape in the high-performance computing (HPC) community, performance-portability solutions are becoming more and more important. One of the most popular choices is Kokkos, which recently became a Linux Foundation project. Most of its development is supported by the US Department of Energy and the French Alternative Energies and Atomic Energy Commission. Kokkos is implemented as a C++ library with multiple backends to support CPUs as well as various GPU architectures. These backends include OpenMP, CUDA, HIP, and also SCYL. This approach enables users to leverage the preferred vendor toolchain for the respective platform (e.g. CUDA, ROCm, OneAPI). The SYCL backend is used to target Intel GPUs, in particular to support the Aurora exascale supercomputer. However, SYCL itself also offers a large degree of portability, and in fact Kokkos’ CI for SYCL has been running on NVIDIA hardware due to a lack of access to Intel GPUs. In this report, we describe our experience with using Kokkos SYCL backend on AMD GPUs targeting the Frontier supercomputer at Oak Ridge National Laboratory. The two major SYCL implementations are DPC++ and AdaptiveCpp. While the Kokkos SYCL backend has been implemented using the former, the latter was the first implementation to target AMD GPUs. We will discuss the experience with both of these SYCL implementations in terms of functionality and performance. Using Kokkos to evaluate SYCL toolchains has a number of benefits. Kokkos’ use of SYCL is fairly complex, exercising features such as graphs, relocatable device functions, atomics – including for non-arithmetic types, as well as pinned and page migratable memory allocations. Kokkos also needs to implement capabilities such as Kokkos’ hierarchical parallelism that are not a straight-forward mapping to SYCL capabilities. Furthermore, a large number of libraries and applications that represent diverse use cases are implemented in Kokkos, providing readily available test cases for a toolchain evaluation. Preliminary results show that support for AMD GPUs in DPC++ is much less mature than for NVIDIA GPUs or Intel GPUs. While the situation has improved significantly over the last year, we still encounter many runtime failures, dispatching problems, and code generation issues. With AdaptiveCpp the challenges arise even earlier in the evaluation process. Since Kokkos’ SYCL implementation is largely focused on supporting Intel GPUs, we opted to leverage SYCL extensions which are available in DPC++ but not in AdaptiveCpp. Furthermore, AdaptiveCpp appears to be less conformant with the SYCL2020 standard which Kokkos relies on. In some cases, we are able to work around the lack of feature support, in other cases we have to disable certain Kokkos capabilities to evaluate the toolchain. Our evaluation will leverage Kokkos’ unit tests to establish basic functionality and feature completeness. We then use simple benchmarks for components of a CG implementation as a measure of usability and performance of the SYCL toolchains.

97 MATHEMATICS AND COMPUTING↗