Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “High performance Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Creating Continuous Integration Infrastructure for Software Development on U.S. Department of Energy High-Performance Computing Systems

The Exascale Computing Project (ECP) software deployment effort developed and advanced DevOps capabilities. One goal was to enable robust continuous integration (CI) workflows that span the protected high performance computing (HPC) environments found within many of the Department of Energy’s (DOE) national laboratories. This article highlights several challenges encountered with enabling automation, such as charging models for CI jobs, and meeting individualized security requirements that revolve around strongly associating running code with a human identity. Here, it also describes how the Jacamar CI tool evolved to meet latter requirements and became a key aspect of the solutions currently offered. Derived from this experience, we offer a conceptual framework for understanding current and future CI challenges at DOE facilities and offer suggestions for long-term solutions.

97 MATHEMATICS AND COMPUTING↗

Software engineering to sustain a high-performance computing scientific application: QMCPACK

We provide an overview of the software engineering efforts and their impact in QMCPACK, a production-level ab-initio Quantum MonteCarlo open-source code targeting high-performance computing (HPC) systems. Aspects included are: (i) strategic expansion ofcontinuous integration (CI) targeting CPU, using GitHub Actions runners, and graphics processing units (GPU) in pre-exascalesystems, using self-hosted hardware; (ii) incremental reduction of memory leaks using sanitizers, (iii) incorporation of Dockercontainers for CI and reproducibility, and (iv) refactoring efforts to improve maintainability, testing coverage, and memory lifetime management. We quantify the value of these improvements by providing metrics to illustrate the shift towards a predictive, rather than reactive, sustainable maintenance approach. Our goal, in documenting the impact of these efforts on QMCPACK, is to contribute to the body of knowledge on the importance of research software engineering (RSE) for the sustainability of community HPC codes and scientific discovery at scale.

Godoy, William↗

Ground Motion Models (GMMs) Improvements Using Earthquake Simulations on High Performance Computers

A computationally efficient simulation platform was developed that can provide representative synthetic ground motions from crustal earthquakes in the Stable Continental Regions of Central and Eastern US (CEUS), using 3D modeling and high-performance computing. The main objective was to use synthetic ground motion to provide constrains to refinements of exiting ergodic Ground Motion Models (GMMs), for large magnitude earthquakes and near-fault distances, for which these models are less reliable. Physics-based broadband (0-5Hz) ground motion simulations were used to estimate the near-fault ground motion amplitudes and within event and between-event variabilities associated with fault rupture characteristics. As part of a strategy for selecting a reginal velocity model and validation of developed rupture modeling technique, ground motions from the moment magnitude Mw5.0 November 7, 2016, Cushing Oklahoma, and Mw5.8 September 3, 2016, Pawnee Oklahoma earthquakes were simulated. In our simulations we used a 3D regional velocity model that was based on Saikia’s 1D velocity model. Saikia’s model demonstrated better performance in modelling high frequency regional wave propagation for CEUS region. The proposed 3D model includes lateral variations added to the 1D background model using the stochastic scheme of Pitarka and Mellors. Comparisons of the simulations with recordings of both earthquakes demonstrated the reliability of our deterministic simulation approach while emphasizing the importance of including small-scale variability in the regional velocity model needed to reproduce the observed high-frequency wave scattering effects. As part of validation analysis, comparisons with different GMMs for a Mw6.5 earthquake in the CESUS region resulted in a very good match between the simulated and empirical ground motion models. Initial investigations of within-event and between-event ground motion variabilities for Mw6.5 scenario earthquakes on a strike-slip fault, suggest that they are strongly related to spatial slip and slip rate variations, average rupture velocity, rupture area and rupture initiation location. For certain scenarios we found that the ground motion variability observed at near-fault distances (< 5 km) also persists at longer distances. Regardless of the rupture scenario, the simulated ground motion tends to fully saturate at short distances and for all periods. The near-fault saturation has to do with the attenuation of waves propagating along the fault and local rupture radiation pattern that also contribute to stronger ground motion variation at such distances. Analysis of effects of rupture initiation location suggest that the peak ground motion (PGV) and spectral acceleration (SA) can be quite variable due to rupture directivity effects. Such effects are stronger at periods longer than 1s. The effect of the 1D velocity models and surface topography on simulated ground motion were investigated by comparing three component synthetic seismograms computed at selected sites. Effect of surface topography was considered using the ratio between spectral accelerations simulated for two 1D models with flat surface topography and realistic model with surface topography. Overall, the topography slightly amplifies (by ~30%) the ground motion amplitude in the frequency range 1-3Hz. The effect of topography is more visible in the surface and coda waves portion of the seismograms.

58 GEOSCIENCES↗

High-Performance Computing Based EMT Simulation of Large PV or Hybrid PV Plants

Faults in the transmission grid have led to reduced power generation from power electronics resources that are typically not connected to the faulted transmission line. In many of the cases, partial loss of power is observed within the power electronics resources like large photovoltaic (PV) power plants. This phenomena is not captured in existing simulation models and/or simulators. High-fidelity switched system electromagnetic transient (EMT) dynamic models of PV power plants can improve the fidelity of models available for accurate analysis of the impact on PV plants during simulation of faults. However, these models are extremely computationally expensive and take a long time to simulate. Long simulation times limit the ability to use these models as larger regions are studied in EMT simulations with more power electronics resources. In this paper, numerical simulation algorithms are combined with high-performance computing techniques and applied to the high-fidelity switched system EMT model of PV plants. Using these techniques, a speed-up of up to 58x is obtained, while preserving the accuracy of the simulation at greater than 98%.

Debnath, Suman↗

A Partitioned -Task Parallel Implementation of the NASA Multiscale Analysis Tool for High Performance Computing

The NASA Multiscale Analysis Tool (NASMAT) is a “plug and play” software package that allows users to conduct massively multiscale modeling of hierarchical and nonlinear materials. This work extends the scalability and improves the High Performance Computing friendliness of NASMAT by adopting a Partitioned Task-Parallel approach. Interoperability of NASMAT with external software is enhanced through preCICE, a open source library for multiphysics coupling in a partitioned manner. Enhancement through preCICE allows for easy integration of NASMAT to other macro solvers and dissociates the parallelization strategy adopted within NASMAT from the macro solver. The task-parallel framework based on Master-Worker approach is implemented as the parallelization scheme. The scheme accounts for hierarchy of multiple scales (task-dependence) and heterogeneous nature (dynamic load balancing) of computations. The applicability and scalability of the framework will be evaluated by analyzing large scale engineering problems through massively multiscale methods.

NASMAT↗

Signal Processing Based Method for Real-Time Anomaly Detection in High-Performance Computing

Performance anomalies can manifest as irregular execution times or abnormal execution events for many reasons, including network congestion and resource contention. Detecting such anomalies in real-time by analyzing the details of performance traces at scale is impractical due to the sheer volume of data High-Performance Computing (HPC) applications produce. In this paper, we propose formulating HPC performance anomaly detection as a signal-processing problem where anomalies can be treated as noise. We evaluate our proposed method in comparison with two other commonly used anomaly detection techniques of varying complexity based on their detection accuracy and scalability. Since real-time in-situ anomaly detection at a large scale requires lightweight methods that can handle a large volume of streaming data, we find that our proposed method provides the best trade-off. We then implement the proposed method in Chimbuko, the first online, distributed, and scalable workflow-level performance trace analysis framework. We compare our proposed signal-based anomaly detection algorithm with two other methods using a function of their accuracy, F1 score, and detection overhead. Our experiments demonstrate that our proposed approach achieves a 99% improvement for the benchmark datasets and a 93% improvement with Chimbuko traces.

99 GENERAL AND MISCELLANEOUS↗

Performance Analysis of an Optimization Algorithm for Metamaterial Design on the Integrated High-Performance Computing and Quantum Systems

Optimizing metamaterials with complex geometries is a big challenge. Although an active learning algorithm, combining machine learning (ML), quantum computing, and optical simulation, has emerged as an efficient optimization tool, it still faces difficulties in optimizing complex structures that have potentially high performance. In this work, we comprehensively analyze the performance of an optimization algorithm for metamaterial design on the integrated HPC and quantum systems. We demonstrate significant time advantages through message-passing interface (MPI) parallelization on the high-performance computing (HPC) system showing approximately 54% faster ML tasks and 67 times faster optical simulation against serial workloads. Furthermore, we analyze the performance of a quantum algorithm designed for optimization, which runs with various quantum simulators on a local computer or HPC-quantum system. Results showcase ~24 times speedup when executing the optimization algorithm on the HPC-quantum hybrid system. This study paves a way to optimize complex metamaterials using the integrated HPC-quantum system.

Kim, Seongmin↗

The Applicability of Unit Systems to High-Performance Computing Applications

Dimensional analysis is a key technique used to verify the soundness of scientific models. Most experts agree that engineering and scientific software would be made more reliable by integrating dimensional analysis in their type system. We explored how High Performance Computing (HPC) applications could integrate compile-time dimensional analysis. We started by investigating various implementation of unit systems for C++. Eventually, selecting the latest (and most advanced) one to apply to our test codes. We worked with code of increasing complexity, from a projectile trajectory calculation to the proxy-application Lulesh. This included our code, Springs-3D, which focuses on demonstrating language features while performing simple physic computations. Finally, our main contribution is a source-code analysis which extracts constraints on the dimension of all variables, functions, and constants in an application. This resulting system of equations is solved using the dimensions of a few of these objects. This analysis has the potential to greatly reduce the time spent performing dimensional analysis when refactoring application to use a representation of units.

97 MATHEMATICS AND COMPUTING↗

Future Generation High Performance Computing Center (FG-HPCC): RFI Technical Considerations

Lawrence Livermore National Security, LLC (LLNS) is interested in receiving information about technologies that could be available in the 2029-2030 timeframe that may serve to enable the vision for a Future Generation High Performance Computing (HPC) Center (FG-HPCC) described in this document. The future HPC Center vision has been conceived to meet the future mission needs of the Advanced Simulation and Computing (ASC) Program within the National Nuclear Security Administration (NNSA). LLNS envisions a center composed not of many independent clusters, but of heterogeneous elements accessible to users as a single system. The capabilities will be integrated to create a scalable, flexible, yet tightly coupled computing center capable of integrated HPC, AI, and cloud-like workloads.

97 MATHEMATICS AND COMPUTING↗

A High Performance Computing Approach to Tree Cover Delineation in 1-m NAIP Imagery Using a Probabilistic Learning Framework

Tree cover delineation is a useful instrument in deriving Above Ground Biomass (AGB) density estimates from Very High Resolution (VHR) airborne imagery data. Numerous algorithms have been designed to address this problem, but most of them do not scale to these datasets, which are of the order of terabytes. In this paper, we present a semi-automated probabilistic framework for the segmentation and classification of 1-m National Agriculture Imagery Program (NAIP) for tree-cover delineation for the whole of Continental United States, using a High Performance Computing Architecture. Classification is performed using a multi-layer Feedforward Backpropagation Neural Network and segmentation is performed using a Statistical Region Merging algorithm. The results from the classification and segmentation algorithms are then consolidated into a structured prediction framework using a discriminative undirected probabilistic graphical model based on Conditional Random Field, which helps in capturing the higher order contextual dependencies between neighboring pixels. Once the final probability maps are generated, the framework is updated and re-trained by relabeling misclassified image patches. This leads to a significant improvement in the true positive rates and reduction in false positive rates. The tree cover maps were generated for the whole state of California, spanning a total of 11,095 NAIP tiles covering a total geographical area of 163,696 sq. miles. The framework produced true positive rates of around 88% for fragmented forests and 74% for urban tree cover areas, with false positive rates lower than 2% for both landscapes. Comparative studies with the National Land Cover Data (NLCD) algorithm and the LiDAR canopy height model (CHM) showed the effectiveness of our framework for generating accurate high-resolution tree-cover maps.

Segments↗

A High-Performance Computing GNSS-aware Path Planning Algorithm for Safe Urban Flight Operations

The emergence and development of advanced technologies and vehicle types have created a growing demand for new forms of flight operations. These new and increasingly complex operational paradigms, such as Advanced and Urban Air Mobility (AAM/UAM), present regulatory authorities and the aviation community with several design-and-implementation challenges – particularly for highly autonomous vehicles. An overarching and daunting task is to develop protocols that can integrate these operations without compromising safety or disrupting traditional airspace operations. A shift toward a more predictive, autonomous, risk mitigation capability becomes critical to meet this challenge. This paper proposes and evaluates a computationally-efficient path planning approach to perform pre-flight planning and autonomous in-flight re-routing to minimize exposures to selected hazards. In our evaluation, hazards associated with degraded and missing critical GPS navigation data are considered. In this paper, we first present a high-performance computing path planning approach based on an adapted Bellman-Ford algorithm, developed in the CUDA programming language. Using the adapted path planning algorithm, we test this algorithm when encountering issues with GPS quality, and deliver an implementation that can produce flight paths that minimize exposure to risks, while maintaining a low computational burden. In our evaluation, the computation of periodic and aperiodic path updates are evaluated, prioritizing specific events as triggers for updates, based on changes to satellite availability. These critical events can lead to significant exposure to navigational hazards if not dealt with correctly.

GNSS↗

A High-Performance Computing GNSS-aware Path Planning Algorithm for Safe Urban Flight Operations

The emergence and development of advanced technologies and vehicle types have created a growing demand for new forms of flight operations. These new and increasingly complex operational paradigms, such as Advanced and Urban Air Mobility (AAM/UAM), present regulatory authorities and the aviation community with several design-and-implementation challenges – particularly for highly autonomous vehicles. An overarching and daunting task is to develop protocols that can integrate these operations without compromising safety or disrupting traditional airspace operations. A shift toward a more predictive, autonomous, risk mitigation capability becomes critical to meet this challenge. This paper proposes and evaluates a computationally-efficient path planning approach to perform pre-flight planning and autonomous in-flight re-routing to minimize exposures to selected hazards. In our evaluation, hazards associated with degraded and missing critical GPS navigation data are considered. In this paper, we first present a high-performance computing path planning approach based on an adapted Bellman-Ford algorithm, developed in the CUDA programming language. Using the adapted path planning algorithm, we test this algorithm when encountering issues with GPS quality, and deliver an implementation that can produce flight paths that minimize exposure to risks, while maintaining a low computational burden. In our evaluation, the computation of periodic and aperiodic path updates are evaluated, prioritizing specific events as triggers for updates, based on changes to satellite availability. These critical events can lead to significant exposure to navigational hazards if not dealt with correctly.

GNSS↗

A High-level Design for Bidirectional Data Streaming to High-Performance Computing Systems from External Science Facilities

Cutting-edge science is increasingly data-driven due to the emergence of scientific machine learning models that can guide scientists toward fruitful areas of exploration. Experimental science facilities such as light and neutron sources, particle colliders, and radio astronomy telescopes are also producing raw measurement data at rates that exceed available data storage and computing capacity at those facilities. As a result, scientific workflows are being developed that concurrently couple experiments at science facilities with high-performance computing (HPC) facilities to enable analysis of experimental data while the experiment is ongoing, and where analysis results are potentially fed back to the experiment for use in experimental control and/or steering in a time-sensitive manner. Our goal is to design, prototype, and deploy a new capability for the Oak Ridge Leadership Computing Facility (OLCF) that enables such workflows through support for bidirectional, memory-based streaming of data from external experiments into and out of OLCF HPC systems. This high-level design document describes the related work and motivating use cases that inform our understanding of the technical requirements for this capability, and describes a proposed architectural solution that meets these requirements and our plans for demonstrating the capability.

97 MATHEMATICS AND COMPUTING↗

The Virtual Blast Furnace - An Integrated High Performance Computing Modeling, Simulation, and Visualization Capability for Steel Manufacturing (Final Report)

Many manufacturing industries require substantial capital and utilize energy intensive processes that involve complex phenomena. One example of such an industry is the steel industry, which is the fourth largest energy consuming industry in the U.S. By harnessing the power of High-Performance Computing (HPC) to enhance current simulation and visualization methods in the steel industry, it should be possible to increase resolution and/or decrease time of these methods by a factor of 1000. In this way, information can be obtained in a time frame that is useful for making business and engineering decisions, optimizing manufacturing processes and, ultimately, improving the completeness of U.S. industries. For example, if coke usage in blast furnaces were optimized such that the average coke rate was reduced from 797 lb/net tonne of hot metal (NTHM) to 604 lb/NTHM, costs could be reduced by $894 million/year. Additionally, members of the steel industry need the flexibility to efficiently operate blast furnaces at a range of production rates in order to meet fluctuating market demands. Large scale parameter studies can be utilized to discover workable operating parameters at a range of production rates.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Layer Time Control for Large Scale Additive Manufacturing Using High Performance Computing

This work proposes to optimize an additive manufacturing AM process to reduce energy and printing cost. The polymer AM process is inherently dependent on the time-temperature history of each layer to maintain geometric tolerances and mechanical integrity. Our preliminary study shows that regression-based layer time control model using thermal images could result in up to 30% build time reduction for simple geometries. This proposed work would use high-performance computing (HPC) to couple the data-driven model with thermal simulation for better predicting layer temperature profiles, improving throughput of large-scale additive manufacturing, and reducing its energy cost. We have developed a method to optimize a layer deposition time (a.k.a. layer time) for large-scale AM via physics-based simulations. A long layer time leads to an over-cooled surface on which a new layer is deposited, and therefore, it may result in a weak bonding or debonding between layers, cracking, or warping. A short layer time leads to a high temperature of the structure due to insufficient cooling, and therefore, the structure may not be stiff enough and may collapse during manufacturing. Therefore, it is important to estimate the optimal layer time in additive manufacturing for a high-quality product. The temperature of a top layer right before deposition is recommended to be slightly higher than the glass temperature of the material. A temperature cooling was approximated to an exponential function of time, and the optimized layer time was obtained based on a target temperature while maintaining a minimal printing time. The material used is carbon fiber-reinforced polycarbonate (CF/PC), and the large-scale deposition system used is LSAM TM from Thermwood Corporation. Three different layer time cases were used for experiments, and a series of thermal images were obtained via an infra-red (IR) camera during the entire AM processes. AM process simulations were performed using a finite element method and the temperature profiles from the simulation were in good agreements with those from experiments. The layer time optimization was performed based on the temperature profiles from the simulations. A layer temperature with the optimal layer time was confirmed as the target temperature through simulation. In addition to the development of a layer time optimization method, we have developed a numerical framework for AM simulation with element activations in sync with toolpath, based on an open source finite element framework, DEAL.II. A major portion of this work was presented at SAMPE 2022 Conference and Exhibition on May 2022, and published in Proceedings of SAMPE 2022.

42 ENGINEERING↗

Parallel-vector unsymmetric Eigen-Solver on high performance computers

The popular QR algorithm for solving all eigenvalues of an unsymmetric matrix is reviewed. Among the basic components in the QR algorithm, it was concluded from this study, that the reduction of an unsymmetric matrix to a Hessenberg form (before applying the QR algorithm itself) can be done effectively by exploiting the vector speed and multiple processors offered by modern high-performance computers. Numerical examples of several test cases have indicated that the proposed parallel-vector algorithm for converting a given unsymmetric matrix to a Hessenberg form offers computational advantages over the existing algorithm. The time saving obtained by the proposed methods is increased as the problem size increased.

Nguyen, Duc T.↗

Simulation of Physics-Based 0-10Hz Strong Motion Using High Performance Computing Supporting Refinements to Regional Ground Motion Models for the Central Eastern US

In collaboration with the U.S. Nuclear Regulatory Commission (NRC) the LLNL has developed a computationally efficient simulation platform designed to perform physics-based ground motion simulations for crustal earthquakes in the Stable Continental Regions of Central and Eastern US (CEUS), using high-performance computing. The main objective of the earthquake simulations was to use synthetic ground motion to provide constrains to refinements of existing ergodic Ground Motion Models (GMMs), for large magnitude earthquakes and near-fault distances, for which these models are less reliable. Physics-based broadband (0-10Hz) ground motion simulations were used to estimate the near-fault ground motion amplitudes and within event and between-event variabilities associated with fault rupture characteristics. In our simulations we used a 3D regional velocity model that was based on Saikia’s 1D velocity model (1994). In simulations performed during the first stage of this project the Saikia’s velocity model demonstrated better performance in modelling high frequency regional wave propagation for the CEUS region recorded during the Mw5.0 November 7, 2016, Cushing Oklahoma (Taylor et al., 2017), and Mw5.8 September 3, 2016, Pawnee Oklahoma earthquakes. The proposed regional 3D model includes random perturbations to the 1D background model using the stochastic scheme of Pitarka and Mellors (2021). In addition, validation analysis of the rupture generator and regional wave propagation models, using comparisons with different GMMs for Mw6.5 and Mw7.0 scenario earthquakes in the CEUS region resulted in a very good match between the simulated and empirical ground motion models. For the purposes of seismic hazard assessment at the existing and planned nuclear power plants, NRC is interested in studies aimed at improving the current ground motion models (GMM) for both Stable Continental Regions (SCR) in the Central and Eastern US and Active Crustal Regions (ACR) in the Western US. Due to lack of recorded data, these improvements require synthetic data for short fault distances and large magnitude earthquakes for which the existing recorded data is not enough to uniquely constrain the GMMs. The need for simulations and strong motion data is especially critical for the CEUS region where we do not have recorded data from potentially large damaging earthquakes with moment magnitudes 6.0 and higher. In this the project, we focused on 10Hz simulations of Mw7.0 scenario earthquakes with strike slip and thrust faulting mechanisms. We used more than 50 Mw7.0 earthquake rupture scenarios to investigate the ground motion uncertainty due to unknown earthquake rupture parameters, in particular, the slip distribution, rupture velocity, and faulting mechanism, and their implication on ground motion amplification due to forward rupture directivity effects.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Classical Preoptimization Approach for ADAPT-VQE: Maximizing the Potential of High-Performance Computing Resources to Improve Quantum Simulation of Chemical Applications

The ADAPT-VQE algorithm is a promising method for generating a compact ansatz based on derivatives of the underlying cost function, and it yields accurate predictions of electronic energies for molecules. In this work, we report the implementation and performance of ADAPT-VQE with our recently developed sparse wave function circuit solver (SWCS) in terms of accuracy and efficiency for molecular systems with up to 52 spin orbitals. The SWCS can be tuned to balance computational cost and accuracy, which extends the application of ADAPT-VQE for molecular electronic structure calculations to larger basis sets and a larger number of qubits. Using this tunable feature of the SWCS, we propose an alternative optimization procedure for ADAPT-VQE to reduce the computational cost of the optimization. Furthermore, by preoptimizing a quantum simulation with a parametrized ansatz generated with ADAPT-VQE/SWCS, we aim to utilize the power of classical high-performance computing in order to minimize the work required on noisy intermediate-scale quantum hardware, which offers a promising path toward demonstrating quantum advantage for chemical applications.

ADAPT-VQE↗