Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25

Development, Verification and Validation of Parallel, Scalable Volume of Fluid CFD Program for Propulsion Applications

There are many instances involving liquid/gas interfaces and their dynamics in the design of liquid engine powered rockets such as the Space Launch System (SLS). Some examples of these applications are: Propellant tank draining and slosh, subcritical condition injector analysis for gas generators, preburners and thrust chambers, water deluge mitigation for launch induced environments and even solid rocket motor liquid slag dynamics. Commercially available CFD programs simulating gas/liquid interfaces using the Volume of Fluid approach are currently limited in their parallel scalability. In 2010 for instance, an internal NASA/MSFC review of three commercial tools revealed that parallel scalability was seriously compromised at 8 cpus and no additional speedup was possible after 32 cpus. Other non-interface CFD applications at the time were demonstrating useful parallel scalability up to 4,096 processors or more. Based on this review, NASA/MSFC initiated an effort to implement a Volume of Fluid implementation within the unstructured mesh, pressure-based algorithm CFD program, Loci-STREAM. After verification was achieved by comparing results to the commercial CFD program CFD-Ace+, and validation by direct comparison with data, Loci-STREAM-VoF is now the production CFD tool for propellant slosh force and slosh damping rate simulations at NASA/MSFC. On these applications, good parallel scalability has been demonstrated for problems sizes of tens of millions of cells and thousands of cpu cores. Ongoing efforts are focused on the application of Loci-STREAM-VoF to predict the transient flow patterns of water on the SLS Mobile Launch Platform in order to support the phasing of water for launch environment mitigation so that vehicle determinantal effects are not realized.

West, Jeff↗

Analyzing Tropical Waves Using the Parallel Ensemble Empirical Model Decomposition Method: Preliminary Results from Hurricane Sandy

In this study, we discuss the performance of the parallel ensemble empirical mode decomposition (EMD) in the analysis of tropical waves that are associated with tropical cyclone (TC) formation. To efficiently analyze high-resolution, global, multiple-dimensional data sets, we first implement multilevel parallelism into the ensemble EMD (EEMD) and obtain a parallel speedup of 720 using 200 eight-core processors. We then apply the parallel EEMD (PEEMD) to extract the intrinsic mode functions (IMFs) from preselected data sets that represent (1) idealized tropical waves and (2) large-scale environmental flows associated with Hurricane Sandy (2012). Results indicate that the PEEMD is efficient and effective in revealing the major wave characteristics of the data, such as wavelengths and periods, by sifting out the dominant (wave) components. This approach has a potential for hurricane climate study by examining the statistical relationship between tropical waves and TC formation.

PEEMD↗

Magnetospheric Multiscale Observations of Large-Amplitude Parallel, Electrostatic Waves Associated with Magnetic Reconnection at the Magnetopause

We report observations from the Magnetospheric Multiscale satellites of large-amplitude, parallel, electrostatic waves associated with magnetic reconnection at the Earth's magnetopause. The observed waves have parallel electric fields (E(sub parallel)) with amplitudes on the order of 100 mV/m and display nonlinear characteristics that suggest a possible net E(sub parallel). These waves are observed within the ion diffusion region and adjacent to (within several electron skin depths) the electron diffusion region. They are in or near the magnetosphere side current layer. Simulation results support that the strong electrostatic linear and nonlinear wave activities appear to be driven by a two stream instability, which is a consequence of mixing cold (less than 10eV) plasma in the magnetosphere with warm (approximately 100eV) plasma from the magnetosheath on a freshly reconnected magnetic field line. The frequent observation of these waves suggests that cold plasma is often present near the magnetopause.

Ergun, R. E.↗

Parallelization of a Six Degree of Freedom Entry Vehicle Trajectory Simulation Using OpenMP and OpenACC

The art and science of writing parallelized software, using methods such as Open Multi-Processing (OpenMP) and Open Accelerators (OpenACC), is dominated by computer scientists. Engineers and non-computer scientists looking to apply these techniques to their project applications face a steep learning curve, especially when looking to adapt their original single threaded software to run multi-threaded on graphics processing units (GPUs). There are significant changes in mindset that must occur; such as how to manage memory, the organization of instructions, and the use of if statements (also known as branching). The purpose of this work is twofold: 1) to demonstrate the applicability of parallelized coding methodologies, OpenMP and OpenACC, to tasks outside of the typical large scale matrix mathematics; and 2) to discuss, from an engineer’s perspective, the lessons learned from parallelizing software using these computer science techniques. This work applies OpenMP, on both multi-core central processing units (CPUs) and Intel® Xeon Phi™ 7210, and OpenACC on GPUs. These parallelization techniques are used to tackle the simulation of thousands of entry vehicle trajectories through the integration of six degree of freedom (DoF) equations of motion (EoM). The forces and moments acting on the entry vehicle, and used by the EoM, are estimated using multiple models of varying levels of complexity. Several benchmark comparisons are made on the execution of six DoF trajectory simulation: single thread Intel® Xeon® E5-2670 CPU, multi-thread CPU using OpenMP, multi-thread Xeon Phi™ 7210 using OpenMP, and multi-thread NVIDIA® Tesla® K40 GPU using OpenACC. These benchmarks are run on the Pleiades Supercomputer Cluster at the National Aeronautics and Space Administration (NASA) Ames Research Center (ARC), and a Xeon Phi™ 7210 node at NASA Langley Research Center (LaRC).

Green, Justin S.↗

Parallelized Quadrupole Simulations of Thermographic Responses of Composites

Thermography has been shown to be a viable technique for inspection of composites. Model inversion of the thermography data requires a fast method for performing the forward problem. Viable numerical methods for the thermal response forward problem are finite element, finite difference and the quadrupole method. Normally both the finite element and finite difference methods solve for the thermal response in the time domain which limits one’s ability to increase the speed of the simulation by parallelization. In contrast, the quadrupole method solves for the Laplace transform of the thermal response. One of the features of the Laplace transform methodology is the solution at any discrete time is independent of the solution at all other times. Therefore, it is easy to separate into a set of independent calculations with each of the times of interest being performed in parallel. Additionally, the numeric inversion of the Laplace transform typically involves numerically solving for the Laplace transform at multiple Laplace frequencies. Each of those solutions are also independent of solutions at other frequencies and can be calculated in parallel. By parallelization of this method, it is possible to perform the simulations of three-dimensional configurations in seconds. When the input stimulus for thermal response is a delta function heat flux (a reasonable approximation for flash heating), the thermal response is smooth. For this case, it is possible to accurately estimate the thermal response at any time within a given time interval from a set of simulations separated by exponentially increasing time steps. From these simulations, it is possible to accurately interpolate to find the response at intermediate times by a spline interpolation of the logarithm of time versus logarithm of temperature. The thermal response with exponential time stepping is shown to produce values for the thermal response which are within 1% of values within the time interval. The simulations are compared to finite element simulations of the same inspection configurations. The simulations are also compared to the thermographic measurements on composites where shape and depth of the delaminations are obtained from other inspection methods.

Thermography↗

Parallelized Quadrupole Simulations of Thermographic Responses of Composites

Thermography has been shown to be a viable technique for inspection of composites. Model inversion of the thermography data requires a fast method for performing the forward problem. Viable numerical methods for the thermal response forward problem are finite element, finite difference and the quadrupole method. Normally both the finite element and finite difference methods solve for the thermal response in the time domain which limits one’s ability to increase the speed of the simulation by parallelization. In contrast, the quadrupole method solves for the Laplace transform of the thermal response. One of the features of the Laplace transform methodology is the solution at any discrete time is independent of the solution at all other times. Therefore, it is easy to separate into a set of independent calculations with each of the times of interest being performed in parallel. Additionally, the numeric inversion of the Laplace transform typically involves numerically solving for the Laplace transform at multiple Laplace frequencies. Each of those solutions are also independent of solutions at other frequencies and can be calculated in parallel. By parallelization of this method, it is possible to perform the simulations of three-dimensional configurations in seconds. When the input stimulus for thermal response is a delta function heat flux (a reasonable approximation for flash heating), the thermal response is smooth. For this case, it is possible to accurately estimate the thermal response at any time within a given time interval from a set of simulations separated by exponentially increasing time steps. From these simulations, it is possible to accurately interpolate to find the response at intermediate times by a spline interpolation of the logarithm of time versus logarithm of temperature. The thermal response with exponential time stepping is shown to produce values for the thermal response which are within 1% of values within the time interval. The simulations are compared to finite element simulations of the same inspection configurations. The simulations are also compared to the thermographic measurements on composites where shape and depth of the delaminations are obtained from other inspection methods.

Thermography↗

UCLA parallel PIC framework

The UCLA Parallel PIC Framework (UPIC) has been developed to provide trusted components for the rapid construction of new, parallel Particle-in-Cell (PIC) codes. The Framework uses object-based ideas in Fortran95, and is designed to provide support for various kinds of PIC codes on various kinds of hardware. The focus is on student programmers. The Framework supports multiple numerical methods, different physics approximations, different numerical optimizations and implementations for different hardware. It is designed with "defensive" programming in mind, meaning that it contains many error checks and debugging helps. Above all, it is designed to hide the complexity of parallel processing. It is currently being used in a number of new Parallel PIC codes.

Norton, Charles D.↗

Parallel Particle Advection Bake-Off for Scientific Visualization Workloads

There are multiple algorithms for parallelizing particle advection for scientific visualization workloads. While many previous studies have contributed to the understanding of individual algorithms, our study aims to provide a holistic understanding of how algorithms perform relative to each other on various workloads. To accomplish this, we consider four popular parallelization algorithms and run a “bake-off” study (i.e., an empirical study) to identify the best matches for each. The study includes 216 tests, going to a concurrency of up to 8192 cores and considering data sets as large as 34 billion cells with 300 million particles. Overall, our study informs three important research questions: (1) which parallelization algorithms perform best for a given workload?, (2) why?, and (3) what are the unsolved problems in parallel particle advection? In terms of findings, we find that the seeding box is the most important factor in choosing the best algorithm, and also that there is a significant opportunity for improvement in execution time, scalability, and efficiency.

Pugmire, Dave↗

POSTER: Automatic Differentiation of Parallel Loops with Formal Methods

The accompanying poster to this short paper presents a combination of reverse mode AD and formal methods to enable efficient differentiation of (or backpropagation through) shared-memory parallel code. Compared to the state of the art, our approach can more often avoid the need for atomic updates or private data copies during the parallel derivative computation, even in the presence of unstructured or data-dependent data access patterns. This is achieved by gathering information about the memory access patterns from the input program, which is assumed to be correctly parallelized. This information is then used to build a model of assertions in a theorem prover, which can be used to check the safety of shared memory accesses during the parallel derivative computation

Automatic Differentiation↗

Speedup of UEDGE Parameter Scans Using Machine-Learning Optimized OpenMP Parallelization and a Continuation Solver

This article presents the OpenMP parallelization of the preconditioning Jacobian assembly and right‐hand side residual evaluation in UEDGE. A continuation algorithm, utilizing the internal NKSOL implicit Jacobian‐Free Newton‐Krylov solver to efficiently scan physical parameters, is also presented. The implemented parallelization reduces the computational time for a benchmark scan run on 32 threads by compared to the serial version when using trained random forest regression models to identify the optimal decomposition of the system of equations. Random forest regression models applied to the UEDGE time‐dependent and continuation solver algorithms did not yield meaningful improvement in computational performance. A benchmark DIII‐D gas injection rate scan in the 0.35–0.75 kA interval, performed on a test cluster using the parallelized code and continuation solver, produced 1066 steady‐state solutions with a 22 s average wall‐clock computational time per steady‐state solution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Massively parallel quantum chemical density matrix renormalization group method

Here, we present, to the best of our knowledge, the first attempt to exploit the supercomputer platform for quantum chemical density matrix renormalization group (QCDMRG) calculations. We have developed the parallel scheme based on the in-house MPI global memory library, which combines operator and symmetry sector parallelisms, and tested its performance on three different molecules, all typical candidates for QC-DMRG calculations. In case of the largest calculation, which is the nitrogenase FeMo cofactor cluster with the active space comprising 113 electrons in 76 orbitals and bond dimension equal to 6000, our parallel approach scales up to approximately 2000 CPU cores.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A parallel hub-and-spoke system for large-scale scenario-based optimization under uncertainty

Practical solution of stochastic programming problems generally requires the use of parallel computing resources. Here, we describe the open source package mpi-sppy, in which efficient and scalable parallelization is a central feature. We report computational experiments that demonstrate the ability to solve very large stochastic programming problems - including mixed-integer variants - in minutes of wall clock time, efficiently leveraging significant parallel computing resources. We report results for the largest publicly available instances of stochastic mixed-integer unit commitment problems, solving to provably tight optimality gaps. In addition, we introduce a novel software architecture that facilitates combinations of methods for accelerating convergence that can be combined in plug-and-play manner. Finally, the mpi-sppy package is written in Python, leverages the widely used Pyomo (http://www.pyomo.org) library for modeling mathematical programs, builds on existing MPI implementations to ensure efficiency and scalability, and is available via http://github.com/Pyomo/mpi-sppy.

97 MATHEMATICS AND COMPUTING↗

Exploring temporal community evolution: algorithmic approaches and parallel optimization for dynamic community detection

Abstract Dynamic (temporal) graphs are a convenient mathematical abstraction for many practical complex systems including social contacts, business transactions, and computer communications. Community discovery is an extensively used graph analysis kernel with rich literature for static graphs. However, community discovery in a dynamic setting is challenging for two specific reasons. Firstly, the notion of temporal community lacks a widely accepted formalization, and only limited work exists on understanding how communities emerge over time. Secondly, the added temporal dimension along with the sheer size of modern graph data necessitates new scalable algorithms. In this paper, we investigate how communities evolve over time based on several graph metrics under a temporal formalization. We compare six different algorithmic approaches for dynamic community detection for their quality and runtime. We identify that a vertex-centric (local) optimization method works as efficiently as the classical modularity-based methods. To its advantage, such local computation allows for the efficient design of parallel algorithms without incurring a significant parallel overhead. Based on this insight, we design a shared-memory parallel algorithm DyComPar , which demonstrates between 4 and 18 fold speed-up on a multi-core machine with 20 threads, for several real-world and synthetic graphs from different domains.

97 MATHEMATICS AND COMPUTING↗

A parallel discrete dislocation dynamics/kinetic Monte Carlo method to study non-conservative plastic processes

Non-conservative processes play a fundamental role in plasticity and are behind important macroscopic phenomena such as creep, dynamic strain aging, loop raft formation, etc. In the most general case, vacancy-induced dislocation climb is the operating unit mechanism. While dislocation/vacancy interactions have been modeled in the literature using a variety of methods, the approaches developed rely on continuum descriptions of both the vacancy population and its fluxes. However, there are numerous situations in physics where point defect populations display heterogeneous concentrations and/or non-smooth kinetics. Here, a kinetic Monte Carlo (kMC) approach for modeling vacancy transport in response to arbitrary stress fields is used. Vacancies are treated as point particles and are coupled to the dislocation substructure representing a deformed material via an advection term defined by the local stress gradients. The stress fields and the dislocation substructure are evolved using a discrete dislocation dynamics (DDD) module. To extend the coupled model to the treatment of large systems, we have implemented it in the massively-parallel DDD code ParaDiS. To avoid numerical incompatibilities associated with merging deterministic (DDD) and stochastic (kMC) integration algorithms, we cast the entire elasto-plastic-diffusive problem within a single stochastic framework, taking advantage of a parallel kMC algorithm to evolve the system as a single event-driven process. The large-scale implementation enables the study of the evolution of a variety of dislocation-defect scenarios governed by non-conservative transport kinetics. After carrying out an exhaustive numerical and computational analysis of our parallel algorithm, we show results that emphasize situations where inhomogeneous vacancy dynamics are of relevance, and compare discrete kinetics to continuum solutions for several cases.

36 MATERIALS SCIENCE↗

Dynamic load balancing with enhanced shared-memory parallelism for particle-in-cell codes

Furthering our understanding of many of today’s interesting problems in plasma physics – including plasma based acceleration and magnetic reconnection with pair production due to quantum electrodynamic effects – requires large-scale kinetic simulations using particle-in-cell (PIC) codes. However, these simulations are extremely demanding, requiring that contemporary PIC codes be designed to efficiently use a new fleet of exascale computing architectures. To this end, the key issue of parallel load balance across computational nodes must be addressed. We discuss the implementation of dynamic load balancing by dividing the simulation space into many small, self-contained regions or ‘‘tiles,’’ along with shared-memory (e.g., OpenMP) parallelism both over many tiles and within single tiles. The load balancing algorithm can be used with three different topologies, including two space-filling curves. Here, we tested this implementation in the code Osiris and show low overhead and improved scalability with OpenMP thread number on simulations with both uniform load and severe load imbalance. Compared to other load-balancing techniques, our algorithm gives order-of-magnitude improvement in parallel scalability for simulations with severe load imbalance issues.

97 MATHEMATICS AND COMPUTING↗

Parallel quantum annealing

Quantum annealers of D-Wave Systems, Inc., offer an efficient way to compute high quality solutions of NP-hard problems. This is done by mapping a problem onto the physical qubits of the quantum chip, from which a solution is obtained after quantum annealing. However, since the connectivity of the physical qubits on the chip is limited, a minor embedding of the problem structure onto the chip is required. In this process, and especially for smaller problems, many qubits will stay unused. We propose a novel method, called parallel quantum annealing, to make better use of available qubits, wherein either the same or several independent problems are solved in the same annealing cycle of a quantum annealer, assuming enough physical qubits are available to embed more than one problem. Although the individual solution quality may be slightly decreased when solving several problems in parallel (as opposed to solving each problem separately), we demonstrate that our method may give dramatic speed-ups in terms of the Time-To-Solution (TTS) metric for solving instances of the Maximum Clique problem when compared to solving each problem sequentially on the quantum annealer. Additionally, we show that solving a single Maximum Clique problem using parallel quantum annealing reduces the TTS significantly.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Unorthodox parallelization for Bayesian quantum state estimation

Quantum state tomography (QST) allows for the reconstruction of quantum states through measurements and some inference technique under the assumption of repeated state preparations. Bayesian inference provides a promising platform to achieve both efficient QST and accurate uncertainty quantification, yet is generally plagued by the computational limitations associated with long Markov chains. In this work, we present a novel Bayesian QST approach that leverages modern distributed parallel computer architectures to efficiently sample a D-dimensional Hilbert space. Using a parallelized preconditioned Crank–Nicholson Metropolis–Hastings algorithm, we demonstrate our approach on simulated data and experimental results from IBM Quantum systems up to four qubits, showing significant speedups through parallelization. Although highly unorthodox in pooling independent Markov chains, our method proves remarkably practical, with validation ex post facto via diagnostics like the intrachain autocorrelation time. We conclude by discussing scalability to higher-dimensional systems, offering a path toward efficient and accurate Bayesian characterization of large quantum systems.

Bayesian inference↗

Towards Low-Overhead Resilience for Data Parallel Deep Learning

Data parallel techniques have been widely adopted both in academia and industry as a tool to enable scalable training of deep learning models. At scale, DL training jobs can fail due to software or hardware bugs, may need to be preempted or terminated due to unexpected events, or may perform suboptimally because they were misconfigured. Under such circumstances, there is a need to recover and/or reconfigure data-parallel DL training jobs on-the-fly, while minimizing the impact on the accuracy of the DNN model and the runtime overhead. In this regard, state-of-art techniques adopted by the HPC community mostly rely on checkpoint-restart, which inevitably leads to loss of progress, thus increasing the runtime overhead. In this paper we explore alternative techniques that exploit the properties of modern deep learning frameworks (overlapping of gradient averaging and weight updates with local gradient computations through pipeline parallelism) to reduce the overhead of resilience/elasticity. To this end we introduce a failure simulation framework and two resilience strategies (immediate mini-batch rollback and lossy forward recovery), which we study compared with checkpoint-restart approaches in a variety of settings in order to understand the trade-offs between the accuracy loss of the DNN model and the runtime overhead.

data-parallel training↗