Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Analysis of Threading Libraries for High Performance Computing

With the appearance of multi-/many core machines, applications and runtime systems have evolved in order to exploit the new on-node concurrency brought by new software paradigms. POSIX threads (Pthreads) was widely-adopted for that purpose and it remains as the most used threading solution in current hardware. Lightweight thread (LWT) libraries emerged as an alternative offering lighter mechanisms to tackle the massive concurrency of current hardware. In this article, we analyze in detail the most representative threading libraries including Pthread- and LWT-based solutions. In addition, to examine the suitability of LWTs for different use cases, we develop a set of microbenchmarks consisting of OpenMP patterns commonly found in current parallel codes, and we compare the results using threading libraries and OpenMP implementations. Moreover, we study the semantics offered by threading libraries in order to expose the similarities among different LWT application programming interfaces and their advantages over Pthreads. This article exposes that LWT libraries outperform solutions based on operating system threads when tasks and nested parallelism are required.

GLT↗

Coupled Lattice Boltzmann Modeling Framework for Pore-Scale Fluid Flow and Reactive Transport

In this paper, we propose a modeling framework for pore-scale fluid flow and reactive transport based on a coupled lattice Boltzmann model (LBM). We develop a modeling interface to integrate the LBM modeling code parallel lattice Boltzmann solver and the PHREEQC reaction solver using multiple flow and reaction cell mapping schemes. The major advantage of the proposed workflow is the high modeling flexibility obtained by coupling the geochemical model with the LBM fluid flow model. Consequently, the model is capable of executing one or more complex reactions within desired cells while preserving the high data communication efficiency between the two codes. Meanwhile, the developed mapping mechanism enables the flow, diffusion, and reactions in complex pore-scale geometries. We validate the coupled code in a series of benchmark numerical experiments, including 2D single-phase Poiseuille flow and diffusion, 2D reactive transport with calcite dissolution, as well as surface complexation reactions. The simulation results show good agreement with analytical solutions, experimental data, and multiple other simulation codes. In addition, we design an AI-based optimization workflow and implement it on the surface complexation model to enable increased capacity of the coupled modeling framework. Compared to the manual tuning results proposed in the literature, our workflow demonstrates fast and reliable model optimization results without incorporating pre-existing domain knowledge.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Parallel Simulation of Quantum Networks with Distributed Quantum State Management

Quantum network simulators offer the opportunity to cost-efficiently investigate potential avenues for building networks that scale with the number of users, communication distance, and application demands by simulating alternative hardware designs and control protocols. Several quantum network simulators have been recently developed with these goals in mind. As the size of the simulated networks increases, however, sequential execution becomes time-consuming. Parallel execution presents a suitable method for scalable simulations of large-scale quantum networks, but the unique attributes of quantum information create unexpected challenges. In this work, we identify requirements for parallel simulation of quantum networks and develop the first parallel discrete-event quantum network simulator by modifying the existing serial simulator SeQUeNCe. Our contributions include the design and development of a quantum state manager (QSM) that maintains shared quantum information distributed across multiple processes. We also optimize our parallel code by minimizing the overhead of the QSM and decreasing the amount of synchronization needed among processes. Using these techniques, we observe a speedup of 2 to 25 times when simulating a 1,024-node linear network topology using 2 to 128 processes. We also observe an efficiency greater than 0.5 for up to 32 processes in a linear network topology of the same size and with the same workload. We repeat this evaluation with a randomized workload on a caveman network. We also introduce several methods for partitioning networks by mapping them to different parallel simulation processes. We have released the parallel SeQUeNCe simulator as an open source tool alongside the existing sequential version.

97 MATHEMATICS AND COMPUTING↗

Speedup of UEDGE Parameter Scans Using Machine-Learning Optimized OpenMP Parallelization and a Continuation Solver

This article presents the OpenMP parallelization of the preconditioning Jacobian assembly and right‐hand side residual evaluation in UEDGE. A continuation algorithm, utilizing the internal NKSOL implicit Jacobian‐Free Newton‐Krylov solver to efficiently scan physical parameters, is also presented. The implemented parallelization reduces the computational time for a benchmark scan run on 32 threads by compared to the serial version when using trained random forest regression models to identify the optimal decomposition of the system of equations. Random forest regression models applied to the UEDGE time‐dependent and continuation solver algorithms did not yield meaningful improvement in computational performance. A benchmark DIII‐D gas injection rate scan in the 0.35–0.75 kA interval, performed on a test cluster using the parallelized code and continuation solver, produced 1066 steady‐state solutions with a 22 s average wall‐clock computational time per steady‐state solution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Exploring the scaling limitations of the variational quantum eigensolver with the bond dissociation of hydride diatomic molecules

Abstract Materials simulations involving strongly correlated electrons pose fundamental challenges to state‐of‐the‐art electronic structure methods but are hypothesized to be the ideal use case for quantum computing algorithms. To date, no quantum computer has simulated a molecule of a size and complexity relevant to real‐world applications, despite the fact that the variational quantum eigensolver (VQE) algorithm can predict chemically accurate total energies. Nevertheless, because of the many applications of moderately sized, strongly correlated systems, such as molecular catalysts, the successful use of the VQE stands as an important waypoint in the advancement toward useful chemical modeling on near‐term quantum processors. In this paper, we take a significant step in this direction. We lay out the steps, write, and run parallel code for an (emulated) quantum computer to compute the bond dissociation curves of the TiH, LiH, NaH, and KH diatomic hydride molecules using the VQE. TiH was chosen as a relatively simple chemical system that incorporates d orbitals and strong electron correlation. Because current VQE implementations on existing quantum hardware are limited by qubit error rates, the number of qubits available, and the allowable gate depth, recent studies using it have focused on chemical systems involving s and p block elements. Through VQE + UCCSD calculations of TiH, we evaluate the near‐term feasibility of modeling a molecule with d‐orbitals on real quantum hardware. We demonstrate that the inclusion of d‐orbitals and the use of the UCCSD ansatz, which are both necessary to capture the correct TiH physics, dramatically increase the cost of this problem. We estimate the approximate error rates necessary to model TiH on current quantum computing hardware using VQE + UCCSD and show them to likely be prohibitive until significant improvements in hardware and error correction algorithms are available.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Investigation of Cycle-to-Cycle Variations in Internal Combustion Engine Using Proper Orthogonal Decomposition

The understanding, modeling and control of the cycle-to-cycle variation (CCV) in the modern internal combustion engine (ICE) is a key scientific challenge to achieve stable engine operation. High CCV in the engine combustion chamber may contribute to partial burn, misfire and knock, which adversely affects the engine performance and may potentially damage the engine. The objective of the current study is to leverage high-fidelity numerical simulations to improve the understanding of the causes of CCV. Using the massively parallel code, Nek5000, multi-cycle, wall-resolved large-eddy simulations (LES) were performed for the General Motors (GM), Transparent Combustion Chamber (TCC-III) optical engine under motored operating conditions. Further, the large-scale structures of the in-cylinder flow were investigated using a triple proper orthogonal decomposition (POD) technique to explore the characteristics of different parts of the flow and their contributions to CCV. The kinetic energy of the subset of flow structures were determined and correlated between the intake and compression strokes. The insights from the analysis of the large-scale flow structures will be used to assist the development of improved engine designs with reduced CCV and enhance the engine performance.

33 ADVANCED PROPULSION SYSTEMS↗

Kinetic Monte Carlo simulations of structural evolution during anneal of additively manufactured materials

Our experiments indicated that upon a post-processing anneal, an additively manufactured 316L stainless steel exhibits cubic grains rather than the conventional equiaxed grains. In this work, we have used kinetic Monte Carlo simulations to explore the origin of these cubic grains. First, we implemented a new kinetic Monte Carlo model in parallel code SPPARKS to simulate grain growth and recrystallization under a residual energy distribution. Our model incorporates physical properties and real-time, as opposed to generic properties and relative time. We further validated that our SPPARKS simulations reproduced the expected kinetic behavior of single-grain evolution. We then used the validated approach to simulate the anneal of an additively manufactured material under the same conditions used in our experiments. We found that the cubic grains can origin from a periodically varying residual energy that may be present in additively manufactured materials.

36 MATERIALS SCIENCE↗

Permutationally Invariant Polynomial Expansions with Unrestricted Complexity

A general strategy is presented for constructing and validating permutationally invariant polynomial (PIP) expansions for chemical systems of any stoichiometry. Demonstrations are made for three categories of gas-phase dynamics and kinetics: collisional energy-transfer trajectories for predicting pressure-dependent kinetics, three-body collisions for describing transient van der Waals adducts relevant to atmospheric chemistry, and nonthermal reactivity via quasiclassical trajectories. In total, 30 systems are considered with up to 15 atoms and 39 degrees of freedom. Permutational invariance is enforced in PIP expansions with as many as 13 million terms and 13 permutationally distinct atom types by taking advantage of petascale computational resources. The quality of the PIP expansions is demonstrated through the systematic convergence of in-sample and out-of-sample errors with respect to both the number of training data and the order of the expansion, and these errors are shown to predict errors in the dynamics for both reactive and nonreactive applications. Here, the parallelized code distributed as part of this work enables the automation of PIP generation for complex systems with multiple channels and flexible user-defined symmetry constraints and for automatically removing unphysical unconnected terms from the basis set expansions, all of which are required for simulating complex reactive systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

tih_vqe [SWR-23-32]

This software supports the paper, "Exploring the scaling limitations of the variational quantum eigensolver with the bond dissociation of hydride diatomic molecules," published in the International Journal of Quantum Chemistry, whose abstract is as follows: Materials simulations involving strongly correlated electrons pose fundamental challenges to state-of-the-art electronic structure methods but are hypothesized to be the ideal use case for quantum computing. To date, no quantum computer has simulated a molecule of a size and complexity relevant to real-world applications, despite the fact that the variational quantum eigensolver (VQE) algorithm can predict chemically accurate total energies. Nevertheless, because of the many applications of moderately-sized, strongly correlated systems, such as molecular catalysts, the successful use of the VQE stands as an important waypoint in the advancement toward useful chemical modeling on near-term quantum processors. In this paper, we take a significant step in this direction. We lay out the steps, write, and run parallel code for an (emulated) quantum computer to compute the bond dissociation curves of the TiH, LiH, NaH, and KH diatomic hydride molecules using VQE. TiH was chosen as a relatively simple chemical system that incorporates d orbitals and strong electron correlation. Because current VQE implementations on existing quantum hardware are limited by qubit error rates, the number of qubits available, and the allowable gate depth, recent studies have focused on chemical systems involving s and p block elements. Through VQE + UCCSD calculations of TiH, we evaluate the near-term feasibility of modeling a molecule with d-orbitals on real quantum hardware. We demonstrate that the inclusion of d-orbitals and the use of the UCCSD ansatz, which are both necessary to capture the correct TiH physics, dramatically increase the cost of this problem. We estimate the approximate error rates necessary to model TiH on current quantum computing hardware using VQE+UCCSD and show them to likely be prohibitive until significant improvements in hardware and error correction algorithms are available.

Graf, Peter↗

Efficient phase-space generation for hadron collider event simulation

We present a simple yet efficient algorithm for phase-space integration at hadron colliders. Individual mappings consist of a single t-channel combined with any number of s-channel decays, and are constructed using diagrammatic information. The factorial growth in the number of channels is tamed by providing an option to limit the number of s-channel topologies. We provide a publicly available, parallelized code in C++ and test its performance in typical LHC scenarios.

47 OTHER INSTRUMENTATION↗

FULL RANGE TUNE SCAN STUDIES USING GRAPHICS PROCESSING UNITS WITH CUDA IN EIC BEAM-BEAM SIMULATIONS

The hadron beam in the Electron-Ion Collider (EIC) suffers high order betatron and synchro-betatron resonances. In this paper, we present a weak-strong full range (0.0 ~ 0.5) fractional tune scan with a step size as small as 0.001. Multiple Graphics Processing Units (GPUs) are used to speed up the simulation. A code parallelized with MPI and CUDA is implemented. The good tune region from weak-strong scan is further checked by the self-consistent strong-strong simulation. This study provides beam dynamics guidance in choosing proper working points for the future EIC.

43 PARTICLE ACCELERATORS↗

PETSc/TAO Users Manual (Rev. 3.19)

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for the implementation of large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication. PETSc/TAO includes a large suite of parallel linear solvers, nonlinear solvers, time integrators, and opti mization that may be used in application codes written in Fortran, C, C++, and Python (via petsc4py; see Getting Started). PETSc provides many of the mechanisms needed within parallel application codes, such as parallel matrix and vector assembly routines. The library is organized hierarchically, enabling users to employ the level of abstraction that is most appropriate for a particular problem. By using techniques of object-oriented programming, PETSc provides enormous flexibility for users. PETSc is a sophisticated set of software tools; as such, for some users it initially has a much steeper learning curve than packages such as MATLAB or a simple subroutine library. In particular, for individuals without some computer science background, experience programming in C, C++, python, or Fortran and experience using a debugger such as gdb or lldb, it may require a significant amount of time to take full advantage of the features that enable efficient software use. However, the power of the PETSc design and the algorithms it incorporates may make the efficient implementation of many application codes simpler than “rolling them” yourself. For many tasks a package such as MATLAB is often the best tool; PETSc is not intended for the classes of problems for which effective MATLAB code can be written. There are several packages, built on PETSc, that may satisfy your needs without requiring directly using PETSc. We recommend reviewing these packages functionality before starting to code directly with PETSc. PETSc can be used to provide a “MPI parallel linear solver” in an otherwise sequential, or OpenMP parallel code. This approach cannot provide extremely large improvements in the application time by utilizing large numbers of MPI processes but can still improve the performance. Certainly all parts of a previously sequential code need not be parallelized but the matrix generation portion must be parallelized to expect true scalability to large numbers of MPI processes. See PCMPI for details on how to utilize the PETSc MPI linear solver server. Since PETSc is under continued development, small changes in usage and calling sequences of routines will occur. PETSc has been supported for twenty-five years; see mailing list information on our website for information on contacting support.

97 MATHEMATICS AND COMPUTING↗

ddcMD.os

ddcMD is a general purpose molecular dynamics (MD) code that supports MPI parallelism. MD codes are used for simulation of particles systems and capture all the many-body effects of the underlying particle potential that defines the physical system. Though MD can be used to model systems from the subatomic to astrological length scales ddcMD is mainly focused on the atomic scale length scale, In this release of ddcMD the support will be mainly for systems using the coarse-grain Martini potential, a particle potential for biological systems.

Glosli, JamesN↗

PETSc Users Manual (Rev. 3.13)

This manual describes the use of PETSc for the numerical solution of partial differential equations and related problems on high-performance computers. The Portable, Extensible Toolkit for Scientific Computation (PETSc) is a suite of data structures and routines that provide the building blocks for the implementation of large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all message-passing communication. PETSc includes an expanding suite of parallel linear solvers, nonlinear solvers, and time integrators that may be used in application codes written in Fortran, C, C++, and Python. PETSc provides many of the mechanisms needed within parallel application codes, such as parallel matrix and vector assembly routines. The library is organized hierarchically, enabling users to employ the level of abstraction that is most appropriate for a particular problem. By using techniques of object-oriented programming, PETSc provides enormous flexibility for users. PETSc is a sophisticated set of software tools; as such, for some users it initially has a much steeper learning curve than a simple subroutine library. In particular, for individuals without some computer science background, experience programming in C, C++, python, or Fortran and experience using a debugger such as gdb or dbx, it may require a significant amount of time to take full advantage of the features that enable efficient software use. However, the power of the PETSc design and the algorithms it incorporates may make the efficient implementation of many application codes simpler than \rolling them" yourself.

97 MATHEMATICS AND COMPUTING↗

TINES - Time Integration, Newton and Eigen Solver v. 1.0

SAND2021-1505 O. TINES is an open source software providing math infrastructure for solving many stiff time ordinary differential equations (ODEs) and/or differential algebraic equations (DAEs) using a batch hierarchical parallelism. The code is written using a parallel programming model (i.e., Kokkos) to future-proof the next generation parallel computing platforms such as GPU accelerators. This code is developed to support Exascale Catalytic Chemistry (ECC) Project. The code provides fundamental math helpers that can aid other research projects. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Kim, Kyungjoo↗

KokkACC: Enhancing Kokkos with OpenACC

Template metaprogramming is gaining popularity as a high-level solution for achieving performance portability on heterogeneous computing resources. Kokkos is a representative approach that offers programmers high-level abstractions for generic programming while most of the device-specific code generation and optimizations are delegated to the compiler through template specializations. For this, Kokkos provides a set of device-specific code specializations in multiple back ends, such as CUDA and HIP. Unlike CUDA or HIP, OpenACC is a high-level and directive-based programming model. This descriptive model allows developers to insert hints (pragmas) into their code that help the compiler to parallelize the code. The compiler is responsible for the transformation of the code, which is completely transparent to the programmer. This paper presents an OpenACC back end for Kokkos: KokkACC. As an alternative to Kokkos’s existing device-specific back ends, KokkACC is a multi-architecture back end providing a high-productivity programming environment enabled by OpenACC’s high-level and descriptive programming model. Moreover, we have observed competitive performance; in some cases, KokkACC is faster (up to 9×) than NVIDIA’s CUDA back end and much faster than OpenMP’s GPU offloading back end. This work also includes implementation details and a detailed performance study conducted with a set of mini-benchmarks (AXPY and DOT product) and three mini-apps (LULESH, miniFE and SNAP, a LAMMPS proxy mini-app).

Valero Lara, Pedro↗