Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

RAPID

Parallel computer code for the simulator for dynamics of power systems which has the capability to initiate the system and create different faults for the dynamic analysis. The code is based on time-parallel method (Parareal) with Adaptive Method Reduction (AMR). The coarse solvers for the Parareal algorithm include several Semi Analytical Solution methods. Also, Integrated simulation of coupled transmission and distribution systems can be studied.

Simunovic, Srdjan [Oak Ridge National Lab. (ORNL),↗

Benchmarking of massively parallel phase-field codes for directional solidification

We present a detailed benchmark comparing two state-of-the-art phase-field implementations for simulating alloy solidification under experimentally relevant conditions. The study investigates the directional solidification of Al-3wt%Cu under high-velocity solidification conditions and SCN-0.46wt% camphor under microgravity conditions from National Aeronautics and Space Administration (NASA) DECLIC-DSI-R experiments. Both codes, one employing finite-difference discretization with uniform mesh and GPU-acceleration (GPU-PF) and the other one employing finite-element discretization with adaptive-mesh and CPU-parallelization (PRISMS-PF), solve the same quantitative phase-field formulation that incorporates an anti-trapping current for the solidification of dilute alloys. We evaluate the predictions of each code for dendritic morphology, primary spacing, and tip dynamics in both 2D and 3D, as well as their numerical convergence and computational performance. While existing benchmark problems have primarily focused on simplified or small-scale simulations, they do not reflect the computational and modeling challenges posed by employing experimentally relevant time and length scales. Our results provide a practical framework for assessing phase-field code performance as well as validating and facilitating their application in integrated computational materials engineering (ICME) workflows that require integration with realistic experimental data.

36 MATERIALS SCIENCE↗

A High-Performance Design for Hierarchical Parallelism in the QMCPACK Monte Carlo code

We introduce a new high-performance design for parallelism within the Quantum Monte Carlo code QMCPACK. We demonstrate that the new design is better able to exploit the hierarchical parallelism of heterogeneous architectures compared to the previous GPU implementation. The new version is able to achieve higher GPU occupancy via the new concept of crowds of Monte Carlo walkers, and by enabling more host CPU threads to effectively offload to the GPU. The higher performance is expected to be achieved independent of the underlying hardware, significantly improving developer productivity and reducing code maintenance costs. Scientific productivity is also improved with full support for fallback to CPU execution when GPU implementations are not available or CPU execution is more optimal.

Luo, Ye↗

3D Coded SUMMA: Communication-Efficient and Robust Parallel Matrix Multiplication

In this paper, we propose a novel fault-tolerant parallel matrix multiplication algorithm called 3D Coded SUMMA that achieves higher failure-tolerance than replication-based schemes for the same amount of redundancy. This work bridges the gap between recent developments in coded computing and fault-tolerance in high-performance computing (HPC). The core idea of coded computing is the same as algorithm-based fault-tolerance (ABFT), which is weaving redundancy in the computation using error-correcting codes. In particular, we show that MatDot codes, an innovative code construction for parallel matrix multiplications, can be integrated into three-dimensional SUMMA (Scalable Universal Matrix Multiplication Algorithm [30]) in a communication-avoiding manner. To tolerate any two node failures, the proposed 3D Coded SUMMA requires ~50% less redundancy than replication, while the overhead in execution time is only about 5–10%.

97 MATHEMATICS AND COMPUTING↗

Parallel simulation via SPPARKS of on-lattice kinetic and Metropolis Monte Carlo models for materials processing

Abstract SPPARKS is an open-source parallel simulation code for developing and running various kinds of on-lattice Monte Carlo models at the atomic or meso scales. It can be used to study the properties of solid-state materials as well as model their dynamic evolution during processing. The modular nature of the code allows new models and diagnostic computations to be added without modification to its core functionality, including its parallel algorithms. A variety of models for microstructural evolution (grain growth), solid-state diffusion, thin film deposition, and additive manufacturing (AM) processes are included in the code. SPPARKS can also be used to implement grid-based algorithms such as phase field or cellular automata models, to run either in tandem with a Monte Carlo method or independently. For very large systems such as AM applications, the Stitch I/O library is included, which enables only a small portion of a huge system to be resident in memory. In this paper we describe SPPARKS and its parallel algorithms and performance, explain how new Monte Carlo models can be added, and highlight a variety of applications which have been developed within the code.

36 MATERIALS SCIENCE↗

Large language model evaluation for high–performance computing software development

We apply AI-assisted large language model (LLM) capabilities of GPT-3 targeting high-performance computing (HPC) kernels for (i) code generation, and (ii) auto-parallelization of serial code in C ++, Fortran, Python and Julia. Our scope includes the following fundamental numerical kernels: AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG, and language/programming models: (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). Kernel implementations are generated using GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code given simple + + prompt variants. To quantify and compare the generated results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. For auto-parallelization, we use ChatGPT interactively giving simple prompts as in a dialogue with another human including simple “prompt engineering” follow ups. Results suggest that correct outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general-purpose Python can benefit from adding language keywords, while Julia prompts perform acceptably well for its Threads and CUDA.jl programming models. Finally, we expect to provide an initial quantifiable point of reference for code generation in each programming model using a state-of-the-art LLM. Overall, understanding the convergence of LLMs, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

97 MATHEMATICS AND COMPUTING↗

PETSc/TAO Users Manual V.3.21

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication. PETSc/TAO includes a large suite of parallel linear solvers, nonlinear solvers, time integrators, and optimizers that may be used in application codes written in Fortran, C, C++, and Python (via petsc4py; see Getting Started ). The library is organized hierarchically, enabling users to employ the abstraction level most appropriate for a particular problem. By using techniques of object-oriented programming, PETSc provides enormous flexibility for users.

97 MATHEMATICS AND COMPUTING↗

Developing an ELM Ecosystem Dynamics Model on GPU with OpenACC

Porting a complex scientific code, such as the E3SM land model (ELM), onto a new computing architecture is challenging. The paper presents design strategies and technical approaches to develop an ELM ecosystem dynamics model with compiler directives (OpenACC) on NVIDIA GPUs. The code has been refactored with advanced OpenACC features (such as deepcopy and routine directives) to reduce memory consumption and to increase the levels of parallelism through parallel loop reconstruction and new data structures. As a result, the optimized parallel implementation achieved more than a 140-time speedup (50 ms vs 7600 ms), compared to a naive implementation that uses OpenACC routine directive and parallelizes the code across existing loops on a single NVIDIA V100. On a fully loaded computing node with 44 CPUs and 6 GPUs, the code achieved over a 3.0-times speedup, compared to the original code on the CPU. Furthermore, the memory footprint of the optimized parallel implementation is 300 MB, which is around 15% of the 2.15 GB of memory consumed by a naive implementation. This study is the first effort to develop the ELM component on GPUs efficiently to support ultra-high-resolution land simulations at continental scales.

Schwartz, Peter↗

Massively parallel phase-field simulations targeting exascale

The interface thickness in the phase-field (PF) method limits its simulation scales. Consequently, large-scale PF simulations become prohibitively expensive for resolving the extremely fine microstructures that typically form during rapid solidification processing. This challenge is significant in predicting microstructure evolution in metal additive manufacturing and has been identified by the United States Department of Energy’s Exascale Computing Project. Here, to address this, we develop a multi-GPU and MPI-based massively parallel simulation code, utilizing state-of-the-art algorithms, software, and libraries, for large-scale three-dimensional (3D) PF simulations. We report the first GPU-parallel PF simulations on Frontier (currently the second TOP500 exascale cluster) and Summit machines, taking dendritic growth as an example problem. We evaluate the parallel performance of our implementation using scaling studies with more than 24 000 GPUs (among the largest known computations to date) and the acceleration performance using large-scale simulations of dendritic growth in 3D. Finally, massively parallel GPUs in these supercomputers enabled the first coupled multiscale simulations of laser melting and subsequent dendritic solidification on the scale of a full melt-pool, demonstrating the feasibility of performing PF simulations with a point total over 2 billion grid points within an acceptable time.

Exascale↗

A parallel discrete dislocation dynamics/kinetic Monte Carlo method to study non-conservative plastic processes

Non-conservative processes play a fundamental role in plasticity and are behind important macroscopic phenomena such as creep, dynamic strain aging, loop raft formation, etc. In the most general case, vacancy-induced dislocation climb is the operating unit mechanism. While dislocation/vacancy interactions have been modeled in the literature using a variety of methods, the approaches developed rely on continuum descriptions of both the vacancy population and its fluxes. However, there are numerous situations in physics where point defect populations display heterogeneous concentrations and/or non-smooth kinetics. Here, a kinetic Monte Carlo (kMC) approach for modeling vacancy transport in response to arbitrary stress fields is used. Vacancies are treated as point particles and are coupled to the dislocation substructure representing a deformed material via an advection term defined by the local stress gradients. The stress fields and the dislocation substructure are evolved using a discrete dislocation dynamics (DDD) module. To extend the coupled model to the treatment of large systems, we have implemented it in the massively-parallel DDD code ParaDiS. To avoid numerical incompatibilities associated with merging deterministic (DDD) and stochastic (kMC) integration algorithms, we cast the entire elasto-plastic-diffusive problem within a single stochastic framework, taking advantage of a parallel kMC algorithm to evolve the system as a single event-driven process. The large-scale implementation enables the study of the evolution of a variety of dislocation-defect scenarios governed by non-conservative transport kinetics. After carrying out an exhaustive numerical and computational analysis of our parallel algorithm, we show results that emphasize situations where inhomogeneous vacancy dynamics are of relevance, and compare discrete kinetics to continuum solutions for several cases.

36 MATERIALS SCIENCE↗

ExaCA: A performance portable exascale cellular automata application for alloy solidification modeling

Modeling the as-solidified grain structures that form during alloy processing is a critical component in understanding process-property relationships, particularly for additive manufacturing (AM) where grain structure is very sensitive to processing conditions. While cellular automata (CA)-based models have proven able to predict aspects of microstructure for several alloys and AM process conditions, long run times and large resource sets required limit the utility and the problem size to which existing CA models can be applied. As part of the ExaAM project, an initiative within the Exascale Computing Project (ECP) to develop, test, and optimize an exascale-capable coupled and self-consistent model of AM parts, we developed ExaCA (https://github.com/LLNL/ExaCA) for the liquid–solid phase transformation in the wake of AM melt pools. The CA-based code is parallelized using MPI and the Kokkos programming model, the latter enabling simulation on both CPUs and GPUs within a single-source implementation. Here, we detail the steps taken to transform a baseline, MPI-based CA code into one that is performant on CPUs and GPUs. Performance testing of ExaCA on Summit (a pre-exascale machine at Oak Ridge National Laboratory) was used to quantify CPU–GPU speedup comparing with equal numbers of nodes. Testing showed comparable CPU performance to the MPI-only CA code and a 5-20x speedup when running AM-based test problems using GPUs. The improved performance of CA through GPU utilization and the performance portable nature of ExaCA will enable accurate part-scale modeling by harnessing the power of current and future generations of high performance computing resources. Future work will include improving the strong scaling of ExaCA on GPUs by reducing load imbalance associated with the locality of the problem, and continuing performance optimization across exascale hardware.

36 MATERIALS SCIENCE↗

Phase Field Dislocation Dynamics (PFDD) version 2.x

This disclosure is for version 2.x of a mesoscale model called Phase Field Dislocation Dynamics (PFDD). PFDD is used for investigating deformation in nanoscale (grain sizes of ~300 nm and less) materials, such as metals and alloys. This approach models the motion and interaction of individual defects, namely dislocations, in the material using scalar-valued phase field variables, also called order parameters. The system is evolved through energy minimization thus the model calculates the total energy density in terms of the phase field variables. The energy minimization is completed using the Ginzburg-Landau equation, and is implemented with explicit time integration. The total system energy can be comprised of several terms, including the strain energy (which describes dislocation-dislocation interactions), the energy due to an applied stress (dislocation interactions with the applied stress), and a core/lattice (perfect dislocations) or generalized stacking fault (partial dislocations) energy (described the dislocation core structure). The latter term in particular may vary based on the crystal structure being modeled and is typically informed using lower length scale (e.g., atomistic) approaches, although no such (atomistic) calculations are completed within the PFDD algorithm. This basic formulation was previously reviewed by Los Alamos National Laboratory and released under license number C17113. This previously reviewed version we will henceforth refer to as PFDD v1.0. PFDD v1.0 consisted of 2 codes (one parallel and one serial) plus input files, all written in the C language. This new disclosure is addressing the next versions of the PFDD, versions 2.x. There have been several enhancements of PFDD v1.0, which are described here and included in the attached code, which we will refer to as PFDD v2.0. There are also several new features described here that are either planned or already in process and are expected to be subsequent releases, i.e., v2.1, v2.2, ...v2.x.

Hunter, Abigail↗

Optimizing the hit finding algorithm for liquid argon TPC neutrino detectors using parallel architectures

Neutrinos are particles that interact rarely, so identifying them requires large detectors which produce lots of data. Processing this data with the computing power available is becoming even more difficult as the detectors increase in size to reach their physics goals. Liquid argon time projection chamber (LArTPC) neutrino experiments are expected to grow in the next decade to have 100 times more wires than in currently operating experiments, and modernization of LArTPC reconstruction code, including parallelization both at data- and instruction-level, will help to mitigate this challenge. The LArTPC hit finding algorithm is used across multiple experiments through a common software framework. In this paper we discuss a parallel implementation of this algorithm. Using a standalone setup we find speedup factors of two times from vectorization and 30–100 times from multi-threading on Intel architectures. The new version has been incorporated back into the framework so that it can be used by experiments. On a serial execution, the integrated version is about 10 times faster than the previous one and, once parallelization is enabled, more speedups comparable to the standalone program are achieved.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

PIAFS: A 2D nonlinear hydrodynamics code to model gaseous optics

The survivability of final optics is expected to be a major challenge for all future inertial fusion energy concepts. Due to their higher damage threshold, gaseous optics have been identified as a promising solution to this problem. Gaseous optics can be created through the photoabsorption of spatially modulated UV light, which induces various chemical processes that heat the gas. This heating leads to a pressure perturbation, which in turn launches a density perturbation that can imprint a refractive index modulation such as a grating. In this article, we introduce a parallel C/C++ code to simulate gaseous optics. PIAFS2D is a high-order conservative finite-difference code to solve the compressible Navier–Stokes equations along with the photochemical heating sources on Cartesian grids. The simulations are validated by the linear theory derived in a previous paper [Michel et al., Phys. Rev. Appl. 22, 024014 (2024)]. For larger perturbations, the behavior of the system—particularly the evolution of the generated acoustic wave—demonstrates strong nonlinearity. PIAFS2D allows the study of nonlinear behaviors and can be used for the design of high-efficiency gaseous optics elements in realistic experimental conditions.

Oudin, A. [Lawrence Livermore National Laboratory ↗

Parallelized real-time physics codes for plasma control on DIII-D

A real-time safe multi-threading library was developed on the DIII-D plasma control system to optimize the real-time TORBEAM and real-time STRIDE physics codes. These physics codes are crucial for future fusion power plant operation as they provide information about electron cyclotron wave propagation and heating as well as inform about ideal plasma stability limits. The real-time TORBEAM code executed consistently in under 20 ms while the real-time STRIDE code computes in 100 ms. The multi-threading library developed in this work can be applied to other real-time physics-based codes that will be crucial for the next generation of fusion devices.

DIII-D↗

The LISE package: solvers for static and time-dependent superfluid local density approximation equations in three dimensions

Nuclear implementation of the density functional theory (DFT) is at present the only microscopic framework applicable to the whole nuclear landscape. The extension of DFT to superfluid systems in the spirit of the Kohn-Sham approach, the superfluid local density approximation (SLDA) and its extension to time-dependent situations, time-dependent superfluid local density ap- proximation (TDSLDA), have been extensively used to describe various static and dynamical problems in nuclear physics, neutron star crust, and cold atom systems. In this paper, we present the codes that solve the static and time-dependent SLDA equations in three-dimensional coordinate space without any symmetry restriction. These codes are fully parallelized with the message passing interface (MPI) library and take advantage of graphic processing units (GPU) for accelerating execution. The dynamic codes have checkpoint/restart capabilities and for initial conditions one can use any generalized Slater determinant type of wave function. The code can describe a large number of physical problems: nuclear fission, collisions of heavy ions, the interaction of quantized vor- tices with nuclei in the nuclear star crust, excitation of superfluid fermion systems by time dependent external fields, quantum shock waves, domain wall generation and propagation, the dynamics of the Anderson-Bogoliubov-Higgs mode, dynamics of fragmented condensates, vortex rings dynamics, generation and dynamics of quantized vortices, their crossing and recombinations and the incipient phases of quantum turbulence.

Jin, Shi↗

Causal explicit algorithm for heat conduction in a plasma

Hyperbolic heat conduction extends standard Spitzer-Harm heat conduction by including a term proportional to the time derivative of the heat flux. The new term arises from a kinetic derivation of the heat flux that includes higher order corrections. Here we present a causal explicit numerical algorithm for solving the nonlinear hyperbolic heat conduction equation in an unmagnetized plasma. The maximum stable timestep for the causal explicit algorithm scales linearly with the cell size, owing to the hyperbolic nature of the problem. This is in contrast to the quadratic scaling of the maximum stable timestep with the cell size for the parabolic forward time centered space algorithm. The favorable scaling of the timestep with the cell size enables a practical explicit implementation of heat conduction in high-performance massively parallel plasma codes. In particular, we have implemented the causal explicit algorithm in the laser plasma interaction code pF3D. We verify the CE algorithm and analyze its convergence rate by simulating a harmonic mode, which has an analytic solution within the context of the HHC model. We also compare simulations using the CE algorithm to those using the forward time centered space algorithm on a pair of test problems: evolution in time of a Gaussian temperature perturbation in a uniform plasma and heat transport in the presence of inverse bremsstrahlung heating by a Gaussian laser speckle.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

EUTERPE: A global gyrokinetic code for stellarator geometry

The current state of the EUTERPE code is described with emphasis on the implemented models and their numerical implementation. The code solves the multi-species electromagnetic gyrokinetic equations in the full volume of a three-dimensional domain. Noise reduction of the particle-in-cell method is achieved by using a δf-method and Fourier filters. The field equations are discretized with B-splines and the resulting system of equations is solved iteratively. For linear simulations a phase-factor transformation is applied in order to strongly reduce the necessary grid resolution. Apart from the full gyrokinetic model, other numerically less expensive hybrid models are also implemented. They are mainly tailored for comparison with fluid theory and for studying the interaction of the bulk plasma with fast particles. The code is parallelized for CPUs by particle and domain decomposition. Good scalability up to several thousand nodes is demonstrated.

97 MATHEMATICS AND COMPUTING↗