Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “massively parallel algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Multiscale Ecosystem for solving Maxwell-Schrodinger equations of open quantum systems (OpenMS)

Light-matter interactions play an important role in many branches of physics, chemistry, energy, and materials science. In the strong coupling regime, light-matter interactions are able to tune the materials properties via the formation of new quasiparticles, such as plasmons and polaritons. However, theoretical and numerical modeling of the light-matter interaction-mediated processes remain a big challenge because light-matter interactions are fundamentally multiscale and multiphysics problems involving multiple interactions between electrons, nuclei, and photons at different time/length scales. In the current community, light-matter interactions were treated at different levels of theoretical complexity in quantum chemistry and quantum optics. In quantum optics or quantum photonics, the matter is usually simplified as a few-level system, and the light is treated quantum-mechanically. On the other hand, quantum chemistry explores first-principles methods, including both single-particle and many-body-based techniques, to describe the electronic properties of matter in detail. However, light is usually prescribed as a classical electromagnetic field, and the light-matter interaction is taken into account as an external potential via classical approximations. This software is designed to fill current modeling shortcomings by delivering a first-ever scalable multiscale platform for simulating light-matter interactions in realistic electromagnetic environments. The software solves Maxwell and Schrodinger equations self-consistent on the heterogeneous platforms. It implements HF/DFT, TDDFT, and coupled-cluster counterparts for light-matter interactions and adopts modular programming to offload massively parallel algorithms on a large number of CPU/GPUs.

Zhang, Yu↗

Virtual Time III, Part 2: Combining Conservative and Optimistic Synchronization

This is Part 2 of a trio of works intended to provide a unifying framework in which conservative and optimistic synchronization for parallel discrete event simulations can be freely and transparently combined in the same logical process on an event-by-event basis. Here, in this article, we continue the outline of an approach called Unified Virtual Time (UVT) that was introduced in Part 1, showing in detail via two extended examples how conservative synchronization can be refactored and combined with optimistic synchronization in the UVT framework. We describe UVT versions of both a basic time windowing algorithm called Unified Simple Time Windows and a refactored version of the Chandy-Misra-Bryant Null Message algorithm called Unified CMB.

97 MATHEMATICS AND COMPUTING↗

High speed two-dimensional event detection and imaging using an analog interface and a massively parallel processor

A quantitative pulse count (event detection) algorithm with linearity to high count rates is accomplished by combining a high-speed, high frame rate camera with simple logic code run on a massively parallel processor such as a GPU or a FPGA. The parallel processor elements examine frames from the camera pixel by pixel to find and tag events or count pulses. The tagged events are combined to form a combined quantitative event image.

Waugh, Justin↗

Massively-parallel Lagrangian particle code and applications

Massively-parallel, distributed-memory algorithms for the Lagrangian particle hydrodynamic method (Samulyak et al., 2018) have been developed, verified, and implemented. The key component of parallel algorithms is a particle management module that includes a parallel construction of octree databases, dynamic adaptation and refinement of octrees, and particle migration between parallel subdomains. The particle management module is based on the p4est (parallel forest of k-trees) library. The massively-parallel Lagrangian particle code has been applied to a variety of fundamental science and applied problems. A summary of Lagrangian particle code applications to the injection of impurities into thermonuclear fusion devices and to the simulation of supersonic hydrogen jets in support of laser-plasma wakefield acceleration research has also been presented.

97 MATHEMATICS AND COMPUTING↗

PeleLMeX: an AMR Low Mach Number Reactive Flow Simulation Code without level sub-cycling

PeleLMeX simulates chemically reacting low Mach number flows with block-structured adaptive mesh refinement (AMR). The code is built upon the AMReX library, which provides the underlying data structures and tools to manage and operate on them across massively parallel computing architectures. PeleLMeX algorithmic features are inherited from its predecessor PeleLM but key improvements allow representation of more complex physical processes. Together with its compressible flow counterpart PeleC, the thermo-chemistry library PelePhysics and the multi-physics library PeleMP, it forms the Pele suite of open-source reactive flow simulation codes.

97 MATHEMATICS AND COMPUTING↗

Extension of the high-resolution thermal-hydraulics code ESCOT to hexagonal core geometries for multi-physics calculations

The extension of the capabilities of the pin-level nuclear reactor core thermal-hydraulics (T/H) code ESCOT to analyze hexagonal fueled cores and its performance are presented. ESCOT is an accurate yet fast core thermal-hydraulics solution aiming at high-fidelity and high-resolution multi-physics core analysis in the framework of massively parallel computing platforms. Its algorithm solution is based on the four-equation drift-flux model for two-phase calculations, these are numerically solved by applying the Finite Volume Method (FVM) and the Semi-Implicit Method for Pressure-Linked Equation (SIMPLE)-like algorithm in a staggered grid system. Constitutive models such as turbulent mixing, pressure drop, and vapor generation are employed to simulate key phenomena in subchannel-scale analysis. ESCOT is parallelized by a double (radial and axial) domain decomposition that enables its highly parallelized execution. The coupling of the code with the neutronics whole core solver for hexagonal geometries nTRACER is described. The newly implemented ESCOT features are validated by comparing single assembly and full core steady state nTRACER-ESCOT solutions with nTRACER standalone internal one-dimensional T/H solver results. The validation problems are based on the VVER 440 and VVER 1000 cores. ESCOT results show differences within an acceptable range with respect to the simple 1D nTRACER built-in solver. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Track reconstruction as a service for collider physics

Optimizing charged-particle track reconstruction algorithms is crucial for efficient event reconstruction in Large Hadron Collider (LHC) experiments due to their significant computational demands. Existing track reconstruction algorithms have been adapted to run on massively parallel coprocessors, such as graphics processing units (GPUs), to reduce processing time. Nevertheless, challenges remain in fully harnessing the computational capacity of coprocessors in a scalable and non-disruptive manner. This paper proposes an inference-as-a-service approach for particle tracking in high energy physics experiments. To evaluate the efficacy of this approach, two distinct tracking algorithms are tested: Patatrack, a rule-based algorithm, and Exa.TrkX, a machine learning-based algorithm. The as-a-service implementations show enhanced GPU utilization and can process requests from multiple CPU cores concurrently without increasing per-request latency. The impact of data transfer is minimal and insignificant compared to running on local coprocessors. This approach greatly improves the computational efficiency of charged particle tracking, providing a solution to the computing challenges anticipated in the High-Luminosity LHC era.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Causal explicit algorithm for heat conduction in a plasma

Hyperbolic heat conduction extends standard Spitzer-Harm heat conduction by including a term proportional to the time derivative of the heat flux. The new term arises from a kinetic derivation of the heat flux that includes higher order corrections. Here we present a causal explicit numerical algorithm for solving the nonlinear hyperbolic heat conduction equation in an unmagnetized plasma. The maximum stable timestep for the causal explicit algorithm scales linearly with the cell size, owing to the hyperbolic nature of the problem. This is in contrast to the quadratic scaling of the maximum stable timestep with the cell size for the parabolic forward time centered space algorithm. The favorable scaling of the timestep with the cell size enables a practical explicit implementation of heat conduction in high-performance massively parallel plasma codes. In particular, we have implemented the causal explicit algorithm in the laser plasma interaction code pF3D. We verify the CE algorithm and analyze its convergence rate by simulating a harmonic mode, which has an analytic solution within the context of the HHC model. We also compare simulations using the CE algorithm to those using the forward time centered space algorithm on a pair of test problems: evolution in time of a Gaussian temperature perturbation in a uniform plasma and heat transport in the presence of inverse bremsstrahlung heating by a Gaussian laser speckle.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

In situ multi-tier auto-ignition detection applied to dual-fuel combustion simulations

Here we use an anomaly detection methodology that is centered on analyzing fourth-order joint moments (co-kurtosis), particularly focusing on its application in auto-ignition of combustion problems with large numbers of species. Unsupervised anomaly detection is challenging to generalize across problem types and domains. A recent technique, centered on analyzing information in the fourth-order joint moment co-kurtosis, has shown promise, especially for high-dimensional scientific data. In this work we present developments to the co-kurtosis based anomaly detection method needed to make it effective and scalable for large-scale distributed scientific data, such as those generated by massively parallel simulations. An in situ co-kurtosis algorithm is employed as the anomaly detection method for identifying ignition kernels in simulations of turbulent combustion. Here, we extend an existing methodology which identifies regions of the domain where anomalies are present, and add another tier of anomaly detection where the individual samples contributing to the anomaly are identified. We apply this algorithm on-the-fly to a variety of turbulent reacting flow problems and compare it to the widely used (but significantly more expensive) chemical explosive mode analysis (CEMA). We demonstrate the ability of the method to detect and identify the onset of low and high temperature ignition which can be used for computational steering, as chemical and combustion anomalies occur intermittently at spatio-temporal locations unknown a priori. Finally, we apply our lightweight in situ algorithm to an exascale high-fidelity simulation with a total of 2.4 Trillion degrees of freedom, performed using an adaptive mesh refinement solver. Furthermore, through a scalability analysis, we show that the relative computational cost of this in-situ anomaly detection algorithm compared to an iteration of the reacting flow solver is negligible.

97 MATHEMATICS AND COMPUTING↗

SPARC-X: Quantum simulations at extreme scale - reactive dynamics from first principles

We have developed the massively parallel electronic structure code SPARC-X: a computational framework for performing Kohn-Sham Density Functional Theory (DFT) calculations that can scale linearly with the number of atoms in the system, while being able to leverage petascale and emerging exascale parallel computers to study chemical phenomena at unprecedented length and time scales. SPARC-X exploits a recent breakthrough in electronic structure methodologies: systematically improvable, strictly local, orthonormal, discontinuous real-space bases that efficiently and systematically capture the local chemistry of the system. With further adaptation using new machine-learning techniques and the use of the massively parallel Spectral Quadrature (SQ) electronic structure method, the algorithmic complexity and prefactor associated with DFT calculations involving semilocal as well as hybrid functionals are dramatically reduced. Using petascale computational resources, SPARC-X enables quantum mechanical simulations at length and time scales previously accessible only by empirical approaches, e.g., 1,000,000 atoms for a few picoseconds using semilocal functionals or 1,000 atoms for a few picoseconds using hybrid functionals. Using exascale resources, the sizes and times targeted are two orders of magnitude larger. Such a capability has applications in a wide variety of chemical sciences, including reactive interfaces where large length- and/or long time-scales are needed and traditional force fields fail. This is particularly important in dynamic catalysis, where bond breaking and formation must be understood in detail. We developed, tested, and employed the SPARC-X framework to understand the photocatalytic properties of TiO 2 nanoparticles, revealing finite size effects that cannot be captured with standard model systems or functionals. This integrated development and application strategy ensures that SPARC-X remains a robust, efficient, and scalable software package for quantum simulations on current petascale and emerging exascale computing resources.

97 MATHEMATICS AND COMPUTING↗

Massively parallel phase-field simulations targeting exascale

The interface thickness in the phase-field (PF) method limits its simulation scales. Consequently, large-scale PF simulations become prohibitively expensive for resolving the extremely fine microstructures that typically form during rapid solidification processing. This challenge is significant in predicting microstructure evolution in metal additive manufacturing and has been identified by the United States Department of Energy’s Exascale Computing Project. Here, to address this, we develop a multi-GPU and MPI-based massively parallel simulation code, utilizing state-of-the-art algorithms, software, and libraries, for large-scale three-dimensional (3D) PF simulations. We report the first GPU-parallel PF simulations on Frontier (currently the second TOP500 exascale cluster) and Summit machines, taking dendritic growth as an example problem. We evaluate the parallel performance of our implementation using scaling studies with more than 24 000 GPUs (among the largest known computations to date) and the acceleration performance using large-scale simulations of dendritic growth in 3D. Finally, massively parallel GPUs in these supercomputers enabled the first coupled multiscale simulations of laser melting and subsequent dendritic solidification on the scale of a full melt-pool, demonstrating the feasibility of performing PF simulations with a point total over 2 billion grid points within an acceptable time.

Exascale↗

Distributed memory, GPU accelerated Fock construction for hybrid, Gaussian basis density functional theory

With the growing reliance of modern supercomputers on accelerator-based architecture such a graphics processing units (GPUs), the development and optimization of electronic structure methods to exploit these massively parallel resources has become a recent priority. While significant strides have been made in the development GPU accelerated, distributed memory algorithms for many modern electronic structure methods, the primary focus of GPU development for Gaussian basis atomic orbital methods has been for shared memory systems with only a handful of examples pursing massive parallelism. Here in this work, we present a set of distributed memory algorithms for the evaluation of the Coulomb and exact exchange matrices for hybrid Kohn–Sham DFT with Gaussian basis sets via direct density-fitted (DF-J-Engine) and seminumerical (sn-K) methods, respectively. The absolute performance and strong scalability of the developed methods are demonstrated on systems ranging from a few hundred to over one thousand atoms using up to 128 NVIDIA A100 GPUs on the Perlmutter supercomputer.

97 MATHEMATICS AND COMPUTING↗

Novel Solver Algorithms for Nearly Singular Linear Systems Arising in Combustion Modelling

Direct Numerical Simulations of realistic combustion devices are extremely challenging due to the wide separation of scales in the simulation, for example an internal combustion (IC) engine chamber, and the flame thickness of a high-pressure flame. The PeleLMeX solver uses adaptive mesh refinement (AMR) to evolve multi-species reacting flows in the low Mach number limit at the Exascale and relies on an embedded boundary (EB) approach to represent complex geometries. In that framework, the EB geometries often give rise to very small cut-cells along the boundary, which translate into extreme ill-conditioning of the pressure-projection, with eigenvalues that span 15-16 orders of magnitude. In this talk, we focus on the case of a typical IC piston bowl geometry for which we present on a novel approach towards solving these nearly singular linear systems with ILU-based, C-AMG smoothers on massively parallel architectures. In particular, we use scaling and equilibration algorithms to handle the non-normality of the upper triangular factors. This enables us to approximate the highly sequential triangular solve algorithm, embedded in the AMG smoothing-solve phase, with Jacobi iterations. This approximation can be written as a convergent Neumann series whose terms are composed of highly parallel sparse matrix vector multiplications. The result is an algorithm that substantially decreases setup and solve time, compared to state-of-the-art, for these challenging linear systems.

combustion modelling↗

Hardware acceleration for HPS algorithms in two and three dimensions

We provide a flexible, open-source framework for hardware acceleration, namely massively-parallel execution on general-purpose graphics processing units (GPUs), applied to the hierarchical Poincaré–Steklov (HPS) family of algorithms for building fast direct solvers for linear elliptic partial differential equations. To take full advantage of the power of hardware acceleration, we propose two variants of HPS algorithms to improve performance on two- and three-dimensional problems. In the two-dimensional setting, we introduce a novel recomputation strategy that minimizes costly data transfers to and from the GPU; in three dimensions, we modify and extend the adaptive discretization technique of Geldermans and Gillman [1] to greatly reduce peak memory usage. We provide an open-source implementation of these methods written in JAX, a high-level accelerated linear algebra package, which allows for the first integration of a high-order fast direct solver with automatic differentiation tools. We conclude with extensive numerical examples showing our methods are fast and accurate on two- and three-dimensional problems.

Fast direct solvers↗

Streaming Matching and Edge Cover in Practice

Graph algorithms with polynomial space and time requirements often become infeasible for massive graphs with billions of edges or more. State-of-the-art approaches therefore employ approximate serial, parallel, and distributed algorithms to tackle these challenges. However, such approaches require storing the entire graph in memory and thus need access to costly computing resources such as clusters and supercomputers. In this paper, we present practical streaming approaches for solving massive graph problems using limited memory for two prototypical graph problems: maximum weighted matching and minimum weighted edge cover. For matching, we conduct a thorough computational study on two of the semi-streaming algorithms including a recent breakthrough result that achieves a $1/(2+\varepsilon)$-approximation of the weight while using $O( n \log W /\epsilon)$ memory (here $n$ is the number of vertices and $W$ is the maximum edge weight), designed by Paz and Schwartzman [SODA, 2017]. Empirically, we show that the semi-streaming algorithms produce matchings whose weight is close to the best $1/2$-approximate offline algorithm while requiring less time and an order-of-magnitude less memory. For minimum weighted edge cover, we develop three novel semi-streaming algorithms. Two of these algorithms require a single pass through the input graph, require $O(n \log n)$ memory, and provide a 2-approximation guarantee on the objective. We also leverage a relationship between approximate maximum weighted matching and approximate minimum weighted edge cover to develop a two-pass $3/2+\epsilon$-approximate algorithm with the memory requirement of Paz and Schwartzman's semi-streaming matching algorithm. These streaming approaches are compared against the state-of-the-art 3/2-approximate offline algorithm. The semi-streaming matching and the novel edge cover algorithms proposed in this paper can process graphs with several billions of edges in under 30 minutes using 6 GB of memory, which is at least an order of magnitude improvement from the offline (non-streaming) algorithms. For the largest graph, the best alternative offline parallel approximation algorithm (GPA+ROMA) could not finish in three hours even while employing hundreds of processors and 1 TB of memory. We also demonstrate an application of the semi-streaming algorithm by computing a matching using linearly bounded memory on item intersection graphs derived from three machine learning datasets, whereas the existing offline algorithms could not complete on one of these datasets since their memory requirements exceeded 1TB.

Ferdous, S M.↗

Programming approaches for scalability, performance, and portability of combustion physics codes

Here, this paper presents the process, strategy, and results associated with porting a typical combustion physics flow solver to current state-of-the-art and future massively-parallel computer architectures. Major focus is placed on the distinct algorithmic structure of these types of codes and how it can be integrated with modern programming paradigms for heterogeneous platforms (i.e., distributed many-core systems with accelerators). An end-to-end case study is presented that exemplifies the process in a generic manner, which then serves as a clear guide with respect to the strategy and best practices leading to a robust and adaptable framework that performs well, is durable over time, is portable, and requires minimal human-effort. This end is accomplished beginning with the use of a mature, validated, structured, multiblock code framework optimized for application of both Large Eddy Simulation (LES) and Direct Numerical Simulation (DNS). This code has been ported to a variety of platforms over the past decade, including most recently the Oak Ridge Leadership Computing Facility’s “Summit” Platform. The experience gained on these multiple platforms provides general insights and thus the results presented are not specific to any one code or platform other than the overarching trend toward distributed many-core systems with accelerators in order to move toward exascale performance. The resultant performance and scalability of the ported code is demonstrated on a real-world application; a state-of-the-art rotating detonation rocket engine simulation that matches the complex geometry and boundary conditions imposed as part of a companion experimental campaign.

97 MATHEMATICS AND COMPUTING↗