Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed memory”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Computing material volume fractions on a superimposed mesh as applied to Monte Carlo particle transport simulations

Here, we present a newly implemented ray tracing algorithm in OpenMC for efficiently computing material volume fractions on superimposed meshes in complex geometries. By firing rays along each coordinate direction through the geometry, the approach accumulates track-length data in each mesh element, thereby determining the fractional composition of each material. Scaling studies on three different models—a random tetrahedra configuration, the Frascati Neutron Generator ITER dose rate benchmark, and a stellarator design—show excellent parallel performance, with nearly linear speedup on modern multi-threaded and distributed-memory systems. An analysis of the residual error relative to high-resolution reference solutions demonstrated that under optimal conditions it decreases as 1/R, where R is the number of rays fired, making it straightforward to achieve user-prescribed accuracy. This new functionality enables practical, mesh-based approaches for detailed nuclear analyses in production Monte Carlo workflows without resorting to expensive, fully conformal or unstructured meshing.

Monte Carlo↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗

Development of a CPU/GPU portable software library for Lagrangian–Eulerian simulations of liquid sprays

The Lagrangian–Eulerian method is widely used for simulations of fuel sprays in turbulent combustion because of the advantage of treating the spray droplets as discrete points. One challenge of the Lagrangian–Eulerian method is the intense computational requirement when tracking the large number of Lagrangian particles needed for high fidelity. We have developed a performance-portable library, Grit, to track the Lagrangian particles in parallel on central processing unit (CPU) and graphics processing unit (GPU) accelerated high performance computing (HPC) architectures. Grit is a C++ library which employs Message Passing Interface (MPI) for distributed memory parallelism and Kokkos programming model for on-node shared memory parallelism with performance portability across different architectures of GPUs and multi-core/manycore CPUs. The parallel algorithms, key parallel kernels, and their performances on the pre-exascale supercomputer, Summit, are presented. Grit is coupled with a direct numerical simulation (DNS) solver, S3D, for multiphase simulations. A conservative formulation has been developed and implemented in Grit for phase coupling with thermodynamic consistency. The formulation separates the conservation of mass, momentum and energy from the physical models to prevent accidental violation of conservation laws due to inconsistent models. The formulation also enforces consistent definitions of enthalpies of the fuel for both phases and the latent heat of evaporation. Finally, simulations of turbulent particle-laden flow and the evaporation of dilute turbulent spray jet are performed to verify the software implementation, and to demonstrate the scalability of Grit for large-scale multiphase simulations.

97 MATHEMATICS AND COMPUTING↗

Spectral quadrature for the first principles study of crystal defects: Application to magnesium

In this work, we present an accurate and efficient finite-difference formulation and parallel implementation of Kohn-Sham Density (Operator) Functional Theory (DFT) for non periodic systems embedded in a bulk environment. Specifically, employing non-local pseudopotentials, local reformulation of electrostatics, and truncation of the spatial Kohn-Sham Hamiltonian, and the Linear Scaling Spectral Quadrature method to solve for the pointwise electronic fields in real-space and the non-local component of the atomic force, we develop a parallel finite difference framework suitable for distributed memory computing architectures to simulate non-periodic systems embedded in a bulk environment. Choosing examples from magnesium-aluminum alloys, we first demonstrate the convergence of energies and forces with respect to spectral quadrature polynomial order, and the width of the spatially truncated Hamiltonian. Next, we demonstrate the parallel scaling of our framework, and show that the computation time and memory scale linearly with respect to the number of atoms. Next, we use the developed framework to simulate isolated point defects and their interactions in magnesium-aluminum alloys. Our findings conclude that the binding energies of divacancies, Al solute-vacancy and two Al solute atoms are anisotropic and are dependent on cell size. Furthermore, the binding is favorable in all three cases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Scalable Implicit Solvers with Dynamic Mesh Adaptation for a Relativistic Drift-Kinetic Fokker–Planck–Boltzmann Model

In this work we consider a relativistic drift-kinetic model for runaway electrons along with a Fokker–Planck operator for small-angle Coulomb collisions, a radiation damping operator, and a secondary knock-on (Boltzmann) collision source. Here, we develop a new scalable fully implicit solver utilizing finite volume and conservative finite difference schemes and dynamic mesh adaptivity. A new data management framework in the PETSc library based on the p4est library is developed to enable simulations with dynamic adaptive mesh refinement (AMR), distributed memory parallelization, and dynamic load balancing of computational work. This framework and the runaway electron solver building on the framework are able to dynamically capture both bulk Maxwellian at the low-energy region and a runaway tail at the high-energy region. To effectively capture features via the AMR algorithm, a new AMR indicator prediction strategy is proposed that is performed alongside the implicit time evolution of the solution. This strategy is complemented by the introduction of computationally cheap feature-based AMR indicators that are analyzed theoretically. Numerical results quantify the advantages of the prediction strategy in better capturing features compared with nonpredictive strategies; and we demonstrate trade-offs regarding computational costs. The robustness with respect to model parameters, algorithmic scalability, and parallel scalability are demonstrated through several benchmark problems including manufactured solutions and solutions of different physics models. We focus on demonstrating the advantages of using implicit time stepping and AMR for runaway electron simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Demonstrating the viability of Lagrangian in situ reduction on supercomputers

Performing exploratory analysis and visualization of large-scale time-varying computational science applications is challenging due to inaccuracies that arise from under-resolved data. In recent years, Lagrangian representations of the vector field computed using in situ processing are being increasingly researched and have emerged as a potential solution to enable exploration. However, prior works have offered limited estimates of the encumbrance on the simulation code as they consider “theoretical” in situ environments. Further, the effectiveness of this approach varies based on the nature of the vector field, benefitting from an in-depth investigation for each application area. With this study, an extended version of Sane et al. (2021), we contribute an evaluation of Lagrangian analysis viability and efficacy for simulation codes executing at scale on a supercomputer. We investigated previously unexplored cosmology and seismology applications as well as conducted a performance benchmarking study by using a hydrodynamics mini-application targeting exascale computing. Here, to inform encumbrance, we integrated in situ infrastructure with simulation codes, and evaluated Lagrangian in situ reduction in representative homogeneous and heterogeneous HPC environments. To inform post hoc accuracy, we conducted a statistical analysis across a range of spatiotemporal configurations as well as a qualitative evaluation. Additionally, our study contributes cost estimates for distributed-memory post hoc reconstruction. In all, we demonstrate viability for each application — data reduction to less than 1% of the total data via Lagrangian representations, while maintaining accurate reconstruction and requiring under 10% of total execution time in over 90% of our experiments.

97 MATHEMATICS AND COMPUTING↗

Online multimedia retrieval on CPU–GPU platforms with adaptive work partition

Nearest neighbors search is a core operation found in several online multimedia services. These services have to handle very large databases, while, at the same time, they must minimize the query response times observed by users. This is specially complex because those services deal with fluctuating query workloads (rates). Consequently, they must adapt at run-time to minimize the response times as the load varies. In this paper, we address the aforementioned challenges with a distributed memory parallelization of the product quantization nearest neighbor search, also known as IVFADC, for hybrid CPU–GPU machines. Overall, our parallel IVFADC implements an out-of-GPU memory execution scheme to use the GPU for databases in which the index does not fit in its memory, which is crucial for searching in very large databases. The careful use of CPU and GPU with work stealing led to an average response time reduction of 2.4 as compared to using the GPU only. Also, our approach to adapt the system to fluctuating loads, called Dynamic Query Processing Policy (DQPP), attained a response time reduction of up to 5 vs. the best static (BS) policy for moderate loads. The system has attained high query processing rates and near-linear scalability in all experiments. We have evaluated our system on a machine with up to 256 NVIDIA V100 GPUs processing a database of 256 billion SIFT features vectors.

97 MATHEMATICS AND COMPUTING↗

Massively-parallel Lagrangian particle code and applications

Massively-parallel, distributed-memory algorithms for the Lagrangian particle hydrodynamic method (Samulyak et al., 2018) have been developed, verified, and implemented. The key component of parallel algorithms is a particle management module that includes a parallel construction of octree databases, dynamic adaptation and refinement of octrees, and particle migration between parallel subdomains. The particle management module is based on the p4est (parallel forest of k-trees) library. The massively-parallel Lagrangian particle code has been applied to a variety of fundamental science and applied problems. A summary of Lagrangian particle code applications to the injection of impurities into thermonuclear fusion devices and to the simulation of supersonic hydrogen jets in support of laser-plasma wakefield acceleration research has also been presented.

97 MATHEMATICS AND COMPUTING↗

Scalable self attraction and loading calculations for unstructured ocean tide models

Self attraction and earth-loading effects are important for accurately modeling global tides. A common approach of handling this forcing is to expand mass anomalies into spherical harmonics, which are scaled by load Love numbers to account for elastic earth deformation. We investigate two different approaches to perform these calculations for ocean models that employ unstructured meshes and distributed memory parallelization. The first approach leverages a highly efficient spherical harmonics library, but requires all-to-one and one-to-all communications and interpolation operations between the unstructured and a structured mesh. This approach is compared to a parallel algorithm that computes the spherical harmonic transformations directly on the unstructured mesh with an all-reduce communication. Here, our results show that although the unstructured mesh calculations are more expensive, the scalability of the unstructured mesh approach allows for more efficient spherical harmonics transforms for high-resolution meshes and large processor counts. This methodology enables the efficient inclusion of tidal dynamics large-scale Earth system model simulations.

54 ENVIRONMENTAL SCIENCES↗

Accelerating error correction in tomographic reconstruction

Abstract Spurred by recent advances in detector technology and X-ray optics, upgrades to scanning-probe-based tomographic imaging have led to an exponential growth in the amount and complexity of experimental data and have created a clear opportunity for tomographic imaging to approach single-atom sensitivity. The improved spatial resolution, however, is highly susceptible to systematic and random experimental errors, such as center of rotation drifts, which may lead to imaging artifacts and prevent reliable data extraction. Here, we present a model-based approach that simultaneously optimizes the reconstructed specimen and sinogram alignment as a single optimization problem for tomographic reconstruction with center of rotation error correction. Our algorithm utilizes an adaptive regularizer that is dynamically adjusted at each alternating iteration step. Furthermore, we describe its implementation in a software package targeting high-throughput workflows for execution on distributed-memory clusters. We demonstrate the performance of our solver on large-scale synthetic problems and show that it is robust to a wide range of noise and experimental drifts with near-ideal throughput.

Ali, Sajid (ORCID:0000000321864636)↗

Massively Parallel Quantum Chemistry: A high-performance research platform for electronic structure

The Massively Parallel Quantum Chemistry (MPQC) program is a 30-year-old project that enables facile development of electronic structure methods for molecules for efficient deployment to massively parallel computing architectures. Here, we describe the historical evolution of MPQC’s design into its latest (fourth) version, the capabilities and modular architecture of today’s MPQC, and how MPQC facilitates rapid composition of new methods as well as its state-of-the-art performance on a variety of commodity and high-end distributed-memory computer platforms.

Peng, Chong↗

The PetscSF Scalable Communication Layer

PetscSF, the communication component of the Portable, Extensible Toolkit for Scientific Computation (PETSc), is designed to provide PETSc's communication infrastructure suitable for exascale computers that utilize GPUs and other accelerators. PetscSF provides a simple application programming interface (API) for managing common communication patterns in scientific computations by using a star-forest graph representation. PetscSF supports several implementations based on MPI and NVSHMEM, whose selection is based on the characteristics of the application or the target architecture. An efficient and portable model for network and intra-node communication is essential for implementing large-scale applications. The Message Passing Interface, which has been the de facto standard for distributed memory systems, has developed into a large complex API that does not yet provide high performance on the emerging heterogeneous CPU-GPU-based exascale systems. Here, we discuss the design of PetscSF, how it can overcome some difficulties of working directly with MPI on GPUs, and we demonstrate its performance, scalability, and novel features.

97 MATHEMATICS AND COMPUTING↗

Simulating the Impact of Dynamic Rerouting on Metropolitan-scale Traffic Systems

The rapid introduction of mobile navigation aides that use real-time road network information to suggest alternate routes to drivers is making it more difficult for researchers and government transportation agencies to understand and predict the dynamics of congested transportation systems. Computer simulation is a key capability for these organizations to analyze hypothetical scenarios; however, the complexity of transportation systems makes it challenging for them to simulate very large geographical regions, such as multi-city metropolitan areas. In this article, we describe enhancements to the Mobiliti parallel traffic simulator to model dynamic rerouting behavior with the addition of vehicle controller actors and vehicle-to-controller reroute requests. The simulator is designed to support distributed-memory parallel execution using discrete event simulation and be scalable on high-performance computing platforms. We demonstrate the potential of the simulator by analyzing the impact of varying the population penetration rate of dynamic rerouting on the San Francisco Bay Area road network. Using high-performance parallel computing, we can simulate a day in the San Francisco Bay Area with 19 million vehicle trips with 50 percent dynamic rerouting penetration over a road network with 0.5 million nodes and 1 million links in less than three minutes. We present a sensitivity study on the dynamic rerouting parameters, discuss the simulator’s parallel scalability, and analyze system-level impacts of changing the dynamic rerouting penetration. Furthermore, we examine the varying effects on different functional classes and geographical regions and present a validation of the simulation results compared to real-world data.

97 MATHEMATICS AND COMPUTING↗

Performance Analysis and Optimal Node-aware Communication for Enlarged Conjugate Gradient Methods

Krylov methods are a key way of solving large sparse linear systems of equations but suffer from poor strong scalability on distributed memory machines. Furthermore, this is due to high synchronization costs from large numbers of collective communication calls alongside a low computational workload. Enlarged Krylov methods address this issue by decreasing the total iterations to convergence, an artifact of splitting the initial residual and resulting in operations on block vectors. In this article, we present a performance study of an enlarged Krylov method, Enlarged Conjugate Gradients (ECG), noting the impact of block vectors on parallel performance at scale. Most notably, we observe the increased overhead of point-to-point communication as a result of denser messages in the sparse matrix-block vector multiplication kernel. Additionally, we present models to analyze expected performance of ECG, as well as motivate design decisions. Most importantly, we introduce a new point-to-point communication approach based on node-aware communication techniques that increases efficiency of the method at scale.

97 MATHEMATICS AND COMPUTING↗

Communication Lower Bounds and Optimal Algorithms for Symmetric Matrix Computations

In this article, we focus on the communication costs of three symmetric matrix computations: (i) multiplying a matrix with its transpose, known as a symmetric rank-k update (SYRK) (ii) adding the result of the multiplication of a matrix with the transpose of another matrix and the transpose of that result, known as a symmetric rank-2k update (SYR2K) (iii) performing matrix multiplication with a symmetric input matrix (SYMM). All three computations appear in the Level 3 Basic Linear Algebra Subroutines (BLAS) and have wide use in applications involving symmetric matrices. We establish communication lower bounds for these kernels using sequential and distributed-memory parallel computational models, and we show that our bounds are tight by presenting communication-optimal algorithms for each setting. Our lower bound proofs rely on applying a geometric inequality for symmetric computations and analytically solving constrained nonlinear optimization problems. As a result, the symmetric matrix and its corresponding computations are accessed and performed according to a triangular block partitioning scheme in the optimal algorithms.

Al Daas, Hussam [Rutherford Appleton Laboratory, D↗

PASTIS v0.1

PASTIS is a distributed-memory library for aligning large protein sequence sets and constructing protein similarity networks.

Buluc, Aydin↗

Reeber v2.0

Reeber is a library for shared- and distributed-memory parallel computation of merge trees

Morozov, Dmitriy↗