Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel and distributed computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

A Scalable Parallel Hypergraph Generator (HyGen)

Graphs are extensively used to model real-world complex systems. An edge in a graph can model pairwise relationships. However, multiway relationships (connections between three or more vertices) are common in many complex systems such as cellular process, image segmentation, and circuit design. A graph edge cannot model multiway relationships. A hypergraph, which can connect more than two vertices, is thus a better option to model multiway relationships. A large-scale hypergraph analysis has the potential to find useful insights from a complex system and assist in knowledge discovery. Currently a limited number of hypergraphs exists that are representative of real-world datasets. Moreover, real-world hypergraph datasets are small in size and inadequate to incorporate future needs. A graph generator that can produce large-scale synthetic hypergraphs can solve the above mentioned problems. In this paper, we present a scalable parallel hypergraph generator (HyGen) based on the Message Passing Interface (MPI) standard. To generate hypergraphs, HyGen takes the following parameter values as inputs: i) number of vertices, ii) number of hyperedges, iii) number of clusters, iv) vertex distribution, v) hyperedge distribution, vi) local cluster cardinality, and vii) global cluster cardinality. We have demonstrated that HyGen can generate hypergraphs of various sizes in a scalable fashion. HyGen takes approximately four minutes to generate a hypergraph with 4.8 million vertices, 1.6 million hyperedges, and 800 clusters using 1,024 processes on a leadership class computing platform. Our strong and weak scaling experiments on supercomputers demonstrate that HyGen can quickly create large-scale hypergraphs in a parallel manner, thus providing a useful capability for hypergraph analysis.

Hasan, S M Shamimul↗

An Adaptive-Mesh-Refinement Based Computational Tool for Simulating Catalysis at Mesoscale

In this work, we present a computational tool for mesoscale applications using open-source exascale- computing compatible adaptive-mesh-refinement (AMR) library, AMReX [2]. AMReX is software library that enables development of application solvers with block-structured Cartesian AMR. Our tool has capabilities to include realistic geometry representation, chemical species transport, reactions and thermodynamics that are critical for capturing mesoscale physics. A significant achievement is the ability of our solver to automatically import electron microscopy data in the form of a stereolithography (STL) or pixelated file format (mrc, tiff) without undergoing the tedious task of unstructured mesh generation. This feature allows for rapid simulation of catalyst particles with complex morphologies using an immersed-boundary formulation. The use of AMR allows for higher resolutions at catalyst surface interfaces, which in turn provides an accurate description of surface reactions and transport. Our solver uses a hybrid distributed and shared memory parallelism (OpenMP/GPU-based) with which strong scaling up to 10,000 processors for realistic catalyst particle simulations have been demonstrated.

BIOMASS FUELS,MATHEMATICS AND COMPUTING↗

Multiscale aperture synthesis imager

Synthetic aperture imaging has enabled breakthrough observations from radar to astronomy. However, optical implementation remains challenging due to stringent wavefield synchronization requirements among multiple receivers. Here we present the multiscale aperture synthesis imager (MASI), which utilizes parallelism to break complex optical challenges into tractable sub-problems. MASI employs a distributed array of coded sensors that operate independently yet coherently to surpass the diffraction limit of single receiver. It combines the propagated wavefields from individual sensors through a computational phase synchronization scheme, eliminating the need for overlapping measurement regions to establish phase coherence. Light diffraction in MASI naturally expands the imaging field, generating phase-contrast visualizations that are substantially larger than sensor dimensions. Without using lenses, MASI resolves sub-micron features at ultralong working distances and reconstructs 3D shapes over centimeter-scale fields. MASI transforms the intractable optical synchronization problem into a computational one, enabling practical deployment of scalable synthetic aperture systems at optical wavelengths.

electrical and electronic engineering↗

Optimizing temperature distributions for training neural quantum states using parallel tempering

Parametrized artificial neural networks (ANNs) can be very expressive ansatzes for variational algorithms, reaching state-of-the-art energies on many quantum many-body Hamiltonians. Nevertheless, the training of the ANN can be slow and stymied by the presence of local minima in the parameter landscape. One approach to mitigate this issue is to use parallel tempering methods, and in this work, we focus on the role played by the temperature distribution of the parallel tempering replicas. Using an adaptive method that adjusts the temperatures in order to equate the exchange probability between neighboring replicas, we show that this temperature optimization can significantly increase the success rate of the variational algorithm with negligible computational cost by eliminating bottlenecks in the replicas' random walk. Furthermore, we demonstrate this using two different neural networks, a restricted Boltzmann machine and a feedforward network, which we use to study a toy problem based on a permutation invariant Hamiltonian with a pernicious local minimum and the 𝐽 1 −𝐽 2 model on a rectangular lattice.

Neural network simulations↗

A parallel p ‐adaptive discontinuous Galerkin method for the Euler equations with dynamic load‐balancing on tetrahedral grids

Abstract A novel p ‐adaptive discontinuous Galerkin (DG) method has been developed to solve the Euler equations on three‐dimensional tetrahedral grids. Hierarchical orthogonal basis functions are adopted for the DG spatial discretization while a third order TVD Runge‐Kutta method is used for the time integration. A vertex‐based limiter is applied to the numerical solution in order to eliminate oscillations in the high order method. An error indicator constructed from the solution of order and is used to adapt degrees of freedom in each computational element, which remarkably reduces the computational cost while still maintaining an accurate solution. The developed method is implemented with under the Charm++ parallel computing framework. Charm++ is a parallel computing framework that includes various load‐balancing strategies. Implementing the numerical solver under Charm++ system provides us with access to a suite of dynamic load balancing strategies. This can be efficiently used to alleviate the load imbalances created by p ‐adaptation. A number of numerical experiments are performed to demonstrate both the numerical accuracy and parallel performance of the developed p ‐adaptive DG method. It is observed that the unbalanced load distribution caused by the parallel p ‐adaptive DG method can be alleviated by the dynamic load balancing from Charm++ system. Due to this, high performance gain can be achieved. For the testcases studied in the current work, the parallel performance gain ranged from 1.5× to 3.7×. Therefore, the developed p ‐adaptive DG method can significantly reduce the total simulation time in comparison to the standard DG method without p ‐adaptation.

97 MATHEMATICS AND COMPUTING↗

Impact of Different Thermal Gradients on the Dynamics of Cylindrical Lithium-ion Cells Subject to Accelerated Aging and on Module Performance

This study investigates the impacts of applying different thermal gradient patterns to cylindrical lithium-ion cells in a module on cell dynamics (temperatures, current flows, state of charge), module performance (evolution of resistance, capacity, and energy versus cycle number), and module lifetime. The thermal gradients were generated using cooling plates (CPs) with three different flow-field designs, namely, straight, perpendicular, and U-turn. The study uses computational fluid dynamics (CFD), the pseudo-two-dimensional (P2D) battery model, capacity loss and increased impedance due to the growth of a solid-electrolyte-interphase, and the electric current distribution from module terminals to cells that depends on the series-parallel electrical connections among the cells. The impact of the thermal gradient (resulting from the CP designs) on the variability in resistance, current, state of charge, and voltage among the cells was analyzed and linked to differences in the module's performance. Applying a thermal gradient to parallel-connected strings of series-connected cells led to variation in the current through each parallel string and an imbalance in the voltage of series-connected cells. Module performance is poorer when the thermal gradient causes a voltage imbalance than when it causes a current imbalance. Module performance becomes the worst when both current variation and voltage imbalance happen together. For instance, the module's lifetime (estimated as reaching 80% of its initial capacity) varied by 5% to 17.5%, depending on the magnitude and pattern of the imposed thermal gradient. As the relative orientation between thermal gradients and cells' electrical connectivity influences the module's performance, appropriate consideration should be given to the choice of the CP, especially if large thermal gradients are allowed.

Battery thermal management↗

Scalable self attraction and loading calculations for unstructured ocean tide models

Self attraction and earth-loading effects are important for accurately modeling global tides. A common approach of handling this forcing is to expand mass anomalies into spherical harmonics, which are scaled by load Love numbers to account for elastic earth deformation. We investigate two different approaches to perform these calculations for ocean models that employ unstructured meshes and distributed memory parallelization. The first approach leverages a highly efficient spherical harmonics library, but requires all-to-one and one-to-all communications and interpolation operations between the unstructured and a structured mesh. This approach is compared to a parallel algorithm that computes the spherical harmonic transformations directly on the unstructured mesh with an all-reduce communication. Here, our results show that although the unstructured mesh calculations are more expensive, the scalability of the unstructured mesh approach allows for more efficient spherical harmonics transforms for high-resolution meshes and large processor counts. This methodology enables the efficient inclusion of tidal dynamics large-scale Earth system model simulations.

54 ENVIRONMENTAL SCIENCES↗

Characterization and Optimization of the Fitting of Quantum Correlation Functions

This case study presents a characterization and optimization of an application code for extracting parton distribution functions from high energy electron-proton scattering data. Profiling this application code reveals that the phase-space density computation accounts for 93% of the overall execution time for a single iteration on a single core. When executing multiple iterations in parallel on a multicore system, the application spends 78% of its overall execution time idling due to load imbalance. We address these issues by first transforming the application code from Python to C++ and then tackling the application load imbalance via a hybrid scheduling strategy that combines dynamic and static scheduling. These techniques result in a 62% reduction in CPU idle time and a 2.46x speedup in overall execution time per node. In addition, the typically enabled power-management mechanisms in supercomputers (e.g., AMD Turbo Core, Intel Turbo Boost, and RAPL) can significantly impact intra-node scalability when more than 50% of the CPU cores are used. This finding underscores the importance of understanding system interactions with power management, as they can adversely impact application performance, and highlights the necessity of intra-node scaling tests to identify performance degradation that inter-node scaling tests might otherwise overlook.

Chuang, Pi-Yueh [Virginia Tech,Dept. of Computer S↗

Improving I/O Performance for Exascale Applications through Online Data Layout Reorganization

The applications being developed within the U.S. Exascale Computing Project (ECP) to run on imminent Exascale computers will generate scientific results with unprecedented fidelity and record turn-around time. Many of these codes are based on particle-mesh methods and use advanced algorithms, especially dynamic load-balancing and mesh-refinement, to achieve high performance on Exascale machines. Yet, as such algorithms improve parallel application efficiency, they raise new challenges for I/O logic due to their irregular and dynamic data distributions. Thus, while the enormous data rates of Exascale simulations already challenge existing file system write strategies, the need for efficient read and processing of generated data introduces additional constraints on the data layout strategies that can be used when writing data to secondary storage. We review these I/O challenges and introduce two online data layout reorganization approaches for achieving good tradeoffs between read and write performance. We demonstrate the benefits of using these two approaches for the ECP particle-in-cell simulation WarpX, which serves as a motif for a large class of important Exascale applications. Here, we show that by understanding application I/O patterns and carefully designing data layouts we can increase read performance by more than 80 percent.

97 MATHEMATICS AND COMPUTING↗

Enabling machine learning-ready HPC ensembles with Merlin

With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. Here, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. As a producer–consumer workflow model, Merlin enables multi-machine, cross-batch job, dynamically allocated yet persistent workflows capable of utilizing surge-compute resources. Key features of Merlin are a flexible HPC-centric interface, low per-task overhead, multi-tiered fault recovery, and a hierarchical sampling algorithm that allows for $\mathscr{O}$(N) task execution and $\mathscr{O}$(N ln N) task queuing to ensembles of millions of tasks. In addition to Merlin’s design, we test the algorithm’s performance in an HPC center and demonstrate the ability to enqueue 40 million simulations in 100 s, with a 30 millisecond per-task overhead that is independent of ensemble size. Finally, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.

97 MATHEMATICS AND COMPUTING↗

ARENA: Asynchronous Reconfigurable Accelerator Ring to Enable Data-Centric Parallel Computing

The next generation HPC and data centers are likely to be reconfigurable and data-centric due to the trend of hardware specialization and the emergence of data-driven applications. In this work, we propose ARENA – an asynchronous reconfigurable accelerator ring architecture as a potential scenario on how the future HPC and data centers will be like. Despite using the coarse-grained reconfigurable arrays (CGRAs) as the substrate platform, our key contribution is not only the CGRA-cluster design itself, but also the ensemble of a new architecture and programming model that enables asynchronous tasking across a cluster of reconfigurable nodes, so as to bring specialized computation to the data rather than the reverse. We presume distributed data storage without asserting any prior knowledge on the data distribution. Hardware specialization occurs at runtime when a task finds the majority of data it requires are available at the present node. In other words, we dynamically generate specialized CGRA accelerators where the data reside. The asynchronous tasking for bringing computation to data is achieved by circulating the task token, which describes the dataflow graphs to be executed for a task, among the CGRA cluster connected by a fast ring network. Evaluations on a set of HPC and data-driven applications across different domains show that ARENA can provide better parallel scalability with reduced data movement (53.9 percent). Compared with contemporary compute-centric parallel models, ARENA can bring on average 4.37× speedup. The synthesized CGRAs and their task-dispatchers only occupy 2.93mm 2 chip area under 45nm process technology and can run at 800MHz with on average 759.8mW power consumption. ARENA also supports the concurrent execution of multi-applications, offering ideal architectural support for future high-performance parallel computing and data analytics systems.

97 MATHEMATICS AND COMPUTING↗

Particlization in fluid dynamical simulations of heavy-ion collisions: The IS3D module

The IS3D particlization module simulates the emission of hadrons from heavy-ion collisions via Monte-Carlo sampling of the Cooper–Frye formula which converts fluid dynamical information into local phase-space distributions for hadrons. The code package includes multiple choices for the non-equilibrium correction to these distribution functions: the 14-moment approximation, first-order Chapman–Enskog expansion, and two types of modified equilibrium distributions. This makes it possible to explore to what extent heavy-ion experimental data are sensitive to different choices for δf n , presently the main source of theoretical uncertainty in the particlization stage. Here, we validate our particle sampler with a high degree of precision by generating several million hadron emission events from a longitudinally boost-invariant hypersurface and comparing the event-averaged particle spectra and space–time distributions to the Cooper–Frye formula.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

The impact of non-local parallel electron transport on plasma-impurity reaction rates in tokamak scrape-off layer plasmas

Abstract Plasma-impurity reaction rates are a crucial part of modelling tokamak scrape-off layer (SOL) plasmas. To avoid calculating the full set of rates for the large number of important processes involved, a set of effective rates are typically derived which assume Maxwellian electrons. However, non-local parallel electron transport may result in non-Maxwellian electrons, particularly close to divertor targets. Here, the validity of using Maxwellian-averaged rates in this context is investigated by computing the full set of rate equations for a fixed plasma background from kinetic and fluid SOL simulations. We consider the effect of the electron distribution as well as the impact of the electron transport model on plasma profiles. Results are presented for lithium, beryllium, carbon, nitrogen, neon and argon. It is found that electron distributions with enhanced high-energy tails can result in significant modifications to the ionisation balance and radiative power loss rates from excitation, on the order of 50%–75% for the latter. Fluid electron models with Spitzer-Härm or flux-limited Spitzer-Härm thermal conductivity, combined with Maxwellian electrons for rate calculations, can increase or decrease this error, depending on the impurity species and plasma conditions. Based on these results, we also discuss some approaches to experimentally observing non-local electron transport in SOL plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Model-parallel Fourier neural operators as learned surrogates for large-scale parametric PDEs

Fourier neural operators (FNOs) are a recently introduced neural network architecture for learning solution operators of partial differential equations (PDEs), which have been shown to perform significantly better than comparable deep learning approaches. Once trained, FNOs can achieve speed-ups of multiple orders of magnitude over conventional numerical PDE solvers. However, due to the high dimensionality of their input data and network weights, FNOs have so far only been applied to two-dimensional or small three-dimensional problems. To remove this limited problem-size barrier, we propose a model-parallel version of FNOs based on domain-decomposition of both the input data and network weights. Here, we demonstrate that our model-parallel FNO is able to predict time-varying PDE solutions of over 2.6 billion variables on Perlmutter using up to 512 A100 GPUs and show an example of training a distributed FNO on the Azure cloud for simulating multiphase CO 2 dynamics in the Earth’s subsurface.

58 GEOSCIENCES↗

Current possibilities and future opportunities for erasure coded computations

The key capability established through the research funded by this award are erasure coded computations for linear systems, in serial and in parallel. This capability enables powerful efficient and scalable alternatives to existing linear system solvers in fault-prone computational systems.

97 MATHEMATICS AND COMPUTING↗

The VTK-m Users' Guide (V.2.0)

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created VTK-m: the visualization toolkit for multi-/many-core architectures. VTK-m supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. VTK-m also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although VTK-m provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction. VTK-m simplifies the development of parallel scientific visualization algorithms by providing a framework of supporting functionality that allows developers to focus on visualization operations. Consider the listings in Figure 1.1 that compares the size of the implementation for the Marching Cubes algorithm in VTK-m with the equivalent reference implementation in the CUDA software development kit. Because VTK-m internally manages the parallel distribution of work and data, the VTK-m implementation is shorter and easier to maintain. Additionally, VTK-m provides data abstractions not provided by other libraries that make code written in VTK-m more versatile.

97 MATHEMATICS AND COMPUTING↗

Dynamic Phasor Modeling of Three Phase Voltage Source Inverters

With the increase in the development and implementation of distributed energy resources, application of parallel connected inverters is increasing and as a result having accurate modeling and simulation tools that can help in the design, analysis and stability assessment of the grid is of great importance. Development of fundamental methods that can achieve accurate, reliable, and computationally efficient results can be very beneficial. Application of Dynamic phasor (DP) modeling method has been limited to study of limited harmonics and small combination of interconnected converters due to the complexity associated with developing models that describe larger systems. In this paper, application of DP modeling method is expanded to model any number of parallel connected three phase voltage source inverter (VSI) with inclusion of a wider harmonic content including fundamental, subharmonics, inter-harmonics, switching frequency and their sidebands. Results achieved from this modeling method is compared with conventional average model as well as detailed switching model and the effect of inclusion of wider harmonic content on accuracy of DP modeling method is demonstrated.

Xue, Yaosuo↗