Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Parallel Runtime Interface for Fortran (PRIF) Design Document (Rev. 0.2)

This design document proposes an interface to support the parallel features of Fortran, named the Parallel Runtime Interface for Fortran (PRIF). PRIF is a proposed solution in which the runtime library is responsible for coarray allocation, deallocation and accesses, image synchronization, atomic operations, events, and teams. In this interface, the compiler is responsible for transforming the invocation of Fortran-level parallel features into procedure calls to the necessary PRIF procedures. The interface is designed for portability across shared- and distributed-memory machines, different operating systems, and multiple architectures. Implementations of this interface are intended as an augmentation for the compiler's own runtime library. With an implementation-agnostic interface, alternative parallel runtime libraries may be developed that support the same interface. One benefit of this approach is the ability to vary the communication substrate. A central aim of this document is to define a parallel runtime interface in standard Fortran syntax, which enables us to leverage Fortran to succinctly express various properties of the procedure interfaces, including argument attributes.

97 MATHEMATICS AND COMPUTING↗

Parallel Runtime Interface for Fortran (PRIF) Specification (Rev. 0.6)

This document specifies an interface to support the multi-image parallelism features of Fortran, named the Parallel Runtime Interface for Fortran (PRIF). PRIF is a solution in which a runtime library is primarily responsible for implementing coarray allocation, deallocation and accesses, image synchronization, atomic operations, events, teams and collective subroutines. The Fortran compiler is responsible for transforming the invocation of Fortran-level multi-image parallelism features into procedure calls to the necessary PRIF subroutines. The interface is designed for portability across shared- and distributed-memory machines, different operating systems, and multiple architectures. Implementations of this interface are intended as an augmentation for the compiler's own runtime library. With an implementation-agnostic interface, alternative parallel runtime libraries may be developed that support the same interface. One benefit of this approach is the ability to vary the communication substrate. A central aim of this document is to define a parallel runtime interface in standard Fortran syntax, which enables us to leverage Fortran to succinctly express various properties of the procedure interfaces, including argument attributes.

97 MATHEMATICS AND COMPUTING↗

Parallel Runtime Interface for Fortran (PRIF): A Multi-Image Solution for LLVM Flang

Fortran compilers that provide support for Fortran’s native parallel features often do so with a runtime library that depends on details of both the compiler implementation and the communication library, while others provide limited or no support at all. This paper introduces a new generalized interface that is both compiler- and runtime-library-agnostic, providing flexibility while fully supporting all of Fortran’s parallel features. The Parallel Runtime Interface for Fortran (PRIF) was developed to be portable across shared- and distributed-memory systems, with varying operating systems, toolchains and architectures. It achieves this by defining a set of Fortran procedures corresponding to each of the parallel features defined in the Fortran standard that may be invoked by a Fortran compiler and implemented by a runtime library. PRIF aims to be used as the solution for LLVM Flang to provide parallel Fortran support. This paper also briefly describes our PRIF prototype implementation: Caffeine.

Bonachea, Dan↗

Automatically parallelizing batch inference on deep neural networks using Fiats and Fortran 2023 `do concurrent`

This paper introduces novel programming strategies that leverage features of the Fortran 2023 standard of the International Standards Organization (ISO) to automatically parallelize computations on deep neural networks. The paper focuses on the interplay of object-oriented, parallel, and functional programming paradigms in the Fiats deep learning library. We demonstrate how several infrequently used language features play a role in enabling efficient, parallel execution. Specifically, the ability to explicitly declare that a procedure is pure facilitates inference in the context of the language’s loop-parallelism construct `do concurrent`. Also, explicitly prohibiting the overriding of a parent type’s type-bound procedures eliminates the need for dynamic dispatch in performance-critical code. Finally, this paper uses batch inference calculations on a neural network surrogate for atmospheric aerosol dynamics to demonstrate that LLVM Flang compiler’s automatic parallelization of `do concurrent` achieves roughly the same performance and scalability as achieved by OpenMP compiler directives. We also demonstrate that double-precision inference costs 37–72% longer runtime than default-real precision with most values in the range 57-60%.

Rouson, Damian↗

Parallel Runtime Interface for Fortran (PRIF) Specification (Rev. 0.4)

This document specifies an interface to support the parallel features of Fortran, named the Parallel Runtime Interface for Fortran (PRIF). PRIF is a proposed solution in which the runtime library is responsible for coarray allocation, deallocation and accesses, image synchronization, atomic operations, events, and teams. In this interface, the compiler is responsible for transforming the invocation of Fortran-level parallel features into procedure calls to the necessary PRIF procedures. The interface is designed for portability across shared- and distributed-memory machines, different operating systems, and multiple architectures. Implementations of this interface are intended as an augmentation for the compiler's own runtime library. With an implementation-agnostic interface, alternative parallel runtime libraries may be developed that support the same interface. One benefit of this approach is the ability to vary the communication substrate. A central aim of this document is to define a parallel runtime interface in standard Fortran syntax, which enables us to leverage Fortran to succinctly express various properties of the procedure interfaces, including argument attributes.

97 MATHEMATICS AND COMPUTING↗

Efficient Parallelization of Irregular Applications on GPU Architectures

With the enlarging computation capacity of general Graphics Processing Units (GPUs), leveraging GPUs to accelerate parallel applications has become a critical topic in academia and industry. However, a wide range of irregular applications with the computation-/memory-intensive nature cannot easily achieve high GPU utilization. The challenges mainly involve the following aspects: first, data dependence leads to coarse-grained kernel and inefficient parallelism; second, heavy GPU memory usage may cause frequent memory evictions and extra overhead of I/O; third, specific computation patterns produce memory redundancies; last, workload balance and data reusability conjunctly benefit the overall performance, but there may exist a dynamic trade-off between them. Targeting these challenges, this dissertation proposes multiple optimizations to accelerate two real-world applications: many-body correlation functions to simulate nuclear physics in a large-scale scientific system; the other is the eALS-based matrix factorization recommendation system. To accelerate the calculations of many-body correlation functions, this dissertation presents three frameworks in GPU memory management and multi-GPU scheduling. Firstly, an optimized systematic GPU memory management framework, MemHC, utilizes a series of new memory reduction designs in GPU memory allocation, CPU/GPU communications, and GPU memory oversubscription. Secondly, an enhanced multi-GPU scheduling framework, MICCO, particularly by taking both data dimension (e.g., data reuse and data eviction) and computation dimension into account. MICCO designs a heuristic scheduling algorithm and a machine learning-based regression model to generate the optimal settings of a proposed new concept to manage the trade-off. Thirdly, a locality-aware multi-GPU scheduling framework. This scheduler leverages pipeline batch generation with a looking-ahead strategy by building local dependency graphs for memory transfer reduction and better data reuse, achieving up to 79.92% memory cost reduction and 1.67x speedup. To parallelize the eALS-based recommendation system, this dissertation proposes an efficient CPU/GPU heterogeneous recommendation system, HEALS. HEALS employs newly designed architecture-adaptive data formats to achieve load balance and good data locality on CPU and GPU. To mitigate the data dependence, HEALS presents a CPU/GPU collaboration model for both task parallelism and data parallelism with multiple kernel computation optimizations. In summary, this dissertation efficiently accelerates two typical irregular applications on GPUs by building four frameworks, including CPU/GPU collaboration, GPU memory management, and multi-GPU scheduling.

Wang, Qihan↗

DyG-DPCD: A Distributed Parallel Community Detection Algorithm for Large-Scale Dynamic Graphs

Dynamic (Temporal) graphs capture the valuable evolution of real-world systems, from the continuously evolving patterns of social interactions and genetic pathways to the dynamic fluctuations of economic forces. Detecting communities for such evolving networks poses unique challenges. Detecting and analyzing the evolution of communities within dynamic graphs unlocks valuable insights into the underlying structural and temporal patterns of real-world systems. However, the sheer volume of modern graph data and the inherent complexity of the temporal dimension pose significant challenges to scalable community detection algorithms. Addressing this gap, our work explores the limited landscape of scalable distributed-memory parallel methods specifically designed for dynamic network community detection. We propose a novel parallel algorithm, DyG-DPCD (Dynamic Graph Distributed Parallel Community Detection), to detect communities in dynamic networks using the Message Passing Interface (MPI) framework. We present a vertex-centric approach, allowing us to detect communities through local optimization. Furthermore, we enhance our baseline algorithm by incorporating three heuristics, which improve the algorithm’s performance significantly while maintaining the quality of the solutions. We demonstrate the efficiency of our algorithm by experimenting on several real-world large-scale networks with hundreds of millions of edges spanning diverse domains. Notably, DyG-DPCD achieves speedups between 25× and 30× for large networks that we experimented on using NERSC compute nodes. In conclusion, our algorithm outperforms the STINGER parallel re-agglomeration algorithm by 30×.

97 MATHEMATICS AND COMPUTING↗

Generation of a Strong Parallel Electric Field and Embedded Electron Jet in the Exhaust of Moderate Guide Field Reconnection

Abstract Magnetospheric Multiscale observed an extended current layer concurrent with a strong parallel electric field. The current layer is embedded in the reconnection exhaust and is consistent with an extended magnetized jet expected in roughly symmetric moderate guide field reconnection. The strong parallel electric field observed at the jet balances a strong gradient in the parallel electron pressure caused by a transition between magnetized electrons with T e ‖ > T e ⊥ on one side of the jet and demagnetized, roughly isotropic electrons across the thin current layer. Simulation results show how this transition can occur and how the associated electron pressure gradients in Ohm's law balance the parallel electric field.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

An Effective Current Balancing Method for Inverters With Paralleled Silicon Carbide Power Modules

Silicon carbide (SiC) MOSFET has many superior characteristics in high power applications. Paralleling the SiC devices is an effective approach to increase the current capacity of the system, which also leads to current sharing issues. In this article, both transient and steady-state current sharing issues are studied for the high-power grid-tied inverters with paralleled SiC modules with an effective current balancing strategy proposed to address both issues simultaneously. The steady-state imbalance current due to the parameter mismatch is suppressed by the self-inductance of power cables used for paralleling connection. With the elimination of steady-state imbalance current, the transient imbalance current sensing is much simplified and can be realized by using low-cost current sensors. With the sensed imbalance current, a closed-loop control approach is proposed to adjust the modulation reference signal for each module to mitigate the remaining transient imbalance current. Furthermore, the proposed method is straightforward to implement with enhanced compatibility and flexibility compared to existing methods. Experimental studies are performed using inverters with 2 and 4 power modules paralleled in each phase to demonstrate the effectiveness of the proposed method.

42 ENGINEERING↗

The Integrated Reference Region Analysis for Parallel DFIGs’ Interfacing Inductors

Although the traditional design of doubly-fed induction generators (DFIGs)’ interfacing inductors consider the peak ripple of the grid-side converter (GSC)’s output current, it does not consider the inductors’ impact on DFIGs’ smallsignal stability. Located in series between the GSC and the stator, the inappropriate selection of the interfacing inductor can easily result to system instability. Therefore, this paper first proposes the integrated reference region analysis for small wind farm's interfacing inductors to improve the traditional design method. Firstly, a new detailed parallel DFIGs’ smallsignal model that considers output currents’ coupling is built in d-q coordinate system with state-space approach. The model focuses on representing the operating states of parallel DFIGs in wind farm. Secondly, the linearized state-space matrix of parallel DFIGs is decomposed into nominal-value matrix and location matrix. Considering the traditional design requirements, the integrated reference region for interfacing inductor is proposed through spectral radius and bialternate matrix sum (BMS). Furthermore, it can provide better guidance for parameter selecting and stabilization method researches. Finally, the simulation and experimental results show that the proposed integrated reference region for parallel DFIGs’ interfacing inductors is accurate and instructive.

47 OTHER INSTRUMENTATION↗

Parallel Randomized Tucker Decomposition Algorithms

The Tucker tensor decomposition is a natural extension of the singular value decomposition (SVD) to multiway data. Here, we propose to accelerate Tucker tensor decomposition algorithms by using randomization and parallelization. We present two algorithms that scale to large data and many processors, significantly reduce both computation and communication cost compared to previous deterministic and randomized approaches, and obtain nearly the same approximation errors. The key idea in our algorithms is to perform randomized sketches with Kronecker-structured random matrices, which reduces computation compared to unstructured matrices and can be implemented using a fundamental tensor computational kernel. We provide probabilistic error analysis of our algorithms and implement a new parallel algorithm for the structured randomized sketch. Our experimental results demonstrate that our combination of randomization and parallelization achieves accurate Tucker decompositions much faster than alternative approaches. We observe up to a 16X speedup over the fastest deterministic parallel implementation on 3D simulation data.

Tucker decompositions↗

Robustness of Deep Learning Classification to Adversarial Input on GPUs: Asynchronous Parallel Accumulation Is a Source of Vulnerability

The ability of machine learning (ML) classification models to resist small, targeted input perturbations—known as adversarial attacks—is a key measure of their safety and reliability. We show that floating-point non associativity (FPNA) coupled with asynchronous parallel programming on GPUs is sufficient to result in misclassification, without any perturbation to the input. Additionally, we show that this misclassification is particularly significant for inputs close to the decision boundary and that standard adversarial robustness results may be overestimated up to 4.6 when not considering machine-level details. We first study a linear classifier, before focusing on standard Graph Neural Network (GNN) architectures and datasets used in robustness assessments. We develop a novel black-box attack using Bayesian optimization to discover external workloads that can change the instruction scheduling which bias the output of reductions on GPUs and reliably lead to misclassification. Motivated by these results, we present a new learnable permutation (LP) gradient-based approach to learning floating-point operation orderings that lead to misclassifications. The LP approach provides a worst-case estimate in a computationally efficient manner, avoiding the need to run identical experiments tens of thousands of times over a potentially large set of possible GPU states or architectures. Finally, using instrumentation-based testing, we investigate parallel reduction ordering across different GPU architectures under external background workloads, when utilizing multi-GPU virtualization, and when applying power capping. Our results demonstrate that parallel reduction ordering varies significantly across architectures under the first two conditions, substantially increasing the search space required to fully test the effects of this parallel scheduler-based vulnerability. These results and the methods developed here can help to include machine-level considerations into adversarial robustness assessments, which can make a difference in safety and mission critical applications.

Shanmugavelu, Sanjif [Maxeler Technologies, a Groq↗

Massively parallel phase-field simulations targeting exascale

The interface thickness in the phase-field (PF) method limits its simulation scales. Consequently, large-scale PF simulations become prohibitively expensive for resolving the extremely fine microstructures that typically form during rapid solidification processing. This challenge is significant in predicting microstructure evolution in metal additive manufacturing and has been identified by the United States Department of Energy’s Exascale Computing Project. Here, to address this, we develop a multi-GPU and MPI-based massively parallel simulation code, utilizing state-of-the-art algorithms, software, and libraries, for large-scale three-dimensional (3D) PF simulations. We report the first GPU-parallel PF simulations on Frontier (currently the second TOP500 exascale cluster) and Summit machines, taking dendritic growth as an example problem. We evaluate the parallel performance of our implementation using scaling studies with more than 24 000 GPUs (among the largest known computations to date) and the acceleration performance using large-scale simulations of dendritic growth in 3D. Finally, massively parallel GPUs in these supercomputers enabled the first coupled multiscale simulations of laser melting and subsequent dendritic solidification on the scale of a full melt-pool, demonstrating the feasibility of performing PF simulations with a point total over 2 billion grid points within an acceptable time.

Exascale↗

Including the parallel mass flow in calculating the steady-state solutions and stability of the momentum balance equations for a quasisymmetric stellarator

The Helically Symmetric Experiment (HSX) is a quasisymmetric stellarator with minimal parallel viscous damping in a helical direction. The parallel flow (Vǁ) along the magnetic field is similarly weakly damped by viscosity. In this paper, the self-consistent steady-state parallel and poloidal momentum balance equations are used to show that a large Vǁ on the order of the ion thermal velocity can increase the ion resonant radial electric field (Er) beyond the value calculated using the typical approximation that Vǁ is zero. By altering the damping of Vǁ, either by degrading the quasisymmetry or varying the neutral density, the ion resonant Er can shift in a controllable fashion. It is shown explicitly that there exist stable and unstable steady-state solutions in the two-dimensional space of Vǁ and Er. A stability analysis of each solution is performed by calculating the eigenvalues and eigenvectors of the Jacobian. The unstable solution corresponds to a saddle point in which the eigenvalues have opposite signs. The analysis leads to the conclusion that unstable solutions occur when the derivative of the total poloidal damping with respect to Er is positive. A hysteresis in Er and Vǁ is observed when the radial current density is linearly increased to a maximum and then decreased back to zero. Jumps in the radial electric field and the parallel flow are observed as the radial current density drives the evolution from one stable point to the next. This result is similar to experimental data observed on several devices.

Physics↗

Kinetic study of shock formation and particle acceleration in laser-driven quasi-parallel magnetized collisionless shocks

Quasi-parallel magnetized collisionless shocks are believed to be one of the most efficient accelerators in the universe. Compared to quasi-perpendicular shocks, quasi-parallel shocks are more difficult to form in the laboratory and to simulate because of their large spatial scales and long formation times. Our two-dimensional particle-in-cell simulations show that the early stages of quasi-parallel shock formation are achievable in experiments planned for the National Ignition Facility and that particles accelerated by diffusive shock acceleration (DSA) are expected to be observable in the experiment. Repetitive ion acceleration by crossings of the shock front, a key feature of DSA, is seen in the simulations. Other characteristic features of quasi-parallel shocks such as upstream wave excitation by energetic ions are also observed, and energy partition between the ions and the electrons in the downstream of the shock is briefly discussed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Analysis of the impact of parallel magnetic fluctuations on linear gyrokinetic stability in NSTX-U and verification of gyro-fluid models

In this work, we use the CGYRO gyrokinetic code to analyze two L- and one H-mode discharges from the National Spherical Torus Experiment (NSTX) and NSTX-Upgrade (NSTX-U) selected due to their different mix of ion-scale driftwaves, ion temperature gradient (ITG) mode and trapped electron mode (TEM), and electromagnetic instabilities, kinetic ballooning mode (KBM), and micro-tearing mode (MTM) in the plasma core. It is found that the effect of parallel magnetic fluctuations is strongly destabilizing to the unstable KBMs compared to calculations with only perpendicular magnetic fluctuations. Two discharges have a mix of ITG/TEM and MTMs that are predicted to be dominant instability across the plasma radius. The parallel magnetic fluctuations are found to have little effect on the MTM stability but are destabilizing to ITG/TEM modes. To test the validity of the gyro-fluid linear stability codes TGLF and GFS at low aspect ratio, a database of linear growth rates has been created using the CGYRO gyrokinetic code. The database is comprised of various parameter scans around a standardized set of NSTX-U core parameters. It contains a group of electrostatic cases and an electromagnetic group that includes the effects of perpendicular and parallel magnetic fluctuations. Comparing the results from the GFS and TGLF models, we find that GFS exhibits the best agreement with the database of CGYRO linear growth rates. Comparing the model results for the electromagnetic scans shows that GFS captures the effects of parallel magnetic fluctuations accurately, while the TGLF model does not, as it lacks sufficient perpendicular energy resolution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Parallel interior-point solver for block-structured nonlinear programs on SIMD/GPU architectures

Here, we investigate how to port the standard interior-point method to new exascale architectures for block-structured nonlinear programs with state equations. Computationally, we decompose the interior-point algorithm into two successive operations: the evaluation of the derivatives and the solution of the associated Karush-Kuhn-Tucker (KKT) linear system. Our method accelerates both operations using two levels of parallelism. First, we distribute the computations on multiple processes using coarse parallelism. Second, each process uses SIMD/GPU accelerators locally to accelerate the operations using fine-grained parallelism. The KKT system is reduced by eliminating the inequalities and the state variables from the corresponding equations. We demonstrate our method's capability on the supercomputer Polaris, a testbed for the future exascale Aurora system. Each node is equipped with four GPUs, a setup amenable to our two-level approach. Our experiments on the stochastic optimal power flow problem show that the reduction method is 50x faster than the sparse linear solver HSL MA57 running in serial on the CPU, and 6x faster than Pardiso running in parallel on CPU on the same number of processes.

97 MATHEMATICS AND COMPUTING↗

Implementation of Pfirsch–Schlüter parallel flow in x-ray imaging crystal spectrometer inversion analysis

The x-ray imaging crystal spectrometer (XICS) tomographic inversion code for Wendelstein 7-X (W7-X) has been modified to consider the effects of parallel flows and has been applied to analyze measurements taken during recent experimental campaigns. Previous analysis neglected the effects of parallel flows due to the primarily perpendicular geometry of the sightlines and the small magnitude predicted by neoclassical theory. To reconsider these effects, the incompressibility condition for plasma flows is used to calculate the parallel Pfirsch–Schlüter flow component for the equilibrium configuration. By incorporating this condition along with the geometry of the sightlines—i.e. the fractional contributions of perpendicular and parallel flows—, an updated expression for the measured flow is used for the profile inversion. Application of this modified inversion code to data from W7-X shows that the magnitude of the radial electric field and the flux surface averaged perpendicular flow are reduced by approximately a factor of 2 and brought into better agreement with neoclassical predictions and charge exchange recombination spectroscopy measurements.

Pfirsch–Schlüter flows↗