Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

User Manual - HydraGNN v5.0: Distributed Implementation of Multi-Tasking Graph Neural Networks

This document serves as the user manual for HydraGNN v5.0, a scalable graph neural network (GNN) architecture for simultaneous prediction of multiple target properties using multi-task learning (MTL). This version of HydraGNN has been developed primarily to support the development, training, and deployment of predictive graph-based deep learning (DL) models for atomistic materials modeling. HydraGNN is templated over 13 message-passing policies, including invariant models (GIN, PNA, PNAPlus, GAT, MFC, CGCNN, SAGE, SchNet, DimeNet) and equivariant models (EGNN, PNAEq, PAINN, MACE), and supports distributed training via distributed data parallelism (DDP), DeepSpeed, and Fully Sharded Data Parallelism (FSDP) on leadership-class supercomputers. Although HydraGNN can be applied to problems beyond atomistic materials modeling, its current use is confined to homogeneous graphs. Additional capabilities include machine-learned interatomic potentials with energy-conserving forces, General, Powerful, and Scalable Graph Transformer (GraphGPS) global attention, periodic boundary conditions, hyperparameter optimization, mixed-precision training, and uncertainty quantification.

97 MATHEMATICS AND COMPUTING↗

Scale-up Unlearnable Examples Learning with High-performance Computing

Recent advancements in AI models, like ChatGPT, are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources (e.g., a single workstation). To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE’s unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. The use of Summit’s high-performance GPUs, along with the efficiency of the DDP framework, facilitated rapid updates of model parameters and consistent training across nodes. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications. The source code is publicly available at https: // github. com/ hrlblab/ UE_ HPC .

Zhu, Yanfan [Vanderbilt University, Nashville, TN,↗

Computing Hypergraph Homology in Chapel

In this paper, we discuss our experience in implementing homology computation, in particular Betti number calculation in Chapel hypergraph Library (CHGL). Given a dataset represented as a hypergraph, a Betti number for a particular dimension $k$ indicates how many $k$-dimensional `voids' are present in the dataset. Computing Betti number involves various array-centric and linear algebra operations. We demonstrate that implementing these operations in Chapel is both concise and intuitive. In addition, we show that Chapel provides language constructs for implementing parallel and distributed execution of the linear algebra kernels with minimal effort. Syntactically, Chapel provides succinctness of Python, while delivering comparable and better performance than C++-based and Julia-based packages for calculating Betti numbers respectively.

hypergraph, topological data analysis↗

Parallel Simulation of Quantum Networks with Distributed Quantum State Management

Quantum network simulators offer the opportunity to cost-efficiently investigate potential avenues for building networks that scale with the number of users, communication distance, and application demands by simulating alternative hardware designs and control protocols. Several quantum network simulators have been recently developed with these goals in mind. As the size of the simulated networks increases, however, sequential execution becomes time-consuming. Parallel execution presents a suitable method for scalable simulations of large-scale quantum networks, but the unique attributes of quantum information create unexpected challenges. In this work, we identify requirements for parallel simulation of quantum networks and develop the first parallel discrete-event quantum network simulator by modifying the existing serial simulator SeQUeNCe. Our contributions include the design and development of a quantum state manager (QSM) that maintains shared quantum information distributed across multiple processes. We also optimize our parallel code by minimizing the overhead of the QSM and decreasing the amount of synchronization needed among processes. Using these techniques, we observe a speedup of 2 to 25 times when simulating a 1,024-node linear network topology using 2 to 128 processes. We also observe an efficiency greater than 0.5 for up to 32 processes in a linear network topology of the same size and with the same workload. We repeat this evaluation with a randomized workload on a caveman network. We also introduce several methods for partitioning networks by mapping them to different parallel simulation processes. We have released the parallel SeQUeNCe simulator as an open source tool alongside the existing sequential version.

97 MATHEMATICS AND COMPUTING↗

Aperture size distribution, length, and preferential location of bed-parallel veins in shale

Bed-parallel, calcite-filled veins (BPVs) are common in shale formations, and although they have been widely described in other studies, little is known about their population aperture size distribution. To address this knowledge gap, we analyzed BPV sizes in outcrops and cores from the Vaca Muerta Formation, Neuquén Basin, Argentina; in two cores from the Marcellus Formation, Appalachian Basin, northeast Pennsylvania; and in one core from the Wolfcamp Shale, Delaware Basin, West Texas. Nine out of ten aperture size populations follow a negative exponential distribution, with one following a weak power law. Bed-parallel vein size distribution and intensity vary among formations and within the same shale. We define three groups of distributions: (1) Vaca Muerta outcrops, with the highest BPV intensity and the largest BPVs (cumulative frequency of 4.9 BPVs per meter [BPVs/m] for apertures 0.265 mm to 8.7 cm); (2) Vaca Muerta cores with a similar BPV intensity overall but with no apertures wider than 1.2 cm; and (3) Vaca Muerta, Wolfcamp, and Marcellus cores with the fewest BPVs (cumulative frequency up to 0.63 BPVs/m) and very few wider than 1 cm. Aperture and length in two outcrop data sets are weakly positively correlated and follow power laws with exponents of 0.44 and 0.49. Mechanical interfaces at boundaries between different lithologies exert a strong control on BPV location, with 65–75% of observed interfaces having BPVs along them. Only 25–30% of the BPVs occur at observed material interfaces, however, and unless subtle, unobserved mechanical layering is present, other factors must also control location. BPV intensity and organic richness (TOC) from Vaca Muerta well logs are correlated in some instances but not in others, indicating TOC is not always a good proxy for BPV location or intensity. Furthermore, these findings provide useful information for modeling of hydraulic fracture treatments where BPVs may influence development of the stimulated fracture network, for example by limiting height growth.

58 GEOSCIENCES↗

Parallel interior-point solver for block-structured nonlinear programs on SIMD/GPU architectures

Here, we investigate how to port the standard interior-point method to new exascale architectures for block-structured nonlinear programs with state equations. Computationally, we decompose the interior-point algorithm into two successive operations: the evaluation of the derivatives and the solution of the associated Karush-Kuhn-Tucker (KKT) linear system. Our method accelerates both operations using two levels of parallelism. First, we distribute the computations on multiple processes using coarse parallelism. Second, each process uses SIMD/GPU accelerators locally to accelerate the operations using fine-grained parallelism. The KKT system is reduced by eliminating the inequalities and the state variables from the corresponding equations. We demonstrate our method's capability on the supercomputer Polaris, a testbed for the future exascale Aurora system. Each node is equipped with four GPUs, a setup amenable to our two-level approach. Our experiments on the stochastic optimal power flow problem show that the reduction method is 50x faster than the sparse linear solver HSL MA57 running in serial on the CPU, and 6x faster than Pardiso running in parallel on CPU on the same number of processes.

97 MATHEMATICS AND COMPUTING↗

EQC: Ensembled Quantum Computing for Variational Quantum Algorithms

Variational quantum algorithms (VQA), which are comprised of a classical optimizer and a parameterized quantum circuit, emerges as one of the most promising approaches of harvesting quantum power in the noisy-intermediate-scale-quantum (NISQ) era. However, the deployment of VQAs on today's NISQ devices often faces considerable system noise and prohibitively slow training speeds. On the other hand, the expensive supporting sources and infrastructure make quantum computers extremely keen on high utilization. In this paper, we propose a novel way of thinking about a quantum backend: rather than relying on one physical device which tends to introduce platform-specific noise and bias, a quantum ensemble, which distributes quantum tasks across parallel devices, can serve as a virtualized quantum computer for offering reduced noise levels through an adaptive mixture and also provide significantly improved training speeds through parallelization. With this idea, we build a distributive VQA optimization framework called DVQA, serving as the first effort in adopting parallel quantum devices for cooperative VQA training. To further constraint noise and speed-up convergence, we design a model for individual NISQ devices concerning their properties and running conditions, and propose a weighting mechanism for regularizing the returned gradients. Extensive evaluations on 10 IBM-Q quantum devices using the VQE example show that the distributive VQA training framework can substantially boost the training speed by 10.5x on average (up to 86x and at least 5.2x) with improved training accuracy.

Stein, Samuel A.↗

Advanced architectures for high-performance quantum networking

As practical quantum networks prepare to serve an ever-expanding number of nodes, there has grown a need for advanced auxiliary classical systems that support the quantum protocols and maintain compatibility with the existing fiber-optic infrastructure. We propose and demonstrate a quantum local area network design that addresses current deployment limitations in timing and security in a scalable fashion using commercial off-the-shelf components. First, we employ White Rabbit switches to synchronize three remote nodes with ultra-low timing jitter, significantly increasing the fidelities of the distributed entangled states over previous work with Global Positioning System clocks. Second, using a parallel quantum key distribution channel, we secure the classical communications needed for instrument control and data management. Therefore, the conventional network that manages our entanglement network is secured using keys generated via an underlying quantum key distribution layer, preserving the integrity of the supporting systems and the relevant data in a future-proof fashion.

97 MATHEMATICS AND COMPUTING↗

CG-Kit: Code Generation Toolkit for performant and maintainable variants of source code applied to Flash-X hydrodynamics simulations

CG-Kit is a new Code Generation tool-Kit that we have developed as a part of the solution for portability and maintainability for multiphysics computing applications. The development of CG-Kit is rooted in the urgent need created by the shifting landscape of high-performance computing platforms and the algorithmic complexities of a particular large-scale multiphysics application: Flash-X. To efficiently use computing resources on a heterogeneous node, an application must have a map of computation to resources and a mechanism to move the data and computation to the resources according to the map. Most existing performance portability solutions are focussed on abstracting the expression of computations so that a unified source code can be specialized to run on different resources. However, such an approach is insufficient for a code like Flash-X, which has a multitude of code components that can be assembled in various permutations and combinations to form different instances of applications. Similar challenges apply to any code that has composability, where a single specified way of apportioning work among devices may not be optimal. Additionally, use cases arise where the optimal control flow of computation may differ for different devices while the underlying numerics remain identical. This combination leads to unique challenges including handling an existing large code base in Fortran and/or C/C++, subdivision of code into a great variety of units supporting a wide range of physics and numerical methods, different parallelization techniques for distributed and shared memory systems and accelerator devices, and heterogeneity of computing platforms requiring coexisting variants of parallel algorithms. All of these challenges demand that scientific software developers apply existing knowledge about domain applications, algorithms, and computing platforms to determine custom abstractions and granularity for code generation. There is a critical lack of tools to tackle those problems. CG-Kit is designed to fill this gap by providing a user with the ability to express their desired control flow and computation-to-resource map in the form a pseudocode-like recipe. It consists of standalone tools that can be combined into highly specific and, we argue, highly effective portability and maintainability toolchains. Here we present the design of our new tools: parametrized source trees, control flow graphs, and recipes. The tools are implemented in Python. They are agnostic to the programming language of the source code targeted for code generation. In conclusion, we demonstrate the capabilities of the toolkit with two examples, first, multithreaded variants of the basic AXPY operation, and second, variants of parallel algorithms within a hydrodynamics solver, called Spark, from Flash-X that operates on block-structured adaptive meshes.

Algorithmic portability↗

Efficient distributed continual learning for steering experiments in real-time

Deep learning has emerged as a powerful method for extracting valuable information from large volumes of data. However, when new training data arrives continuously (i.e., is not fully available from the beginning), incremental training suffers from catastrophic forgetting (i.e., new patterns are reinforced at the expense of previously acquired knowledge). Training from scratch each time new training data becomes available would result in extremely long training times and massive data accumulation. Rehearsal-based continual learning has shown promise for addressing the catastrophic forgetting challenge, but research to date has not addressed performance and scalability. To fill this gap, we propose an approach based on a distributed rehearsal buffer that efficiently complements data-parallel training on multiple GPUs to achieve high accuracy, short runtime, and scalability. It leverages a set of buffers (local to each GPU) and uses several asynchronous techniques for updating these local buffers in an embarrassingly parallel fashion, all while handling the communication overheads necessary to augment input minibatches using unbiased, global sampling. We further propose a generalization of rehearsal buffers to support both classification and generative learning tasks, as well as more advanced rehearsal strategies (notably Dark Experience Replay, leveraging knowledge distillation). We illustrate this approach with a real-life HPC streaming application from the domain of ptychographic image reconstruction. Furthermore, we run extensive experiments on up to 128 GPUs of the ThetaGPU supercomputer to compare our approach with baselines representative of training-from-scratch (the upper bound in terms of accuracy) and incremental training (the lower bound). Results show that rehearsal-based continual learning achieves a top-5 validation accuracy close to the upper bound, while simultaneously exhibiting a runtime close to the lower bound.

Asynchronous data management↗

FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters

In this paper, we propose a framework called FPDeep, which uses a hybrid of model and layer paral- lelism to configure distributed reconfigurable clusters to train DNNs. This approach has numerous benefits. First, the design does not suffer from batch size growth. Second, novel workload and weight partitioning leads to balanced loads of both among nodes. And third, the entire system is fine-grained pipeline. This leads to high parallelism and utilization and also minimizes the time features need to be cached while waiting for back-propagation.

Wang, Tianqi↗

A Massively Parallel Implementation of the CCSD(T) Method Using the Resolution-of-the-Identity Approximation and a Hybrid Distributed/Shared Memory Parallelization Model

In this work, a parallel algorithm is described for the coupled-cluster singles and doubles method augmented with a perturbative correction for triple excitations [CCSD(T)] using the resolution-of-the-identity (RI) approximation for two-electron repulsion integrals (ERIs). The algorithm bypasses the storage of four-center ERIs by adopting an integral-direct strategy. The CCSD amplitude equations are given in a compact quasi-linear form by factorizing them in terms of amplitude-dressed three-center intermediates. A hybrid MPI/OpenMP parallelization scheme is employed, which uses the OpenMP-based shared memory model for intranode parallelization and the MPI-based distributed memory model for internode parallelization. Parallel efficiency has been optimized for all terms in the CCSD amplitude equations. Two different algorithms have been implemented for the rate-limiting terms in the CCSD amplitude equations that entail and -scaling computational costs, where N O and N V denote the number of correlated occupied and virtual orbitals, respectively. One of the algorithms assembles the four-center ERIs requiring N V 4 and N O 2 N V 2 -scaling memory costs in a distributed manner on a number of MPI ranks, while the other algorithm completely bypasses the assembling of quartic memory-scaling ERIs and thus largely reduces the memory demand. It is demonstrated that the former memory-expensive algorithm is faster on a few hundred cores, while the latter memory-economic algorithm shows a better strong scaling in the limit of a few thousand cores. The program is shown to exhibit a near-linear scaling, in particular for the compute-intensive triples correction step, on up to 8000 cores. The performance of the program is demonstrated via calculations involving molecules with 24–51 atoms and up to 1624 atomic basis functions. As the first application, the complete basis set (CBS) limit for the interaction energy of the π-stacked uracil dimer from the S66 data set has been investigated. This work reports the first calculation of the interaction energy at the CCSD(T)/aug-cc-pVQZ level without local orbital approximation. The CBS limit for the CCSD correlation contribution to the interaction energy was found to be -8.01 kcal/mol, which agrees very well with the value -7.99 kcal/mol reported by Schmitz, Hättig, and Tew [ Phys. Chem. Chem. Phys. 2014 , 16 , 22167-22178]. The CBS limit for the total interaction energy was estimated to be -9.64 kcal/mol.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

PREEMPT: Scalable Epidemic Interventions Using Submodular Optimization on Multi-GPU Systems

Preventing and slowing the spread of epidemics is achieved through techniques such as vaccination and social distancing. Given practical limitations on the number of vaccines and cost of administration, optimization becomes a necessity. Previous approaches using mathematical programming methods have shown to be effective but are limited by computational costs. In this work, we make several contributions: First, we present a new approach for intervention via maximizing the influence of vaccinated nodes on the network. We call this method \preempt. Next, we prove submodular properties associated with the objective function of our method so that it aids in construction of an efficient greedy approximation strategy. Consequently, we present a new parallel algorithm based on greedy hill climbing for \preempt, and present an efficient parallel implementation for distributed CPU-GPU heterogeneous platforms. Our results demonstrate that \preempt{} is able to achieve a significant reduction (up to 6.75$\times$) in the percentage of people infected on a city-scale network. We also show strong scaling results of \preempt{} on 128 nodes of the Summit supercomputer. Our parallel implementation is able to significantly reduce time to solution, from hours to minutes on large networks. This work represents a first-of-its-kind effort in parallelizing greedy hill climbing and applying it toward devising effective interventions for epidemics.

Minutoli, Marco↗

GronOR: Massively Parallel and GPU-Accelerated Non-Orthogonal Configuration Interaction for Large Molecular Systems

GronOR is a program package for non-orthogonal configuration interaction calculations for an electronic wave function built in terms of anti-symmetrized products of multi-configuration molecular fragment wave functions. The two-electron integrals that have to be processed may be expressed in terms of atomic orbitals or in terms of an orbital basis determined from the molecular orbitals of the fragments. The code has been specifically designed for execution on distributed memory massively parallel and Graphics Processing Unit (GPU)-accelerated computer architectures, using an MPI+OpenACC/OpenMP programming approach. The task-based execution model used in the implementation allows for linear scaling with the number of nodes on the largest pre-exascale architectures available, provides hardware fault resiliency, and enables effective execution on systems with distinct central processing unit-only and GPU-accelerated partitions. The code interfaces with existing multi-configuration electronic structure codes that provide optimized molecular fragment orbitals, configuration interaction coefficients, and the required integrals. Algorithm and implementation details, parallel and accelerated performance benchmarks, and an analysis of the sensitivity of the accuracy of results and computational performance to thresholds used in the calculations are presented.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Exploring Anomalous Photoelectron Angular Distributions in the Photoelectron Spectra of Gd 3 O 3 – : Study of Gd 3 O 2 – and Gd 3 O 3 – Using Photoelectron Spectroscopy and Density Functional Theory Calculations

Anion photoelectron (PE) spectra of lanthanide oxide clusters obtained previously have exhibited anomalous photoelectron angular distributions which were attributed to strong PE–valence electron (PEVE) interactions. Here, to further explore this effect, we have obtained the PE spectra of Gd 3 O 2 – and Gd 3 O 3 – , two clusters that have similarly complex electronic structures but contrasting symmetries. The spectra exhibit manifolds of detachment transitions at similar binding energies in a 0.5 eV window of energy. The electron affinity of Gd 3 O 2 is measured to be 1.29 ± 0.05 eV, and that of Gd 3 O 3 is 1.31 ± 0.05 eV. As seen in previous studies on lanthanide oxide cluster anions in lower than conventional oxidation states, transitions in spectra obtained lower photon energies are more congested than those obtained with higher photon energy, a signature of strong PEVE interactions. While the detachment transitions have predominantly parallel photoelectron angular distributions (PAD), the PAD varies across the manifold of transitions in the PE spectrum of Gd 3 O 3 – in a way that suggests four different subgroups of transitions. Results of calculations on Gd 3 O 2 – suggest kite or V-shape structures with antiferromagnetic coupling between one of the 4f 7 subshells with the two others. Calculations on Gd 3 O 3 – more definitively point to ring structures with a nearly isoenergetic ferromagnetically coupled high spin (24-tet) state and a dectet state in which one of the 4f 7 subshells is antiferromagnetically coupled with the other two. Taking these results as qualitative, we propose that strong mixing between the unperturbed states predicted computationally leads to overlapping transitions with different PADs.

anions↗

ArborX

ArborX library tackles a problem of efficiently finding geometric objects that are close in space. Variations of this problem, such as finding the nearest neighbors of a point, or finding all objects within a certain distance, are inherent components of applications in many fields. The data may be large so that solving the problem efficiently may require significant computational resources, such as multiple processors or accelerators such as general purpose GPUs. ArborX' main advantage in its ability to solve large problems efficiently utilizing a combination of distributed and on-node parallelism. ArborX can be run efficiently on a wide variety of hardware, including GPUs from different vendors, which distinguishes it from other available libraries which typically choose only few of these. The other advantage is that it supports both types of user problems: spatial problems (useful for intersections and finding objects within certain distance), and nearest neighbor problems. ArborX also supports flexible interface in its interaction with a user. Particularly, it allows a user to call user's own function on a positive match, a functionality not rarely available in other libraries. ArborX implements construction and traversal algorithms using efficient tree structures, such as bounding volume hierarchy (BVH). At its core, it uses linear BVH for its low construction cost and sufficient quality. ArborX is written using C++, and is parallelized using the message passing interface (MPI) for the distributed communication, and the Kokkos library for on-node parallelism. This approach allows ArborX to be run on a wide variety of hardware, from common laptops and desktops to supercomputers while using the same codebase. ArborX also implements several advanced algorithms using geometric search, such as density-based clustering algorithm DBSCAN.

ECP↗