Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

ArborX 2.0

ArborX library tackles a problem of efficiently finding geometric objects that are close in space. Variations of this problem, such as finding the nearest neighbors of a point, or finding all objects within a certain distance, are inherent components of applications in many fields. The data may be large so that solving the problem efficiently may require significant computational resources, such as multiple processors or accelerators such as general purpose GPUs. ArborX' main advantage in its ability to solve large problems efficiently utilizing a combination of distributed and on-node parallelism. ArborX can be run efficiently on a wide variety of hardware, including GPUs from different vendors, which distinguishes it from other available libraries which typically choose only few of these. The other advantage is that it supports both types of user problems: spatial problems (useful for intersections and finding objects within certain distance), and nearest neighbor problems. ArborX also supports flexible interface in its interaction with a user. Particularly, it allows a user to call user's own function on a positive match, a functionality not rarely available in other libraries. ArborX implements construction and traversal algorithms using efficient tree structures, such as bounding volume hierarchy (BVH). At its core, ArborX uses linear BVH for its low construction cost and sufficient quality. ArborX implements both spatial and nearest-neighbor traversal algorithms. ArborX also provides several clustering algorithms (minimum spanning tree, DBSCAN, HDBSCAN*), interpolation using minimum least squares and ray tracing. ArborX is written using C++, and is parallelized using the message passing interface (MPI) for the distributed communication, and the Kokkos library for on-node parallelism. This approach allows ArborX to be run on a wide variety of hardware, from common laptops and desktops to supercomputers while using the same codebase.

Prokopenko, Andrey [Oak Ridge National Laboratory ↗

Multi-task Parallelism for Robust Pre-training of Graph Foundation Models on Multi-source, Multi-fidelity Atomistic Modeling Data

Graph foundation models using graph neural networks promise sustainable, efficient atomistic modeling. To tackle challenges of processing multi-source, multi-fidelity data during pre-training, recent studies employ multi-task learning, in which shared message passing layers initially process input atomistic structures regardless of source, then route them to multiple decoding heads that predict data-specific outputs. This approach stabilizes pre-training and enhances a model’s transferability to unexplored chemical regions. Preliminary results on approximately four million structures are encouraging, yet questions remain about generalizability to larger, more diverse datasets and scalability on supercomputers. We propose a multi-task parallelism method that distributes each head across computing resources with GPU acceleration. Implemented in the open-source HydraGNN architecture, our method was trained on over 24 million structures from five datasets and tested on the Perlmutter, Aurora, and Frontier supercomputers, demonstrating efficient scaling on all three highly heterogeneous super-computing architectures.

Lupo Pasini, Massimiliano [ORNL] (ORCID:0000000249↗

Bifurcation-like transition of divertor conditions induced by X-point radiation in KSTAR L-mode plasmas *

Abstract Density ramps with ion grad B drift directed into lower single null KSTAR L-mode plasmas are associated with a simultaneous and abrupt reduction of the divertor particle flux on both low- and high-field-side targets when the mid-plane line averaged electron density reaches a given level. Target embedded Langmuir probe signals show a clear ‘cliff edge’ behavior similar to that observed in the divertor target electron temperature in DIII-D H-mode plasmas (Eldon et al 2017 Nucl. Fusion 57 066039; McLean et al 2015 J. Nucl. Mater. 463 533–6). The collapse of the particle flux is observed along the whole divertor target area (from private flux region to the far scrape-off layer (SOL)). The critical upstream density of this target flux cliff is invariant under fuel gas throughput modulation. The transition along the cliff occurs in tens of milliseconds. With the cliff, carbon impurities and deuterium neutrals transported through the X-point to the core produce a strong radiation spot near the X-point, seen on bolometric signals, and increase the upstream density. The experimental observations are consistent with time-dependent SOLPS-ITER simulations, which also demonstrate an abrupt transition of the target flux and upstream density with the increase in X-point radiation. The timescale of the cliff predicted by SOLPS-ITER is consistent with the experiment, although, it is influenced by gas throughput or time-dependent numerical methods. In the L-mode phase space of separatrix electron density and temperature, branches are divided based on target temperature, because the latter is strongly coupled to the radiation front and ionization front due to the monotonic characteristic of the parallel electron temperature distribution. Since the H-mode condition operates at a much higher upstream density and electron temperature in phase space, dissipation from sputtered carbon alone leads to the density limit before reaching the X-point radiation condition. This is therefore consistent with the fact that cliffs have never been observed in H-mode KSTAR experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A semi-automated algorithm for designing stellarator divertor and limiter plates and application to HSX

We present a semi-automated algorithm for designing three-dimensional divertor or limiter plates targeting low heat loads. The algorithm designs the plates in two stages: firstly, the parallel heat flux distribution is caught on vertically-inclined plates at one or several toroidal locations. Secondly, the power per unit area is reduced by stretching, tilting and bending the plates toroidally. Heat transport is modelled using the EMC3-Lite code, which uses an anisotropic diffusion model. We apply this scheme to HSX, a medium-sized stellarator located at the University of Wisconsin–Madison. Starting from the current machine with an extended vessel wall, we construct plates which are able to effectively catch and spread the heat for three different magnetic configurations. The scheme has a computational cost in the order of tens of CPU-minutes, making it a powerful tool for semi-automated plasma-facing component design in three-dimensional environments.

anisotropic diffusion↗

Grid-Forming and Grid-Following Inverter Comparison of Droop Response

With the increase in penetration of inverter-based resources (IBRs) in the electrical power system, the ability of these devices to provide grid support to the system has become a necessity. With standards previously developed for the interconnection requirements of grid-following inverters (GFLI) (most commonly photovoltaic inverters), it has been well documented how these inverters “should” respond to changes in voltage and frequency. However, with other IBRs such as grid-forming inverters (GFMIs) (used for energy storage systems, standalone systems, and as uninterruptable power supplies) these requirements are either: not yet documented, or require a more in deep analysis. With the increased interest in microgrids, GFMIs that can be paralleled onto a distribution system have become desired. With the proper control schemes, a GFMI can help maintain grid stability through fast response compared to rotating machines. This paper will present an experimental comparison of commercially available GFMI and GFLI ' responses to voltage and frequency deviation, as well as the GFMI operating as a standalone system and subjected to various changes in loads.

Grid Support, Inverter, Droop Control, Volt-VAR, F↗

Optimizing the Weather Research and Forecasting Model with OpenMP Offload and Codee

Currently, the Weather Research and Forecasting model (WRF) utilizes shared memory (OpenMP) and distributed memory (MPI) parallelisms. To take advantage of GPU resources on the Perlmutter supercomputer at NERSC, we port parts of the computationally expensive routine Fast Spectral Bin Microphysics (FSBM) to NVIDIA GPUs using OpenMP device offloading directives. To facilitate this process, we explore a workflow for optimization which uses both runtime profilers and a static code inspection tool Codee to refactor the subroutine. We observe an 2.24x overall speedup for the CONUS-12km storm test case.

Wichitrnithed, Chayanon (Namo) [Odin Institute]↗

Quandary

Quandary numerically simulates and optimizes the time-evolution of open quantum systems. The underlying dynamics are modelled by Lindblad's master equation, a linear ordinary differential equation (ODE) describing quantum systems interacting with the environment. Quandary solves this ODE numerically by applying a time-stepping integration scheme, and utilizes a gradient-based optimization approach to determine optimal control pulses that drive the quantum system to a desired target state. Two optimization objectives are considered: (a) Unitary gate optimization that finds controls to realize a unitary gate transformation, and (b) optimal reset that aims to drive the quantum system to the ground states. Gradient-based optimization schemes utilizing Petsc's Tao optimization package are applied to generate control pulses that minimize the respective measure. To evaluate the gradient of the objective function, the discrete adjoint method is used while leveraging techniques from Algorithmic Differentiation to produce exact and consistent gradients. To mitigate excessive execution run times, the software can be build together with the XBraid software library which provides a parallelization strategy to distribute the time-evolution of the underlying dynamics onto multiple processor.

Petersson, NilsA.↗

Decomposition and Algorithmic Approaches for Solving Large-Scale Process Family Design Problems

Our most recent work expands the water desalination case study from 76 variants to 10,897 variants using the equation-oriented model built in Pyomo as part of the PARETO project. Using the discretization formulation presented in Stinchfield (2024a), rather than solving for all 10,897 variants simultaneously, we decompose the formulation into subproblems containing subsets of variants from the process family. We solve the overall problem with Progressive Hedging (PH) deployed in parallel on a distributed HPC cluster using the open-source Python package mpi-sppy (Knueven et al., 2023). This approach allowed us to solve this process family design problem to ~1.5% relative optimality gap in about 5 hours; in comparison, Gurobi reached ~50% relative optimality gap in about 6 hours (Stinchfield et al., 2024b). However, this approach still requires discretization of the common unit module design ranges; additionally, PH acts as a heuristic for MILP’s with gap-closing capabilities. Ideally, we would not have to use ML surrogates or discretization to solve this problem, instead solving the process family design problem with the equation-oriented model directly to achieve the most accurate results. However, recall that we did not consider solving the MINLP directly due to complexity and size. In this work, we aim to decompose and solve this large-scale MINLP using a Structured Nonlinear Global Optimization algorithm presented by Cao and Zavala (2019).

Stinchfield, Georgia↗

A parallel p ‐adaptive discontinuous Galerkin method for the Euler equations with dynamic load‐balancing on tetrahedral grids

Abstract A novel p ‐adaptive discontinuous Galerkin (DG) method has been developed to solve the Euler equations on three‐dimensional tetrahedral grids. Hierarchical orthogonal basis functions are adopted for the DG spatial discretization while a third order TVD Runge‐Kutta method is used for the time integration. A vertex‐based limiter is applied to the numerical solution in order to eliminate oscillations in the high order method. An error indicator constructed from the solution of order and is used to adapt degrees of freedom in each computational element, which remarkably reduces the computational cost while still maintaining an accurate solution. The developed method is implemented with under the Charm++ parallel computing framework. Charm++ is a parallel computing framework that includes various load‐balancing strategies. Implementing the numerical solver under Charm++ system provides us with access to a suite of dynamic load balancing strategies. This can be efficiently used to alleviate the load imbalances created by p ‐adaptation. A number of numerical experiments are performed to demonstrate both the numerical accuracy and parallel performance of the developed p ‐adaptive DG method. It is observed that the unbalanced load distribution caused by the parallel p ‐adaptive DG method can be alleviated by the dynamic load balancing from Charm++ system. Due to this, high performance gain can be achieved. For the testcases studied in the current work, the parallel performance gain ranged from 1.5× to 3.7×. Therefore, the developed p ‐adaptive DG method can significantly reduce the total simulation time in comparison to the standard DG method without p ‐adaptation.

97 MATHEMATICS AND COMPUTING↗

Design and implementation of dynamic I/O control scheme for large scale distributed file systems

In this paper, we have analyzed the input/output (I/O) activities of Cori, which is a high-performance computing system at the National Energy Research Scientific Computing Center at Lawrence Berkeley National Laboratory. Our analysis results indicate that most users do not adjust storage configurations but rather use the default settings. In addition, owing to the interference from many applications running simultaneously, the performance varies based on the system status. To configure file systems autonomously in complex environments, we developed DCA-IO, a dynamic distributed file system configuration adjustment algorithm that utilizes the system log information to adjust storage configurations automatically. Our scheme aims to improve the application performance and avoid interference from other applications without user intervention. Moreover, DCA-IO uses the existing system logs and does not require code modifications, an additional library, or user intervention. To demonstrate the effectiveness of DCA-IO, we performed experiments using I/O kernels of real applications in both an isolated small-sized Lustre environment and Cori. Our experimental results shows that our scheme can improve the performance of HPC applications by up to 263% with the default Lustre configuration.

97 MATHEMATICS AND COMPUTING↗

Distributed non-negative matrix factorization with determination of the number of latent features

The holistic analysis and understanding of the latent (that is, not directly observable) variables and patterns buried in large datasets is crucial for data-driven science, decision making and emergency response. Such exploratory analyses require devising unsupervised learning methods for data mining and extraction of the latent features, and non-negative matrix factorization (NMF) is one of the prominent such methods. NMF is based on compute-intense non-convex constrained minimization, which, for large datasets requires fast and distributed algorithms. However, current parallel implementations of NMF fail to estimate the number of latent features. In practice, identifying these features is both difficult and significant for pattern recognition and latent feature analysis, especially for large dense matrices. Here, we introduce a distributed NMF algorithm coupled with distributed custom clustering followed by a stability analysis on dense data, which we call DnMFk, to determine the number of latent variables. The results on synthetic data and the classical Swimmer data set demonstrate the accuracy of model determination while scaling nearly linearly across multiple processors for large data. Further, we employ DnMFk to determine the number of hidden features from a terabyte matrix.

97 MATHEMATICS AND COMPUTING↗

Distributed Training for High Resolution Images: A Domain and Spatial Decomposition Approach

In this work we developed two Pytorch libraries using the PyTorch RPC interface for distributed deep learning approaches on high resolution images. The spatial decomposition library allows for distributed training on very large images, which otherwise wouldn’t be possible on a single GPU. The domain parallelism library allows for distributed training across multiple domain unlabeled data, by leveraging the domain separation architecture. Both of those libraries where tested on the Summit supercomputer at Oak Ridge National Laboratory at a moderate scale.

Tsaris, Aristeidis (aris)↗

TuckerMPI: A Parallel C++/MPI Software Package for Large-scale Data Compression via the Tucker Tensor Decomposition

With this study, our goal is compression of massive-scale grid-structured data, such as the multi-terabyte output of a high-fidelity computational simulation. For such data sets, we have developed a new software package called TuckerMPI, a parallel C++/MPI software package for compressing distributed data. The approach is based on treating the data as a tensor, i.e., a multidimensional array, and computing its truncated Tucker decomposition, a higher-order analogue to the truncated singular value decomposition of a matrix. The result is a low-rank approximation of the original tensor-structured data. Compression efficiency is achieved by detecting latent global structure within the data, which we contrast to most compression methods that are focused on local structure. In this work, we describe TuckerMPI, our implementation of the truncated Tucker decomposition, including details of the data distribution and in-memory layouts, the parallel and serial implementations of the key kernels, and analysis of the storage, communication, and computational costs. We test the software on 4.5 and 6.7 terabyte data sets distributed across 100 s of nodes (1,000 s of MPI processes), achieving compression ratios between 100 and 200,000×, which equates to 99--99.999% compression (depending on the desired accuracy) in substantially less time than it would take to even read the same dataset from a parallel file system. Moreover, we show that our method also allows for reconstruction of partial or down-sampled data on a single node, without a parallel computer so long as the reconstructed portion is small enough to fit on a single machine, e.g., in the instance of reconstructing/visualizing a single down-sampled time step or computing summary statistics. The code is available at https://gitlab.com/tensors/TuckerMPI.

97 MATHEMATICS AND COMPUTING↗

Performance Evaluation of Multi-Vendor Grid-Forming Inverters for Grid-Connected Operation Through Hardware Experimentation: Preprint

Existing real-world projects of GFM inverters that operate in parallel to power grids typically are sized between dozens and a few hundreds Megawatt (MW) scale according to a recent NERC GFM inverter white paper. These large systems are often difficult to evaluate prior to deployments because of their large size. The performance of smaller GFM inverters (dozens to a few hundreds MW) that operate parallel with power grids (distribution systems) is even less understood. There is an opportunity to better understand these systems through hardware testing under controlled laboratory conditions. Therefore, this paper presents the functional performance evaluation tests of multiple (three) commercial GFM inverters when they operate in parallel with the grid through hardware experiments. The goal of these tests is to explore the GFM inverters' functionalities and dynamic response when in parallel with power grids to eventually develop universal specifications for GFM inverters. Both steady state (changing the inverter's frequency and voltage droop) and transient (adding step change in grid's frequency/voltage) tests are performed for each GFM inverter with the same testing circuit and testing protocol. The experimental results indicate the bench-marked performance that: 1) the GFM inverters can be dispatched through frequency and voltage droop intercepts to output the target power when paralleled to the grid; 2) the GFM inverters automatically respond to system frequency and voltage events to output the needed power, however, the GFM inverters all show stability issues when absorbing reactive power from the grid.

grid-forming inverters↗