Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

Local time stepping for the shallow water equations in MPAS

In this work we assess the performance of a set of local time-stepping (LTS) schemes for the shallow water equations implemented in the Model for Prediction Across Scales (MPAS). The goal of LTS is to speed up the simulation by allowing different time-steps on different regions of the computational grid. The LTS schemes considered here were originally introduced by Hoang et al. (2019) [26], who laid out the mathematical foundation of the methods. Here, the authors take on the task of presenting a fast, efficient and scalable parallel implementation of these LTS methods on high performance computing machines, with the aim to provide a recipe for other climate modeling groups that may be interested in employing LTS algorithms in their codes. As a matter of fact, even if MPAS is our framework of choice, our approach is general enough and could be of interest to other groups beyond the MPAS community. Due to their nature, LTS methods possess an inherent load imbalance that needs to be carefully addressed in order to obtain efficient scalability. Even more important is the far from trivial task of computing the right-hand side terms only on specific LTS regions during the time-stepping procedure. An inefficient handling of this task causes a drastic decay of the CPU time performance, making the LTS algorithms practically of no use. The emphasis of the present work is therefore on the computational and parallel aspects of the LTS methods, whose proper treatment is crucial to make the methods run faster against existing strategies, such as for instance high-order explicit global time-stepping schemes. This is in fact the ultimate goal of using an LTS procedure and it is the one to which we direct all our optimization efforts.

97 MATHEMATICS AND COMPUTING↗

HipBone: A performance-portable graphics processing unit-accelerated C++ version of the NekBone benchmark

We present hipBone, an open-source performance-portable proxy application for the Nek5000 (and NekRS) computational fluid dynamics applications. HipBone is a fully GPU-accelerated C++ implementation of the original NekBone CPU proxy application with several novel algorithmic and implementation improvements which optimize its performance on modern fine-grain parallel GPU accelerators. Our optimizations include a conversion to store the degrees of freedom of the problem in assembled form in order to reduce the amount of data moved during the main iteration and a portable implementation of the main Poisson operator kernel. We demonstrate near-roofline performance of the operator kernel on three different modern GPU accelerators from two different vendors. We present a novel algorithm for splitting the application of the Poisson operator on GPUs which aggressively hides MPI communication required for both halo exchange and assembly. Our implementation of nearest-neighbor MPI communication then leverages several different routing algorithms and GPU-Direct RDMA capabilities, when available, which improves scalability of the benchmark. We demonstrate the performance of hipBone on three different clusters housed at Oak Ridge National Laboratory, namely, the Summit supercomputer and the Frontier early-access clusters, Spock and Crusher. Our tests demonstrate both portability across different clusters and very good scaling efficiency, especially on large problems.

Computer Science↗

Narrow-Band Least-Squares Infrasound Array Processing

Infrasound data from arrays can be used to detect, locate, and quantify a variety of natural and anthropogenic sources from local to remote distances. However, many array processing methods use a single broad frequency range to process the data, which can lead to signals of interest being missed due to the choice of frequency limits or simultaneous clutter sources. In this work, we introduce a new open-source Python code that processes infrasound array data in multiple sequential narrow frequency bands using the least-squares approach. We test our algorithm on a few examples of natural sources (volcanic eruptions, mass movements, and bolides) for a variety of array configurations. Our method reduces the need to choose frequency limits for processing, which may result in missed signals, and it is parallelized to decrease the computational burden. Improvements of our narrow-band least-squares algorithm over broad-band least-squares processing include the ability to distinguish between multiple simultaneous sources if distinct in their frequency content (e.g., microbarom or surf vs. volcanic eruption), the ability to track changes in frequency content of a signal through time, and a decreased need to fine-tune frequency limits for processing. We incorporate a measure of planarity of the wavefield across the array (sigma tau, στ) as well as the ability to utilize the robust least trimmed squares algorithm to improve signal processing and insight into array performance. Our implementation allows for more detailed characterization of infrasound signals recorded at arrays that can improve monitoring and enhance research capabilities.

58 GEOSCIENCES↗

Observation of Skewed Electromagnetic Wakefields in an Asymmetric Structure Driven by Flat Electron Bunches

Relativistic charged -particle beams that generate intense longitudinal fields in accelerating structures also inherently couple to transverse modes. The effects of this coupling may lead to beam breakup instability and thus must be countered to preserve beam quality in applications such as linear colliders. Beams with highly asymmetric transverse sizes (flat beams) have been shown to suppress the initial instability in slab -symmetric structures. However, as the coupling to transverse modes remains, this solution serves only to delay instability. In order to understand the hazards of transverse coupling in such a case, we describe here an experiment characterizing the transverse effects on a flat beam, traversing near a planar dielectric lined structure. Further, the measurements reveal the emergence of a previously unobserved skew-quadrupolelike interaction when the beam is canted transversely, which is not present when the flat beam travels parallel to the dielectric surface. We deploy a multipole field fitting algorithm to reconstruct the projected transverse wakefields from the data. We generate the effective kick vector map using a simple two -particle theoretical model, with particle -in -cell simulations used to provide further insight for realistic particle distributions.

43 PARTICLE ACCELERATORS↗

Software Control Program For Transportable Microgrid State-of-charge Balancing And Frequency Stability Controls

A deterministic state-of-charge (SOC) balancing approach software control code is introduced as an integral secondary management to primary control layer of an islanded small microgrid or nanogrid system made up of multiple grid-forming inverter/battery/solar combination systems, where each set of batteries with each inverter are on independent DC buses (i.e. non-paralleled on the DC sides). A DERMS-level control approach, algorithm and automation controller program was developed to improve coordination and enable microgrid asset compliance and SOC balancing, enabling provision of a system-level power stability support architecture, load support, and asset scalability. The architecture is configured to treat each unit or micro/nano-grid as a node in a microgrid network, allowing for autonomous DERMS control regarding load and SOC balancing and power stability. As the network grows with the addition of units, greater coordination efforts may be required. The ideal small network microgrid ranges from 2-10 inverter/battery units before additional control parameters must be considered in the existing architecture. The control approach focuses on a deterministic state-of-charge analysis as the primary level control process followed by a secondary control loop using a forced frequency-watt droop strategy to conform off-the-shelf components into behaving under a leader-follower configuration. Adopting this control scheme has been shown to allow for a balanced, unit-coordinated microgrid network, enabling stable power flow. The deterministic state-of-charge approach is introduced as an integral primary control layer of an islanded small network microgrid. A standard strategy for SOC balancing is implementing a battery management system (BMS) to control SOC on the DC side. An alternative approach is to determine how to coordinate sending and receiving power on the AC side with multiple units. The latter approach assesses all the integrated units in the microgrid network. Once the individual units are identified, further system data is required to calculate each unit's total kWh, provided information about its capability to supply or consume kWh and availability. The secondary control layer in the multi-layered small network microgrid methodology uses the primary layer’s decision to initiate frequency setpoint changes, initializing the SOC balancing. The secondary control layer considers numerous system-dependent variables to enable a charging and discharging profile based on adjustable frequency setpoints. The combined architecture will result in stable, coordinated power flow enhancing an AC microgrid's functionalities.

Myers, KurtS [Idaho National Laboratory (INL), Ida↗

System and method of storing and analyzing information

A system and method of storing and analyzing information is disclosed. The system includes a compiler layer to convert user queries to data parallel executable code. The system further includes a library of multithreaded algorithms, processes, and data structures. The system also includes a multithreaded runtime library for implementing compiled code at runtime. The executable code is dynamically loaded on computing elements and contains calls to the library of multithreaded algorithms, processes, and data structures and the multithreaded runtime library.

Feo, John T.↗

Development of MGMC: A proxy Multi-Group Monte Carlo Particle Transport Application

This document details the development and current state of MGMC: a proxy Multi-Group Monte Carlo transport application. The goal of this work was to develop a light-weight Monte Carlo transport solver that could be easily reconfigured to test parallelization strategies targeting advanced architectures and heterogeneous computing environments. Additionally, MGMC is able to act as a proxy-app for testing the development of algorithms and the integration of other libraries and into Monte Carlo transport codes. MGMC is currently able to run in parallel using OpenMP for CPU threads or CUDA for NVIDIA GPUs. MGMC was also used to investigate the using SYCL for portable parallelization targeting either CPU threads, vender-agnostic GPUs, and ARM chips, detailed in Section II. In the process of developing MGMC, several other C++ libraries have been created with the intent for re-use in future research projects and potential inclusion in production-level codes, detailed in Section III. The physics capabilities in MGMC are detailed in Section IV, code verification is detailed in Section V, and performance results are detailed in Section VI.

97 MATHEMATICS AND COMPUTING↗

Estimating the randomness of quantum circuit ensembles up to 50 qubits

Random quantum circuits have been utilized in the contexts of quantum supremacy demonstrations, variational quantum algorithms for chemistry and machine learning, and blackhole information. The ability of random circuits to approximate any random unitaries has consequences on their complexity, expressibility, and trainability. To study this property of random circuits, we develop numerical protocols for estimating the frame potential, the distance between a given ensemble and the exact randomness. Our tensor-network-based algorithm has polynomial complexity for shallow circuits and is high-performing using CPU and GPU parallelism. We study 1. local and parallel random circuits to verify the linear growth in complexity as stated by the Brown–Susskind conjecture, and; 2. hardware-efficient ansätze to shed light on its expressibility and the barren plateau problem in the context of variational algorithms. Our work shows that large-scale tensor network simulations could provide important hints toward open problems in quantum information science.

97 MATHEMATICS AND COMPUTING↗

Practical Implementation of GPU-based Computing at the Grid Edge for Resilience Scenarios

This paper presents a practical implementation of GPU-accelerated computing at the grid edge to enhance power system resilience through next-generation smart meters. Advanced Metering Infrastructure (AMI) systems rely predominantly on centralized processing architectures, which limit real-time response capabilities during grid disturbances. This work proposes the integration of GPU-enabled computational platforms directly within smart meter to enable local execution support for power system analytics, fault detection algorithms, and optimization routines. The proposed framework uses the Julia programming language to leverage highperformance parallel computing capabilities while maintaining code portability and development efficiency. We use two experimental scenarios to benchmark the computational feasibility of this approach: sparse linear system solutions representative of power flow analyses, and multi-stage production cost simulations incorporating unit commitment and economic dispatch operations. Results demonstrate that computationally intensive power system algorithms, such as those supporting resilience scenario calculations, can be effectively executed at the distribution edge using commercially available embedded GPU hardware. Keywords—GPU acceleration, edge computing, smart meters, grid resilience, AMI, resilience.

De Souza, Reubun [School of Electrical Engineering↗

Quantum Simulators and Applications on Quantum Framework

Simulating quantum circuits is essential for validating quantum algorithms. However, no single simulator consistently performs best - efficiency depends on circuit structure, entanglement, and depth. In this work, we integrate Qiskit-Aer (state-vector and matrix product state) and QTensor, a tree-tensor-network based simulator, into the Quantum Framework (QFw), a modular platform that supports multiple quantum backends via a unified interface. We also enable distributed quantum approximate optimization algorithm (DQAOA) application compatibility with QFw, allowing sub-problems to be solved in parallel at scale. We then benchmark DQAOA and TFIM (transverse field Ising model) circuits across supported simulators, showing how performance varies significantly with problem type. All simulations are deployed on the Frontier supercomputer using QFw's MPI-based orchestration for distributed, multinode execution. These results underscore the need for simulatoragnostic infrastructure to enable systematic evaluation and highperformance scaling of quantum workloads. QFw provides a practical and extensible path toward reproducible quantum algorithm development across diverse application domains.

Chundury, Srikar [ORNL] (ORCID:0009000183359259)↗

PySpawn: Software for Nonadiabatic Quantum Molecular Dynamics

The ab initio multiple spawning (AIMS) method enables nonadiabatic quantum molecular dynamics simulations in an arbitrary number of dimensions, with potential energy surfaces provided by electronic structure calculations performed on-the-fly. However, the intricacy of the AIMS algorithm complicates software development, deployment on modern shared computer resources, and post-simulation data analysis. PySpawn is a nonadiabatic molecular dynamics software package that addresses these issues. Here, the program is designed to be easily interfaced with electronic structure software, and an interface to the TeraChem software package is described here. PySpawn introduces a task-based reorganization of the AIMS algorithm, allowing fine-grained restart capability and setting the stage for efficient parallelization in a future release. PySpawn includes a user-friendly and interactive Python analysis module that will enable novice users to painlessly adopt AIMS. As a demonstration of PySpawn’s simulation capability and analysis module, we report complete active space self-consistent field–based AIMS simulations of the 1,2- dithienyl-1,2-dicyanoethene molecule, a promising molecular photoswitch.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

C-SAW: a framework for graph sampling and random walk on GPUs

Many applications require to learn, mine, analyze and visualize large-scale graphs. These graphs are often too large to be addressed efficiently using conventional graph processing technologies. Fortunately, recent research efforts find out graph sampling and random walk, which significantly reduce the size of original graphs, can benefit the tasks of learning, mining, analyzing and visualizing large graphs by capturing the desirable graph properties. This paper introduces C-SAW, the first framework that accelerates Sampling and Random Walk framework on GPUs. Particularly, C-SAW makes three contributions: First, our framework provides a generic API which allows users to implement a wide range of sampling and random walk algorithms with ease. Second, offloading this framework on GPU, we introduce warp-centric parallel selection, and two novel optimizations for collision migration. Third, towards supporting graphs that exceed the GPU memory capacity, we introduce efficient data transfer optimizations for out-of-memory and multi-GPU sampling, such as workload-aware scheduling and batched multi-instance sampling. Taken together, our framework constantly outperforms the state of the art projects in addition to the capability of supporting a wide range of sampling and random walk algorithms.

97 MATHEMATICS AND COMPUTING↗

A supernodal all-pairs shortest path algorithm

We show how to exploit graph sparsity in the Floyd-Warshall algorithm for the all-pairs shortest path (Apsp) problem. Floyd-Warshall is an attractive choice for Apsp on high-performing systems due to its structural similarity to solving dense linear systems and matrix multiplication. However, if sparsity of the input graph is not properly exploited, Floyd-Warshall will perform unnecessary asymptotic work and thus may not be a suitable choice for many input graphs. To overcome this limitation, the key idea in our approach is to use the known algebraic relationship between Floyd-Warshall and Gaussian elimination, and import several algorithmic techniques from sparse Cholesky factorization, namely, fill-in reducing ordering, symbolic analysis, supernodal traversal, and elimination tree parallelism. When combined, these techniques reduce computation, improve locality and enhance parallelism. We implement these ideas in an efficient shared memory parallel prototype that is orders of magnitude faster than an efficient multi-threaded baseline Floyd-Warshall that does not exploit sparsity. Our experiments suggest that the Floyd-Warshall algorithm can compete with Dijkstra's algorithm (the algorithmic core of Johnson's algorithm) for several classes sparse graphs.

Sao, Piyush↗

Grain structure and texture selection regimes in metal powder bed fusion

Additive manufacturing (AM) offers opportunities to produce complex part geometries not possible with conventional processing and in some cases even improve part performance. However, adoption has been slowed by difficulties assessing microstructure variability and there is no straightforward approach to relate processing to grain structure characteristics. In this study, datasets from AdditiveFOAM heat transport simulations of laser powder bed fusion (LPBF) are used to drive ExaCA simulations of grain structure. The GPU utilization of ExaCA and an algorithmic update for modeling melt pool overlap region solidification enabled rapid and parallel simulation across a wider range of process conditions than previously explored with cellular automata-based solidification models. A texture selection angle $θ_s$ is defined based on melt pool overlap geometry, and the range of $θ_s$ over which a commonly observed texture transition occurs in characterized AM builds was well-reproduced by ExaCA simulations over a wide range of melt pool shape, hatch spacing, and layer height. ExaCA simulations with 90 degree rotation of the scan direction on every other layer reproduced a number of trends from the AM literature including grain refinement, the dominance of layers with larger melt pools on the final grain structure, and the weakening or strengthening of texture depending on odd and even layer melt pool overlap geometry. EBSD data from a benchmark AM part is used to validate the simulated mechanism of a layer rotation-induced texture strengthening effect. Importantly, these results expand the understanding of the mechanisms for texture selection in alloys with cubic crystal symmetry and offer an approach to easily evaluate processing conditions. With this new understanding, these modeling tools will enable anticipation of previously unexpected variations in grain structure and target specific microstructures and properties.

36 MATERIALS SCIENCE↗

Refining HPCToolkit for application performance analysis at exascale

As part of the US Department of Energy’s Exascale Computing Project (ECP), Rice University has been refining its HPCToolkit performance tools to better support measurement and analysis of applications executing on exascale supercomputers. To efficiently collect performance measurements of GPU-accelerated applications, HPCToolkit employs novel non-blocking data structures to communicate performance measurements between tool threads and application threads. To attribute performance information in detail to source lines, loop nests, and inlined call chains, HPCToolkit performs parallel analysis of large CPU and GPU binaries involved in the execution of an exascale application to rapidly recover mappings between machine instructions and source code. To analyze terabytes of performance measurements gathered during executions at exascale, HPCToolkit employs distributed-memory parallelism, multithreading, sparse data structures, and out-of-core streaming analysis algorithms. To support interactive exploration of profiles up to terabytes in size, HPCToolkit’s hpcviewer graphical user interface uses out-of-core methods to visualize performance data. The result of these efforts is that HPCToolkit now supports collection, analysis, and presentation of profiles and traces of GPU-accelerated applications at exascale. These improvements have enabled HPCToolkit to efficiently measure, analyze and explore terabytes of performance data for executions using as many as 64K MPI ranks and 64K GPU tiles on ORNL’s Frontier supercomputer. HPCToolkit’s support for measurement and analysis of GPU-accelerated applications has been employed to study a collection of open-science applications developed as part of ECP. This paper reports on these experiences, which provided insight into opportunities for tuning applications, strengths and weaknesses of HPCToolkit itself, as well as unexpected behaviors in executions at exascale.

Adhianto, Laksono↗

DEEP Solar: Data DrivEn Modeling and Analytics for Enhanced System Layer ImPlementation

Realizing the SETO 2030 mission of reducing solar energy costs to 3-5 c/kWh will require innovative enabling research on effective, cost-efficient integration of local PV within distribution systems. However, the intermittent and variable nature of PVs compels operators to impose conservative hosting capacity constraints. Given the extremely high variability of (intermittent and unpredictable) solar energy generation, relaxing the capacity constraints (which are currently around 15%) and achieving 100% or greater integration of renewables will require a fundamental transformation of the power grid via the utilization of exponentially larger amounts of AMI enabled fine-grained data. To address the challenges in increasing the penetration of renewable energy based DERs, this project envisions an Enhanced System Layer (ESL) at the distribution network level that is reliable, cost-effective and scalable to millions of Distributed Energy Resources (DERs)/devices. This includes developing: 1) Transformative and highly scalable machine learning based predictive analytics tools that plug into distribution system planning and provide real-time situational awareness at the distribution level for short and long-term operational planning. The tools will be built using novel data-driven energy models of millions of active nodes with AMI, 2) Adaptive stochastic analysis and optimization algorithms for real-time grid operations, 3) Dynamic Scenario Analysis using parallel Cloudenabled implementations with < 1 minute computational cycle times.

14 SOLAR ENERGY↗

Efficient Scalable Contact Network Generation from Population Data

Modeling the contacts among a population is critical to understanding the dynamics of a disease outbreak. Contact networks, where nodes are individuals and edges are contacts among them, are used to represent these complex individual-level interactions. In this work, we are given the daily activity schedules of an urban population that represent the activity location and time of individuals in a population during a single twenty four hour period over multiple days. Using collocation to determine contact between individuals, our goal is to extract hourly contact networks from large-scale activity data. We improve upon the existing adjacency matrix-based method by implementing our custom sparse matrix multiplication algorithm. Starting with a Python implementation, we achieve a 1600x speed up in the computation with a fast custom designed sparse matrix multiplier algorithm implemented in the C++ language. This work is central to future parallel designs of the problem.

97 MATHEMATICS AND COMPUTING↗

Tensor Extraction of Latent Features (TELF)

Tensor ELF is a user-friendly parallel tensor decomposition Python toolbox that includes a suite of machine learning algorithms for CPU and GPU architectures for the analysis of sparse and dense data including utility tools for pre-processing and post-processing.

Eren, Maksim↗