Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “asynchronous algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

36 records · Page 2

Framework for Extensible, Asynchronous Task Scheduling (FEATS) in Fortran

Most parallel scientific programs contain compiler directives (pragmas) such as those from OpenMP, explicit calls to runtime library procedures such as those implementing the Message Passing Interface (MPI), or compiler-specific language extensions such as those provided by CUDA. By contrast, the recent Fortran standards empower developers to express parallel algorithms without directly referencing lower-level parallel programming models. Fortran’s parallel features place the language within the Partitioned Global Address Space (PGAS) class of programming models. When writing programs that exploit data-parallelism, application developers often find it straightforward to develop custom parallel algorithms. Problems involving complex, heterogeneous, staged calculations, however, pose much greater challenges. Such applications require careful coordination of tasks in a manner that respects dependencies prescribed by a directed acyclic graph. When rolling one’s own solution proves difficult, extending a customizable framework becomes attractive. The paper presents the design, implementation, and use of the Framework for Extensible Asynchronous Task Scheduling (FEATS), which we believe to be the first task-scheduling tool written in modern Fortran. We describe the benefits and compromises associated with choosing Fortran as the implementation language, and we propose ways in which future Fortran standards can best support the use case in this paper.

Richardson, Brad↗

Distributed approximate minimal Steiner trees with millions of seed vertices on billion-edge graphs

In this report, we present a parallel 2-approximation Steiner minimal tree algorithm and its MPI-based distributed implementation. In place of expensive distance computations between all pairs of seed vertices, the solution we employ exploits a cheaper Voronoi cell computation. Our design leverages asynchronous processing and message prioritization to accelerate convergence of distance computations, and harnesses vertex and edge centric processing to offer fast time-to-solution. We demonstrate scalability and performance using real-world graphs with up to 128 billion edges and 512 compute nodes, and show the ability to find Steiner trees with up to one million seed vertices. Using 12 data instances, we present comparison with the state-of-the-art exact solver, SCIP-Jack, and two sequential 2-approximate algorithms. We empirically show that, on average, the total distance of the Steiner tree identified by our solution is 1.1290 times greater than the Steiner minimal tree – well within the theoretical approximation bound of 2.

97 MATHEMATICS AND COMPUTING↗

Rethinking Programming Paradigms in the QC-HPC Context

Programming for today’s quantum computers is making significant strides toward modern workflows compatible with high performance computing (HPC), but fundamental challenges still remain in the integration of these vastly different technologies. Quantum computing (QC) programming languages share some common ground, as well as their emerging runtimes and algorithmic modalities. In this short paper, we explore avenues of refinement for the quantum processing unit (QPU) in the context of many-tasks management, asynchronous or otherwise, in order to understand the value it can play in linking QC with HPC. Through examples, we illustrate how its potential for scientific discovery might be realized.

Wong, Elaine↗

Envisioning the Future Renewable and Resilient Energy Grids—A Power Grid Revolution Enabled by Renewables, Energy Storage, and Energy Electronics

Today’s power grids are facing tremendous challenges because of the ever-increasing power demand, system complexity, infrastructure cost, knowledge base, and policy and regulatory issues to achieve supply–demand power balance and resiliency with respect to more frequent extreme weather events and cyberattacks. It is particularly challenging when the transition toward 100% intermittent renewable energy sources is considered. Many countries are calling for building up more transmission and distribution lines to increase power delivery capacities. This article is an attempt to answer two urgent questions: Is more transmission and distribution infrastructure really needed to meet the increasing power demand? What kind of future grid infrastructure should we envision and build? This article attempts to answer these questions and proposes the concept of community-centric asynchronous renewable and resilient energy grids. By clearly differentiating the concepts of grid resilience and reliability, the importance of building resilient power electronics’ devices and robust system-level control algorithms to achieve 100% renewable energy integrated resilient grids is presented. To identify the shortcomings and propose advancements, power electronics’ technologies are categorized using the proposed concepts of natural source frequencies (NSf), energy storage, direct energy conversion/control and fault protection (DeCaFp), and high-efficiency energy consumption and buffering (heECaB) technology. The ability of networked microgrids to greatly reduce power outages and power system restoration time is demonstrated by leveraging robust decentralized and centralized control algorithms, identified through a comprehensive literature review. Future research areas are proposed to further enhance grid stability, controllability, cybersecurity, and protection against faults in the presence of 100% renewable sources by leveraging the advanced capabilities of NSf, DeCaFp, and heECaB devices and system-level control algorithms.

14 SOLAR ENERGY↗

Introduction: Neuromorphic Materials

The explosive growth in data collection and the need to process it efficiently, as well as the desire to automate increasingly complex tasks in transportation, medical care, manufacturing, security and many other fields have motivated a growing interest in neuromorphic computing. Unlike the binary, transistorbased ON/OFF logic gates and separate logic and memory functionalities employed in digital computing, neuromorphic computing is inspired by animal brains that use interconnected synapses and neurons to perform processing, storage and transmission of information at the same location, while only consuming ~20 W or less of power. Motivated by the brain’s efficiency, adaptability, self-learning and resiliency qualities, neuromorphic computing can be broadly defined as an approach to processing and storing information using hardware and algorithms inspired by models of biological neural systems. Present research in neuromorphic computing encompasses approaches that vary significantly in their degree of neuro-inspiration, from systems that only incorporate features such as asynchronous, event-driven operation or use crossbar arrays of non-volatile memory (NVM) elements to accelerate deep neural networks (DNNs), to designs that embrace the extreme parallelism, sparsity, reconfigurability, adaptability, complexity and stochasticity observed in nervous systems. The term ‘neuromorphic’ computing is often credited to Carver Mead, who in the 1980s investigated Si-based analog electronics to replicate functions of the animal retina. Earlier important advances in this field include the work of Frank Rosenblatt, who proposed the concept of the perceptron, Bernard Widrow, who used this concept to build one of the first analog neural networks, the Adaline and many other researchers (see ref. 6 for an historical perspective on neuromorphic computing). With the recent increase in the use of artificial intelligence and large language models, and rising concerns over the associated energy costs, interest in neuromorphic hardware has expanded rapidly. According to some estimates, driven largely by the drastic growth in the training use of artificial intelligence (AI) models using the current computing architectures, the energy cost of computing is projected to reach the energy supply worldwide by 2045. Furthermore, while this is not a realistic outcome, it means that, if more efficient computing technologies are not developed -- soon -- the world will soon become one where demand for energy and market constraints limit the continued increase of societal access to AI and cloud services from data centers. Data centers used for training and use of these models consume hundreds of terawatt hours of electricity, already past 4% of the US electricity demand.

Circuits↗

A Surrogate-Based Asynchronous Decomposition Technique for Realistic Security-Constrained Optimal Power Flow Problems

Here we present a decomposition approach for obtaining good feasible solutions for the security-constrained, alternating-current, optimal power flow (SC-AC-OPF) problem at an industrial scale and under real-world time and computational limits. The approach was designed while preparing and participating in ARPA-E’s Grid Optimization Competition (GOC) Challenge 1. The challenge focused on a near-real-time version of the SC-AC-OPF problem, where a base operating point is optimized, taking into account possible single-element contingencies, after which the system adapts its operating point following the response of automatic frequency droop controllers and voltage regulators. Our solution approach for this problem relies on state-of-the-art nonlinear programming algorithms, and it employs nonconvex relaxations for complementarity constraints, a specialized two-stage decomposition technique with sparse approximations of recourse terms and contingency ranking and prescreening. The paper describes and justifies our approach and outlines the features of its implementation, including functions and derivatives evaluation, warm-starting strategies, and asynchronous parallelism. We discuss the results of the independent benchmark of our approach by ARPA-E’s GOC team in Challenge 1, where it was found to consistently produce high-quality solutions across a wide range of network sizes and difficulty, and conclude by outlining future extensions of the approach.

97 MATHEMATICS AND COMPUTING↗

Scheduling and Performance of Asynchronous Tasks in Fortran 2018 with FEATS

Most parallel scientific programs contain compiler directives (pragmas) such as those from OpenMP (Hermanns in Parallel programming in Fortran 95 using openMP, 2002. School of Aeronautical Engineering, Universidad Politécnica de Madrid, España, 2011), explicit calls to runtime library procedures such as those implementing the Message Passing Interface (MPI) (in A message-passing interface standard version 4.0, 2021. https://www.mpi-forum.org/docs/mpi-4.0/mpi40-report.pdf), or compiler-specific language extensions such as those provided by CUDA (Ruetsch and Fatica in CUDA Fortran for scientists and engineers: best practices for efficient CUDA Fortran programming, Elsevier, 2013). By contrast, the recent Fortran standards empower developers to express parallel algorithms without directly referencing lower-level parallel programming models (Numrich in Parallel programming with co-arrays, CRC Press, 2018, and Curcic in Modern Fortran: building efficient parallel applications, Manning Publications, 2020). Fortran’s parallel features place the language within the Partitioned Global Address Space (PGAS) class of programming models. When writing programs that exploit data parallelism, application developers often find it straightforward to develop custom parallel algorithms. Problems involving complex, heterogeneous, staged calculations, however, pose much greater challenges. Such applications require careful coordination of tasks in a manner that respects dependencies prescribed by a directed acyclic graph. When rolling one’s own solution proves difficult, extending a customizable framework becomes attractive. Further, the paper presents the design, implementation, and use of the Framework for Extensible Asynchronous Task Scheduling (FEATS), which we believe to be the first task scheduling tool written in modern Fortran. We describe the benefits and compromises associated with choosing Fortran as the implementation language, and we propose ways in which future Fortran standards can best support the use case in this paper.

97 MATHEMATICS AND COMPUTING↗

Image Gradient Decomposition for Parallel and Memory-Efficient Ptychographic Reconstruction

Ptychography is a popular microscopic imaging modality for many scientific discoveries and sets the record for highest image resolution. Unfortunately, the high image resolution for ptychographic reconstruction requires significant amount of memory and computations, forcing many applications to compromise their image resolution in exchange for a smaller memory footprint and a shorter reconstruction time. In this paper, we propose a novel image gradient decomposition method that significantly reduces the memory footprint for ptychographic reconstruction by tessellating image gradients and diffraction measurements into tiles. In addition, we propose a parallel image gradient decomposition method that enables asynchronous point-to-point communications and parallel pipelining with minimal overhead on a large number of GPUs. Our experiments on a Titanate material dataset (PbTiO3) with 16632 probe locations show that our Gradient Decomposition algorithm reduces memory footprint by 51 times. In addition, it achieves time-to-solution within 2.2 minutes by scaling to 4158 GPUs with a super-linear strong scaling efficiency at 364% compared to runtimes at 6 GPUs. This performance is 2.7 times more memory efficient, 9 times more scalable and 86 times faster than the state-of-the-art algorithm.

Wang, Xiao↗

State Estimation for Distribution Networks with Asynchronous Sensors Using Stochastic Descent: Preprint

This paper investigates the problem of state estimation for distribution networks with asynchronous sensors comprising of a mix of smart meters and phasor measurement units (PMUs) with multiple sampling and reporting rates. We consider two independent scenarios of state estimation and tracking, with either voltages or currents as states. With these two sets, we investigate estimation under (a) full data, assuming all measurements are available and (b) limited data, where an online algorithmic approach is adopted to estimate the possibly time-varying states by processing measurements as and when available. The proposed algorithm, inspired by the classical Stochastic Gradient Descent (SGD) approach updates the states based on the previous estimate and the newly available measurements. Finally, we demonstrate the estimation and tracking efficacy through numerical simulations on the IEEE-37 test network, while also highlighting how estimation with currents as states leads to faster convergence.

asynchronous sensors↗

Radiation-Induced Noise Resilience of Neuromorphic Architectures

Neuromorphic event-based networks use asynchronous time-dependent information to extract features from input data that can allow for edge-based distributed applications such as object recognition. The noise resilience properties of such networks, especially in the context of space applications, are yet to be explored. In this paper, we use the hierarchy of time surfaces (HOTS) algorithm, which is one of the neuromorphic algorithms, to understand the least and most resilient modules in a neuromorphic network. The HOTS algorithm relies on the computing of time surfaces that maps the temporal delays between neighboring pixels into normalized features that involve many computations that are also found in other neuromorphic networks such as exponential decays, distance computations, etcetera. We implemented HOTS on a Digilent PYNQ board with a Xilinx Zynq 7020 system on a chip, and we subjected the boards running the HOTS network inference to neutron radiation at the Los Alamos Neutron Science Center. Furthermore, we used simulation models from our previous similar experiments on the event-based sensor to create a neutron induced noise model to quantify the effect of this noise on the overall performance of the network. This experiment provides the preliminary measurements of the reliability of the HOTS algorithm and proposes methods to create a more reliable HOTS architecture in future spacecraft missions.

Engineering↗

Asynchronous GPU-based DEM solver embedded in commercial CFD software with polyhedral mesh support

A novel graphical processing unit-based discrete element method solver is introduced to improve stability, performance, and provide seamless integration into commercial or open-source computational fluid dynamics software. A key innovation is eliminating a need for network communication between solvers, which was previously required for cross-platform coupling. This is accomplished by a direct coupling method that employs dynamic-linked libraries. Furthermore, the solver optimizes memory usage by streamlining the particle-cell search algorithm by eliminating the cells' searching grid. This ensures the solver is compatible with a wide range of mesh types, providing high geometric flexibility. The approach simplifies the simulation process by directly incorporating computational fluid dynamics mesh information into the discrete element method solver. The performance analysis indicates about sixteen times boost in computational speed compared to benchmark central processing unit-based solvers. Finally, the solver's compatibility with polyhedral meshes, a vital advantage for complex geometries, is tested against a referenced study regarding the simulation of an immersed-tube fluidized bed.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Online State Estimation for Time-Varying Systems

The paper investigates the problem of estimating the state of a time-varying system with a linear measurement model; in particular, the paper considers the case where the number of measurements available can be smaller than the number of states. In lieu of a batch linear least-squares (LS) approach well-suited for static networks, where a sufficient number of measurements could be collected to obtain a full-rank design matrix the paper proposes an online algorithm to estimate the possibly time-varying state by processing measurements as and when available. The design of the algorithm hinges on a generalized LS cost augmented with a proximal-point-type regularization. With the solution of the regularized LS problem available in closed-form, the online algorithm is written as a linear dynamical system where the state is updated based on the previous estimate and based on the new available measurements. Conditions under which the algorithmic steps are in fact a contractive mapping are shown, and bounds on the estimation error are derived for different noise models. Numerical simulations are provided to corroborate the analytical findings.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Distributed out-of-memory NMF on CPU/GPU architectures

We propose an efficient distributed out-of-memory implementation of the non-negative matrix factorization (NMF) algorithm for heterogeneous high-performance-computing systems. The proposed implementation is based on prior work on NMFk, which can perform automatic model selection and extract latent variables and patterns from data. In this work, we extend NMFk by adding support for dense and sparse matrix operation on multi-node, multi-GPU systems. The resulting algorithm is optimized for out-of-memory problems where the memory required to factorize a given matrix is greater than the available GPU memory. Memory complexity is reduced by batching/tiling strategies, and sparse and dense matrix operations are significantly accelerated with GPU cores (or tensor cores when available). Input/output latency associated with batch copies between host and device is hidden using CUDA streams to overlap data transfers and compute asynchronously, and latency associated with collective communications (both intra-node and inter-node) is reduced using optimized NVIDIA Collective Communication Library (NCCL) based communicators. Benchmark results show significant improvement, from 32X to 76x speedup, with the new implementation using GPUs over the CPU-based NMFk. Good weak scaling was demonstrated on up to 4096 multi-GPU cluster nodes with approximately 25,000 GPUs when decomposing a dense 340 Terabyte-size matrix and an 11 Exabyte-size sparse matrix of density 10 -6 .

97 MATHEMATICS AND COMPUTING↗

tomoCAM : fast model-based iterative reconstruction via GPU acceleration and non-uniform fast Fourier transforms

X-ray-based computed tomography is a well established technique for determining the three-dimensional structure of an object from its two-dimensional projections. In the past few decades, there have been significant advancements in the brightness and detector technology of tomography instruments at synchrotron sources. These advancements have led to the emergence of new observations and discoveries, with improved capabilities such as faster frame rates, larger fields of view, higher resolution and higher dimensionality. These advancements have enabled the material science community to expand the scope of tomographic measurements towards increasingly in situ and in operando measurements. In these new experiments, samples can be rapidly evolving, have complex geometries and restrictions on the field of view, limiting the number of projections that can be collected. In such cases, standard filtered back-projection often results in poor quality reconstructions. Iterative reconstruction algorithms, such as model-based iterative reconstructions (MBIR), have demonstrated considerable success in producing high-quality reconstructions under such restrictions, but typically require high-performance computing resources with hundreds of compute nodes to solve the problem in a reasonable time. Here, tomoCAM , is introduced, a new GPU-accelerated implementation of model-based iterative reconstruction that leverages non-uniform fast Fourier transforms to efficiently compute Radon and back-projection operators and asynchronous memory transfers to maximize the throughput to the GPU memory. The resulting code is significantly faster than traditional MBIR codes and delivers the reconstructive improvement offered by MBIR with affordable computing time and resources. tomoCAM has a Python front-end, allowing access from Jupyter -based frameworks, providing straightforward integration into existing workflows at synchrotron facilities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

IPC-Fusion (Infrastructure Perception and Control (IPC): Multisensor Data Fusion Software) [SWR-25-153]

As part of the National Laboratory of the Rockies' (NLR’s) Infrastructure Perception and Control Laboratory, the IPC-Fusion toolkit provides a probabilistic, scalable, multi-sensor fusion framework that integrates (late-stage fusion) heterogeneous object detection data from traffic sensors to enable robust, real-time tracking of roadway occupants. The algorithmic design of the toolkit is motivated by the need for creating a digital twin of traffic at the edge in a scalable and affordable manner. The software operates by combining object-level measurements (such as position and velocity) from a suite of sensors (such as radar, lidar, camera) using Kalman filtering and probabilistic data association techniques to overcome individual sensor limitations and achieve superior tracking performance in complex traffic zones. The framework addresses key challenges including heterogeneous measurement uncertainties, asynchronous data streams, varying spatiotemporal data resolutions, robust data association, and adaptive object lifecycle management. Validated on real-world traffic intersection data including vehicles and pedestrians, IPC-Fusion demonstrates enhanced tracking reliability across scenarios involving occlusions, sensor failures, and varying traffic densities, supporting the broader IPC initiative's goal of transforming transportation infrastructure through advanced perception capabilities for intelligent transportation systems, traffic safety applications, and autonomous vehicle support.

Sandhu, Rimple [National Laboratory of the Rockies↗

Neutron transport methods for multiphysics heterogeneous reactor core simulation in Griffin

Griffin is a reactor physics application based on the Multiphysics Object-Oriented Simulation Environment (MOOSE). This work discloses the methods, algorithms, and implementation for simulating heterogeneous reactor dynamics models. Griffin utilizes a discontinuous finite-element method with discrete ordinates (DFEM-S ) to discretize the field variable of the multigroup neutron transport equation. Multiphysics feedback is handled using two-step tabulated cross-section methodology. Feedback quantities are evaluated using the MOOSE-MultiApp system to couple various engineering phenomena, such as heat conduction and thermal fluids. The multiphysics DFEM-S system is solved using fixed-point iteration with a fully asynchronous parallel sweeper, unstructured coarse-mesh finite difference acceleration, and a multi-timescale improved quasi-static method scheme. The implementation is applied to a multiphysics microreactor model, with two transients: one initiated by a single heat-pipe failure and another by control drum rotation. Importantly, these examples demonstrate the ability of Griffin to tractably solve the neutron transport equation considering seven independent variables and feedback.

97 MATHEMATICS AND COMPUTING↗

AMReX v2024

The software framework, AMReX, supports the development of block-structured adaptive mesh refinement (AMR) algorithms for solving systems of partial differential equations. AMR reduces the computational cost and memory footprint compared to a uniform mesh while preserving the essential local descriptions of different physical processes in complex multiphysics algorithms. AMR uses a hierarchical representation of the solution at multiple levels of resolution where the solution on each level is defined on the union of data containers at that resolution. These data containers, which represent the solution over a logically rectangular subregion of the domain, can contain field data defined on a mesh, Lagrangian particles or combinations of both. In addition to these basic data types, AMReX supports a multilevel embedded boundary representation of complex geometry; linear solvers for cell-centered and nodal data; asynchronous I/O in a native format readable by ParaView, VisIt and yt; and interfaces to hypre and PETSc solvers. AMReX enables applications to run on distributed memory architectures with multicore CPUs and with GPU accelerators. AMReX uses a lightweight abstraction layer that effectively hides the details of the architecture from the application. The framework currently supports CUDA, HIP and SYCL for GPU acceleration and OpenMP for multi-core CPU architectures.

Almgren, Ann↗

A High-Fidelity Electromagnetic Transient Model of Inverter-based Resources Integrated to an IEEE-9 Bus System for Benchmarking Studies

An inverter-based resource (IBR) is a source of electricity that is asynchronously connected to the electrical grid via an electronic power converter. These power sources lack the intrinsic behaviors of the standard power plants, presenting specific challenges to system stability. These power sources lack the intrinsic behaviors of the standards power plants and their features are almost entirely defined by the control strategy used on them, presenting specific challenges to system stability as their penetration increases into the already fragile Bulk Power System (BPS). Utilizing renewable energy has a great upside but the electrical industry needs to understand the requirements like performing electromagnetic transient (EMT) simulation studies in planning and/or in interconnection studies. This paper presents the development of a high-fidelity EMT model based on commercial equipment and field parameters, which at the same time aims to study the increasing penetration of IBRs and their impact on the actual BPS. The open-source EMT model developed will be a benchmarking example that can be used to evaluate algorithms in the power grid (including evaluating grid-forming controls, power system monitoring and operation, etc.).

Martinez Montejano, Misael↗