Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “asynchronous”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Real-Time Distribution System State Estimation with Asynchronous Measurements

We report state estimation is a fundamental task in power systems. Although distribution systems are increasingly equipped with sensing devices and smart meters, measurements are typically reported at different rates and asynchronously; these aspects pose severe strains on workhorse state estimation algorithms, which are designed to process batches of data collected in a synchronous manner from all the measurement units. In this paper, we develop a novel state estimation algorithm to continuously update the estimate of the state based on measurements received in an asynchronous manner from measurement units. The synthesis of the algorithm hinges on a proximal-point type method, implemented in an online fashion, and capable of processing measurements received sequentially from sensors. A performance analysis is presented by providing bounds on the estimation error in terms of the mean and variance that hold at each iteration and asymptotically. The scheme is also compared with a more traditional Weighted Least Squares estimator that compensates for the lack of measurement data by using, as pseudo measurements, the measurement retrieved during a certain time window. Numerical simulations on the IEEE 37-bus feeder corroborate the analytical findings.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Modifying the Asynchronous Jacobi Method for Data Corruption Resilience

Moving scientific computation from high-performance computing (HPC) and cloud computing (CC) environments to devices on the edge, i.e., physically near instruments of interest, has received tremendous interest in recent years. Such edge computing environments can operate on data in situ, offering enticing benefits over data aggregation to HPC and CC facilities that include avoiding costs of transmission, increased data privacy, and real-time data analysis. Because of the inherent unreliability of edge computing environments, new fault-tolerant approaches must be developed before the benefits of edge computing can be realized. Motivated by algorithm-based fault tolerance, a variant of the asynchronous Jacobi (ASJ) method is developed that achieves resilience to data corruption by rejecting solution approximations from neighbor devices according to a bound derived from convergence theory. Numerical results on a two-dimensional Poisson problem show that the new rejection criterion, along with a novel approximation to the shortest path length on which the criterion depends, restores convergence for the ASJ variant in the presence of certain types data corruption. Numerical results are obtained for when the singular values in the analytic bound are approximated. Additional linear systems are also explored, one with a more dense sparsity pattern and one that includes advection. All results indicate that successful resilience to data corruption depends on whether the bound tightens fast enough to reject corrupted data before the iteration evolution deviates significantly from that predicted by the convergence theory defining the bound. This observation generalizes to future work on algorithm-based fault tolerance for other asynchronous algorithms, including upcoming approaches that leverage Krylov subspaces.

97 MATHEMATICS AND COMPUTING↗

Asynchronous domain decomposition methods for nonlinear PDEs

One- and two-level parallel asynchronous methods for the numerical solution of nonlinear systems of equations, especially those arising from (nonlinear) partial differential equations, are studied. The proposed methods are based on domain decomposition techniques. Local convergence theorems are presented in several cases, with appropriate hypotheses. Computational results on a shared memory multiprocessor machine for various problems exhibiting nonlinearities are reported, illustrating the potential of these asynchronous methods, especially for heterogeneous clusters.

97 MATHEMATICS AND COMPUTING↗

Ability to Simulate Absorption and Melt Pool Dynamics for Laser Melting of Bare Aluminum Plate: Results and Insights from the 2022 Asynchronous AM-Bench Challenge

The 2022 Asynchronous AM-Bench challenge was designed to test the ability of simulations to accurately predict laser power absorption as well as various melt pool behaviors (width, depth, and solidification) during laser melting of solid metal during stationary and scanned laser illumination. In this challenge, participants were asked to predict a series of experimental outcomes. Experimental data were obtained from a series of experiments performed at the Advanced Photon Source at Argonne National Laboratories in 2019. These experiments combined integrating sphere radiometry with high-speed X-ray imaging, allowing for the simultaneous recording of absolute laser power absorption and two-dimensional, projected images of the melt pool. All challenge problems were based on experiments using bare aluminum solid metal. Participants were provided with pertinent experimental information like laser power, scan speed, laser spot size, and material composition. Additionally, participants were given absorptance and X-ray imaging data from stationary and scanned laser experiments on solid Ti–6Al–4V that could be used for testing their models before attempting challenge problems. In total, this challenge received 56 submissions from eight different research groups for eight individual challenge problems. The data for this challenge, and associated information, are available for download from the NIST Public Data Repository. This paper summarizes the results from the 2022 Asynchronous AM-Bench challenge as well as discusses the lessons learned to help inform future challenges.

36 MATERIALS SCIENCE↗

Asynchronous distributed-memory task-parallel algorithm for compressible flows on unstructured 3D Eulerian grids

Here, we discuss the implementation of a finite element method, used to numerically solve the Euler equations of compressible flows, using an asynchronous runtime system (RTS). The algorithm is implemented for distributed-memory machines, using stationary unstructured 3D meshes, combining data-, and task-parallelism on top of the Charm++ RTS. Charm++’s execution model is asynchronous by default, allowing arbitrary overlap of computation and communication. Task-parallelism allows scheduling parts of an algorithm independently of, or dependent on, each other. Built-in automatic load balancing enables continuous redistribution of computational load by migration of work units based on real-time CPU load measurement. The RTS also features automatic checkpointing, fault tolerance, resilience against hardware failure, and supports power-, and energy-aware computation. We demonstrate scalability up to 25 x 10 9 cells at $\mathscr{O}$10 4 compute cores and the benefits of automatic load balancing for irregular workloads. The full source code with documentation is available at https://quinoacomputing.org.

42 ENGINEERING↗

Asynchronous quadratic control for constrained hidden markov jump linear systems with incomplete MTPM and MOCPM

Abstract This paper investigates the quadratic optimal control problem for constrained Markov jump linear systems with incomplete mode transition probability matrix (MTPM). Considering original system mode is not accessible, observed mode is utilized for asynchronous controller design where mode observation conditional probability matrix (MOCPM), which characterizes the emission between original modes and observed modes is assumed to be partially known. An LMI optimization problem is formulated for such constrained hidden Markov jump linear systems with incomplete MTPM and MOCPM. Based on this, a feasible state-feedback controller can be designed with the application of free-connection weighting matrix method. The desired controller, dependent on observed mode, is an asynchronous one which can minimize the upper bound of quadratic cost and satisfy restrictions on system states and control variables. Furthermore, clustering observation where observed modes recast into several clusters, is explored for simplifying the computational complexity. Numerical examples are provided to illustrate the validity.

Zhu, Jin↗

Asynchronous I/O VOL Connector (AsyncVOL) v0.1

Asynchronous I/O is becoming increasingly popular with the large amount of data access required by scientific applications. They can take advantage of an asynchronous interface by scheduling I/O as early as possible and overlap computation or communication with I/O operations, which hides the cost associated with I/O and improves the overall performance.

Tang, Houjun↗

Highly Asynchronous Visitor Queue Graph Toolkit

HavoqGT (Highly Asynchronous Visitor Queue Graph Toolkit) is a framework for expressing asynchronous vertex-centric graph algorithms, and executing them on High Performance Computing (HPC) systems. It provides a vertex 'visitor' interface, where actions are defined at an individual vertex level, and contains a suite of classic graph algorithms. HavoqGT is capable of processing large graphs stored in NVRAM (SSDs) using a memory mapped interface.

Reza, TahsinA.↗

Projective Hedging Algorithms for Multistage Stochastic Programming, Supporting Distributed and Asynchronous Implementation

Here we propose a decomposition algorithm for multistage stochastic programming that resembles the progressive hedging method of Rockafellar and Wets but is provably capable of several forms of asynchronous operation. We derive the method from a class of projective operator splitting methods fairly recently proposed by Combettes and Eckstein, significantly expanding the known applications of those methods. Our derivation assures convergence for convex problems whose feasible set is compact, subject to some standard regularity conditions and a mild “fairness” condition on subproblem selection. The method’s convergence guarantees are deterministic and do not require randomization, in contrast to other proposed asynchronous variations of progressive hedging. Computational experiments described in an online appendix show the method to outperform progressive hedging on large-scale problems in a highly parallel computing environment.

97 MATHEMATICS AND COMPUTING↗

Asynchronous Iterative Solvers for Extreme-Scale Computing (Final Report)

This is the final report for the project: Asynchronous Iterative Solvers for Extreme-Scale Computing. This was a collaborative project. This report only covers the activities specific to Georgia Institute of Technology. The project investigated and developed iterative solvers that operate asynchronously, thereby avoiding the high cost of synchronization that is apparent when using standard, synchronous iterative solvers at extreme levels of parallelism.

97 MATHEMATICS AND COMPUTING↗

Coordinated frequency regulation among asynchronous AC grids with an MTDC system

With the increasing penetration level of renewable power, the power systems’ inertias are decreasing. To ensure that the frequency is above the threshold of the under frequency load shedding, one effective way is to share the spinning reserves among asynchronous AC grids through HVDC systems. In this paper, a coordinated frequency regulation scheme that enables spinning reserves sharing is proposed for such hybrid AC/DC grids. This scheme is a communication-free method that consists of the droop control, the inertia emulation control and the frequency safety control. The combined response of the three controls can increase not only the frequency nadir but also the settled frequency of the disturbed grid. In addition, the hysteresis comparators and the power injection limits are adopted in the frequency regulation scheme to reduce the coupling effect among AC grids. Thus, under small and harmless disturbance, the frequency regulation scheme will be inactive; the disturbance in strong AC grids would not affect the weak AC systems. The frequency regulation scheme coordinates well with the automatic generation control of the grids. Furthermore, the feasibility and effectiveness of the proposed frequency regulation scheme are verified through simulation studies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

SAGIPS: a physics-inspired scalable asynchronous generative inverse-problem solver

Abstract Solving large-scale inverse problems using deep-learning algorithms have become an essential part of modern research and industrial applications. The complexity of the underlying inverse problem may require the utilization of high performance computing systems which poses a challenge on the algorithmic design of the inverse problem solver. Most deep learning algorithms require, due to their design, custom parallelization techniques in order to be resource efficient while showing a reasonable convergence. In this paper we introduce a S calable A synchronous G enerative I nverse P roblem S olver (SAGIPS) on high-performance computing systems. We present a workflow that utilizes an asynchronous ring-allreduce algorithm to transfer the gradients of the generator network across multiple GPUs. Experiments with a scientific proxy application demonstrate that SAGIPS shows near linear weak scaling, together with a convergence quality that is comparable to traditional methods. The approach presented here allows leveraging Generative Adverserial Network across multiple GPUs, promising advancements in solving complex inverse problems at scale.

97 MATHEMATICS AND COMPUTING↗

State Estimation for Distribution Networks with Asynchronous Sensors Using Stochastic Descent: Preprint

This paper investigates the problem of state estimation for distribution networks with asynchronous sensors comprising of a mix of smart meters and phasor measurement units (PMUs) with multiple sampling and reporting rates. We consider two independent scenarios of state estimation and tracking, with either voltages or currents as states. With these two sets, we investigate estimation under (a) full data, assuming all measurements are available and (b) limited data, where an online algorithmic approach is adopted to estimate the possibly time-varying states by processing measurements as and when available. The proposed algorithm, inspired by the classical Stochastic Gradient Descent (SGD) approach updates the states based on the previous estimate and the newly available measurements. Finally, we demonstrate the estimation and tracking efficacy through numerical simulations on the IEEE-37 test network, while also highlighting how estimation with currents as states leads to faster convergence.

asynchronous sensors↗

Asynchronous Grid Connections Providing Fast-Frequency Response: System Integration Study

This paper presents an integration study for the recent power electronic-based fast-frequency response technology, "asynchronous grid connection" which operates as an aggregator for behind-the-meter resources and distributed generators. Both technical feasibility and techno-economic viability studies are presented. The fast-frequency response characteristics, validated against Power Hardware-in-the-Loop experiments, are integrated into an IEEE 9- bus system in DigSilent PowerFactory for system-level dynamic analysis. It demonstrates that droop-based control enhancements to local distributed generators allow their aggregation to provide grid-supporting functionalities and participate in the ancillary service markets. To this end, a long-term simulation embedding the system within the ancillary service market framework of PJM has been performed. The fast-frequency response regulation is subsequently used to calculate the potential revenue and project the results on a 15-year investment horizon. Finally, the techno-economic analysis provides recommendations for enhancements to access the full potential of distributed generators on a technical and regulatory level.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Avoiding Excess Computation in Asynchronous Evolutionary Algorithms

Asynchronous evolutionary algorithms are becoming increasingly popular as a means of making full use of many processors while solving computationally expensive search and optimization problems. These algorithms excel at keeping large clusters fully utilized, but may sometimes inefficiently sample an excess of fast-evaluating solutions at the expense of higher-quality, slow-evaluating ones. We introduce a steady-state parent selection strategy, SWEET (“Selection whilE EvaluaTing”), that sometimes selects individuals that are still being evaluated and allows them to reproduce early. This gives slow-evaluating individuals that have higher fitnesses an increased ability to multiply in the population. We find that SWEET appears effective in simulated take-over time analysis, but that its benefit is confined mostly to early in the run, and our preliminary study on an autonomous vehicle controller problem that involves tuning a spiking neural network proves inconclusive.

Scott, Eric↗

Robustness of Deep Learning Classification to Adversarial Input on GPUs: Asynchronous Parallel Accumulation Is a Source of Vulnerability

The ability of machine learning (ML) classification models to resist small, targeted input perturbations—known as adversarial attacks—is a key measure of their safety and reliability. We show that floating-point non associativity (FPNA) coupled with asynchronous parallel programming on GPUs is sufficient to result in misclassification, without any perturbation to the input. Additionally, we show that this misclassification is particularly significant for inputs close to the decision boundary and that standard adversarial robustness results may be overestimated up to 4.6 when not considering machine-level details. We first study a linear classifier, before focusing on standard Graph Neural Network (GNN) architectures and datasets used in robustness assessments. We develop a novel black-box attack using Bayesian optimization to discover external workloads that can change the instruction scheduling which bias the output of reductions on GPUs and reliably lead to misclassification. Motivated by these results, we present a new learnable permutation (LP) gradient-based approach to learning floating-point operation orderings that lead to misclassifications. The LP approach provides a worst-case estimate in a computationally efficient manner, avoiding the need to run identical experiments tens of thousands of times over a potentially large set of possible GPU states or architectures. Finally, using instrumentation-based testing, we investigate parallel reduction ordering across different GPU architectures under external background workloads, when utilizing multi-GPU virtualization, and when applying power capping. Our results demonstrate that parallel reduction ordering varies significantly across architectures under the first two conditions, substantially increasing the search space required to fully test the effects of this parallel scheduler-based vulnerability. These results and the methods developed here can help to include machine-level considerations into adversarial robustness assessments, which can make a difference in safety and mission critical applications.

Shanmugavelu, Sanjif [Maxeler Technologies, a Groq↗

An asynchronous parallel high-throughput model calibration framework for crystal plasticity finite element constitutive models

Crystal plasticity finite element model (CPFEM) is a powerful numerical simulation in the integrated computational materials engineering toolboxes that relates microstructures to homogenized materials properties and establishes the structure–property linkages in computational materials science. However, to establish the predictive capability, one needs to calibrate the underlying constitutive model, verify the solution and validate the model prediction against experimental data. Bayesian optimization (BO) has stood out as a gradient-free efficient global optimization algorithm that is capable of calibrating constitutive models for CPFEM. Here in this paper, we apply a recently developed asynchronous parallel constrained BO algorithm to calibrate phenomenological constitutive models for stainless steel 304 L, Tantalum, and Cantor high-entropy alloy.

304L stainless steel↗

Scheduling and Performance of Asynchronous Tasks in Fortran 2018 with FEATS

Most parallel scientific programs contain compiler directives (pragmas) such as those from OpenMP (Hermanns in Parallel programming in Fortran 95 using openMP, 2002. School of Aeronautical Engineering, Universidad Politécnica de Madrid, España, 2011), explicit calls to runtime library procedures such as those implementing the Message Passing Interface (MPI) (in A message-passing interface standard version 4.0, 2021. https://www.mpi-forum.org/docs/mpi-4.0/mpi40-report.pdf), or compiler-specific language extensions such as those provided by CUDA (Ruetsch and Fatica in CUDA Fortran for scientists and engineers: best practices for efficient CUDA Fortran programming, Elsevier, 2013). By contrast, the recent Fortran standards empower developers to express parallel algorithms without directly referencing lower-level parallel programming models (Numrich in Parallel programming with co-arrays, CRC Press, 2018, and Curcic in Modern Fortran: building efficient parallel applications, Manning Publications, 2020). Fortran’s parallel features place the language within the Partitioned Global Address Space (PGAS) class of programming models. When writing programs that exploit data parallelism, application developers often find it straightforward to develop custom parallel algorithms. Problems involving complex, heterogeneous, staged calculations, however, pose much greater challenges. Such applications require careful coordination of tasks in a manner that respects dependencies prescribed by a directed acyclic graph. When rolling one’s own solution proves difficult, extending a customizable framework becomes attractive. Further, the paper presents the design, implementation, and use of the Framework for Extensible Asynchronous Task Scheduling (FEATS), which we believe to be the first task scheduling tool written in modern Fortran. We describe the benefits and compromises associated with choosing Fortran as the implementation language, and we propose ways in which future Fortran standards can best support the use case in this paper.

97 MATHEMATICS AND COMPUTING↗