Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed Asynchronous”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Accelerating shared file checkpoint with local burst buffers

A data management system and method for accelerating shared file checkpointing. Written application data is aggregated in an application data file created in a local burst buffer memory at a compute node, and an associated data mapping built index to maintain information related to the offsets into a shared file at which segments of the application data is to be stored in a parallel file system, and where in the buffer those segments are located. The node asynchronously transfers a data file containing the application data and the associated data mapping index to a file server for shared file storage. The data management system and method further accelerates shared file checkpointing in which a shared file, together with a map file that specifies how the shared file is to be distributed, is asynchronously transferred to local burst buffer memories at the nodes to accelerate reading of the shared file.

Gooding, Thomas↗

Three practical workflow schedulers for easy maximum parallelism

Runtime scheduling and workflow systems are an increasingly popular algorithmic component in HPC because they allow full system utilization with relaxed synchronization requirements. There are so many special-purpose tools for task scheduling, one might wonder why more are needed. Use cases seen on the Summit supercomputer needed better integration with MPI and greater flexibility in job launch configurations. Preparation, execution, and analysis of computational chemistry simulations at the scale of tens of thousands of processors revealed three distinct workflow patterns. A separate job scheduler was implemented for each one using extremely simple and robust designs: file-based, task-list based, and bulk-synchronous. Comparing to existing methods shows unique benefits of this work, including simplicity of design, suitability for HPC centers, short startup time, and well-understood per-task overhead. All three new tools have been shown to scale to full utilization of Summit, and have been made publicly available with tests and documentation. This work presents a complete characterization of the minimum effective task granularity for efficient scheduler usage scenarios. Here, these schedulers have the same bottlenecks, and hence similar task granularities as those reported for existing tools following comparable paradigms.

97 MATHEMATICS AND COMPUTING↗

Online State Estimation for Time-Varying Systems

The paper investigates the problem of estimating the state of a time-varying system with a linear measurement model; in particular, the paper considers the case where the number of measurements available can be smaller than the number of states. In lieu of a batch linear least-squares (LS) approach well-suited for static networks, where a sufficient number of measurements could be collected to obtain a full-rank design matrix the paper proposes an online algorithm to estimate the possibly time-varying state by processing measurements as and when available. The design of the algorithm hinges on a generalized LS cost augmented with a proximal-point-type regularization. With the solution of the regularized LS problem available in closed-form, the online algorithm is written as a linear dynamical system where the state is updated based on the previous estimate and based on the new available measurements. Conditions under which the algorithmic steps are in fact a contractive mapping are shown, and bounds on the estimation error are derived for different noise models. Numerical simulations are provided to corroborate the analytical findings.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Beam Synchronous for the Rest of Us!

Fermilab’s Tevatron Clock (TCLK) infrastructure has been an integral part of the accelerator control network since the 1980’s. This 10MHz Manchester encoded protocol has enabled flexible, real-time event distribution for thousands of devices connected to the timing network with a high degree of reliability. Forthcoming upgrades to the Fermilab complex (PIP-II, LBNF, ACORN) necessitate higher levels of precision to maintain inter-bunch timing for Instrumentation and Control purposes. This presents as an opportunity to refine the event distribution protocol for tighter synchronization between machines, experiments, and eventually far-site operations. This paper outlines a method by which beam-synchronous events may be distributed through asynchronous serial protocols via integration with local LLRF and global PPS reference signals. This method is ideal for synchrotron machines with aggressive frequency sweeps (such as Fermilab's 38~53MHz Booster) and allows for precision timing to be maintained across machines without specialized hardware.

43 PARTICLE ACCELERATORS↗

Secure State Estimation with Asynchronous Measurements for Coordinated Cyber Attack Detection in Active Distribution Systems

Coordinated cyber attacks tamper with measurement data to disrupt the situational awareness of active distribution systems. Various sensors report measurements asynchronously at different rates, which introduces challenges during state estimation. In addition, this forces cyber intruders to exert greater effort to compromise multiple communication channels and launch coordinated attacks. Therefore, multi-channel and asynchronous measurements could be harnessed to develop more secure cyber defense strategies. In this paper, a prediction-correction-based multi-rate observer is designed to exploit the value of asynchronous measurements for the detection of coordinated false data injection (FDI) attacks. First, a time-function-dependent prediction-correction strategy is proposed to adjust the sampling interval for each sensor’s measurement. Then, an observer is designed based on the trade-off between estimation error and the optimal period of the most recent sampling instant, with the convergence of estimation error with the maximum permitted sampling interval. Moreover, the conditions for exponential stability are developed using the Lyapunov–Krasovskii functional technique. Next, a coordinated FDI attack detection strategy is developed based on the dual nonlinear minimization problem. The proposed attack detection and secure state estimation strategies are tested on the IEEE 13-node system. Simulation results show that these schemes are effective in enhancing attack detection based on asynchronous measurements or compromised data.

asynchronous measurements↗

Asynchronous and Load-Balanced Union-Find for Distributed and Parallel Scientific Data Visualization and Analysis

We present a novel distributed union-find algorithm that features asynchronous parallelism and k-d tree based load balancing for scalable visualization and analysis of scientific data. Applications of union-find include level set extraction and critical point tracking, but distributed union-find can suffer from high synchronization costs and imbalanced workloads across parallel processes. In this study, we prove that global synchronizations in existing distributed union-find can be eliminated without changing final results, allowing overlapped communications and computations for scalable processing. We also use a k-d tree decomposition to redistribute inputs, in order to improve workload balancing. We benchmark the scalability of our algorithm with up to 1,024 processes using both synthetic and application data. Here, we demonstrate the use of our algorithm in critical point tracking and super-level set extraction with high-speed imaging experiments and fusion plasma simulations, respectively.

97 MATHEMATICS AND COMPUTING↗

Optimized asynchronous training of neural networks using a distributed parameter server with eager updates

A method of training a neural network includes, at a local computing node, receiving remote parameters from a set of one or more remote computing nodes, initiating execution of a forward pass in a local neural network in the local computing node to determine a final output based on the remote parameters, initiating execution of a backward pass in the local neural network to determine updated parameters for the local neural network, and prior to completion of the backward pass, transmitting a subset of the updated parameters to the set of remote computing nodes.

Hamidouche, Khaled↗

Observations of particle number size distributions and new particle formation in six Indian locations

Atmospheric new particle formation (NPF) is a crucial process driving aerosol number concentrations in the atmosphere; it can significantly impact the evolution of atmospheric aerosol and cloud processes. This study analyses at least 1 year of asynchronous particle number size distributions from six different locations in India. We also analyze the frequency of NPF and its contribution to cloud condensation nuclei (CCN) concentrations. We found that the NPF frequency has a considerable seasonal variability. At the measurement sites analyzed in this study, NPF frequently occurs in March–May (pre-monsoon, about 21% of the days) and is the least common in October–November (post-monsoon, about 7% of the days). Considering the NPF events in all locations, the particle formation rate (J SDS ) varied by more than 2 orders of magnitude (0.001–0.6 cm –3 s –1 ) and the growth rate between the smallest detectable size and 25nm (GR SDS-25nm ) by about 3 orders of magnitude (0.2–17.2nm h –1 ). We found that J SDS was higher by nearly 1 order of magnitude during NPF events in urban areas than mountain sites. GR SDS did not show a systematic difference. Our results showed that NPF events could significantly modulate the shape of particle number size distributions and CCN concentrations in India. The contribution of a given NPF event to CCN concentrations was the highest in urban locations (4.3 × 10 3 cm –3 per event and 1.2 × 10 3 cm –3 per event for 50 and 100nm, respectively) as compared to mountain background sites (2.7 × 10 3 cm –3 per event and 1.0 × 10 3 cm –3 per event, respectively). We emphasize that the physical and chemical pathways responsible for NPF and factors that control its contribution to CCN production require in situ field observations using recent advances in aerosol and its precursor gaseous measurement techniques.

54 ENVIRONMENTAL SCIENCES↗

Cooperative fault management for resilient integration of renewable energy

Cooperative fault management (CFM) is designed herein to control different types of renewable energy resources cooperatively during electrical faults. This paper studies systems with a high penetration of photovoltaic (PV) energy and wind energy. First, CFM leverages power converters of PV farms to boost the ride-through capability of nearby doubly-fed induction generators (DFIGs). By controlling PV farms’ output voltages to change smoothly during both fault initiation and fault clearance, the widely used crowbar in DFIGs is less likely to be activated. Crowbar activation adversely makes DFIGs lose controllability and absorb reactive power. The second contribution is the development of a software-defined CFM controller and a controller in-the-loop demonstration of the real-time performance of this optimization-based CFM. CFM capitalizes on distributed optimization formulation to enable flexibility, plug-and-play, and privacy-preserving. Computation time, however, is a major concern for optimization-based dynamics control. Here, real-time controller-in-the-loop simulation results show optimization-based CFM can output reference values around 60 ms and is quick enough for dynamic control.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Asynchronous Grid Connections Providing Fast-Frequency Response: System Integration Study

This paper presents an integration study for the recent power electronic-based fast-frequency response technology, "asynchronous grid connection" which operates as an aggregator for behind-the-meter resources and distributed generators. Both technical feasibility and techno-economic viability studies are presented. The fast-frequency response characteristics, validated against Power Hardware-in-the-Loop experiments, are integrated into an IEEE 9- bus system in DigSilent PowerFactory for system-level dynamic analysis. It demonstrates that droop-based control enhancements to local distributed generators allow their aggregation to provide grid-supporting functionalities and participate in the ancillary service markets. To this end, a long-term simulation embedding the system within the ancillary service market framework of PJM has been performed. The fast-frequency response regulation is subsequently used to calculate the potential revenue and project the results on a 15-year investment horizon. Finally, the techno-economic analysis provides recommendations for enhancements to access the full potential of distributed generators on a technical and regulatory level.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Characteristics of Vertical Ground Motions and Their Effect on the Seismic Response of Bridges in the Near-Field: A State-of-the-Art Review

Despite the evidence from past earthquakes and several numerical investigations demonstrating the detrimental impact of vertical ground motions (VGMs) on the integrity of bridge structures, incorporating their effects into seismic assessment and design procedures has traditionally been given limited consideration. Current codes utilize rather simplistic approaches to account for the concurrent effects of vertical and horizontal motions in structural performance evaluations, potentially leading to unconservative estimates of structural demands. This paper reviews the main features of VGMs and their effect on the seismic response of bridges. The methods and empirical models available to estimate vertical motions for design purposes are discussed, and research gaps and related research needs are identified. Finally, the emerging role of physics-based ground-motion simulations, as well as their limitations, in supporting future research and informing the development of simplified design procedures is examined. The main areas of interest for future research are identified in the need to carry out systematic sensitivity studies to gain insight into the main earthquake parameters that influence key VGMs features, understand the influence of soil nonlinearities on VGMs amplitude and frequency content, inform the development of empirical models that cover a range of site conditions and source-to-site distances where current models are poorly constrained, generate arrays of motions to update coherency models to properly inform the analysis of distributed infrastructure, investigate the impulsive character of VGMs, and assess the approximations made in estimating VGMs with 1D site response analyses. Furthermore, specific focus is laid on large-magnitude earthquakes in the near-field.

Asynchronous ground motion↗

A Fine-grained Asynchronous Bulk Synchronous parallelism model for PGAS applications

The Partitioned Global Address Space (PGAS) model is well suited for executing irregular applications on cluster-based systems, due to its efficient support for short, one-sided messages. Separately, the actor model has been gaining popularity as a productive asynchronous message-passing approach for distributed objects in enterprise and cloud computing platforms, typically implemented in languages such as Erlang, Scala or Rust. To the best of our knowledge, there has been no past work on using the actor model to deliver both productivity and scalability to irregular PGAS applications with large number of small messages. In this paper, we introduce a new programming system for PGAS applications, in which point-to-point remote operations can be expressed as fine-grained asynchronous actor messages. In our approach, the programmer does not need to worry about programming complexities related to message aggregation and termination detection. Our approach can be viewed as extending the classical Bulk Synchronous Parallelism model with fine-grained asynchronous communications within a phase or superstep. Here, we believe that our approach offers a desirable point in the productivity-performance space for PGAS applications, with more scalable performance and higher productivity relative to past approaches. Specifically, for seven irregular mini-applications from the Bale Kernels and three graph kernels executed using 2048 cores in the NERSC Cori system, our approach shows geometric mean performance improvements of ≥ 20X relative to standard PGAS versions (UPC and OpenSHMEM) while maintaining comparable productivity to those versions.

97 MATHEMATICS AND COMPUTING↗

Packetized energy management control systems and methods of using the same

Aspects of the present disclosure include anonymous, asynchronous, and randomized control schemes for distributed energy resources (DERs). Such control schemes may include packetized energy management (PEM) control schemes for managing DERs that may provide near-optimal tracking performance under imperfect information and consumer quality of service (QoS) constraints.

Frolik, Jeffrey↗

Modifying the Asynchronous Jacobi Method for Data Corruption Resilience

Moving scientific computation from high-performance computing (HPC) and cloud computing (CC) environments to devices on the edge, i.e., physically near instruments of interest, has received tremendous interest in recent years. Such edge computing environments can operate on data in situ, offering enticing benefits over data aggregation to HPC and CC facilities that include avoiding costs of transmission, increased data privacy, and real-time data analysis. Because of the inherent unreliability of edge computing environments, new fault-tolerant approaches must be developed before the benefits of edge computing can be realized. Motivated by algorithm-based fault tolerance, a variant of the asynchronous Jacobi (ASJ) method is developed that achieves resilience to data corruption by rejecting solution approximations from neighbor devices according to a bound derived from convergence theory. Numerical results on a two-dimensional Poisson problem show that the new rejection criterion, along with a novel approximation to the shortest path length on which the criterion depends, restores convergence for the ASJ variant in the presence of certain types data corruption. Numerical results are obtained for when the singular values in the analytic bound are approximated. Additional linear systems are also explored, one with a more dense sparsity pattern and one that includes advection. All results indicate that successful resilience to data corruption depends on whether the bound tightens fast enough to reject corrupted data before the iteration evolution deviates significantly from that predicted by the convergence theory defining the bound. This observation generalizes to future work on algorithm-based fault tolerance for other asynchronous algorithms, including upcoming approaches that leverage Krylov subspaces.

97 MATHEMATICS AND COMPUTING↗