Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “asynchronous algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Saturation: An efficient iteration strategy for symbolic state-space generation

This paper presents a novel algorithm for generating state spaces of asynchronous systems using Multi-valued Decision Diagrams. In contrast to related work, the next-state function of a system is not encoded as a single Boolean function, but as cross-products of integer functions. This permits the application of various iteration strategies to build a system's state space. In particular, this paper introduces a new elegant strategy, called saturation, and implements it in the tool SMART. On top of usually performing several orders of magnitude faster than existing BDD-based state-space generators, the algorithm's required peak memory is often close to the nal memory needed for storing the overall state spaces.

Ciardo, Gianfranco↗

SAGIPS: a physics-inspired scalable asynchronous generative inverse-problem solver

Abstract Solving large-scale inverse problems using deep-learning algorithms have become an essential part of modern research and industrial applications. The complexity of the underlying inverse problem may require the utilization of high performance computing systems which poses a challenge on the algorithmic design of the inverse problem solver. Most deep learning algorithms require, due to their design, custom parallelization techniques in order to be resource efficient while showing a reasonable convergence. In this paper we introduce a S calable A synchronous G enerative I nverse P roblem S olver (SAGIPS) on high-performance computing systems. We present a workflow that utilizes an asynchronous ring-allreduce algorithm to transfer the gradients of the generator network across multiple GPUs. Experiments with a scientific proxy application demonstrate that SAGIPS shows near linear weak scaling, together with a convergence quality that is comparable to traditional methods. The approach presented here allows leveraging Generative Adverserial Network across multiple GPUs, promising advancements in solving complex inverse problems at scale.

97 MATHEMATICS AND COMPUTING↗

Redundant asynchronous microprocessor system for fault tolerant flight control and navigation

Unlike their synchronized counterparts, redundant channels in an asynchronous flight system can, under no-fault conditions, exhibit cross-channel data disparities. Sources of these errors are examined in terms of the general, individual functions of the flight control and navigation application in the asynchronous digital environment. The effects of asynchronism on trajectory programmers, dynamic control algorithms and data reconstruction processes are examined in terms of data skews, data latencies, and clock rate uncertainties. An example is presented in which time corrections are applied to reduce the data disparities. Practical limitations of the approach of the example are discussed.

Dunn, W. R.↗

The Use of Efficient Broadcast Protocols in Asynchronous Distributed Systems

Reliable broadcast protocols are important tools in distributed and fault-tolerant programming. They are useful for sharing information and for maintaining replicated data in a distributed system. However, a wide range of such protocols has been proposed. These protocols differ in their fault tolerance and delivery ordering characteristics. There is a tradeoff between the cost of a broadcast protocol and how much ordering it provides. It is, therefore, desirable to employ protocols that support only a low degree of ordering whenever possible. This dissertation presents techniques for deciding how strongly ordered a protocol is necessary to solve a given application problem. It is shown that there are two distinct classes of application problems: problems that can be solved with efficient, asynchronous protocols, and problems that require global ordering. The concept of a linearization function that maps partially ordered sets of events to totally ordered histories is introduced. How to construct an asynchronous implementation that solves a given problem if a linearization function for it can be found is shown. It is proved that in general the question of whether a problem has an asynchronous solution is undecidable. Hence there exists no general algorithm that would automatically construct a suitable linearization function for a given problem. Therefore, an important subclass of problems that have certain commutativity properties are considered. Techniques for constructing asynchronous implementations for this class are presented. These techniques are useful for constructing efficient asynchronous implementations for a broad range of practical problems.

Schmuck, Frank Bernhard↗

An asynchronous parallel high-throughput model calibration framework for crystal plasticity finite element constitutive models

Crystal plasticity finite element model (CPFEM) is a powerful numerical simulation in the integrated computational materials engineering toolboxes that relates microstructures to homogenized materials properties and establishes the structure–property linkages in computational materials science. However, to establish the predictive capability, one needs to calibrate the underlying constitutive model, verify the solution and validate the model prediction against experimental data. Bayesian optimization (BO) has stood out as a gradient-free efficient global optimization algorithm that is capable of calibrating constitutive models for CPFEM. Here in this paper, we apply a recently developed asynchronous parallel constrained BO algorithm to calibrate phenomenological constitutive models for stainless steel 304 L, Tantalum, and Cantor high-entropy alloy.

304L stainless steel↗

IRIS-MASH: Efficient Multi-device Asynchronous Multi-Stream Heterogeneous Computing

In the rapidly evolving field of high-performance computing (HPC), effectively leveraging heterogeneous devices through asynchronous task programming is paramount. This paper presents a robust asynchronous task programming model tailored for a multi-device, multi-stream execution environment that incorporates a diverse array of heterogeneous computing units, including GPUs from various vendors and other accelerators. Current state-of-the-art task programming models provide methodologies to support asynchronous task executions, but they typically handle homogeneous devices using native programming languages, while support for heterogeneous devices is limited to frameworks like OpenCL. This gap presents significant challenges in abstracting heterogeneous devices to harness their true asynchronous capabilities effectively using their native programming languages. By implementing asynchronous task execution, our model significantly boosts the performance of tiled algorithm task graphs through overlapping data transfers with computation and enabling the simultaneous execution of multiple kernels. We integrate this approach into a heterogeneous Intelligent Runtime System (IRIS) and assess its performance using a suite of tiled algorithm benchmarks from the heterogeneous math kernels library (MatRIS) based on IRIS. Experimental results demonstrate a performance improvement ranging from 1.6 × to 2 × over IRIS without asynchronous support, and a notable 22% performance enhancement compared to established runtime systems such as StarPU and PaRSEC. This approach significantly improves computation efficiency of HPC workflows and provides a solid base for future exploration and development in the area of asynchronous task programming in heterogeneous systems.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

Seven open problems in applied combinatorics

We present and discuss seven different open problems in applied combinatorics. Additionally, the application areas relevant to this compilation include quantum computing, algorithmic differentiation, topological data analysis, iterative methods, hypergraph cut algorithms, and power systems.

97 MATHEMATICS AND COMPUTING↗

A Hardware and Software Co-design Framework for Energy Efficient Neuromorphic Systems

Neuromorphic systems can be realized by a variety of algorithms and architectures. A common understanding is that spiking neuromorphic designs, which encode information into spatio-temporal spiking events, are both a biologically-accurate and efficient way of processing information. However, representing the information through timing relationships induces sophisticated circuit designs in traditional CMOS-based implementations. In recent years, high-capacity resistive memory (RRAM, aka, memristor) has demonstrated great potential in mimicking synaptic behaviors. Several RRAM-based spiking neuromorphic designs exist, most of which focus on rate coding schemes. These designs simplify circuit implementations of neuron models and explore challenges such as unsatisfactory speed, resolution, and performance. As an alternative, we will explore temporal coding spiking neuromorphic systems that encode information as the relative timing of neuron activations (spikes), which have been proven to be more adaptive and energy-efficient. Developing a neuromorphic system for spiking neural network (SNN) inference and online training, however, faces some major technical challenges: (1) It lacks circuit implementation support for temporal-coding SNN to achieve satisfying power efficiency and accuracy; (2) Although existing research works have investigated memristive synapse and neuron designs for spike-timing-dependent plasticity, the non-ideal conditions in implementation, such as device variations and signal degradation, degrade online learning accuracy of large scale systems; and (3) Non-optimized, inter-layer data traffic in SNNs, leads to unnecessary data communication costs. In this project, we plan to address these challenges by a hardware and software co-design framework that incorporates solutions at the circuit, architecture, and algorithm levels. At the circuit-level, we will elaborate on the in-situ SNN processing element designs for supporting both inference and online training modes. Variation-aware schemes will be studied to improve reliability. At the architecture level, we propose a pipelined, asynchronous architecture to retain the timing resolution of spikes. At the algorithm level, we will investigate an innovative SNN training algorithm for enabling activation sparsification and reducing unnecessary data communication costs. This neuromorphic system will provide an effective solution to real-life energy-constrained applications and significantly contribute to the exploration of next-generation high-performance computing systems under the DOE context.

97 MATHEMATICS AND COMPUTING↗

Where are the parallel algorithms?

Four paradigms that can be useful in developing parallel algorithms are discussed. These include computational complexity analysis, changing the order of computation, asynchronous computation, and divide and conquer. Each is illustrated with an example from scientific computation, and it is shown that computational complexity must be used with great care or an inefficient algorithm may be selected.

Voigt, R. G.↗

Fast tree-based algorithms for DBSCAN for low-dimensional data on GPUs

DBSCAN is a well-known density-based clustering algorithm to discover arbitrary shape clusters. While conceptually simple in serial, the algorithm is challenging to efficiently parallelize on manycore GPU architectures. Common pitfalls, such as asynchronous range query calls, result in high thread execution divergence in many implementations. In this paper, we propose a new framework for GPU-accelerated DBSCAN, and describe two tree-based algorithms within that framework. Both algorithms fuse the search for neighbors with updating cluster information, but differ in their treatment of dense regions of the data. We show that the time taken to compute clusters is at most twice that of determination of the neighbors. We compare the proposed algorithms with existing CPU and GPU implementations, and demonstrate their competitiveness and performance using a fast traversal structure (bounding volume hierarchy) for low dimensional data. We also show that the memory usage can be reduced by processing object neighbors dynamically without storing them.

Prokopenko, Andrey↗

Self arbitrated VLSI asynchronous sequential circuits

A new class of asynchronous sequential circuits is introduced in this paper. The new design procedures are oriented towards producing asynchronous sequential circuits that are implemented with CMOS VLSI and take advantage of pass transistor technology. The first design algorithm utilizes a standard Single Transition Time (STT) state assignment. The second method introduces a new class of self synchronizing asynchronous circuits which eliminates the need for critical race free state assignments. These circuits arbitrate the transition path action by forcing the circuit to sequence through proper unstable states. These methods result in near minimum hardware since only the transition paths associated with state variable changes need to be implemented with pass transistor networks.

Whitaker, S.↗

Dynamic grid refinement for partial differential equations on parallel computers

The fast adaptive composite grid method (FAC) is an algorithm that uses various levels of uniform grids to provide adaptive resolution and fast solution of PDEs. An asynchronous version of FAC, called AFAC, that completely eliminates the bottleneck to parallelism is presented. This paper describes the advantage that this algorithm has in adaptive refinement for moving singularities on multiprocessor computers. This work is applicable to the parallel solution of two- and three-dimensional shock tracking problems.

Mccormick, S.↗

Online Distribution System State Estimation via Stochastic Gradient Algorithm

Distribution network operation is becoming more challenging because of the growing integration of intermittent and volatile distributed energy resources (DERs). This motivates the development of new distribution system state estimation (DSSE) paradigms that can operate at fast timescale based on real-time data stream of asynchronous measurements enabled by modern information and communications technology. To solve the real-time DSSE with asynchronous measurements effectively and accurately, this paper formulates a weighted least squares DSSE problem and proposes an online stochastic gradient algorithm to solve it. The performance of the proposed scheme is analytically guaranteed and is numerically corroborated with realistic data on IEEE 123-bus feeder.

distribution system state estimation↗

Asynchronous Truncated Multigrid-Reduction-in-Time

In this paper, we present the new “asynchronous truncated multigrid-reduction-in-time” (AT-MGRIT) algorithm for introducing time parallelism to the solution of discretized time-dependent problems. The new algorithm is based on the multigrid-reduction-in-time (MGRIT) approach, which, in certain settings, is equivalent to another common multilevel parallel-in-time method, Parareal. In contrast to Parareal and MGRIT that both consider a global temporal grid over the entire time interval on the coarsest level, the AT-MGRIT algorithm uses truncated local time grids on the coarsest level, each grid covering certain temporal subintervals. Further, these local grids can be solved completely in an independent way from each other, which reduces the sequential part of the algorithm and, thus, increases parallelism in the method. Here, we study the effect of using truncated local coarse grids on the convergence of the algorithm, both theoretically and numerically, and show, using challenging nonlinear problems, that the new algorithm consistently outperforms classical Parareal/MGRIT in terms of time to solution.

97 MATHEMATICS AND COMPUTING↗

Optimization of the computational load of a hypercube supercomputer onboard a mobile robot

A combinatorial optimization methodology is developed, which enables the efficient use of hypercube multiprocessors onboard mobile intelligent robots dedicated to time-critical missions. The methodology is implemented in terms of large-scale concurrent algorithms based either on fast simulated annealing, or on nonlinear asynchronous neural networks. In particular, analytic expressions are given for the effect of single-neuron perturbations on the systems' configuration energy. Compact neuromorphic data structures are used to model effects such as precedence constraints, processor idling times, and task-schedule overlaps. Results for a typical robot-dynamics benchmark are presented.

Barhen, Jacob↗

Simulation and Verification of Synchronous Set Relations in Rewriting Logic

This paper presents a mathematical foundation and a rewriting logic infrastructure for the execution and property veri cation of synchronous set relations. The mathematical foundation is given in the language of abstract set relations. The infrastructure consists of an ordersorted rewrite theory in Maude, a rewriting logic system, that enables the synchronous execution of a set relation provided by the user. By using the infrastructure, existing algorithm veri cation techniques already available in Maude for traditional asynchronous rewriting, such as reachability analysis and model checking, are automatically available to synchronous set rewriting. The use of the infrastructure is illustrated with an executable operational semantics of a simple synchronous language and the veri cation of temporal properties of a synchronous system.

Rocha, Camilo↗

Asynchronous interactive control systems

A class of interactive control systems is derived by generalizing interactive manipulator control systems. The general structural properties of such systems are discussed and an appropriate general software implementation is proposed. This is based on the fact that tasks of interactive control systems can be represented as a network of a finite set of actions which have specific operational characteristics and specific resource requirements, and which are of limited duration. This has enabled the decomposition of the overall control algorithm into a set of subalgorithms, called subcontrollers, which can operate simultaneously and asynchronously. Coordinate transformations of sensor feedback data and actuator set-points have enabled the further simplification of the subcontrollers and have reduced their conflicting resource requirements. The modules of the decomposed control system are implemented as parallel processes with disjoint memory space communicating only by I/O. The synchronization mechanisms for dynamic resource allocation among subcontrollers and other synchronization mechanisms are also discussed in this paper. Such a software organization is suitable for the general form of multiprocessing using computer networks with distributed storage.

Vuskovic, M. I.↗