Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “asynchronous”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Accelerating Flash-X Simulations with Asynchronous I/O

Most high-fidelity physics simulation codes, such as Flash-X, need to save intermediate results (checkpoint files) to restart or gain insights into the evolution of the simulation. These simulation codes save such intermediate files synchronously, where computation is stalled while the data is written to storage. Depending on the problem size and computational requirements, this file write time can be a substantial portion of the total simulation time. In order to hide the I/O latency of checkpointing, asynchronous I/O methods have been introduced. These methods use background threads for performing I/O while the main threads continue with the simulation. The usage of background threads can compete for resources on the node as well as with communication. In this paper, we evaluate the overheads and the overall benefit of asynchronous I/O in HDF5 to simulations. Results from real-world high-fidelity simulations on the Summit supercomputer show that I/O operation is overlapped with application communication or computation or both, effectively hiding some or all of the I/O latency. Our evaluation shows that while using asynchronous I/O adds overhead to the application, the I/O time reduction is more significant, resulting in overall up to 1.5X performance speedup.

Jain, Rajeev↗

Divergence Reduction in Monte Carlo Neutron Transport with On-GPU Asynchronous Scheduling

While Monte Carlo Neutron Transport (MCNT) is near-embarrasingly parallel, the effectively unpredictable lifetime of neutrons can lead to divergence when MCNT is evaluated on GPUs. Divergence is the phenomenon of adjacent threads in a warp executing different control flow paths; on GPUS, it reduces performance because each work group may only execute one path at a time. The process of Thread Data Remapping (TDR) resolves these discrepancies by moving data across hardware such that data in the same warp will be processed through similar paths. A common issue among prior implementations of TDR is the synchronous nature of its remapping and processing cycles, which exhaustively sort data produced by prior processing passes and exhaustively evaluate the sorted data. In another work, we defined a method of remapping data through an asynchronous scheduler which allows for work to be stored in shared memory and deferred arbitrarily until that work is a viable option for low-divergence evaluation. This article surveys a wider set of cases, with the goal of characterizing performance trends across a more comprehensive set of parameters. These parameters include cross sections of scattering/capturing/fission, use of implicit capture, source neutron counts, simulation time spans, and tuned memory allocations. Across these cases, we have recorded minimum and average execution times, as well as a heuristically tuned near-optimal memory allocation size for both synchronous and asynchronous scheduling. Across the collected data, it is shown that the asynchronous method is faster and more memory efficient in the majority of cases, and that it requires less tuning to achieve competitive performance.

Computer Science↗

Quasi-perfect FIFO: Synchronous or asynchronous with application in controller design for the UNICON laser memory

The first-in-first-out memory buffer (FIFO), is an elastic digital memory whose main application is in data buffering between devices operating at different rates. Data written into the top is moved autonomously down toward the bottom of the FIFO to the lowest unoccupied location, and data read from the bottom of the FIFO will cause data from the top to move autonomously down toward the bottom. The FIFO is available in MOS LSI asynchronous form with data rate in the 1 MHz region. The FIFO described yields a simple high-speed iterative implementation, either synchronous of asynchronous. Because of this simple iterative structure, the FIFO is expandable in both number of words and bits per word, and it is attractive from the viewpoint of integrated-circuit production. For the synchronous FIFO, a model was built and successfully used in the controller for the UNICON laser memory. For the asynchronous FIFO, a model was built and also successfully used in a high-performance magnetic tape controller.

Lim, R. S.↗

Asynchronous vibration problem of centrifugal compressor

An unstable asynchronous vibration problem in a high pressure centrifugal compressor and the remedial actions against it are described. Asynchronous vibration of the compressor took place when the discharge pressure (Pd) was increased, after the rotor was already at full speed. The typical spectral data of the shaft vibration indicate that as the pressure Pd increases, pre-unstable vibration appears and becomes larger, and large unstable asynchronous vibration occurs suddenly (Pd = 5.49MPa). A computer program was used which calculated the logarithmic decrement and the damped natural frequency of the rotor bearing systems. The analysis of the log-decrement is concluded to be effective in preventing unstable vibration in both the design stage and remedial actions.

Fujikawa, T.↗

Distributed asynchronous microprocessor architectures in fault tolerant integrated flight systems

The paper discusses the implementation of fault tolerant digital flight control and navigation systems for rotorcraft application. It is shown that in implementing fault tolerance at the systems level using advanced LSI/VLSI technology, aircraft physical layout and flight systems requirements tend to define a system architecture of distributed, asynchronous microprocessors in which fault tolerance can be achieved locally through hardware redundancy and/or globally through application of analytical redundancy. The effects of asynchronism on the execution of dynamic flight software is discussed. It is shown that if the asynchronous microprocessors have knowledge of time, these errors can be significantly reduced through appropiate modifications of the flight software. Finally, the papear extends previous work to show that through the combined use of time referencing and stable flight algorithms, individual microprocessors can be configured to autonomously tolerate intermittent faults.

Dunn, W. R.↗

Redundant asynchronous microprocessor system for fault tolerant flight control and navigation

Unlike their synchronized counterparts, redundant channels in an asynchronous flight system can, under no-fault conditions, exhibit cross-channel data disparities. Sources of these errors are examined in terms of the general, individual functions of the flight control and navigation application in the asynchronous digital environment. The effects of asynchronism on trajectory programmers, dynamic control algorithms and data reconstruction processes are examined in terms of data skews, data latencies, and clock rate uncertainties. An example is presented in which time corrections are applied to reduce the data disparities. Practical limitations of the approach of the example are discussed.

Dunn, W. R.↗

Carrying Synchronous Voice Data On Asynchronous Networks

Buffers restore synchronism for internal use and permit asynchronism in external transmission. Proposed asynchronous local-area digital communication network (LAN) carries synchronous voice, data, or video signals, or non-real-time asynchronous data signals. Network uses double buffering scheme that reestablishes phase and frequency references at each node in network. Concept demonstrated in token-ring network operating at 80 Mb/s, pending development of equipment operating at planned data rate of 200 Mb/s. Technique generic and used with any LAN as long as protocol offers deterministic (or bonded) access delays and sufficient capacity.

Bergman, Larry A.↗

A formal model of asynchronous communication and its use in mechanically verifying a biphase mark protocol

In this paper we present a formal model of asynchronous communication as a function in the Boyer-Moore logic. The function transforms the signal stream generated by one processor into the signal stream consumed by an independently clocked processor. This transformation 'blurs' edges and 'dilates' time due to differences in the phases and rates of the two clocks and the communications delay. The model can be used quantitatively to derive concrete performance bounds on asynchronous communications at ISO protocol level 1 (physical level). We develop part of the reusable formal theory that permits the convenient application of the model. We use the theory to show that a biphase mark protocol can be used to send messages of arbitrary length between two asynchronous processors. We study two versions of the protocol, a conventional one which uses cells of size 32 cycles and an unconventional one which uses cells of size 18. We conjecture that the protocol can be proved to work under our model for smaller cell sizes and more divergent clock rates but the proofs would be harder.

Moore, J. Strother↗

Photon detection with parallel asynchronous processing

An approach to photon detection with a parallel asynchronous signal processor is described. The visible or IR photon-detection capability of the silicon p(+)-n-n(+) detectors and the parallel asynchronous processing are addressed separately. This approach would permit an independent analog processing channel to be dedicated to every pixel. A laminar architecture consisting of a stack of planar arrays of the devices would form a 2D array processor with a 2D array of inputs located directly behind a focal-plane detector array. A 2D image data stream would propagate in neuronlike asynchronous pulse-coded form through the laminar processor. Such systems can integrate image acquisition and image processing. Acquisition and processing would be performed concurrently as in natural vision systems. The possibility of multispectral image processing is addressed.

Coon, D. D.↗

Self arbitrated VLSI asynchronous sequential circuits

A new class of asynchronous sequential circuits is introduced in this paper. The new design procedures are oriented towards producing asynchronous sequential circuits that are implemented with CMOS VLSI and take advantage of pass transistor technology. The first design algorithm utilizes a standard Single Transition Time (STT) state assignment. The second method introduces a new class of self synchronizing asynchronous circuits which eliminates the need for critical race free state assignments. These circuits arbitrate the transition path action by forcing the circuit to sequence through proper unstable states. These methods result in near minimum hardware since only the transition paths associated with state variable changes need to be implemented with pass transistor networks.

Whitaker, S.↗

Implications of Tracey's theorem to asynchronous sequential circuit design

Tracey's Theorem has long been recognized as essential in generating state assignments for asynchronous sequential circuits. This paper shows that Tracey's Theorem also has a significant impact in generating the design equations. Moreover, this theorem is important to the fundamental understanding of asynchronous sequential operation. The results of this work simplify asynchronous logic design. Moreover, detection of safe circuits is made easier.

Gopalakrishnan, S.↗

Partition algebraic design of asynchronous sequential circuits

Tracey's Theorem has long been recognized as essential in generating state assignments for asynchronous sequential circuits. This paper shows that partitioning variables derived from Tracey's Theorem also has a significant impact in generating the design equations. Moreover, this theorem is important to the fundamental understanding of asynchronous sequential operation. The results of this work simplify asynchronous logic design. Moreover, detection of safe circuits is made easier.

Maki, Gary K.↗

Simulating fail-stop in asynchronous distributed systems

The fail-stop failure model appears frequently in the distributed systems literature. However, in an asynchronous distributed system, the fail-stop model cannot be implemented. In particular, it is impossible to reliably detect crash failures in an asynchronous system. In this paper, we show that it is possible to specify and implement a failure model that is indistinguishable from the fail-stop model from the point of view of any process within an asynchronous system. We give necessary conditions for a failure model to be indistinguishable from the fail-stop model, and derive lower bounds on the amount of process replication needed to implement such a failure model. We present a simple one-round protocol for implementing one such failure model, which we call simulated fail-stop.

Sabel, Laura↗

Asynchronous Laser Transponders for Precise Interplanetary Ranging and Time Transfer

The feasibility of a two-way asynchronous (i.e. independently firing) interplanetary laser transponder pair, capable of decimeter ranging and subnanosecond time transfer from Earth to a spacecraft anywhere within the inner Solar System, is discussed. In the Introduction, we briefly discuss the current state-of-the-art in Satellite Laser Ranging (SLR) and Lunar Laser Ranging (LLR) which use single-ended range measurements to a passive optical reflector, and the limitations of this approach in ranging beyond the Moon to the planets. In Section 2 of this paper, we describe two types of transponders (echo and asynchronous), introduce the transponder link equation and the concept of "balanced" transponders, describe how range and time can be transferred between terminals, and preview the potential advantages of photon counting asynchronous transponders for interplanetary applications. In Section 3, we discuss and provide mathematical models for the various sources of noise in an interplanetary transponder link including planetary albedo, solar or lunar illumination of the local atmosphere, and laser backscatter off the local atmosphere. In Section 4, we introduce the key engineering elements of an interplanetary laser transponder and develop an operational scenario for the acquisition and tracking of the opposite terminal. In Section 5, we use the theoretical models of th~ previous sections to perform an Earth-Mars link analysis over a full synodic period of 780 days under the simplifying assumption of coaxial, coplanar, circular orbits. We demonstrate that, using slightly modified versions of existing space and ground based laser systems, an Earth-Mars transponder link is not only feasible but quite robust. We also demonstrate through analysis the advantages and feasibility of compact, low output power (<300 mW photon-counting transponders using NASA's developmental SLR2000 satellite laser ranging system as the Earth terminal. Section 6 provides a summary of the results and some concluding remarks regarding future applications.

Degnan, John J.↗

Modeling and Analysis of Asynchronous Systems Using SAL and Hybrid SAL

We present formal models and results of formal analysis of two different asynchronous systems. We first examine a mid-value select module that merges the signals coming from three different sensors that are each asynchronously sampling the same input signal. We then consider the phase locking protocol proposed by Daly, Hopkins, and McKenna. This protocol is designed to keep a set of non-faulty (asynchronous) clocks phase locked even in the presence of Byzantine-faulty clocks on the network. All models and verifications have been developed using the SAL model checking tools and the Hybrid SAL abstractor.

Tiwari, Ashish↗

Asynchronicity in opposed-piston RCMs: Does it matter?

Rapid Compression Machines (RCMs) are widely utilized to study combustion phenomena at engine-relevant conditions, and significant efforts are typically made to create a quiescent environment, particularly for investigations of autoignition chemistry. Opposed-piston configurations can be advantageous due to shorter compression times and reduced surface area to volume ratios. Each side must be actuated simultaneously, but this can be challenging in practice. These devices, like most RCMs, utilize hydraulics for actuation, speed control and arrestation of the piston at the end of the stroke; there is no mechanical control or linkage of the two piston trajectories. To quantify the magnitudes and effects of piston asynchronous behavior, this work employs both detailed experimental measurements and, for the first time, high-fidelity, Direct Numerical Simulation (DNS). The boundary conditions are carefully considered applying insight from high-resolution linear variable differential transformer (LVDT) measurements of the piston trajectory and a zero-dimensional kinematics model of the piston-shaft assembly. Sufficient resolution in the piston crevice region is used. The complicated fluid dynamical behavior that can evolve during piston compression and the ensuing delay processes due to offset timings from t offset = 0-10 ms is elucidated. It is found that near t offset = 6 ms and beyond, the boundary layer on the face of the first-seating piston can be sufficiently perturbed, due initially to reemergence of gas from the crevice of the firstseating piston, so that the adiabatic core can become degraded at long ignition delay times. Substantial mixing of colder gas into the interior of the reaction chamber can alter the measurements, similar to effects previously observed for improper piston crevice configuration. In conclusion, experimental techniques to mitigate asynchronous behavior are discussed and demonstrated.

33 ADVANCED PROPULSION SYSTEMS↗

TomocuPy – efficient GPU-based tomographic reconstruction with asynchronous data processing

Fast 3D data analysis and steering of a tomographic experiment by changing environmental conditions or acquisition parameters require fast, close to real-time, 3D reconstruction of large data volumes. Here a performance-optimized TomocuPy package is presented as a GPU alternative to the commonly used central processing unit (CPU) based TomoPy package for tomographic reconstruction. TomocuPy utilizes modern hardware capabilities to organize a 3D asynchronous reconstruction involving parallel read/write operations with storage drives, CPU–GPU data transfers, and GPU computations. In the asynchronous reconstruction, all the operations are timely overlapped to almost fully hide all data management time. Since most cameras work with less than 16-bit digital output, the memory usage and processing speed are furthermore optimized by using 16-bit floating-point arithmetic. As a result, 3D reconstruction with TomocuPy became 20–30 times faster than its multi-threaded CPU equivalent. Full reconstruction (including read/write operations and methods initialization) of a 2048 3 tomographic volume takes less than 7 s on a single Nvidia Tesla A100 and PCIe 4.0 NVMe SSD, and scales almost linearly increasing the data size. To simplify operation at synchrotron beamlines, TomocuPy provides an easy-to-use command-line interface. Efficacy of the package was demonstrated during a tomographic experiment on gas-hydrate formation in porous samples, where a steering option was implemented as a lens-changing mechanism for zooming to regions of interest.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

An 802pW 93% Peak Efficiency Buck Converter with 5.5×10 6 Dynamic Range Featuring Fast DVFS and Asynchronous Load-Transient Control

Here this paper presents a buck converter with sub-nW quiescent power, high efficiency, and a wide dynamic range for ultra-low-power (ULP) IoT SoCs. To optimize the SoC power consumption, the buck converter supports fast dynamic voltage and frequency scaling (DVFS) and enables fast load-transient response (FLTR) through asynchronous control. In addition, the buck converter is fully self-contained with all features integrated on chip including a proposed adaptive deadtime controller. Fabricated in 65nm CMOS, measurement results show the buck converter has an 802pW quiescent power at 1.5V input voltage and a 93% peak efficiency. The measured dynamic range is from 0.5nW to 2.75mW, which is over 6 orders of magnitude. The measured voltage droop is 54mV for a 45nA-to-1mA load current step thanks to the asynchronous load-transient detector. The buck converter achieves the highest efficiency and widest dynamic range among all the state-of-the-art sub-nW switching voltage regulators, which makes it well suited for power management in ULP SoCs.

42 ENGINEERING↗