Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “rollback”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

On the feasibility of a spaceborne fault-tolerant hypercube

The feasibility of implementing a fault-tolerant hypercube architecture for space applications is discussed. Node-level architectures and designs are considered and a first-order reliability model is presented. It is shown how error recovery can be implemented using program rollback or roll-forward techniques. Shared memory augmentations to the message-passing structure can be used to get around the inefficiencies of multicomputers to provide efficient use of hardware to achieve the needed reliabilities while maintaining performance.

Rennels, David A.↗

Recoverable distributed shared virtual memory - Memory coherence and storage structures

This paper examines the problem of implementing rollback recovery in multicomputer distributed shared virtual memory environments, in which the shared memory is implemented in software and exists only virtually. A user-transparent checkpointing recovery scheme and new twin-page disk storage management are presented to implement a recoverable distributed shared virtual memory. The checkpointing scheme is integrated with the shared virtual memory management. The twin-page disk approach allows incremental checkpointing without an explicit undo at the time of recovery. A single consistent checkpoint state is maintained on stable disk storage. The recoverable distributed shared virtual memory allows the system to restart computation from a previous checkpoint due to a processor failure without a global restart.

Wu, Kun-Lung↗

Cache-based error recovery for shared memory multiprocessor systems

A multiprocessor cache-based checkpointing and recovery scheme for of recovering from transient processor errors in a shared-memory multiprocessor with private caches is presented. New implementation techniques that use checkpoint identifiers and recovery stacks to reduce performance degradation in processor utilization during normal execution are examined. This cache-based checkpointing technique prevents rollback propagation, provides for rapid recovery, and can be integrated into standard cache coherence protocols. An analytical model is used to estimate the relative performance of the scheme during normal execution. Extensions that take error latency into account are presented.

Wu, Kun-Lung↗

Error recovery in shared memory multiprocessors using private caches

The problem of recovering from processor transient faults in shared memory multiprocesses systems is examined. A user-transparent checkpointing and recovery scheme using private caches is presented. Processes can recover from errors due to faulty processors by restarting from the checkpointed computation state. Implementation techniques using checkpoint identifiers and recovery stacks are examined as a means of reducing performance degradation in processor utilization during normal execution. This cache-based checkpointing technique prevents rollback propagation, provides rapid recovery, and can be integrated into standard cache coherence protocols. An analytical model is used to estimate the relative performance of the scheme during normal execution. Extensions to take error latency into account are presented.

Wu, Kun-Lung↗

Recoverable distributed shared virtual memory

The problem of rollback recovery in distributed shared virtual environments, in which the shared memory is implemented in software in a loosely coupled distributed multicomputer system, is examined. A user-transparent checkpointing recovery scheme and a new twin-page disk storage management technique are presented for implementing recoverable distributed shared virtual memory. The checkpointing scheme can be integrated with the memory coherence protocol for managing the shared virtual memory. The twin-page disk design allows checkpointing to proceed in an incremental fashion without an explicit undo at the time of recovery. The recoverable distributed shared virtual memory allows the system to restart computation from a checkpoint without a global restart.

Wu, Kun-Lung↗

A conservative approach to parallelizing the Sharks World simulation

Parallelizing a benchmark problem for parallel simulation, the Sharks World, is described. The described solution is conservative, in the sense that no state information is saved, and no 'rollbacks' occur. The used approach illustrates both the principal advantage and principal disadvantage of conservative parallel simulation. The advantage is that by exploiting lookahead an approach was found that dramatically improves the serial execution time, and also achieves excellent speedups. The disadvantage is that if the model rules are changed in such a way that the lookahead is destroyed, it is difficult to modify the solution to accommodate the changes.

Nicol, David M.↗

Building a generalized distributed system model

A number of topics related to building a generalized distributed system model are discussed. The effects of distributed database modeling on evaluation of transaction rollbacks, the measurement of effects of distributed database models on transaction availability measures, and a performance analysis of static locking in replicated distributed database systems are covered.

Mukkamala, Ravi↗

Mission safety evaluation report for STS-39, postflight edition

After a delay of approximately 2 months due to a rollback from the pad to replace the External Tank door lug housing, Space Shuttle Discovery was launched from NASA-Kennedy at 7:33 a.m. Eastern Daylight Time on 28 April 1991. STS-39 was the first unclassified DoD Shuttle mission. On 28 April, countdown proceeded normally through the T-20 minute hold. No significant problems were encountered except for the Operations Sequence-2 recorder starting unexpectedly; it was stopped by an uplink command. Discovery landed on KSC runway 15 at 2:55 p.m. EDT on 6 May 1991. This was the second time in 6 months that the Space Shuttle was diverted to KSC for landing because of high winds at Edwards AFB, Calif. This was also the 7th of 40 Shuttle missions to land at KSC in the history of the Space Shuttle Program. The Main Landing Gear outer right tire shredded 3 of the 16 cords due to either an uneven landing or a maximum force breaking test during rollout. Contributing factors to the tire cord shredding were the development of last minute crosswinds and reluctance of the ground controllers to distract the Shuttle pilots with warnings of the low flight path. As a corrective action, communication procedures will be modified for future flights.

Hardie, Kenneth O.↗

Design and scheduling for periodic concurrent error detection and recovery in processor arrays

Periodic application of time-redundant error checking provides the trade-off between error detection latency and performance degradation. The goal is to achieve high error coverage while satisfying performance requirements. We derive the optimal scheduling of checking patterns in order to uniformly distribute the available checking capability and maximize the error coverage. Synchronous buffering designs using data forwarding and dynamic reconfiguration are described. Efficient single-cycle diagnosis is implemented by error pattern analysis and direct-mapped recovery cache. A rollback recovery scheme using start-up control for local recovery is also presented.

Wang, Yi-Min↗

Time Warp Operating System (TWOS)

Designed to support parallel discrete-event simulation, TWOS is complete implementation of Time Warp mechanism - distributed protocol for virtual time synchronization based on process rollback and message annihilation.

Bellenot, Steven F.↗

Time Warp Operating System, Version 2.5.1

Time Warp Operating System, TWOS, is special purpose computer program designed to support parallel simulation of discrete events. Complete implementation of Time Warp software mechanism, which implements distributed protocol for virtual synchronization based on rollback of processes and annihilation of messages. Supports simulations and other computations in which both virtual time and dynamic load balancing used. Program utilizes underlying resources of operating system. Written in C programming language.

Bellenot, Steven F.↗

Branch recovery with compiler-assisted multiple instruction retry

In processing systems where rapid recovery from transient faults is important, schemes for multiple instruction rollback recovery may be appropriate. Multiple instruction retry has been implemented in hardware by researchers and also in mainframe computers. This paper extends compiler-assisted instruction retry to a broad class of code execution failures. Five benchmarks were used to measure the performance penalty of hazard resolution. Results indicate that the enhanced pure software approach can produce performance penalties consistent with existing hardware techniques. A combined compiler/hardware resolution strategy is also described and evaluated. Experimental results indicate a lower performance penalty than with either a totally hardware or totally software approach.

Alewine, N. J.↗

Relaxing consistency in recoverable distributed shared memory

Relaxed memory consistency models have recently been proposed to tolerate memory access latency in both hardware and software distributed shared memory systems. In recoverable shared memory multiprocessors, relaxing consistency has the added benefit of reducing the number of checkpoints needed to avoid rollback propagation. In this paper, we introduce new checkpointing algorithms that take advantage of relaxed consistency to reduce the performance overhead of checkpointing. We also introduce a scheme based on lazy relaxed consistency, that reduces both checkpointing overhead and the overhead of avoiding error propagation in systems with error latency. Multiprocessor address traces are used to evaluate the relaxed consistency approach to checkpointing with distributed shared memory.

Janssens, Bob↗

Flexure and the role of inplane force around coronae on Venus

Large coronae on Venus, such as Artemis and Latona, are rimmed by conspicuous trenches and associated outer rises. Sandwell and Schubert have observed that these systems resemble terrestrial subduction zones in planform and have succeeded in fitting an elastic plate bending equation to the inferred flexural topography. However, the first zero crossing bending moments required are -2.5 x 10(exp 17) N for Artemis and -5.0 x 10(exp 16) N for Latona. Since these moments are similar in magnitude to those of subducting slabs on Earth, a rollback subduction mechanism was proposed to explain the flexure around the largest coronae, although a differential thermal subsidence model is sufficient to account for the topography around some coronae. The purpose is to investigate the effect of inplane force as a possible alternative to large applied moments in producing flexure at Artemis and Latona. The close correlation of gravity to topography on Venus implies the absence of a low viscosity zone and the strong coupling of the lithosphere to mantle convection. If coronae are the surface manifestations of mantle plumes, they may be the sites of active convective stress coupling. As the upwelling reaches the lithosphere, it spreads radially outward, inducing shear tractions on the base of the plate. In addition, the hot, expanding corona may load the surrounding plate horizonally. Both the basal shear stresses and radial loading can be treated as an equivalent compressive inplane force in the mechanical lithosphere, which contributes to the bending of the outlying plate. Using a model that relates inplane force to the measured gravity anomalies, a rough value of the inplane force at Artemis was calculated. Recent Pioneer Venus spherical harmonic gravity models indicate a geoid anomaly of about 75 m over Artemis, which corresponds to an estimated inplane force on the order of -1x10(exp 13) N/m. The gravity model is unable to resolve Latona, but an inplane force of similar dimensions is assumed. The maximum possible inplane force based on the expected rheology can be constrained by using the approximate 5 K/km thermal gradient inferred from the best fit 30 km elastic plate at Artemis and Latona. For a dry olivine flow law in the upper mantle, the compressional load limit of the 60 km thick mechanical lithosphere is -4 x 10(exp 13) N/m. This value is equivalent to a load of -8 x 10(exp 13) N/m on a 30 km thick elastic plate.

Brown, C. David↗

Efficient massively parallel simulation of dynamic channel assignment schemes for wireless cellular communications

Fast, efficient parallel algorithms are presented for discrete event simulations of dynamic channel assignment schemes for wireless cellular communication networks. The driving events are call arrivals and departures, in continuous time, to cells geographically distributed across the service area. A dynamic channel assignment scheme decides which call arrivals to accept, and which channels to allocate to the accepted calls, attempting to minimize call blocking while ensuring co-channel interference is tolerably low. Specifically, the scheme ensures that the same channel is used concurrently at different cells only if the pairwise distances between those cells are sufficiently large. Much of the complexity of the system comes from ensuring this separation. The network is modeled as a system of interacting continuous time automata, each corresponding to a cell. To simulate the model, conservative methods are used; i.e., methods in which no errors occur in the course of the simulation and so no rollback or relaxation is needed. Implemented on a 16K processor MasPar MP-1, an elegant and simple technique provides speedups of about 15 times over an optimized serial simulation running on a high speed workstation. A drawback of this technique, typical of conservative methods, is that processor utilization is rather low. To overcome this, new methods were developed that exploit slackness in event dependencies over short intervals of time, thereby raising the utilization to above 50 percent and the speedup over the optimized serial code to about 120 times.

Greenberg, Albert G.↗

Relaxing consistency in recoverable distributed shared memory

Relaxed memory consistency models tolerate increased memory access latency in both hardware and software distributed shared memory systems. In recoverable systems, relaxing consistency has the added benefit of reducing the number of checkpoints needed to avoid rollback propagation. In this paper, we introduce new checkpointing algorithms that take advantage of relaxed consistency to reduce the performance overhead of checkpointing. We also introduce a scheme based on lazy relaxed consistency, that reduces both checkpointing overhead and the overhead of avoiding error propagation in systems with error latency. We use multiprocessor address traces to evaluate the relaxed consistency approach to checkpointing with distributed shared memory.

Janssens, Bob↗

A Scheduling Algorithm for Replicated Real-Time Tasks

We present an algorithm for scheduling real-time periodic tasks on a multiprocessor system under fault-tolerant requirement. Our approach incorporates both the redundancy and masking technique and the imprecise computation model. Since the tasks in hard real-time systems have stringent timing constraints, the redundancy and masking technique are more appropriate than the rollback techniques which usually require extra time for error recovery. The imprecise computation model provides flexible functionality by trading off the quality of the result produced by a task with the amount of processing time required to produce it. It therefore permits the performance of a real-time system to degrade gracefully. We evaluate the algorithm by stochastic analysis and Monte Carlo simulations. The results show that the algorithm is resilient under hardware failures.

Yu, Albert C.↗

Reducing Interprocessor Dependence in Recoverable Distributed Shared Memory

Checkpointing techniques in parallel systems use dependency tracking and/or message logging to ensure that a system rolls back to a consistent state. Traditional dependency tracking in distributed shared memory (DSM) systems is expensive because of high communication frequency. In this paper we show that, if designed correctly, a DSM system only needs to consider dependencies due to the transfer of blocks of data, resulting in reduced dependency tracking overhead and reduced potential for rollback propagation. We develop an ownership timestamp scheme to tolerate the loss of block state information and develop a passive server model of execution where interactions between processors are considered atomic. With our scheme, dependencies are significantly reduced compared to the traditional message-passing model.

Janssens, Bob↗