Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “failure recovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Transforming AdaPT to Ada

This paper describes how the main features of the proposed Ada language extensions intended to support distribution, and offered as possible solutions for Ada9X can be implemented by transformation into standard Ada83. We start by summarizing the features proposed in a paper (Gargaro et al, 1990) which constitutes the definition of the extensions. For convenience we have called the language in its modified form AdaPT which might be interpreted as Ada with partitions. These features were carefully chosen to provide support for the construction of executable modules for execution in nodes of a network of loosely coupled computers, but flexibly configurable for different network architectures and for recovery following failure, or adapting to mode changes. The intention in their design was to provide extensions which would not impact adversely on the normal use of Ada, and would fit well in style and feel with the existing standard. We begin by summarizing the features introduced in AdaPT.

Goldsack, Stephen J.↗

Performance of redundant disk array organizations in transaction processing environments

A performance evaluation is conducted for two redundant disk-array organizations in a transaction-processing environment, relative to the performance of both mirrored disk organizations and organizations using neither striping nor redundancy. The proposed parity-striping alternative to striping with rotated parity is shown to furnish rapid recovery from failure at the same low storage cost without interleaving the data over multiple disks. Both noncached systems and systems using a nonvolatile cache as the controller are considered.

Mourad, Antoine N.↗

Performance Characterization and Simulation of Amine-Based Vacuum Swing Sorption Units for Spacesuit Carbon Dioxide and Humidity Control

Controlling carbon dioxide (CO2) and water (H2O) vapor concentrations in a space suit is critical to ensuring an astronauts safety, comfort, and capability to perform extra-vehicular activity (EVA) tasks. Historically, this has been accomplished using lithium hydroxide (LiOH) and metal oxide (MetOx) canisters. Lithium hydroxide is a consumable material that requires priming with water before it becomes effective at removing carbon dioxide. MetOx is regenerable through a power-intensive thermal cycle but is significantly heavier on a volume basis than LiOH. As an alternative, amine-based vacuum swing beds are under aggressive development for EVA applications. The vacuum swing units control atmospheric concentrations of both CO2 and H2O through fully-regenerative process. The current concept, referred to as the rapid cycle amine (RCA), has resulted in numerous laboratory prototypes. Performance of these prototypes have been assessed experimentally and documented in previous reports. To support developmental e orts, a first principles model has also been established for the vacuum swing sorption technology. For the first time in several decades, a major re-design of Portable Life Support System (PLSS) for the extra-vehicular mobility unit (EMU) is underway. NASA at Johnson Space Center built and tested an integrated PLSS test bed of all subsystems under a variety of simulated EVA conditions of which the RCA prototype played a significant role. The efforts documented herein summarize RCA test performance and simulation results for single and variable metabolic rate experiments in an integrated context. In addition, a variety of off-nominal tests were performed to assess the capability of the RCA to function under challenging circumstances. Tests included high water production experiments, degraded vacuum regeneration, and deliberate valve/power failure and recovery.

Swickrath, Michael J.↗

A failure management prototype: DR/Rx

This failure management prototype performs failure diagnosis and recovery management of hierarchical, distributed systems. The prototype, which evolved from a series of previous prototypes following a spiral model for development, focuses on two functions: (1) the diagnostic reasoner (DR) performs integrated failure diagnosis in distributed systems; and (2) the recovery expert (Rx) develops plans to recover from the failure. Issues related to expert system prototype design and the previous history of this prototype are discussed. The architecture of the current prototype is described in terms of the knowledge representation and functionality of its components.

Hammen, David G.↗

Implementing Journaling in a Linux Shared Disk File System

In computer systems today, speed and responsiveness is often determined by network and storage subsystem performance. Faster, more scalable networking interfaces like Fibre Channel and Gigabit Ethernet provide the scaffolding from which higher performance computer systems implementations may be constructed, but new thinking is required about how machines interact with network-enabled storage devices. In this paper we describe how we implemented journaling in the Global File System (GFS), a shared-disk, cluster file system for Linux. Our previous three papers on GFS at the Mass Storage Symposium discussed our first three GFS implementations, their performance, and the lessons learned. Our fourth paper describes, appropriately enough, the evolution of GFS version 3 to version 4, which supports journaling and recovery from client failures. In addition, GFS scalability tests extending to 8 machines accessing 8 4-disk enclosures were conducted: these tests showed good scaling. We describe the GFS cluster infrastructure, which is necessary for proper recovery from machine and disk failures in a collection of machines sharing disks using GFS. Finally, we discuss the suitability of Linux for handling the big data requirements of supercomputing centers.

Preslan, Kenneth W.↗

Application of compiler-assisted multiple instruction rollback recovery to speculative execution

Speculative execution is a method to increase instruction level parallelism which can be exploited by both super-scalar and VLIW architectures. The key to a successful general speculation strategy is a repair mechanism to handle mispredicted branches and accurate reporting of exceptions for speculated instructions. Multiple instruction rollback is a technique developed for recovery from transient processor failure. Many of the difficulties encountered during recovery from branch misprediction or from instruction re-execution due to exception in a speculative execution architecture are similar to those encountered during multiple instruction rollback. The applicability of a recently developed compiler-assisted multiple instruction rollback scheme to aid in speculative execution repair is investigated. Extensions to the compiler-assisted scheme to support branch and exception repair are presented along with performance measurements across ten application programs.

Alewine, N. J.↗

Application of compiler-assisted multiple instruction rollback recovery to speculative execution

Speculative execution is a method to increase instruction level parallelism which can be exploited by both super-scalar and VLIW architectures. The key to a successful general speculation strategy is a repair mechanism to handle mispredicted branches and accurate reporting of exceptions for speculated instructions. Multiple instruction rollback is a technique developed for recovery from transient processor failure. Many of the difficulties encountered during recovery from branch misprediction or from instruction re-execution due to exception in a speculative execution architecture are similar to those encountered during multiple instruction rollback. The applicability of a recently developed compiler-assisted multiple instruction rollback scheme to aid in speculative execution repair is investigated. Extensions to the compiler-assisted scheme to support branch and exception repair are presented along with performance measurements across ten application programs.

Alewine, N. J.↗

IOPS advisor: Research in progress on knowledge-intensive methods for irregular operations airline scheduling

Our research focuses on the problem of recovering from perturbations in large-scale schedules, specifically on the ability of a human-machine partnership to dynamically modify an airline schedule in response to unanticipated disruptions. This task is characterized by massive interdependencies and a large space of possible actions. Our approach is to apply the following: qualitative, knowledge-intensive techniques relying on a memory of stereotypical failures and appropriate recoveries; and quantitative techniques drawn from the Operations Research community's work on scheduling. Our main scientific challenge is to represent schedules, failures, and repairs so as to make both sets of techniques applicable to the same data. This paper outlines ongoing research in which we are cooperating with United Airlines to develop our understanding of the scientific issues underlying the practicalities of dynamic, real-time schedule repair.

Borse, John E.↗

A neural network-based estimator for the mixture ratio of the Space Shuttle Main Engine

In order to properly utilize the available fuel and oxidizer of a liquid propellant rocket engine, the mixture ratio is closed loop controlled during main stage (65 percent - 109 percent power) operation. However, because of the lack of flight-capable instrumentation for measuring mixture ratio, the value of mixture ratio in the control loop is estimated using available sensor measurements such as the combustion chamber pressure and the volumetric flow, and the temperature and pressure at the exit duct on the low pressure fuel pump. This estimation scheme has two limitations. First, the estimation formula is based on an empirical curve fitting which is accurate only within a narrow operating range. Second, the mixture ratio estimate relies on a few sensor measurements and loss of any of these measurements will make the estimate invalid. In this paper, we propose a neural network-based estimator for the mixture ratio of the Space Shuttle Main Engine. The estimator is an extension of a previously developed neural network based sensor failure detection and recovery algorithm (sensor validation). This neural network uses an auto associative structure which utilizes the redundant information of dissimilar sensors to detect inconsistent measurements. Two approaches have been identified for synthesizing mixture ratio from measurement data using a neural network. The first approach uses an auto associative neural network for sensor validation which is modified to include the mixture ratio as an additional output. The second uses a new network for the mixture ratio estimation in addition to the sensor validation network. Although mixture ratio is not directly measured in flight, it is generally available in simulation and in test bed firing data from facility measurements of fuel and oxidizer volumetric flows. The pros and cons of these two approaches will be discussed in terms of robustness to sensor failures and accuracy of the estimate during typical transients using simulation data.

Guo, T. H.↗

Multiversion software reliability through fault-avoidance and fault-tolerance

In this project we have proposed to investigate a number of experimental and theoretical issues associated with the practical use of multi-version software in providing dependable software through fault-avoidance and fault-elimination, as well as run-time tolerance of software faults. In the period reported here we have working on the following: We have continued collection of data on the relationships between software faults and reliability, and the coverage provided by the testing process as measured by different metrics (including data flow metrics). We continued work on software reliability estimation methods based on non-random sampling, and the relationship between software reliability and code coverage provided through testing. We have continued studying back-to-back testing as an efficient mechanism for removal of uncorrelated faults, and common-cause faults of variable span. We have also been studying back-to-back testing as a tool for improvement of the software change process, including regression testing. We continued investigating existing, and worked on formulation of new fault-tolerance models. In particular, we have partly finished evaluation of Consensus Voting in the presence of correlated failures, and are in the process of finishing evaluation of Consensus Recovery Block (CRB) under failure correlation. We find both approaches far superior to commonly employed fixed agreement number voting (usually majority voting). We have also finished a cost analysis of the CRB approach.

Vouk, Mladen A.↗

DMTN-260: Failure Modes and Error Handling for Prompt Processing

The Prompt Processing system will be responsible for processing roughly a thousand visits per night, and distributing the results in near real time, for at least ten years of Rubin Observatory operations. As such, it must be highly robust to algorithmic, network, and infrastructure failures, ranging from momentary glitches to extended downtimes. DMTN-219 introduced the initial design for the Prompt Processing framework; this document expands on the design to address expected failure modes and recovery strategies for each.

79 ASTRONOMY AND ASTROPHYSICS↗

Simulation of the XV-15 tilt rotor research aircraft

The effective use of simulation from issuance of the request for proposal through conduct of a flight test program for the XV-15 Tilt Rotor Research Aircraft is discussed. From program inception, simulation complemented all phases of XV-15 development. The initial simulation evaluations during the source evaluation board proceedings contributed significantly to performance and stability and control evaluations. Eight subsequent simulation periods provided major contributions in the areas of control concepts; cockpit configuration; handling qualities; pilot workload; failure effects and recovery procedures; and flight boundary problems and recovery procedures. The fidelity of the simulation also made it a valuable pilot training aid, as well as a suitable tool for military and civil mission evaluations. Simulation also provided valuable design data for refinement of automatic flight control systems. Throughout the program, fidelity was a prime issue and resulted in unique data and methods for fidelity evaluation which are presented and discussed.

Churchill, G. B.↗

Synthetic bounds for semi-Markov reliability models

Upper and lower bounds are derived for the probability of failure for a class of highly reliable process control computers. The bounds are synthetic in the sense that the descriptions of component failure and system recovery are assumed to be obtained from different sources. The reliability model is constructed under the assumption that the processes are independent.

White, A. L.↗

Engineering challenges of in-flight spacecraft - Voyager: A case study

Some of the engineering problems encountered during the post-launch phase of interplanetary space missions are described, with emphasis given to the Voyager missions. The major in-flight modifications in Voyager spacecraft's operational capability with respect to communications, payload, and navigation systems are discussed. Attention is given to the instances of 'failure workaround' including: recovery from a failed receiver, recovery from a seized scan platform actuator, and modifications to the Attitude Articulation and Control Subsystem (AACS) software during the Saturn encounter. A detailed line drawing of the Voyager spacecraft is provided.

Jones, C. P.↗

An approximation formula for a class of fault-tolerant computers

An approximation formula is derived for the probability of failure for fault-tolerant process-control computers. These computers use redundancy and reconfiguration to achieve high reliability. Finite-state Markov models capture the dynamic behavior of component failure and system recovery, and the approximation formula permits an estimation of system reliability by an easy examination of the model.

White, A. L.↗