Engineering PapersSearch

SEARCH · Engineering Papers

Results for “faults”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Abnormal fault-recovery characteristics of the fault-tolerant multiprocessor uncovered using a new fault-injection methodology

An investigation was made in AIRLAB of the fault handling performance of the Fault Tolerant MultiProcessor (FTMP). Fault handling errors detected during fault injection experiments were characterized. In these fault injection experiments, the FTMP disabled a working unit instead of the faulted unit once in every 500 faults, on the average. System design weaknesses allow active faults to exercise a part of the fault management software that handles Byzantine or lying faults. Byzantine faults behave such that the faulted unit points to a working unit as the source of errors. The design's problems involve: (1) the design and interface between the simplex error detection hardware and the error processing software, (2) the functional capabilities of the FTMP system bus, and (3) the communication requirements of a multiprocessor architecture. These weak areas in the FTMP's design increase the probability that, for any hardware fault, a good line replacement unit (LRU) is mistakenly disabled by the fault management software.

Padilla, Peter A.

Model-based fault detection and isolation for intermittently active faults with application to motion-based thruster fault detection and isolation for spacecraft

The present invention is a method for detecting and isolating fault modes in a system having a model describing its behavior and regularly sampled measurements. The models are used to calculate past and present deviations from measurements that would result with no faults present, as well as with one or more potential fault modes present. Algorithms that calculate and store these deviations, along with memory of when said faults, if present, would have an effect on the said actual measurements, are used to detect when a fault is present. Related algorithms are used to exonerate false fault modes and finally to isolate the true fault mode. This invention is presented with application to detection and isolation of thruster faults for a thruster-controlled spacecraft. As a supporting aspect of the invention, a novel, effective, and efficient filtering method for estimating the derivative of a noisy signal is presented.

Wilson, Edward

Rapid recovery from transient faults in the fault-tolerant processor with fault-tolerant shared memory

The Draper fault-tolerant processor with fault-tolerant shared memory (FTP/FTSM), which is designed to allow application tasks to continue execution during the memory alignment process, is described. Processor performance is not affected by memory alignment. In addition, the FTP/FTSM incorporates a hardware scrubber device to perform the memory alignment quickly during unused memory access cycles. The FTP/FTSM architecture is described, followed by an estimate of the time required for channel reintegration.

Harper, Richard E.

Fault Propagation, EMI Propagation, and Fault Containment in Aerospace Systems

The occurrence of faults in aerospace system hardware and software have consequences ranging from minor effects to catastrophic effects, and such faults can directly affect the safety of hardware and personnel. There are many origins to fault conditions, and the hardware that is capable of still meeting its performance requirements after experiencing itself a fault is said to be fault tolerant. A fault tolerant hardware is capable of detecting, isolating, and recovering from a fault condition; and this is a subfield of control engineering. An aerospace system that has been shown to have electromagnetic compatibility (EMC) in all its subsystems and systems cannot induced faults caused by electromagnetic interference (EMI). It can be proposed that the presence of EMI (or lack of EMC) is analogous to a potential fault initiator and the effects can likewise range from minor to severe. This paper starts by addressing the consequences of hardware failure in aerospace systems from a fault perspective, because the design of fault tolerant system is a major endeavor in aerospace. To arrive to this goal the paper starts with the concepts of fault, fault propagation, and a new concept called fault containment region. The paper then proceeds to provide two very recent examples in the aircraft industry of fault propagation with catastrophic effects. The paper proceeds to introduce the concept of EMI fault containment and a brief introduction to another new concept called the EMI containment region. The paper proceeds with an example of EMI fault containment region. The paper ends with a lesson learned conclusions.

Perez, Reinaldo

Predeployment validation of fault-tolerant systems through software-implemented fault insertion

Fault injection-based automated testing (FIAT) environment, which can be used to experimentally characterize and evaluate distributed realtime systems under fault-free and faulted conditions is described. A survey is presented of validation methodologies. The need for fault insertion based on validation methodologies is demonstrated. The origins and models of faults, and motivation for the FIAT concept are reviewed. FIAT employs a validation methodology which builds confidence in the system through first providing a baseline of fault-free performance data and then characterizing the behavior of the system with faults present. Fault insertion is accomplished through software and allows faults or the manifestation of faults to be inserted by either seeding faults into memory or triggering error detection mechanisms. FIAT is capable of emulating a variety of fault-tolerant strategies and architectures, can monitor system activity, and can automatically orchestrate experiments involving insertion of faults. There is a common system interface which allows ease of use to decrease experiment development and run time. Fault models chosen for experiments on FIAT have generated system responses which parallel those observed in real systems under faulty conditions. These capabilities are shown by two example experiments each using a different fault-tolerance strategy.

Czeck, Edward W.

Fault recovery characteristics of the fault tolerant multi-processor

The fault handling performance of the fault tolerant multiprocessor (FTMP) was investigated. Fault handling errors detected during fault injection experiments were characterized. In these fault injection experiments, the FTMP disabled a working unit instead of the faulted unit once every 500 faults, on the average. System design weaknesses allow active faults to exercise a part of the fault management software that handles byzantine or lying faults. It is pointed out that these weak areas in the FTMP's design increase the probability that, for any hardware fault, a good LRU (line replaceable unit) is mistakenly disabled by the fault management software. It is concluded that fault injection can help detect and analyze the behavior of a system in the ultra-reliable regime. Although fault injection testing cannot be exhaustive, it has been demonstrated that it provides a unique capability to unmask problems and to characterize the behavior of a fault-tolerant system.

Padilla, Peter A.

Fault pattern at the northern end of the Death Valley - Furnace Creek fault zone, California and Nevada

The author has identified the following significant results. The pattern of faulting associated with the termination of the Death Valley-Furnace Creek Fault Zone in northern Fish Lake Valley, Nevada was studied in ERTS-1 MSS color composite imagery and color IR U-2 photography. Imagery analysis was supported by field reconnaissance and low altitude aerial photography. The northwest-trending right-lateral Death Valley-Furnace Creek Fault Zone changes northward to a complex pattern of discontinuous dip slip and strike slip faults. This fault pattern terminates to the north against an east-northeast trending zone herein called the Montgomery Fault Zone. No evidence for continuation of the Death Valley-Furnace Creek Fault Zone is recognized north of the Montgomery Fault Zone. Penecontemporaneous displacement in the Death Valley-Furnace Creek Fault Zone, the complex transitional zone, and the Montgomery Fault Zone suggests that the systems are genetically related. Mercury mineralization appears to have been localized along faults recognizable in ERTS-1 imagery within the transitional zone and the Montgomery Fault Zone.

Liggett, M. A.

Transform fault earthquakes in the North Atlantic: Source mechanisms and depth of faulting

The centroid depths and source mechanisms of 12 large earthquakes on transform faults of the northern Mid-Atlantic Ridge were determined from an inversion of long-period body waveforms. The earthquakes occurred on the Gibbs, Oceanographer, Hayes, Kane, 15 deg 20 min, and Vema transforms. The depth extent of faulting during each earthquake was estimated from the centroid depth and the fault width. The source mechanisms for all events in this study display the strike slip motion expected for transform fault earthquakes; slip vector azimuths agree to 2 to 3 deg of the local strike of the zone of active faulting. The only anomalies in mechanism were for two earthquakes near the western end of the Vema transform which occurred on significantly nonvertical fault planes. Secondary faulting, occurring either precursory to or near the end of the main episode of strike-slip rupture, was observed for 5 of the 12 earthquakes. For three events the secondary faulting was characterized by reverse motion on fault planes striking oblique to the trend of the transform. In all three cases, the site of secondary reverse faulting is near a compression jog in the current trace of the active transform fault zone. No evidence was found to support the conclusions of Engeln, Wiens, and Stein that oceanic transform faults in general are either hotter than expected from current thermal models or weaker than normal oceanic lithosphere.

Bergman, Eric A.

Transform fault earthquakes in the North Atlantic - Source mechanisms and depth of faulting

The centroid depths and source mechanisms of 12 large earthquakes on transform faults of the northern Mid-Atlantic Ridge were determined from an inversion of long-period body waveforms. The earthquakes occurred on the Gibbs, Oceanographer, Hayes, Kane, 15 deg 20 min, and Vema transforms. The depth extent of faulting during each earthquake was estimated from the centroid depth and the fault width. The source mechanisms for all events in this study display the strike slip motion expected for transform fault earthquakes; slip vector azimuths agree to 2 to 3 deg of the local strike of the zone of active faulting. The only anomalies in mechanism were for two earthquakes near the western end of the Vema transform which occurred on significantly nonvertical fault planes. Secondary faulting, occurring either precursory to or near the end of the main episode of strike-slip rupture, was observed for 5 of the 12 earthquakes. For three events the secondary faulting was characterized by reverse motion on fault planes striking oblique to the trend of the transform. In all three cases, the site of secondary reverse faulting is near a compression jog in the current trace of the active transform fault zone. No evidence was found to support the conclusions of Engeln, Wiens, and Stein that oceanic transform faults in general are either hotter than expected from current thermal models or weaker than normal oceanic lithosphere.

Bergman, Eric A.

On Identifiability of Bias-Type Actuator-Sensor Faults in Multiple-Model-Based Fault Detection and Identification

This paper explores a class of multiple-model-based fault detection and identification (FDI) methods for bias-type faults in actuators and sensors. These methods employ banks of Kalman-Bucy filters to detect the faults, determine the fault pattern, and estimate the fault values, wherein each Kalman-Bucy filter is tuned to a different failure pattern. Necessary and sufficient conditions are presented for identifiability of actuator faults, sensor faults, and simultaneous actuator and sensor faults. It is shown that FDI of simultaneous actuator and sensor faults is not possible using these methods when all sensors have biases.

Joshi, Suresh M.

Fault Scarp Detection Beneath Dense Vegetation Cover: Airborne Lidar Mapping of the Seattle Fault Zone, Bainbridge Island, Washington State

The emergence of a commercial airborne laser mapping industry is paying major dividends in an assessment of earthquake hazards in the Puget Lowland of Washington State. Geophysical observations and historical seismicity indicate the presence of active upper-crustal faults in the Puget Lowland, placing the major population centers of Seattle and Tacoma at significant risk. However, until recently the surface trace of these faults had never been identified, neither on the ground nor from remote sensing, due to cover by the dense vegetation of the Pacific Northwest temperate rainforests and extremely thick Pleistocene glacial deposits. A pilot lidar mapping project of Bainbridge Island in the Puget Sound, contracted by the Kitsap Public Utility District (KPUD) and conducted by Airborne Laser Mapping in late 1996, spectacularly revealed geomorphic features associated with fault strands within the Seattle fault zone. The features include a previously unrecognized fault scarp, an uplifted marine wave-cut platform, and tilted sedimentary strata. The United States Geologic Survey (USGS) is now conducting trenching studies across the fault scarp to establish ages, displacements, and recurrence intervals of recent earthquakes on this active fault. The success of this pilot study has inspired the formation of a consortium of federal and local organizations to extend this work to a 2350 square kilometer (580,000 acre) region of the Puget Lowland, covering nearly the entire extent (approx. 85 km) of the Seattle fault. The consortium includes NASA, the USGS, and four local groups consisting of KPUD, Kitsap County, the City of Seattle, and the Puget Sound Regional Council (PSRC). The consortium has selected Terrapoint, a commercial lidar mapping vendor, to acquire the data.

Harding, David J.

Fault type predictions from stress distributions on planetary surfaces - Importance of fault initiation depth

The prediction of fault type on planetary surfaces from model stresses calculated at depth is discussed. These fault-type predictions yield different faults than those predicted using the surface criteria commonly employed in geophysical models. For elastic-plate flexure models of mascon loading on the moon, stresses calculated at the surface predict the occurrence of strike-slip faulting at the radial distance where grabens are found. Normal faults bounding lunar grabens and thrust faults responsible for wrinkle ridges are analyzed. It is found that the former initiate at the mechanical discontinuity that separates the breccia of the megaregolith from in situ fractured rock and that the latter initiate at the mechanical discontinuity between basalt layers and the underlying basin floor. The difference between elastic constants for the outer few kilometers of brecciated megaregolith and the underlying lunar lithosphere are evaluated. Superposing nonisotropic stresses resulting from the weight of overburden to the depth of the relevant mechanical discontinuity yield stresses that predict wrinkle ridges in the basin centers and grabens outside the basin margin, and eliminate the predicted zone of strike-slip faults.

Golombek, M. P.

Characterization of fault recovery through fault injection on FTMP

The development of fault-injection procedures and statistical analysis techniques to characterize the fault recovery of fault-tolerant systems is described. Pin-level fault-injection was conducted on a fault-tolerant microprocessor computer in order to generate data to assess the utility of current fault-injection sampling methods. The validity of common reliability-modeling assumptions concerning the statistical distribution of recovery times is investigated. A multiple comparison analysis for detecting behavior variations, and a distribution fitting for determining the best fit for the data were conducted. It is observed that the detection behavior is not homogeneous across all data sets, and that none of the factors under experimental control can account for the observed groupings of behavior. It is determined that no single distribution fits all the data sets, and that stratified random sampling and statistically robust parameter-estimation techniques are required to characterize fault detection time.

Finelli, George B.

Multi-version software reliability through fault-avoidance and fault-tolerance

A number of experimental and theoretical issues associated with the practical use of multi-version software to provide run-time tolerance to software faults were investigated. A specialized tool was developed and evaluated for measuring testing coverage for a variety of metrics. The tool was used to collect information on the relationships between software faults and coverage provided by the testing process as measured by different metrics (including data flow metrics). Considerable correlation was found between coverage provided by some higher metrics and the elimination of faults in the code. Back-to-back testing was continued as an efficient mechanism for removal of un-correlated faults, and common-cause faults of variable span. Software reliability estimation methods was also continued based on non-random sampling, and the relationship between software reliability and code coverage provided through testing. New fault tolerance models were formulated. Simulation studies of the Acceptance Voting and Multi-stage Voting algorithms were finished and it was found that these two schemes for software fault tolerance are superior in many respects to some commonly used schemes. Particularly encouraging are the safety properties of the Acceptance testing scheme.

Vouk, Mladen A.

Flight elements: Fault detection and fault management

Fault management for an intelligent computational system must be developed using a top down integrated engineering approach. An approach proposed includes integrating the overall environment involving sensors and their associated data; design knowledge capture; operations; fault detection, identification, and reconfiguration; testability; causal models including digraph matrix analysis; and overall performance impacts on the hardware and software architecture. Implementation of the concept to achieve a real time intelligent fault detection and management system will be accomplished via the implementation of several objectives, which are: Development of fault tolerant/FDIR requirement and specification from a systems level which will carry through from conceptual design through implementation and mission operations; Implementation of monitoring, diagnosis, and reconfiguration at all system levels providing fault isolation and system integration; Optimize system operations to manage degraded system performance through system integration; and Lower development and operations costs through the implementation of an intelligent real time fault detection and fault management system and an information management system.

Lum, H.

An empirical comparison of software fault tolerance and fault elimination

Reliability is an important concern in the development of software for modern systems. Some researchers have hypothesized that particular fault-handling approaches or techniques are so effective that other approaches or techniques are superfluous. The authors have performed a study that compares two major approaches to the improvement of software, software fault elimination and software fault tolerance, by examination of the fault detection obtained by five techniques: run-time assertions, multi-version voting, functional testing augmented by structural testing, code reading by stepwise abstraction, and static data-flow analysis. This study has focused on characterizing the sets of faults detected by the techniques and on characterizing the relationships between these sets of faults. The results of the study show that none of the techniques studied is necessarily redundant to any combination of the others. Further results reveal strengths and weakness in the fault detection by the techniques studied and suggest directions for future research.

Shimeall, Timothy J.

Stress near geometrically complex strike-slip faults - Application to the San Andreas fault at Cajon Pass, southern California

A model is presented to rationalize the state of stress near a geometrically complex major strike-slip fault. Slip on such a fault creates residual stresses that, with the occurrence of several slip events, can dominate the stress field near the fault. The model is applied to the San Andreas fault near Cajon Pass. The results are consistent with the geological features, seismicity, the existence of left-lateral stress on the Cleghorn fault, and the in situ stress orientation in the scientific well, found to be sinistral when resolved on a plane parallel to the San Andreas fault. It is suggested that the creation of residual stresses caused by slip on a wiggle San Andreas fault is the dominating process there.

Saucier, Francois

Multiversion software reliability through fault-avoidance and fault-tolerance

In this project we have proposed to investigate a number of experimental and theoretical issues associated with the practical use of multi-version software in providing dependable software through fault-avoidance and fault-elimination, as well as run-time tolerance of software faults. In the period reported here we have working on the following: We have continued collection of data on the relationships between software faults and reliability, and the coverage provided by the testing process as measured by different metrics (including data flow metrics). We continued work on software reliability estimation methods based on non-random sampling, and the relationship between software reliability and code coverage provided through testing. We have continued studying back-to-back testing as an efficient mechanism for removal of uncorrelated faults, and common-cause faults of variable span. We have also been studying back-to-back testing as a tool for improvement of the software change process, including regression testing. We continued investigating existing, and worked on formulation of new fault-tolerance models. In particular, we have partly finished evaluation of Consensus Voting in the presence of correlated failures, and are in the process of finishing evaluation of Consensus Recovery Block (CRB) under failure correlation. We find both approaches far superior to commonly employed fixed agreement number voting (usually majority voting). We have also finished a cost analysis of the CRB approach.

Vouk, Mladen A.