Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Fault-tolerance”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Fault-tolerance experiments with the JPL STAR computer.

Results of fault-tolerance experiments performed using an experimental computer with dynamic (standby) redundancy, including replaceable subsystems and a 'program rollback' provision to eliminate transient-caused errors. After a brief review of the specification of fault-tolerance with respect to transient faults, including a description of the method of injection of transient faults in software and system tests, fault-tolerance experiments carried out with this computer with regard to the determination of fault classes, software verification, system verification, and recovery stability are summarized. A test and repair processor is described which constitutes a special monitor unit of the computer and is used to obtain information for fault detection in the other subsystems of the computer and to ensure that proper recovery occurs when a fault is detected.

Avizienis, A.

Enabling Reliable, Fault-Tolerant Autonomous Lunar Habitats with High-Performance Spaceflight Computing

The lunar surface presents unfavorable constraints and harsh living conditions. To address these challenges, autonomous habitats will require complex integrated systems that combine advanced software, high-performance hardware, and cutting-edge sensors to ensure sustainability, safety, and operational efficiency. Consequently, maintaining a sustainable presence on the Moon requires reliable infrastructure and efficient development, precise monitoring, and utilization of resources within a lunar installation. These elements are essential not only to ensure that lunar settlement can be long-term, self-sustaining, and resource-efficient, but also to serve as a foundation for future missions and eventual human habitation on Mars. Humans are not native to the Moon; therefore, our survival and ability to thrive will depend on autonomous systems that can foster safety and resilience through high-availability architectures, graceful degradation, and highly fault-tolerant spaceflight hardware capable of continuing operation during failures. This requires advanced human-rated distributed systems architectures with specialized electronics, scalable capabilities, and an integrated design approach. Unlike current practices focused on short-term missions and regularly maintained components, permanent lunar compute systems must be designed for extended operations beyond mission durations. This paper explores the necessity of transitioning toward fault- tolerant, highly autonomous hardware systems designed for multi-year missions. It also identifies critical subsystems that require high levels of autonomy, supported by radiation-hardened processors and extreme thermal loads, which are essential to mitigate long-term degradation and ensure sustainable lunar habitation. Finally, the paper aligns with NASA’s identified Civil Space Shortfalls, particularly in high-performance onboard computing, advanced data acquisition, extreme-environment avionics, radiation monitoring and countermeasures, and autonomous health management. It proposes NASA’s new High-Performance Spaceflight Computing (HPSC) processor as a turnkey solution, delivering 100 times the performance-per-watt of legacy rad-hard CPUs and enabling onboard AI, edge computing, and fault-tolerant features essential for sustained lunar autonomy and beyond.

Sarkis S Mikaelian

The STAR /self-testing and repairing/ computer - An investigation of the theory and practice of fault-tolerant computer design.

This paper presents the results obtained in a continuing investigation of fault-tolerant computing which is being conducted at the Jet Propulsion Laboratory. Initial studies led to the decision to design and construct an experimental computer with dynamic (standby) redundancy, including replaceable subsystems and a program rollback provision to eliminate transient errors. This system, called the STAR computer, began operation in 1969. The following aspects of the STAR system are described: architecture, reliability analysis, software, automatic maintenance of peripheral systems, and adaptation to serve as the central computer of an outer-planet exploration spacecraft.

A Avizienis

On reliability modeling and analysis of ultrareliable fault-tolerant digital systems.

The processes of protective redundancy, namely, standby replacement (SR) redundancy and hybrid redundancy (a combination of SR and multiple-line voting redundancy), find application in the architecture of fault-tolerant digital computers and enable them to be ultrareliable and self-repairing. The claims to ultrareliability lead to the challenge of quantitatively evaluating and assigning a value to the probability of survival as a function of the mission durations intended. This note presents various mathematical models, and derives and displays quantitative evaluations of system reliability as a function of various mission parameters of interest to the system designer.

Mathur, F. P.

A fault-tolerant information processing concept for space vehicles.

A distributed fault-tolerant information processing system is proposed, comprising a central multiprocessor, dedicated local processors, and multiplexed input-output buses connecting them together. The processors in the multiprocessor are duplicated for error detection, which is felt to be less expensive than using coded redundancy of comparable effectiveness. Error recovery is made possible by a triplicated scratchpad memory in each processor. The main multiprocessor memory uses replicated memory for error detection and correction. Local processors use any of three conventional redundancy techniques: voting, duplex pairs with backup, and duplex pairs in independent subsystems.

Hopkins, A. L., Jr.

A simple executive for a fault-tolerant, real-time multiprocessor.

Description of a simple executive for operation with a fault-tolerant multiprocessor that is oriented toward application in an environment where the primary function is to provide real-time control. The primary executive function is to accept requests for jobs placed by other jobs or from peripheral equipment and then schedule their initiation in accordance with the request parameters. The executive is also brought into action when a processor fails, so that appropriate disposition may be made of the job that was running on the failed processor. Many architectural features intended to support this executive concept are included.

Filene, R. J.

Reliability Models and Demonstration of a Fault-Tolerant Motor Concept for Vertical Takeoff and Landing Vehicles

This report documents the completion of the Revolutionary Vertical Lift Technology Project Annual Performance Indicator 24-3.2.4.1: “Apply and document reliability prediction for high reliability motor concept.” Two modeling tools were completed for calculation of reliability of fault-tolerant (FT) motors, and key FT operations of a modular FT motor were demonstrated experimentally. The two models are complementary tools for the stakeholder and user community. Both models employ Markov chain theory. The first model is a time-homogeneous Markov chain model, and the second is a time-inhomogeneous Markov-Weibull model. This report’s main sections are as follows: 1.0 Introduction, 2.0 Theory, 3.0 Motor Reliability Models, 4.0 Validation of FT Operation by Hardware Demonstration, and 5.0 Concluding Remarks. Novel contributions to the field include development of a modular FT motor concept for electrified vertical takeoff and landing (eVTOL) application, solution methods to solve the reliability calculations, development of figures of merit, and the introduction of “linked chains” to formulate a building-block approach for time-inhomogeneous Markov-Weibull modeling of motor reliability. Example case studies have been completed, and results are provided and discussed herein. A four-module FT motor concept was developed to a preliminary-design level of detail. This eVTOL FT motor concept was designed for galvanic, magnetic, and thermal isolation of stator winding faults. The reliability of the concept motor was calculated using a time-inhomogeneous Markov chain model. Employing average failure rate as a metric, 570 times greater reliability was achieved as compared to a baseline motor without fault tolerance. A demonstrator motor was built and tested. The testing demonstrated the key features of FT operation and validated the essential premises of the FT motor concepts presented herein. The experiments included successful demonstration of the feasibility of the following four key FT features: (1) terminal open-circuit operation, (2) thermal isolation after fault, (3) terminal short-circuit operation, and (4) internal short-circuit operation. These works indicate that FT modular motor drives offer promise for addressing the daunting reliability gap that electric aircraft propulsor drives are facing relative to the best conventional motor drive technology that is available today.

Electric Motor

Reliability estimation procedures and CARE: The Computer-Aided Reliability Estimation Program

Ultrareliable fault-tolerant onboard digital systems for spacecraft intended for long mission life exploration of the outer planets are under development. The design of systems involving self-repair and fault-tolerance leads to the companion problem of quantifying and evaluating the survival probability of the system for the mission under consideration and the constraints imposed upon the system. Methods have been developed to (1) model self-repair and fault-tolerant organizations; (2) compute survival probability, mean life, and many other reliability predictive functions with respect to various systems and mission parameters; (3) perform sensitivity analysis of the system with respect to mission parameters; and (4) quantitatively compare competitive fault-tolerant systems. Various measures of comparison are offered. To automate the procedures of reliability mathematical modeling and evaluation, the CARE (computer-aided reliability estimation) program was developed. CARE is an interactive program residing on the UNIVAC 1108 system, which makes the above calculations and facilitates report preparation by providing output in tabular form, graphical 2-dimensional plots, and 3-dimensional projections. The reliability estimation of fault-tolerant organization by means of the CARE program is described.

Mathur, F. P.

Logic design for dynamic and interactive recovery.

Recovery in a fault-tolerant computer means the continuation of system operation with data integrity after an error occurs. This paper delineates two parallel concepts embodied in the hardware and software functions required for recovery; detection, diagnosis, and reconfiguration for hardware, data integrity, checkpointing, and restart for the software. The hardware relies on the recovery variable set, checking circuits, and diagnostics, and the software relies on the recovery information set, audit, and reconstruct routines, to characterize the system state and assist in recovery when required. Of particular utility is a handware unit, the recovery control unit, which serves as an interface between error detection and software recovery programs in the supervisor and provides dynamic interactive recovery.

Carter, W. C.

Braid read-only memory

Transformer-type memory is fault-tolerant array of independent read-only memory units. Information pattern in each unit is written by weaving wires through array of linear (nonswitching) transformers. Presence or absence of a bit is determined by whether a given wire threads or bypasses given transformer.

Mckenna, J. F.

A survey of an introduction to fault diagnosis algorithms

This report surveys the field of diagnosis and introduces some of the key algorithms and heuristics currently in use. Fault diagnosis is an important and a rapidly growing discipline. This is important in the design of self-repairable computers because the present diagnosis resolution of its fault-tolerant computer is limited to a functional unit or processor. Better resolution is necessary before failed units can become partially reuseable. The approach that holds the greatest promise is that of resident microdiagnostics; however, that presupposes a microprogrammable architecture for the computer being self-diagnosed. The presentation is tutorial and contains examples. An extensive bibliography of some 220 entries is included.

Mathur, F. P.

Program for computer aided reliability estimation

A computer program for estimating the reliability of self-repair and fault-tolerant systems with respect to selected system and mission parameters is presented. The computer program is capable of operation in an interactive conversational mode as well as in a batch mode and is characterized by maintenance of several general equations representative of basic redundancy schemes in an equation repository. Selected reliability functions applicable to any mathematical model formulated with the general equations, used singly or in combination with each other, are separately stored. One or more system and/or mission parameters may be designated as a variable. Data in the form of values for selected reliability functions is generated in a tabular or graphic format for each formulated model.

Mathur, F. P.