Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

The implementation and use of Ada on distributed systems with reliability requirements

The issues involved in the use of the programming language Ada on distributed systems are discussed. The effects of Ada programs on hardware failures such as loss of a processor are emphasized. It is shown that many Ada language elements are not well suited to this environment. Processor failure can easily lead to difficulties on those processors which remain. As an example, the calling task in a rendezvous may be suspended forever if the processor executing the serving task fails. A mechanism for detecting failure is proposed and changes to the Ada run time support system are suggested which avoid most of the difficulties. Ada program structures are defined which allow programs to reconfigure and continue to provide service following processor failure.

Reynolds, P. F.↗

Subsystem testing of Galileo's attitude and articulation control fault protection

This paper discusses the fault protection tesing of the Attitude and Articulation Control Subsystem (AACS) of the Galileo spacecraft. The need for an autonomous fault protection system on an interplanetary spacecraft is discussed. Galileo requirements for the detection and response of specific hardware failures is discussed along with the fault protection software design and implementation. The test beds and test methods used for fault protection testing are described. The paper concludes with a presentation of the results of this testing with an emphasis on requirement and design changes that were made as a result of these tests.

Anderson, L. L.↗

Engineering the Voyager Uranus mission

Several factors that affected the continuing expedition of Voyager 2 to Uranus are discussed. The hydrazine propellant and electrical power supplies on the spacecraft, and the radio signal strength and sun sensor and star tracker sensitivities are examined. The use of the on-board computer to steady the spacecraft and reduce image smear is described. The Voyager employs radiometric and optical data to navigate. The changes required to improve the transmission capabilities of the spacecraft and to eliminate the on-board hardware failures are analyzed.

Laeser, R. P.↗

Simulation evaluation of the control system command monitoring concept for the NASA V/STOL research aircraft (VSRA)

A control-system monitoring concept is described that has the potential of rapidly detecting computer command failures (hardware or software) in fly-by-wire control systems. The concept has been successfully tested on the NASA Vertical/Short Takeoff and Landing Research Aircraft (VSRA) in the Ames Research Center's Vertical Motion Simulator. The test was particularly stringent, since the VSRA is required to operate in a hazardous environment. The fidelity of the aircraft model used in the simulation was verified by flying both the simulated and actual aircraft in a precision hover task using specially designed targets.

Schroeder, J. A.↗

Diagnostic emulation: Implementation and user's guide

The Diagnostic Emulation Technique was developed within the System Validation Methods Branch as a part of the development of methods for the analysis of the reliability of highly reliable, fault tolerant digital avionics systems. This is a general technique which allows for the emulation of a digital hardware system. The technique is general in the sense that it is completely independent of the particular target hardware which is being emulated. Parts of the system are described and emulated at the logic or gate level, while other parts of the system are described and emulated at the functional level. This algorithm allows for the insertion of faults into the system, and for the observation of the response of the system to these faults. This allows for controlled and accelerated testing of system reaction to hardware failures in the target machine. This document describes in detail how the algorithm was implemented at NASA Langley Research Center and gives instructions for using the system.

Becher, Bernice↗

REDEX: The ranging equipment diagnostic expert system

REDEX, an advanced prototype expert system that diagnoses hardware failures in the Ranging Equipment (RE) at NASA's Ground Network tracking stations is described. REDEX will help the RE technician identify faulty circuit cards or modules that must be replaced, and thereby reduce troubleshooting time. It features a highly graphical user interface that uses color block diagrams and layout diagrams to illustrate the location of a fault. A semantic network knowledge representation technique was used to model the design structure of the RE. A catalog of generic troubleshooting rules was compiled to represent heuristics that are applied in diagnosing electronic equipment. Specific troubleshooting rules were identified to represent additional diagnostic knowledge that is unique to the RE. Over 50 generic and 250 specific troubleshooting rules have been derived. REDEX is implemented in Prolog on an IBM PC AT-compatible workstation. Block diagram graphics displays are color-coded to identify signals that have been monitored or inferred to have nominal values, signals that are out of tolerance, and circuit cards and functions that are diagnosed as faulty. A hypertext-like scheme is used to allow the user to easily navigate through the space of diagrams and tables. Over 50 graphic and tabular displays have been implemented. REDEX is currently being evaluated in a stand-alone mode using simulated RE fault scenarios. It will soon be interfaced to the RE and tested in an online environment. When completed and fielded, REDEX will be a concrete example of the application of expert systems technology to the problem of improving performance and reducing the lifecycle costs of operating NASA's communications networks in the 1990's.

Luczak, Edward C.↗

Voyager programmability - Experiences in control system adaptation

The two Voyager spacecraft were designed with reprogrammable attitude and articulation control system flight control processors to permit inflight modification of the software. These software modifications were designed to: (1) correct for hardware failures that occurred during the interstellar journey, and (2) improve spacecraft dynamic performance for better control so that, by minimizing the use of propellant, the life of the mission could be extended. It is suggested that, in the future, spacecraft be launched with additional unused memory to deal with anomalies unforeseen in the decision-making process.

Patel, Keyur↗

Hardware and software fault tolerance - A unified architectural approach

The loss of hardware fault tolerance which often arises when design diversity is used to improve the fault tolerance of computer software is considered analytically, and a unified design approach is proposed to avoid the problem. The fundamental theory of fault-tolerant (FT) architectures is reviewed; the current status of design-diversity software development is surveyed; and the FT-processor/attached-processor (FTP/AP) architecture developed by Lala et al. (1986) is described in detail and illustrated with diagrams. FTP/AP is shown to permit efficient implementation of N-version FT software while still tolerating random hardware failures with very high coverage; the reliability is found to be significantly higher than that of conventional majority-vote N-version software.

Lala, Jaynarayan H.↗

Knowledge representation and user interface concepts to support mixed-initiative diagnosis

The Remote Maintenance Monitoring System (RMMS) provides automated support for the maintenance and repair of ModComp computer systems used in the Launch Processing System (LPS) at Kennedy Space Center. RMMS supports manual and automated diagnosis of intermittent hardware failures, providing an efficient means for accessing and analyzing the data generated by catastrophic failure recovery procedures. This paper describes the design and functionality of the user interface for interactive analysis of memory dump data, relating it to the underlying declarative representation of memory dumps.

Sobelman, Beverly H.↗

REDEX - The ranging equipment diagnostic expert system

REDEX, an advanced prototype expert system that diagnoses hardware failures in the Ranging Equipment (RE) at NASA's Ground Network tracking stations is described. REDEX will help the RE technician identify faulty circuit cards or modules that must be replaced, and thereby reduce troubleshooting time. It features a highly graphical user interface that uses color block diagrams and layout diagrams to illustrate the location of a fault. A semantic network knowledge representation technique was used to model the design structure of the RE. A catalog of generic troubleshooting rules was compiled to represent heuristics that are applied in diagnosing electronic equipment. Specific troubleshooting rules were identified to represent additional diagnostic knowledge that is unique to the RE. Over 50 generic and 250 specific troubleshooting rules have been derived. REDEX is implemented in Prolog on an IBM PC AT-compatible workstation. Block diagram graphics displays are color-coded to identify signals that have been monitored or inferred to have nominal values, signals that are out of tolerance, and circuit cards and functions that are diagnosed as faulty. A hypertext-like scheme is used to allow the user to easily navigate through the space of diagrams and tables. Over 50 graphic and tabular displays have been implemented. REDEX is currently being evaluated in a stand-alone mode using simulated RE fault scenarios. It will soon be interfaced to the RE and tested in an online environment. When completed and fielded, REDEX will be a concrete example of the application of expert systems technology to the problem of improving performance and reducing the lifecycle costs of operating NASA's communications networks in the 1990s.

Luczak, Edward C.↗

Fault-tolerant multichannel demultiplexer subsystems

Fault tolerance in future processing and switching communication satellites is addressed by showing new methods for detecting hardware failures in the first major subsystem, the multichannel demultiplexer. An efficient method for demultiplexing frequency slotted channels uses multirate filter banks which contain fast Fourier transform processing. All numerical processing is performed at a lower rate commensurate with the small bandwidth of each bandbase channel. The integrity of the demultiplexing operations is protected by using real number convolutional codes to compute comparable parity values which detect errors at the data sample level. High rate, systematic convolutional codes produce parity values at a much reduced rate, and protection is achieved by generating parity values in two ways and comparing them. Parity values corresponding to each output channel are generated in parallel by a subsystem, operating even slower and in parallel with the demultiplexer that is virtually identical to the original structure. These parity calculations may be time shared with the same processing resources because they are so similar.

Redinbo, Robert↗

Evaluation of spacecraft product assurance requirements and flight performance history

The Jet Propulsion Laboratory (JPL) has undertaken a task to relate long duration, space flight hardware performance to specific product assurance requirements established during the hardware development process. This paper describes the approach that JPL is using to correlate in-flight and ground test hardware failures to critical product assurance practices implemented on a given flight project. The first step in this effort has been to collect, and convert into a convenient format, in-flight problem, failure, and anomaly data. A characterization of a subset of this anomaly data base, focusing on anomaly causes, time dependence, and subsystem affected is presented.

Gonzalez, Charles C.↗

FORTH as the basis for an integrated operations environment for a Space Shuttle scientific experiment

Over a period of three years, a FORTH-based system was developed by JPL for the operations of a major scientific instrument onboard the Space Shuttle. The software had to meet a very demanding operations environment where the interactiveness of the software was not merely desirable but essential to the success of the mission. Forth was chosen for its capability of integrating divergent software needs into an interactive package. The mission flown in October 1984 was beset with numerous hardware failures and challenged the capability of the system to its fullest.

Harris, Henry M.↗

Progressive retry for software error recovery in distributed systems

In this paper, we describe a method of execution retry for bypassing software errors based on checkpointing, rollback, message reordering and replaying. We demonstrate how rollback techniques, previously developed for transient hardware failure recovery, can also be used to recover from software faults by exploiting message reordering to bypass software errors. Our approach intentionally increases the degree of nondeterminism and the scope of rollback when a previous retry fails. Examples from our experience with telecommunications software systems illustrate the benefits of the scheme.

Wang, Yi-Min↗

High-performance reactionless scan mechanism

A high-performance reactionless scan mirror mechanism was developed for space applications to provide thermal images of the Earth. The design incorporates a unique mechanical means of providing reactionless operation that also minimizes weight, mechanical resonance operation to minimize power, combined use of a single optical encoder to sense coarse and fine angular position, and a new kinematic mount of the mirror. A flex pivot hardware failure and current project status are discussed.

Williams, Ellen I.↗

A Scheduling Algorithm for Replicated Real-Time Tasks

We present an algorithm for scheduling real-time periodic tasks on a multiprocessor system under fault-tolerant requirement. Our approach incorporates both the redundancy and masking technique and the imprecise computation model. Since the tasks in hard real-time systems have stringent timing constraints, the redundancy and masking technique are more appropriate than the rollback techniques which usually require extra time for error recovery. The imprecise computation model provides flexible functionality by trading off the quality of the result produced by a task with the amount of processing time required to produce it. It therefore permits the performance of a real-time system to degrade gracefully. We evaluate the algorithm by stochastic analysis and Monte Carlo simulations. The results show that the algorithm is resilient under hardware failures.

Yu, Albert C.↗

Logic Design Pathology and Space Flight Electronics

This paper presents a look at logic design from early in the US Space Program and examines faults in recent logic designs. Most examples are based on flight hardware failures and analysis of new tools and techniques. The paper is presented in viewgraph form.

Katz, Richard B.↗

ESTAR Measurements During SGP-99

The synthetic aperture radiometer, ESTAR, provided L-band brightness temperature maps of the experiment site during the Southern Great Plains Experiment in 1999. ESTAR flew on the NASA P-3 aircraft at an altitude of 7.6 km and mapped a swath about 50 km wide and about 300 km long. The area mapped extended west from Oklahoma City to El Reno and north from the Little Washita River watershed to the Kansas border. The flight lines and mapping were done in a manner similar to that used in 1997 during SGP-97. ESTAR is a prototype designed to develop the technology of aperture synthesis for passive microwave remote sensing. It is actually a hybrid that uses real aperture to obtain resolution along track and synthetic aperture to obtain resolution across track. ESTAR operates in the band at 1.413 GHz set aside for passive use and images at horizontal polarization in the equivalent of a cross-track scan. The scan is done in software as part of image reconstruction. Calibration consists of blackbody observed before and after each flight and a local body of water (Lake Kaw). ESTAR arrived in Oklahoma on July 7, 1999 and flew mapping missions on July 8,9, 11, 14, 15, 19 and 20. The aircraft was down on July 10 by design and was down again on July 12-13 and 16-18 due to mechanical problems. ESTAR data for July I I is questionable because of a hardware failure. The brightness temperature maps reflect the patterns of soil moisture observed in the SGP99 study area and compare well with measurements made by ESTAR in other experiments in this region (e.g. SGP -97 and Washita-92).

LeVine, D. M.↗