Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “transient faults”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Branch recovery with compiler-assisted multiple instruction retry

In processing systems where rapid recovery from transient faults is important, schemes for multiple instruction rollback recovery may be appropriate. Multiple instruction retry has been implemented in hardware by researchers and also in mainframe computers. This paper extends compiler-assisted instruction retry to a broad class of code execution failures. Five benchmarks were used to measure the performance penalty of hazard resolution. Results indicate that the enhanced pure software approach can produce performance penalties consistent with existing hardware techniques. A combined compiler/hardware resolution strategy is also described and evaluated. Experimental results indicate a lower performance penalty than with either a totally hardware or totally software approach.

Alewine, N. J.↗

Provable Transient Recovery for Frame-Based, Fault-Tolerant Computing Systems

We present a formal verification of the transient fault recovery aspects of the Reliable Computing Platform (RCP), a fault-tolerant computing system architecture for digital flight control applications. The RCP uses NMR-style redundancy to mask faults and internal majority voting to purge the effects of transient faults. The system design has been formally specified and verified using the EHDM verification system. Our formalization accommodates a wide variety of voting schemes for purging the effects of transients.

DiVito, Ben L.↗

Secured fault detection in a power substation

Systems and methods for fault detection and protection in electric power systems that evaluates electromagnetic transients caused by faults. A fault can be detected using sampled data from a first monitored point in the power system. Detection of fault transients and associated characteristics, including transient direction, can also be extracted through evaluation of sample data from other monitored points in the power system. A monitoring device can evaluate whether to trip a switching device in response to the detection of the fault and based on confirmation of an indication of detection of fault transients at the other monitored points of the power system. The determination of whether to trip or activate the switching device can also be based on other factors, including the timing of receipt of an indication of the detection of the fault transients and/or an evaluation of the characteristics of the detected transients.

Cui, Tao↗

Formal specification and verification of a fault-masking and transient-recovery model for digital flight-control systems

The formal specification and mechanically checked verification for a model of fault-masking and transient-recovery among the replicated computers of digital flight-control systems are presented. The verification establishes, subject to certain carefully stated assumptions, that faults among the component computers are masked so that commands sent to the actuators are the same as those that would be sent by a single computer that suffers no failures.

Rushby, John↗

Control and protection system for paralleled modular static inverter-converter systems

A control and protection system was developed for use with a paralleled 2.5-kWe-per-module static inverter-converter system. The control and protection system senses internal and external fault parameters such as voltage, frequency, current, and paralleling current unbalance. A logic system controls contactors to isolate defective power conditioners or loads. The system sequences contactor operation to automatically control parallel operation, startup, and fault isolation. Transient overload protection and fault checking sequences are included. The operation and performance of a control and protection system, with detailed circuit descriptions, are presented.

Birchenough, A. G.↗

Comparative Study of Nonlinear Black-Box Modeling for Power Electronics Converters

With the increasing penetration level of renewable sources and power electronics loads in modern power systems, accurate and computationally efficient models are needed. Black-box model (BBM) could be a useful method in such systems. However, not very extensive research efforts have been made for power electronics BBM so far, and existing works mostly focus on steady-state operation, neglecting the important transient behaviors such as load transients, voltage transients, and faults. This paper presents a comparative study of three commonly used nonlinear BBM approaches for transient behaviors of power electronics converters. Comparison methods are proposed, and the evaluations are conducted under different transients using a grid-connected single-phase photovoltaic inverter. The findings of this study provide valuable references for further feasibility investigations on implementing BBMs in large-scale power electronics-rich power systems.

Qiao, Liang↗

Towards Precision-Aware Fault Tolerance Approaches for Mixed-Precision Applications

Graphics Processing Units (GPUs), the dominantly adopted accelerators in HPC systems, are susceptible to transient hardware fault. New generation of GPUs feature mixed-precision architectures such as NVIDIA Tensor Cores to accelerate matrix multiplications. While being widely adapted, how would they behave under transient hardware faults remain unclear. In this study, we conduct a large-scale fault injection experiments on GEMM kernels implemented with different floating-point data types on the V100 and A100 Tensor Cores, and show distinct error resilience characteristics for the GEMMS with different formats. In the future, we plan to explore this space by building precision-aware floating-point fault tolerance techniques for applications such as DNNs that exercise low-precision computations.

Fang, Bo↗

Fault characterization of a multilayered perceptron network

The results of a set of simulation experiments conducted to quantify the effects of faults in a classification network implemented as a three-layered perception model are reported. The percentage of vectors misclassified by the classification network, the time taken for the network to stabilize, and the output values are measured. The results show that both transient and permanent faults have a significant impact on the performance of the network. Transient faults are also found to cause the network to be increasingly unstable as the duration of a transient is increased. The average percentage of the vectors misclassified is about 25 percent; after relearning, this is reduced to 10 percent. The impact of link faults is relatively insignificant in comparison with node faults (1 percent versus 19 percent misclassified after relearning). A study of the impact of hardware redundancy shows a linear increase in misclassifications with increasing hardware size.

Tan, Chang H.↗

Experimental fault characterization of a neural network

The effects of a variety of faults on a neural network is quantified via simulation. The neural network consists of a single-layered clustering network and a three-layered classification network. The percentage of vectors mistagged by the clustering network, the percentage of vectors misclassified by the classification network, the time taken for the network to stabilize, and the output values are all measured. The results show that both transient and permanent faults have a significant impact on the performance of the measured network. The corresponding mistag and misclassification percentages are typically within 5 to 10 percent of each other. The average mistag percentage and the average misclassification percentage are both about 25 percent. After relearning, the percentage of misclassifications is reduced to 9 percent. In addition, transient faults are found to cause the network to be increasingly unstable as the duration of a transient is increased. The impact of link faults is relatively insignificant in comparison with node faults (1 versus 19 percent misclassified after relearning). There is a linear increase in the mistag and misclassification percentages with decreasing hardware redundancy. In addition, the mistag and misclassification percentages linearly decrease with increasing network size.

Tan, Chang-Huong↗

Postseismic Transient after the 2002 Denali Fault Earthquake from VLBI Measurements at Fairbanks

The VLBI antenna (GILCREEK) at Fairbanks, Alaska observes in networks routinely twice a week with operational networks and on additional days with other networks on a more uneven basis. The Fairbanks antenna position is about 150 km north of the Denali fault and from the earthquake epicenter. We examine the transient behavior of the estimated VLBI position during the year following the earthquake to determine how the rate of change of postseismic deformation has changed. This is compared with what is seen in the GPS site position series.

MacMillan, Daniel↗

PARIS: Predicting application resilience using machine learning

The traditional method to study application resilience to errors in HPC applications uses fault injection (FI), a time-consuming approach. Furthermore, while analytical models have been built to overcome the inefficiencies of FI, they lack accuracy. In this paper, we present PARIS, a machine-learning method to predict application resilience that avoids the time-consuming process of random FI and provides higher prediction accuracy than analytical models. PARIS captures the implicit relationship between application characteristics and application resilience, which is difficult to capture using most analytical models. We overcome many technical challenges for feature construction, extraction, and selection to use machine learning in our prediction approach. Our evaluation on 16 HPC benchmarks shows that PARIS achieves high prediction accuracy. PARIS is up to 450x faster than random FI (49x on average). Compared to the state-of-the-art analytical model, PARIS is at least 63% better in terms of accuracy and has comparable execution time on average.

97 MATHEMATICS AND COMPUTING↗

R2U2 in Space: System and Software Health Management for Small Satellites

In order for small but complex systems like rovers, SmallSats, or Unmanned Aircraft (UAS) to operate autonomously, they must have a real-time solution for assessing their own system health. System and Software Health Management (SHM) enables better detection of faulty sensors and software problems, and enables better fault management including mitigation of unpredicted fault scenarios in the absence of a human on-board. In recent work, we have developed a Responsive, Realizable, Unobtrusive Unit (R2U2) for on-board SHM of autonomous UAS and demonstrated its ability to detect faults during flight time. These faults, from sensor failures, to software problems, to malicious security attacks, can present as transient temporal faults that even humans are challenged to find. An R2U2 congfiuration is a modular combination of multiple types of temporal logic runtime observers with fault-specic Bayesian Nets and sensor filters. R2U2 reasons about both on-board hardware and software components; R2U2 itself can be instantiated as an independent FPGA (Field-Programmable Gate Array)-based conguration or as a software component running independently from other software on-board. Small satellites, such as CubeSats, also require on-board SHM and failure mitigation, as limited telemetry bandwidth does not allow the transmission of the entire system state for ground-based health management. However, the autonomous operation of satellites brings a set of challenges different from UAS, including the effects of radiation on non-rad-hard, low-cost components, and the harsher environment of space. We surmise that a new extension of R2U2 could be adapted to help better detect, for example, radiation errors in cheaper COTS (Commercial Off the Shelf) (not rad-hard) components often used in small space systems. Since small satellites often operate in coordination, we will also examine new ways of distributed monitoring of their communication and cooperation and real-time detection of off-nominal situations utilizing multiple satellites. This talk will discuss preliminary work and ideas for building on terrestrial success of system and software health management for the harsher, and differently challenging, environment of space.

Runtime Verification & Validation↗

Fault-tolerance experiments with the JPL STAR computer.

Results of fault-tolerance experiments performed using an experimental computer with dynamic (standby) redundancy, including replaceable subsystems and a 'program rollback' provision to eliminate transient-caused errors. After a brief review of the specification of fault-tolerance with respect to transient faults, including a description of the method of injection of transient faults in software and system tests, fault-tolerance experiments carried out with this computer with regard to the determination of fault classes, software verification, system verification, and recovery stability are summarized. A test and repair processor is described which constitutes a special monitor unit of the computer and is used to obtain information for fault detection in the other subsystems of the computer and to ensure that proper recovery occurs when a fault is detected.

Avizienis, A.↗

Nuclear Thermal Rocket Emulator for a Hardware-in-the-Loop Test Bed

To support NASA’s mission to use nuclear thermal rockets for future Mars missions, an instrumentation and control test bed has been built at Oak Ridge National Laboratory. The system is designed as a hardware-in-the-loop test bed for testing control elements and autonomous control algorithms for nuclear thermal propulsion rockets. The mock reactor system consists of a modular and scalable framework, using inexpensive components and open-source software. The hardware system consists of a two-phase flow loop and a mock reactor with six control drums. A single-board computer (NVIDIA Jetson) handles reactor core emulation and hosts a message queuing telemetry transport broker that allows user-deployed control algorithms to interact with the system hardware. The reactor emulator receives sensor data from the hardware and provides the simulated performance of the reactor under steady-state, transient, and fault conditions. The emulator uses a reactivity lookup table and the point kinetics equations to solve for the reactor dynamics in real time. Emulated reactor dynamics and sensor input inform the autonomous control algorithm’s decision-making in a closed-loop manner. The current system is capable of operating at 10 Hz, but faster cycle rates are an area of ongoing research. This test bed will enable NASA and other space vendors to rigorously test their autonomous control systems for NTP rockets under transient (reactor startup and shutdown), steady-state, and fault conditions to reduce development time and risk for autonomous control systems in future missions.

autonomous control↗