Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “fault injection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Fault Injection for TensorFlow Applications

As machine learning (ML) has seen increasing adoption in safety-critical domains (e.g., autonomous vehicles), the reliability of ML systems has also grown in importance. While prior studies have proposed techniques to enable efficient error-resilience (e.g., selective instruction duplication), a fundamental requirement for realizing these techniques is a detailed understanding of the application’s resilience. In this work, we present TensorFI 1 and TensorFI 2, high-level fault injection (FI) frameworks for TensorFlow-based applications. TensorFI 1 and 2 are able to inject both hardware and software faults in any general TensorFlow 1 and 2 program respectively. Both are configurable FI tools that are flexible, easy to use, and portable. They can be integrated into existing TensorFlow programs to assess their resilience for different fault types (e.g., bit-flips in particular operations or layers). We use the TensorFI 1 and TensorFI 2 to evaluate the resilience of 12 and 10 ML programs written in TensorFlow, including DNNs used in the autonomous vehicle domain. The results give us insights into why some of the models are more resilient. We also measure the performance overheads of the two injectors, and present 4 case studies, two for each tool, to demonstrate their utility.

97 MATHEMATICS AND COMPUTING↗

Measuring fault tolerance with the FTAPE fault injection tool

This paper describes FTAPE (Fault Tolerance And Performance Evaluator), a tool that can be used to compare fault-tolerant computers. The major parts of the tool include a system-wide fault-injector, a workload generator, and a workload activity measurement tool. The workload creates high stress conditions on the machine. Using stress-based injection, the fault injector is able to utilize knowledge of the workload activity to ensure a high level of fault propagation. The errors/fault ratio, performance degradation, and number of system crashes are presented as measures of fault tolerance.

Tsai, Timothy K.↗

A study of fault injection in multichannel spacecraft power systems

NASA/Marshall Space Flight Center proposes to implement fault injection into an electrical power system breadboard to study the reactions of the various control elements of this breadboard. Among the elements to be studied are the remote power controllers, the algorithms in the control computers, and the artificially intelligent control programs resident in this breadboard. To this end, a study of electrical power is being performed to yield a list of the most common power system faults. The results of this study are being applied to a multichannel high-voltage DC spacecraft power system called the Large Autonomous Spacecraft Electrical Power System Breadboard. Some of the reactions of the breadboard to some of the faults which have been encountered are presented along with the results of this study.

Dugal-Whitehead, Norma R.↗

Fault Injection Techniques and Tools

Dependability evaluation involves the study of failures and errors. The destructive nature of a crash and long error latency make it difficult to identify the causes of failures in the operational environment. It is particularly hard to recreate a failure scenario for a large, complex system. To identify and understand potential failures, we use an experiment-based approach for studying the dependability of a system. Such an approach is applied not only during the conception and design phases, but also during the prototype and operational phases. To take an experiment-based approach, we must first understand a system's architecture, structure, and behavior. Specifically, we need to know its tolerance for faults and failures, including its built-in detection and recovery mechanisms, and we need specific instruments and tools to inject faults, create failures or errors, and monitor their effects.

Hsueh, Mei-Chen↗

Experimental analysis of computer system dependability

This paper reviews an area which has evolved over the past 15 years: experimental analysis of computer system dependability. Methodologies and advances are discussed for three basic approaches used in the area: simulated fault injection, physical fault injection, and measurement-based analysis. The three approaches are suited, respectively, to dependability evaluation in the three phases of a system's life: design phase, prototype phase, and operational phase. Before the discussion of these phases, several statistical techniques used in the area are introduced. For each phase, a classification of research methods or study topics is outlined, followed by discussion of these methods or topics as well as representative studies. The statistical techniques introduced include the estimation of parameters and confidence intervals, probability distribution characterization, and several multivariate analysis methods. Importance sampling, a statistical technique used to accelerate Monte Carlo simulation, is also introduced. The discussion of simulated fault injection covers electrical-level, logic-level, and function-level fault injection methods as well as representative simulation environments such as FOCUS and DEPEND. The discussion of physical fault injection covers hardware, software, and radiation fault injection methods as well as several software and hybrid tools including FIAT, FERARI, HYBRID, and FINE. The discussion of measurement-based analysis covers measurement and data processing techniques, basic error characterization, dependency analysis, Markov reward modeling, software-dependability, and fault diagnosis. The discussion involves several important issues studies in the area, including fault models, fast simulation techniques, workload/failure dependency, correlated failures, and software fault tolerance.

Iyer, Ravishankar, K.↗

Understanding Mixed Precision GEMM with MPGemmFI: Insights into Fault Resilience

Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). Thus, one of the critical features of machine-learning-specific accelerators such as NVIDIA Tensor Cores, AMD Matrix Cores, and Google TPUs is the support of mixed-precision enabled GEMM. For DNN models, lower-precision FP data formats and computation offer acceptable correctness but significant performance, area, and memory footprint improvement. While promising, the mixed-precision computation on error resilience remains unexplored. To this end, we develop a fault injection framework that systematically injects fault into the mixed-precision computation results. We investigate how the faults affect the accuracy of machine learning applications. Based on the characteristics of error resilience, we offer lightweight error detection and correction solutions that significantly improve the overall model accuracy by 75% if the models experience hardware faults. The solutions can be efficiently integrated into the accelerator's pipelines.

Fang, Bo↗

Control Mechanisms for Self‐Sealing in Activated Clay‐Rich Faults Through Controlled Hydraulic Injection Experiment

Abstract In a high‐pressure injection fault activation experiment conducted at the Mont Terri underground research laboratory in Switzerland, the transmissivity of the Opalinus Clay fault significantly increased due to opening and shearing. The fluid injection, spanning a few hours, generated a 10 m radius fault activation patch. Subsequent pressure pulse tests conducted bi‐weekly for a year revealed the gradual return of fault transmissivity to its initial state. The study utilized fluid pressure decay analysis, optical fiber monitoring, continuous active source seismic measurements and borehole displacement sensors for measuring fault displacements. The fault zone exhibited a dilation of approximately 1.4 mm, associated with both normal and tangential movements during activation, resulting in a sudden transmissivity increase from 1 × 10 −12 to 3.2 × 10 −7 m 2 /s. Early post‐activation, transient compaction and the subsequent slow compaction were observed, transitioning to an extension regime. The pressure pulse tests demonstrated a rapid transmissivity drop by more than two orders of magnitude within the first 10 days, followed by a gradual and less pronounced decrease. Plastic shear and compaction dominated the transmissivity evolution until 70 days after injection ended, followed by a period where additional factors, such as clay mineral swelling, influenced the behavior. Extrapolation suggested a sealing process taking at least 50 years after the initial activation. Plain Language Summary A field‐scale fault activation experiment offers valuable insights into the elasto‐plastic processes governing the sealing of shale faults. The experiment reveals a rapid increase in the fault's transmissivity by approximately five orders of magnitude during activation. Subsequent observations show a gradual transmissivity decrease by about three orders of magnitude post‐activation, with slow long‐term plastic shear and compaction of the fault competing against secondary processes, notably clay mineral swelling. All conceptual models employed to interpret these field data converge on the estimation that the fault's return to its initial low transmissivity state would require a minimum of 50 years. Key Points High‐pressure injection fault activation experiment at the Mont Terri underground research laboratory Continuous transmissivity measurements record self‐sealing inside a clay‐rich fault zone Transmissivity undergoes a phase of domination by slow plastic compaction and shearing during the initial post‐activation period, with mineral swelling exerting its influence over the long term

Guglielmi, Yves↗

Fault recovery characteristics of the fault tolerant multi-processor

The fault handling performance of the fault tolerant multiprocessor (FTMP) was investigated. Fault handling errors detected during fault injection experiments were characterized. In these fault injection experiments, the FTMP disabled a working unit instead of the faulted unit once every 500 faults, on the average. System design weaknesses allow active faults to exercise a part of the fault management software that handles byzantine or lying faults. It is pointed out that these weak areas in the FTMP's design increase the probability that, for any hardware fault, a good LRU (line replaceable unit) is mistakenly disabled by the fault management software. It is concluded that fault injection can help detect and analyze the behavior of a system in the ultra-reliable regime. Although fault injection testing cannot be exhaustive, it has been demonstrated that it provides a unique capability to unmask problems and to characterize the behavior of a fault-tolerant system.

Padilla, Peter A.↗

ARETE: Accurate Error Assessment via Machine Learning-Guided Dynamic-Timing Analysis

Nanometer circuits are increasingly prone to timing errors, escalating the need for fault injection frameworks to accurately evaluate their impact on applications. Here in this paper, we propose ARETE, a novel cross-layer, fault-injection framework that combines dynamic-binary instrumentation with machine learning-guided dynamic-timing analysis. ARETE enables accurate fault-injection into any application by estimating the location of the injecting errors via dynamic-timing analysis. To accelerate fault-injection, we develop a novel, data-aware, machine learning-based mechanism that dynamically pre-selects the error-prone instructions and limits the application of the costly dynamic-timing analysis only to them. To evaluate ARETE's accuracy, our fully automated toolflow is configured to support fault-injection based on detailed post-layout gate-level simulations as well as via existing workload-agnostic error models. Our results for various workloads, including an autonomous-driving library, show that the location and time of injected errors performed by ARETE, is 89.9% consistent with fault-injection based on full gate-level simulation. On average, ARETE executes 84.6x faster than gate-level simulation and at a cost of 3.4% loss in the program output quality estimation. When compared to the existing statistical fault-injection tools that are based on workload-agnostic error models, ARETE improves the accuracy of fault-injection rate and output quality estimation by 143.9% and 40.4% on average, respectively.

97 MATHEMATICS AND COMPUTING↗

Board Level Proton Testing Book of Knowledge for NASA Electronic Parts and Packaging Program

This book of knowledge (BoK) provides a critical review of the benefits and difficulties associated with using proton irradiation as a means of exploring the radiation hardness of commercial-off-the-shelf (COTS) systems. This work was developed for the NASA Electronic Parts and Packaging (NEPP) Board Level Testing for the COTS task. The fundamental findings of this BoK are the following. The board-level test method can reduce the worst case estimate for a board's single-event effect (SEE) sensitivity compared to the case of no test data, but only by a factor of ten. The estimated worst case rate of failure for untested boards is about 0.1 SEE/board-day. By employing the use of protons with energies near or above 200 MeV, this rate can be safely reduced to 0.01 SEE/board-day, with only those SEEs with deep charge collection mechanisms rising this high. For general SEEs, such as static random-access memory (SRAM) upsets, single-event transients (SETs), single-event gate ruptures (SEGRs), and similar cases where the relevant charge collection depth is less than 10 μm, the worst case rate for SEE is below 0.001 SEE/board-day. Note that these bounds assume that no SEEs are observed during testing. When SEEs are observed during testing, the board-level test method can establish a reliable event rate in some orbits, though all established rates will be at or above 0.001 SEE/board-day. The board-level test approach we explore has picked up support as a radiation hardness assurance technique over the last twenty years. The approach originally was used to provide a very limited verification of the suitability of low cost assemblies to be used in the very benign environment of the International Space Station (ISS), in limited reliability applications. Recently the method has been gaining popularity as a way to establish a minimum level of SEE performance of systems that require somewhat higher reliability performance than previous applications. This sort of application of the method suggests a critical analysis of the method is in order. This is also of current consideration because the primary facility used for this type of work, the Indiana University Cyclotron Facility (IUCF) (also known as the Integrated Science and Technology (ISAT) hall), has closed permanently, and the future selection of alternate test facilities is critically important. This document reviews the main theoretical work on proton testing of assemblies over the last twenty years. It augments this with review of reported data generated from the method and other data that applies to the limitations of the proton board-level test approach. When protons are incident on a system for test they can produce spallation reactions. From these reactions, secondary particles with linear energy transfers (LETs) significantly higher than the incident protons can be produced. These secondary particles, together with the protons, can simulate a subset of the space environment for particles capable of inducing single event effects (SEEs). The proton board-level test approach has been used to bound SEE rates, establishing a maximum possible SEE rate that a test article may exhibit in space. This bound is not particularly useful in many cases because the bound is quite loose. We discuss the established limit that the proton board-level test approach leaves us with. The remaining possible SEE rates may be as high as one per ten years for most devices. The situation is actually more problematic for many SEE types with deep charge collection. In cases with these SEEs, the limits set by the proton board-level test can be on the order of one per 100 days. Because of the limited nature of the bounds established by proton testing alone, it is possible that tested devices will have actual SEE sensitivity that is very low (e.g., fewer than one event in 1 × 10(exp 4) years), but the test method will only be able to establish the limits indicated above. This BoK further examines other benefits of proton board-level testing besides hardness assurance. The primary alternate use is the injection of errors. Error injection, or fault injection, is something that is often done in a simulation environment. But the proton beam has the benefit of injecting the majority of actual SEEs without risk of something being missed, and without the risk of simulation artifacts misleading the SEE investigation.

Guertin, Steven M.↗

Characterizing Impacts of Storage Faults on HPC Applications: A methodology and insights

In recent years, the increasing complexity in scientific simulations and emerging demands for training heavy artificial intelligence models require massive and fast data accesses, which urges high-performance computing (HPC) platforms to equip with more advanced storage infrastructures such as solid-state disks (SSDs). While SSDs offer high-performance I/O, it remains unclear about the reliability challenges faced by the HPC applications under the SSD-related failures, in particular, failures resulting in data corruptions. The goal of this paper is to understand the impact of SSD-related data corruptions on the behaviors of complex HPC applications. To this end, we propose FFIS, a FUSE-based fault injection framework that systematically introduces storage faults into the application layer to model the errors originated from SSDs. FFIS is able to plant different I/O related faults into the data returned from underlying file systems, which also enables the investigation on the error resilience characteristics of the scientific file format for the first time. We demonstrate the use of FFIS with three representative real HPC applications, show how each application reacts to the data corruptions, and provide insights on the error resilience of the widely-adopted HDF5 file format for the HPC applications.

Fang, Bo↗

Estimating the distribution of fault latency in a digital processor

Presented is a statistical approach to measuring fault latency in a digital processor. The method relies on the use of physical fault injection where the duration of the fault injection can be controlled. Although a specific fault's latency period is never directly measured, the method indirectly determines the distribution of fault latency.

Ellis, Erik L.↗

In-circuit fault injector user's guide

A fault injector system, called an in-circuit injector, was designed and developed to facilitate fault injection experiments performed at NASA-Langley's Avionics Integration Research Lab (AIRLAB). The in-circuit fault injector (ICFI) allows fault injections to be performed on electronic systems without special test features, e.g., sockets. The system supports stuck-at-zero, stuck-at-one, and transient fault models. The ICFI system is interfaced to a VAX-11/750 minicomputer. An interface program has been developed in the VAX. The computer code required to access the interface program is presented. Also presented is the connection procedure to be followed to connect the ICFI system to a circuit under test and the ICFI front panel controls which allow manual control of fault injections.

Padilla, Peter A.↗

Experimental evaluation of the certification-trail method

Certification trails are a recently introduced and promising approach to fault-detection and fault-tolerance. A comprehensive attempt to assess experimentally the performance and overall value of the method is reported. The method is applied to algorithms for the following problems: huffman tree, shortest path, minimum spanning tree, sorting, and convex hull. Our results reveal many cases in which an approach using certification-trails allows for significantly faster overall program execution time than a basic time redundancy-approach. Algorithms for the answer-validation problem for abstract data types were also examined. This kind of problem provides a basis for applying the certification-trail method to wide classes of algorithms. Answer-validation solutions for two types of priority queues were implemented and analyzed. In both cases, the algorithm which performs answer-validation is substantially faster than the original algorithm for computing the answer. Next, a probabilistic model and analysis which enables comparison between the certification-trail method and the time-redundancy approach were presented. The analysis reveals some substantial and sometimes surprising advantages for ther certification-trail method. Finally, the work our group performed on the design and implementation of fault injection testbeds for experimental analysis of the certification trail technique is discussed. This work employs two distinct methodologies, software fault injection (modification of instruction, data, and stack segments of programs on a Sun Sparcstation ELC and on an IBM 386 PC) and hardware fault injection (control, address, and data lines of a Motorola MC68000-based target system pulsed at logical zero/one values). Our results indicate the viability of the certification trail technique. It is also believed that the tools developed provide a solid base for additional exploration.

Sullivan, Gregory F.↗

Modeling injection-induced fault slip using long short-term memory networks

Stress changes due to changes in fluid pressure and temperature in a faulted formation may lead to the opening/shearing of the fault. This can be due to subsurface (geo)engineering activities such as fluid injections and geologic disposal of nuclear waste. Such activities are expected to rise in the future making it necessary to assess their short- and long-term safety. Here, a new machine learning (ML) approach to model pore pressure and fault displacements in response to high-pressure fluid injection cycles is developed. The focus is on fault behavior near the injection borehole. To capture the temporal dependencies in the data, long short-term memory (LSTM) networks are utilized. To prevent error accumulation within the forecast window, four critical measures to train a robust LSTM model for predicting fault response are highlighted: (i) setting an appropriate value of LSTM lag, (ii) calibrating the LSTM cell dimension, (iii) learning rate reduction during weight optimization, and (iv) not adopting an independent injection cycle as a validation set. Several numerical experiments were conducted, which demonstrated that the ML model can capture peaks in pressure and associated fault displacement that accompany an increase in fluid injection. The model also captured the decay in pressure and displacement during the injection shut-in period. Further, the ability of an ML model to highlight key changes in fault hydromechanical activation processes was investigated, which shows that ML can be used to monitor risk of fault activation and leakage during high pressure fluid injections.

58 GEOSCIENCES↗

Fail-Safe Logic Design Strategies Within Modern FPGA Architectures

Fail-safe computing refers to computing systems that revert to a non-operational safe state when a fault occurs. In this paper, we investigate a circuit level technique as mitigation for single event upsets (SEUs) and fault injection attacks on field programmable gate arrays (FPGAs), and analyze the effectiveness of the technique as a fail-safe monitor for an encryption algorithm. The propagation of fault effects through FPGA primitives including lookup tables (LUTs) and programmable interconnect points (PIPs) is assessed within an FPGA architecture created using an open source tool, and validated using fault injection experiments on an FPGA. The analysis reveals additional vulnerabilities exist within reconfigurable architectures over those in equivalent fail-safe application specific integrated circuit (ASIC), thus requiring a more elaborate network of redundant circuits and checking logic. The configuration memory bits (CMBs), which configure routing and designate logic functions within the LUTs of the FPGA, add complexity to fail-safe design strategies by introducing additional fault conditions and fault propagation paths. A resource-efficient fail-safe circuit design technique called DEsign for Fail-safe in reCONfigurable systems (DEFCON) is proposed. The benefits and limitations associated with DEFCON are described in the context of fault injection experiments carried out as simulations and in FPGA hardware.

Bhakta, Priya A. [Univ. of New Mexico, Albuquerque↗