Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “fault injection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Fault Injection for TensorFlow Applications

As machine learning (ML) has seen increasing adoption in safety-critical domains (e.g., autonomous vehicles), the reliability of ML systems has also grown in importance. While prior studies have proposed techniques to enable efficient error-resilience (e.g., selective instruction duplication), a fundamental requirement for realizing these techniques is a detailed understanding of the application’s resilience. In this work, we present TensorFI 1 and TensorFI 2, high-level fault injection (FI) frameworks for TensorFlow-based applications. TensorFI 1 and 2 are able to inject both hardware and software faults in any general TensorFlow 1 and 2 program respectively. Both are configurable FI tools that are flexible, easy to use, and portable. They can be integrated into existing TensorFlow programs to assess their resilience for different fault types (e.g., bit-flips in particular operations or layers). We use the TensorFI 1 and TensorFI 2 to evaluate the resilience of 12 and 10 ML programs written in TensorFlow, including DNNs used in the autonomous vehicle domain. The results give us insights into why some of the models are more resilient. We also measure the performance overheads of the two injectors, and present 4 case studies, two for each tool, to demonstrate their utility.

97 MATHEMATICS AND COMPUTING↗

Understanding Mixed Precision GEMM with MPGemmFI: Insights into Fault Resilience

Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). Thus, one of the critical features of machine-learning-specific accelerators such as NVIDIA Tensor Cores, AMD Matrix Cores, and Google TPUs is the support of mixed-precision enabled GEMM. For DNN models, lower-precision FP data formats and computation offer acceptable correctness but significant performance, area, and memory footprint improvement. While promising, the mixed-precision computation on error resilience remains unexplored. To this end, we develop a fault injection framework that systematically injects fault into the mixed-precision computation results. We investigate how the faults affect the accuracy of machine learning applications. Based on the characteristics of error resilience, we offer lightweight error detection and correction solutions that significantly improve the overall model accuracy by 75% if the models experience hardware faults. The solutions can be efficiently integrated into the accelerator's pipelines.

Fang, Bo↗

Control Mechanisms for Self‐Sealing in Activated Clay‐Rich Faults Through Controlled Hydraulic Injection Experiment

Abstract In a high‐pressure injection fault activation experiment conducted at the Mont Terri underground research laboratory in Switzerland, the transmissivity of the Opalinus Clay fault significantly increased due to opening and shearing. The fluid injection, spanning a few hours, generated a 10 m radius fault activation patch. Subsequent pressure pulse tests conducted bi‐weekly for a year revealed the gradual return of fault transmissivity to its initial state. The study utilized fluid pressure decay analysis, optical fiber monitoring, continuous active source seismic measurements and borehole displacement sensors for measuring fault displacements. The fault zone exhibited a dilation of approximately 1.4 mm, associated with both normal and tangential movements during activation, resulting in a sudden transmissivity increase from 1 × 10 −12 to 3.2 × 10 −7 m 2 /s. Early post‐activation, transient compaction and the subsequent slow compaction were observed, transitioning to an extension regime. The pressure pulse tests demonstrated a rapid transmissivity drop by more than two orders of magnitude within the first 10 days, followed by a gradual and less pronounced decrease. Plastic shear and compaction dominated the transmissivity evolution until 70 days after injection ended, followed by a period where additional factors, such as clay mineral swelling, influenced the behavior. Extrapolation suggested a sealing process taking at least 50 years after the initial activation. Plain Language Summary A field‐scale fault activation experiment offers valuable insights into the elasto‐plastic processes governing the sealing of shale faults. The experiment reveals a rapid increase in the fault's transmissivity by approximately five orders of magnitude during activation. Subsequent observations show a gradual transmissivity decrease by about three orders of magnitude post‐activation, with slow long‐term plastic shear and compaction of the fault competing against secondary processes, notably clay mineral swelling. All conceptual models employed to interpret these field data converge on the estimation that the fault's return to its initial low transmissivity state would require a minimum of 50 years. Key Points High‐pressure injection fault activation experiment at the Mont Terri underground research laboratory Continuous transmissivity measurements record self‐sealing inside a clay‐rich fault zone Transmissivity undergoes a phase of domination by slow plastic compaction and shearing during the initial post‐activation period, with mineral swelling exerting its influence over the long term

Guglielmi, Yves↗

ARETE: Accurate Error Assessment via Machine Learning-Guided Dynamic-Timing Analysis

Nanometer circuits are increasingly prone to timing errors, escalating the need for fault injection frameworks to accurately evaluate their impact on applications. Here in this paper, we propose ARETE, a novel cross-layer, fault-injection framework that combines dynamic-binary instrumentation with machine learning-guided dynamic-timing analysis. ARETE enables accurate fault-injection into any application by estimating the location of the injecting errors via dynamic-timing analysis. To accelerate fault-injection, we develop a novel, data-aware, machine learning-based mechanism that dynamically pre-selects the error-prone instructions and limits the application of the costly dynamic-timing analysis only to them. To evaluate ARETE's accuracy, our fully automated toolflow is configured to support fault-injection based on detailed post-layout gate-level simulations as well as via existing workload-agnostic error models. Our results for various workloads, including an autonomous-driving library, show that the location and time of injected errors performed by ARETE, is 89.9% consistent with fault-injection based on full gate-level simulation. On average, ARETE executes 84.6x faster than gate-level simulation and at a cost of 3.4% loss in the program output quality estimation. When compared to the existing statistical fault-injection tools that are based on workload-agnostic error models, ARETE improves the accuracy of fault-injection rate and output quality estimation by 143.9% and 40.4% on average, respectively.

97 MATHEMATICS AND COMPUTING↗

Characterizing Impacts of Storage Faults on HPC Applications: A methodology and insights

In recent years, the increasing complexity in scientific simulations and emerging demands for training heavy artificial intelligence models require massive and fast data accesses, which urges high-performance computing (HPC) platforms to equip with more advanced storage infrastructures such as solid-state disks (SSDs). While SSDs offer high-performance I/O, it remains unclear about the reliability challenges faced by the HPC applications under the SSD-related failures, in particular, failures resulting in data corruptions. The goal of this paper is to understand the impact of SSD-related data corruptions on the behaviors of complex HPC applications. To this end, we propose FFIS, a FUSE-based fault injection framework that systematically introduces storage faults into the application layer to model the errors originated from SSDs. FFIS is able to plant different I/O related faults into the data returned from underlying file systems, which also enables the investigation on the error resilience characteristics of the scientific file format for the first time. We demonstrate the use of FFIS with three representative real HPC applications, show how each application reacts to the data corruptions, and provide insights on the error resilience of the widely-adopted HDF5 file format for the HPC applications.

Fang, Bo↗

Modeling injection-induced fault slip using long short-term memory networks

Stress changes due to changes in fluid pressure and temperature in a faulted formation may lead to the opening/shearing of the fault. This can be due to subsurface (geo)engineering activities such as fluid injections and geologic disposal of nuclear waste. Such activities are expected to rise in the future making it necessary to assess their short- and long-term safety. Here, a new machine learning (ML) approach to model pore pressure and fault displacements in response to high-pressure fluid injection cycles is developed. The focus is on fault behavior near the injection borehole. To capture the temporal dependencies in the data, long short-term memory (LSTM) networks are utilized. To prevent error accumulation within the forecast window, four critical measures to train a robust LSTM model for predicting fault response are highlighted: (i) setting an appropriate value of LSTM lag, (ii) calibrating the LSTM cell dimension, (iii) learning rate reduction during weight optimization, and (iv) not adopting an independent injection cycle as a validation set. Several numerical experiments were conducted, which demonstrated that the ML model can capture peaks in pressure and associated fault displacement that accompany an increase in fluid injection. The model also captured the decay in pressure and displacement during the injection shut-in period. Further, the ability of an ML model to highlight key changes in fault hydromechanical activation processes was investigated, which shows that ML can be used to monitor risk of fault activation and leakage during high pressure fluid injections.

58 GEOSCIENCES↗

Fail-Safe Logic Design Strategies Within Modern FPGA Architectures

Fail-safe computing refers to computing systems that revert to a non-operational safe state when a fault occurs. In this paper, we investigate a circuit level technique as mitigation for single event upsets (SEUs) and fault injection attacks on field programmable gate arrays (FPGAs), and analyze the effectiveness of the technique as a fail-safe monitor for an encryption algorithm. The propagation of fault effects through FPGA primitives including lookup tables (LUTs) and programmable interconnect points (PIPs) is assessed within an FPGA architecture created using an open source tool, and validated using fault injection experiments on an FPGA. The analysis reveals additional vulnerabilities exist within reconfigurable architectures over those in equivalent fail-safe application specific integrated circuit (ASIC), thus requiring a more elaborate network of redundant circuits and checking logic. The configuration memory bits (CMBs), which configure routing and designate logic functions within the LUTs of the FPGA, add complexity to fail-safe design strategies by introducing additional fault conditions and fault propagation paths. A resource-efficient fail-safe circuit design technique called DEsign for Fail-safe in reCONfigurable systems (DEFCON) is proposed. The benefits and limitations associated with DEFCON are described in the context of fault injection experiments carried out as simulations and in FPGA hardware.

Bhakta, Priya A. [Univ. of New Mexico, Albuquerque↗

Node Monitoring as a Fault Detection Countermeasure against Information Leakage within a RISC-V Microprocessor

Advanced, superscalar microprocessors (μP) are highly susceptible to wear-out failures because of their highly complex, densely packed circuit structure and extreme operational frequencies. Although many types of fault detection and mitigation strategies have been proposed, none have addressed the specific problem of detecting faults that lead to information leakage events on I/O channels of the μP. Information leakage can be defined very generally as any type of output that the executing program did not intend to produce. In this work, we restrict this definition to output that represents a security concern, and in particular, to the leakage of plaintext or encryption keys, and propose a counter-based countermeasure to detect faults that cause this type of leakage event. Fault injection (FI) experiments are carried out on two RISC-V microprocessors emulated as soft cores on a Xilinx multi-processor System-on-chip (MPSoC) FPGA. The μP designs are instrumented with a set of counters that records the number of transitions that occur on internal nodes. The transition counts are collected from all internal nodes under both fault-free and faulty conditions, and are analyzed to determine which counters provide the highest fault coverage and lowest latency for detecting leakage faults. We show that complete coverage of all leakage faults is possible using only a single counter strategically placed within the branch compare logic of the μPs.

42 ENGINEERING↗

Estimation of radiation fields generated by injected beam losses at the EIC's RCS

This technical note provides a general estimate of radiation fields generated by injection fault events at the electron-ion collider´s (EIC) Rapid Cycling Synchrotron (RCS), calculated with the Monte Carlo particle transport and interaction code FLUKA. Calculations were performed for two major injection loss scenarios that involve iron targets and featured different electron beam energy and current values. The results presented here constitute a first order assessment of several radiological quantities associated with these electromagnetic showers and their potential effect on environmental safety and health (ESH) systems in the vicinity of injection areas.

43 PARTICLE ACCELERATORS↗

ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training

Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensive and susceptible to faults, particularly in the attention mechanism, which is a critical component of transformer-based LLMs. In this paper, we investigate the impact of faults on LLM training, focusing on INF, NaN, and near-INF values in the computation results with systematic fault injection experiments. We observe the propagation patterns of these errors, which can trigger non-trainable states in the model and disrupt training, forcing the procedure to load from checkpoints. To mitigate the impact of these faults, we propose ATTNChecker, the first Algorithm-Based Fault Tolerance (ABFT) technique tailored for the attention mechanism in LLMs. ATTNChecker is designed based on fault propagation patterns of LLM and incorporates performance optimization to adapt to both system reliability and model vulnerability while providing lightweight protection for fast LLM training. Evaluations on four LLMs show that ATTNChecker on average incurs on average 7% overhead on training while detecting and correcting all extreme errors. Compared with the state-of-the-art checkpoint/restore approach, ATTNChecker reduces recovery overhead by up to 49×.

Liang, Yuhang [University of Alabama - Birmingham]↗

Towards Precision-Aware Fault Tolerance Approaches for Mixed-Precision Applications

Graphics Processing Units (GPUs), the dominantly adopted accelerators in HPC systems, are susceptible to transient hardware fault. New generation of GPUs feature mixed-precision architectures such as NVIDIA Tensor Cores to accelerate matrix multiplications. While being widely adapted, how would they behave under transient hardware faults remain unclear. In this study, we conduct a large-scale fault injection experiments on GEMM kernels implemented with different floating-point data types on the V100 and A100 Tensor Cores, and show distinct error resilience characteristics for the GEMMS with different formats. In the future, we plan to explore this space by building precision-aware floating-point fault tolerance techniques for applications such as DNNs that exercise low-precision computations.

Fang, Bo↗

HAPPA: A Modular Platform for HPC Application Resilience Analysis with LLMs Embedded

High-performance computing (HPC) systems are increasingly vulnerable to soft errors, which pose significant challenges in maintaining computational accuracy and reliability. Predicting the resilience of HPC applications to these errors is crucial for robust code protection and detailed resilience analysis. In this study, we present HAppA, a modular platform designed for HPC Application Resilience Analysis. Embedding Large Language Models (LLMs), HAppA addresses understanding the context information of long code sequences typical in HPC applications. HAppA implements a novel code representation module that chunks the code into fixed-size segments and aggregates the embeddings of these segments. Three aggregation methods have been explored: MeanPooling, MaxPooling, and LSTM-based techniques. We built a DAtaset for REsilience analysis using Fault Injection (FI), named DARE. Using our DARE dataset, HAppA is trained for regression prediction tasks. Our evaluation results demonstrate the predictive accuracy of HAppA compared to other models, particularly noting that the LSTM-based aggregation method -- HAppA-LSTM -- achieves a mean squared error (MSE) of 0.078 for SDC prediction, surpassing the existing state-of-the-art PARIS model, which recorded an MSE of 0.1172. Additionally, HAppA with the KeyBERT model extracts a list of keywords representing the source code. A comprehensive importance analysis of these keywords further elucidates the code patterns contributing to the error rate. These findings highlight the effectiveness of HAppA in analyzing the resilience of HPC applications and establish a new benchmark for predictive accuracy in resilience.

Jiang, Hailong [Kent State University]↗

Reconfigurable Framework for Resilient Semantic Segmentation for Space Applications

Deep learning (DL) presents new opportunities for enabling spacecraft autonomy, onboard analysis, and intelligent applications for space missions. However, DL applications are computationally intensive and often infeasible to deploy on radiation-hardened (rad-hard) processors, which traditionally harness a fraction of the computational capability of their commercial-off-the-shelf counterparts. Commercial FPGAs and system-on-chips present numerous architectural advantages and provide the computation capabilities to enable onboard DL applications; however, these devices are highly susceptible to radiation-induced single-event effects (SEEs) that can degrade the dependability of DL applications. In this article, we propose Reconfigurable ConvNet (RECON), a reconfigurable acceleration framework for dependable, high-performance semantic segmentation for space applications. In RECON, we propose both selective and adaptive approaches to enable efficient SEE mitigation. In our selective approach, control-flow parts are selectively protected by triple-modular redundancy to minimize SEE-induced hangs, and in our adaptive approach, partial reconfiguration is used to adapt the mitigation of dataflow parts in response to a dynamic radiation environment. Combined, both approaches enable RECON to maximize system performability subject to mission availability constraints. We perform fault injection and neutron irradiation to observe the susceptibility of RECON and use dependability modeling to evaluate RECON in various orbital case studies to demonstrate a 1.5–3.0× performability improvement in both performance and energy efficiency compared to static approaches.

97 MATHEMATICS AND COMPUTING↗

Reusable Verification Components for High-Energy Physics readout ASICs

Verification is a critical aspect of designing front-end (FE) readout ASICs for High-Energy Physics (HEP) experiments. These ASICs share several similar functional features, resulting in similar verification objectives, which can be addressed using comparable verification strategies. This contribution presents a set of re-usable verification components for addressing common verification tasks, such as clock generation, reset handling, configuration, as well as hit and fault injections. The components were developed as part of the CHIPS initiative and they have been successfully used in the verification of multiple HEP ASICs.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Factors controlling injection-induced rupture of intersecting faults during geological sequestration of CO 2

This study addresses coupled multiphase fluid flow and geomechanics effects on potential fault activation associated with subsurface CO 2 injection around intersecting faults. An enhanced fault-representation model is used to capture geomechanical responses of two intersecting faults with finite length during CO 2 injection. The faults are embedded in a strike-slip stress regime of a caprock-reservoir-basement system with the faults represented by zero-thickness interfaces with adjacent finite-thickness damage zones. A sensitivity analysis is conducted to study the effect of fault permeability, slip-weakening behavior, well location relative to the orientation of faults, and well placement (the number and location of injection wells). Five metrics (pressure, CO 2 plume, shear state on the fault, as well as shear displacement and stress path at selected fault monitoring points) are selected to assess CO 2 migration and reactivation of intersecting faults. The results show that induced ruptures are favored by low permeability faults due to high pressure buildup and by slip-weakening behavior resulting from fault strength reduction. The location of one injection well relative to fault orientation determines the magnitude of changes in effective normal stress and shear stress, affecting the location of induced ruptures. Well placement (two injection wells used in the paper) dominates pressure diffusion around the intersection and tips of faults. This redistributes changes in effective normal stress caused by each injection well, influencing the spatial distribution of ruptures along faults. A larger injection volume induces far-field ruptures that are controlled by stress transfer within the injection layer. The findings presented here can provide valuable insights into engineering operations for a long-term, safe, and reliable geologic CO 2 storage.

Fault permeability↗

Utah FORGE: Fault Reactivation Through Fluid Injection Induced Seismicity Laboratory Experiments

Included are results from shear reactivation experiments on laboratory faults pre-loaded close to failure and reactivated by the injection of fluid into the fault. The sample comprises a single-inclined-fracture (SIF) transecting a cylindrical sample of Westerly granite. All experiments are conducted at ambient temperature and follow a similar protocol: (i) application of confining stresses (3MPa) on the fault fully saturated with DI water, (ii) shear-mobilization through the increase of axial loading at a constant displacement rate until a post-peak steady-state condition is reached, (iii) reduction of axial loading and related shear stress to a prescribed fraction of the peak steady-state frictional strength (typically 60% to 90%, representing intermediate to high magnitudes) and (iv), fault reactivation triggered by a stepwise increase of pore pressure on the fault in 0.1 MPa increments held constant for 1-5 minutes. Mechanical data from three ISCO pumps connected to a Temco pressure vessel measure axial, confining, and fault-related parameters, including fluid pressure (kPa), fluid flow rate (mL/min), and axial displacement (mm). See included code for initial data analysis and visualization for select experiments. Resource names represent experiment numbers found in the "Read Me" file, which describes each experimental setup and parameters.

15 GEOTHERMAL ENERGY↗