Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “faults”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Robust Fault Detection

This research used mixed structured singular value theory to develop new estimator (or observer) based approaches to fault detection for dynamic systems. The initial developments were based on minimizing the H-infinity, I-1 and H2 system norms. The resultant fault detection algorithms were each shown to be successful, but the fault detection algorithm based on the I-1 norm was best able to detect abrupt faults. This latter technique was further improved by using fuzzy logic for the fault evaluation. Based on an anomaly observed in this research and apparently ignored in the literature, current research focuses on the determination of a fault using a norm of the change in the residual (the difference between the output of the system and observer) and not simply a norm of the residual itself. This research may lead to a fundamental contribution to research in fault detection and isolation.

Guo, Ten-Huei↗

Application of a Bank of Kalman Filters for Aircraft Engine Fault Diagnostics

In this paper, a bank of Kalman filters is applied to aircraft gas turbine engine sensor and actuator fault detection and isolation (FDI) in conjunction with the detection of component faults. This approach uses multiple Kalman filters, each of which is designed for detecting a specific sensor or actuator fault. In the event that a fault does occur, all filters except the one using the correct hypothesis will produce large estimation errors, thereby isolating the specific fault. In the meantime, a set of parameters that indicate engine component performance is estimated for the detection of abrupt degradation. The proposed FDI approach is applied to a nonlinear engine simulation at nominal and aged conditions, and the evaluation results for various engine faults at cruise operating conditions are given. The ability of the proposed approach to reliably detect and isolate sensor and actuator faults is demonstrated.

Kobayashi, Takahisa↗

Secondary Fracturing of Europa's Crust in Response to Combined Slip and Dilation Along Strike-Slip Faults

A commonly observed feature in faulted terrestrial rocks is the occurrence of secondary fractures alongside faults. Depending on exact morphology, such fractures have been termed tail cracks, wing cracks, kinks, or horsetail fractures, and typically form at the tip of a slipping fault or around small jogs or steps along a fault surface. The location and orientation of secondary fracturing with respect to the fault plane or the fault tip can be used to determine if fault motion is left-lateral or right-lateral.

Kattenhorn, S. A.↗

Parameter Transient Behavior Analysis on Fault Tolerant Control System

In a fault tolerant control (FTC) system, a parameter varying FTC law is reconfigured based on fault parameters estimated by fault detection and isolation (FDI) modules. FDI modules require some time to detect fault occurrences in aero-vehicle dynamics. This paper illustrates analysis of a FTC system based on estimated fault parameter transient behavior which may include false fault detections during a short time interval. Using Lyapunov function analysis, the upper bound of an induced-L2 norm of the FTC system performance is calculated as a function of a fault detection time and the exponential decay rate of the Lyapunov function.

Belcastro, Christine↗

Immunity-Based Aircraft Fault Detection System

In the study reported in this paper, we have developed and applied an Artificial Immune System (AIS) algorithm for aircraft fault detection, as an extension to a previous work on intelligent flight control (IFC). Though the prior studies had established the benefits of IFC, one area of weakness that needed to be strengthened was the control dead band induced by commanding a failed surface. Since the IFC approach uses fault accommodation with no detection, the dead band, although it reduces over time due to learning, is present and causes degradation in handling qualities. If the failure can be identified, this dead band can be further A ed to ensure rapid fault accommodation and better handling qualities. The paper describes the application of an immunity-based approach that can detect a broad spectrum of known and unforeseen failures. The approach incorporates the knowledge of the normal operational behavior of the aircraft from sensory data, and probabilistically generates a set of pattern detectors that can detect any abnormalities (including faults) in the behavior pattern indicating unsafe in-flight operation. We developed a tool called MILD (Multi-level Immune Learning Detection) based on a real-valued negative selection algorithm that can generate a small number of specialized detectors (as signatures of known failure conditions) and a larger set of generalized detectors for unknown (or possible) fault conditions. Once the fault is detected and identified, an adaptive control system would use this detection information to stabilize the aircraft by utilizing available resources (control surfaces). We experimented with data sets collected under normal and various simulated failure conditions using a piloted motion-base simulation facility. The reported results are from a collection of test cases that reflect the performance of the proposed immunity-based fault detection algorithm.

Dasgupta, D.↗

Arc burst pattern analysis fault detection system

A method and apparatus are provided for detecting an arcing fault on a power line carrying a load current. Parameters indicative of power flow and possible fault events on the line, such as voltage and load current, are monitored and analyzed for an arc burst pattern exhibited by arcing faults in a power system. These arcing faults are detected by identifying bursts of each half-cycle of the fundamental current. Bursts occurring at or near a voltage peak indicate arcing on that phase. Once a faulted phase line is identified, a comparison of the current and voltage reveals whether the fault is located in a downstream direction of power flow toward customers, or upstream toward a generation station. If the fault is located downstream, the line is de-energized, and if located upstream, the line may remain energized to prevent unnecessary power outages.

Russell, B. Don↗

Designing Fault-Injection Experiments for the Reliability of Embedded Systems

This paper considers the long-standing problem of conducting fault-injections experiments to establish the ultra-reliability of embedded systems. There have been extensive efforts in fault injection, and this paper offers a partial summary of the efforts, but these previous efforts have focused on realism and efficiency. Fault injections have been used to examine diagnostics and to test algorithms, but the literature does not contain any framework that says how to conduct fault-injection experiments to establish ultra-reliability. A solution to this problem integrates field-data, arguments-from-design, and fault-injection into a seamless whole. The solution in this paper is to derive a model reduction theorem for a class of semi-Markov models suitable for describing ultra-reliable embedded systems. The derivation shows that a tight upper bound on the probability of system failure can be obtained using only the means of system-recovery times, thus reducing the experimental effort to estimating a reasonable number of easily-observed parameters. The paper includes an example of a system subject to both permanent and transient faults. There is a discussion of integrating fault-injection with field-data and arguments-from-design.

White, Allan L.↗

An Integrated Approach for Aircraft Engine Performance Estimation and Fault Diagnostics

A Kalman filter-based approach for integrated on-line aircraft engine performance estimation and gas path fault diagnostics is presented. This technique is specifically designed for underdetermined estimation problems where there are more unknown system parameters representing deterioration and faults than available sensor measurements. A previously developed methodology is applied to optimally design a Kalman filter to estimate a vector of tuning parameters, appropriately sized to enable estimation. The estimated tuning parameters can then be transformed into a larger vector of health parameters representing system performance deterioration and fault effects. The results of this study show that basing fault isolation decisions solely on the estimated health parameter vector does not provide ideal results. Furthermore, expanding the number of the health parameters to address additional gas path faults causes a decrease in the estimation accuracy of those health parameters representative of turbomachinery performance deterioration. However, improved fault isolation performance is demonstrated through direct analysis of the estimated tuning parameters produced by the Kalman filter. This was found to provide equivalent or superior accuracy compared to the conventional fault isolation approach based on the analysis of sensed engine outputs, while simplifying online implementation requirements. Results from the application of these techniques to an aircraft engine simulation are presented and discussed.

imon, Donald L.↗

Fault Management Metrics

This paper describes the theory and considerations in the application of metrics to measure the effectiveness of fault management. Fault management refers here to the operational aspect of system health management, and as such is considered as a meta-control loop that operates to preserve or maximize the system's ability to achieve its goals in the face of current or prospective failure. As a suite of control loops, the metrics to estimate and measure the effectiveness of fault management are similar to those of classical control loops in being divided into two major classes: state estimation, and state control. State estimation metrics can be classified into lower-level subdivisions for detection coverage, detection effectiveness, fault isolation and fault identification (diagnostics), and failure prognosis. State control metrics can be classified into response determination effectiveness and response effectiveness. These metrics are applied to each and every fault management control loop in the system, for each failure to which they apply, and probabilistically summed to determine the effectiveness of these fault management control loops to preserve the relevant system goals that they are intended to protect.

Johnson, Stephen B.↗

Model-Based Investigation of Multi-Fault Interactions and Performance Degradation in Residential Heat Pump Systems

Faults in heat pump systems can significantly degrade performance, reduce efficiency, and accelerate component wear, leading to higher operating costs and maintenance demands. While numerous studies have investigated the impact of individual faults, the interactions between multiple concurrent faults remain insufficiently understood, despite their common occurrence in real-world operation. This study conducts a comprehensive simulation analysis of multiple simultaneous faults in a vapor compression heat pump using a validated heat pump design model (HPDM) tool. Detailed component-level modeling methods are implemented to examine performance sensitivity under combinations of refrigerant flow and heat exchanger faults. The results reveal complex fault interactions that can mask or amplify system deviations, challenging conventional diagnostic approaches. Findings from this work provide meaningful insights for the development of more robust fault detection and diagnosis algorithms, supporting improved reliability and energy efficiency in next-generation heat pump technologies.

Hu, Yifeng [ORNL] (ORCID:0000000242875185)↗

Deep learning model for fast, science-based forecasting of fluid migration along faults in geologic carbon storage scenarios

Effective long-term geologic storage depends on robust site selection and credible, science-based forecasting of subsurface behavior to ensure storage integrity. For this work, we develop a deep learning–based reduced-order model (ROM) to quantify potential carbon dioxide (CO₂) and brine migration through geological faults. The ROM combines a Transformer model for binary classification and a Stacked Ensemble for regression, trained on a comprehensive dataset generated from 1400 physics-based reservoir simulations. Key geologic and operational parameters—including fault geometry, reservoir structure, and injection conditions—were systematically varied to capture a wide range of fluid migration scenarios. The ROM accurately predicts the onset of migration, cumulative migration volumes of both CO₂ and brine, and associated migration rates, as compared to an independent set of validation simulations, while significantly reducing computational cost compared to traditional simulation methods. Model performance was evaluated across diverse fault configurations, revealing that shallow reservoir geometry and fault angle are among the most influential factors governing migration behavior. Sensitivity analysis using SHapley Additive exPlanations (SHAP) provided interpretability, revealing distinct patterns in how geological and operational features drive transient versus cumulative migration outcomes. The ROM’s ability to rapidly simulate fault migration scenarios enables efficient sensitivity analyses, scenario evaluations, and decision support for site selection and monitoring design. This approach enhances the safety, scalability, and long-term operational performance of geologic carbon storage (GCS) systems by providing a robust, interpretable tool for predicting subsurface fluid migration and assessing fault-related migration potential.

42 ENGINEERING↗

Fault Mitigation Schemes for Future Spaceflight Multicore Processors

Future planetary exploration missions demand significant advances in on-board computing capabilities over current avionics architectures based on a single-core processing element. The state-of-the-art multi-core processor provides much promise in meeting such challenges while introducing new fault tolerance problems when applied to space missions. Software-based schemes are being presented in this paper that can achieve system-level fault mitigation beyond that provided by radiation-hard-by-design (RHBD). For mission and time critical applications such as the Terrain Relative Navigation (TRN) for planetary or small body navigation, and landing, a range of fault tolerance methods can be adapted by the application. The software methods being investigated include Error Correction Code (ECC) for data packet routing between cores, virtual network routing, Triple Modular Redundancy (TMR), and Algorithm-Based Fault Tolerance (ABFT). A robust fault tolerance framework that provides fail-operational behavior under hard real-time constraints and graceful degradation will be demonstrated using TRN executing on a commercial Tilera(R) processor with simulated fault injections.

software based↗

Analyzing and Predicting Effort Associated with Finding and Fixing Software Faults

Context: Software developers spend a significant amount of time fixing faults. However, not many papers have addressed the actual effort needed to fix software faults. Objective: The objective of this paper is twofold: (1) analysis of the effort needed to fix software faults and how it was affected by several factors and (2) prediction of the level of fix implementation effort based on the information provided in software change requests. Method: The work is based on data related to 1200 failures, extracted from the change tracking system of a large NASA mission. The analysis includes descriptive and inferential statistics. Predictions are made using three supervised machine learning algorithms and three sampling techniques aimed at addressing the imbalanced data problem. Results: Our results show that (1) 83% of the total fix implementation effort was associated with only 20% of failures. (2) Both safety critical failures and post-release failures required three times more effort to fix compared to non-critical and pre-release counterparts, respectively. (3) Failures with fixes spread across multiple components or across multiple types of software artifacts required more effort. The spread across artifacts was more costly than spread across components. (4) Surprisingly, some types of faults associated with later life-cycle activities did not require significant effort. (5) The level of fix implementation effort was predicted with 73% overall accuracy using the original, imbalanced data. Using oversampling techniques improved the overall accuracy up to 77%. More importantly, oversampling significantly improved the prediction of the high level effort, from 31% to around 85%. Conclusions: This paper shows the importance of tying software failures to changes made to fix all associated faults, in one or more software components and/or in one or more software artifacts, and the benefit of studying how the spread of faults and other factors affect the fix implementation effort.

software fix implementation effort↗

On the distribution of stacking faults at dissociated medium-angle grain boundaries: Crystallographic geometry and metastability

Grain boundaries in FCC metals with low stacking-fault energy can form in dissociated configurations of stacking faults. Perhaps the most studied example is 9R stacking at boundaries near Σ3{112}, where the distribution of stacking faults is related to the emission of Shockley partial dislocations. Here, we combine atomic-resolution electron microscopy, atomistic simulations, and dislocation theory to demonstrate that boundaries vicinal to Σ33a support the stabilization of dissociated 9R stacking within a narrow range of inclinations. This boundary is interesting since its misorientation (20.05°) lies in the medium-angle regime, just past the upper misorientation limit for low-angle boundaries, motivating questions for how best to describe it in terms of dislocations. Our HAADF-STEM observations of thin film bicrystals, supported by atomistic modeling, reveal that this inclination dependence arises from specific geometric constraints on the arrangement of Shockley partial dislocations at the interface. Quantification of stacking-fault distributions across multiple boundaries indicates that the density and spacing of faults closely follow the ideal 9R motif, with subtle variations reflecting the complex energy landscape of these boundaries. Through energetic analysis, we establish the presence of competing metastable states enabled by variations in stacking sequences, emphasizing the significant role of crystallographic geometry. We generalize our analysis as a function of misorientation, showing how 9R at the Σ33a boundary is related to previous observations and calculations of HCP at a 29.7° boundary. This study provides a crystallographically grounded framework connecting dislocation structures, stacking-fault distributions, and metastability at grain boundaries in FCC metals.

Atomistic modeling↗

Fault Network Geometry Modulates Earthquake Source Spectra Across Scales

Earthquake source spectra provide unique insights into the earthquake rupture process. Motivated by previous research suggesting that complex fault geometries enhance high‐frequency seismic radiation, we study the influence of fault network geometry on earthquake source spectra using multiple independent observations. At regional scales, we examine correlations of stress drop measurements with surface fault trace misalignment in Southern California, Japan, and Central Italy. At a global scale, we examine correlations of moment‐rate function complexity of large earthquakes with focal mechanism variability, a proxy for local fault complexity. Despite significant scatter in the observations, we find overall consistent positive correlations. The concept that elastic interactions of discrete fault structures during the earthquake rupture process generates high‐frequency ground motions offers a coherent framework for interpreting our observations. These findings suggest that variations in fault complexity explain why some earthquakes produce stronger high‐frequency ground motions than others.

Lee, Jaeseok [Brown Univ., Providence, RI (United ↗

A Boundary Element Model for Assessing Large‐Scale Pressurization in Faulted Geological Storage Systems

Assessing large-scale pressurization at the regional scale—a possible outcome of large subsurface storage applications such as wastewater injection and geological carbon sequestration—presents significant computational challenges. These challenges are particularly pronounced when accounting for complex geologic structures with multiple reservoir and caprock layers, fault zones, and wells. This study introduces a computationally efficient model that integrates single-phase semi-analytical solutions with a boundary element (BE) approach. The model simulates pressure propagation in multilayered 3D systems, including vertical faults, caprock, basement, and confining units. We apply this new model to a representative scenario involving CO 2 injection near a partially sealing fault with verification against an independent two-phase flow model. Results demonstrate that our model accurately captures far-field pressure responses and that, outside the CO 2 plume zone, pressure predictions from single-phase and two-phase models are nearly identical. This supports the use of single-phase models like ours for efficient estimation of far-field pressure changes. Additionally, we demonstrate its effectiveness at a large scale, incorporating multiple wells and faults. With its ability to represent multiple wells, fault zones, and geological heterogeneity, our model is well suited for assessments of basin-scale pressurization. Its computational efficiency also makes it a promising tool for integration with optimization frameworks aimed at designing and managing injection strategies in faulted storage systems.

Cihan, A. [Lawrence Berkeley National Laboratory (↗

ATTNChecker: Highly-Optimized Fault Tolerant Attention for Large Language Model Training

Large Language Models (LLMs) have demonstrated remarkable performance in various natural language processing tasks. However, the training of these models is computationally intensive and susceptible to faults, particularly in the attention mechanism, which is a critical component of transformer-based LLMs. In this paper, we investigate the impact of faults on LLM training, focusing on INF, NaN, and near-INF values in the computation results with systematic fault injection experiments. We observe the propagation patterns of these errors, which can trigger non-trainable states in the model and disrupt training, forcing the procedure to load from checkpoints. To mitigate the impact of these faults, we propose ATTNChecker, the first Algorithm-Based Fault Tolerance (ABFT) technique tailored for the attention mechanism in LLMs. ATTNChecker is designed based on fault propagation patterns of LLM and incorporates performance optimization to adapt to both system reliability and model vulnerability while providing lightweight protection for fast LLM training. Evaluations on four LLMs show that ATTNChecker on average incurs on average 7% overhead on training while detecting and correcting all extreme errors. Compared with the state-of-the-art checkpoint/restore approach, ATTNChecker reduces recovery overhead by up to 49×.

Liang, Yuhang [University of Alabama - Birmingham]↗

An Advanced Synchronized Time Digital Grid Twin Testbed for Relay Misoperation Analysis of Electrical Fault Type Detection Algorithms

Distributed energy resources and the number of relays are expected to rise in modern electrical grids; consequently, relay misoperations are also expected to grow. Relays can detect electrical fault types using an internal algorithm and can display the result using light indicators on the front of the relay. However, some relays’ internal algorithms for predicting types of electrical faults could be improved. This study assesses a relay’s external and internal algorithms with an Advanced Synchronized Time Digital Grid Twin (ASTDGT) testbed with paired relays. A misoperation relay analysis focused on measuring the accuracy of using the boundary admittance (the external algorithm) versus the set-default (the internal algorithm) relay method to determine the electrical fault types was performed. In this study, the internal and external relay algorithms were assessed with a synchronized time digital grid twin testbed using a real-time simulator. This testbed evaluated two sets of logic at the same time with the digital grid twin and paired relays in the loop. Different types of electrical faults were simulated, and the relays’ recorded events and electrical fault light indicator states were collected from the human–machine interfaces. This ASTDGT testbed with paired relays successfully evaluated the relay algorithm misoperations. The boundary admittance method had an accuracy of 100% for line-to-line, line-to-ground, and line-to-line ground faults.

24 POWER TRANSMISSION AND DISTRIBUTION↗