Engineering PapersSearch

SEARCH · Engineering Papers

Results for “fault mitigation methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Fault Mitigation Schemes for Future Spaceflight Multicore Processors

Future planetary exploration missions demand significant advances in on-board computing capabilities over current avionics architectures based on a single-core processing element. The state-of-the-art multi-core processor provides much promise in meeting such challenges while introducing new fault tolerance problems when applied to space missions. Software-based schemes are being presented in this paper that can achieve system-level fault mitigation beyond that provided by radiation-hard-by-design (RHBD). For mission and time critical applications such as the Terrain Relative Navigation (TRN) for planetary or small body navigation, and landing, a range of fault tolerance methods can be adapted by the application. The software methods being investigated include Error Correction Code (ECC) for data packet routing between cores, virtual network routing, Triple Modular Redundancy (TMR), and Algorithm-Based Fault Tolerance (ABFT). A robust fault tolerance framework that provides fail-operational behavior under hard real-time constraints and graceful degradation will be demonstrated using TRN executing on a commercial Tilera(R) processor with simulated fault injections.

software based

Evolutionary Based Techniques for Fault Tolerant Field Programmable Gate Arrays

The use of SRAM-based Field Programmable Gate Arrays (FPGAs) is becoming more and more prevalent in space applications. Commercial-grade FPGAs are potentially susceptible to permanently debilitating Single-Event Latchups (SELs). Repair methods based on Evolutionary Algorithms may be applied to FPGA circuits to enable successful fault recovery. This paper presents the experimental results of applying such methods to repair four commonly used circuits (quadrature decoder, 3-by-3-bit multiplier, 3-by-3-bit adder, 440-7 decoder) into which a number of simulated faults have been introduced. The results suggest that evolutionary repair techniques can improve the process of fault recovery when used instead of or as a supplement to Triple Modular Redundancy (TMR), which is currently the predominant method for mitigating FPGA faults.

Larchev, Gregory V.

Extended Testability Analysis Tool

The Extended Testability Analysis (ETA) Tool is a software application that supports fault management (FM) by performing testability analyses on the fault propagation model of a given system. Fault management includes the prevention of faults through robust design margins and quality assurance methods, or the mitigation of system failures. Fault management requires an understanding of the system design and operation, potential failure mechanisms within the system, and the propagation of those potential failures through the system. The purpose of the ETA Tool software is to process the testability analysis results from a commercial software program called TEAMS Designer in order to provide a detailed set of diagnostic assessment reports. The ETA Tool is a command-line process with several user-selectable report output options. The ETA Tool also extends the COTS testability analysis and enables variation studies with sensor sensitivity impacts on system diagnostics and component isolation using a single testability output. The ETA Tool can also provide extended analyses from a single set of testability output files. The following analysis reports are available to the user: (1) the Detectability Report provides a breakdown of how each tested failure mode was detected, (2) the Test Utilization Report identifies all the failure modes that each test detects, (3) the Failure Mode Isolation Report demonstrates the system s ability to discriminate between failure modes, (4) the Component Isolation Report demonstrates the system s ability to discriminate between failure modes relative to the components containing the failure modes, (5) the Sensor Sensor Sensitivity Analysis Report shows the diagnostic impact due to loss of sensor information, and (6) the Effect Mapping Report identifies failure modes that result in specified system-level effects.

Melcher, Kevin

NASA Tech Briefs, March 2014

Topics include: Data Fusion for Global Estimation of Forest Characteristics From Sparse Lidar Data; Debris and Ice Mapping Analysis Tool - Database; Data Acquisition and Processing Software - DAPS; Metal-Assisted Fabrication of Biodegradable Porous Silicon Nanostructures; Post-Growth, In Situ Adhesion of Carbon Nanotubes to a Substrate for Robust CNT Cathodes; Integrated PEMFC Flow Field Design for Gravity-Independent Passive Water Removal; Thermal Mechanical Preparation of Glass Spheres; Mechanistic-Based Multiaxial-Stochastic-Strength Model for Transversely-Isotropic Brittle Materials; Methods for Mitigating Space Radiation Effects, Fault Detection and Correction, and Processing Sensor Data; Compact Ka-Band Antenna Feed with Double Circularly Polarized Capability; Dual-Leadframe Transient Liquid Phase Bonded Power Semiconductor Module Assembly and Bonding Process; Quad First Stage Processor: A Four-Channel Digitizer and Digital Beam-Forming Processor; Protective Sleeve for a Pyrotechnic Reefing Line Cutter; Metabolic Heat Regenerated Temperature Swing Adsorption; CubeSat Deployable Log Periodic Dipole Array; Re-entry Vehicle Shape for Enhanced Performance; NanoRacks-Scale MEMS Gas Chromatograph System; Variable Camber Aerodynamic Control Surfaces and Active Wing Shaping Control; Spacecraft Line-of-Sight Stabilization Using LWIR Earth Signature; Technique for Finding Retro-Reflectors in Flash LIDAR Imagery; Novel Hemispherical Dynamic Camera for EVAs; 360 deg Visual Detection and Object Tracking on an Autonomous Surface Vehicle; Simulation of Charge Carrier Mobility in Conducting Polymers; Observational Data Formatter Using CMOR for CMIP5; Propellant Loading Physics Model for Fault Detection Isolation and Recovery; Probabilistic Guidance for Swarms of Autonomous Agents; Reducing Drift in Stereo Visual Odometry; Future Air-Traffic Management Concepts Evaluation Tool; Examination and A Priori Analysis of a Direct Numerical Simulation Database for High-Pressure Turbulent Flows; and Resource-Constrained Application of Support Vector Machines to Imagery.

Source record

Wide bandgap optical switch circuit breaker for controlling propagation of current therethrough a wide bandgap optical device

A high-voltage switch is adapted for use as a medium-voltage direct current circuit breaker, which provides a low-cost, small-footprint device to mitigate system faults. In one example, a method for operating a wideband optical device includes illuminating the wide bandgap optical device with a light within a first range of wavelengths and a first average intensity, allowing a current to propagate therethrough without substantial absorption of the current, illuminating the wide bandgap optical device with light within the first range of wavelengths and a second average intensity that is lower than the first average intensity to allow a sustained current flow though the wide bandgap optical device, and illuminating the wide bandgap optical device with light within a second range of wavelengths to stop or substantially restrict propagation of the current through the wide gap material.

Voss, Lars F.

Fault Tree Analysis Application for Safety and Reliability

Many commercial software tools exist for fault tree analysis (FTA), an accepted method for mitigating risk in systems. The method embedded in the tools identifies a root as use in system components, but when software is identified as a root cause, it does not build trees into the software component. No commercial software tools have been built specifically for development and analysis of software fault trees. Research indicates that the methods of FTA could be applied to software, but the method is not practical without automated tool support. With appropriate automated tool support, software fault tree analysis (SFTA) may be a practical technique for identifying the underlying cause of software faults that may lead to critical system failures. We strive to demonstrate that existing commercial tools for FTA can be adapted for use with SFTA, and that applied to a safety-critical system, SFTA can be used to identify serious potential problems long before integrator and system testing.

Wallace, Dolores R.

Early Fault Detection in Nuclear Systems: A Digital Engineering Approach

Nuclear energy systems present unique challenges in terms of ensuring safety, reliability, and efficiency during their design and operation. Early fault detection is critical for mitigating risks and fostering system resilience. However, current methods often fall short at identifying faults during early stages, potentially leading to costly delays and safety risks. The present work proposes a comprehensive digital engineering approach that leverages digital twins, digital threads, model-based systems engineering, artificial intelligence, and immersive extended reality to support early fault detection in nuclear systems. Through a series of case studies, we highlight specific gaps in the fault detection mechanisms of traditional nuclear design and operation processes, then demonstrate a suite of solutions we are working to implement to address these shortcomings in similar projects. Our findings suggest that a digital engineering approach to design and operation can significantly improve fault detection, ultimately leading to reductions in risk.

42 - ENGINEERING

Seismicity-constrained fault detection and characterization with a multitask machine learning model

Geological fault detection and characterization are crucial for understanding subsurface dynamics across scales. While methods for fault delineation based on either seismicity location analysis or seismic image reflector discontinuity are well-established, a systematic approach that integrates both data types remains absent. We develop a novel machine learning model that unifies seismic reflector images and seismicity location information to automatically identify geological faults and characterize their geometrical properties. The model encodes a seismic image and a seismicity location image separately, and fuses the encoded features with a spatial-channel attention fusion module to improve the learning of important features in both inputs. We design an automated strategy to generate high-quality synthetic training data and labels. To improve the realism of the seismicity location image, we include random seismicity noise and missing seismicity location associated with some of the faults. We validate the model’s efficacy and accuracy using synthetic data examples and two field data examples. Moreover, we show that fine-tuning the trained model with a small, domain-specific dataset enhances its fidelity for field data applications. The results demonstrate that integrating seismicity location and seismic images into a unified framework allows the end-to-end neural network to achieve higher fidelity and accuracy in delineating subsurface faults and their geometrical properties compared with image-only fault detection methods. Our approach offers an adaptive data-driven tool for geological fault characterization and seismic hazard mitigation, bridging the gap between seismicity location and image-based fault detection methods.

58 GEOSCIENCES

Logical Shadow Tomography: Efficient Estimation of Error-mitigated Observables

In near-term quantum applications, reducing errors and improving device reliability is an essential task. Towards these ends, various techniques have been introduced in recent literature, collectively referred to as quantum error mitigation techniques, for reducing errors in pre-fault-tolerant devices. Here, we introduce logical shadow tomography as a versatile error mitigation method. Our technique uses a stabilizer code to encode information in a logical state. Instead of doing active error correction, quantum states will be measured at the end of computation via shadow tomography and non-logical errors are projected out in the classical post-processing. Relative to quantum subspace expansion which requires O(2(M-1)L) experiments to estimate an logical Pauli observable encoded by an [[M, L, d]] code, our technique only requires 2L experiments, an important practical reduction in resources.

Hong-Ye Hu

Design and Testing of a Hard-Fault Protection Circuit for a 1 kV SiC MOSFET Inverter

Due to increasingly high DC link voltages and further advancements in the current density of silicon carbide (SiC) MOSFETs, it has become evident that conventional IGBT protection methods are not sufficient to prevent exceeding the current rating of these devices during low-inductance fault events. This paper explores the use of an air core Rogowski coil topology to mitigate these hard fault events. The design of this circuit resulted in safe shutdown of a low impedance phase-to-phase fault in under one microsecond, tested up to DC link voltages of 1 kV. This paper details the theory, design, simulation, and successful test results of this method.

hard fault protection

Human Error Analysis for Human-Rated Space Systems

Humans bring unique capabilities to space systems and contribute to mission success in a manner that cannot be matched by machines. Nevertheless, from time to time, human error can present a threat to system performance, and system designers must anticipate and manage this risk. NASA’s Human-Rating Requirements for Space Systems call for program managers to conduct a human error analysis (HEA) during system development but does not specify how to do this. In 2018, NASA’s Engineering and Safety Center asked the authors to develop a guidance document on HEA. The resulting position paper outlines a suggested method for HEA and makes it clear that error analysis is about identifying and mitigating problems at a system level, and not about finding fault with individuals. Error management strategies must be directed at error-producing conditions, thereby reducing the likelihood of human error, while retaining the positive contribution that humans make to system operations.

human error human-rated space

NASA Space Flight Vehicle Fault Isolation Challenges

The Space Launch System (SLS) is the new NASA heavy lift launch vehicle and is scheduled for its first mission in 2017. The goal of the first mission, which will be uncrewed, is to demonstrate the integrated system performance of the SLS rocket and spacecraft before a crewed flight in 2021. SLS has many of the same logistics challenges as any other large scale program. Common logistics concerns for SLS include integration of discrete programs geographically separated, multiple prime contractors with distinct and different goals, schedule pressures and funding constraints. However, SLS also faces unique challenges. The new program is a confluence of new hardware and heritage, with heritage hardware constituting seventy-five percent of the program. This unique approach to design makes logistics concerns such as testability of the integrated flight vehicle especially problematic. The cost of fully automated diagnostics can be completely justified for a large fleet, but not so for a single flight vehicle. Fault detection is mandatory to assure the vehicle is capable of a safe launch, but fault isolation is another issue. SLS has considered various methods for fault isolation which can provide a reasonable balance between adequacy, timeliness and cost. This paper will address the analyses and decisions the NASA Logistics engineers are making to mitigate risk while providing a reasonable testability solution for fault isolation.

Bramon, Christopher

Hard Fault Protection for a Silicon Carbide-Based Aerospace Motor Drive

Due to increasingly high DC link voltages and further advancements in the current density of silicon carbide (SiC) MOSFETs, it has become evident that conventional IGBT protection methods are not sufficient to protect these devices from overcurrent during low-inductance fault events. The use of an air core Rogowski coil topology was explored to see if it could mitigate these hard fault events. The design of this circuit resulted in safe shutdown of a low impedance phase-tophase fault, tested up to DC link voltages of 1 kV.

High Voltage

Quantum utility-scale error mitigation for quantum quench dynamics in Heisenberg spin chains

Here, we implement a quantum error mitigation method termed self-mitigation, which is comparable to zero-noise extrapolation, at large scales to achieve quantum utility on near-term, noisy quantum computers. We investigate the effectiveness of several quantum error mitigation strategies, including self-mitigation, by simulating quantum quench dynamics for Heisenberg spin chains with system sizes up to 104 qubits using IBM quantum processors. In particular, we discuss the limitations of zero-noise extrapolation and the advantages offered by self-mitigation at large scales. The self-mitigation method demonstrates stable accuracy with large systems of 104 qubits comprising more than 3,000 CNOT gates. Also, we combine the discussed quantum error mitigation methods with practical entanglement entropy measuring methods, and it shows a good agreement with the theoretical estimation. Our study illustrates the usefulness of near-term noisy quantum hardware in examining the quantum quench dynamics of many-body systems at large scales and lays the groundwork for surpassing classical simulations with quantum methods prior to the development of fault-tolerant quantum computers.

97 MATHEMATICS AND COMPUTING

Risk Mitigation for Managing On-Orbit Anomalies

This slide presentation reviews strategies for managing risk mitigation that occur with anomalies in on-orbit spacecraft. It reviews the risks associated with mission operations, a diagram of the method used to manage undesirable events that occur which is a closed loop fault analysis and until corrective action is successful. It also reviews the fish bone diagram which is used if greater detail is required and aids in eliminating possible failure factors.

La, Jim

Overview of the SCEC/USGS Community Stress Drop Validation Study Using the 2019 Ridgecrest Earthquake Sequence

We present initial findings from the ongoing Community Stress Drop Validation Study to compare spectral stress-drop estimates for earthquakes in the 2019 Ridgecrest, California, sequence. This study uses a unified dataset to independently estimate earthquake source parameters through various methods. Stress drop, which denotes the change in average shear stress along a fault during earthquake rupture, is a critical parameter in earthquake science, impacting ground motion, rupture simulation, and source physics. Spectral stress drop is commonly derived by fitting the amplitude-spectrum shape, but estimates can vary substantially across studies for individual earthquakes. Sponsored jointly by the U.S. Geological Survey and the Statewide (previously, Southern) California Earthquake Center our community study aims to elucidate sources of variability and uncertainty in earthquake spectral stress-drop estimates through quantitative comparison of submitted results from independent analyses. The dataset includes nearly 13,000 earthquakes ranging from M 1 to 7 during a two-week period of the 2019 Ridgecrest sequence, recorded within a 1° radius. Here, in this article, we report on 56 unique submissions received from 20 different groups, detailing spectral corner frequencies (or source durations), moment magnitudes, and estimated spectral stress drops. Methods employed encompass spectral ratio analysis, spectral decomposition and inversion, finite-fault modeling, ground-motion-based approaches, and combined methods. Initial analysis reveals significant scatter across submitted spectral stress drops spanning over six orders of magnitude. However, we can identify between-method trends and offsets within the data to mitigate this variability. Averaging submissions for a prioritized subset of 56 events shows reduced variability of spectral stress drop, indicating overall consistency in recovered spectral stress-drop values.

58 GEOSCIENCES

Damage Characterization Using the Extended Finite Element Method for Structural Health Management

The development of validated multidisciplinary Integrated Vehicle Health Management (IVHM) tools, technologies, and techniques to enable detection, diagnosis, prognosis, and mitigation in the presence of adverse conditions during flight will provide effective solutions to deal with safety related challenges facing next generation aircraft. The adverse conditions include loss of control caused by environmental factors, actuator and sensor faults or failures, and damage conditions. A major concern in these structures is the growth of undetected damage/cracks due to fatigue and low velocity foreign impact that can reach a critical size during flight, resulting in loss of control of the aircraft. Hence, development of efficient methodologies to determine the presence, location, and severity of damage/cracks in critical structural components is highly important in developing efficient structural health management systems.

Krishnamurthy, Thiagarajan

ByzSec — A Multi-layered Byzantine Resilient Architecture for Bulk Power System Protective Relays

Reliability, selectivity, and sensitivity are the fundamental attributes of any protection system, acting as the main drivers in the selection of schemes, and equipment. In high-voltage systems, microprocessor-based relays represent the industry’s preferred solution, providing engineers with a vast array of benefits. However, they remain vulnerable to cybersecurity events that may compromise their functionality. To help mitigate against potential cybersecurity risks, this paper presents a fault-tolerant, Byzantine Resilient (BR) architecture that significantly increases the cybersecurity attributes of a protection system while minimizing the amount of performance impacts and integration overheads introduced. The solution relies on an array of independent relays that utilize robust consensus methods (based on Spire [1], [2]) to ensure correct system behavior is achieved even when a relay has been compromised. Furthermore, the solution has been complemented with a custom-built Situational Awareness engine that can be used to detect and identify potential threats. The implemented solution has been developed in consultation with three hardware vendors and has been tested to comply with the performance requirements of a 345kV differential protection scheme (87T). The results indicate that the proposed architecture is a comprehensive solution that: supports the strict correctness and performance requirements of the bulk power grid while providing a cost-effective alternative that offers a seamless, long-term solution.

byzantine security, Fault Tolerant Application Sof