Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A bilinear failure criterion for mixed-mode delamination

Many different failure criteria have been suggested for mixed-mode delamination toughness, but few sets of mixed-mode data exist that are consistent over the full range of Mode 1 opening load to Mode 2 shear load range. The mixed-mode bending (MMB) test was used to measure the delamination toughness of a brittle epoxy composite, a state-of-the-art toughened epoxy composite, and a tough thermoplastic composite over the full mixed-mode range. To gain insight into the different failure responses of the different materials, the delamination fracture surfaces were also examined. An evaluation of several failure criteria that have been reported in the literature was performed, and the range of responses modeled by each criterion was analyzed. A bilinear failure criterion was introduced based on a change in the failure mechanism observed from the delamination surfaces. The different criteria were compared to the failure response of the three materials tested. The responses of the two epoxies were best modeled with the new bilinear failure criterion. The failure response of the tough thermoplastic composite could be modeled well with the bilinear criterion but could also be modeled with the more simple linear failure criterion. Since the materials differed in their mixed-mode failure response, mixed-mode delamination testing will be needed to characterize a composite material. This paper presents consistent sets of mixed-mode data, provides a critical evaluation of the mixed-mode failure criteria, and should provide general guidance for selecting an appropriate criterion for other materials.

Reeder, James R.

Progressive Failure Analysis Methodology for Laminated Composite Structures

A progressive failure analysis method has been developed for predicting the failure of laminated composite structures under geometrically nonlinear deformations. The progressive failure analysis uses C(exp 1) shell elements based on classical lamination theory to calculate the in-plane stresses. Several failure criteria, including the maximum strain criterion, Hashin's criterion, and Christensen's criterion, are used to predict the failure mechanisms and several options are available to degrade the material properties after failures. The progressive failure analysis method is implemented in the COMET finite element analysis code and can predict the damage and response of laminated composite structures from initial loading to final failure. The different failure criteria and material degradation methods are compared and assessed by performing analyses of several laminated composite structures. Results from the progressive failure method indicate good correlation with the existing test data except in structural applications where interlaminar stresses are important which may cause failure mechanisms such as debonding or delaminations.

Sleight, David W.

Analyses of Transistor Punchthrough Failures

The failure of two transistors in the Altitude Switch Assembly for the Solid Rocket Booster followed by two additional failures a year later presented a challenge to failure analysts. These devices had successfully worked for many years on numerous missions. There was no history of failures with this type of device. Extensive checks of the test procedures gave no indication for a source of the cause. The devices were manufactured more than twenty years ago and failure information on this lot date code was not readily available. External visual exam, radiography, PEID, and leak testing were performed with nominal results Electrical testing indicated nearly identical base-emitter and base-collector characteristics (both forward and reverse) with a low resistance short emitter to collector. These characteristics are indicative of a classic failure mechanism called punchthrough. In failure analysis punchthrough refers to an condition where a relatively low voltage pulse causes the device to conduct very hard producing localized areas of thermal runaway or "hot spots". At one or more of these hot spots, the excessive currents melt the silicon. Heavily doped emitter material diffuses through the base region to the collector forming a diffusion pipe shorting the emitter to base to collector. Upon cooling, an alloy junction forms between the pipe and the base region. Generally, the hot spot (punch-through site) is under the bond and no surface artifact is visible. The devices were delidded and the internal structures were examined microscopically. The gold emitter lead was melted on one device, but others had anomalies in the metallization around the in-tact emitter bonds. The SEM examination confirmed some anomalies to be cosmetic defects while other anomalies were artifacts of the punchthrough site. Subsequent to these analyses, the contractor determined that some irregular testing procedures occurred at the time of the failures heretofore unreported. These testing irregularities involved the use of a breakout box and were the likely cause of the failures. There was no evidence to suggest a generic failure mechanism was responsible for the failure of these transistors.

Nicolas, David P.

Evaluation of a Multi-Axial, Temperature, and Time Dependent (MATT) Failure Model

To obtain a better understanding the response of the structural adhesives used in the Space Shuttle's Reusable Solid Rocket Motor (RSRM) nozzle, an extensive effort has been conducted to characterize in detail the failure properties of these adhesives. This effort involved the development of a failure model that includes the effects of multi-axial loading, temperature, and time. An understanding of the effects of these parameters on the failure of the adhesive is crucial to the understanding and prediction of the safety of the RSRM nozzle. This paper documents the use of this newly developed multi-axial, temperature, and time (MATT) dependent failure model for modeling failure for the adhesives TIGA 321, EA913NA, and EA946. The development of the mathematical failure model using constant load rate normal and shear test data is presented. Verification of the accuracy of the failure model is shown through comparisons between predictions and measured creep and multi-axial failure data. The verification indicates that the failure model performs well for a wide range of conditions (loading, temperature, and time) for the three adhesives. The failure criterion is shown to be accurate through the glass transition for the adhesive EA946. Though this failure model has been developed and evaluated with adhesives, the concepts are applicable for other isotropic materials.

Richardson, D. E.

Decomposition-Based Failure Mode Identification Method for Risk-Free Design of Large Systems

When designing products, it is crucial to assure failure and risk-free operation in the intended operating environment. Failures are typically studied and eliminated as much as possible during the early stages of design. The few failures that go undetected result in unacceptable damage and losses in high-risk applications where public safety is of concern. Published NASA and NTSB accident reports point to a variety of components identified as sources of failures in the reported cases. In previous work, data from these reports were processed and placed in matrix form for all the system components and failure modes encountered, and then manipulated using matrix methods to determine similarities between the different components and failure modes. In this paper, these matrices are represented in the form of a linear combination of failures modes, mathematically formed using Principal Components Analysis (PCA) decomposition. The PCA decomposition results in a low-dimensionality representation of all failure modes and components of interest, represented in a transformed coordinate system. Such a representation opens the way for efficient pattern analysis and prediction of failure modes with highest potential risks on the final product, rather than making decisions based on the large space of component and failure mode data. The mathematics of the proposed method are explained first using a simple example problem. The method is then applied to component failure data gathered from helicopter, accident reports to demonstrate its potential.

Tumer, Irem Y.

Direct Adaptive Control of Systems with Actuator Failures: State of the Art and Continuing Challenges

In this paper, the problem of controlling systems with failures and faults is introduced, and an overview of recent work on direct adaptive control for compensation of uncertain actuator failures is presented. Actuator failures may be characterized by some unknown system inputs being stuck at some unknown (fixed or varying) values at unknown time instants, that cannot be influenced by the control signals. The key task of adaptive compensation is to design the control signals in such a manner that the remaining actuators can automatically and seamlessly take over for the failed ones, and achieve desired stability and asymptotic tracking. A certain degree of redundancy is necessary to accomplish failure compensation. The objective of adaptive control design is to effectively use the available actuation redundancy to handle failures without the knowledge of the failure patterns, parameters, and time of occurrence. This is a challenging problem because failures introduce large uncertainties in the dynamic structure of the system, in addition to parametric uncertainties and unknown disturbances. The paper addresses some theoretical issues in adaptive actuator failure compensation: actuator failure modeling, redundant actuation requirements, plant-model matching, error system dynamics, adaptation laws, and stability, tracking, and performance analysis. Adaptive control designs can be shown to effectively handle uncertain actuator failures without explicit failure detection. Some open technical challenges and research problems in this important research area are discussed.

Tao, Gang

Spacecraft Parachute Recovery System Testing from a Failure Rate Perspective

Spacecraft parachute recovery systems, especially those with a parachute cluster, require testing to identify and reduce failures. This is especially important when the spacecraft in question is human-rated. Due to the recent effort to make spaceflight affordable, the importance of determining a minimum requirement for testing has increased. The number of tests required to achieve a mature design, with a relatively constant failure rate, can be estimated from a review of previous complex spacecraft recovery systems. Examination of the Apollo parachute testing and the Shuttle Solid Rocket Booster recovery chute system operation will clarify at which point in those programs the system reached maturity. This examination will also clarify the risks inherent in not performing a sufficient number of tests prior to operation with humans on-board. When looking at complex parachute systems used in spaceflight landing systems, a pattern begins to emerge regarding the need for a minimum amount of testing required to wring out the failure modes and reduce the failure rate of the parachute system to an acceptable level for human spaceflight. Not only a sufficient number of system level testing, but also the ability to update the design as failure modes are found is required to drive the failure rate of the system down to an acceptable level. In addition, sufficient data and images are necessary to identify incipient failure modes or to identify failure causes when a system failure occurs. In order to demonstrate the need for sufficient system level testing prior to an acceptable failure rate, the Apollo Earth Landing System (ELS) test program and the Shuttle Solid Rocket Booster Recovery System failure history will be examined, as well as some experiences in the Orion Capsule Parachute Assembly System will be noted.

Stewart, Christine E.

Carbon Fiber Strand Tensile Failure Dynamic Event Characterization

There are few if any clear, visual, and detailed images of carbon fiber strand failures under tension useful for determining mechanisms, sequences of events, different types of failure modes, etc. available to researchers. This makes discussion of physics of failure difficult. It was also desired to find out whether the test article-to-test rig interface (grip) played a part in some failures. These failures have nothing to do with stress rupture failure, thus representing a source of waste for the larger 13-00912 investigation into that specific failure type. Being able to identify or mitigate any competing failure modes would improve the value of the 13-00912 test data. The beginnings of the solution to these problems lay in obtaining images of strand failures useful for understanding physics of failure and the events leading up to failure. Necessary steps include identifying imaging techniques that result in useful data, using those techniques to home in on where in a strand and when in the sequence of events one should obtain imaging data.

Johnson, Kenneth L.

Achieving Improved Reliability with Failure Analysis

Reliability is the ability of a product to properly function, within specified performance limits, for a specified period of time, under the life cycle application conditions. Failure analysis is a vital tool in the effort to ensure reliability of electronic products and systems throughout their product lifecycle. Today, organizations involved in activities within the electronics supply chain are facing new challenges, not just from complex assembly styles, harsher lifecycle environments, and sophisticated supply chains, but also from customers who are demanding a quicker turn-around. Unfortunately, root cause failure analysis is often performed incompletely, leading to a poor understanding of failure mechanisms and causes and, customer dissatisfaction due to recurring failures. The PDC starts with an introduction to reliability concepts, physics of failure and an overview of failure mechanisms that affect PCBs, PCBAs and components. The PDC then dives into root cause hypothesizing techniques (Pareto, FMEA, fishbone, FTA), non-destructive and destructive analysis and, materials characterization will be discussed. Numerous failure analysis case studies will be used to illustrate the techniques and analysis principles to arrive at the root cause(s) of field failures on printed circuit boards, active components, and assemblies. What Will You Learn: Topics include: Overview of Reliability Concepts Failure mechanisms of electronic products Root cause analysis Failure analysis techniques -Non-destructive techniques (optical, CSAM etc.) -Destructive analysis (DPA, Decap, FIB etc.) -Materials characterization (XRF, EDS, TMA/DSC etc.) Who Will Benefit: Reliability engineers, failure analysis engineers, engineering managers, design engineers, component engineers, quality assurance functions and, personnel involved with reliability activities within their company.

non-destructive techniques

Achieving Improved Reliability with Failure Analysis

Reliability is the ability of a product to properly function, within specified performance limits, for a specified period of time, under the life cycle application conditions. Failure analysis is a vital tool in the effort to ensure reliability of electronic products and systems throughout their product lifecycle. Today, organizations involved in activities within the electronics supply chain are facing new challenges, not just from complex assembly styles, harsher lifecycle environments, and sophisticated supply chains, but also from customers who are demanding a quicker turn-around. Unfortunately, root cause failure analysis is often performed incompletely, leading to a poor understanding of failure mechanisms and causes and, customer dissatisfaction due to recurring failures. The PDC (Professional Development Course) starts with an introduction to reliability concepts, physics of failure and an overview of failure mechanisms that affect PCBs (Printed Circuit Boards), PCBAs (Printed Circuit Board Assembly) and components. The PDC then dives into root cause hypothesizing techniques (Pareto, FMEA (Failure Modes and Effects Analysis), fishbone (Cause-And-Effect Diagram), FTA (Fault Tree Analysis)), non-destructive and destructive analysis and, materials characterization will be discussed. Numerous failure analysis case studies will be used to illustrate the techniques and analysis principles to arrive at the root cause(s) of field failures on printed circuit boards, active components, and assemblies. What Attendees will Learn: Topics include: Overview of Reliability Concepts Failure mechanisms of electronic products Root cause analysis Failure analysis techniques -Non-destructive techniques (optical, CSAM (Confocal Scanning Electron Microscopy) etc.) -Destructive analysis (DPA (Destructive Physical Analysis), Decap (Decapsulation), FIB (Focused Ion Beam) etc.) -Materials characterization (XRF (X-Ray Fluorescence) , EDS (Error Detection Sequential), TMA/DSC (Thermal Mechanical Analysis/Differential Scanning Calorimetry) etc.)

PCB quality

Knowledge Discovery for Early Failure Assessment of Complex Engineered Systems Using Natural Language Processing

Emerging complex engineered systems may have unexpected safety issues due to novel operational environments, increasing autonomy, human-machine interaction, and other factors. To prevent failures in operation or testing that necessitate costly redesign, it is desirable to predict likely failure modes early in the design process. Information about past engineering failures in natural language format presents one possible solution by enabling the retrieval of information that can inform new designs. However, identifying documents containing usable information and extracting the required information can be prohibitively time-consuming when implemented at scale. In this research, an automated natural language processing (NLP) framework is proposed to discover relevant knowledge from documents containing failure-related design information. The framework is applied to NASA’s Lessons Learned Information System (LLIS),which is publicly available. Documents containing usable information are filtered using two different NLP-based models. Next, from the identified usable documents, a failure taxonomy is extracted using a partitioned hierarchical topic modeling approach. Partitions of the document describe different sections of the failure taxonomy – i.e., failure, cause of failure, and recommendations – as indicated by the structure of the original document. The extracted failure taxonomy can be leveraged in early design failure assessment methods. Moreover, the framework can be used to identify documents containing usable failure-related design information from other databases and extract relevant information from these documents.

Documentation and Information Science

Knowledge Discovery for Early Failure Assessment of Complex Engineered Systems Using Natural Language Processing

Emerging complex engineered systems may have unexpected safety issues due to novel operational environments, increasing autonomy, human-machine interaction, and other factors. To prevent failures in operation or testing that necessitate costly redesign, it is desirable to predict likely failure modes early in the design process. Information about past engineering failures in natural language format presents one possible solution by enabling the retrieval of information that can inform new designs. However, identifying documents containing usable information and extracting the required information can be prohibitively time-consuming when implemented at scale. In this research, an automated natural language processing (NLP) framework is proposed to discover relevant knowledge from documents containing failure-related design information. The framework is applied to NASA’s Lessons Learned Information System (LLIS),which is publicly available. Documents containing usable information are filtered using two different NLP-based models. Next, from the identified usable documents, a failure taxonomy is extracted using a partitioned hierarchical topic modeling approach. Partitions of the document describe different sections of the failure taxonomy – i.e., failure, cause of failure, and recommendations – as indicated by the structure of the original document. The extracted failure taxonomy can be leveraged in early design failure assessment methods. Moreover, the framework can be used to identify documents containing usable failure-related design information from other databases and extract relevant information from these documents.

Documentation and Information Science

Localization of Ad-Hoc Lunar Constellations in Communication Failure Modes for Distributed Spacecraft Autonomy

As lunar missions increase in complexity inspired by NASA’s Artemis Program, they will require reliable and sufficient capability of the Position, Navigation, and Timing (PNT) system to support their scientific objectives. In addition, NASA's Commercial Lunar Payload Services (CLPS) program initiates the proliferation of public and private exploration partnerships using small satellites from commercial and private organizations, expanding traditionally confined low Earth orbit to be used for missions beyond geosynchronous orbit (Zucherman et al., 2022). Therefore, the Lunar PNT system is also required to provide navigation services compatible with the smaller platforms being sent by the public and private sectors, like CubeSats. However, traditional approaches to deep space missions’ navigation based on ground radio facilities have difficulties in providing sufficient support for the increasing number of users and communication at a distance from the Earth (Kaplev et al., 2022). In particular, the existing Lunar navigation technologies such as weak signal global positioning system (GPS) and deep space network (DSN) are not able to ensure operations of the upcoming small-scale Lunar missions due to their limitations in localization performance as well as capacity aspects. Another way to provide Lunar PNT service is to create a dedicated Lunar global navigation satellite system (GNSS) constellation, like GNSS systems on Earth. Space agencies like NASA, ESA, and JAXA are now developing the lunar communications relay and navigation systems (LCRNS) and Lunar navigation satellite systems (LNSS). In their systems, satellites will be deployed in moon orbits to provide the communication, positioning, navigation, and timing (CPNT) service at the lunar south pole region where the Artemis base camp will be expected (Murata et al., 2022). Meanwhile, common challenges considered in lunar PNT research arise from poor geometry of the terrestrial GNSS satellites when seen from the lunar user, highly perturbed lunar orbits, and limitations in power, size, and cost of the equipment on lunar satellites (Iiyama et al., 2023). It is also not clear if there will be enough Lunar users to support the cost and resources this would require as the Low-cost surface missions may not be able to support the large power, mass, and weight requirements that these navigation solutions entail (Niemoeller et al., 2022). As an alternative, existing Lunar science and exploration assets could be used to create a low-cost, autonomous, ad-hoc, and on-demand mission-centric Lunar PNT swarm capable of providing PNT services to these low-cost lunar missions (Hagenau et al., 2021). Introducing the non-dedicated and ad-hoc Lunar navigation constellation gives a way to provide PNT services on-demand. The non-dedicated swarm assets of Lunar constellations are designed to localize themselves with minimal interaction with Earth by adding cooperative autonomous localization to lunar missions, freeing up valuable bandwidth and ground segment resources. An autonomous localization of Lunar constellations is based on the concept of the decentralized PNT system with a distributed extended Kalman filter (DEKF) approach to state estimation for minimal onboard operating costs. In the distributed data processing algorithm, computation is broken down and assigned to each satellite, resulting in a considerably decreased computational amount while maintaining the accuracy of the orbit ephemeris and clock offsets as the result of centralized data processing (Wen et al., 2019). The DEKF requires spacecraft to perform two-way ranging operations with each other to communicate simultaneously, leveraging neighbor two-way intersatellite link (ISL) measurements such as pseudoranges to, and relative velocities between, visible satellites as sensor values (Frank et al., 2021). The Lunar autonomous PNT simulation (LAPS) demonstrated the feasibility of orbital asset localization among ad-hoc Lunar small-sat constellations based on the DEKF in Hagenau et al. (2021) and evaluated the matching algorithm proposed by Frank et al. (2021) in scheduling position estimation updates. In previous papers, all assets and measurements are assumed to be always available without consideration of the impact of intermittent and permanent communication failure. This study presents localization performance with increasing levels of network degradation for swarm assets and users to demonstrate the robustness of the decentralized Lunar PNT service in more realistic scenarios. Main issues arising from communication failure include spacecraft permanent or transient loss, antenna failures, message delays, etc. We tested four possible reasons for network degradation for 7 days in 21 satellites frozen with an altitude of 5500 km, evenly spaced around 3 circular, 40 inclination orbital planes where each spacecraft has two directional antennas. As anchor nodes with an independent estimate of their position are required in the DEKF approach, two ground nodes in each pole and one node in the gateway were implemented in the simulation. First, the most probable failure scenario involves the loss of a single spacecraft due to solar interference and technical malfunctions of the assets. Losing the availability of a single spacecraft means losing the two-way ISL measurement of the asset in the DEKF update. In order to provide the best possible quality of PNT service with limited time and resources, the distributed Lunar constellations must schedule the communication activities. The scheduler leverages mixed-integer linear programming (MILP) for the coordination and scheduling of the desired “as-needed” localization service (Niemoeller et al., 2022). We assume the scheduler has completely excluded the spacecraft information before the DEKF update in the failure scenario. When a random spacecraft has been turned off at a specific time, the robustness of the autonomous Lunar PNT system is evaluated. The simulation results give an 11.5% degradation in median position accuracy compared to the idealized performance excluding the asset loss. Second, a large number of assets may vanish due to major hardware problems or meteor strikes around the moon. A multiple spacecraft loss can degrade the localization performance very fast by losing the communication ability to do cross-plane measurements and in-plane measurements in a 3-plane constellation. When the matching-based scheduler is aware of ISL availability, we investigate a large number of in-plane and cross-plane asset vanishments both in close proximity and equally spaced throughout the orbital plane. According to the simulations, the loss of in-plane measurements gives 40.2% degradation while cross-plane measurements degrade 50.5% of asset localization performance among available assets. Therefore, it is concluded that cross-plane measurements are more important in improving the position estimation accuracy. Third, spacecraft failure information can be lost due to the internal message delay, resulting in the DEKF update scheduler to solve the matching problem with unavailable assets. The DEKF update cycle is comprised of network setup, communication, and computations where a global broadcast network and a 2-way ISL network setup take 6 minutes in total (Frank et al., 2021). Once the broadcast network successfully transmits and receives information, a random spacecraft may lose its availability right before solving the matching problem. This means the matching solution is no longer optimal, resulting in degradation in the localization performance. A numerical assessment shows the matching-based scheduler with knowing failure holds 11.5% of position accuracy degradation, whereas the scheduler without knowing failure gives 34% degraded localization performance without asset loss. Fourth, a transient loss of a single or multiple spacecraft may occur due to their antenna outages. After losing the two-way ISL availability for a few DEKF update cycles, the availability of spacecraft can easily be recovered as their states have been independently updated using measurements from anchor nodes. It is likely that the longer failure will result in worse localization performance. We have tested the transient failure of a random single asset for 30 min in the simulation, which is losing 3 update cycles in the DEKF system. From the simulation results, the position accuracy has been degraded to 4.84% which is better than the degraded localization performance of 11.5% from the permanent loss scenario among available assets. In conclusion, the autonomous Lunar PNT system based on the DEKF approach shows the ability to maintain resilience and robustness in the possible communication failure scenarios, ensuring that localization accuracy is preserved across various network degradation and outages. Future studies on investigating user localization performance near the South Pole and the broadcast network system will be continued in the following months.

Yeji Kim

Printed Circuit Board Inspection and Quality Control - PCB Failure Causes and Cures

This two day workshop will discuss a range of topics, including root cause analysis, physics-of-failure principles and failure mechanisms in printed circuit boards. Printed circuit boards (PCBs) are the baseline for electronics manufacturing upon which electronic components are mounted and formed into electronic systems. PCBs are used in a variety of electronic circuits from simple one-transistor amplifiers to large super computers. A PCB serves three main functions: 1) it provides the necessary mechanical support for the components in the circuit 2) it provides the necessary electrical interconnections, and 3) it bears some form of legend which identifies the components it carries. The failure modes on the PCBs can be categorized in a hierarchical structure, in which the mechanisms and causes are site or location dependant. Specimen preparation techniques, non-destructive and destructive analysis, and materials characterization will also be discussed. The first day of the workshop will present methodologies for identifying potential failure mechanisms in electronics based on the failure history and, systematic approaches to root cause analysis. The second day will cover failure analysis techniques geared towards various failure mechanisms, along with numerous component and PCB assembly failure analysis case studies that illustrate the techniques and analysis. Failure analysis case studies will be used to illustrate the techniques and analysis principles to arrive at the root cause(s) of field failures on printed circuit boards, active components, and assemblies.

Printed circuit boards

More Data Needed for Failure Rate Estimation, Validation, and Uncertainty Reduction

Current Environmental Control and Life Support System (ECLSS) development and test activities are not generating data fast enough to provide statistically-supportable precise Orbital Replacement Unit (ORU) failure rate estimates for future missions. Accurate and precise failure rate estimates are critical for missions beyond Low Earth Orbit (LEO) because current risk mitigation approaches – namely regular resupply and rapid abort capabilities – will not be available. Safe operations will depend on mission planners’ ability to accurately forecast spares demand and efficiently provide the necessary resources. However, even after more than a decade of operations on board the International Space Station (ISS), a significant amount of uncertainty remains in failure rate estimates. Uncertain or inaccurate failure rates result in increased risk and spares mass for future missions. A Bayesian failure rate estimation approach, such as the one currently implemented by the ISS Program, can help reduce uncertainty by incorporating engineering judgement into failure rate estimates. However, experience on the ISS and with other complex systems shows that these prior failure rate estimates are often inaccurate. In addition, prior failure rate estimates are typically point values; some level of uncertainty must be added to convert these into probability distributions for Bayesian updating, and there are several potential methods for doing so. Due to the low rate of data collection, these subjective (and often inaccurate) prior estimates currently have a strong influence on the end result. This paper examines the challenges associated with failure rate estimation, validation, and uncertainty reduction in the context of ECLSS development for beyond-LEO missions. A variety of techniques for generating and updating Bayesian priors are discussed and evaluated using both real-world and simulated data. Potential solutions for improving failure rate estimation, including testing additional units, are analyzed and discussed, and a set of recommendations are provided for next-generation system development activities.

Supportability

Discovery and Analysis of Rare High-Impact Failure Modes using Adversarial RL-Informed Sampling

Adaptive learning agents have tremendous potential to handle critical tasks currently performed by humans. Unfortunately, due to their complexity, it can be difficult to verify that these learning agents do not have critical failure modes. Standard verification and validation methods often do not apply directly to learning agents and Monte Carlo methods have difficulty covering even a small fraction of the state space, especially in multiagent systems or over long time horizons. To overcome this difficulty, we demonstrate an adaptive stress-testing method based on reinforcement learning of correlations that raise the probability of failure. This approach has three key properties: (1) it is able to find rare failure modes with far greater sample efficiency than Monte Carlo methods, (2) it can estimate the true probability of a failure mode despite the inherent bias in the learning method, and (3) it is capable of learning and resampling compact representations of multimodal failure spaces. These properties are important in practice as we need to find disparate failure modes while accounting for their actual relevance. This is a significant advantage over traditional adaptive stress testing methods that give abstract likelihoods of particular failure instances, but cannot estimate the probability of a broader failure mode. We test our algorithm on a simple problem from the aviation domain where an autonomous aircraft lands in gusty wind conditions. The results suggest that we can find failure modes with far fewer samples than the Monte Carlo approach and simultaneously estimate the probability of failure.

reinforcement learning

Correlation Between Weather Alerts and Grid Component Failures for Grid Alert

Weather events cause most grid failures. Often, we even get notifications on our phones to take cover or be prepared for an imminent event. If electric grid utilities had a similar warning that also included probable scenarios and the equipment involved, they could prepare and minimize the effects. Recent research at Idaho National Laboratory into electric grid risk analysis methods resulted in a tool that allows for the development of the most likely scenarios given failure probabilities of grid components. INL has a project with the U.S. Department of Energy’s Cybersecurity, Energy Security, and Emergency Response (CESER) program to develop a Grid Alert application that receives messages from the existing emergency alert system, filters and determines components possibly affected by the emergency event, calculates probable scenarios uses MASTERRI and then notifies the utility if there is significant risk. Historical failure data of elements that comprise the U.S. electric grid have been compiled by utilities and organizations such as the international regulatory body North American Electric Reliability Corporation (NERC). Nominal failure rates are obtained from this data. To make this tool possible, estimated failure rates are needed for different component types given the alert type, severity, and location. Historic weather-related grid element failures are correlated with historic weather events from Integrated Public Alert & Warning System (IPAWS). These correlated events and failures are used along with Bayesian updates from the historical norms to provide a modified failure rate for grid elements in the alert areas and calculate probable scenarios. This discusses the Grid Alert project plan but focuses on the data gathered and process used in determining failure rates for possible grid failure scenarios.

24 - POWER TRANSMISSION AND DISTRIBUTION

Sensor failure and multivariable control for airbreathing propulsion systems

A new sensor/actuator failure analysis technique for turbofan jet engines was developed. Three phases of failure analysis, namely detection, isolation, and accommodation are considered. Failure detection and isolation techniques are developed by utilizing the concept of Generalized Likelihood Ratio (GLR) tests. These techniques are applicable to both time varying and time invariant systems. Three GLR detectors are developed for: (1) hard-over sensor failure; (2) hard-over actuator failure; and (3) brief disturbances in the actuators. The probability distribution of the GLR detectors and the detectability of sensor/actuator failures are established. Failure type is determined by the maximum of the GLR detectors. Failure accommodation is accomplished by extending the Multivariable Nyquest Array (MNA) control design techniques to nonsquare system designs. The performance and effectiveness of the failure analysis technique are studied by applying the technique to a turbofan jet engine, namely the Quiet Clean Short Haul Experimental Engine (QCSEE). Single and multiple sensor/actuator failures in the QCSEE are simulated and analyzed and the effects of model degradation are studied.

Behbehani, K.