Engineering PapersSearch

SEARCH · Engineering Papers

Results for “failure cause”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A Markov chain model for reliability growth and decay

A mathematical model is developed to describe a complex system undergoing a sequence of trials in which there is interaction between the internal states of the system and the outcomes of the trials. For example, the model might describe a system undergoing testing that is redesigned after each failure. The basic assumptions for the model are that the state of the system after a trial depends probabilistically only on the state before the trial and on the outcome of the trial and that the outcome of a trial depends probabilistically only on the state of the system before the trial. It is shown that under these basic assumptions, the successive states form a Markov chain and the successive states and outcomes jointly form a Markov chain. General results are obtained for the transition probabilities, steady-state distributions, etc. A special case studied in detail describes a system that has two possible state ('repaired' and 'unrepaired') undergoing trials that have three possible outcomes ('inherent failure', 'assignable-cause' 'failure' and 'success'). For this model, the reliability function is computed explicitly and an optimal repair policy is obtained.

Siegrist, K.

An Approach to Automate tools for the Risk Assessment of Digital Instrumentation and Control Systems

Reliable digital instrumentation and control systems (DI&C) are integral for sustaining the continued operation of nuclear power plants. These systems ensure that nuclear reactors operate safely, efficiently, and within regulatory requirements. Yet, the cost of designing and licensing new nuclear DI&C can be prohibitively expensive. Under the U.S. Department of Energy Light Water Reactor Sustainability Program, Idaho National Laboratory has developed a framework for supporting the risk-informed design of DI&C systems by offering methods to support the identification, quantification, and evaluation of risks for various DI&C design architectures. The framework indicates potential software failure modes and provides pathways for quantifying the potential for these software failures, including common cause failures. Using the framework’s systematic approach, challenges for assessing risks within new and existing nuclear DI&C systems can be reduced. Nevertheless, the current framework can be further improved using the convenience of automation. This paper introduces the development of Software for the Hazard Identification and Evaluation of Digital Systems (SHIELDS). SHIELDS is an engineering software package that enables the identification, elimination, and mitigation of potential risks and reduces the burden of deploying reliable DI&C systems. This work introduces plans and techniques to digitize and improve the manual risk assessment modules of the framework. These improvements will save time and increase the repeatability and usability of the framework, making it more accessible to a wider range of users. Ultimately, this introduces SHIELDS and how its modules support efficient development of safe and reliable DI&C systems.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN

Solving Component Structural Dynamic Failures Due to Extremely High Frequency Structural Response on the Space Shuttle Program

For many years, the capabilities to determine the root-cause failure of component failures have been limited to the analytical tools and the state of the art data acquisition systems. With this limited capability, many anomalies have been resolved by adding material to the design to increase robustness without the ability to determine if the design solution was satisfactory until after a series of expensive test programs were complete. The risk of failure and multiple design, test, and redesign cycles were high. During the Space Shuttle Program, many crack investigations in high energy density turbomachines, like the SSME turbopumps and high energy flows in the main propulsion system, have led to the discovery of numerous root-cause failures and anomalies due to the coexistences of acoustic forcing functions, structural natural modes, and a high energy excitation, such as an edge tone or shedding flow, leading the technical community to understand many of the primary contributors to extremely high frequency high cycle fatique fluid-structure interaction anomalies. These contributors have been identified using advanced analysis tools and verified using component and system tests during component ground tests, systems tests, and flight. The structural dynamics and fluid dynamics communities have developed a special sensitivity to the fluid-structure interaction problems and have been able to adjust and solve these problems in a time effective manner to meet budget and schedule deadlines of operational vehicle programs, such as the Space Shuttle Program over the years.

Frady, Greg

Reliability Growth Modeling and Testing

Reliability growth has been modelled as an exponential decline in the cumulative failure rate that continues indefinitely as long as testing continues. Contrary to this, most reliability growth data show a brief high initial failure rate due to infant mortality followed by a long period of constant low failure rate. A two part failure rate model with an initial exponential decline followed by a constant failure rate usually fits the data and provides a more realistic description of reliability growth. The reliability growth process consists of testing, experiencing failures, finding the failure causes, and redesigning the system to remove them. The cost of reliability growth increases with the number of inherent failure modes and the time needed for them to occur and be removed. The failure modes with the lower failure rates will tend to occur later, as their Mean Time Before Failure (MTBF) is the inverse of the failure rate. Reliability growth testing has diminishing returns, since it takes longer to find and remove the less probable failures.This paper first discusses the reliability bathtub curve and then explains that reliability growth is produced by testing, identifying failure causes, and designing to remove them. A simple model of reliability growth is introduced, with a brief group of early failures followed by a constant failure rate. The cumulative failure rate n(t)/t can decline as rapidly as1/t or t-1butdeclines more slowly if additiona lfailures occur. The 56-failure Crow data seti s used to demonstrate the two-phase model of reliability growth followed by a constant failure rate. 13 additional data sets are modeled, with 9 of the 14 data sets showing reliability growth approximately as n(t)/t =1/t or t-1and substantial final failure rates. The model fits most of the data sets, but 4of the 14 show no reliability growth. The reliability growth period typically includes six failures and extends one-quarter or half the total test time. As reliability growth testing continues, the cumulative failure rate should be tracked to estimate the reliability growth exponent and the final failure rate.

reliability growth modeling

Modeling Reliability Growth

Reliability growth has been modelled as an exponential decline in the cumulative failure rate that continues indefinitely as long as testing continues. Contrary to this, most reliability growth data show a brief high initial failure rate due to infant mortality followed by a long period of constant low failure rate. A two part failure rate model with an initial exponential decline followed by a constant failure rate usually fits the data and provides a more realistic description of reliability growth. The reliability growth process consists of testing, experiencing failures, finding the failure causes, and redesigning the system to remove them. The cost of reliability growth increases with the number of inherent failure modes and the time needed for them to occur and be removed. The failure modes with the lower failure rates will tend to occur later, as their Mean Time Before Failure (MTBF) is the inverse of the failure rate. Reliability growth testing has diminishing returns, since it takes longer to find and remove the less probable failures.This paper first discusses the reliability bathtub curve and then explains that reliability growth is produced by testing, identifying failure causes, and designing to remove them. A simple model of reliability growth is introduced, with a brief group of early failures followed by a constant failure rate. The cumulative failure rate n(t)/t can decline as rapidly as1/t or t-1butdeclines more slowly if additiona lfailures occur. The 56-failure Crow data seti s used to demonstrate the two-phase model of reliability growth followed by a constant failure rate. 13 additional data sets are modeled, with 9 of the 14 data sets showing reliability growth approximately as n(t)/t =1/t or t-1and substantial final failure rates. The model fits most of the data sets, but 4of the 14 show no reliability growth. The reliability growth period typically includes six failures and extends one-quarter or half the total test time. As reliability growth testing continues, the cumulative failure rate should be tracked to estimate the reliability growth exponent and the final failure rate.

reliability growth modeling

Diverse Redundant Systems for Reliable Space Life Support

Reliable life support systems are required for deep space missions. The probability of a fatal life support failure should be less than one in a thousand in a multi-year mission. It is far too expensive to develop a single system with such high reliability. Using three redundant units would require only that each have a failure probability of one in ten over the mission. Since the system development cost is inverse to the failure probability, this would cut cost by a factor of one hundred. Using replaceable subsystems instead of full systems would further cut cost. Using full sets of replaceable components improves reliability more than using complete systems as spares, since a set of components could repair many different failures instead of just one. Replaceable components would require more tools, space, and planning than full systems or replaceable subsystems. However, identical system redundancy cannot be relied on in practice. Common cause failures can disable all the identical redundant systems. Typical levels of common cause failures will defeat redundancy greater than two. Diverse redundant systems are required for reliable space life support. Three, four, or five diverse redundant systems could be needed for sufficient reliability. One system with lower level repair could be substituted for two diverse systems to save cost.

life support

On the Use of Resilience Models as Digital Twins for Operational Support and In time Decision Making

Human error is a major contributor to accidents and performance losses in complex engineered systems. If one examines these human error caused failures further, a specific cause, the lack of situation awareness, has dominated as a major cause of human errors that instigate latent or catastrophic failures in complex systems. Studies of aviation accidents involving major air carriers revealed that situation awareness was the root cause of around 90% of accidents involving pilot error. Another study explored offshore drilling accidents involving human error and found that 40% of accidents were directly attributed to the loss of situation awareness. Studies of human errors in other domains such as nuclear power, air traffic control, process industry, and advanced driving show that loss of SA was a root cause in a majority of the events. Situation awareness-related failures are not only common but also costly and fatal (e.g., Bhopal Gas Leak, Air France 447 Flight Crash). Thus, the concept of situation awareness has emerged as an important construct in human factors, resulting in numerous models and measurement methods to aid in promoting appropriate levels of situation awareness.

Lukman Irshad

Conical Seat Shut-Off Valve

A moveable valve for controlling flow of a pressurized working fluid was designed. This valve consists of a hollow, moveable floating piston pressed against a stationary solid seat, and can use the working fluid to seal the valve. This open/closed, novel valve is able to use metal-to-metal seats, without requiring seat sliding action; therefore there are no associated damaging effects. During use, existing standard high-pressure ball valve seats tend to become damaged during rotation of the ball. Additionally, forces acting on the ball and stem create large amounts of friction. The combination of these effects can lead to system failure. In an attempt to reduce damaging effects and seat failures, soft seats in the ball valve have been eliminated; however, the sliding action of the ball across the highly loaded seat still tends to scratch the seat, causing failure. Also, in order to operate, ball valves require the use of large actuators. Positioning the metal-to-metal seats requires more loading, which tends to increase the size of the required actuator, and can also lead to other failures in other areas such as the stem and bearing mechanisms, thus increasing cost and maintenance. This novel non-sliding seat surface valve allows metal-to-metal seats without the damaging effects that can lead to failure, and enables large seating forces without damaging the valve. Additionally, this valve design, even when used with large, high-pressure applications, does not require large conventional valve actuators and the valve stem itself is eliminated. Actuation is achieved with the use of a small, simple solenoid valve. This design also eliminates the need for many seals used with existing ball valve and globe valve designs, which commonly cause failure, too. This, coupled with the elimination of the valve stem and conventional valve actuator, improves valve reliability and seat life. Other mechanical liftoff seats have been designed; however, they have only resulted in increased cost, and incurred other reliability issues. With this novel design, the seat is lifted by simply removing the working fluid pressure that presses it against the seat and no external force is required. By eliminating variables associated with existing ball and globe configurations that can have damaging effects upon a valve, this novel design reduces downtime in rocket engine test schedules and maintenance costs.

Farner, Bruce

Current Emergency Locator Transmitter (ELT) deficiencies and potential improvements utilizing TSO-C91a ELTs

An analysis was conducted of current ELT problems and potential improvements that could be made by employing the TSO-C91a ELTs to replace the current TSO-C91 ELTs. The scope of the study included the following: (1) validate the problems; (2) determine specific failure causes; (3) determine false alarm causes; (4) estimate improvements from TSO-C91a; (5) estimate benefits from replacement of the current ELTs; and (6) determine need and benefits for improved ELT inspection and maintenance. A detailed comparison between the two requirements documents (TSO-C91 and -91a) was made to assess improved performance of the ELT in each category of failure cause and each cause of false alarms. The comparison and analysis resulted in projecting a success of operation rate approximately 3 times the current rate and a reduction in false alarms to 0.25 of those generated by TSO-C91 ELTs. These improvements led to a projection of benefits of approximately 25 additional lives to be saved each year with TSO-C91a ELTs and an improved inspection and maintenance program.

Trudell, Bernard J.

STS-3 main parachute failure

A failure analysis of the parachute on the Space Transportation System 3 flight's solid rocket booster's is presented. During the reentry phase of the two Solid Rocket Boosters (SRBs), one 115 ft diameter main parachute failed on the right hand SRB (A12). This parachute failure caused the SRB to impact the Ocean at 110 ft/sec in lieu of the expected 3 parachute impact velocity of 88 ft/sec. This higher impact velocity relates directly to more SRB aft skirt and more motor case damage. The cause of the parachute failure, the potential risks of losing an SRB as a result of this failure, and recommendations to ensure that the probability of chute failures of this type in the future will be low are discussed.

Runkle, R.

An investigation of the causes of failure of flexible thermal protection materials in an aerodynamic environment

Tests of small panels of advanced flexible reusable surface insulation (AFRSI) were conducted using a small wind tunnel that was designed to simulate Space Shuttle Orbiter entry mean-flow and pulsating aerodynamic loads. The wind tunnel, with a 3 inch wide by 1.75 inch high by 7.5 inch long test section, proved to be capable of continuous flow at dynamic pressures q near 580 psf with fluctuating pressures over 2 psi RMS at an excitation frequency f sub E of 200 Hz. For this investigation, however, the wind tunnel was used to test entry-temperature preconditioned and heat-cleaned AFRSI at q = 280 psf, Prms was nearly equal to 1.2 psi and f sub E = 200 Hz. The objective of these tests was to determine the mechanism of failure of AFRSI at Orbiter entry conditions. Details of the test apparatus and test results are presented.

Coe, Charles F.

Failure-Modes-And-Effects Analysis Of Software Logic

Rigorous analysis applied early in design effort. Method of identifying potential inadequacies and modes and effects of failures caused by inadequacies (failure-modes-and-effects analysis or "FMEA" for short) devised for application to software logic.

Garcia, Danny

Spacecraft Parachute Recovery System Testing from a Failure Rate Perspective

Spacecraft parachute recovery systems, especially those with a parachute cluster, require testing to identify and reduce failures. This is especially important when the spacecraft in question is human-rated. Due to the recent effort to make spaceflight affordable, the importance of determining a minimum requirement for testing has increased. The number of tests required to achieve a mature design, with a relatively constant failure rate, can be estimated from a review of previous complex spacecraft recovery systems. Examination of the Apollo parachute testing and the Shuttle Solid Rocket Booster recovery chute system operation will clarify at which point in those programs the system reached maturity. This examination will also clarify the risks inherent in not performing a sufficient number of tests prior to operation with humans on-board. When looking at complex parachute systems used in spaceflight landing systems, a pattern begins to emerge regarding the need for a minimum amount of testing required to wring out the failure modes and reduce the failure rate of the parachute system to an acceptable level for human spaceflight. Not only a sufficient number of system level testing, but also the ability to update the design as failure modes are found is required to drive the failure rate of the system down to an acceptable level. In addition, sufficient data and images are necessary to identify incipient failure modes or to identify failure causes when a system failure occurs. In order to demonstrate the need for sufficient system level testing prior to an acceptable failure rate, the Apollo Earth Landing System (ELS) test program and the Shuttle Solid Rocket Booster Recovery System failure history will be examined, as well as some experiences in the Orion Capsule Parachute Assembly System will be noted.

Stewart, Christine E.

Full Stall Simulations of a Redesigned Ventilation Fan for the ISS

The concept of a stall is studied rigorously in the aerospace industry. From a design standpoint, instabilities such as stall are undesirable– operation in the stall regime has a tremendous impact on aerodynamic performance as well as structural integrity. In extreme cases, operating in stall conditions can cause failure. Further, stall can cause a loss of lift on aircraft wings or a loss of thrust in aircraft engines. In any case, stall continues to be a topic of interest in the aerospace industry. The present work aims to analyze the stall characteristics of a ventilation fan that was recently designed for the International Space Station (ISS). Although the ventilation fan has a rotor-stator design, this paper considers a rotor-only configuration. The FUN3D Computational Fluid Dynamic (CFD) solver developed by NASA Langley Research Center was used to simulate the operational characteristics of the ventilation fan. FUN3D solves the Unsteady Reynolds-Averaged Naiver-Stokes (URANS) equations using implicit time marching and a dynamic overset grid. The FUN3D solver was originally written for exterior flow fields; however, this work represents an extension of the FUN3D solver to turbomachinery or interior flow fields. The FUN3D results for the rotor-only ventilation fan accurately captured the operational characteristics inherent to compressors– the results shared similar performance trends when compared to the experimental results for the rotor-stator case. A peak adiabatic efficiency of 96% occurred at a MFR of 105.2 CFM, the minimum aerodynamically stable point. A computationally stable stall occurred at a mass flow rate of 43.8 CFM where the adiabatic efficiency dropped to 69%.

CFD

Fault Management Algorithm Risk Assessment for the NASA Space Launch System

This paper presents the false positive (FP) and false negative (FN) risk assessment process currently being conducted for the Space Launch System (SLS) Artemis II Fault Management (FM) detection functions. The analysis scope, general assumptions and guide rules, and key modeling concepts were discussed to establish the basis of the risk assessments conducted. Initial analyses indicated a dominance in the total risk by software and firmware failures. This paper presents efforts applied to refine the software risks and the overall impact of implementing those modifications. Current analyses conducted on the detection functions implemented for the SLS Artemis II mission indicate primary risk drivers for the individual FM detection functions are flight software failures, firmware design failures, and hardware Common Cause Failures (CCFs). There still remains issues of how to account for time and redundancy in the software risk estimations.

probability risk analysis

Fault Management Algorithm Risk Assessment for the NASA Space Launch System

This presentation describes the false positive (FP) and false negative (FN) risk assessment process currently being conducted for the Space Launch System (SLS) Artemis II Fault Management (FM) detection functions. The analysis scope, general assumptions and guide rules, and key modeling concepts were discussed to establish the basis of the risk assessments conducted. Initial analyses indicated a dominance in the total risk by software and firmware failures. This paper presents efforts applied to refine the software risks and the overall impact of implementing those modifications. Current analyses conducted on the detection functions implemented for the SLS Artemis II mission indicate primary risk drivers for the individual FM detection functions are flight software failures, firmware design failures, and hardware Common Cause Failures (CCFs). There still remains issues of how to account for time and redundancy in the software risk estimations.

probability risk analysis

Achieving Improved Reliability with Failure Analysis

Reliability is the ability of a product to properly function, within specified performance limits, for a specified period of time, under the life cycle application conditions. Failure analysis is a vital tool in the effort to ensure reliability of electronic products and systems throughout their product lifecycle. Today, organizations involved in activities within the electronics supply chain are facing new challenges, not just from complex assembly styles, harsher lifecycle environments, and sophisticated supply chains, but also from customers who are demanding a quicker turn-around. Unfortunately, root cause failure analysis is often performed incompletely, leading to a poor understanding of failure mechanisms and causes and, customer dissatisfaction due to recurring failures. The PDC (Professional Development Course) starts with an introduction to reliability concepts, physics of failure and an overview of failure mechanisms that affect PCBs (Printed Circuit Boards), PCBAs (Printed Circuit Board Assembly) and components. The PDC then dives into root cause hypothesizing techniques (Pareto, FMEA (Failure Modes and Effects Analysis), fishbone (Cause-And-Effect Diagram), FTA (Fault Tree Analysis)), non-destructive and destructive analysis and, materials characterization will be discussed. Numerous failure analysis case studies will be used to illustrate the techniques and analysis principles to arrive at the root cause(s) of field failures on printed circuit boards, active components, and assemblies. What Attendees will Learn: Topics include: Overview of Reliability Concepts Failure mechanisms of electronic products Root cause analysis Failure analysis techniques -Non-destructive techniques (optical, CSAM (Confocal Scanning Electron Microscopy) etc.) -Destructive analysis (DPA (Destructive Physical Analysis), Decap (Decapsulation), FIB (Focused Ion Beam) etc.) -Materials characterization (XRF (X-Ray Fluorescence) , EDS (Error Detection Sequential), TMA/DSC (Thermal Mechanical Analysis/Differential Scanning Calorimetry) etc.)

PCB quality