Engineering PapersSearch

SEARCH · Engineering Papers

Results for “failure cause”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Common Cause Failure Modeling

Common Cause Failures (CCFs) are a known and documented phenomenon that defeats system redundancy. CCFS are a set of dependent type of failures that can be caused by: system environments; manufacturing; transportation; storage; maintenance; and assembly, as examples. Since there are many factors that contribute to CCFs, the effects can be reduced, but they are difficult to eliminate entirely. Furthermore, failure databases sometimes fail to differentiate between independent and CCF (dependent) failure and data is limited, especially for launch vehicles. The Probabilistic Risk Assessment (PRA) of NASA's Safety and Mission Assurance Directorate at Marshall Space Flight Center (MFSC) is using generic data from the Nuclear Regulatory Commission's database of common cause failures at nuclear power plants to estimate CCF due to the lack of a more appropriate data source. There remains uncertainty in the actual magnitude of the common cause risk estimates for different systems at this stage of the design. Given the limited data about launch vehicle CCF and that launch vehicles are a highly redundant system by design, it is important to make design decisions to account for a range of values for independent and CCFs. When investigating the design of the one-out-of-two component redundant system for launch vehicles, a response surface was constructed to represent the impact of the independent failure rate versus a common cause beta factor effect on a system's failure probability. This presentation will define a CCF and review estimation calculations. It gives a summary of reduction methodologies and a review of examples of historical CCFs. Finally, it presents the response surface and discusses the results of the different CCFs on the reliability of a one-out-of-two system.

Hark, Frank

Common Cause Failure Modeling

Common Cause Failures (CCFs) are a known and documented phenomenon that defeats system redundancy. CCFS are a set of dependent type of failures that can be caused by: system environments; manufacturing; transportation; storage; maintenance; and assembly, as examples. Since there are many factors that contribute to CCFs, the effects can be reduced, but they are difficult to eliminate entirely. Furthermore, failure databases sometimes fail to differentiate between independent and CCF (dependent) failure and data is limited, especially for launch vehicles. The Probabilistic Risk Assessment (PRA) of NASA's Safety and Mission Assurance Directorate at Marshal Space Flight Center (MFSC) is using generic data from the Nuclear Regulatory Commission's database of common cause failures at nuclear power plants to estimate CCF due to the lack of a more appropriate data source. There remains uncertainty in the actual magnitude of the common cause risk estimates for different systems at this stage of the design. Given the limited data about launch vehicle CCF and that launch vehicles are a highly redundant system by design, it is important to make design decisions to account for a range of values for independent and CCFs. When investigating the design of the one-out-of-two component redundant system for launch vehicles, a response surface was constructed to represent the impact of the independent failure rate versus a common cause beta factor effect on a system's failure probability. This presentation will define a CCF and review estimation calculations. It gives a summary of reduction methodologies and a review of examples of historical CCFs. Finally, it presents the response surface and discusses the results of the different CCFs on the reliability of a one-out-of-two system.

Hark, Frank

Common Cause Failures Dominate and Defeat Redundancy

Common cause failures occur when several malfunctions are produced by a single event or process. They are especially damaging when they eliminate an entire set of redundant systems and disable their intended function. Redundancy is used when the individual system failure probability is unacceptably high. Redundancy can improve the overall system failure probability if the failures are independent, but the reliability gain is limited if there are dependent failures having a common cause. No amount of redundancy can reduce the total failure probability below the common cause failure probability. Common cause failures defeat redundancy. Systems with high reliability requirements often use extensive redundancy. These highly redundant systems rarely fail unless all the redundant components providing a particular function fail. Complete failures of such highly redundant systems are then usually common cause failures. Common cause failures are prevalent in highly redundant, high reliability systems. Common cause failures dominate redundancy. Redundant systems may fail due to specification, design, manufacturing, operations, or maintenance problems that disable all the identical redundant systems. Common cause failures typically account for one tenth of all failures. If the failure probability is relatively low and common cause failures are significant, adding more than two or three redundant identical units usually gives little added reliability improvement. Common cause failures can be reduced by using diverse components with different technologies and manufacturers, by separating and shielding subsystems, and by avoiding shared control, power, or location. External events and shared vulnerabilities may still cause common cause failures.

common cause failures

Common Cause Failures Dominate and Defeat Redundancy

Common cause failures occur when several malfunctions are produced by a single event or process. They are especially damaging when they eliminate an entire set of redundant systems and disable their intended function. Redundancy is used when the individual system failure probability is unacceptably high. Redundancy can improve the overall system failure probability if the failures are independent, but the reliability gain is limited if there are dependent failures having a common cause. No amount of redundancy can reduce the total failure probability below the common cause failure probability. Common cause failures defeat redundancy. Systems with high reliability requirements often use extensive redundancy. These highly redundant systems rarely fail unless all the redundant components providing a particular function fail. Complete failures of such highly redundant systems are then usually common cause failures. Common cause failures are prevalent in highly redundant, high reliability systems. Common cause failures dominate redundancy. Redundant systems may fail due to specification, design, manufacturing, operations, or maintenance problems that disable all the identical redundant systems. Common cause failures typically account for one tenth of all failures. If the failure probability is relatively low and common cause failures are significant, adding more than two or three redundant identical units usually gives little added reliability improvement. Common cause failures can be reduced by using diverse components with different technologies and manufacturers, by separating and shielding subsystems, and by avoiding shared control, power, or location. External events and shared vulnerabilities may still cause common cause failures.

common cause failures

Common Cause Failures and Ultra Reliability

A common cause failure occurs when several failures have the same origin. Common cause failures are either common event failures, where the cause is a single external event, or common mode failures, where two systems fail in the same way for the same reason. Common mode failures can occur at different times because of a design defect or a repeated external event. Common event failures reduce the reliability of on-line redundant systems but not of systems using off-line spare parts. Common mode failures reduce the dependability of systems using off-line spare parts and on-line redundancy.

reliability

Common Cause Failure Modeling in Space Launch Vehicles

Common Cause Failures (CCFs) are a known and documented phenomenon that defeats system redundancy. CCFs are a set of dependent type of failures that can be caused for example by system environments, manufacturing, transportation, storage, maintenance, and assembly. Since there are many factors that contribute to CCFs, they can be reduced, but are difficult to eliminate entirely. Furthermore, failure databases sometimes fail to differentiate between independent and dependent CCF. Because common cause failure data is limited in the aerospace industry, the Probabilistic Risk Assessment (PRA) Team at Bastion Technology Inc. is estimating CCF risk using generic data collected by the Nuclear Regulatory Commission (NRC). Consequently, common cause risk estimates based on this database, when applied to other industry applications, are highly uncertain. Therefore, it is important to account for a range of values for independent and CCF risk and to communicate the uncertainty to decision makers. There is an existing methodology for reducing CCF risk during design, which includes a checklist of 40+ factors grouped into eight categories. Using this checklist, an approach to produce a beta factor estimate is being investigated that quantitatively relates these factors. In this example, the checklist will be tailored to space launch vehicles, a quantitative approach will be described, and an example of the method will be presented.

Hark, Frank

A systems engineering approach to automated failure cause diagnosis in space power systems

Automatic failure-cause diagnosis is a key element in autonomous operation of space power systems such as Space Station's. A rule-based diagnostic system has been developed for determining the cause of degraded performance. The knowledge required for such diagnosis is elicited from the system engineering process by using traditional failure analysis techniques. Symptoms, failures, causes, and detector information are represented with structured data; and diagnostic procedural knowledge is represented with rules. Detected symptoms instantiate failure modes and possible causes consistent with currently held beliefs about the likelihood of the cause. A diagnosis concludes with an explanation of the observed symptoms in terms of a chain of possible causes and subcauses.

Dolce, James L.

Common Cause Failure Modeling: Aerospace Versus Nuclear

Aggregate nuclear plant failure data is used to produce generic common-cause factors that are specifically for use in the common-cause failure models of NUREG/CR-5485. Furthermore, the models presented in NUREG/CR-5485 are specifically designed to incorporate two significantly distinct assumptions about the methods of surveillance testing from whence this aggregate failure data came. What are the implications of using these NUREG generic factors to model the common-cause failures of aerospace systems? Herein, the implications of using the NUREG generic factors in the modeling of aerospace systems are investigated in detail and strong recommendations for modeling the common-cause failures of aerospace systems are given.

Stott, James E.

Modeling Common Cause Failures of Thrusters on ISS Visiting Vehicles

This paper discusses the methodology used to model common cause failures of thrusters on the International Space Station (ISS) Visiting Vehicles. The ISS Visiting Vehicles each have as many as 32 thrusters, whose redundancy makes them susceptible to common cause failures. The Global Alpha Model (as described in NUREG/CR‐5485) can be used to represent the system common cause contribution, but NUREG/CR‐5496 supplies global alpha parameters for groups only up to size six. Because of the large number of redundant thrusters on each vehicle, regression is used to determine parameter values for groups of size larger than six. An additional challenge is that Visiting Vehicle thruster failures must occur in specific combinations in order to fail the propulsion system; not all failure groups of a certain size are critical.

Haught, Megan

Modeling Common Cause Failures of Thrusters on ISS Visiting Vehicles

This paper discusses the methodology used to model common cause failures of thrusters on the International Space Station (ISS) Visiting Vehicles. The ISS Visiting Vehicles each have as many as 32 thrusters, whose redundancy and similar design make them susceptible to common cause failures. The Global Alpha Model (as described in NUREG/CR-5485) can be used to represent the system common cause contribution, but NUREG/CR-5496 supplies global alpha parameters for groups only up to size six. Because of the large number of redundant thrusters on each vehicle, regression is used to determine parameter values for groups of size larger than six. An additional challenge is that Visiting Vehicle thruster failures must occur in specific combinations in order to fail the propulsion system; not all failure groups of a certain size are critical.

Haught, Megan

Common Cause Failure Modeling

Space Launch System (SLS) Agenda: Objective; Key Definitions; Calculating Common Cause; Examples; Defense against Common Cause; Impact of varied Common Cause Failure (CCF) and abortability; Response Surface for various CCF Beta; Takeaways.

Hark, Frank

Common Cause Failure Modes

High technology industries with high failure costs commonly use redundancy as a means to reduce risk. Redundant systems, whether similar or dissimilar, are susceptible to Common Cause Failures (CCF). CCF is not always considered in the design effort and, therefore, can be a major threat to success. There are several aspects to CCF which must be understood to perform an analysis which will find hidden issues that may negate redundancy. This paper will provide definition, types, a list of possible causes and some examples of CCF. Requirements and designs from NASA projects will be used in the paper as examples.

Wetherholt, Jon

Printed Circuit Board Inspection and Quality Control - PCB Failure Causes and Cures

This two day workshop will discuss a range of topics, including root cause analysis, physics-of-failure principles and failure mechanisms in printed circuit boards. Printed circuit boards (PCBs) are the baseline for electronics manufacturing upon which electronic components are mounted and formed into electronic systems. PCBs are used in a variety of electronic circuits from simple one-transistor amplifiers to large super computers. A PCB serves three main functions: 1) it provides the necessary mechanical support for the components in the circuit 2) it provides the necessary electrical interconnections, and 3) it bears some form of legend which identifies the components it carries. The failure modes on the PCBs can be categorized in a hierarchical structure, in which the mechanisms and causes are site or location dependant. Specimen preparation techniques, non-destructive and destructive analysis, and materials characterization will also be discussed. The first day of the workshop will present methodologies for identifying potential failure mechanisms in electronics based on the failure history and, systematic approaches to root cause analysis. The second day will cover failure analysis techniques geared towards various failure mechanisms, along with numerous component and PCB assembly failure analysis case studies that illustrate the techniques and analysis. Failure analysis case studies will be used to illustrate the techniques and analysis principles to arrive at the root cause(s) of field failures on printed circuit boards, active components, and assemblies.

Printed circuit boards

Pyrotechnic system failures: Causes and prevention

Although pyrotechnics have successfully accomplished many critical mechanical spacecraft functions, such as ignition, severance, jettisoning and valving (excluding propulsion), failures continue to occur. Provided is a listing of 84 failures of pyrotechnic hardware with completed design over a 23-year period, compiled informally by experts from every NASA Center, as well as the Air Force Space Division and the Naval Surface Warfare Center. Analyses are presented as to when and where these failures occurred, their technical source or cause, followed by the reasons why and how these kinds of failures persist. The major contributor is a fundamental lack of understanding of the functional mechanisms of pyrotechnic devices and systems, followed by not recognizing pyrotechnics as an engineering technology, insufficient manpower with hands-on experience, too few test facilities, and inadequate guidelines and specifications for design, development, qualification and acceptance. Recommendations are made on both a managerial and technical basis to prevent failures, increase reliability, improve existing and future designs, and develop the technology to meet future requirements.

Bement, Laurence J.

Natural Language Processing Techniques for Intelligent Knowledge Management of Safety Reports

Safety, failure, and incident reports are common artifacts across various domains, including aviation and wildfire response. These reports are often mandatory to submit, resulting in the culmination of large repositories of text-based documents. Simultaneously, these reports and corresponding repositories are often only manually analyzed and queried by users via out-of-date search engines. As a consequence, we have been developing the Manager for Intelligent Knowledge Access (MIKA) toolkit, which uses natural language processing to improve information access and reuse. In this presentation, we discuss natural language processing techniques for knowledge discovery and apply these methods to a repository of aerial wildfire mishap reports. Two methods are used for knowledge discovery: topic modeling and named-entity recognition. We use topic modeling to identify hazards and perform a trend analysis to produce a data-driven risk matrix. A custom named-entity recognition model, build from fine tuning a pre-trained language model, is used to identify failure modes, failure causes, failure effects, control processes, and recommendations to aid in failure modes and effects analysis (FMEA). Throughout the presentation, we discuss and apply natural language processing techniques to better leverage the vast amount of information contained in report repositories.

Machine learning

Analysis of strain gage reliability in F-100 jet engine testing at NASA Lewis Research Center

A reliability analysis was performed on 64 strain gage systems mounted on the 3 rotor stages of the fan of a YF-100 engine. The strain gages were used in a 65 hour fan flutter research program which included about 5 hours of blade flutter. The analysis was part of a reliability improvement program. Eighty-four percent of the strain gages survived the test and performed satisfactorily. A post test analysis determined most failure causes. Five failures were caused by open circuits, three failed gages showed elevated circuit resistance, and one gage circuit was grounded. One failure was undetermined.

Holanda, R.