Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Common Cause Failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Oversimplification of Systems Engineering Goals, Processes, and Criteria in NASA Space Life Support

This paper investigates the oversimplification of the inherently complex systems engineering process in space life support. The standard systems engineering process steps are described. The International Space Station (ISS) life support system is explained with its goals and performance criteria. Although it is not usually emphasized, the essential function of developing a hierarchy of systems and subsystems is to simplify the design process. The System Complexity Metric (SCM) shows how this di-vide-and-conquer approach also reduces the system complexity. The complete systems engineering process has many detailed steps. It is often simplified because of the effort required and the human limitations on working memory and decision span. Systems analysis demands slow, logical, and fo-cused thinking but is often bypassed in favor of quick, intuitive, subconscious “gut feel.” A study of 100 system designs found examples of 12 specific mental mistakes, such as ignoring stakeholder needs, and these mistakes are essentially oversimplifications of the systems engineering process. An analysis of space life support goals, options, criteria, and processes found 11 examples of oversimplifications in systems engineering, such as neglecting safety and cost. All these 11 oversimplifications could be traced to one or more of the 12 previously identified mental mistakes or other well-known ones, such as ig-noring sunk costs. Oversimplification of the systems engineering process is rarely noticed but is a common and harmful problem. A study of failures in 50 different space systems found that problems in systems engineering caused failures and often led to errors in design, development, and test that further contributed to failure. It seems that more diligent systems engineering could prevent many project problems and failures, but projects seem to be more guided by “gut feel” based on tradition, authority, and consensus than on the logical, rational systems engineering approach.

Simplified systems engineering↗

Common Cause Case Study: An Estimated Probability of Four Solid Rocket Booster Hold-Down Post Stud Hang-ups

Until Solid Rocket Motor ignition, the Space Shuttle is mated to the Mobil Launch Platform in part via eight (8) Solid Rocket Booster (SRB) hold-down bolts. The bolts are fractured using redundant pyrotechnics, and are designed to drop through a hold-down post on the Mobile Launch Platform before the Space Shuttle begins movement. The Space Shuttle program has experienced numerous failures where a bolt has hung up. That is, it did not clear the hold-down post before liftoff and was caught by the SRBs. This places an additional structural load on the vehicle that was not included in the original certification requirements. The Space Shuttle is currently being certified to withstand the loads induced by up to three (3) of eight (8) SRB hold-down experiencing a "hang-up". The results of loads analyses performed for (4) stud hang-ups indicate that the internal vehicle loads exceed current structural certification limits at several locations. To determine the risk to the vehicle from four (4) stud hang-ups, the likelihood of the scenario occurring must first be evaluated. Prior to the analysis discussed in this paper, the likelihood of occurrence had been estimated assuming that the stud hang-ups were completely independent events. That is, it was assumed that no common causes or factors existed between the individual stud hang-up events. A review of the data associated with the hang-up events, showed that a common factor (timing skew) was present. This paper summarizes a revised likelihood evaluation performed for the four (4) stud hang-ups case considering that there are common factors associated with the stud hang-ups. The results show that explicitly (i.e. not using standard common cause methodologies such as beta factor or Multiple Greek Letter modeling) taking into account the common factor of timing skew results in an increase in the estimated likelihood of four (4) stud hang-ups of an order of magnitude over the independent failure case.

Cross, Robert↗

Common Cause Case Study: An Estimated Probability of Four Solid Rocket Booster Hold-down Post Stud Hang-ups

Until Solid Rocket Motor ignition, the Space Shuttle is mated to the Mobil Launch Platform in part via eight (8) Solid Rocket Booster (SRB) hold-down bolts. The bolts are fractured using redundant pyrotechnics, and are designed to drop through a hold-down post on the Mobile Launch Platform before the Space Shuttle begins movement. The Space Shuttle program has experienced numerous failures where a bolt has "hung-up." That is, it did not clear the hold-down post before liftoff and was caught by the SRBs. This places an additional structural load on the vehicle that was not included in the original certification requirements. The Space Shuttle is currently being certified to withstand the loads induced by up to three (3) of eight (8) SRB hold-down post studs experiencing a "hang-up." The results af loads analyses performed for four (4) stud-hang ups indicate that the internal vehicle loads exceed current structural certification limits at several locations. To determine the risk to the vehicle from four (4) stud hang-ups, the likelihood of the scenario occurring must first be evaluated. Prior to the analysis discussed in this paper, the likelihood of occurrence had been estimated assuming that the stud hang-ups were completely independent events. That is, it was assumed that no common causes or factors existed between the individual stud hang-up events. A review of the data associated with the hang-up events, showed that a common factor (timing skew) was present. This paper summarizes a revised likelihood evaluation performed for the four (4) stud hang-ups case considering that there are common factors associated with the stud hang-ups. The results show that explicitly (i.e. not using standard common cause methodologies such as beta factor or Multiple Greek Letter modeling) taking into account the common factor of timing skew results in an increase in the estimated likelihood of four (4) stud hang-ups of an order of magnitude over the independent failure case.

Cross, Robert↗

What Reliability Engineers Should Know about Space Radiation Effects

Space radiation in space systems present unique failure modes and considerations for reliability engineers. Radiation effects is not a one size fits all field. Threat conditions that must be addressed for a given mission depend on the mission orbital profile, the technologies of parts used in critical functions and on application considerations, such as supply voltages, temperature, duty cycle, and redundancy. In general, the threats that must be addressed are of two types-the cumulative degradation mechanisms of total ionizing dose (TID) and displacement damage (DD). and the prompt responses of components to ionizing particles (protons and heavy ions) falling under the heading of single-event effects. Generally degradation mechanisms behave like wear-out mechanisms on any active components in a system: Total Ionizing Dose (TID) and Displacement Damage: (1) TID affects all active devices over time. Devices can fail either because of parametric shifts that prevent the device from fulfilling its application or due to device failures where the device stops functioning altogether. Since this failure mode varies from part to part and lot to lot, lot qualification testing with sufficient statistics is vital. Displacement damage failures are caused by the displacement of semiconductor atoms from their lattice positions. As with TID, failures can be either parametric or catastrophic, although parametric degradation is more common for displacement damage. Lot testing is critical not just to assure proper device fi.mctionality throughout the mission. It can also suggest remediation strategies when a device fails. This paper will look at these effects on a variety of devices in a variety of applications. This paper will look at these effects on a variety of devices in a variety of applications. (2) On the NEAR mission a functional failure was traced to a PIN diode failure caused by TID induced high leakage currents. NEAR was able to recover from the failure by reversing the current of a nearby Thermal Electric Cooler (turning the TEC into a heater). The elevated temperature caused the PIN diode to anneal and the device to recover. It was by lot qualification testing that NEAR knew the diode would recover when annealed. This paper will look at these effects on a variety of devices in a variety of applications. Single Event Effects (SEE): (1) In contrast to TID and displacement damage, Single Event Effects (SEE) resemble random failures. SEE modes can range from changes in device logic (single-event upset, or SEU). temporary disturbances (single-event transient) to catastrophic effects such as the destructive SEE modes, single-event latchup (SEL). single-event gate rupture (SEGR) and single-event burnout (SEB) (2) The consequences of nondestructive SEE modes such as SEU and SET depend critically on their application--and may range from trivial nuisance errors to catastrophic loss of mission. It is critical not just to ensure that potentially susceptible devices are well characterized for their susceptibility, but also to work with design engineers to understand the implications of each error mode. -For destructive SEE, the predominant risk mitigation strategy is to avoid susceptible parts, or if that is not possible. to avoid conditions under which the part may be susceptible. Destructive SEE mechanisms are often not well understood, and testing is slow and expensive, making rate prediction very challenging. (3) Because the consequences of radiation failure and degradation modes depend so critically on the application as well as the component technology, it is essential that radiation, component. design and system engineers work togetherpreferably starting early in the program to ensure critical applications are addressed in time to optimize the probability of mission success.

DiBari, Rebecca↗

A diagnosis system using object-oriented fault tree models

Spaceborne computing systems must provide reliable, continuous operation for extended periods. Due to weight, power, and volume constraints, these systems must manage resources very effectively. A fault diagnosis algorithm is described which enables fast and flexible diagnoses in the dynamic distributed computing environments planned for future space missions. The algorithm uses a knowledge base that is easily changed and updated to reflect current system status. Augmented fault trees represented in an object-oriented form provide deep system knowledge that is easy to access and revise as a system changes. Given such a fault tree, a set of failure events that have occurred, and a set of failure events that have not occurred, this diagnosis system uses forward and backward chaining to propagate causal and temporal information about other failure events in the system being diagnosed. Once the system has established temporal and causal constraints, it reasons backward from heuristically selected failure events to find a set of basic failure events which are a likely cause of the occurrence of the top failure event in the fault tree. The diagnosis system has been implemented in common LISP using Flavors.

Iverson, David L.↗

Phenomena associated with bench and thermal-vacuum testing of super conductors - Heat pipes.

Test failures of heat pipes occur when the functional performance is unable to match the expected design limits or when the power applied to the heat pipe (in the form of heat) is distributed unevenly through the system, yielding a large thermal gradient. When a thermal gradient larger than expected is measured, it normally occurs in the evaporator or condenser sections of the pipe. Common causes include evaporator overheating, condenser dropout, noncondensable gas formation, surge and partial recovery of evaporator temperatures, masking of thermal profiles, and simple malfunctions due to leaks and mechanical failures or flaws. Examples of each of these phenomena are described along with corresponding failure analyses and corrective measures.

Marshburn, J. P.↗

Detecting and Characterizing Patterns of Failure in Complex Engineered Systems: an Ontology Development and Clustering Approach

While the causes of failures in complex engineered systems are often clear in hindsight, it can be challenging to predict failures proactively during the design of novel engineered products or systems. Identifying patterns can be useful for capturing common characteristics that may lead to failure. In this paper, we present a methodology for identifying patterns of failure from NASA’s publicly available Lessons Learned Information System (LLIS). We apply an ontology development and clustering approach to identify representative patterns leading to failures in historical lessons learned. A joint inductive-deductive approach reveals the key themes in lessons that lead to failure, which are formalized and recorded as an ontology of complex systems failure causes. Documents from the LLIS are manually tagged with relevant characteristics from the ontology. From the tagged set, clustering is used to capture co-occurring sets of characteristics that lead to failure. The primary contribution of this work is a method for extracting a set of generic failure patterns in complex engineered systems and characteristics for these patterns that can be identified at design time, knowledge of which can be used to plan mitigation strategies.

Systems Engineering↗

Detecting and Characterizing Patterns of Failure in Complex Systems: An Ontology Development and Clustering Approach

While the causes of failures in complex engineered systems are often clear in hindsight, it can be challenging to predict failures proactively during the design of novel engineered products or systems. Identifying patterns can be useful for capturing common characteristics that may lead to failure. In this paper, we present a methodology for identifying patterns of failure from NASA’s publicly available Lessons Learned Information System (LLIS). We apply an ontology development and clustering approach to identify representative patterns leading to failures in historical lessons learned. A joint inductive-deductive approach reveals the key themes in lessons that lead to failure, which are formalized and recorded as an ontology of complex systems failure causes. Documents from the LLIS are manually tagged with relevant characteristics from the ontology. From the tagged set, clustering is used to capture co-occurring sets of characteristics that lead to failure. The primary contribution of this work is a method for extracting a set of generic failure patterns in complex engineered systems and characteristics for these patterns that can be identified at design time, knowledge of which can be used to plan mitigation strategies.

Systems Engineering↗

Parts, Materials, and Processes Experience Summary

The ALERT program, a system for communicating common problems with parts, materials, and processes, is condensed and catalogued. Expanded information on selected topics is provided by relating the problem area (failure) to the cause, the investigations and findings, the suggestions for avoidance (inspections, screening tests, proper part applications), and failure analysis procedures. The basic objective of ALERT is the avoidance of the recurrence of parts, materials, and processed problems, thus improving the reliability of equipment produced for and used by the government.

Source record↗

Z-2 Threaded Insert Design and Testing

NASA's Z-2 prototype space suit contains several components fabricated from an advanced hybrid composite laminate consisting of IM10 carbon fiber and fiber glass. One requirement was to have removable, replaceable helicoil inserts to which other suit components would be fastened. An approach utilizing bonded in inserts with helicoils inside of them was implemented. During initial assembly, cracking sounds were heard followed by the lifting of one of the blind inserts out of its hole when the screws were torqued. A failure investigation was initiated to understand the mechanism of the failure. Ultimately, it was determined that the pre-tension caused by torqueing the fasteners is a much larger force than induced from the pressure loads of the suit which was not considered in the insert design. Bolt tension is determined by dividing the torque on the screw by a k value multiplied by the thread diameter of the bolt. The k value is a factor that accounts for friction in the system. A common value used for k for a non-lubricated screw is 0.2. The k value can go down by as much as 0.1 if the screw is lubricated which means for the same torque, a much larger tension could be placed on the bolt and insert. This paper summarizes the failure investigation that was performed to identify the root cause of the suit failure and details how the insert design was modified to resist a higher pull out tension.

Ross, Amy↗

Failure Modes Experienced on Spacecraft Nicd Batteries

A review was made of failures and irregularities experienced on nickel cadmium batteries for 31 spacecraft. Only rarely did batteries fail completely. In many cases, poorly performing batteries were compensated for by a reduction in loads or by continuing to operate in spite of out-of-voltage conditions. Low discharge voltage was the most common problem observed in flight spacecraft (42%). Spacecraft batteries are often designed to protect against cell shorts, but cell shorts accounted for only 16% of the failures. Other causes of problems were high charge voltage (16%), battery problems caused by other elements of the spacecraft (10%), and open circuit failures (6%). Problems of miscellaneous or unknown causes occurred in 10% of the cases.

Gross, S.↗

Multiversion software reliability through fault-avoidance and fault-tolerance

In this project we have proposed to investigate a number of experimental and theoretical issues associated with the practical use of multi-version software in providing dependable software through fault-avoidance and fault-elimination, as well as run-time tolerance of software faults. In the period reported here we have working on the following: We have continued collection of data on the relationships between software faults and reliability, and the coverage provided by the testing process as measured by different metrics (including data flow metrics). We continued work on software reliability estimation methods based on non-random sampling, and the relationship between software reliability and code coverage provided through testing. We have continued studying back-to-back testing as an efficient mechanism for removal of uncorrelated faults, and common-cause faults of variable span. We have also been studying back-to-back testing as a tool for improvement of the software change process, including regression testing. We continued investigating existing, and worked on formulation of new fault-tolerance models. In particular, we have partly finished evaluation of Consensus Voting in the presence of correlated failures, and are in the process of finishing evaluation of Consensus Recovery Block (CRB) under failure correlation. We find both approaches far superior to commonly employed fixed agreement number voting (usually majority voting). We have also finished a cost analysis of the CRB approach.

Vouk, Mladen A.↗

Improved Processing Techniques for Inclusion-Free Steel for Bearing and Mechanical Component Applications

High hardness, high carbide powder metallurgy tools steels such as M62 enable the operation of ball bearings at extremely high load and stress levels. Operation under such conditions increases the potential for rolling contact fatigue failure attributed to ceramic particle inclusions. To address this challenge, industry has sought steel made from ever increasing levels of cleanliness but the results have been uneven owing to the random nature of the occurrence of material flaws. One common approach is to rely upon careful ingot inspections prior to bearing manufacture. By selecting the cleanest portion of an ingot, it is expected that bearings relatively free from material flaws will result. This approach is not always successful because detrimental flaws that exist deep within an ingot can pass inspection undetected potentially causing subsequent failure. Recent efforts to commercialize an intermetallic material, 60NiTi, for rolling element bearings demonstrates a pathway to produce bearing steel that is free from unwanted ceramic particle inclusions. In this paper, the process used to make bearing grade ceramic-free NiTi alloys is described and applied to steelmaking. At its core, the NiTi process differs from steel making in one key aspect. NiTi alloys are made from elementally pure starting materials that are melted, blended and processed in equipment absolutely free from exposure to oxygen and ceramics ensuring a ceramic particle-free product. In contrast, the predominant method to make bearing steel is to employ a successive series of purification steps to reduce contamination levels below required thresholds. This paper describes the processes developed and applied to high carbide tool steel, M62. The resulting material and microstructures are evaluated and compared to M62 prepared by conventional powder metallurgy techniques. It is hoped that the application of materials manufacturing techniques used for fracture sensitive ceramics and intermetallic materials like NiTi can provide a pathway to Ultra-Clean, Ceramic-Inclusion free steels for rolling element bearings and other failure critical applications.

Steel↗

Electrochemical Impedance Spectroscopy of Alloys in a Simulated Space Shuttle Launch Environment

Type 304L stainless steel (304L SS) tubing is currently used in various supply lines that service the Orbiter at NASA's John F. Kennedy Space Center Launch Pads in Florida (USA). The atmosphere at the Space Shuffle launch site is very corrosive due to a combination of factors, such as the proximity of the Atlantic Ocean and the concentrated hydrochloric acid produced by the fuel combustion reaction in the solid rocket boosters. The acidic chloride environment is aggressive to most metals and causes severe pitting in many of the common stainless steel alloys such as 304L SS. Stainless steel tubing is susceptible to pitting corrosion that can cause cracking and rupture of both high-pressure gas and fluid systems. Outages in the systems where failures occur can impact the normal operation of the shuttle and launch schedules. The use of a more corrosion resistant tubing alloy for launch pad applications would greatly reduce the probability of failure, improve safety, lessen maintenance costs, and reduce downtime. A study which included ten alloys was undertaken to find a more corrosion resistant material to replace the existing 304L SS tubing. The study included atmospheric exposure at NASA's John F. Kennedy Space Center outdoor corrosion test site near the launch pads and electrochemical measurements in the laboratory which included DC techniques and electrochemical impedance spectroscopy (EIS). This paper presents the results from EIS measurements on three of the alloys: AL6XN (UN N08367), 254SMO (UNS S32l54), and 304L SS (UNS S30403). Type 304L SS was included in the study as a control. The alloys were tested in three electrolyte solutions which consisted of neutral 3.55% NaC1, 3.55% NaCl in O.1N HC1, and 3.55% NaCl in 1.ON HC1. The solutions were chosen to simulate environments that were expected to be less, similar, and more aggressive, respectively, than those present at the Space Shuttle launch pads. The results from the EIS measurements were analyzed to evaluate the corrosion susceptibility of the alloys and to predict the long-term corrosion performance of the subject materials. The results from the EIS measurements for the three alloys indicated that the higher-alloyed 254SMO and AL6XN exhibited a significantly improved resistance to corrosion than the 304L SS as the concentration of hydrochloric acid in the 3.55% NaC1 solution was increased. The polarization resistance values obtained from the EIS measurements were consistent with those from linear polarization measurements, and were indicative of the actual long-term corrosion performance of the alloys during a two-year atmospheric exposure study.

Calle, L. M.↗

Temperature Effects in Elastohydrodynamically Lubricated Contacts

This paper gives an overview of our current understanding of thermal phenomena in elastohydrodynamic contacts and suggests some avenues for fruitful research in the next decade. Typical measured temperatures are presented for representative conditions and ranges of operating parameters. Temperatures can range from bulk ambient temperature to several hundred degrees centigrade in fully separated elastohydrodynamic films. Although attention in the past decade has been on the full film for the purposes of understanding film thickness and traction phenomena, the more interesting conditions are in the mixed elastohydrodynamic films. These mixed conditions are both common in tribological systems and they are the conditions that border on unsuccessful run-in and failure of the elastohydrodynamic contact. In mixed film conditions local hotspots can have temperatures of the order of 1888 C which cause increased reactivity of the surfaces with surrounding materials as well as changes of the surface physical properties so important to the operation of concentrated contacts. An additional area discussed is that of the bulk system thermal transients which occur in tribological systems. These transients are frequently long in duration and have a direct bearing on the elastohydrodynamic film thickness and traction.

Ward O Winer↗

Understanding How Kurtosis Is Transferred from Input Acceleration to Stress Response and Its Influence on Fatigue Llife

High cycle fatigue of metals typically occurs through long term exposure to time varying loads which, although modest in amplitude, give rise to microscopic cracks that can ultimately propagate to failure. The fatigue life of a component is primarily dependent on the stress amplitude response at critical failure locations. For most vibration tests, it is common to assume a Gaussian distribution of both the input acceleration and stress response. In real life, however, it is common to experience non-Gaussian acceleration input, and this can cause the response to be non-Gaussian. Examples of non-Gaussian loads include road irregularities such as potholes in the automotive world or turbulent boundary layer pressure fluctuations for the aerospace sector or more generally wind, wave or high amplitude acoustic loads. The paper first reviews some of the methods used to generate non-Gaussian excitation signals with a given power spectral density and kurtosis. The kurtosis of the response is examined once the signal is passed through a linear time invariant system. Finally an algorithm is presented that determines the output kurtosis based upon the input kurtosis, the input power spectral density and the frequency response function of the system. The algorithm is validated using numerical simulations. Direct applications of these results include improved fatigue life estimations and a method to accelerate shaker tests by generating high kurtosis, non-Gaussian drive signals.

Kihm, Frederic↗

Faults Discovery By Using Mined Data

Fault discovery in the complex systems consist of model based reasoning, fault tree analysis, rule based inference methods, and other approaches. Model based reasoning builds models for the systems either by mathematic formulations or by experiment model. Fault Tree Analysis shows the possible causes of a system malfunction by enumerating the suspect components and their respective failure modes that may have induced the problem. The rule based inference build the model based on the expert knowledge. Those models and methods have one thing in common; they have presumed some prior-conditions. Complex systems often use fault trees to analyze the faults. Fault diagnosis, when error occurs, is performed by engineers and analysts performing extensive examination of all data gathered during the mission. International Space Station (ISS) control center operates on the data feedback from the system and decisions are made based on threshold values by using fault trees. Since those decision-making tasks are safety critical and must be done promptly, the engineers who manually analyze the data are facing time challenge. To automate this process, this paper present an approach that uses decision trees to discover fault from data in real-time and capture the contents of fault trees as the initial state of the trees.

Lee, Charles↗

Integrated Systems Engineering, Safety, Reliability and Risk Management – Minimizing Black Swan Events

This paper examines key barriers that can possibly inhibit safe and reliable mission execution and, in the worst case, result in loss of human life due to many unknown contributory factors that can lead to Black Swan events. Some of the representative contributory factors include decision errors, overconfidence and a host of common causes including cultural and human factors. Decisions are always easy to criticize in hindsight when more information is available after a major accident. Depending on the type and complexity of the project and/or mission, the catastrophic risks of drifting into failure can be alleviated by implementing uniquely and strategically tailored Integrated-System-of-Systems, dynamic, risk-informed decision management processes. This paper presents some of the lessons learned from James Webb Space Telescope (JWST), NASA’s Human Space Flight program, and industry that provide motivation to organizations working on mega-complex missions to prudently accomplish targeted mission success. These lessons are important for future human Lunar, Mars and Beyond missions planned to be pursued by NASA through a public-private partnership using nimble but effective safety-conscious, proven sound engineering practices including implementation of integrated risk mitigation practices.

SLS↗