Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Common Cause Failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Delamination durability of composite materials for rotorcraft

Delamination is the most commonly observed failure mode in composite rotorcraft dynamic components. Although delamination may not cause immediate failure of the composite part, it often precipitates component repair or replacement, which inhibits fleet readiness, and results in increased life cycle costs. A fracture mechanics approach for analyzing, characterizing, and designing against delamination will be outlined. Examples of delamination problems will be illustrated where the strain energy release rate associated with delamination growth was found to be a useful generic parameter, independent of thickness, layup, and delamination source, for characterizing delamination failure. Several analysis techniques for calculating strain energy release rates for delamination from a variety of sources will be outlined. Current efforts to develop ASTM standard test methods for measuring interlaminar fracture toughness and developing delamination failure criteria will be reviewed. A technique for quantifying delamination durability due to cyclic loading will be presented. The use of this technique for predicting fatigue life of composite laminates and developing a fatigue design philosophy for composite structural components will be reviewed.

Obrien, T. Kevin↗

Implementation and Qualifications Lessons Learned for Space Flight Photonic Components

This slide presentation reviews the process for implementation and qualification of space flight photonic components. It discusses the causes for most common anomalies for the space flight components, design compatibility, a specific failure analysis of optical fiber that occurred in a cable in 1999-2000, and another ExPCA connector anomaly involving pins that broke off. It reviews issues around material selection, quality processes and documentation, and current projects that the Photonics group is involved in. The importance of good documentation is stressed.

Ott, Melanie N.↗

Reliability Effects of Surge Current Testing of Solid Tantalum Capacitors

Solid tantalum capacitors are widely used in space applications to filter low-frequency ripple currents in power supply circuits and stabilize DC voltages in the system. Tantalum capacitors manufactured per military specifications (MIL-PRF-55365) are established reliability components and have less than 0.001% of failures per 1000 hours (the failure rate is less than 10 FIT) for grades D or S, thus positioning these parts among electronic components with the highest reliability characteristics. Still, failures of tantalum capacitors do happen and when it occurs it might have catastrophic consequences for the system. This is due to a short-circuit failure mode, which might be damaging to a power supply, and also to the capability of tantalum capacitors with manganese cathodes to self-ignite when a failure occurs in low-impedance applications. During such a failure, a substantial amount of energy is released by exothermic reaction of the tantalum pellet with oxygen generated by the overheated manganese oxide cathode, resulting not only in destruction of the part, but also in damage of the board and surrounding components. A specific feature of tantalum capacitors, compared to ceramic parts, is a relatively large value of capacitance, which in contemporary low-size chip capacitors reaches dozens and hundreds of microfarads. This might result in so-called surge current or turn-on failures in the parts when the board is first powered up. Such a failure, which is considered as the most prevalent type of failures in tantalum capacitors [I], is due to fast changes of the voltage in the circuit, dV/dt, producing high surge current spikes, I(sub sp) = Cx(dV/dt), when current in the circuit is unrestricted. These spikes can reach hundreds of amperes and cause catastrophic failures in the system. The mechanism of surge current failures has not been understood completely yet, and different hypotheses were discussed in relevant literature. These include a sustained scintillation breakdown model [1-3]; electrical oscillations in circuits with a relatively high inductance [4-6]; local overheating of the cathode [5,7, 8]; mechanical damage to tantalum pentoxide dielectric caused by the impact of MnO2 crystals [2,9, 10]; or stress-induced-generation of electron traps caused by electromagnetic forces developed during current spikes [11]. A commonly accepted explanation of the surge current failures is that at unlimited current supply during surge current conditions, the self-healing mechanism in tantalum capacitors does not work, and what would be a minor scintillation spike if the current were limited, becomes a catastrophic failure of the part [l, 12]. However, our data show that the scintillation breakdown voltages are significantly greater that the surge current breakdown voltages, so it is still not clear why the part, which has no scintillations, would fail at the same voltage during surge current testing (SCT).

Teverovsky, Alexander↗

On the Use of Resilience Models as Digital Twins for Operational Support and In time Decision Making

Human error is a major contributor to accidents and performance losses in complex engineered systems. If one examines these human error caused failures further, a specific cause, the lack of situation awareness, has dominated as a major cause of human errors that instigate latent or catastrophic failures in complex systems. Studies of aviation accidents involving major air carriers revealed that situation awareness was the root cause of around 90% of accidents involving pilot error. Another study explored offshore drilling accidents involving human error and found that 40% of accidents were directly attributed to the loss of situation awareness. Studies of human errors in other domains such as nuclear power, air traffic control, process industry, and advanced driving show that loss of SA was a root cause in a majority of the events. Situation awareness-related failures are not only common but also costly and fatal (e.g., Bhopal Gas Leak, Air France 447 Flight Crash). Thus, the concept of situation awareness has emerged as an important construct in human factors, resulting in numerous models and measurement methods to aid in promoting appropriate levels of situation awareness.

Lukman Irshad↗

Software For Fault-Tree Diagnosis Of A System

Fault Tree Diagnosis System (FTDS) computer program is automated-diagnostic-system program identifying likely causes of specified failure on basis of information represented in system-reliability mathematical models known as fault trees. Is modified implementation of failure-cause-identification phase of Narayanan's and Viswanadham's methodology for acquisition of knowledge and reasoning in analyzing failures of systems. Knowledge base of if/then rules replaced with object-oriented fault-tree representation. Enhancement yields more-efficient identification of causes of failures and enables dynamic updating of knowledge base. Written in C language, C++, and Common LISP.

Iverson, Dave↗

Failure Modes and Effects Analysis (FMEA) Assistant Tool Feasibility Study

An effort to determine the feasibility of a software tool to assist in Failure Modes and Effects Analysis (FMEA) has been completed. This new and unique approach to FMEA uses model based systems engineering concepts to recommend failure modes, causes, and effects to the user after they have made several selections from pick lists about a component s functions and inputs/outputs. Recommendations are made based on a library using common failure modes identified over the course of several major human spaceflight programs. However, the tool could be adapted for use in a wide range of applications from NASA to the energy industry.

Flores, Melissa↗

Electrical Submersible Pumps: A System Modeling Approach for Power Quality Analysis With Variable Frequency Drives

This article proposes a high-level engineering guide to develop an integrated system model for power quality analysis in electrical submersible pump (ESP) applications with variable frequency drives (VFDs). Such analyses are troublesome in the industry due to the complexity of these systems. In this article, simple steps to perform system integration analyses of such arrangements, including torsional analysis, are developed. A simplified VFD-ESP model suitable for coupled electrical and mechanical analysis in steady state is proposed in all their configurations of practical interest. Such a model can be easily implemented in common simulation software, significantly reducing engineering efforts for implementation and analysis. The focus of the proposed model is the prediction of ESP failures that might result from a poor power quality caused by VFDs. Analytical expressions of different types of harmonics in these systems, as well as their accurate locations in the frequency domain, including their interharmonics and common-mode harmonics, are derived for this purpose. Further, the effectiveness of the proposed model is verified through offline and real-time hybrid simulation results. Finally, a comparison between simulation results obtained using the proposed model and measurements collected on a down-scale laboratory prototype is carried out to demonstrate the accuracy of the suggested modeling approach.

42 ENGINEERING↗

Industrial Ergonomics And Safe Lifting Practices: Food Depot [Slides]

Injuries happen when we go beyond what tissues can handle. It's repeated loading and subsequent failure of soft tissue that seem to account for most low back pain. This presentation discusses some common causes of workplace injuries and ways to reduce the likelihood of them happening.

99 GENERAL AND MISCELLANEOUS↗

Oversimplification of Systems Engineering Goals, Processes, and Criteria in NASA Space Life Support

This paper investigates the oversimplification of the inherently complex systems engineering process in space life support. The standard systems engineering process steps are described. The International Space Station (ISS) life support system is explained with its goals and performance criteria. Although it is not usually emphasized, the essential function of developing a hierarchy of systems and subsystems is to simplify the design process. The System Complexity Metric (SCM) shows how this di-vide-and-conquer approach also reduces the system complexity. The complete systems engineering process has many detailed steps. It is often simplified because of the effort required and the human limitations on working memory and decision span. Systems analysis demands slow, logical, and fo-cused thinking but is often bypassed in favor of quick, intuitive, subconscious “gut feel.” A study of 100 system designs found examples of 12 specific mental mistakes, such as ignoring stakeholder needs, and these mistakes are essentially oversimplifications of the systems engineering process. An analysis of space life support goals, options, criteria, and processes found 11 examples of oversimplifications in systems engineering, such as neglecting safety and cost. All these 11 oversimplifications could be traced to one or more of the 12 previously identified mental mistakes or other well-known ones, such as ig-noring sunk costs. Oversimplification of the systems engineering process is rarely noticed but is a common and harmful problem. A study of failures in 50 different space systems found that problems in systems engineering caused failures and often led to errors in design, development, and test that further contributed to failure. It seems that more diligent systems engineering could prevent many project problems and failures, but projects seem to be more guided by “gut feel” based on tradition, authority, and consensus than on the logical, rational systems engineering approach.

Simplified systems engineering↗

Common Cause Case Study: An Estimated Probability of Four Solid Rocket Booster Hold-Down Post Stud Hang-ups

Until Solid Rocket Motor ignition, the Space Shuttle is mated to the Mobil Launch Platform in part via eight (8) Solid Rocket Booster (SRB) hold-down bolts. The bolts are fractured using redundant pyrotechnics, and are designed to drop through a hold-down post on the Mobile Launch Platform before the Space Shuttle begins movement. The Space Shuttle program has experienced numerous failures where a bolt has hung up. That is, it did not clear the hold-down post before liftoff and was caught by the SRBs. This places an additional structural load on the vehicle that was not included in the original certification requirements. The Space Shuttle is currently being certified to withstand the loads induced by up to three (3) of eight (8) SRB hold-down experiencing a "hang-up". The results of loads analyses performed for (4) stud hang-ups indicate that the internal vehicle loads exceed current structural certification limits at several locations. To determine the risk to the vehicle from four (4) stud hang-ups, the likelihood of the scenario occurring must first be evaluated. Prior to the analysis discussed in this paper, the likelihood of occurrence had been estimated assuming that the stud hang-ups were completely independent events. That is, it was assumed that no common causes or factors existed between the individual stud hang-up events. A review of the data associated with the hang-up events, showed that a common factor (timing skew) was present. This paper summarizes a revised likelihood evaluation performed for the four (4) stud hang-ups case considering that there are common factors associated with the stud hang-ups. The results show that explicitly (i.e. not using standard common cause methodologies such as beta factor or Multiple Greek Letter modeling) taking into account the common factor of timing skew results in an increase in the estimated likelihood of four (4) stud hang-ups of an order of magnitude over the independent failure case.

Cross, Robert↗

Common Cause Case Study: An Estimated Probability of Four Solid Rocket Booster Hold-down Post Stud Hang-ups

Until Solid Rocket Motor ignition, the Space Shuttle is mated to the Mobil Launch Platform in part via eight (8) Solid Rocket Booster (SRB) hold-down bolts. The bolts are fractured using redundant pyrotechnics, and are designed to drop through a hold-down post on the Mobile Launch Platform before the Space Shuttle begins movement. The Space Shuttle program has experienced numerous failures where a bolt has "hung-up." That is, it did not clear the hold-down post before liftoff and was caught by the SRBs. This places an additional structural load on the vehicle that was not included in the original certification requirements. The Space Shuttle is currently being certified to withstand the loads induced by up to three (3) of eight (8) SRB hold-down post studs experiencing a "hang-up." The results af loads analyses performed for four (4) stud-hang ups indicate that the internal vehicle loads exceed current structural certification limits at several locations. To determine the risk to the vehicle from four (4) stud hang-ups, the likelihood of the scenario occurring must first be evaluated. Prior to the analysis discussed in this paper, the likelihood of occurrence had been estimated assuming that the stud hang-ups were completely independent events. That is, it was assumed that no common causes or factors existed between the individual stud hang-up events. A review of the data associated with the hang-up events, showed that a common factor (timing skew) was present. This paper summarizes a revised likelihood evaluation performed for the four (4) stud hang-ups case considering that there are common factors associated with the stud hang-ups. The results show that explicitly (i.e. not using standard common cause methodologies such as beta factor or Multiple Greek Letter modeling) taking into account the common factor of timing skew results in an increase in the estimated likelihood of four (4) stud hang-ups of an order of magnitude over the independent failure case.

Cross, Robert↗

What Reliability Engineers Should Know about Space Radiation Effects

Space radiation in space systems present unique failure modes and considerations for reliability engineers. Radiation effects is not a one size fits all field. Threat conditions that must be addressed for a given mission depend on the mission orbital profile, the technologies of parts used in critical functions and on application considerations, such as supply voltages, temperature, duty cycle, and redundancy. In general, the threats that must be addressed are of two types-the cumulative degradation mechanisms of total ionizing dose (TID) and displacement damage (DD). and the prompt responses of components to ionizing particles (protons and heavy ions) falling under the heading of single-event effects. Generally degradation mechanisms behave like wear-out mechanisms on any active components in a system: Total Ionizing Dose (TID) and Displacement Damage: (1) TID affects all active devices over time. Devices can fail either because of parametric shifts that prevent the device from fulfilling its application or due to device failures where the device stops functioning altogether. Since this failure mode varies from part to part and lot to lot, lot qualification testing with sufficient statistics is vital. Displacement damage failures are caused by the displacement of semiconductor atoms from their lattice positions. As with TID, failures can be either parametric or catastrophic, although parametric degradation is more common for displacement damage. Lot testing is critical not just to assure proper device fi.mctionality throughout the mission. It can also suggest remediation strategies when a device fails. This paper will look at these effects on a variety of devices in a variety of applications. This paper will look at these effects on a variety of devices in a variety of applications. (2) On the NEAR mission a functional failure was traced to a PIN diode failure caused by TID induced high leakage currents. NEAR was able to recover from the failure by reversing the current of a nearby Thermal Electric Cooler (turning the TEC into a heater). The elevated temperature caused the PIN diode to anneal and the device to recover. It was by lot qualification testing that NEAR knew the diode would recover when annealed. This paper will look at these effects on a variety of devices in a variety of applications. Single Event Effects (SEE): (1) In contrast to TID and displacement damage, Single Event Effects (SEE) resemble random failures. SEE modes can range from changes in device logic (single-event upset, or SEU). temporary disturbances (single-event transient) to catastrophic effects such as the destructive SEE modes, single-event latchup (SEL). single-event gate rupture (SEGR) and single-event burnout (SEB) (2) The consequences of nondestructive SEE modes such as SEU and SET depend critically on their application--and may range from trivial nuisance errors to catastrophic loss of mission. It is critical not just to ensure that potentially susceptible devices are well characterized for their susceptibility, but also to work with design engineers to understand the implications of each error mode. -For destructive SEE, the predominant risk mitigation strategy is to avoid susceptible parts, or if that is not possible. to avoid conditions under which the part may be susceptible. Destructive SEE mechanisms are often not well understood, and testing is slow and expensive, making rate prediction very challenging. (3) Because the consequences of radiation failure and degradation modes depend so critically on the application as well as the component technology, it is essential that radiation, component. design and system engineers work togetherpreferably starting early in the program to ensure critical applications are addressed in time to optimize the probability of mission success.

DiBari, Rebecca↗

A diagnosis system using object-oriented fault tree models

Spaceborne computing systems must provide reliable, continuous operation for extended periods. Due to weight, power, and volume constraints, these systems must manage resources very effectively. A fault diagnosis algorithm is described which enables fast and flexible diagnoses in the dynamic distributed computing environments planned for future space missions. The algorithm uses a knowledge base that is easily changed and updated to reflect current system status. Augmented fault trees represented in an object-oriented form provide deep system knowledge that is easy to access and revise as a system changes. Given such a fault tree, a set of failure events that have occurred, and a set of failure events that have not occurred, this diagnosis system uses forward and backward chaining to propagate causal and temporal information about other failure events in the system being diagnosed. Once the system has established temporal and causal constraints, it reasons backward from heuristically selected failure events to find a set of basic failure events which are a likely cause of the occurrence of the top failure event in the fault tree. The diagnosis system has been implemented in common LISP using Flavors.

Iverson, David L.↗

Phenomena associated with bench and thermal-vacuum testing of super conductors - Heat pipes.

Test failures of heat pipes occur when the functional performance is unable to match the expected design limits or when the power applied to the heat pipe (in the form of heat) is distributed unevenly through the system, yielding a large thermal gradient. When a thermal gradient larger than expected is measured, it normally occurs in the evaporator or condenser sections of the pipe. Common causes include evaporator overheating, condenser dropout, noncondensable gas formation, surge and partial recovery of evaporator temperatures, masking of thermal profiles, and simple malfunctions due to leaks and mechanical failures or flaws. Examples of each of these phenomena are described along with corresponding failure analyses and corrective measures.

Marshburn, J. P.↗

Cryo-EM confirms a common fibril fold in the heart of four patients with ATTRwt amyloidosis

ATTR amyloidosis results from the conversion of transthyretin into amyloid fibrils that deposit in tissues causing organ failure and death. This conversion is facilitated by mutations in ATTRv amyloidosis, or aging in ATTRwt amyloidosis. ATTRv amyloidosis exhibits extreme phenotypic variability, whereas ATTRwt amyloidosis presentation is consistent and predictable. Previously, we found unique structural variabilities in cardiac amyloid fibrils from polyneuropathic ATTRv-I84S patients. In contrast, cardiac fibrils from five genotypically different patients with cardiomyopathy or mixed phenotypes are structurally homogeneous. To understand fibril structure’s impact on phenotype, it is necessary to study the fibrils from multiple patients sharing genotype and phenotype. Here we show the cryo-electron microscopy structures of fibrils extracted from four cardiomyopathic ATTRwt amyloidosis patients. Our study confirms that they share identical conformations with minimal structural variability, consistent with their homogenous clinical presentation. Our study contributes to the understanding of ATTR amyloidosis biopathology and calls for further studies.

59 BASIC BIOLOGICAL SCIENCES↗

Detecting and Characterizing Patterns of Failure in Complex Engineered Systems: an Ontology Development and Clustering Approach

While the causes of failures in complex engineered systems are often clear in hindsight, it can be challenging to predict failures proactively during the design of novel engineered products or systems. Identifying patterns can be useful for capturing common characteristics that may lead to failure. In this paper, we present a methodology for identifying patterns of failure from NASA’s publicly available Lessons Learned Information System (LLIS). We apply an ontology development and clustering approach to identify representative patterns leading to failures in historical lessons learned. A joint inductive-deductive approach reveals the key themes in lessons that lead to failure, which are formalized and recorded as an ontology of complex systems failure causes. Documents from the LLIS are manually tagged with relevant characteristics from the ontology. From the tagged set, clustering is used to capture co-occurring sets of characteristics that lead to failure. The primary contribution of this work is a method for extracting a set of generic failure patterns in complex engineered systems and characteristics for these patterns that can be identified at design time, knowledge of which can be used to plan mitigation strategies.

Systems Engineering↗

Detecting and Characterizing Patterns of Failure in Complex Systems: An Ontology Development and Clustering Approach

While the causes of failures in complex engineered systems are often clear in hindsight, it can be challenging to predict failures proactively during the design of novel engineered products or systems. Identifying patterns can be useful for capturing common characteristics that may lead to failure. In this paper, we present a methodology for identifying patterns of failure from NASA’s publicly available Lessons Learned Information System (LLIS). We apply an ontology development and clustering approach to identify representative patterns leading to failures in historical lessons learned. A joint inductive-deductive approach reveals the key themes in lessons that lead to failure, which are formalized and recorded as an ontology of complex systems failure causes. Documents from the LLIS are manually tagged with relevant characteristics from the ontology. From the tagged set, clustering is used to capture co-occurring sets of characteristics that lead to failure. The primary contribution of this work is a method for extracting a set of generic failure patterns in complex engineered systems and characteristics for these patterns that can be identified at design time, knowledge of which can be used to plan mitigation strategies.

Systems Engineering↗

Parts, Materials, and Processes Experience Summary

The ALERT program, a system for communicating common problems with parts, materials, and processes, is condensed and catalogued. Expanded information on selected topics is provided by relating the problem area (failure) to the cause, the investigations and findings, the suggestions for avoidance (inspections, screening tests, proper part applications), and failure analysis procedures. The basic objective of ALERT is the avoidance of the recurrence of parts, materials, and processed problems, thus improving the reliability of equipment produced for and used by the government.

Source record↗