Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Common Cause Failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

HRA Aerospace Challenges

Compared to equipment designed to perform the same function over and over, humans are just not as reliable. Computers and machines perform the same action in the same way repeatedly getting the same result, unless equipment fails or a human interferes. Humans who are supposed to perform the same actions repeatedly often perform them incorrectly due to a variety of issues including: stress, fatigue, illness, lack of training, distraction, acting at the wrong time, not acting when they should, not following procedures, misinterpreting information or inattention to detail. Why not use robots and automatic controls exclusively if human error is so common? In an emergency or off normal situation that the computer, robotic element, or automatic control system is not designed to respond to, the result is failure unless a human can intervene. The human in the loop may be more likely to cause an error, but is also more likely to catch the error and correct it. When it comes to unexpected situations, or performing multiple tasks outside the defined mission parameters, humans are the only viable alternative. Human Reliability Assessments (HRA) identifies ways to improve human performance and reliability and can lead to improvements in systems designed to interact with humans. Understanding the context of the situation that can lead to human errors, which include taking the wrong action, no action or making bad decisions provides additional information to mitigate risks. With improved human reliability comes reduced risk for the overall operation or project.

DeMott, Diana↗

What is “First Ply Failure” in a Unidirectionally Loaded Quasi-Isotropic Composite (and does it exist)?

This study presents experimental results of unidirectional tension tests on quasi-isotropic laminates to observe so called “first ply failure” (FPF). The modulus of the laminates was measured throughout the test to observe any ply degradation. Quasi-isotropic (QI) laminates were tested since this is the most common layup sequence of a laminate in industry and FPF, if it exists to any practical extent, should be observed. The authors are unaware of any claim that only bi-axial loads are needed to cause FPF, thus uniaxial testing was performed since this is much simpler than bi-axial testing. Laminates can be engineered to cause anomalies (“knees” or abrupt change in modulus) in a stress-strain curve (such as sandwiching multiple 90° plies between a couple of 0° plies) but such laminates should never be used in practice. The search for FPF was examined in simple 8 ply QI laminates by stopping tension tests at various levels post predicted FPF and examining the specimen for “ply failure” and ascertaining its effect (if any) on the modulus of the laminate. The ultimate strength values were found for some specimens. Results showed that for 8 ply QI laminates, no FPF of note was found and challenges the whole notion of FPF and thus “progressive failure”.

First ply failure↗

Variational mechanics analysis of the stresses in microdrop debond specimens

A recently derived variational mechanics analysis of stresses in single-fiber model composites has been applied to the analysis of the stresses in the microdrop debond specimen. The new analysis is more accurate than the commonly applied shear-lag or elastic-plastic analyses. The results from a sample stress state calculation suggest that interfacial failure between the fiber and the microdrop is by mode I or opening mode failure at the beginning of the microdrop. The opening mode failure is caused by a large tensile radial stress at the fiber/matrix interface. Previous analyses of microdrop debond data have been in terms of a shear strength. It is suggested that these analyses misrepresent microdrop debond results and recommend instead a failure analysis based on energy release rate and interfacial fracture toughness. A procedure for calculating the energy release rate for the growth of an interfacial crack is described.

Scheer, Robert J.↗

Airglow response to vertically standing gravity waves

There is currently much interest in fluctuations of airglow emissions caused by atmospheric gravity waves. The fluctuations of brightness tend to be found in phase (or occasionally in antiphase) with the fluctuations of measured temperature, whereas current theory tends to anticipate substantial phase differences. We suggest here that the discrepancy results from failure of the common theoretical assumption that the relevant gravity waves are dominated by a single upgoing component: that, instead, there is an accompanying downgoing component of comparable magnitude, produced by reflection. In the case of total reflection, simple, steady state chemistry and vertical viewing, the phase difference is necessarily zero (or 180 deg).

Hines, Colin O.↗

Fault Tolerant State Machines

State machines are commonly used to control sequential logic in FPGAs and ASKS. An errant state machine can cause considerable damage to the device it is controlling. For example in space applications, the FPGA might be controlling Pyros, which when fired at the wrong time will cause a mission failure. Even a well designed state machine can be subject to random errors us a result of SEUs from the radiation environment in space. There are various ways to encode the states of a state machine, and the type of encoding makes a large difference in the susceptibility of the state machine to radiation. In this paper we compare 4 methods of state machine encoding and find which method gives the best fault tolerance, as well as determining the resources needed for each method.

encoding↗

Recovering from On-orbit Anomalies on the Astrobee Free Flyers and its Systems

Since 2019, NASA has been operating three Astrobee free flying robots on board the International Space Station (ISS) providing an autonomous and flexible research platform for national and international payload developers in microgravity and serving as a robotic assistant for astronauts on the ISS. During its use on the ISS, in particular with over 750 hours of free-flyer operation as of March 2022, Astrobee and its Docking Station have encountered multiple software and hardware anomalies. These anomalies were either resolved remotely via software and firmware updates, or, where not possible, with hardware replacements on orbit or by the return of the faulty unit to NASA’s ground facilities for its repair. Despite being inherently designed to be repaired or replaced on orbit, Astrobee and its systems can still suffer anomalies that would be complex enough to disassemble, cause risks of hardware damage, or use excessive crew time to perform the repair in orbit. That was the case for the anomaly the Astrobee unit ‘Honey’ encountered, reason why it needed to be down-massed for repair. One of the most common points of failure was found to be the SD card, which is used for the different Astrobee processors and for the Dock Station. Other comparable SD card anomalies were found also on the Astrobee ground units, which provided useful data in the effort of upgrading their systems. This presentation will focus on 1) The overview of the different faults and anomalies on Astrobee and its systems on orbit and on the ground 2) The processes and procedures implemented to resolve the anomalies 3) The implementation of software updates and hardware upgrades in order to reduce the risk on returning anomalies 4) The lessons learned in increasing Astrobee’s robustness and resilience to such anomalies.

International Space Station↗

Faulty behavior of asynchronous storage elements

It is often assumed that the faults in storage elements (SE's) can be modeled as output/input stuck-at-faults of the element. They are implicitly considered equivalent to the stuck-at faults in the combinational logic surrounding the SE cells. A more accurate higher level fault model for elementary SE's used in asynchronous circuits is presented. This model offers better representation of the physical failures. It is shown that the stuck-at model may be adequate if only modest fault coverage is desired. The enhanced model includes some common fault behaviors of SE's that are not covered by the stuck-at model. These include data-feed-through behaviors that cause the SE to be combinational. Fault models for complex SE cells can be obtained without a significant loss of information about the structure of the circuit.

Al-Assadi, Waleed K.↗

Comparing the Identification of Recommendations by Different Accident Investigators Using a Common Methodology

Accident reports play a key role in the safety of complex systems. These reports present the recommendations that are intended to help avoid any recurrence of past failures. However, the value of these findings depends upon the causal analysis that helps to identify the reasons why an accident occurred. Various techniques have been developed to help investigators distinguish root causes from contributory factors and contextual information. This paper presents the results from a study into the individual differences that can arise when a group of investigators independently apply the same technique to identify the causes of an accident. This work is important if we are to increase the consistency and coherence of investigations following major accidents.

Johnson, Chris W.↗

Orbital Anomalies in Goddard Spacecraft for Calendar Year 1994

This report summarizes and updates the annual on-orbit performance between January I and December 31, 1994, for spacecraft built by or managed by the Goddard Space Flight Center (GSFC). During 1994, GSFC had 27 active orbiting satellites and I Shuttle-launched and retrieved 'free flyer.' There were 310 reported anomalies among 21 satellites and one GSFC instrument (TOMS). GOES-8 accounted for 66 anomalies, and SAMPES reported 155 'anomalies'. Of the 155 anomalies reported for all but SAMPEX, only 4 affected the spacecraft missions 'substantially' or greater, that is, presented a loss of more than 33% of the total missions. The most frequent subsystem anomalies were Instrument/Payload(44), Timing Command and Control(40), and Attitude Control Systems(33). Of the non-SAMPEX anomalies, 29% had no effect on the missions and 28% caused subsystem or instrument degradation and, for another 28%, no anomaly effect on the mission could be determined. Fifty-three percent of non-SAMPEX anomalies could not be classified according to 'type'; the other most common types were 'systemic'(35), 'random'(19), and 'normal or expected operation'(15). Forty percent of the anomalies were not classified according to failure category; the remaining most frequent occurrences were 'design problems'(50) and 'other known problems'(35).

Thomas, Walter B.↗

Timing issues in the distributed execution of Ada programs

This paper examines, in the context of distributed execution, the meaning of Ada constructs involving time. In the process, unresolved questions of interpretation and problems with the implementation of a consistent notion of time across a network are uncovered. It is observed that there are two Ada mechanisms that can involve a distributed sense of time: the conditional entry call, and the timed entry call. It is shown that a recent interpretation by the Language Maintenance Committee resolves the questions for the conditional entry calls but results in an anomaly for timed entry calls. A detailed discussion of alternative implementations for the timed entry call is made, and it is aruged that: (1) timed entry calls imply a common sense of time between the machines holding the calling and called tasks; and (2) the measurement of time for the expiration of the delay and the decision of whether or not to perform the rendezvous should be made on the machine holding the called task. The need to distinguish the unreadiness of the called task from timeouts caused by network failure is pointed out. Finally, techniques for realizing a single sense of time across the distributed system (at least to within an acceptable degree of uncertainty) are also discussed.

Volz, Richard A.↗

Analysis of SSEM Sensor Data Using BEAM

A report describes analysis of space shuttle main engine (SSME) sensor data using Beacon-based Exception Analysis for Multimissions (BEAM) [NASA Tech Briefs articles, the two most relevant being Beacon-Based Exception Analysis for Multimissions (NPO- 20827), Vol. 26, No.9 (September 2002), page 32 and Integrated Formulation of Beacon-Based Exception Analysis for Multimissions (NPO- 21126), Vol. 27, No. 3 (March 2003), page 74] for automated detection of anomalies. A specific implementation of BEAM, using the Dynamical Invariant Anomaly Detector (DIAD), is used to find anomalies commonly encountered during SSME ground test firings. The DIAD detects anomalies by computing coefficients of an autoregressive model and comparing them to expected values extracted from previous training data. The DIAD was trained using nominal SSME test-firing data. DIAD detected all the major anomalies including blade failures, frozen sense lines, and deactivated sensors. The DIAD was particularly sensitive to anomalies caused by faulty sensors and unexpected transients. The system offers a way to reduce SSME analysis time and cost by automatically indicating specific time periods, signals, and features contributing to each anomaly. The software described here executes on a standard workstation and delivers analyses in seconds, a computing time comparable to or faster than the test duration itself, offering potential for real-time analysis.

Zak, Michail↗

Report of the Odyssey FPGA Independent Assessment Team

An independent assessment team (IAT) was formed and met on April 2, 2001, at Lockheed Martin in Denver, Colorado, to aid in understanding a technical issue for the Mars Odyssey spacecraft scheduled for launch on April 7, 2001. An RP1280A field-programmable gate array (FPGA) from a lot of parts common to the SIRTF, Odyssey, and Genesis missions had failed on a SIRTF printed circuit board. A second FPGA from an earlier Odyssey circuit board was also known to have failed and was also included in the analysis by the IAT. Observations indicated an abnormally high failure rate for flight RP1280A devices (the first flight lot produced using this flow) at Lockheed Martin and the causes of these failures were not determined. Standard failure analysis techniques were applied to these parts, however, additional diagnostic techniques unique for devices of this class were not used, and the parts were prematurely submitted to a destructive physical analysis, making a determination of the root cause of failure difficult. Any of several potential failure scenarios may have caused these failures, including electrostatic discharge, electrical overstress, manufacturing defects, board design errors, board manufacturing errors, FPGA design errors, or programmer errors. Several of these mechanisms would have relatively benign consequences for disposition of the parts currently installed on boards in the Odyssey spacecraft if established as the root cause of failure. However, other potential failure mechanisms could have more dire consequences. As there is no simple way to determine the likely failure mechanisms with reasonable confidence before Odyssey launch, it is not possible for the IAT to recommend a disposition for the other parts on boards in the Odyssey spacecraft based on sound engineering principles.

Mayer, Donald C.↗

Influence of Design Variations on Systems Performance

High-risk aerospace components have to meet very stringent quality, performance, and safety requirements. Any source of variation is a concern, as it may result in scrap or rework. poor performance, and potentially unsafe flying conditions. The sources of variation during product development, including design, manufacturing, and assembly, and during operation are shown. Sources of static and dynamic variation during development need to be detected accurately in order to prevent failure when the components are placed in operation. The Systems' Health and Safety (SHAS) research at the NASA Ames Research Center addresses the problem of detecting and evaluating the statistical variation in helicopter transmissions. In this work, we focus on the variations caused by design, manufacturing, and assembly of these components, prior to being placed in operation (DMV). In particular, we aim to understand and represent the failure and variation information, and their correlation to performance and safety and feed this information back into the development cycle at an early stage. The feedback of such critical information will assure the development of more reliable components with less rework and scrap. Variations during design and manufacturing are a common source of concern in the development and production of such components. Accounting for these variations, especially those that have the potential to affect performance, is accomplished in a variety ways, including Taguchi methods, FMEA, quality control, statistical process control, and variation risk management. In this work, we start with the assumption that any of these variations can be represented mathematically, and accounted for by using analytical tools incorporating these mathematical representations. In this paper, we concentrate on variations that are introduced during design. Variations introduced during manufacturing are investigated in parallel work.

Tumer, Irem Y.↗

Debonding Stress Concentrations in a Pressurized Lobed Sandwich-Walled Generic Cryogenic Tank

A finite-element stress analysis has been conducted on a lobed composite sandwich tank subjected to internal pressure and cryogenic cooling. The lobed geometry consists of two obtuse circular walls joined together with a common flat wall. Under internal pressure and cryogenic cooling, this type of lobed tank wall will experience open-mode (a process in which the honeycomb is stretched in the depth direction) and shear stress concentrations at the junctures where curved wall changes into flat wall (known as a curve-flat juncture). Open-mode and shear stress concentrations occur in the honeycomb core at the curve-flat junctures and could cause debonding failure. The levels of contributions from internal pressure and temperature loading to the open-mode and shear debonding failure are compared. The lobed fuel tank with honeycomb sandwich walls has been found to be a structurally unsound geometry because of very low debonding failure strengths. The debonding failure problem could be eliminated if the honeycomb core at the curve-flat juncture is replaced with a solid core.

Ko, William L.↗

Synthesizing a New Launch Vehicle Failure Probability Based on Historical Flight Data

New launch vehicles have historically had significantly higher failure probabilities in early flights than what has been predicted using Probabilistic Risk Assessment. Work on a new methodology originally started with ARES I-X and the Common Standards Working Group (CSWG) for range safety applications. CSWG consists of the Federal Aviation Administration (FAA), Air Force, and NASA. Historical launch vehicle data was viewed as the best predictor of success/failure for launches of new vehicles. A launch vehicle database was developed that includes all launches from 1980-2017 (both US and foreign). Entries to the database include: Vehicle by model type; Launch dates; Failure description; Failure Result (Loss Of Vehicle (LOV)/Loss Of Mission (LOM); Failure cause (when available); Vehicle designs (stages/engines/etc.)

Early flight risk↗

Lessons Learned Entry: Hypergolic Propellant Related Spills and Fires

The attached report is a compilation of all credible, unintentional hypergolic fluid related spills, fires, and explosions from the Apollo Program, the Space Shuttle Program, Titan Program, and a few other programs. Spill sites include the following government facilities: KSC, JSC, WSTF, VAFB, CCAFS, EAFB, Little Rock AFB, and McConnell AFB. The root causes and consequences of the incidents contained in this document vary drastically; however, certain "themes" can be deduced and utilized for future hypergolic propellant handling. Some of those common "themes" are summarized below: (1) Improper configuration control and complacency can lead to being falsely comfortable with a system (2) Communication breakdown can escalate an incident to a level where injuries occur and/or hardware is damaged (3) Improper propulsion system and ground support system designs can destine a system for failure (4) Improper training of technicians, engineers, and safety personnel can put lives in danger (5) Improper PPE, spill protection, and staging of fire extinguishing equipment can result in unnecessary injuries or hardware damage if an incident occurs (6) Improper procedural oversight, development, and adherence to the procedure can be detrimental and quickly lead to an undesirable incident (7) Improper local cleanliness or compatibility can result in fires or explosions The items listed above are only a short list of the issues that should be recognized prior to handling of hypergolic fluids or processing of vehicles containing hypergolic propellants. The summary of incidents in this report is intended to cover many more issues than those listed above that have been found during nearly the entire spectrum. of hypergolic propellant and/or vehicle processing.

Nufer, Brian↗

A Summary of NASA and USAF Hypergolic Propellant Related Spills and Fires

Several unintentional hypergolic fluid related spills, fires, and explosions from the Apollo Program, the Space Shuttle Program, the Titan Program, and a few others have occurred over the past several decades. Spill sites include the following government facilities: Kennedy Space Center (KSC), Johnson Space Center (JSC), White Sands Test Facility (WSTF), Vandenberg Air Force Base (VAFB), Cape Canaveral Air Force Station (CCAFS), Edwards Air Force Base (EAFB), Little Rock AFB, and McConnell AFB. Until now, the only method of capturing the lessons learned from these incidents has been "word of mouth" or by studying each individual incident report. The root causes and consequences of the incidents vary drastically; however, certain "themes" can be deduced and utilized for future hypergolic propellant handling. Some of those common "themes" are summarized below: (1) Improper configuration control and internal or external human performance shaping factors can lead to being falsely comfortable with a system (2) Communication breakdown can escalate an incident to a level where injuries occur and/or hardware is damaged (3) Improper propulsion system and ground support system designs can destine a system for failure (4) Improper training of technicians, engineers, and safety personnel can put lives in danger (5) Improper PPE, spill protection, and staging of fire extinguishing equipment can result in unnecessary injuries or hardware damage if an incident occurs (6) Improper procedural oversight, development, and adherence to the procedure can be detrimental and quickly lead to an undesirable incident (7) Improper materials cleanliness or compatibility and chemical reactivity can result in fires or explosions (8) Improper established "back-out" and/or emergency safing procedures can escalate an event The items listed above are only a short list of the issues that should be recognized prior to handling hypergolic fluids or processing vehicles containing hypergolic propellants. The summary of incidents in this report is intended to cover many more issues than those listed above.

Nufer, Brian M.↗

Fault Management Techniques in Human Spaceflight Operations

This paper discusses human spaceflight fault management operations. Fault detection and response capabilities available in current US human spaceflight programs Space Shuttle and International Space Station are described while emphasizing system design impacts on operational techniques and constraints. Preflight and inflight processes along with products used to anticipate, mitigate and respond to failures are introduced. Examples of operational products used to support failure responses are presented. Possible improvements in the state of the art, as well as prioritization and success criteria for their implementation are proposed. This paper describes how the architecture of a command and control system impacts operations in areas such as the required fault response times, automated vs. manual fault responses, use of workarounds, etc. The architecture includes the use of redundancy at the system and software function level, software capabilities, use of intelligent or autonomous systems, number and severity of software defects, etc. This in turn drives which Caution and Warning (C&W) events should be annunciated, C&W event classification, operator display designs, crew training, flight control team training, and procedure development. Other factors impacting operations are the complexity of a system, skills needed to understand and operate a system, and the use of commonality vs. optimized solutions for software and responses. Fault detection, annunciation, safing responses, and recovery capabilities are explored using real examples to uncover underlying philosophies and constraints. These factors directly impact operations in that the crew and flight control team need to understand what happened, why it happened, what the system is doing, and what, if any, corrective actions they need to perform. If a fault results in multiple C&W events, or if several faults occur simultaneously, the root cause(s) of the fault(s), as well as their vehicle-wide impacts, must be determined in order to maintain situational awareness. This allows both automated and manual recovery operations to focus on the real cause of the fault(s). An appropriate balance must be struck between correcting the root cause failure and addressing the impacts of that fault on other vehicle components. Lastly, this paper presents a strategy for using lessons learned to improve the software, displays, and procedures in addition to determining what is a candidate for automation. Enabling technologies and techniques are identified to promote system evolution from one that requires manual fault responses to one that uses automation and autonomy where they are most effective. These considerations include the value in correcting software defects in a timely manner, automation of repetitive tasks, making time critical responses autonomous, etc. The paper recommends the appropriate use of intelligent systems to determine the root causes of faults and correctly identify separate unrelated faults.

O'Hagan, Brian↗