Engineering PapersSearch

SEARCH · Engineering Papers

Results for “FMEA”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Independent Orbiter Assessment (IOA): Assessment of the reaction control system, volume 5

The results of the Independent Orbiter Assessment (IOA) of the Failure Modes and Effects Analysis (FMEA) and Critical Items List (CIL) are presented. The IOA effort first completed an analysis of the aft and forward Reaction Control System (RCS) hardware and Electrical Power Distribution and Control (EPD and C), generating draft failure modes and potential critical items. The IOA results were then compared to the proposed Post 51-L NASA FMEA/CIL baseline. This report documents the results of that comparison for the Orbiter RCS hardware and EPD and C systems. Volume 5 contains detailed analysis and superseded analysis worksheets and the NASA FMEA to IOA worksheet cross reference and recommendations.

Prust, Chet D.

Space Shuttle Main Engine Quantitative Risk Assessment: Illustrating Modeling of a Complex System with a New QRA Software Package

During 1997, a team from Hernandez Engineering, MSFC, Rocketdyne, Thiokol, Pratt & Whitney, and USBI completed the first phase of a two year Quantitative Risk Assessment (QRA) of the Space Shuttle. The models for the Shuttle systems were entered and analyzed by a new QRA software package. This system, termed the Quantitative Risk Assessment System(QRAS), was designed by NASA and programmed by the University of Maryland. The software is a groundbreaking PC-based risk assessment package that allows the user to model complex systems in a hierarchical fashion. Features of the software include the ability to easily select quantifications of failure modes, draw Event Sequence Diagrams(ESDs) interactively, perform uncertainty and sensitivity analysis, and document the modeling. This paper illustrates both the approach used in modeling and the particular features of the software package. The software is general and can be used in a QRA of any complex engineered system. The author is the project lead for the modeling of the Space Shuttle Main Engines (SSMEs), and this paper focuses on the modeling completed for the SSMEs during 1997. In particular, the groundrules for the study, the databases used, the way in which ESDs were used to model catastrophic failure of the SSMES, the methods used to quantify the failure rates, and how QRAS was used in the modeling effort are discussed. Groundrules were necessary to limit the scope of such a complex study, especially with regard to a liquid rocket engine such as the SSME, which can be shut down after ignition either on the pad or in flight. The SSME was divided into its constituent components and subsystems. These were ranked on the basis of the possibility of being upgraded and risk of catastrophic failure. Once this was done the Shuttle program Hazard Analysis and Failure Modes and Effects Analysis (FMEA) were used to create a list of potential failure modes to be modeled. The groundrules and other criteria were used to screen out the many failure modes that did not contribute significantly to the catastrophic risk. The Hazard Analysis and FMEA for the SSME were also used to build ESDs that show the chain of events leading from the failure mode occurence to one of the following end states: catastrophic failure, engine shutdown, or siccessful operation( successful with respect to the failure mode under consideration).

Smart, Christian

A Framework for Creating a Function-based Design Tool for Failure Mode Identification

Knowledge of potential failure modes during design is critical for prevention of failures. Currently industries use procedures such as Failure Modes and Effects Analysis (FMEA), Fault Tree analysis, or Failure Modes, Effects and Criticality analysis (FMECA), as well as knowledge and experience, to determine potential failure modes. When new products are being developed there is often a lack of sufficient knowledge of potential failure mode and/or a lack of sufficient experience to identify all failure modes. This gives rise to a situation in which engineers are unable to extract maximum benefits from the above procedures. This work describes a function-based failure identification methodology, which would act as a storehouse of information and experience, providing useful information about the potential failure modes for the design under consideration, as well as enhancing the usefulness of procedures like FMEA. As an example, the method is applied to fifteen products and the benefits are illustrated.

Arunajadai, Srikesh G.

Integrated System Health Management (ISHM) for Test Stand and J-2X Engine: Core Implementation

ISHM capability enables a system to detect anomalies, determine causes and effects, predict future anomalies, and provides an integrated awareness of the health of the system to users (operators, customers, management, etc.). NASA Stennis Space Center, NASA Ames Research Center, and Pratt & Whitney Rocketdyne have implemented a core ISHM capability that encompasses the A1 Test Stand and the J-2X Engine. The implementation incorporates all aspects of ISHM; from anomaly detection (e.g. leaks) to root-cause-analysis based on failure mode and effects analysis (FMEA), to a user interface for an integrated visualization of the health of the system (Test Stand and Engine). The implementation provides a low functional capability level (FCL) in that it is populated with few algorithms and approaches for anomaly detection, and root-cause trees from a limited FMEA effort. However, it is a demonstration of a credible ISHM capability, and it is inherently designed for continuous and systematic augmentation of the capability. The ISHM capability is grounded on an integrating software environment used to create an ISHM model of the system. The ISHM model follows an object-oriented approach: includes all elements of the system (from schematics) and provides for compartmentalized storage of information associated with each element. For instance, a sensor object contains a transducer electronic data sheet (TEDS) with information that might be used by algorithms and approaches for anomaly detection, diagnostics, etc. Similarly, a component, such as a tank, contains a Component Electronic Data Sheet (CEDS). Each element also includes a Health Electronic Data Sheet (HEDS) that contains health-related information such as anomalies and health state. Some practical aspects of the implementation include: (1) near real-time data flow from the test stand data acquisition system through the ISHM model, for near real-time detection of anomalies and diagnostics, (2) insertion of the J-2X predictive model providing predicted sensor values for comparison with measured values and use in anomaly detection and diagnostics, and (3) insertion of third-party anomaly detection algorithms into the integrated ISHM model.

Figueroa, Jorge F.

Modeling Off-Nominal Behavior in SysML

Fault Management is an essential part of the system engineering process that is limited in its effectiveness by the ad hoc nature of the applied approaches and methods. Providing a rigorous way to develop and describe off-nominal behavior is a necessary step in the improvement of fault management, and as a result, will enable safe, reliable and available systems even as system complexity increases... The basic concepts described in this paper provide a foundation to build a larger set of necessary concepts and relationships for precise modeling of off-nominal behavior, and a basis for incorporating these ideas into the overall systems engineering process.. The simple FMEA example provided applies the modeling patterns we have developed and illustrates how the information in the model can be used to reason about the system and derive typical fault management artifacts.. A key insight from the FMEA work was the utility of defining failure modes as the "inverse of intent", and deriving this from the behavior models.. Additional work is planned to extend these ideas and capabilities to other types of relevant information and additional products.

Soil Moisture Active-Passive (SMAP) Mission

Automated Generation of Fault Management Artifacts from a Simple System Model

Our understanding of off-nominal behavior - failure modes and fault propagation - in complex systems is often based purely on engineering intuition; specific cases are assessed in an ad hoc fashion as a (fallible) fault management engineer sees fit. This work is an attempt to provide a more rigorous approach to this understanding and assessment by automating the creation of a fault management artifact, the Failure Modes and Effects Analysis (FMEA) through querying a representation of the system in a SysML model. This work builds off the previous development of an off-nominal behavior model for the upcoming Soil Moisture Active-Passive (SMAP) mission at the Jet Propulsion Laboratory. We further developed the previous system model to more fully incorporate the ideas of State Analysis, and it was restructured in an organizational hierarchy that models the system as layers of control systems while also incorporating the concept of "design authority". We present software that was developed to traverse the elements and relationships in this model to automatically construct an FMEA spreadsheet. We further discuss extending this model to automatically generate other typical fault management artifacts, such as Fault Trees, to efficiently portray system behavior, and depend less on the intuition of fault management engineers to ensure complete examination of off-nominal behavior.

Spinup and Orient

Reliability and Maintainability Analysis for the Amine Swingbed Carbon Dioxide Removal System

I have performed a reliability & maintainability analysis for the Amine Swingbed payload system. The Amine Swingbed is a carbon dioxide removal technology that has gone through 2,400 hours of International Space Station on-orbit use between 2013 and 2016. While the Amine Swingbed is currently an experimental payload system, the Amine Swingbed may be converted to system hardware. If the Amine Swingbed becomes system hardware, it will supplement the Carbon Dioxide Removal Assembly (CDRA) as the primary CO2 removal technology on the International Space Station. NASA is also considering using the Amine Swingbed as the primary carbon dioxide removal technology for future extravehicular mobility units and for the Orion, which will be used for the Asteroid Redirect and Journey to Mars missions. The qualitative component of the reliability and maintainability analysis is a Failure Modes and Effects Analysis (FMEA). In the FMEA, I have investigated how individual components in the Amine Swingbed may fail, and what the worst case scenario is should a failure occur. The significant failure effects are the loss of ability to remove carbon dioxide, the formation of ammonia due to chemical degradation of the amine, and loss of atmosphere because the Amine Swingbed uses the vacuum of space to regenerate the Amine Swingbed. In the quantitative component of the reliability and maintainability analysis, I have assumed a constant failure rate for both electronic and nonelectronic parts. Using this data, I have created a Poisson distribution to predict the failure rate of the Amine Swingbed as a whole. I have determined a mean time to failure for the Amine Swingbed to be approximately 1,400 hours. The observed mean time to failure for the system is between 600 and 1,200 hours. This range includes initial testing of the Amine Swingbed, as well as software faults that are understood to be non-critical. If many of the commercial parts were switched to military-grade parts, the expected mean time to failure would be 2,300 hours. Both calculated mean times to failure for the Amine Swingbed use conservative failure rate models. The observed mean time to failure for CDRA is 2,500 hours. Working on this project and for NASA in general has helped me gain insight into current aeronautics missions, reliability engineering, circuit analysis, and different cultures. Prior my internship, I did not have a lot knowledge about the work being performed at NASA. As a chemical engineer, I had not really considered working for NASA as a career path. By engaging in interactions with civil servants, contractors, and other interns, I have learned a great deal about modern challenges that NASA is addressing. My work has helped me develop a knowledge base in safety and reliability that would be difficult to find elsewhere. Prior to this internship, I had not thought about reliability engineering. Now, I have gained a skillset in performing reliability analyses, and understanding the inner workings of a large mechanical system. I have also gained experience in understanding how electrical systems work while I was analyzing the electrical components of the Amine Swingbed. I did not expect to be exposed to as many different cultures as I have while working at NASA. I am referring to both within NASA and the Houston area. NASA employs individuals with a broad range of backgrounds. It has been great to learn from individuals who have highly diverse experiences and outlooks on the world. In the Houston area, I have come across individuals from different parts of the world. Interacting with such a high number of individuals with significantly different backgrounds has helped me to grow as a person in ways that I did not expect. My time at NASA has opened a window into the field of aeronautics. After earning a bachelor's degree in chemical engineering, I plan to go to graduate school for a PhD in engineering. Prior to coming to NASA, I was not aware of the graduate Pathways program. I intend to apply for the graduate Pathways program as positions are opened up. I would like to pursue future opportunities with NASA, especially as my engineering career progresses.

Dunbar, Tyler

Uncovering Hazards Using a Multi-Objective Optimization to Explore the Faulty State-Space

Considering resilience when designing complex engineered systems is crucial to ensure the system is safe under unexpected hazardous scenarios. Traditional risk-based approaches, such as Failure Modes and Effects Analysis (FMEA) are useful for designing the system to mitigate hazardous scenarios that can be identified by the designer, but often require experience or prior knowledge of system failures to generate. More recently, researchers have developed simulation tools that enable the designer to model large sets of hazardous scenarios (driven by both internal faults and external factors) through simulation. While these tools enable a wider scope of fault modes to be evaluated (e.g., by injecting combined set of fault modes or injecting modes at different times), the resulting assessments (like FMEA) still require knowledge of the specific modes to be evaluated. However, failure to analyze a wide variety of fault scenarios can lead to an incomplete picture of the system resilience, especially to "surprise events'' which may be difficult for the designer to identify and predict beforehand. To overcome this challenge, previous work developed a fault sampling approach for resilience simulations which would procedurally-generate a wide variety of potential faults by systematically perturbing the health states of the system. While the resulting fault modes generated covered a much larger space hazards than would be otherwise considered (and identified many unique failure trajectories which would not have otherwise been identified), it also significantly increased the computational cost of the analysis and resulted in the simulation and analysis of a large set of essentially duplicate scenarios. Additionally, as the number of dimensions in the faulty state-space increases, the full elaboration of possible modes becomes computationally infeasible, justifying the use of a more targeted search. To resolve this limitation, this work proposes the use of a multiobjective optimization algorithm to search the health state space for potential fault modes that are both (1) hazardous and (2) unique. To solve this type of problem, this work proposes the use of a cooperative co-evolutionary algorithm. To demonstrate this approach, it will be applied to a model of an autonomous rover which uses line markings to navigate, focusing on potential hazards in the drive system which could cause the rover to crash. To determine the merit of the approach, it will further be compared with the previously-presented range elaboration approach and a random mode generation approach on the basis of computational efficiency and found modes.

Resilience

ECAR-7517 Rev 1 MARVEL I&C Failure Modes and Effects Analysis

The purpose of this document is to perform a single-failure analysis through the use of a Failure Modes and Effects Analysis (FMEA) for the Safety Related components of the MARVEL Instrumentation and Control (I&C) System. The intent is that this document will meet the requirements for a single-failure analysis described in IEEE-379, “IEEE Standard for Application of the Single-Failure Criterion to Nuclear Power Generating Station Safety Systems”, to verify that this design does indeed meet the single failure criterion. Principles of IEEE-352, “IEEE Guide for General Principles of Reliability Analysis of Nuclear Power Generating Station Safety Systems” are followed to ensure the analysis is consistent with industry standards.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN

Toward the application of the Risk-Based Maintenance approach to a hydrogen-fueled manufacturing plant

Safety, reliability, and maintenance of hydrogen-based equipment, including Risk-Based Maintenance (RBM), have recently gained increasing attention, as hydrogen, while key to decarbonizing hard-to-abate sectors, raises safety concerns. However, two literature gaps limit the development of risk-based strategies for hydrogen technologies. First, existing RBM methodologies do not account for hydrogen-specific components, such as electrolyzers, which differ from conventional equipment. Second, available studies on electrolyzer reliability remain largely qualitative or laboratory-scale. This study addresses these gaps by proposing an adapted RBM methodology tailored to hydrogen-fueled manufacturing facilities. Integrating qualitative tools (e.g., FMEA) with probabilistic models (e.g., Bayesian Networks) enables comprehensive hazardous scenarios identification and likelihood estimation. A conceptual layout of a glass furnace supplied by a 3 MW PEM electrolyzer is considered as a case study to demonstrate the feasibility of the adapted RBM framework in identifying high-risk components and enabling maintenance prioritization, potentially improving plant safety.

08 HYDROGEN

Adapting Traditional Hazards Analysis Methods to Address Cyber Risks

Traditional hazards analysis (HA) methods, originally developed to address physical and operational risks, often fall short when it comes to identifying and mitigating cyber threats. These cyber threats pose unique and evolving risks to critical infrastructure and industrial control systems (ICS). This report explores the integration of Cyber-Informed Engineering (CIE) principles into existing HA methods to enhance their ability to address cyber-induced risks. CIE provides organizations with a practical, cost-effective approach to closing the gap between traditional HA methods and the need for cyber risk mitigation. By leveraging existing safety processes and controls, CIE allows users to examine and mitigate cyber vulnerabilities without overhauling existing HA methods. This report identifies areas where HA and CIE naturally align and where their approaches diverge. It emphasizes how CIE principles can be used to adapt HA methods, broadening their scope to include cyber risks and enabling the mitigation of cyber- induced impacts alongside traditional hazards and failure scenarios. This report examines how CIE can be applied across various HA methods—such as Hazard and Operability Studies (HAZOP), Probabilistic Risk Assessment (PRA), Failure Modes and Effects Analysis (FMEA), Systems-Theoretic Process Analysis (STPA), Hazard and Consequence Analysis for Digital Systems (HAZCADS), and Layers of Protection Analysis (LOPA). It provides strategies for integrating CIE to strengthen the identification, assessment, and mitigation of cyber-induced risks. The findings offer a structured entry point for organizations to embed CIE concepts into hazards and safety analyses, as well as broader engineering processes, ultimately supporting the design and operation of a more resilient infrastructure.

42 ENGINEERING

Adapting Traditional Hazards Analysis Methods to Address Cyber Risks

Traditional hazards analysis (HA) methods, originally developed to address physical and operational risks, often fall short when it comes to identifying and mitigating cyber threats. These cyber threats pose unique and evolving risks to critical infrastructure and industrial control systems (ICS). This report explores the integration of Cyber-Informed Engineering (CIE) principles into existing HA methods to enhance their ability to address cyber-induced risks. CIE provides organizations with a practical, cost-effective approach to closing the gap between traditional HA methods and the need for cyber risk mitigation. By leveraging existing safety processes and controls, CIE allows users to examine and mitigate cyber vulnerabilities without overhauling existing HA methods. This report identifies areas where HA and CIE naturally align and where their approaches diverge. It emphasizes how CIE principles can be used to adapt HA methods, broadening their scope to include cyber risks and enabling the mitigation of cyber- induced impacts alongside traditional hazards and failure scenarios. This report examines how CIE can be applied across various HA methods—such as Hazard and Operability Studies (HAZOP), Probabilistic Risk Assessment (PRA), Failure Modes and Effects Analysis (FMEA), Systems-Theoretic Process Analysis (STPA), Hazard and Consequence Analysis for Digital Systems (HAZCADS), and Layers of Protection Analysis (LOPA). It provides strategies for integrating CIE to strengthen the identification, assessment, and mitigation of cyber-induced risks. The findings offer a structured entry point for organizations to embed CIE concepts into hazards and safety analyses, as well as broader engineering processes, ultimately supporting the design and operation of a more resilient infrastructure.

42 - ENGINEERING

Nuclear Reactor Heat Extraction for Synthetic Fuel Plants

This report looks at the viability and best approach when producing synthetic fuels using heat and power supplied by advanced nuclear energy systems. The report looks at a low temperature integration pathway with four types of advanced reactors: a pressurized water reactor (PWR), an advanced light water reactor (A-LWR), a sodium fast reactor (SFR), and a high-temperature gas-cooled reactor (HTGR). A failure modes and effects analysis (FMEA) of the coupling system between the nuclear plant and the synthetic fuel systems is performed with the goal of identifying the reliability of such a thermal delivery system. Furthermore, a high temperature pathway is investigated for synfuel coupling to determine if this is more efficient and more cost effective as a coupling approach.

10 - SYNTHETIC FUELS

System safety as applied to Skylab

Procedural and organizational guidelines used in accordance with NASA safety policy for the Skylab missions are outlined. The basic areas examined in the safety program for Skylab were the crew interface, extra-vehicular activity (EVA), energy sources, spacecraft interface, and hardware complexity. Fire prevention was a primary goal, with firefighting as backup. Studies of the vectorcardiogram and sleep monitoring experiments exemplify special efforts to prevent fire and shock. The final fire control study included material review, fire detection capability, and fire extinguishing capability. Contractors had major responsibility for system safety. Failure mode and effects analysis (FMEA) and equipment criticality categories are outlined. Redundancy was provided on systems that were critical to crew survival (category I). The five key checkpoints in Skylab hardware development are explained. Skylab rescue capability was demonstrated by preparations to rescue the Skylab 3 crew after their spacecraft developed attitude control problems.

Kleinknecht, K. S.

Photovoltaic power system reliability considerations

An example of how modern engineering and safety techniques can be used to assure the reliable and safe operation of photovoltaic power systems is presented. This particular application is for a solar cell power system demonstration project designed to provide electric power requirements for remote villages. The techniques utilized involve a definition of the power system natural and operating environment, use of design criteria and analysis techniques, an awareness of potential problems via the inherent reliability and FMEA methods, and use of fail-safe and planned spare parts engineering philosophy.

Lalli, V. R.

Mod 1 wind turbine generator failure modes and effects analysis

A failure modes and effects analysis (FMEA) was directed primarily at identifying those critical failure modes that would be hazardous to life or would result in major damage to the system. Each subsystem was approached from the top down, and broken down to successive lower levels where it appeared that the criticality of the failure mode warranted more detail analysis. The results were reviewed by specialists from outside the Mod 1 program, and corrective action taken wherever recommended.

Source record

Modified aerospace R&QA method for wind turbines

This paper describes the Safety, Reliability and Quality Assurance (SR&QA) approach developed for the first large wind turbine generator project, MOD-OA. The SR&QA approach to be used had to assure that the machine would not be hazardous, would operate unattended on a utility grid, would demonstrate reliable operation, and would help establish the quality assurance and maintainability requirements for wind turbine projects. The final approach consisted of a modified Failure Modes and Effects Analysis (FMEA) during the design phase, minimal hardware inspections during parts fabrication, and three documents to control activities during machine construction and operation.

Klein, W. E.

Photovoltaic power system reliability considerations

This paper describes an example of how modern engineering and safety techniques can be used to assure the reliable and safe operation of photovoltaic power systems. This particular application was for a solar cell power system demonstration project in Tangaye, Upper Volta, Africa. The techniques involve a definition of the power system natural and operating environment, use of design criteria and analysis techniques, an awareness of potential problems via the inherent reliability and FMEA methods, and use of a fail-safe and planned spare parts engineering philosophy.

Lalli, V. R.