Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Failure mode and effect analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A Framework for Creating a Function-based Design Tool for Failure Mode Identification

Knowledge of potential failure modes during design is critical for prevention of failures. Currently industries use procedures such as Failure Modes and Effects Analysis (FMEA), Fault Tree analysis, or Failure Modes, Effects and Criticality analysis (FMECA), as well as knowledge and experience, to determine potential failure modes. When new products are being developed there is often a lack of sufficient knowledge of potential failure mode and/or a lack of sufficient experience to identify all failure modes. This gives rise to a situation in which engineers are unable to extract maximum benefits from the above procedures. This work describes a function-based failure identification methodology, which would act as a storehouse of information and experience, providing useful information about the potential failure modes for the design under consideration, as well as enhancing the usefulness of procedures like FMEA. As an example, the method is applied to fifteen products and the benefits are illustrated.

Arunajadai, Srikesh G.↗

Enhancing Fault Isolation for Health Monitoring of Electric Aircraft Propulsion by Embedding Failure Mode and Effect Analysis into Bayesian Networks

This paper describes a fault isolation approach for electric powertrains of unmanned aerial vehicles. The approach leverages the combination of failure mode and effect analysis (FMEA) and Bayesian networks, thus introducing depend-ability structures into a diagnostic framework. Faults and failure events from the FMEA are mapped within a Bayesian network, where network edges replicate the links embedded within FMEAs. This framework helps the fault isolation process by identifying the probability of occurrence of specific faults or root causes given evidence observed through sensor signals. The framework is applied to an electric power-train system of a small, rotary-wing unmanned aerial vehicle, demonstrating how a Bayesian network enhanced by FMEA helps disambiguate between root causes of incipient failures, which would otherwise be considered as equally probable.

Fault Isolation↗

Failure mode and effects analysis (FMEA) for the Space Shuttle solid rocket motor

The recertification of the Space Shuttle Solid Rocket Booster (SRB) and Solid Rocket Motor (SRM) has included an extensive rewriting of the Failure Mode and Effects Analysis (FMEA) and Critical Items List (CIL). The evolution of the groundrules and methodology used in the analysis is discussed and compared to standard FMEA techniques. Especially highlighted are aspects of the FMEA/CIL which are unique to the analysis of an SRM. The criticality category definitions are presented and the rationale for assigning criticality is presented. The various data required by the CIL and contribution of this data to the retention rationale is also presented. As an example, the FMEA and CIL for the SRM nozzle assembly is discussed in detail. This highlights some of the difficulties associated with the analysis of a system with the unique mission requirements of the Space Shuttle.

Russell, D. L.↗

Failure Mode and Effects Analysis (FMEA) for Photovoltaic Inverter

Photovoltaic (PV) inverters are critical yet vulnerable components in modern energy systems, often acting as reliability bottlenecks that increase the levelized cost of energy (LCOE). To address this, this paper presents a comprehensive Failure Mode and Effects Analysis (FMEA) tailored for PV inverters. Leveraging field data and literature, we identify failure-prone components, such as capacitors,, and relays, and prioritize their risks based on quantitative Risk Priority Numbers (RPNs). The analysis reveals that surge-induced MOV short circuits, capacitor degradation, and environmental cooling fan failures dominate the risk profile. These findings provide a targeted framework for reliability improvement, guiding future efforts in predictive diagnostics, design optimization, and accelerated life testing strategies.

14 SOLAR ENERGY↗

Failure Mode and Effects Analysis for a Photovoltaic Inverter

While PV panel reliability continues to increase, PV inverters become the limiting factor for PV system reliability. Consequently, it is critical to have a generic tool from a third party for PV inverter reliability assessment to help 1) utilities/PV farm operators schedule maintenance in advance, and 2) inverter developers improve the next-generation design. However, these two things cannot be accomplished without first understanding the reasons behind inverter failure. Following this idea, as the first step, it is essential to identify and investigate the most failure-prone components within a PV inverter system. After all, any system is only as reliable as the components that are contained within it. This motivates the failure mode and effects analysis (FMEA) work presented for this workshop. The FMEA is conducted as follows: first, the overview of the methodology on the development of the FMEA is presented; then, based on a top-down approach starting from the PV inverter system, critical inverter components with high failure rates are identified and summarized; afterward, a thorough FMEA study at a component-level is performed and its results, including failure modes, failure mechanisms, and critical stressors, are tabulated; finally, according to three rankings (chance of occurrence, severity of occurrence, and ease of detection prior to failure) for each failure mechanism provided by the FMEA, risk priority numbers are calculated and the failure mechanisms along with the critical stressors are ranked in terms of their potentially detrimental effect on the PV inverter.

Brown, Buck↗

Failure-Modes-And-Effects Analysis Of Software Logic

Rigorous analysis applied early in design effort. Method of identifying potential inadequacies and modes and effects of failures caused by inadequacies (failure-modes-and-effects analysis or "FMEA" for short) devised for application to software logic.

Garcia, Danny↗

General Failure Modes and Effects Analysis for Accelerator and Detector Magnet Design at JLab

The aim of this article is to develop a risk management procedure, which could be applied to the magnet design process, for both superconducting and normal magnets at the Jefferson Laboratory (JLab). This procedure allowed us to identify the key risks at each of the critical phases of design and propose procedures, tests, and checks to mitigate each risk. In this article, we present a qualitative and quantitative risk management procedure commonly referred to a “failure modes and effects analysis.” As part of this procedure, we calculated a risk priority number (RPN) for each activity of the process, identified the most critical activities and proposed mitigation activities, which in turn resulted in a revised RPN. Additionally, another benefit of this procedure was the identification of appropriate “control and hold” points within the design process, which allowed one to review and approve a particular outcome before proceeding to the next sequential activity.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Procedure for Failure Mode, Effects, and Criticality Analysis (FMECA)

This document provides guidelines for the accomplishment of Failure Mode, Effects, and Criticality Analysis (FMECA) on the Apollo program. It is a procedure for analysis of hardware items to determine those items contributing most to system unreliability and crew safety problems.

Source record↗

Contamination Sources Effects Analysis (CSEA) - A Tool to Balance Cost/Schedule While Managing Facility Availability

A CSEA is similar to a Failure Modes Effects Analysis (FMEA). A CSEA tracks risk, deterrence, and occurrence of sources of contamination and their mitigation plans. Documentation is provided spanning mechanical and electrical assembly, precision cleaning, thermal vacuum bake-out, and thermal vacuum testing. These facilities all may play a role in contamination budgeting and reduction ultimately affecting test and flight. With a CSEA, visibility can be given to availability of these facilities, test sequencing and trade-offs. A cross-functional team including specialty engineering, contamination control, electrostatic dissipation, manufacturing, testing, and material engineering participate in an exercise that identifies contaminants and minimizes the complexity of scheduling these facilities considering their volatile schedules. Care can be taken in an efficient manner to insure correct cleaning processes are employed. The result is reduction in cycle time ("schedule hits"), reduced cost due to rework, reduced risk and improved communication and quality while achieving adherence to the Contamination Control Plan.

Wilcox, Margaret↗

Service Life Extension of the Propulsion System of Long-Term Manned Orbital Stations

One of the critical non-replaceable systems of a long-term manned orbital station is the propulsion system. Since the propulsion system operates beginning with the launch of station elements into orbit, its service life determines the service life of the station overall. Weighing almost a million pounds, the International Space Station (ISS) is about four times as large as the Russian space station Mir and about five times as large as the U.S. Skylab. Constructed over a span of more than a decade with the help of over 100 space flights, elements and modules of the ISS provide more research space than any spacecraft ever built. Originally envisaged for a service life of fifteen years, this Earth orbiting laboratory has been in orbit since 1998. Some elements that have been launched later in the assembly sequence were not yet built when the first elements were placed in orbit. Hence, some of the early modules that were launched at the inception of the program were already nearing the end of their design life when the ISS was finally ready and operational. To maximize the return on global investments on ISS, it is essential for the valuable research on ISS to continue as long as the station can be sustained safely in orbit. This paper describes the work performed to extend the service life of the ISS propulsion system. A system comprises of many components with varying failure rates. Reliability of a system is the probability that it will perform its intended function under encountered operating conditions, for a specified period of time. As we are interested in finding out how reliable a system would be in the future, reliability expressed as a function of time provides valuable insight. In a hypothetical bathtub shaped failure rate curve, the failure rate, defined as the number of failures per unit time that a currently healthy component will suffer in a given future time interval, decreases during infant-mortality period, stays nearly constant during the service life and increases at the end when the design service life ends and wear-out phase begins. However, the component failure rates do not remain constant over the entire cycle life. The failure rate depends on various factors such as design complexity, current age of the component, operating conditions, severity of environmental stress factors, etc. Development, qualification and acceptance test processes provide rigorous screening of components to weed out imperfections that might otherwise cause infant mortality failures. If sufficient samples are tested to failure, the failure time versus failure quantity can be analyzed statistically to develop a failure probability distribution function (PDF), a statistical model of the probability of failure versus time. Driven by cost and schedule constraints however, spacecraft components are generally not tested in large numbers. Uncertainties in failure rate and remaining life estimates increase when fewer units are tested. To account for this, spacecraft operators prefer to limit useful operations to a period shorter than the maximum demonstrated service life of the weakest component. Running each component to its failure to determine the maximum possible service life of a system can become overly expensive and impractical. Spacecraft operators therefore, specify the required service life and an acceptable factor of safety (FOS). The designers use these requirements to limit the life test duration. Midway through the design life, when benefits justify additional investments, supplementary life test may be performed to demonstrate the capability to safely extend the service life of the system. An innovative approach is required to evaluate the entire system, without having to go through an elaborate test program of propulsion system elements. Evaluating every component through a brute force test program would be a cost prohibitive and time consuming endeavor. ISS propulsion system components were designed and built decades ago. There are no representative ground test articles for some of the components. A 'test everything' approach would require manufacturing new test articles. The paper outlines some of the techniques used for selective testing, by way of cherry picking candidate components based on failure mode effects analysis, system level impacts, hazard analysis, etc. The type of testing required for extending the service life depends on the design and criticality of the component, failure modes and failure mechanisms, life cycle margin provided by the original certification, operational and environmental stresses encountered, etc. When specific failure mechanism being considered and the underlying relationship of that mode to the stresses provided in the test can be correlated by supporting analysis, time and effort required for conducting life extension testing can be significantly reduced. Exposure to corrosive propellants over long periods of time, for instance, lead to specific failure mechanisms in several components used in the propulsion system. Using Arrhenius model, which is tied to chemically dependent failure mechanisms such as corrosion or chemical reactions, it is possible to subject carefully selected test articles to accelerated life test. Arrhenius model reflects the proportional relationship between time to failure of a component and the exponential of the inverse of absolute temperature acting on the component. The acceleration factor is used to perform tests at higher stresses that allow direct correlation between the times to failure at a high test temperature to the temperatures to be expected in actual use. As long as the temperatures are such that new failure mechanisms are not introduced, this becomes a very useful method for testing to failure a relatively small sample of items for a much shorter amount of time. In this article, based on the example of the propulsion system of the first ISS module Zarya, theoretical approaches and practical activities of extending the service life of the propulsion system are reviewed with the goal of determining the maximum duration of its safe operation.

Kamath, Ulhas↗

Combining System Safety and Reliability to Ensure NASA CoNNeCT's Success

Hazard Analysis, Failure Modes and Effects Analysis (FMEA), the Limited-Life Items List (LLIL), and the Single Point Failure (SPF) List were applied by System Safety and Reliability engineers on NASA's Communications, Navigation, and Networking reConfigurable Testbed (CoNNeCT) Project. The integrated approach involving cross reviews of these reports by System Safety, Reliability, and Design engineers resulted in the mitigation of all identified hazards. The outcome was that the system met all the safety requirements it was required to meet.

Havenhill, Maria↗

CompactPCI(Registered TradeMark) Connectors in Space Flight Applications

This report documents the current status of CompactPCI(Registered TradeMark) connectors in GSFC spaceflight applications. To the extent the information is known, this report summarizes to what component quality level each NASA contractor (referred to as OEM in this report) procured the parts, and what board level and system level testing was performed. The report also provides the current status of the reliability assessment for each GSFC project based on the results of testing and FMEA (Failure Mode Effects Analysis). This report addresses how the CompactPCI(Registered TradeMark) connectors came into existence, and how these became the connector style chosen by many designers of space flight hardware. It identifies the design philosophy and the lack of robustness which has led to several known failure modes. These failure modes include fretting of connector pins during vibration, shock and thermal cycling, exposure of underplating, and increased resistance, including brief excursions to very high resistance. Each of these are signs of aging, which becomes an increasing concern for long duration orbiting space flight applications. This report addresses the mitigation strategy to replace CompactPCI(Registered TradeMark) connectors with space qualified Hypertronics 2mm cPCI connectors. The Hypertronics 2mm cPCI connectors are pin-to-pin compatible with the CompactPCI(Registered TradeMark) connectors and meet all of the same technical requirements, except the ability to hot mate, and to mate directly with a CompactPCI of the opposite gender. A detailed comparison of the CompactPCI(Registered TradeMark) connector and the Hypertronics 2mm cPCI connector is provided to describe the ruggedness of Hypertronics connector for space flight applications. Finally, this report makes recommendations for flight hardware for the future missions where the hardware is yet to be built, as well as for the hardware which has already been built with CompactPCI(Registered TradeMark) connectors.

Williams, Richard↗

Control, Fault Management, and Grid Support Functionality of an MV AC-DC Solid State Transformer based EV Extreme Fast Charging Station

Electric vehicles (EVs) have become increasingly popular in recent times while revolutionizing the consumer and commercial transportation market. The development of charging infrastructure has become one of the priorities for increasing the adoption of EVs. Extreme fast charging (XFC) technology can reduce the so-called ’range anxiety’ of consumers as they significantly reduce the charging time. With the advent of wide band-gap (WBG) power devices and improvement in power electronic converters, medium voltage (MV) solid state transformer (SST) based XFC system has the potential to replace the traditional XFC stations because of the lower footprint, ease of installation, enhanced control feature, and better system efficiency. The control system design is one of the critical aspects of the SST development process. Careful consideration and detailed analysis are required to find out suitable control method for the SST based on its topology among different centralized and decentralized control architectures. Also, the control parameters selection and potential improvement to the transient response of the controller ought to be investigated. Another major concern of the SST is different types of internal fault which reduces the overall reliability of the XFC system. As a result, designing a robust protection system is essential. Among different fault modes, open circuit switch faults have received significant attention as an active research area because of their likelihood and severe effects on converters. Therefore, the power stages used in the XFC system require functional and accurate open circuit switch fault management methods. An equally significant aspect of this SST based XFC is its compatibility in a microgrid where there is no synchronous generator present. When the grid is not available, the XFC SSTs can provide grid forming capability and continue supplying the critical loads in islanded mode. The transition between grid connected and islanded mode, especially the grid resynchronization process has to be carefully performed for the safety of the microgrid components. The challenges posed by the aforementioned issues have inspired the work done in this dissertation. Here, a 13.2 kV, 1 MVA, AC/DC SST for the XFC system is examined and a comparative analysis is conducted to select the control architecture based on feasibility of implementation and performance. A detailed control parameter design process is demonstrated considering the sensor dynamics and delay. The selected decentralized control method is augmented by introducing a novel sensor-less load current feedforward method to provide better voltage regulation at the DC bus during a change of load. Next, in the fault management section, a hierarchical failure mode effect analysis (FMEA) is proposed to enable a systematic design of the internal fault protection of the XFC SST as there are limited examples in the literature regarding the analysis of the safety and design of the protection of a power electronic converter system. Novel open circuit switch fault management methods for the converters in the system are presented. Finally, XFC SST based MV microgrid operations in grid connected mode and islanded mode are explored. A secondary control method for grid resynchronization is presented and a design process of control parameters is shown to ensure the stability of the secondary voltage and frequency regulation.

30 DIRECT ENERGY CONVERSION↗

A Review of Diagnostic Techniques for ISHM Applications

System diagnosis is an integral part of any Integrated System Health Management application. Diagnostic applications make use of system information from the design phase, such as safety and mission assurance analysis, failure modes and effects analysis, hazards analysis, functional models, fault propagation models, and testability analysis. In modern process control and equipment monitoring systems, topological and analytic , models of the nominal system, derived from design documents, are also employed for fault isolation and identification. Depending on the complexity of the monitored signals from the physical system, diagnostic applications may involve straightforward trending and feature extraction techniques to retrieve the parameters of importance from the sensor streams. They also may involve very complex analysis routines, such as signal processing, learning or classification methods to derive the parameters of importance to diagnosis. The process that is used to diagnose anomalous conditions from monitored system signals varies widely across the different approaches to system diagnosis. Rule-based expert systems, case-based reasoning systems, model-based reasoning systems, learning systems, and probabilistic reasoning systems are examples of the many diverse approaches ta diagnostic reasoning. Many engineering disciplines have specific approaches to modeling, monitoring and diagnosing anomalous conditions. Therefore, there is no "one-size-fits-all" approach to building diagnostic and health monitoring capabilities for a system. For instance, the conventional approaches to diagnosing failures in rotorcraft applications are very different from those used in communications systems. Further, online and offline automated diagnostic applications are integrated into an operations framework with flight crews, flight controllers and maintenance teams. While the emphasis of this paper is automation of health management functions, striking the correct balance between automated and human-performed tasks is a vital concern.

Patterson-Hine, Ann↗

Weapon Systems Risk-Assessment Tool Review

To anticipate, and potentially mitigate, future problems in aging weapon systems, four unique riskassessment techniques were analyzed, including root-cause analysis tools, six-sigma problemsolving approaches, and lean six-sigma tools. Identifying the most efficient process, or tool, is crucial for successful application to current and future weapon systems, subsystems, and components. The following processes were reviewed: Fault Tree Analysis, Failure Modes and Effects Analysis, Bow-Tie Analysis, and Hazard and Operability Study. A systematic assessment was performed to determine the most desirable method, and included investigating qualitative versus quantitative characteristics, scope, process durations, advantages, and limitations. Presented results will outline the study findings and further illustrate a capacity to identify future issues and/or concerns, and ultimately, reduce risk.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Assessment of Potential Failure Modes and Effects for On-Board Components for Hydrogen-Powered Locomotives

Hydrogen fuel sources offer alternatives to conventional fuels in the rail transportation industry. Hydrogen powered locomotive designs utilizing either a fuel cell or an internal combustion engine can make migration to alternative fuels possible for rail transportation. Codes and standards are still in development for rail application of hydrogen and safety risks must be assessed for hydrogen locomotive applications. This report utilizes a failure mode & effects analysis framework to help qualitatively understand possible risks from a hydrogen locomotive system. Findings illustrate how a combination of three mitigations greatly reduces the risks from a hydrogen locomotive system.

08 HYDROGEN↗