Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Failure mode and effect analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Procedure for Failure Mode, Effects, and Criticality Analysis (FMECA)

This document provides guidelines for the accomplishment of Failure Mode, Effects, and Criticality Analysis (FMECA) on the Apollo program. It is a procedure for analysis of hardware items to determine those items contributing most to system unreliability and crew safety problems.

Source record↗

Contamination Sources Effects Analysis (CSEA) - A Tool to Balance Cost/Schedule While Managing Facility Availability

A CSEA is similar to a Failure Modes Effects Analysis (FMEA). A CSEA tracks risk, deterrence, and occurrence of sources of contamination and their mitigation plans. Documentation is provided spanning mechanical and electrical assembly, precision cleaning, thermal vacuum bake-out, and thermal vacuum testing. These facilities all may play a role in contamination budgeting and reduction ultimately affecting test and flight. With a CSEA, visibility can be given to availability of these facilities, test sequencing and trade-offs. A cross-functional team including specialty engineering, contamination control, electrostatic dissipation, manufacturing, testing, and material engineering participate in an exercise that identifies contaminants and minimizes the complexity of scheduling these facilities considering their volatile schedules. Care can be taken in an efficient manner to insure correct cleaning processes are employed. The result is reduction in cycle time ("schedule hits"), reduced cost due to rework, reduced risk and improved communication and quality while achieving adherence to the Contamination Control Plan.

Wilcox, Margaret↗

Service Life Extension of the Propulsion System of Long-Term Manned Orbital Stations

One of the critical non-replaceable systems of a long-term manned orbital station is the propulsion system. Since the propulsion system operates beginning with the launch of station elements into orbit, its service life determines the service life of the station overall. Weighing almost a million pounds, the International Space Station (ISS) is about four times as large as the Russian space station Mir and about five times as large as the U.S. Skylab. Constructed over a span of more than a decade with the help of over 100 space flights, elements and modules of the ISS provide more research space than any spacecraft ever built. Originally envisaged for a service life of fifteen years, this Earth orbiting laboratory has been in orbit since 1998. Some elements that have been launched later in the assembly sequence were not yet built when the first elements were placed in orbit. Hence, some of the early modules that were launched at the inception of the program were already nearing the end of their design life when the ISS was finally ready and operational. To maximize the return on global investments on ISS, it is essential for the valuable research on ISS to continue as long as the station can be sustained safely in orbit. This paper describes the work performed to extend the service life of the ISS propulsion system. A system comprises of many components with varying failure rates. Reliability of a system is the probability that it will perform its intended function under encountered operating conditions, for a specified period of time. As we are interested in finding out how reliable a system would be in the future, reliability expressed as a function of time provides valuable insight. In a hypothetical bathtub shaped failure rate curve, the failure rate, defined as the number of failures per unit time that a currently healthy component will suffer in a given future time interval, decreases during infant-mortality period, stays nearly constant during the service life and increases at the end when the design service life ends and wear-out phase begins. However, the component failure rates do not remain constant over the entire cycle life. The failure rate depends on various factors such as design complexity, current age of the component, operating conditions, severity of environmental stress factors, etc. Development, qualification and acceptance test processes provide rigorous screening of components to weed out imperfections that might otherwise cause infant mortality failures. If sufficient samples are tested to failure, the failure time versus failure quantity can be analyzed statistically to develop a failure probability distribution function (PDF), a statistical model of the probability of failure versus time. Driven by cost and schedule constraints however, spacecraft components are generally not tested in large numbers. Uncertainties in failure rate and remaining life estimates increase when fewer units are tested. To account for this, spacecraft operators prefer to limit useful operations to a period shorter than the maximum demonstrated service life of the weakest component. Running each component to its failure to determine the maximum possible service life of a system can become overly expensive and impractical. Spacecraft operators therefore, specify the required service life and an acceptable factor of safety (FOS). The designers use these requirements to limit the life test duration. Midway through the design life, when benefits justify additional investments, supplementary life test may be performed to demonstrate the capability to safely extend the service life of the system. An innovative approach is required to evaluate the entire system, without having to go through an elaborate test program of propulsion system elements. Evaluating every component through a brute force test program would be a cost prohibitive and time consuming endeavor. ISS propulsion system components were designed and built decades ago. There are no representative ground test articles for some of the components. A 'test everything' approach would require manufacturing new test articles. The paper outlines some of the techniques used for selective testing, by way of cherry picking candidate components based on failure mode effects analysis, system level impacts, hazard analysis, etc. The type of testing required for extending the service life depends on the design and criticality of the component, failure modes and failure mechanisms, life cycle margin provided by the original certification, operational and environmental stresses encountered, etc. When specific failure mechanism being considered and the underlying relationship of that mode to the stresses provided in the test can be correlated by supporting analysis, time and effort required for conducting life extension testing can be significantly reduced. Exposure to corrosive propellants over long periods of time, for instance, lead to specific failure mechanisms in several components used in the propulsion system. Using Arrhenius model, which is tied to chemically dependent failure mechanisms such as corrosion or chemical reactions, it is possible to subject carefully selected test articles to accelerated life test. Arrhenius model reflects the proportional relationship between time to failure of a component and the exponential of the inverse of absolute temperature acting on the component. The acceleration factor is used to perform tests at higher stresses that allow direct correlation between the times to failure at a high test temperature to the temperatures to be expected in actual use. As long as the temperatures are such that new failure mechanisms are not introduced, this becomes a very useful method for testing to failure a relatively small sample of items for a much shorter amount of time. In this article, based on the example of the propulsion system of the first ISS module Zarya, theoretical approaches and practical activities of extending the service life of the propulsion system are reviewed with the goal of determining the maximum duration of its safe operation.

Kamath, Ulhas↗

Combining System Safety and Reliability to Ensure NASA CoNNeCT's Success

Hazard Analysis, Failure Modes and Effects Analysis (FMEA), the Limited-Life Items List (LLIL), and the Single Point Failure (SPF) List were applied by System Safety and Reliability engineers on NASA's Communications, Navigation, and Networking reConfigurable Testbed (CoNNeCT) Project. The integrated approach involving cross reviews of these reports by System Safety, Reliability, and Design engineers resulted in the mitigation of all identified hazards. The outcome was that the system met all the safety requirements it was required to meet.

Havenhill, Maria↗

CompactPCI(Registered TradeMark) Connectors in Space Flight Applications

This report documents the current status of CompactPCI(Registered TradeMark) connectors in GSFC spaceflight applications. To the extent the information is known, this report summarizes to what component quality level each NASA contractor (referred to as OEM in this report) procured the parts, and what board level and system level testing was performed. The report also provides the current status of the reliability assessment for each GSFC project based on the results of testing and FMEA (Failure Mode Effects Analysis). This report addresses how the CompactPCI(Registered TradeMark) connectors came into existence, and how these became the connector style chosen by many designers of space flight hardware. It identifies the design philosophy and the lack of robustness which has led to several known failure modes. These failure modes include fretting of connector pins during vibration, shock and thermal cycling, exposure of underplating, and increased resistance, including brief excursions to very high resistance. Each of these are signs of aging, which becomes an increasing concern for long duration orbiting space flight applications. This report addresses the mitigation strategy to replace CompactPCI(Registered TradeMark) connectors with space qualified Hypertronics 2mm cPCI connectors. The Hypertronics 2mm cPCI connectors are pin-to-pin compatible with the CompactPCI(Registered TradeMark) connectors and meet all of the same technical requirements, except the ability to hot mate, and to mate directly with a CompactPCI of the opposite gender. A detailed comparison of the CompactPCI(Registered TradeMark) connector and the Hypertronics 2mm cPCI connector is provided to describe the ruggedness of Hypertronics connector for space flight applications. Finally, this report makes recommendations for flight hardware for the future missions where the hardware is yet to be built, as well as for the hardware which has already been built with CompactPCI(Registered TradeMark) connectors.

Williams, Richard↗

Space Shuttle Main Engine Quantitative Risk Assessment: Illustrating Modeling of a Complex System with a New QRA Software Package

During 1997, a team from Hernandez Engineering, MSFC, Rocketdyne, Thiokol, Pratt & Whitney, and USBI completed the first phase of a two year Quantitative Risk Assessment (QRA) of the Space Shuttle. The models for the Shuttle systems were entered and analyzed by a new QRA software package. This system, termed the Quantitative Risk Assessment System(QRAS), was designed by NASA and programmed by the University of Maryland. The software is a groundbreaking PC-based risk assessment package that allows the user to model complex systems in a hierarchical fashion. Features of the software include the ability to easily select quantifications of failure modes, draw Event Sequence Diagrams(ESDs) interactively, perform uncertainty and sensitivity analysis, and document the modeling. This paper illustrates both the approach used in modeling and the particular features of the software package. The software is general and can be used in a QRA of any complex engineered system. The author is the project lead for the modeling of the Space Shuttle Main Engines (SSMEs), and this paper focuses on the modeling completed for the SSMEs during 1997. In particular, the groundrules for the study, the databases used, the way in which ESDs were used to model catastrophic failure of the SSMES, the methods used to quantify the failure rates, and how QRAS was used in the modeling effort are discussed. Groundrules were necessary to limit the scope of such a complex study, especially with regard to a liquid rocket engine such as the SSME, which can be shut down after ignition either on the pad or in flight. The SSME was divided into its constituent components and subsystems. These were ranked on the basis of the possibility of being upgraded and risk of catastrophic failure. Once this was done the Shuttle program Hazard Analysis and Failure Modes and Effects Analysis (FMEA) were used to create a list of potential failure modes to be modeled. The groundrules and other criteria were used to screen out the many failure modes that did not contribute significantly to the catastrophic risk. The Hazard Analysis and FMEA for the SSME were also used to build ESDs that show the chain of events leading from the failure mode occurence to one of the following end states: catastrophic failure, engine shutdown, or siccessful operation( successful with respect to the failure mode under consideration).

Smart, Christian↗

A Review of Diagnostic Techniques for ISHM Applications

System diagnosis is an integral part of any Integrated System Health Management application. Diagnostic applications make use of system information from the design phase, such as safety and mission assurance analysis, failure modes and effects analysis, hazards analysis, functional models, fault propagation models, and testability analysis. In modern process control and equipment monitoring systems, topological and analytic , models of the nominal system, derived from design documents, are also employed for fault isolation and identification. Depending on the complexity of the monitored signals from the physical system, diagnostic applications may involve straightforward trending and feature extraction techniques to retrieve the parameters of importance from the sensor streams. They also may involve very complex analysis routines, such as signal processing, learning or classification methods to derive the parameters of importance to diagnosis. The process that is used to diagnose anomalous conditions from monitored system signals varies widely across the different approaches to system diagnosis. Rule-based expert systems, case-based reasoning systems, model-based reasoning systems, learning systems, and probabilistic reasoning systems are examples of the many diverse approaches ta diagnostic reasoning. Many engineering disciplines have specific approaches to modeling, monitoring and diagnosing anomalous conditions. Therefore, there is no "one-size-fits-all" approach to building diagnostic and health monitoring capabilities for a system. For instance, the conventional approaches to diagnosing failures in rotorcraft applications are very different from those used in communications systems. Further, online and offline automated diagnostic applications are integrated into an operations framework with flight crews, flight controllers and maintenance teams. While the emphasis of this paper is automation of health management functions, striking the correct balance between automated and human-performed tasks is a vital concern.

Patterson-Hine, Ann↗

The ac propulsion system for an electric vehicle, phase 1

A functional prototype of an electric vehicle ac propulsion system was built consisting of a 18.65 kW rated ac induction traction motor, pulse width modulated (PWM) transistorized inverter, two speed mechanically shifted automatic transmission, and an overall drive/vehicle controller. Design developmental steps, and test results of individual components and the complex system on an instrumented test frame are described. Computer models were developed for the inverter, motor and a representative vehicle. A preliminary reliability model and failure modes effects analysis are given.

Geppert, S.↗

Demonstration Advanced Avionics System (DAAS). Phase 1 report

An integrated avionics system which provides expanded functional capabilities that significantly enhance the utility and safety of general aviation at a cost commensurate with the general aviation market is discussed. Displays and control were designed so that the pilot can use the system after minimum training. Functional and hardware descriptions, operational evaluation and failure modes effects analysis are included.

Source record↗

NASA GRC Technology Development Project for a Stirling Radioisotope Power System

NASA Glenn Research Center (GRC), the Department of Energy (DOE), and Stirling Technology Company (STC) are developing a Stirling convertor for an advanced radioisotope power system to provide spacecraft on-board electric power for NASA deep space missions. NASA GRC is conducting an in-house project to provide convertor, component, and materials testing and evaluation in support of the overall power system development. A first characterization of the DOE/STC 55-We Stirling Technology Demonstration Convertor (TDC) under the expected launch random vibration environment was recently completed in the NASA GRC Structural Dynamics Laboratory. Two TDCs also completed an initial electromagnetic interference (EMI) characterization at NASA GRC while being tested in a synchronized, opposed configuration. Materials testing is underway to support a life assessment of the heater head, and magnet characterization and aging tests have been initiated. Test facilities are now being established for an independent convertor performance verification and technology development. A preliminary Failure Mode Effect Analysis (FMEA), initial finite element analysis (FEA) for the linear alternator, ionizing radiation survivability assessment, and radiator parametric study have also been completed. This paper will discuss the status, plans, and results to date for these efforts.

Thieme, Lanny G.↗

Identifying, Assessing, and Mitigating Risk of Single-Point Inspections on the Space Shuttle Reusable Solid Rocket Motor

In the production of each Space Shuttle Reusable Solid Rocket Motor (RSRM), over 100,000 inspections are performed. ATK Thiokol Inc. reviewed these inspections to ensure a robust inspection system is maintained. The principal effort within this endeavor was the systematic identification and evaluation of inspections considered to be single-point. Single-point inspections are those accomplished on components, materials, and tooling by only one person, involving no other check. The purpose was to more accurately characterize risk and ultimately address and/or mitigate risk associated with single-point inspections. After the initial review of all inspections and identification/assessment of single-point inspections, review teams applied risk prioritization methodology similar to that used in a Process Failure Modes Effects Analysis to derive a Risk Prioritization Number for each single-point inspection. After the prioritization of risk, all single-point inspection points determined to have significant risk were provided either with risk-mitigating actions or rationale for acceptance. This effort gave confidence to the RSRM program that the correct inspections are being accomplished, that there is appropriate justification for those that remain as single-point inspections, and that risk mitigation was applied to further reduce risk of higher risk single-point inspections. This paper examines the process, results, and lessons learned in identifying, assessing, and mitigating risk associated with single-point inspections accomplished in the production of the Space Shuttle RSRM.

Greenhalgh, Phillip O.↗

Space tug propulsion system failure mode, effects and criticality analysis

For purposes of the study, the propulsion system was considered as consisting of the following: (1) main engine system, (2) auxiliary propulsion system, (3) pneumatic system, (4) hydrogen feed, fill, drain and vent system, (5) oxygen feed, fill, drain and vent system, and (6) helium reentry purge system. Each component was critically examined to identify possible failure modes and the subsequent effect on mission success. Each space tug mission consists of three phases: launch to separation from shuttle, separation to redocking, and redocking to landing. The analysis considered the results of failure of a component during each phase of the mission. After the failure modes of each component were tabulated, those components whose failure would result in possible or certain loss of mission or inability to return the Tug to ground were identified as critical components and a criticality number determined for each. The criticality number of a component denotes the number of mission failures in one million missions due to the loss of that component. A total of 68 components were identified as critical with criticality numbers ranging from 1 to 2990.

Boyd, J. W.↗

Failure modes, effects and criticality analyses.

Failure mode, effects and criticality analyses were developed by NASA as a means of assuring that hardware built for space applications has the desired reliability characteristics. The failure mode and effects analysis is a qualitative reliability technique for systematically analyzing each possible failure mode within a hardware system, and identifying the resulting effect on that system, the mission and personnel. The criticality analysis is a quantitative procedure which ranks the critical failure modes according to their probability of occurrence. This paper describes the failure modes, effects analysis and the criticality analysis. It employs a simple hardware system, not related to the aerospace field, to illustrate the method. It encourages application of this type of analysis to industrial development programs outside the aerospace and defense complex.

Jordan, W. E.↗

Independent Orbiter Assessment (IOA): Analysis of the manned maneuvering unit

Results of the Independent Orbiter Assessment (IOA) of the Failure Modes and Effects Analysis (FMEA) and Critical Items List (CIL) are presented. The IOA approach features a top-down analysis of the hardware to determine failure modes, criticality, and potential critical items (PCIs). To preserve indepedence, this analysis was accomplished without reliance upon the results contained within the NASA FMEA/CIL documentation. This report documents the independent analysis results corresponding to the Manned Maneuvering Unit (MMU) hardware. The MMU is a propulsive backpack, operated through separate hand controllers that input the pilot's translational and rotational maneuvering commands to the control electronics and then to the thrusters. The IOA analysis process utilized available MMU hardware drawings and schematics for defining hardware subsystems, assemblies, components, and hardware items. Final levels of detail were evaluated and analyzed for possible failure modes and effects. Criticality was assigned based upon the worst case severity of the effect for each identified failure mode. The IOA analysis of the MMU found that the majority of the PCIs identified are resultant from the loss of either the propulsion or control functions, or are resultant from inability to perform an immediate or future mission. The five most severe criticalities identified are all resultant from failures imposed on the MMU hand controllers which have no redundancy within the MMU.

Bailey, P. S.↗

Making the Hubble Space Telescope servicing mission safe

The implementation of the HST system safety program is detailed. Numerous safety analyses are conducted through various phases of design, test, and fabrication, and results are presented to NASA management for discussion during dedicated safety reviews. Attention is given to the system safety assessment and risk analysis methodologies used, i.e., hazard analysis, fault tree analysis, and failure modes and effects analysis, and to how they are coupled with engineering and test analysis for a 'synergistic picture' of the system. Some preliminary safety analysis results, showing the relationship between hazard identification, control or abatement, and finally control verification, are presented as examples of this safety process.

Bahr, N. J.↗

Manned testing in a simulated space environment

A view of the facility and operational requirements involved in performing a manned thermal vacuum test is presented. The requirements fall into two major categories. The first category deals with placing the suited crewmen in a hazardous environment and assuring their safety. The second category deals with the constraints and special requirements involved with a suited crewman operating flight hardware in a 1-G environment. Design areas that deal with man rating a chamber, including fire suppression, emergency repress, emergency power, backups, reliable instrumentation and data systems, communications, television monitoring, biomedical monitoring, material compatibilities, and equipment supporting the Extravehicular Mobility Unit (EMU) are discussed. The operational issues that are peculiar to manned testing such as test rules, test procedures, test protocol, emergency drills, availability of hyperbaric facilities, test team training and certification engineering concerns for a safe mechanical and instrumentation buildup, hazard analysis, and Failure Modes and Effects Analysis are discussed. The constraints and special requirements involved with a suited crewman operating flight hardware in a 1-G environment are addressed.

Fender, Donna L.↗