Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Failure Rate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Using the System Complexity Metric (SCM) to Compare CO2 Removal Systems

A fundamental cause of difficulty in large engineering projects is their inherent complexity. An impression of complexity occurs if a system is simply difficult to understand, where there is no obvious mental model that correctly predicts its behavior. Higher system complexity is usually associated with higher cost and higher failure rate. Complexity is perceived if a system has many diverse components, multiple interactions and feedback loops, transients and dynamic behavior, and unanticipated failure modes. Identifying and removing these signs of complexity should improve performance and reduce the cost and failure rate. Complexity can be directly measured by the number of components and their interactions. The System Complexity Metric (SCM) is defined as the sum of the number of parts in a system, N, plus the number of the one-way interconnections between them, I. SCM = N + I. The SCM is easily determined by direct inspection of the system block diagram. SCM can be used to compare systems and to guide their redesign to reduce cost and failure rate. Carbon dioxide removal systems are analyzed using SCM, cost, and failure rate. As in previous work, cost is directly proportional to SCM and that failure rate increases as a power of SCM for large differences in SCM. The SCM ranking of carbon dioxide removal systems is the same as their ranking in detailed analysis and practice.

Harry W. Jones↗

Using the System Complexity Metric (SCM) to Compare CO2 Reduction Systems

A fundamental cause of difficulty in large engineering projects is their inherent complexity. An impression of complexity occurs if a system is simply difficult to understand, where there is no obvious mental model that correctly predicts its behavior. Higher system complexity is usually associated with higher cost and higher failure rate. Complexity is perceived if a system has many diverse components, multiple interactions and feedback loops, transients and dynamic behavior, and unanticipated failure modes. Identifying and removing these signs of complexity should improve performance and reduce the cost and failure rate. Complexity can be directly measured by the number of components and their interactions. The System Complexity Metric (SCM) is defined as the sum of the number of parts in a system, N, plus the number of the one-way interconnections between them, I. SCM = N + I. The SCM is easily determined by direct inspection of the system block diagram. SCM can be used to compare systems and to guide their redesign to reduce cost and failure rate. Carbon dioxide reduction systems are analyzed using SCM, cost, and failure rate. As in previous work, cost is directly proportional to SCM and that failure rate increases as a power of SCM for large differences in SCM. The SCM ranking of carbon dioxide reduction systems is the same as their ranking in detailed analysis and practice.

Harry W. Jones↗

Redundancy: How Many Unreliable Spares are Needed for High Reliability and Confidence?

This paper investigates the number of redundant units needed to achieve high reliability with high confidence. The approach is developed for the case when the system failure rate is too high for a single unit to provide the required reliability over the mission duration. To achieve high reliability, N redundant units can be used, one operating unit and N – 1 spares. If the unit failure rate is f, the mission length is L, and f * L is small (not the case assumed here), the unit failure probability over the mission duration is F1 = f * L << 1. In this case, the probability that all N units will fail is Ffail = F1 N , and the needed redundancy N = LN(F)/LN(F1). For the case of large f * L assumed here, F1 = f * L > 1, and F1 is the expected number of failures during the mission. (When F1 = f * L << 1, F1 is the probability that a unit will fail during the mission. When F1 = f * L > 1, F1 is the expected number of failures during the mission.) The needed redundancy, N, to achieve the required N redundant unit reliability, FN, can be computed using the cumulative Poisson distribution with mean equal to F1. The number of spares, N - 1, is increased until the probability - that the total number of failures will be less than N -1 - is equal to the required reliability. The confidence that this reliability can be achieved can be computed using the cumulative Poisson distribution or the chi-square distribution. Since the measured unit failure rate, f, has some probabilistic uncertainty, the actual failure rate will be randomly higher or lower. This means that the reliability of the N redundant systems will be overestimated about half the time. Adding more redundant units increases the confidence that the required reliability will be achieved. For a fixed number of redundant units, the expected reliability and confidence can be traded off, since lower reliability goals will be achieved with higher confidence. Both the desired reliability and confidence can be specified as initial requirements and the needed number of redundant units estimated using the measured failure rate.

Redundancy↗

A system debugging model.

Consideration of the nature of the 'debugging' process applied to a new complex system during the initial period of its life. During this period failures and errors are corrected as they occur, with resulting improvement in the subsequent performance of the system. One mathematical idealization of this process leads to the assumption that system failure rate is decreasing with time. In practice, the debugging phase is considered completed when the failure rate reaches an equilibrium or constant value. Models are formulated for this phenomenon. Maximum likelihood estimates are obtained for relevant failure rate functions and for the end of the debugging period. A conservative upper confidence bound on the stable failure rate is obtained.

Barlow, R. E.↗

A Causal Approach to Integrate Component Health Data into System Reliability Models

Two of the challenges of current plant reliability approaches are the ability to integrate plant health data, and to support decision making. Condition based data and diagnostic/prognostic information are in fact not considered into plant reliability models to inform system engineers on the most critical components. Currently, the propagation of quantitative health data from the component to the system level is a challenge given the diverse nature/structure of the data. On the other hand, plant reliability methods (which are typically based on fault-trees or reliability block diagrams) can effectively propagate data from the component to the system level, but values of failure rates or failure probabilities are an approximated integral representation of the past industry-wide operational experience, and it neglects the present component health status (e.g., diagnostic and condition-based data) and health projection (when available from prognostic data). Our first claim is that system reliability models should propagate health information from the component to the system/plant level in order to provide a quantitative snapshot of system/plant health and identify the most critical components. Our second claim is that component health should be informed solely by that specific component current and historical performance data and should not be an approximated integral representation of the past industry-wide operational experience. This paper is directly supporting these two claims by proposing a different approach to perform reliability modeling which relies on available component diagnostic, prognostic and condition-based data to measure component health, and it propagates this information through fault tree models. The propagation of health data from the component to the system level is performed not in terms of probability, but in terms of margins where margin is defined as the “distance” between the present actual status and an undesired event (e.g., failure or unacceptable performance). Through a cause-effect lens, while classical reliability models target the effect associated to a component performance, a margin-based approach focuses on the cause of an undesired component performance (i.e., component health). Hence, thinking of reliability in terms of margins implies decision making based on causal reasoning. We will show how fault tree models can be solved using a margin language and how this process can effectively assist system engineers to identify the most critical components.

97 MATHEMATICS AND COMPUTING↗

The fault-tree compiler

The Fault Tree Compiler Program is a new reliability tool used to predict the top event probability for a fault tree. Five different gate types are allowed in the fault tree: AND, OR, EXCLUSIVE OR, INVERT, and M OF N gates. The high level input language is easy to understand and use when describing the system tree. In addition, the use of the hierarchical fault tree capability can simplify the tree description and decrease program execution time. The current solution technique provides an answer precise (within the limits of double precision floating point arithmetic) to the five digits in the answer. The user may vary one failure rate or failure probability over a range of values and plot the results for sensitivity analyses. The solution technique is implemented in FORTRAN; the remaining program code is implemented in Pascal. The program is written to run on a Digital Corporation VAX with the VMS operation system.

Martensen, Anna L.↗

The Fault Tree Compiler (FTC): Program and mathematics

The Fault Tree Compiler Program is a new reliability tool used to predict the top-event probability for a fault tree. Five different gate types are allowed in the fault tree: AND, OR, EXCLUSIVE OR, INVERT, AND m OF n gates. The high-level input language is easy to understand and use when describing the system tree. In addition, the use of the hierarchical fault tree capability can simplify the tree description and decrease program execution time. The current solution technique provides an answer precisely (within the limits of double precision floating point arithmetic) within a user specified number of digits accuracy. The user may vary one failure rate or failure probability over a range of values and plot the results for sensitivity analyses. The solution technique is implemented in FORTRAN; the remaining program code is implemented in Pascal. The program is written to run on a Digital Equipment Corporation (DEC) VAX computer with the VMS operation system.

Butler, Ricky W.↗

Reliability and Safety Assessment of Urban Air Mobility Concept Vehicles

The following work was commissioned by the National Aeronautics and Space Administration (NASA) to guide industry and future regulation related to urban air mobility (UAM). Prior studies compared the relative safety of NASA concept vehicles, Figure ES1, designed for UAM and provided recommendations for industry research, aircraft architectural improvements, and regulatory updates. After the prior study completed, the European Aviation Safety Agency (EASA) released regulatory guidance in the form of a special condition (SC), SC-VTOL-01, for multirotors with distributed propulsion and flight controls (DPFC). The objective of the current work was to develop DPFC architectures that will comply with SC-VTOL-01. Vehicle designs, DPFC architectures, and stability & control (S&C) models were developed to find limitations and trends to guide industry. To guide this task, NASA developed quad, hex, and an octorotor to better define vehicle attributes and trade space. Aircraft for study included electric, hybrid-electric, and turboshaft powerplants and collective and RPM control schemes. Assessments are in terms of the safety level achieved, and/or aircraft component/features needed to meet SC-VTOL-01. The most challenging criteria being the catastrophic failure rate, ≤10-9 catastrophic failures per flight hour, and that no single failures may result in a catastrophic event. A disciplined process was followed, similar to that in Aerospace Recommended Practice (ARP) 4761. A preliminary system safety assessment (PSSA) leveraged prior work as a basis of creating failure rate budgets for system design teams. System designs were updated and iterated upon, working with reliability and safety subject matter experts to develop SC-VTOL-01 compliant designs. Design changes were reflected in updated PSSAs for initial verification of compliance. The DPFC architecture was broken into four system design teams, the (1) flight control system (FCS), (2) drive and power system, (3) thermal management system (TMS), and (4) electrical power and distribution system. The FCS including elements necessary to control the aircraft, drive and power including elements necessary to generate and transmit shaft power, the TMS including elements necessary to maintain temperature limits in all operating environments, and electrical pow-er and distribution including equipment necessary to store and transmit electrical energy. Results found that all aircraft evaluated may have paths to comply with SC-VTOL-01, Figure ES2. However, S&C models showed large power transients that must be addressed and PSSA results show that future work is needed in single load path structures, high voltage power storage and distribution, and in motor/rotor overspeed protection.

Boeing↗

High Reliability Requires More than Providing Spares

It is sometimes optimistically hoped that a space life support system can be kept working throughout a long duration mission by repairing failed components, as long as sufficient spares are flown. It is usually assumed that the components have constant known failure rates. Then the needed numbers of spares can be computed to have any particular probability that all failed components can be replaced by available spares. This approach can provide high reliability if its favorable assumptions, including constant known failure rates, are satisfied. Other favorable assumptions are that the failures are statistically independent, repair will be successful without causing further failures, and all failures are due to internal component failures. These assumptions are not usually justified. The failure rates may be estimates that are inadequately verified because of insufficient testing. Failure rates may change due to materials substitutions, manufacturing changes, redesigns to fix failures, and new failures caused by redesigns. Failures that are not statistically independent may result from one common cause, such as a design or manufacturing error or a cascade of cause and effect, possibly caused by an external event such as a power outage. Repair may be unsuccessful or cause damage. Many failures occur at component interfaces or at the overall systems level, not within isolated components. Other failures causes are completely external to the system, due to assembly, maintenance, and operational errors or to unexpected environmental challenges. Replacement with sufficient spares can compensate for expected internal component failures but may not be able to cope with unpredictable design and manufacturing flaws, human errors, and environmental impacts. Reliability estimates based on providing sufficient spares to compensate for expected failures may be far too high. They are essentially upper bounds on reliability that might be approached if many frequent but often unconsidered failure causes can be eliminated.

spares↗

How Much Testing is Needed to Manage Supportability Risks for Beyond-LEO Missions?

Supportability will be a significantly greater driver of cost and risk for future deep-space crewed missions than it has been in the past. Spares requirements and maintenance risk mitigation in particular present an unprecedented challenge for missions beyond Low Earth Orbit (LEO), since, for the first time in human spaceflight history, crews will be weeks or months away from resupply or a safe return to Earth in the event of an abort. Under these conditions, failure rates are a critical parameter that must be well-understood in order to manage logistics and risk effectively. However, failure rates cannot be measured directly, and can only be estimated based on past experience and test results. Previous research has shown that International Space Station (ISS) operational experience has provided significant benefits to future missions by reducing uncertainty and improving accuracy in failure rate estimates, resulting in significant reductions in mass and risk for beyond-LEO missions. This paper updates and expands on that research and quantifies the potential value of continued testing for future mission supportability. Frequentist and Bayesian models for evaluating, validating, and updating failure rate estimates are described, and are combined with supportability models to examine potential impacts of additional operating experience for future missions in terms of logistics mass reduction. The implications of these results for technology development, system design, and program planning are discussed along with lessons learned and recommendations for future system development. In the end, there is no simple answer to the question of how much testing is required, but the models described in this paper provide a way to evaluate the potential impacts of testing in order to inform test planning.

Andrew C Owens↗

Applying the System Complexity Metric (SCM)

A fundamental cause of difficulty in larger engineering projects is their inherent complexity. An impression of complexity occurs if a system is simply difficult to understand, so that there is no obvious mental model that correctly predicts its behavior. Higher complexity is usually associated with higher cost and higher failure rate. Complexity is indicated by a system having more and diverse components, multiple interactions and feedback loops, transients and dynamic behavior, and often the emergence of unanticipated failure modes. Identifying and removing these signs of complexity should reduce complexity and improve performance. Here we limit complexity measurement to the number of components and their interactions. A System Complexity Metric (SCM) is defined as equal to the sum of the number of parts in a system, N, plus the sum of the one-way interconnections between them, I. SCM = N + I. The SCM is easily determined by direct inspection of system block diagrams. Previous work found that life support system cost was directly proportional to SCM and that failure rate increased faster than SCM squared. SCM can be used to compare systems or to guide their redesign to reduce cost and failure rate. Carbon dioxide removal systems will be analyzed using SCM, cost, and failure rate.

Harry W Jones↗

Payload maintenance cost model for the space telescope

An optimum maintenance cost model for the space telescope for a fifteen year mission cycle was developed. Various documents and subsequent updates of failure rates and configurations were made. The reliability of the space telescope for one year, two and one half years, and five years were determined using the failure rates and configurations. The failure rates and configurations were also used in the maintenance simulation computer model which simulate the failure patterns for the fifteen year mission life of the space telescope. Cost algorithms associated with the maintenance options as indicated by the failure patterns were developed and integrated into the model.

White, W. L.↗

Reliability analysis and fault-tolerant system development for a redundant strapdown inertial measurement unit

A methodology is developed and applied for quantitatively analyzing the reliability of a dual, fail-operational redundant strapdown inertial measurement unit (RSDIMU). A Markov evaluation model is defined in terms of the operational states of the RSDIMU to predict system reliability. A 27 state model is defined based upon a candidate redundancy management system which can detect and isolate a spectrum of failure magnitudes. The results of parametric studies are presented which show the effect on reliability of the gyro failure rate, both the gyro and accelerometer failure rates together, false alarms, probability of failure detection, probability of failure isolation, and probability of damage effects and mission time. A technique is developed and evaluated for generating dynamic thresholds for detecting and isolating failures of the dual, separated IMU. Special emphasis is given to the detection of multiple, nonconcurrent failures. Digital simulation time histories are presented which show the thresholds obtained and their effectiveness in detecting and isolating sensor failures.

Motyka, P.↗

A Novel Solution-Technique Applied to a Novel WAAS Architecture

The Federal Aviation Administration has embarked on an historic task of modernizing and significantly improving the national air transportation system. One system that uses the Global Positioning System (GPS) to determine aircraft navigational information is called the Wide Area Augmentation System (WAAS). This paper describes a reliability assessment of one candidate system architecture for the WAAS. A unique aspect of this study regards the modeling and solution of a candidate system that allows a novel cold sparing scheme. The cold spare is a WAAS communications satellite that is fabricated and launched after a predetermined number of orbiting satellite failures have occurred and after some stochastic fabrication time transpires. Because these satellites are complex systems with redundant components, they exhibit an increasing failure rate with a Weibull time to failure distribution. Moreover, the cold spare satellite build-time is Weibull and upon launch is considered to be a good-as-new system with an increasing failure rate and a Weibull time to failure distribution as well. The reliability model for this system is non-Markovian because three distinct system clocks are required: the time to failure of the orbiting satellites, the build time for the cold spare, and the time to failure for the launched spare satellite. A powerful dynamic fault tree modeling notation and Monte Carlo simulation technique with importance sampling are shown to arrive at a reliability prediction for a 10 year mission.

Bavuso, J.↗

Natural Hazard Forecast Alert Grid Risk System

Weather events cause most power outages. Often, we even get notifications on our phones to take cover or be prepared for an imminent event. If electric grid utilities had a similar warning that also included probable scenarios and the equipment involved, they could prepare and minimize the effects. Idaho National Laboratory had a project with the U.S. Department of Energy’s Cybersecurity, Energy Security, and Emergency Response program to develop a grid alert application that receives messages from the existing emergency alert system, filters and determines components possibly affected by the emergency event, calculates probable scenarios using MASTERRI (Modeling And Simulation for Targeted Reliability and Resilience Improvement). For high-risk events, the application can then send alert links to subscribed electric distribution utility operations staff to allow them to see and evaluate the scenarios and the impact in a web based interactive map tool. This proof of concept application used data from utilities and organizations, such as the international regulatory body North American Electric Reliability Corporation, which have complied historical failure data of elements that comprise the U.S. electric grid. Nominal failure rates are obtained from this data. To make this tool possible, estimated failure rates were calculated for different component types given the alert type, severity, and location. Historic weather-related grid element failures were correlated with historic weather events from the Integrated Public Alert & Warning System. These correlated events and failures are used along with Bayesian updates from the historical norms to provide a modified failure rate for grid elements in the alert areas and calculate probable scenarios. Working with an industry collaborator, actual grid models and data were used for demonstration cases. This report outlines the work performed for this project.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Enhanced Component Performance Study: Turbine Driven Pumps 1998–2022

This report presents an enhanced performance evaluation of turbine driven pumps (TDPs) at U.S. commercial nuclear power plants. The data used in this study are based on the operating experience failure reports from calendar year 1998 through 2022 as reported in the Institute of Nuclear Power Operations (INPO) Industry Reporting and Information System (IRIS). The TDP failure modes considered for standby systems are fail to start (FTS), fail to run (FTR) for one hour of operation (FTR=1H), FTR after one hour of operation (FTR>1H), and for normally running systems FTS and FTR. An eight hour unreliability estimate is also calculated and trended. The component reliability estimates and the reliability data are trended for the most recent 10 year period while yearly estimates for reliability are provided for the entire study period. No increasing trends were identified for TDPs for the most recent 10 year period: The following decreasing trends were identified for TDPs for the most recent 10 year period: • Standby TDP FTR>1H failure rate • Normally running TDP FTR failure rate • Standby TDP unavailability • Standby MDP total unreliability (8-hour mission) • Standby TDP frequency of start demands (demands per reactor year) • Standby TDP frequency of FTR=1H hours (hours per reactor year) • Standby TDP frequency of FTR>1H events (failures per reactor year) • Normally running TDP frequency of start demands • Normally running TDP frequency of run hours • Normally running TDP frequency of FTR events.

99 GENERAL AND MISCELLANEOUS↗

Reliability and Maintainability Analysis for the Amine Swingbed Carbon Dioxide Removal System

I have performed a reliability & maintainability analysis for the Amine Swingbed payload system. The Amine Swingbed is a carbon dioxide removal technology that has gone through 2,400 hours of International Space Station on-orbit use between 2013 and 2016. While the Amine Swingbed is currently an experimental payload system, the Amine Swingbed may be converted to system hardware. If the Amine Swingbed becomes system hardware, it will supplement the Carbon Dioxide Removal Assembly (CDRA) as the primary CO2 removal technology on the International Space Station. NASA is also considering using the Amine Swingbed as the primary carbon dioxide removal technology for future extravehicular mobility units and for the Orion, which will be used for the Asteroid Redirect and Journey to Mars missions. The qualitative component of the reliability and maintainability analysis is a Failure Modes and Effects Analysis (FMEA). In the FMEA, I have investigated how individual components in the Amine Swingbed may fail, and what the worst case scenario is should a failure occur. The significant failure effects are the loss of ability to remove carbon dioxide, the formation of ammonia due to chemical degradation of the amine, and loss of atmosphere because the Amine Swingbed uses the vacuum of space to regenerate the Amine Swingbed. In the quantitative component of the reliability and maintainability analysis, I have assumed a constant failure rate for both electronic and nonelectronic parts. Using this data, I have created a Poisson distribution to predict the failure rate of the Amine Swingbed as a whole. I have determined a mean time to failure for the Amine Swingbed to be approximately 1,400 hours. The observed mean time to failure for the system is between 600 and 1,200 hours. This range includes initial testing of the Amine Swingbed, as well as software faults that are understood to be non-critical. If many of the commercial parts were switched to military-grade parts, the expected mean time to failure would be 2,300 hours. Both calculated mean times to failure for the Amine Swingbed use conservative failure rate models. The observed mean time to failure for CDRA is 2,500 hours. Working on this project and for NASA in general has helped me gain insight into current aeronautics missions, reliability engineering, circuit analysis, and different cultures. Prior my internship, I did not have a lot knowledge about the work being performed at NASA. As a chemical engineer, I had not really considered working for NASA as a career path. By engaging in interactions with civil servants, contractors, and other interns, I have learned a great deal about modern challenges that NASA is addressing. My work has helped me develop a knowledge base in safety and reliability that would be difficult to find elsewhere. Prior to this internship, I had not thought about reliability engineering. Now, I have gained a skillset in performing reliability analyses, and understanding the inner workings of a large mechanical system. I have also gained experience in understanding how electrical systems work while I was analyzing the electrical components of the Amine Swingbed. I did not expect to be exposed to as many different cultures as I have while working at NASA. I am referring to both within NASA and the Houston area. NASA employs individuals with a broad range of backgrounds. It has been great to learn from individuals who have highly diverse experiences and outlooks on the world. In the Houston area, I have come across individuals from different parts of the world. Interacting with such a high number of individuals with significantly different backgrounds has helped me to grow as a person in ways that I did not expect. My time at NASA has opened a window into the field of aeronautics. After earning a bachelor's degree in chemical engineering, I plan to go to graduate school for a PhD in engineering. Prior to coming to NASA, I was not aware of the graduate Pathways program. I intend to apply for the graduate Pathways program as positions are opened up. I would like to pursue future opportunities with NASA, especially as my engineering career progresses.

Dunbar, Tyler↗

Repeated Induction of Inattentional Blindness in a Simulated Aviation Environment

The study reported herein is a subset of a larger investigation on the role of automation in the context of the flight deck and used a fixed-based, human-in-the-loop simulator. This paper explored the relationship between automation and inattentional blindness (IB) occurrences in a repeated induction paradigm using two types of runway incursions. The critical stimuli for both runway incursions were directly relevant to primary task performance. Sixty non-pilot participants performed the final five minutes of a landing scenario twice in one of three automation conditions: full automation (FA), partial automation (PA), and no automation (NA). The first induction resulted in a 70 percent (42 of 60) detection failure rate with those in the PA condition significantly more likely to detect the incursion compared to the FA condition or the NA condition. The second induction yielded a 50 percent detection failure rate. Although detection improved (detection failure rates declined) in all conditions, those in the FA condition demonstrated the greatest improvement with doubled detection rates. The detection behavior in the first trial did not preclude a failed detection in the second induction. Group membership (IB vs. Detection) in the FA condition showed a greater improvement than those in the NA condition and rated the Mental Demand and Effort subscales of the NASA-TLX (NASA Task Load Index) significantly higher for Time 2 compared Time 1. Participants in the FA condition used the experience of IB exposure to improve task performance whereas those in the NA condition did not, indicating the availability and reallocation of attentional resources in the FA condition. These findings support the role of engagement in operational attention detriment and the consideration of attentional failure causation to determine appropriate mitigation strategies.

Kennedy, Kellie D.↗