Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Failure Rate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

AN INITIAL LOOK AT THE HEAT PIPE RESILIENCY TO REACTOR OPERATION

Several special purpose reactor designs are aiming to utilize heat pipes due to its inherently passive features allowing safe heat removal to the power conversion systems. There is insufficient data on heat pipe survivability in reactor environments, although heat-pipe failures are predicted to have low failure rates. In this paper, we perform a coupled neutronics and thermo-mechanics analyses for a 45 kWth HALEU fueled, hydride moderated homogenous design to evaluate the transient effects during both startup and a heatpipe-failure scenario. The heatpipe-failure study includes both a single failure case and a failure-propagation (cascading) case. While the model only evaluates a homogenous core, the approach could be applied to more complex models for any heat pipe cooled small reactor.

99 GENERAL AND MISCELLANEOUS↗

Designing to Mitigate Food Growing Failures in Space

Future space life support systems may use crop plants to grow most of the crew s food. A harvest failure can reduce the food available for future consumption. If the previously stored food is insufficient to reach the next harvest, the crew may go hungry. This paper considers how the overall food supply system should be modified to cope with food production failures. The food supply concept for a mission will use grown food, or stored food, cIr both. The optimum food supply mix depends on the costs and failure probabilities of stored and grown food. A simple food system model assumes that either we obtain the nominal harvest or a failure occurs and no food is harvested. Given the probability that any particular harvest fails, it is easy to compute the expected number of failures and the total food shortfall over a mission. If some food is grown and the probability of harvest failure is high, a non-redundant system has an unacceptable likelihood that the crew will have no food for a full harvest period. Food supply reliability must be increased either by supplying more food initially or by increasing food production capacity. We can obtain a very reliable food supply even when the harvest failure rate is high. If the cost of growing food is much less than the cost of providing stored food, it is better to provide redundant food growing capacity than to increase initial storage. A more realistic biomass production failure model allows the harvest amount or time to vary around the nominal values, using stochastic modeling with repeated Monte Carlo simulation, but such failures have minor impact compared to a complete harvest failure.

Jones, Harry↗

AGR-5/6/7 Post-Irradiation Examination Plan

This plan describes the PIE activities to be carried out on AGR-5/6/7 fuel and the non-fuel components of the AGR-5/6/7 irradiation test train. PIE of AGR-5/6/7 fuel and test train components will provide data on fuel performance over a range of irradiation conditions (e.g., burnup, neutron fluence, and temperature), test fuel under postulated accident conditions, and support development of fuel performance and fission product transport models. Examination of the non-fuel test train components will give information on TRISO fuel and compact performance and information on releases of radioactive fission products from the fuel compacts. PIE of the fuel compacts is used for many purposes including, but not limited to: updating the irradiation thermal calculations, measuring burnup, determining retention/release of fission products from the compacts and the TRISO particles, determining TRISO layer failure rates, determining fuel performance at elevated accident temperatures, determining fuel performance under oxidizing atmospheres, determining the distribution of fission products within the compact and fuel particles, and observing compact and particle morphology to explain mechanisms of TRISO layer failure and/or fission product release/retention.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Reliability measurement during software development

During the development of data base software for a multi-sensor tracking system, reliability was measured. The failure ratio and failure rate were found to be consistent measures. Trend lines were established from these measurements that provided good visualization of the progress on the job as a whole as well as on individual modules. Over one-half of the observed failures were due to factors associated with the individual run submission rather than with the code proper. Possible application of these findings for line management, project managers, functional management, and regulatory agencies is discussed. Steps for simplifying the measurement process and for use of these data in predicting operational software reliability are outlined.

Hecht, H.↗

The determination of measures of software reliability

Measurement of software reliability was carried out during the development of data base software for a multi-sensor tracking system. The failure ratio and failure rate were found to be consistent measures. Trend lines could be established from these measurements that provide good visualization of the progress on the job as a whole as well as on individual modules. Over one-half of the observed failures were due to factors associated with the individual run submission rather than with the code proper. Possible application of these findings for line management, project managers, functional management, and regulatory agencies is discussed. Steps for simplifying the measurement process and for use of these data in predicting operational software reliability are outlined.

Maxwell, F. D.↗

Parts and Components Reliability Assessment: A Cost Effective Approach

System reliability assessment is a methodology which incorporates reliability analyses performed at parts and components level such as Reliability Prediction, Failure Modes and Effects Analysis (FMEA) and Fault Tree Analysis (FTA) to assess risks, perform design tradeoffs, and therefore, to ensure effective productivity and/or mission success. The system reliability is used to optimize the product design to accommodate today?s mandated budget, manpower, and schedule constraints. Stand ard based reliability assessment is an effective approach consisting of reliability predictions together with other reliability analyses for electronic, electrical, and electro-mechanical (EEE) complex parts and components of large systems based on failure rate estimates published by the United States (U.S.) military or commercial standards and handbooks. Many of these standards are globally accepted and recognized. The reliability assessment is especially useful during the initial stages when the system design is still in the development and hard failure data is not yet available or manufacturers are not contractually obliged by their customers to publish the reliability estimates/predictions for their parts and components. This paper presents a methodology to assess system reliability using parts and components reliability estimates to ensure effective productivity and/or mission success in an efficient manner, low cost, and tight schedule.

Lee, Lydia↗

Lessons Learned in Space Life Support System Testing

The earlier problems can be found and corrected, the easier and cheaper it is to fix them. Doing less testing saves cost and time but doing too little testing increases the risk of operational failures causing large costs and delays. Integrated test is necessary to determine if the subsystems work together and the overall architecture performs as intended. This report reviews the testing lessons learned from the NASA Systems Engineering Handbook, a National Research Council report, and five reviews of International Space Station (ISS) lessons learned. The five reviews all mention two important points. First, that testing should be performed on the final integrated system, one as close as possible to the intended flight system. Second, “test as you fly,” while operating as planned in an environment as close as possible to the expected flight environment. Other lessons are the need for extensive preflight ground testing, the need to establish and defend an adequate budget, the problems using protoflight hardware on ISS, and the benefit of having ISS as a zero gravity test bed. The major ISS life support systems, carbon dioxide, water recycling, and oxygen recovery, were protoflight systems with little testing before launch to ISS. The failure rates these systems have been much greater than predicted and this has caused dissatisfaction with the protoflight approach. The more costly traditional approach is building qualification and test units in addition to flight units. The test units are used to test, analyze, and fix failure modes. Other work shows that there is an optimum cost-effective intuitive appeal of a human ecosystem in space.

Life support↗

Lessons Learned in Space Life Support System Testing

The earlier problems can be found and corrected, the easier and cheaper it is to fix them. Doing less testing saves cost and time but doing too little testing increases the risk of operational failures causing large costs and delays. Integrated test is necessary to determine if the subsystems work together and the overall architecture performs as intended. This report reviews the testing lessons learned from the NASA Systems Engineering Handbook, a National Research Council report, and five reviews of International Space Station (ISS) lessons learned. The five reviews all mention two important points. First, that testing should be performed on the final integrated system, one as close as possible to the intended flight system. Second, “test as you fly,” while operating as planned in an environment as close as possible to the expected flight environment. Other lessons are the need for extensive preflight ground testing, the need to establish and defend an adequate budget, the problems using protoflight hardware on ISS, and the benefit of having ISS as a zero gravity test bed. The major ISS life support systems, carbon dioxide removal, water recycling, and oxygen recovery, were protoflight systems with little testing before launch to ISS. The failure rates of these systems have been much greater than predicted and this has caused dissatisfaction with the protoflight approach. The more costly traditional approach builds qualification and test units in addition to flight units. The test units are used to find, analyze, and fix failure modes. Other work shows that there is an optimum cost-effective amount of testing when redundant systems must have a specified reliability and confidence.

Harry W. Jones↗

NASA Physics of Failure (PoF) for Reliability

An item’s reliability or longevity is dependent not only on its design but also on how it is used, manufactured, tested, and the stresses it has or will experience. Stresses include operational and environmental exposures to thermal, voltage, current, age/exposure, mechanical, and radiation mechanisms. Therefore, in reliability analysis, it is important to consider the contributions of all of these factors when predicting the failure rates of components. Historically, there has been a reliance on handbook data (e.g., MIL-HDBK-217), but experience has shown that these values and distributions are not representative of actual performance (1,2). Therefore, to make more credible reliability and risk assessments for its missions, NASA must transition to estimating likelihoods of failure based on an item’s reliability/longevity factors (or the physical susceptibilities and strengths impacting the design’s performance) has or will experience, whenever possible. To facilitate this transition a “Handbook on Methodology for Physics of Failure Based Reliability Assessments” has been developed by NASA to assist in applying physics experiences or experiment physics for empirical analysis and conceptualized physics exposures or theoretical physics for deterministic analysis, to develop and aggregate realistic likelihoods of failure leading to more credible forecasts of item performance and longevity. In addition, since it is NASA’s intention that this document continues to evolve based on community lessons learned and the introduction of new assessment methodologies, NASA is encouraging and appreciates the contributions of current and future authors to maintain and enhance this handbook and its supporting case studies.

Physics of Failure↗

Comparison Modeling of System Reliability for Future NASA Projects

A National Aeronautics and Space Administration (NASA) supported Reliability, Maintainability, and Availability (RMA) analysis team developed a unique RMA analysis methodology using cut set and importance measure analysis in order to comparison model proposed avionics computing architectures. In this paper we will present this efficient application of the RMA analysis methodology for importance measures that includes Reliability Block Diagram (RED) Analysis, Comparison modeling, Cut Set Analysis, and Importance Measure Analysis. We will also demonstrate that integrating RMA early in the system design process as a key to success by providing a fundamental decision metric supporting design selection. The RMA analysis methodology presented in this paper and applied to the avionics architectures enhances the usual way of predicting the need for redundancy based on failure rates or subject matter expert opinion. Using the REDs and the minimal cut sets, along with the Fussell-Vesely (FV) factors, importance measures are calculated for each functional element in the architectures. This paper presents an application of the FV importance measures and presents an improved methodology for using importance measures in success space (instead of failure space) to compare architectures. These importance measures are used to determine which functional element would be most likely to cause a system failure, thus, quickly identifying the path to increase the overall system reliability by either procuring more reliable functional elements or adding redundancy. This application of the RMA analysis methodology, using RBD analysis, cut set analysis, and the importance measure analysis, allows the avionics design team to better understand and compare the vulnerabilities in each of the architectures, enabling them to address the deficiencies in the design architectures more efficiently, while balancing the need to design for optimum weight and space allocations.

Gillespie, Amanda M.↗

AC and DC Fault Management for Megawatt Electrified Aircraft Electrical Powertrains Task 3: Lifetime and Reliability of Electrical Insulators

This research project was a collaborative investigation between researchers at the RTX Technology Research Center (RTRC) and the University of Texas at Austin and made a significant contribution to enabling electric aircraft. The transport of electric power between the points of generation and use requires power cables. These cables must be smaller, lighter and provide a more predictable life than power cables used in stationary applications. Consequently, this investigation provided heretofore unavailable information supporting the safety and reliability of smaller lighter power cables for electrified aircraft. In addition, the research identified key additional engineering data needed to support quantitative reliability assessments. Important advances included: • Demonstrated that at least one manufacturer can make a novel, smaller, lighter power cable that is free from serious defects. • Developed and published an appropriate analytical construct to describe the life of this novel cable. This is a necessary step for use in aviation where the understanding of remaining life is critical. • Demonstrated thermal-mechanical aging that suggested 1000+ flights before the thermal-mechanical processes produced defects large enough that the defect growth was accelerated electrically. • Showed that electrical aging took place at two rates. The first possibly lasting weeks to months and the second possibly days to weeks. If robust, this provides a good diagnostic for cable replacement. • Demonstrated that the traditional electrical testing of cable materials using manufactured voids can be misleading due to the size of the voids. Emerging laser drilling technology permitted demonstration that the physics of failure in realistically small voids is different from that in the unrealistically large voids used in earlier research, which is very important for high-quality, high-performance, small aircraft cables. Although this project represents a significant contribution to the specifics of cable aging in the aircraft environment, important additional research remains to be completed, including: • Non-uniform thermal cycling by applying the heat from the center conductor to maximize thermal stress next to the core area where the electric gradient is the strongest. This builds on the uniform thermal cycling that has been completed. • The augmentation of the thermal-mechanical failure rate by electrical processes. Better understanding of these time constants strongly affects the ability to predict life. • Termination design: Terminations provide not only electrical reflection potential, but a location for a series arc fault and an area where ozone can diffuse into the center conductor and negatively affect cable insulation. • The abrasion and ozone resistance of the cable jacket. • Pressure cycling as an accelerant of thermal, mechanical, and/or electrical aging. • Possible methods for online PD detection and offline PD localization

model↗

AGR-5/6/7 Irradiation Summary as of the End of Cycle 167A

Overall, the first three cycles of irradiation proceeded as planned Items of note: Failure rate of thermocouples was higher than expected Downstream needle valves were throttled resulting in part of the exhaust flow being diverted prior to reaching the fission product monitors (this was corrected near the end of cycle 164B). This was more of annoyance than anything and did not prevent proper analysis of fission product data. No evidence of failures on fission product monitoring system.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Probabilistic Analysis of Space Shuttle Body Flap Actuator Ball Bearings

A probabilistic analysis, using the 2-parameter Weibull-Johnson method, was performed on experimental life test data from space shuttle actuator bearings. Experiments were performed on a test rig under simulated conditions to determine the life and failure mechanism of the grease lubricated bearings that support the input shaft of the space shuttle body flap actuators. The failure mechanism was wear that can cause loss of bearing preload. These tests established life and reliability data for both shuttle flight and ground operation. Test data were used to estimate the failure rate and reliability as a function of the number of shuttle missions flown. The Weibull analysis of the test data for a 2-bearing shaft assembly in each body flap actuator established a reliability level of 99.6 percent for a life of 12 missions. A probabilistic system analysis for four shuttles, each of which has four actuators, predicts a single bearing failure in one actuator of one shuttle after 22 missions (a total of 88 missions for a 4-shuttle fleet). This prediction is comparable with actual shuttle flight history in which a single actuator bearing was found to have failed by wear at 20 missions.

Oswald, Fred B.↗

Reliability metrics and their management implications for open pond algae cultivation

The prevalence of contaminating organisms in outdoor algae cultivation, with the often-associated dramatic crop failures, necessitates the need for metrics describing production system reliability. Standard metrics for algae cultivation reliability are critically needed to be able to map improvements in operational parameters, but do not currently exist. In this work, we present a set of standard metrics including mean time to failure (MTTF) and mean time between failures (MTBF) as the basis for the calculation of pond failure rate (FR) and reliability coefficient (RC). Metrics associated with the numbers of contaminating organisms such as abundance ratio (AR), prevalence of infection (PI), and mean intensity of infection (MII) are also relevant. These metrics, based on measured experimental values of outdoor pond performance during the Algae Testbed Public-Private Partnership (ATP3) Unified Field Studies (UFS), provide a basis for quantifying pond failure and provide insight into potential pond management and contaminant mitigation strategies. From these reliability metrics applied to this dataset, we are able to link operational parameters, such as harvest frequency and inoculum source, to pond reliability for different algae strains and seasons, and from an assessment of AR, provide contamination thresholds beyond which a culture may be unrecoverable. Ultimately, the implementation of widely-practiced and simple-to-calculate algae pond reliability metrics calculated from rapid and easy to collect data or observations will reduce risk and uncertainty in large-scale algae deployment and aid in the development of integrated pest management (IPM) strategies.

09 BIOMASS FUELS↗

Probabilistic Analysis of Space Shuttle Body Flap Actuator Ball Bearings

A probabilistic analysis, using the 2-parameter Weibull-Johnson method, was performed on experimental life test data from space shuttle actuator bearings. Experiments were performed on a test rig under simulated conditions to determine the life and failure mechanism of the grease lubricated bearings that support the input shaft of the space shuttle body flap actuators. The failure mechanism was wear that can cause loss of bearing preload. These tests established life and reliability data for both shuttle flight and ground operation. Test data were used to estimate the failure rate and reliability as a function of the number of shuttle missions flown. The Weibull analysis of the test data for the four actuators on one shuttle, each with a 2-bearing shaft assembly, established a reliability level of 96.9 percent for a life of 12 missions. A probabilistic system analysis for four shuttles, each of which has four actuators, predicts a single bearing failure in one actuator of one shuttle after 22 missions (a total of 88 missions for a 4-shuttle fleet). This prediction is comparable with actual shuttle flight history in which a single actuator bearing was found to have failed by wear at 20 missions.

Oswald, Fred B.↗

Redundant disk arrays: Reliable, parallel secondary storage

During the past decade, advances in processor and memory technology have given rise to increases in computational performance that far outstrip increases in the performance of secondary storage technology. Coupled with emerging small-disk technology, disk arrays provide the cost, volume, and capacity of current disk subsystems, by leveraging parallelism, many times their performance. Unfortunately, arrays of small disks may have much higher failure rates than the single large disks they replace. Redundant arrays of inexpensive disks (RAID) use simple redundancy schemes to provide high data reliability. The data encoding, performance, and reliability of redundant disk arrays are investigated. Organizing redundant data into a disk array is treated as a coding problem. Among alternatives examined, codes as simple as parity are shown to effectively correct single, self-identifying disk failures.

Gibson, Garth Alan↗

MLEC-Sim: A Simulator for Evaluating Multi-Level Erasure Coding

We present MLEC-Sim, a sophisticated simulator for Multi-Level Erasure Coding (MLEC), developed in approximately 13 KLOC. The simulator is engineered to analyze the impact of various system configurations and erasure coding policies on system durability and network overhead. It supports a comprehensive range of parameters including disk capacity, disk I/O bandwidth, failure rates, network bandwidth, and system scale, accommodating various erasure coding approaches such as Single-Level Erasure Coding (SLEC), Multi-Level Erasure Coding (MLEC), and Local Reconstruction Codes (LRC). MLEC-Sim provides support for multiple chunk placement policies, including clustered parity and declustered parity, and encompasses a variety of repair methods like Repair-ALL, Repair-FCO, Repair-HYB, and Repair-MIN. It is capable of simulating disk failures through a variety of means, including distribution-based or trace-based mechanisms, and can handle complex multi-level (de)clustered placements and repair processes. A key feature of MLEC-Sim is its adoption of the splitting simulation method for evaluating system durabilities at extremely high levels, which are challenging to assess with traditional simulation approaches. This feature allows for a detailed evaluation of system resilience under a range of conditions, aiding in the selection of appropriate erasure coding solutions for enhancing system durability. MLEC-Sim contributes to the field of data storage and reliability by providing a tool for the detailed evaluation of the durability and efficiency of erasure coding configurations, intended for use by researchers and practitioners in the design and optimization of storage systems.

Wang, Meng↗