Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Failure Rate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Assuring reliability program effectiveness.

An attempt is made to provide simple identification and description of techniques that have proved to be most useful either in developing a new product or in improving reliability of an established product. The first reliability task is obtaining and organizing parts failure rate data. Other tasks are parts screening, tabulation of general failure rates, preventive maintenance, prediction of new product reliability, and statistical demonstration of achieved reliability. Five principal tasks for improving reliability involve the physics of failure research, derating of internal stresses, control of external stresses, functional redundancy, and failure effects control. A final task is the training and motivation of reliability specialist engineers.

Ball, L. W.↗

Developing Ultra Reliable Life Support for the Moon and Mars

Recycling life support systems can achieve ultra reliability by using spares to replace failed components. The added mass for spares is approximately equal to the original system mass, provided the original system reliability is not very low. Acceptable reliability can be achieved for the space shuttle and space station by preventive maintenance and by replacing failed units, However, this maintenance and repair depends on a logistics supply chain that provides the needed spares. The Mars mission must take all the needed spares at launch. The Mars mission also must achieve ultra reliability, a very low failure rate per hour, since it requires years rather than weeks and cannot be cut short if a failure occurs. Also, the Mars mission has a much higher mass launch cost per kilogram than shuttle or station. Achieving ultra reliable space life support with acceptable mass will require a well-planned and extensive development effort. Analysis must define the reliability requirement and allocate it to subsystems and components. Technologies, components, and materials must be designed and selected for high reliability. Extensive testing is needed to ascertain very low failure rates. Systems design should segregate the failure causes in the smallest, most easily replaceable parts. The systems must be designed, produced, integrated, and tested without impairing system reliability. Maintenance and failed unit replacement should not introduce any additional probability of failure. The overall system must be tested sufficiently to identify any design errors. A program to develop ultra reliable space life support systems with acceptable mass must start soon if it is to produce timely results for the moon and Mars.

Jones, Harry W.↗

Module Hipot and ground continuity test results

Hipot (high voltage potential) and module frame continuity tests of solar energy conversion modules intended for deployment into large arrays are discussed. The purpose of the tests is to reveal potentially hazardous voltage conditions in installed modules, and leakage currents that may result in loss of power or cause ground fault system problems, i.e., current leakage potential and leakage voltage distribution. The tests show a combined failure rate of 36% (69% when environmental testing is included). These failure rates are believed easily corrected by greater care in fabrication.

Griffith, J. S.↗

The effect of initial velocity on manually controlled remote docking of an orbital maneuvering vehicle (OMV) to a space station

Simulated docking maneuvers were performed to assess the effect of initial velocity on docking failure rate, mission duration, and total impulse (fuel consumption). The effect of the removal of the range and rate displays was also examined. Since duration and impulse decrease and increase respectively with increases in initial velocity, two parameters were created by subtracting a reference value from each. These values were termed 'reserve time' and 'radial impulse'. Naive subjects were capable of achieving a high success rate in performing simulated docking maneuvers without extensive experience, and failure rate did not significantly increase with increased velocity. The amount of time pilots reserved for final approach increased with starting velocity. Piloting of docking maneuvers was not significantly affected in any way by the removal of range and rate displays. Values for reserve time, and radial impulse were lowest for docking maneuvers begun at the lowest initial velocity.

Brody, Adam R.↗

Ultra Reliable Closed Loop Life Support for Long Space Missions

Spacecraft human life support systems can achieve ultra reliability by providing sufficient spares to replace all failed components. The additional mass of spares for ultra reliability is approximately equal to the original system mass, provided that the original system reliability is not too low. Acceptable reliability can be achieved for the Space Shuttle and Space Station by preventive maintenance and by replacing failed units. However, on-demand maintenance and repair requires a logistics supply chain in place to provide the needed spares. In contrast, a Mars or other long space mission must take along all the needed spares, since resupply is not possible. Long missions must achieve ultra reliability, a very low failure rate per hour, since they will take years rather than weeks and cannot be cut short if a failure occurs. Also, distant missions have a much higher mass launch cost per kilogram than near-Earth missions. Achieving ultra reliable spacecraft life support systems with acceptable mass will require a well-planned and extensive development effort. Analysis must determine the reliability requirement and allocate it to subsystems and components. Ultra reliability requires reducing the intrinsic failure causes, providing spares to replace failed components and having "graceful" failure modes. Technologies, components, and materials must be selected and designed for high reliability. Long duration testing is needed to confirm very low failure rates. Systems design should segregate the failure causes in the smallest, most easily replaceable parts. The system must be designed, developed, integrated, and tested with system reliability in mind. Maintenance and reparability of failed units must not add to the probability of failure. The overall system must be tested sufficiently to identify any design errors. A program to develop ultra reliable space life support systems with acceptable mass should start soon since it must be a long term effort.

Jones, Harry W.↗

Reliability measurement during software development

Measurement of software reliability was carried out during the development of data base software for a multi-sensor tracking system. Every run made during this project was scored as success or failure, and supporting data were collected on forms for further analysis. The failure ratio (number of failures per calendar interval divided by total number of runs) and failure rate (number of failures divided by CPU time for the interval) were found to be consistent measures, on a month-to-month basis as well as from module to module, and therefore considered valid indicators of reliability in this environment. Trend lines could be established from these measurements that provide good visualization of the progress on the job as a whole as well as on individual modules. Over one-half of the observed failures were due to factors associated with the specific run submission rather than with the code proper.

Hecht, H.↗

JANTX1N2970B zener diode

Report evaluates effects of power and temperature overstress on General Semi-conductor and Siemens devices. Excessive failure rates limited testing. Failure modes are described.

Source record↗

Data Applicability of Heritage and New Hardware For Launch Vehicle Reliability Models

Bayesian reliability requires the development of a prior distribution to represent degree of belief about the value of a parameter (such as a component's failure rate) before system specific data become available from testing or operations. Generic failure data are often provided in reliability databases as point estimates (mean or median). A component's failure rate is considered a random variable where all possible values are represented by a probability distribution. The applicability of the generic data source is a significant source of uncertainty that affects the spread of the distribution. This presentation discusses heuristic guidelines for quantifying uncertainty due to generic data applicability when developing prior distributions mainly from reliability predictions.

Al Hassan, Mohammad↗

Enhanced Component Performance Study: Air-Operated Valves 1998–2020

This report presents an enhanced performance evaluation of air-operated valves (AOVs) at U.S. commercial nuclear power plants. The data used in this study are based on the operating experience failure reports from calendar year 1998 through 2020 as reported in the Institute of Nuclear Power Operations (INPO) Industry Reporting and Information System (IRIS). The AOV failure modes considered are failure-to-open/close (FTOC), failure to operate or control (FTOP), and spurious operation (SO). The component reliability estimates and the reliability data are trended for the most recent 10-year period while yearly estimates for reliability are provided for the entire study period. The following trends were identified for the most recent 10-year period: o Extremely statistically significant increasing trend for the frequency of FTOC demands (demands per reactor year) for low-demand (= 20 demands per year) AOVs o Extremely statistically significant increasing trend for the frequency of FTOC demands for high-demand (> 20 demands per year) AOVs o Highly statistically significant decreasing trend for the failure rate of FTOP for low-demand AOVs o Highly statistically significant decreasing trend for the frequency of FTOP events (failures per reactor year) for low-demand AOVs o Statistically significant decreasing trend for the failure rate of SO for low-demand AOVs o Statistically significant decreasing trend for the frequency of SO events (failures per reactor year) for low-demand AOVs.

99 GENERAL AND MISCELLANEOUS↗

Orbital performance of communication satellite microwave power amplifiers (MPAs)

This paper presents background data on the performance of microwave power amplifiers (MPAs) used as transmitters in currently operating commercial communication satellites. Specifically aspects of two competing MPA types are discussed. These are well known TWTA (travelling wave tube amplifier) and the SSPA (solid state power amplifier). Extensive in-orbit data has been collected from over 2000 MPAs in 1991 and 1993. The study in 1991 invovlved 75 S/C (spacecraft) covering 463 S/C years. The 1993 'second-look' study encompassed a slightly different population of 72 S/C with 497 S/C years of operation. A surprising result of both studies was that SSPAs, although quite reliable, did not achieve the reliability of TWTAs were one-third more reliable in the 1993 study. This was at C-band with comparable power amplifiers, e.g. 6-16W of RF output power and similar gains. Data at K(sub u)-band is for TWTAs only since there are no SSPAs in the current S/C inventory. The other complementary result was that the projected failure rates used as S/C payload design guidelines were, on average, somewhat higher for TWTAs than the actual failure rates uncovered by this study. SSPA rates were as projected.

Strauss, R.↗

Rare events and Griffiths phases in topological quantum error correction

The performance of quantum error correcting (QEC) codes is often studied under the assumption of spatiotemporally uniform error rates. On the other hand, experimental implementations almost always produce heterogeneous error rates, in either space or time, as a result of effects such as imperfect fabrication and/or cosmic rays. It is therefore important to understand if and how their presence can affect the performance of QEC in qualitative ways. Here, in this work, we study the effects of nonuniform error rates in the representative examples of the 1D repetition code and the 2D toric code, focusing on when they have extended spatiotemporal correlations; these may arise, for instance, from rare events (such as cosmic rays) that temporarily elevate error rates over the entire code patch. These effects can be described in the corresponding statistical mechanics models for decoding, where long-range correlations in the error rates lead to extended rare regions of weaker coupling. For the 1D repetition code where the rare regions are linear, we find two distinct decodable phases: a conventional ordered phase in which logical failure rates decay exponentially with the code distance, and a rare-region dominated Griffiths phase in which failure rates are parametrically larger and decay as a stretched exponential. In particular, the latter phase is present when the error rates in the rare regions are above the bulk threshold. For the 2D toric code where the rare regions are planar, we find no decodable Griffiths phase: rare events which boost error rates above the bulk threshold lead to an asymptotic loss of threshold and failure to decode. Unpacking the failure mechanism implies that techniques for suppressing extended sequences of repeated rare events (which, without intervention, will be statistically present with high probability) will be crucial for QEC with the toric code.

classical statistical mechanics↗

Data Applicability of Heritage and New Hardware for Launch Vehicle System Reliability Models

Many launch vehicle systems are designed and developed using heritage and new hardware. In most cases, the heritage hardware undergoes modifications to fit new functional system requirements, impacting the failure rates and, ultimately, the reliability data. New hardware, which lacks historical data, is often compared to like systems when estimating failure rates. Some qualification of applicability for the data source to the current system should be made. Accurately characterizing the reliability data applicability and quality under these circumstances is crucial to developing model estimations that support confident decisions on design changes and trade studies. This presentation will demonstrate a data-source classification method that ranks reliability data according to applicability and quality criteria to a new launch vehicle. This method accounts for similarities/dissimilarities in source and applicability, as well as operating environments like vibrations, acoustic regime, and shock. This classification approach will be followed by uncertainty-importance routines to assess the need for additional data to reduce uncertainty.

Al Hassan Mohammad↗

Evaluation of Power Transmission Lines Hardening Scenarios Using a Machine Learning Approach

The power transmission infrastructure is vulnerable to extreme weather events, particularly hurricanes and tropical storms. A recent example is the damage caused by Hurricane Maria (H-Maria) in the archipelago of Puerto Rico in September 2017, where major failures in the transmission infrastructure led to a total blackout. Numerous studies have been conducted to examine strategies to strengthen the transmission system, including burying the power lines underground or increasing the frequency of tree trimming. However, few studies focus on the direct hardening of the transmission towers to accomplish an increase in resiliency. This machine learning-based study fills this need by analyzing three direct hardening scenarios and determining the effectiveness of these changes in the context of H-Maria. A methodology for estimating transmission tower damage is presented here in this study as well as an analysis of impact of replacing structures with a high failure rate with more resilient ones. We found the steel self-support-pole to be the best replacement option for the towers with high failure rate. Furthermore, the third hardening scenario, where all wooden poles were replaced, exhibited a maximum reduction in damaged towers in a single line of 66% while lowering the mean number of damaged towers per line by 10%.

54 ENVIRONMENTAL SCIENCES↗

Software reliability: Repetitive run experimentation and modeling

A software experiment conducted with repetitive run sampling is reported. Independently generated input data was used to verify that interfailure times are very nearly exponentially distributed and to obtain good estimates of the failure rates of individual errors and demonstrate how widely they vary. This fact invalidates many of the popular software reliability models now in use. The log failure rate of interfailure time was nearly linear as a function of the number of errors corrected. A new model of software reliability is proposed that incorporates these observations.

Nagel, P. M.↗

Single event induced transients in I/O devices - A characterization

The results of single-event upset (SEU) testing performed to evaluate the parametric transients, i.e., amplitude and duration, in several I/O devices, and the impact of these transients are discussed. The failure rate of these devices is dependent on the susceptibility of interconnected devices to the resulting transient change in the output of the I/O device. This failure rate, which is a function of the susceptibility of the interconnected device as well as the SEU response of the I/O device itself, may be significantly different from an upset rate calculated without taking these factors into account. The impact at the system level is discussed by way of an example.

Newberry, D. M.↗

Hierarchical memories: Simulating quantum LDPC codes with local gates

Constant-rate low-density parity-check (LDPC) codes are promising candidates for constructing efficient fault-tolerant quantum memories. However, if physical gates are subject to geometric-locality constraints, it becomes challenging to realize these codes. In this paper, we construct a new family of [[N,K,D]] codes, referred to as hierarchical codes, that encode a number of logical qubits K=Ω(N/log(N) 2 ). The N th element of this code family is obtained by concatenating a constant-rate quantum LDPC code with a surface code; nearest-neighbor gates in two dimensions are sufficient to implement the corresponding syndrome-extraction circuit and achieve a threshold. Below threshold the logical failure rate vanishes superpolynomially as a function of the distance D(N). We present a bilayer architecture for implementing the syndrome-extraction circuit, and estimate the logical failure rate for this architecture. Under conservative assumptions, we find that the hierarchical code outperforms the basic encoding where all logical qubits are encoded in the surface code.

Pattison, Christopher A. [California Institute of ↗

Sensitivity of a critical tracking task to alcohol impairment

A first order critical tracking task is evaluated for its potential to discriminate between sober and intoxicated performances. Mean differences between predrink and postdrink performances as a function of BAC are analyzed. Quantification of the results shows that intoxicated failure rates of 50% for blood alcohol concentrations (BACs) at or above 0.1%, and 75% for BACs at or above 0.14%, can be attained with no sober failure rates. A high initial rate of learning is observed, perhaps due to the very nature of the task whereby the operator is always pushed to his limit, and the scores approach a stable asymptote after approximately 50 trials. Finally, the implementation of the task as an ignition interlock system in the automobile environment is discussed. It is pointed out that lower critical performance limits are anticipated for the mechanized automotive units because of the introduction of larger hardware and neuromuscular lags. Whether such degradation in performance would reduce the effectiveness of the device or not will be determined in a continuing program involving a broader based sample of the driving population and performance correlations with both BACs and driving proficiency.

Tennant, J. A.↗

Design for Reliability (DfR) in Space Life Support

The engineering process of Design for Reliability (DfR) is well established in the automotive and aerospace industries. DfR should be useful in the future development of space life support systems. DfR is a sequence of tasks that develop system requirements and plan reliability analysis and testing. First and fundamentally, the reliability requirement is defined. Next the system reliability model is developed, often using a reliability block diagram. The overall system reliability requirement is allocated to the subsystems and an estimate of the attainable reliability is made. This expected reliability can be improved by simplifying the design by removing components or by replacing less reliable components. Improving reliability can require difficult compromises, such as reducing performance requirements, increasing budget, or extending testing. The actual system reliability can be determined only by testing, which should continue long enough to provide the required confidence in the measured value. New systems often have unexpected design errors that cause failures in early testing. The usual reliability improvement process of testing, finding the failure modes, and redesigning to remove them reduces the failure rate and is referred to as “reliability growth.” After redesign has been completed, the system should be further tested to determine the actual achieved reliability more accurately. If the final system failure rate is too high, redundant systems can be used to improve overall operational reliability. Adding redundancy simply to increase the one- or two-fault tolerance metric may sometimes reduce reliability. Reliability can be improved in three ways: redesigning the system to include more reliable subsystems and components, reliability growth testing and failure mode removal, and by using parallel redundant systems. DfR should combine these approaches to achieve the required reliability while managing performance, cost, and schedule.

Reliability↗