Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “probability of failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

General Multifidelity Surrogate Models: Framework and Active-Learning Strategies for Efficient Rare Event Simulation

Estimating the probability of failure for complex real-world systems using high-fidelity computational models is often prohibitively expensive, especially when the probability is small. Exploiting low-fidelity models can make this process more feasible, but merging information from multiple low-fidelity and high-fidelity models poses several challenges. Here, this paper presents a robust multi-fidelity surrogate modeling strategy in which the multi-fidelity surrogate is assembled using an active learning strategy using an on-the-fly model adequacy assessment set within a subset simulation framework for efficient reliability analysis. The multi-fidelity surrogate is assembled by first applying a Gaussian process correction to each low-fidelity model and assigning a model probability based on the model's local predictive accuracy and cost. Three strategies are proposed to fuse these individual surrogates into an overall surrogate model based on model averaging and deterministic/stochastic model selection. The strategies also dictate which model evaluations are necessary. No assumptions are made about the relationships between low-fidelity models, while the high-fidelity model is assumed to be the most accurate and most computationally expensive model. Through two analytical and two numerical case studies, including a case study evaluating the failure probability of Tristructural isotropic-coated (TRISO) nuclear fuels, the algorithm is shown to be highly accurate while drastically reducing the number of high-fidelity model calls (and hence computational cost).

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Uncertainty Quantification Enabled by Automatic Differentiation for Hydrodynamic Simulation of Shock‐to‐Detonation Transition in High Explosives

Quantifying the effects of uncertainty in a reactive burn model on the run-to-detonation time in high explosives (HEs) provides a robust methodology for assessing the probability of an HE failing the IHE qualification standard. Moreover, uncertainty quantification helps evaluate whether the model calibration accurately represents data outside the calibration set. This study uses a specialized hydrodynamic simulation code for modeling detonation to determine the run-to-detonation time of the HE PBX 9502 for various impact velocities. To quickly approximate uncertainties in the model, a surrogate was constructed using a Taylor series expansion centered at the mean of the input parameters. To obtain the sensitivities required for constructing the Taylor series, HYP-percomplex Automatic Differentiation (HYPAD) was implemented. HYPAD is a methodology for infusing existing codes with automatic differentiation capabilities by augmenting variables with one or more imaginary units to compute step-size independent partial derivatives. These derivatives are accurate to machine precision with respect to the implemented numerical algorithm, meaning their accuracy reflects that of the underlying method (e.g., integration or discretization schemes). Using reduced order modeling techniques, the mean and standard deviation of the run-to-detonation time of a shock within PBX 9502 were computed for a number of initial impact velocities. A weighted least squares regression was then performed to obtain a best fit curve and prediction interval for the computed statistics. Historical data points from explosively driven wedge tests were utilized to validate the prediction interval, ensuring its reliability in predicting future outcomes. With this prediction interval and a known safety constraint curve, the most probable point of failure and the probability of failure for the HE PBX 9502 were determined.

97 MATHEMATICS AND COMPUTING↗

A Comprehensive Reliability Methodology for Assessing Risk of Reusing Failed Hardware Without Corrective Actions with and Without Redundancy

This paper deals with the development of a reliability methodology to assess the consequences of using hardware, without failure analysis or corrective action, that has previously demonstrated that it did not perform per specification. The subject of this paper arose from the need to provide a detailed probabilistic analysis to calculate the change in probability of failures with respect to the base or non-failed hardware. The methodology used for the analysis is primarily based on principles of Monte Carlo simulation. The random variables in the analysis are: Maximum Time of Operation (MTO) and operation Time of each Unit (OTU) The failure of a unit is considered to happen if (OTU) is less than MTO for the Normal Operational Period (NOP) in which this unit is used. NOP as a whole uses a total of 4 units. Two cases are considered. in the first specialized scenario, the failure of any operation or system failure is considered to happen if any of the units used during the NOP fail. in the second specialized scenario, the failure of any operation or system failure is considered to happen only if any two of the units used during the MOP fail together. The probability of failure of the units and the system as a whole is determined for 3 kinds of systems - Perfect System, Imperfect System 1 and Imperfect System 2. in a Perfect System, the operation time of the failed unit is the same as that of the MTO. In an Imperfect System 1, the operation time of the failed unit is assumed as 1 percent of the MTO. In an Imperfect System 2, the operation time of the failed unit is assumed as zero. in addition, simulated operation time of failed units is assumed as 10 percent of the corresponding units before zero value. Monte Carlo simulation analysis is used for this study. Necessary software has been developed as part of this study to perform the reliability calculations. The results of the analysis showed that the predicted change in failure probability (P(sub F)) for the previously failed units is as high as 49 percent above the baseline (perfect system) for the worst case. The predicted change in system P(sub F) for the previously failed units is as high as 36% for single unit failure without any redundancy. For redundant systems, with dual unit failure, the predicted change in P(sub F) for the previously failed units is as high as 16%. These results will help management to make decisions regarding the consequences of using previously failed units without adequate failure analysis or corrective action.

Putcha, Chandra S.↗

A methodology for validating software reliability

A significant problem associated with fault tolerant computer system design is how to insure that there are no embedded software errors, so that an avionics computer system meets the required reliability level. To accomplish this, it is necessary to associate a 'probability of failure' with the operational flight program. It would be more correct to say that the probability of excitation of existing latent design errors within the program is required. In this sense, latent software errors are like latent hardware faults, and techniques that were previously used to measure the probability of failure of hardware due to fault latency can be used to measure the probability of failure of the software. A methodology was developed and applied to a flight control program that was known to operate in a well defined environment. The results indicated that the technique could be used to provide a final validation of the software to a specified reliability level and to evaluate the role of flight test in software validation.

Swern, Frederic L.↗

Reducing the Risk of Human Space Missions with INTEGRITY

The INTEGRITY Program will design and operate a test bed facility to help prepare for future beyond-LEO missions. The purpose of INTEGRITY is to enable future missions by developing, testing, and demonstrating advanced human space systems. INTEGRITY will also implement and validate advanced management techniques including risk analysis and mitigation. One important way INTEGRITY will help enable future missions is by reducing their risk. A risk analysis of human space missions is important in defining the steps that INTEGRITY should take to mitigate risk. This paper describes how a Probabilistic Risk Assessment (PRA) of human space missions will help support the planning and development of INTEGRITY to maximize its benefits to future missions. PRA is a systematic methodology to decompose the system into subsystems and components, to quantify the failure risk as a function of the design elements and their corresponding probability of failure. PRA provides a quantitative estimate of the probability of failure of the system, including an assessment and display of the degree of uncertainty surrounding the probability. PRA provides a basis for understanding the impacts of decisions that affect safety, reliability, performance, and cost. Risks with both high probability and high impact are identified as top priority. The PRA of human missions beyond Earth orbit will help indicate how the risk of future human space missions can be reduced by integrating and testing systems in INTEGRITY.

Jones, Harry W.↗

Meteoroid and Orbital Debris Threats to NASA's Docking Seals: Initial Assessment and Methodology

The Crew Exploration Vehicle (CEV) will be exposed to the Micrometeoroid Orbital Debris (MMOD) environment in Low Earth Orbit (LEO) during missions to the International Space Station (ISS) and to the micrometeoroid environment during lunar missions. The CEV will be equipped with a docking system which enables it to connect to ISS and the lunar module known as Altair; this docking system includes a hatch that opens so crew and supplies can pass between the spacecrafts. This docking system is known as the Low Impact Docking System (LIDS) and uses a silicone rubber seal to seal in cabin air. The rubber seal on LIDS presses against a metal flange on ISS (or Altair). All of these mating surfaces are exposed to the space environment prior to docking. The effects of atomic oxygen, ultraviolet and ionizing radiation, and MMOD have been estimated using ground based facilities. This work presents an initial methodology to predict meteoroid and orbital debris threats to candidate docking seals being considered for LIDS. The methodology integrates the results of ground based hypervelocity impacts on silicone rubber seals and aluminum sheets, risk assessments of the MMOD environment for a variety of mission scenarios, and candidate failure criteria. The experimental effort that addressed the effects of projectile incidence angle, speed, mass, and density, relations between projectile size and resulting crater size, and relations between crater size and the leak rate of candidate seals has culminated in a definition of the seal/flange failure criteria. The risk assessment performed with the BUMPER code used the failure criteria to determine the probability of failure of the seal/flange system and compared the risk to the allotted risk dictated by NASA's program requirements.

deGroh, Henry C., III↗

Meteoroid and Orbital Debris Threats to NASA's Docking Seals: Initial Assessment and Methodology

The Crew Exploration Vehicle (CEV) will be exposed to the Micrometeoroid Orbital Debris (MMOD) environment in Low Earth Orbit (LEO) during missions to the International Space Station (ISS) and to the micrometeoroid environment during lunar missions. The CEV will be equipped with a docking system which enables it to connect to ISS and the lunar module known as Altair; this docking system includes a hatch that opens so crew and supplies can pass between the spacecrafts. This docking system is known as the Low Impact Docking System (LIDS) and uses a silicone rubber seal to seal in cabin air. The rubber seal on LIDS presses against a metal flange on ISS (or Altair). All of these mating surfaces are exposed to the space environment prior to docking. The effects of atomic oxygen, ultraviolet and ionizing radiation, and MMOD have been estimated using ground based facilities. This work presents an initial methodology to predict meteoroid and orbital debris threats to candidate docking seals being considered for LIDS. The methodology integrates the results of ground based hypervelocity impacts on silicone rubber seals and aluminum sheets, risk assessments of the MMOD environment for a variety of mission scenarios, and candidate failure criteria. The experimental effort that addressed the effects of projectile incidence angle, speed, mass, and density, relations between projectile size and resulting crater size, and relations between crater size and the leak rate of candidate seals has culminated in a definition of the seal/flange failure criteria. The risk assessment performed with the BUMPER code used the failure criteria to determine the probability of failure of the seal/flange system and compared the risk to the allotted risk dictated by NASA s program requirements.

deGroh, Henry C., III↗

Acceleration Factors for Reliability Assessment of Polymer Tantalum Capacitors

Using polymer tantalum capacitors in Hi-Rel systems requires an assessment of the reliability characteristics of the parts. For this assessment, tantalum capacitors are typically subjected to reliability testing at temperatures and voltages exceeding their specified values, and the failure rate (FR) — or the probability of failure during use conditions — is calculated based on voltage and temperature acceleration factors. In this work, various types and lots of polymer tantalum capacitors have been tested at highly accelerated life test (HALT) conditions and the acceleration factors have been determined using different techniques. It has been shown that the behavior of capacitors under HALT conditions can be described based on the time-dependent dielectric breakdown (TDDB) model, which explains the presence of infant mortality (IM) and wear-out failures using the same failure mechanisms and allows for an assessment of the acceleration factors. The difference in acceleration factors obtained using exponential and power models is discussed. Analysis shows that with proper derating, screening, and qualification testing, the reliability of frameless Hi-Rel COTS polymer tantalum capacitors is adequate for space missions.

Alex Eidelman↗

Acceleration Factors for Reliability Assessment of Polymer Tantalum Capacitors

Using polymer tantalum capacitors in Hi-Rel systems requires an assessment of the reliability characteristics of the parts. For this assessment, tantalum capacitors are typically subjected to reliability testing at temperatures and voltages exceeding their specified values, and the failure rate (FR) — or the probability of failure during use conditions — is calculated based on voltage and temperature acceleration factors. In this work, various types and lots of polymer tantalum capacitors have been tested at highly accelerated life test (HALT) conditions and the acceleration factors have been determined using different techniques. It has been shown that the behavior of capacitors under HALT conditions can be described based on the time-dependent dielectric breakdown (TDDB) model, which explains the presence of infant mortality (IM) and wear-out failures using the same failure mechanisms and allows for an assessment of the acceleration factors. The difference in acceleration factors obtained using exponential and power models is discussed. Analysis shows that with proper derating, screening, and qualification testing, the reliability of frameless Hi-Rel COTS polymer tantalum capacitors is adequate for space missions.

Alex Eidelman↗

Discovery and Analysis of Rare High-Impact Failure Modes using Adversarial RL-Informed Sampling

Adaptive learning agents have tremendous potential to handle critical tasks currently performed by humans. Unfortunately, due to their complexity, it can be difficult to verify that these learning agents do not have critical failure modes. Standard verification and validation methods often do not apply directly to learning agents and Monte Carlo methods have difficulty covering even a small fraction of the state space, especially in multiagent systems or over long time horizons. To overcome this difficulty, we demonstrate an adaptive stress-testing method based on reinforcement learning of correlations that raise the probability of failure. This approach has three key properties: (1) it is able to find rare failure modes with far greater sample efficiency than Monte Carlo methods, (2) it can estimate the true probability of a failure mode despite the inherent bias in the learning method, and (3) it is capable of learning and resampling compact representations of multimodal failure spaces. These properties are important in practice as we need to find disparate failure modes while accounting for their actual relevance. This is a significant advantage over traditional adaptive stress testing methods that give abstract likelihoods of particular failure instances, but cannot estimate the probability of a broader failure mode. We test our algorithm on a simple problem from the aviation domain where an autonomous aircraft lands in gusty wind conditions. The results suggest that we can find failure modes with far fewer samples than the Monte Carlo approach and simultaneously estimate the probability of failure.

reinforcement learning↗

Discovery and Analysis of Rare High-Impact Failure Modes using Adversarial RL-Informed Sampling

Adaptive learning agents have tremendous potential to handle critical tasks currently performed by humans. Unfortunately, due to their complexity, it can be difficult to verify that these learning agents do not have critical failure modes. Standard verification and validation methods often do not apply directly to learning agents and Monte Carlo methods have difficulty covering even a small fraction of the state space, especially in multiagent systems or over long time horizons. To overcome this difficulty, we demonstrate an adaptive stress-testing method based on reinforcement learning of correlations that raise the probability of failure. This approach has three key properties: (1) it is able to find rare failure modes with far greater sample efficiency than Monte Carlo methods, (2) it can estimate the true probability of a failure mode despite the inherent bias in the learning method, and (3) it is capable of learning and resampling compact representations of multimodal failure spaces. These properties are important in practice as we need to find disparate failure modes while accounting for their actual relevance. This is a significant advantage over traditional adaptive stress testing methods that give abstract likelihoods of particular failure instances, but cannot estimate the probability of a broader failure mode. We test our algorithm on a simple problem from the aviation domain where an autonomous aircraft lands in gusty wind conditions. The results suggest that we can find failure modes with far fewer samples than the Monte Carlo approach and simultaneously estimate the probability of failure.

Validation↗

Discovery and Analysis of Rare High-Impact Failure Modes using Adversarial RL-Informed Sampling

Adaptive learning agents have tremendous potential to handle critical tasks currently performed by humans. Unfortunately, due to their complexity, it can be difficult to verify that these learning agents do not have critical failure modes. Standard verification and validation methods often do not apply directly to learning agents and Monte Carlo methods have difficulty covering even a small fraction of the state space, especially in multiagent systems or over long time horizons. To overcome this difficulty, we demonstrate an adaptive stress-testing method based on reinforcement learning of correlations that raise the probability of failure. This approach has three key properties: (1) it is able to find rare failure modes with far greater sample efficiency than Monte Carlo methods, (2) it can estimate the true probability of a failure mode despite the inherent bias in the learning method, and (3) it is capable of learning and resampling compact representations of multimodal failure spaces. These properties are important in practice as we need to find disparate failure modes while accounting for their actual relevance. This is a significant advantage over traditional adaptive stress testing methods that give abstract likelihoods of particular failure instances, but cannot estimate the probability of a broader failure mode. We test our algorithm on a simple problem from the aviation domain where an autonomous aircraft lands in gusty wind conditions. The results suggest that we can find failure modes with far fewer samples than the Monte Carlo approach and simultaneously estimate the probability of failure.

Validation↗

An expert system for reliability modeling

A method is presented for computing the probability of failure of a unit during the next usage or mission based on sound statistical techniques. With the use of logit regression embedded in Lotus 1-2-3 macros, this package computes the probability of unit failure based upon a historical data base. Using this computed probability of failure, the package then makes a replacement recommendation based upon the experts acceptable risk parameters.

Goss, Ernst↗

Optimized Vertex Method and Hybrid Reliability

A method of calculating the fuzzy response of a system is presented. This method, called the Optimized Vertex Method (OVM), is based upon the vertex method but requires considerably fewer function evaluations. The method is demonstrated by calculating the response membership function of strain-energy release rate for a bonded joint with a crack. The possibility of failure of the bonded joint was determined over a range of loads. After completing the possibilistic analysis, the possibilistic (fuzzy) membership functions were transformed to probability density functions and the probability of failure of the bonded joint was calculated. This approach is called a possibility-based hybrid reliability assessment. The possibility and probability of failure are presented and compared to a Monte Carlo Simulation (MCS) of the bonded joint.

Smith, Steven A.↗

Thermo-Mechanical Modeling and Evaluation of the Cracking Response of Additively Manufactured Monolithic SiC Lattice Structures Subjected to Laser Heating

Increasing operating temperatures of solar receivers is paramount to the efficiency of concentrated solar thermal (CST) and solar power (CSP) systems. Owing to its high temperature stability combined with excellent thermal and optical properties, SiC has been the material of choice for application in high-temperature solar receivers. The state-of-the-art SiC volumetric concentrating solar air receivers such as honeycomb design have been demonstrated in field tests to achieve exit air temperatures approaching 800oC. However, successful application of CST systems for decarbonization requires significant increase in temperature capability of SiC receiver technology. Supported by an award from the Solar Technology Office (SETO), US Department of Energy (DOE), GE Research in collaboration with Heliogen Holdings Inc, is engaged in the development of ultra-High Operating Temperature SiC-matrix Solar Thermal Air Receiver (HOTSSTAR) enabled by additive manufacturing. The program objective is to design and to demonstrate a techno-economically viable SiC air receiver technology to achieve exit air temperature of 1100oC for CST applications. This report summarizes the learnings of a computational study on additively manufactured SiC lattice structures subjected to laser heating. The results from this study aim to help guide the design of a SiC receiver through a better understanding of lattice structure geometric parameters and their implications on the cracking response of the structure. In the study, a heat flux was applied to the surface of a 2”-diameter cylindrical lattice structure to simulate a 4kW CO2 laser. The resulting temperature distribution was applied to a structural model to approximate the stress distribution within the lattice structure and a Weibull analysis was performed to gain insight into the probability of failure and to evaluate the cracking response of the structure. With this approach, the influence of lattice density on temperature, stress, and probability of failure was explored for two different lattice beam spacings. A discussion on modeling assumptions, a comparison with experimental results, and an evaluation of the lattice cracking response is provided.

14 SOLAR ENERGY↗

Reliability-based failure analysis of brittle materials

The reliability of brittle materials under a generalized state of stress is analyzed using the Batdorf model. The model is modified to include the reduction in shear due to the effect of the compressive stress on the microscopic crack faces. The combined effect of both surface and volume flaws is included. Due to the nature of fracture of brittle materials under compressive loading, the component is modeled as a series system in order to establish bounds on the probability of failure. A computer program was written to determine the probability of failure employing data from a finite element analysis. The analysis showed that for tensile loading a single crack will be the cause of total failure but under compressive loading a series of microscopic cracks must join together to form a dominant crack.

Powers, Lynn M.↗

Efficient Subset Simulation using Hamiltonian Neural Network enhanced Markov Chain Monte Carlo Methods

The Monte Carlo method delivers an unbiased estimate of the probability of failure. However, the variance of the estimate depends on the number of evaluated samples. This number must be very large for estimations of a low probability of failure. If the evaluation of each sample is computationally expensive, the crude Monte Carlo simulation strategy is impracticable. Therefore, subset simulations are used to reduce the required number of evaluations. Subset simulations require a Markov Chain Monte Carlo sampler, such as the random walk Metropolis-Hastings algorithm. The algorithm, however, struggles with sampling in low-probability regions, especially if they are narrow. As a consequence, advanced Markov Chain Monte Carlo simulations have been developed. In particular, the Hamiltonian Monte Carlo method explores the target distribution rapidly. Driven by the idea of Hamiltonian dynamics, this sampler provides a non-random walk through the target distribution. The incorporation of subset simulation and Hamiltonian Monte Carlo methods has shown promising results for reliability analysis. One downside of the Hamiltonian Monte Carlo method is that gradient evaluations are computationally expensive, especially when dealing with high-dimensional problems and evaluating long trajectories. We show that integrating Hamiltonian neural networks in Hamiltonian Monte Carlo simulations significantly speeds up the sampling task. Furthermore, the enhancement of adaptive trajectory length within the Hamiltonian Monte Carlo results in the efficient proposal of the following states. Based on this recent enhancement, we provide a fast sampling strategy for subset simulations using Hamiltonian neural networks to replace the evaluation of the gradient and significantly speed up the Hamiltonian Monte Carlo simulation.

97 MATHEMATICS AND COMPUTING↗

The application of structural reliability techniques to plume impingement loading of the Space Station Freedom Photovoltaic Array

A new aerospace application of structural reliability techniques is presented, where the applied forces depend on many probabilistic variables. This application is the plume impingement loading of the Space Station Freedom Photovoltaic Arrays. When the space shuttle berths with Space Station Freedom it must brake and maneuver towards the berthing point using its primary jets. The jet exhaust, or plume, may cause high loads on the photovoltaic arrays. The many parameters governing this problem are highly uncertain and random. An approach, using techniques from structural reliability, as opposed to the accepted deterministic methods, is presented which assesses the probability of failure of the array mast due to plume impingement loading. A Monte Carlo simulation of the berthing approach is used to determine the probability distribution of the loading. A probability distribution is also determined for the strength of the array. Structural reliability techniques are then used to assess the array mast design. These techniques are found to be superior to the standard deterministic dynamic transient analysis, for this class of problem. The results show that the probability of failure of the current array mast design, during its 15 year life, is minute.

Yunis, Isam S.↗