COMPARISON OF VALVE FAILURE RATES ESTIMATED FROM FIELD FAILURE DATA TO THOSE PREDICTED BY CALIBRATED FMEDA
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Advancements in nuclear system designs with automated control features provide many benefits, but can lead to complex coupled systems and dynamic failure scenarios. This is especially true for microreactor designs where components are not expected to be replaced during the reactor’s lifetime. Hence, the life of the system, in addition to the safety, needs to be evaluated. Modeling these sequences of time-dependent events requires addressing cyclical processes and changing failure rates in ways that represent the actual system dynamics in contrast to a single sampling for a component’s time to failure. This research presents two distinct analytical methods for several failure distributions that evaluate a final time to failure used for different scenarios where the time to failure must be sampled multiple times. The first method is used when evaluating a component whose failure rate increases due to an outside event after the initial sampling but before the initially sampled time to failure. The second method is used when evaluating multiple identical components or a component that has been replaced with a new identical version before the second sampling. The two methods were implemented in a few representative case studies developed in the dynamic probabilistic risk assessment tool Event Modeling Risk Assessment using Linked Diagrams. Overall, this paper provides guidelines on how these approaches give a more realistic and accurate dynamic probabilistic risk assessment of complex systems.
The Hydrogen Component Reliability Database (HyCReD) is a collaborative project between the National Renewable Energy Laboratory, the University of Maryland, and hydrogen stakeholders to improve safety and reliability for hydrogen facilities by implementing component reliability data taxonomies that support hydrogen infrastructure failure rate analysis. The project aims to quantify failure rates of hydrogen components through high-quality data collection and analysis on root causes and maintenance needed. HyCReD provides a common database for cataloging hydrogen component failures which exists for reliability research in many other mature industries [2]. The database fills a gap for the hydrogen community by providing a scientifically rigorous approach to quantitative risk assessment (QRA), prognostic health management (PHM), and reliability-centered maintenance (RCM) analysis. High level results will be aggregated and anonymized to protect company sensitive information; detailed results will be used to help address issues of hydrogen components. These advanced analytics will support accelerated deployment of hydrogen infrastructure by enabling better: design and safety of projects (safety codes and standards development), infrastructure reliability and cost (component failure rates, maintenance protocols), and component R&D needs (robust supply chain). A key to a successful HyCReD implementation is facilitating the ease of reporting and data quality in the database that can be used for analysis. Maintenance data was a previously identified gap in initial efforts to populate and validate the database taxonomies [3]. Collection of maintenance data will be instrumental in identifying failure modes and rates, identifying incipient component failures or reduced performance, cataloging best practices for maintenance routines and methods for prognostic health management, and quantifying the risk and effect of different failure modes. Several key priorities are identified for streamlined data collection to achieve quality and detailed failure data: Applicability, Ease of Use, Accessibility, and Information Security. The HyCReD team has now begun deployment of the database to several companies and groups that have signed non-disclosure agreements to facilitate the data collection of failures in industry hydrogen refueling station infrastructure. This paper will provide an update into the process of HyCReD deployment including the development of a coding guide for facility personnel to reference and ensure data quality and consistency from one station to another as well as implementation of contextually dependent data fields of system taxonomy and formatted entries to provide ease of use. The goal is to communicate the lessons learned from the roll-out to technicians and engineers in the field, and the addition of need for high level of security to protect all stakeholders.
Weather events cause most grid failures. Often, we even get notifications on our phones to take cover or be prepared for an imminent event. If electric grid utilities had a similar warning that also included probable scenarios and the equipment involved, they could prepare and minimize the effects. Recent research at Idaho National Laboratory into electric grid risk analysis methods resulted in a tool that allows for the development of the most likely scenarios given failure probabilities of grid components. INL has a project with the U.S. Department of Energy’s Cybersecurity, Energy Security, and Emergency Response (CESER) program to develop a Grid Alert application that receives messages from the existing emergency alert system, filters and determines components possibly affected by the emergency event, calculates probable scenarios uses MASTERRI and then notifies the utility if there is significant risk. Historical failure data of elements that comprise the U.S. electric grid have been compiled by utilities and organizations such as the international regulatory body North American Electric Reliability Corporation (NERC). Nominal failure rates are obtained from this data. To make this tool possible, estimated failure rates are needed for different component types given the alert type, severity, and location. Historic weather-related grid element failures are correlated with historic weather events from Integrated Public Alert & Warning System (IPAWS). These correlated events and failures are used along with Bayesian updates from the historical norms to provide a modified failure rate for grid elements in the alert areas and calculate probable scenarios. This discusses the Grid Alert project plan but focuses on the data gathered and process used in determining failure rates for possible grid failure scenarios.
Two of the challenges of current plant reliability approaches are the ability to integrate plant health data, and to support decision making. Condition based data and diagnostic/prognostic information are in fact not considered into plant reliability models to inform system engineers on the most critical components. Currently, the propagation of quantitative health data from the component to the system level is a challenge given the diverse nature/structure of the data. On the other hand, plant reliability methods (which are typically based on fault-trees or reliability block diagrams) can effectively propagate data from the component to the system level, but values of failure rates or failure probabilities are an approximated integral representation of the past industry-wide operational experience, and it neglects the present component health status (e.g., diagnostic and condition-based data) and health projection (when available from prognostic data). Our first claim is that system reliability models should propagate health information from the component to the system/plant level in order to provide a quantitative snapshot of system/plant health and identify the most critical components. Our second claim is that component health should be informed solely by that specific component current and historical performance data and should not be an approximated integral representation of the past industry-wide operational experience. This paper is directly supporting these two claims by proposing a different approach to perform reliability modeling which relies on available component diagnostic, prognostic and condition-based data to measure component health, and it propagates this information through fault tree models. The propagation of health data from the component to the system level is performed not in terms of probability, but in terms of margins where margin is defined as the “distance” between the present actual status and an undesired event (e.g., failure or unacceptable performance). Through a cause-effect lens, while classical reliability models target the effect associated to a component performance, a margin-based approach focuses on the cause of an undesired component performance (i.e., component health). Hence, thinking of reliability in terms of margins implies decision making based on causal reasoning. We will show how fault tree models can be solved using a margin language and how this process can effectively assist system engineers to identify the most critical components.
Weather events cause most power outages. Often, we even get notifications on our phones to take cover or be prepared for an imminent event. If electric grid utilities had a similar warning that also included probable scenarios and the equipment involved, they could prepare and minimize the effects. Idaho National Laboratory had a project with the U.S. Department of Energy’s Cybersecurity, Energy Security, and Emergency Response program to develop a grid alert application that receives messages from the existing emergency alert system, filters and determines components possibly affected by the emergency event, calculates probable scenarios using MASTERRI (Modeling And Simulation for Targeted Reliability and Resilience Improvement). For high-risk events, the application can then send alert links to subscribed electric distribution utility operations staff to allow them to see and evaluate the scenarios and the impact in a web based interactive map tool. This proof of concept application used data from utilities and organizations, such as the international regulatory body North American Electric Reliability Corporation, which have complied historical failure data of elements that comprise the U.S. electric grid. Nominal failure rates are obtained from this data. To make this tool possible, estimated failure rates were calculated for different component types given the alert type, severity, and location. Historic weather-related grid element failures were correlated with historic weather events from the Integrated Public Alert & Warning System. These correlated events and failures are used along with Bayesian updates from the historical norms to provide a modified failure rate for grid elements in the alert areas and calculate probable scenarios. Working with an industry collaborator, actual grid models and data were used for demonstration cases. This report outlines the work performed for this project.
This report presents an enhanced performance evaluation of turbine driven pumps (TDPs) at U.S. commercial nuclear power plants. The data used in this study are based on the operating experience failure reports from calendar year 1998 through 2022 as reported in the Institute of Nuclear Power Operations (INPO) Industry Reporting and Information System (IRIS). The TDP failure modes considered for standby systems are fail to start (FTS), fail to run (FTR) for one hour of operation (FTR=1H), FTR after one hour of operation (FTR>1H), and for normally running systems FTS and FTR. An eight hour unreliability estimate is also calculated and trended. The component reliability estimates and the reliability data are trended for the most recent 10 year period while yearly estimates for reliability are provided for the entire study period. No increasing trends were identified for TDPs for the most recent 10 year period: The following decreasing trends were identified for TDPs for the most recent 10 year period: • Standby TDP FTR>1H failure rate • Normally running TDP FTR failure rate • Standby TDP unavailability • Standby MDP total unreliability (8-hour mission) • Standby TDP frequency of start demands (demands per reactor year) • Standby TDP frequency of FTR=1H hours (hours per reactor year) • Standby TDP frequency of FTR>1H events (failures per reactor year) • Normally running TDP frequency of start demands • Normally running TDP frequency of run hours • Normally running TDP frequency of FTR events.
A team from the National Renewable Energy Laboratory (NREL) visited Guam in August 2023 to assess failure modes of solar photovoltaic (PV) systems after Typhoon Mawar and to provide recommendations to increase the resilience of PV systems on Guam. The team visited 30 systems: commercial and utility scale, and rooftop and ground-mounted. The team observed systems with no apparent damage, as well as systems that were completely lost. Systems fared very well overall. The average failure rate of rooftop systems was 18%, with a median failure rate of 2%, meaning the few systems that suffered total loss pulled up the average. Only eight 8 of the 25 rooftop systems suffered more than 5% damage. All ground-mounted systems suffered less than 0.5% damage, aside from a carport that lost 16% of its modules. PV systems at Andersen Air Force Base suffered 5% damage on average, with a median system failure of 0.6%. In almost all cases, failures were the result of: (1) Inadequate clamping of the module frame to the mount, (2) Module mounting clamps rotating out of underlying support rail (i.e., T-bolt that rotates free at less than 60 degrees of rotation), (3) An object hitting the panel resulting in a fracture, and in some cases leading to a cascading failure of several more panels, and (4) Excessive tilt angle (in Guam, greater than 5 degrees can be a risk due to wind speed, and power production trade-offs are insignificant).
This report presents an enhanced performance evaluation of air-operated valves (AOVs) at U.S. commercial nuclear power plants. The data used in this study are based on the operating experience failure reports from calendar year 1998 through 2020 as reported in the Institute of Nuclear Power Operations (INPO) Industry Reporting and Information System (IRIS). The AOV failure modes considered are failure-to-open/close (FTOC), failure to operate or control (FTOP), and spurious operation (SO). The component reliability estimates and the reliability data are trended for the most recent 10-year period while yearly estimates for reliability are provided for the entire study period. The following trends were identified for the most recent 10-year period: o Extremely statistically significant increasing trend for the frequency of FTOC demands (demands per reactor year) for low-demand (= 20 demands per year) AOVs o Extremely statistically significant increasing trend for the frequency of FTOC demands for high-demand (> 20 demands per year) AOVs o Highly statistically significant decreasing trend for the failure rate of FTOP for low-demand AOVs o Highly statistically significant decreasing trend for the frequency of FTOP events (failures per reactor year) for low-demand AOVs o Statistically significant decreasing trend for the failure rate of SO for low-demand AOVs o Statistically significant decreasing trend for the frequency of SO events (failures per reactor year) for low-demand AOVs.
The performance of quantum error correcting (QEC) codes is often studied under the assumption of spatiotemporally uniform error rates. On the other hand, experimental implementations almost always produce heterogeneous error rates, in either space or time, as a result of effects such as imperfect fabrication and/or cosmic rays. It is therefore important to understand if and how their presence can affect the performance of QEC in qualitative ways. Here, in this work, we study the effects of nonuniform error rates in the representative examples of the 1D repetition code and the 2D toric code, focusing on when they have extended spatiotemporal correlations; these may arise, for instance, from rare events (such as cosmic rays) that temporarily elevate error rates over the entire code patch. These effects can be described in the corresponding statistical mechanics models for decoding, where long-range correlations in the error rates lead to extended rare regions of weaker coupling. For the 1D repetition code where the rare regions are linear, we find two distinct decodable phases: a conventional ordered phase in which logical failure rates decay exponentially with the code distance, and a rare-region dominated Griffiths phase in which failure rates are parametrically larger and decay as a stretched exponential. In particular, the latter phase is present when the error rates in the rare regions are above the bulk threshold. For the 2D toric code where the rare regions are planar, we find no decodable Griffiths phase: rare events which boost error rates above the bulk threshold lead to an asymptotic loss of threshold and failure to decode. Unpacking the failure mechanism implies that techniques for suppressing extended sequences of repeated rare events (which, without intervention, will be statistically present with high probability) will be crucial for QEC with the toric code.
The power transmission infrastructure is vulnerable to extreme weather events, particularly hurricanes and tropical storms. A recent example is the damage caused by Hurricane Maria (H-Maria) in the archipelago of Puerto Rico in September 2017, where major failures in the transmission infrastructure led to a total blackout. Numerous studies have been conducted to examine strategies to strengthen the transmission system, including burying the power lines underground or increasing the frequency of tree trimming. However, few studies focus on the direct hardening of the transmission towers to accomplish an increase in resiliency. This machine learning-based study fills this need by analyzing three direct hardening scenarios and determining the effectiveness of these changes in the context of H-Maria. A methodology for estimating transmission tower damage is presented here in this study as well as an analysis of impact of replacing structures with a high failure rate with more resilient ones. We found the steel self-support-pole to be the best replacement option for the towers with high failure rate. Furthermore, the third hardening scenario, where all wooden poles were replaced, exhibited a maximum reduction in damaged towers in a single line of 66% while lowering the mean number of damaged towers per line by 10%.
Constant-rate low-density parity-check (LDPC) codes are promising candidates for constructing efficient fault-tolerant quantum memories. However, if physical gates are subject to geometric-locality constraints, it becomes challenging to realize these codes. In this paper, we construct a new family of [[N,K,D]] codes, referred to as hierarchical codes, that encode a number of logical qubits K=Ω(N/log(N) 2 ). The N th element of this code family is obtained by concatenating a constant-rate quantum LDPC code with a surface code; nearest-neighbor gates in two dimensions are sufficient to implement the corresponding syndrome-extraction circuit and achieve a threshold. Below threshold the logical failure rate vanishes superpolynomially as a function of the distance D(N). We present a bilayer architecture for implementing the syndrome-extraction circuit, and estimate the logical failure rate for this architecture. Under conservative assumptions, we find that the hierarchical code outperforms the basic encoding where all logical qubits are encoded in the surface code.
In this paper, we introduce Surf-Deformer, a code deformation framework that seamlessly integrates adaptive defect mitigation functionality into the current surface code workflow. It crafts several basic deformation instructions based on fundamental gauge transformations, which can be combined to explore a larger design space than previous methods. This enables more optimized deformation processes tailored to specific defect situations, restoring the QEC capability of deformed codes more efficiently with minimal qubit resources. Additionally, we design an adaptive code layout that accommodates our defect mitigation strategy while ensuring efficient execution of logical operations. Our evaluation shows that Surf-Deformer outperforms previous methods by significantly reducing the end-to-end failure rate of various quantum programs by 35× to 70×, while requiring only about 50% of the qubit resources compared to the previous method to achieve the same level of failure rate. Ablation studies show that Surf-Deformer surpasses previous defect removal methods in preserving QEC capability and facilitates surface code communication by achieving nearly optimal throughput.
SRNL was funded in Mid-Year FY20 by NNSA NA-231 to continue evaluation of alternate valves for use in tritium service to support domestic Mo-99 production. The focus of the effort was to identify valve cycle life as a function of actuator size and stem tip material. Using the minimum size actuator to reliability open and close valves can reduce glovebox size and thus support domestic companies to “come to market” faster in supplying Mo-99 to the US market. This report serves as a continuation to the FY19 report and summarizes the task activities completed in FY20 after authorization to start work was obtained on May 5 th , 2020. Copper stem tips were tested on the Swagelok 1C and 5C actuated metal bellows valves. With ambitions to cycle each set 150,000 times, both 1C and 5C valves were met with high failure rates. The smaller 1C actuated valves required additional closing pressure to form a seal with the Cu stem tips installed however, excessive stem tip deformation may be the root cause of the majority of the valves failing before 500 cycles. The larger 5C actuated valves were cycled 150,000 times but still resulted in 80% failure rate, suspected of metal fatigue in the bellows due to high cycling frequencies.
Dispensers are the top cause of maintenance events and down-time at hydrogen fueling stations. In an effort to help characterize and enable improvements in dispenser reliability, an extensive accelerated lifetime testing set-up was designed and built at NREL involving components typically part of dispensing operations at fueling stations. Device Under Test (DUTs) included different components such as normally open valves, normally closed valves, fueling nozzles, breakaways devices and filters. Conditions of testing included pressures, and flow rates similar to light duty fuel cell electric vehicles fueling at -40°C, and -20°C for thousands of cycles in hydrogen. Tested components (failed and non-failed) were disassembled at SNL and polymeric O-rings were carefully retrieved and cataloged for chemical and physical characterization. Data collected was compared to similar O-rings from unexposed or non-tested components for hydrogen effects, and failure modes. Degradation analyses, based on select polymer chemistries common across all component types, their location within components, visual assessment of damage coupled with strong hydrogen effects from chemical characterization, was completed and presented to NREL and DOE. Overall, the failure rate amongst the components was not as high as expected for the test conditions. Among the component types tested, breakaways were the most susceptible to damage under these test conditions, with fueling nozzles a close second. The proper combination of selection of the right polymer and optimum component design was found to make a strong difference in component reliability under severe dispenser operating conditions. Physical degradation of polymers, rather than chemical changes due to low temperature hydrogen exposure, is more prevalent as failure mode for these test conditions. The nature and the extent of the degradation was much less at -20°C as compared to -40°C. The damage and failure rates were higher at lower temperatures than at higher test temperatures. As expected, increasing the number of cycles at the lowest test temperature (-40°C) increased damage. This indicates that cycling at the low temperature of -40°C required by SAE J2601 can reduce component life in fuel dispensing operations
The reliability, cost and performance of electrical connectors are a concern in all types of electrical systems, and demands on connectors used on photovoltaic (PV) systems include that connectors maintain electrical conductivity and physical strength, endure ultraviolet sunlight and high ambient temperature, and resist moisture and chemical intrusion over a very long (>25 year) performance period. Connector failures increase operation and maintenance (O&M) costs and reduce plant production, but connector failure can also cause safety and liability problems, which are of greater concern. This work results from a three-year collaboration between Sandia National Laboratories (SNL), the Electric Power Research Institute (EPRI), and the National Renewable Energy Laboratory (NREL) and funded by the U.S. Department of Energy (DOE) Solar Energy Technology Office (SETO) under Agreements #39035 and #38531 "Connector Reliability Across the US Solar Sector." a multi-pronged investigation of PV connector health across the US (see https://energy.sandia.gov/pvconnectors/). This report presents derivation of a Techno-Economic Analysis (TEA) that models failure modes and frequencies (how often failure occurs), estimates O&M costs and lost production associated with connector failures, and then calculates the effect that PV module connectors can have on Levelized Cost of Energy (LCOE). The model is informed with initial data from quantitative assessment of failure rates, root causes and mechanisms, in-situ diagnostics and data collection, lab-based forensics, and interviews with PV connector manufacturers and plant operators. SNL conducted site inspections at multiple utility-scale sites in different climates and subjected field samples of new, used, and degraded connectors to visual and electrical characterization. EPRI conducted metallurgical analysis of the pin and sleeve conductors to study failure-induced morphological and compositional changes. There is in general a shortage of statistically valid data, but data from PVROM database maintained by SNL was sufficient to ascertain failure rates and lost production as well as provide qualitative insight in its curated maintenance records. This report details the structure of the mathematical model but the sources of data to inform the model will continue to evolve. Analysis of a 100 MW PV plant is provided as an example of the use of the model, with results indicating that connectors are responsible for Annualized O&M Costs of $\$$71,933/year; Annualized Unit O&M Costs of $\$$0.72/kW/year; that a Reserve Account of $\$$187,220 should be available to fund repairs related to connectors; that connectors add $\$$1,494,004 to the Net Present Value of the O&M Costs (project life); and that O&M related to connectors adds about $\$$0.00088/kWh to the Levelized Cost of Energy. The impact of this model is to provide a tool to make the US solar sector more robust by quantifying and monetizing the reliability risks to utility-scale PV systems posed by poorly installed, mismatched and/or poorly designed and manufactured connectors. The TEA provides a model incorporating failure statistics, O&M cost data, and lost production into a single figure of merit, informing decisions and enabling practitioners to optimize cost and performance trade-offs. Stakeholders include connector manufacturers, system designers and equipment specifiers, standards bodies, installers and O&M providers, investors and insurance underwriters. This report supports continued growth of PV predicated on assurances that properly installed and maintained PV system connectors are safe and reliable. The project team is proposing future work including accelerated testing of connectors and expanding the approach taken here to other PV system components, such as TEA for rapid shut-down devices.
The project goal is to significantly improve the current ocean wave energy harvesting through innovative Power Take-off (PTO) design, advanced power electronics, and novel wave capture structures. The objective of the project is to design and demonstrate system-agnostic components for application across multiple MHK systems, and complete component designs, build scaled prototypes, and perform testing and analysis for metric validation of 25% increase in component rating/per unit cost and 50% reduction in failure rate. The major innovation of the PTO is the Mechanical Motion Rectifier (MMR) mechanism that rectifies the bi-directional oscillatory motion of the input from waves into a steady unidirectional rotation output to directly drive the electrical generator. Through this mechanism, the efficiency and the fatigue life of the PTO can be significantly improved to benefit the energy absorption and lifespan of the wave energy converters (WEC). During the period of performance, the component and system design are completed, the scaled prototypes are developed and performed testing. It is validated that the 25% increase in a component rating/per unit cost. The 50% reduction in failure rate is not directly validated by experiments, however, it can be explained qualitatively with analysis. Besides, 8 journal articles, 13 conference proceedings, 1 patent, 3 Master thesis and two Ph.D. dissertations are published based on the work related to this project. The list of all the publications can be found at the end of the project as an appendix. Over 30 students and postdocs were trained through this project. Three prototypes of 100W and 500W WECs and 10KW PTO were designed, built, and tested in ocean wave tank and using the NREL dynamometer. This project demonstrated 50-80% PTO efficiency, 90-98% power electronics efficiency, up to 66% capture width ratio in irregular waves, and 34% overall efficiency in regular waves.
There is currently no widely agreed, detailed general method for licensing a novel plant incorporating novel materials (or materials being deployed in novel environments); in many such situations, there are no directly applicable engineering code cases for decision-makers (including regulators) to rely on. This paper discusses a framework for solving this problem that is based on the Reliability and Integrity Management (RIM) approach delineated in ASME BPVC Section XI Division 2. NRC Regulatory Guide 1.246, Rev. 0, endorses, with conditions, the subject portion of the ASME Code. The proposed framework is meant to support development of a licensing case by addressing certain technical challenges. The framework discussed here is compatible with the Licensing Modernization Project, but applying it in a specific case will call for advances in the state of practice, if not the state of the art. The RIM approach calls for applicants to (a) allocate reliability targets to plant structures, systems, and components (SSCs), (b) show that they are able to relate the currently observed physical condition of each SSC in the program to its failure probability well enough to determine whether the target reliability allocations are being satisfied, allowing for uncertainty related to the novelty of the materials/designs/operating environments, and (c) be able to demonstrate that the proposed program of surveillances will reliably detect unacceptable degradation of an SSC before SSC failure occurs. These challenges are discussed in the paper, and a potentially applicable modeling approach based on cumulative damage rather than failure rates is briefly illustrated.