Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Common Cause Failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Investigation of the Use of Dynamic Probabilistic Risk Assessment Methodologies for Identifying Digital I&C System Common Cause Failures

Digital Instrumentation and Control (I&C) systems have a key role in nuclear power plants in the upgrade of aging analog systems. Digital systems improve plant safety and reliability through features such as increased hardware reliability and stability and improved failure detection capability. There is no consensus on which of the current probabilistic risk assessment methods are most suitable for use in the reliability analysis of digital I&C systems. While the traditional event-tree/fault-tree (ET/FT) approach is still used for their reliability modeling, there are concerns regarding this approach in properly accounting for dynamic interactions among system components since potentially significant dependencies among failure events may not be identified and/or their likelihood may not be properly quantified. Dynamic methodologies are expected to provide a much more accurate representation of probabilistic evolution of the I&C systems in time due to their capability to more properly account for complex interactions than the static approach. The applicability of dynamic PRA methodologies for digital I&C system is investigated using the criteria presented in the NUREG/CR-6901, and the comparisons made in NUREG/CR-6901 are updated in light of the latest studies. The Dynamic Event Tree (DET) approach has been identified as one of the top dynamic methods when evaluated against the requirements for the reliability modeling of digital I&C systems. The DET method is a strong candidate for integration into existing PRA studies, as it bears many similarities to the traditional ET approach. In this study, the DET approach has been applied to the Plant Protection System of the APR1400 design, and the results are compared to results from its available traditional ET/FT analysis. Possible approaches to evaluate and quantify the effects of common cause failures on system safety using dynamic methods are also examined.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Historical Aerospace Software Errors Categorized to Influence Fault Tolerance

- Motivation - Very little literature exists characterizing software errors in real-time avionic systems - How, where, and why is software most likely to fail? - Purpose - Raise awareness of how software fails through historical study - Recommend improvements to software fault tolerant design based on historical study - Outline - Discuss Software Failures - Common Cause, Failure Classes, Mitigation strategies - Review NASA requirements for Software Fault Tolerance - Review Historical Software Failures - Analyze failures and provide statistics - Erroneous vs. fail-Silent - Reboot recoverability likelihood - Code Location - Missing or unknown code?

Flilght↗

We Can't Count on Repairing All Failures Going to Mars

Reliability analysis often assumes that a complex system can be kept operating indefinitely with scheduled maintenance and emergency repair using a stock of spare parts, as long as the spare parts are not depleted. This assumption seems justified for well-tested, widely used, long operational systems with a multigenerational history of failure, redesign, and reliability growth. It seems doubtful that newer, relatively untried, high technology space systems can always be repaired. We cannot assume space systems will have a low rate of random failures that can all be repaired with a few identical spares. New untried systems usually have a high initial failure rate, called infant mortality, due to errors in requirements, design, parts, materials, and operations planning. These problems can cause groups of related failures called Common Cause Failures (CCFs). The practical definition of a CCF is any failure mode that cannot be cured using identical redundant systems or spare parts. Systems with CCFs may fail repeatedly for the same reason. Can a life support system be kept operating on the way to Mars using only redundant systems and spare parts? The failure history of International Space Station (ISS) life support systems suggests that CCFs are likely to occur and will probably require design changes rather than being reparable with spare parts.

Mars↗

An Approach to Automate tools for the Risk Assessment of Digital Instrumentation and Control Systems

Reliable digital instrumentation and control systems (DI&C) are integral for sustaining the continued operation of nuclear power plants. These systems ensure that nuclear reactors operate safely, efficiently, and within regulatory requirements. Yet, the cost of designing and licensing new nuclear DI&C can be prohibitively expensive. Under the U.S. Department of Energy Light Water Reactor Sustainability Program, Idaho National Laboratory has developed a framework for supporting the risk-informed design of DI&C systems by offering methods to support the identification, quantification, and evaluation of risks for various DI&C design architectures. The framework indicates potential software failure modes and provides pathways for quantifying the potential for these software failures, including common cause failures. Using the framework’s systematic approach, challenges for assessing risks within new and existing nuclear DI&C systems can be reduced. Nevertheless, the current framework can be further improved using the convenience of automation. This paper introduces the development of Software for the Hazard Identification and Evaluation of Digital Systems (SHIELDS). SHIELDS is an engineering software package that enables the identification, elimination, and mitigation of potential risks and reduces the burden of deploying reliable DI&C systems. This work introduces plans and techniques to digitize and improve the manual risk assessment modules of the framework. These improvements will save time and increase the repeatability and usability of the framework, making it more accessible to a wider range of users. Ultimately, this introduces SHIELDS and how its modules support efficient development of safe and reliable DI&C systems.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

An Efficient Approach for the Reliability Analysis of Phased-Mission Systems with Dependent Failures

We consider the reliability analysis of phased-mission systems with common-cause failures in this paper. Phased-mission systems (PMS) are systems supporting missions characterized by multiple, consecutive, and nonoverlapping phases of operation. System components may be subject to different stresses as well as different reliability requirements throughout the course of the mission. As a result, component behavior and relationships may need to be modeled differently from phase to phase when performing a system-level reliability analysis. This consideration poses unique challenges to existing analysis methods. The challenges increase when common-cause failures (CCF) are incorporated in the model. CCF are multiple dependent component failures within a system that are a direct result of a shared root cause, such as sabotage, flood, earthquake, power outage, or human errors. It has been shown by many reliability studies that CCF tend to increase a system's joint failure probabilities and thus contribute significantly to the overall unreliability of systems subject to CCF.We propose a separable phase-modular approach to the reliability analysis of phased-mission systems with dependent common-cause failures as one way to meet the above challenges in an efficient and elegant manner. Our methodology is twofold: first, we separate the effects of CCF from the PMS analysis using the total probability theorem and the common-cause event space developed based on the elementary common-causes; next, we apply an efficient phase-modular approach to analyze the reliability of the PMS. The phase-modular approach employs both combinatorial binary decision diagram and Markov-chain solution methods as appropriate. We provide an example of a reliability analysis of a PMS with both static and dynamic phases as well as CCF as an illustration of our proposed approach. The example is based on information extracted from a Mars orbiter project. The reliability model for this orbiter considers the various phases of Launch, Cruise, Mars Orbit Insertion, and Orbit. Some of the CCF for the orbiter in this mission include environmental effects, such as micrometeoroids, human operator errors, and software errors.

reliability analysis↗

Diverse Redundant Systems for Reliable Space Life Support

Reliable life support systems are required for deep space missions. The probability of a fatal life support failure should be less than one in a thousand in a multi-year mission. It is far too expensive to develop a single system with such high reliability. Using three redundant units would require only that each have a failure probability of one in ten over the mission. Since the system development cost is inverse to the failure probability, this would cut cost by a factor of one hundred. Using replaceable subsystems instead of full systems would further cut cost. Using full sets of replaceable components improves reliability more than using complete systems as spares, since a set of components could repair many different failures instead of just one. Replaceable components would require more tools, space, and planning than full systems or replaceable subsystems. However, identical system redundancy cannot be relied on in practice. Common cause failures can disable all the identical redundant systems. Typical levels of common cause failures will defeat redundancy greater than two. Diverse redundant systems are required for reliable space life support. Three, four, or five diverse redundant systems could be needed for sufficient reliability. One system with lower level repair could be substituted for two diverse systems to save cost.

life support↗

Causal CCF Parameter Estimations 2020

This report documents the quantitative results of the causal common-cause failure (CCF) parameter estimations for the failure cause groups “component,” “design,” “environment,” “human,” and “other,” based on CCF data through 2020 in the U.S. Nuclear Regulatory Commission (NRC) CCF database: https://rads.inl.gov/Pages/CCF.aspx. This report utilizes the same data period (2006–2020) and CCF templates as INL/EXT-21-62940, Revision 1, CCF Parameter Estimations, 2020 Update. The 2015 causal CCF prior distributions for the specific failure cause groups (instead of the 2015 generic CCF prior distributions) were used in this report to estimate the associated causal CCF parameters. All the 2015 causal CCF prior distributions and generic CCF prior distributions were developed in INL/EXT-21-43723, Developing Generic Prior Distributions for Common Cause Failure Alpha Factors and Causal Alpha Factors, using CCF data from 1997 to 2015. These quantitative results were developed to support the causal alpha factor model and should be used as appropriate in probabilistic risk assessment (PRA) studies such as the NRC Significance Determination Process for commercial nuclear power plants in the United States.

99 GENERAL AND MISCELLANEOUS↗

Fault Management Algorithm Risk Assessment for the NASA Space Launch System

This paper presents the false positive (FP) and false negative (FN) risk assessment process currently being conducted for the Space Launch System (SLS) Artemis II Fault Management (FM) detection functions. The analysis scope, general assumptions and guide rules, and key modeling concepts were discussed to establish the basis of the risk assessments conducted. Initial analyses indicated a dominance in the total risk by software and firmware failures. This paper presents efforts applied to refine the software risks and the overall impact of implementing those modifications. Current analyses conducted on the detection functions implemented for the SLS Artemis II mission indicate primary risk drivers for the individual FM detection functions are flight software failures, firmware design failures, and hardware Common Cause Failures (CCFs). There still remains issues of how to account for time and redundancy in the software risk estimations.

probability risk analysis↗

Fault Management Algorithm Risk Assessment for the NASA Space Launch System

This presentation describes the false positive (FP) and false negative (FN) risk assessment process currently being conducted for the Space Launch System (SLS) Artemis II Fault Management (FM) detection functions. The analysis scope, general assumptions and guide rules, and key modeling concepts were discussed to establish the basis of the risk assessments conducted. Initial analyses indicated a dominance in the total risk by software and firmware failures. This paper presents efforts applied to refine the software risks and the overall impact of implementing those modifications. Current analyses conducted on the detection functions implemented for the SLS Artemis II mission indicate primary risk drivers for the individual FM detection functions are flight software failures, firmware design failures, and hardware Common Cause Failures (CCFs). There still remains issues of how to account for time and redundancy in the software risk estimations.

probability risk analysis↗

An Integrated Risk Assessment Process of Safety-Related Digital I&C Systems in Nuclear Power Plants

Upgrading the existing analog instrumentation and control (I&C) systems to state-of-the-art digital I&C (DI&C) systems will greatly benefit existing light water reactors. However, the issue of software common cause failure (CCF) remains an obstacle in terms of qualification for digital technologies. Existing analyses of CCFs in I&C systems mainly focus on hardware failures. With the application and upgrading of new DI&C systems, design flaws could cause software CCFs to become a potential threat to plant safety, considering that most redundancy designs use similar digital platforms or software in their operating and application systems. With complex multilayer redundancy designs to meet the single failure criterion, these I&C safety systems are of particular concern in U.S. Nuclear Regulatory Commission licensing procedures. In Fiscal Year 2019, the Risk-Informed Systems Analysis (RISA) Pathway of the U.S. Department of Energy’s Light Water Reactor Sustainability Program initiated a project to develop a risk assessment strategy for delivering a strong technical basis to support effective, licensable, and secure DI&C technologies for digital upgrades and designs. An integrated risk assessment for the DI&C process was proposed for this strategy to identify potential key digital-induced failures, implement reliability analyses of related digital safety I&C systems, and evaluate the unanalyzed sequences introduced by these failures (particularly software CCFs) at the plant level. Here this paper summarizes these RISA efforts in the risk analysis of safety-related DI&C systems at Idaho National Laboratory.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Operating Experience Data Analysis for Digital Instrumentation and Control System Reliability and Risk Assessment in Nuclear Power Plants

The implementation of advanced digital instrumentation and control (DI&C) systems in U.S. nuclear power plants (NPPs) can bring significant advancements in reliability, monitoring, and control capabilities. However, these systems also introduce new challenges, particularly in assessing risks such as common-cause failures (CCFs) and establishing robust reliability estimates for DI&C components. Addressing these challenges is critical for ensuring the safe and efficient operation of NPPs. Recently, Idaho National Laboratory was tasked by the U.S. Nuclear Regulatory Commission (NRC) to conduct a DI&C reliability study using operating experience data from the nuclear industry. The two operating experience data sources for the study are the Institute of Nuclear Power Operations’ Industry Reporting and Information System (IRIS) and the NRC’s Licensee Event Report database which is hosted at Idaho National Laboratory at https://lersearch.inl.gov/LERSearchCriteria.aspx. This report provides a comprehensive examination of DI&C systems, including their architecture, operational advantages, and associated challenges. It reviews existing industry DI&C studies and failure mode taxonomies, along with reliability data from various industries. Through a detailed analysis of these databases, the study provides insights into DI&C system performance. Considerations should be given to incorporate DI&C failure data into the NRC's Integrated Data Collection and Coding System and updating the Reliability and Availability Data System to support ongoing DI&C reliability studies. Recommendations are also provided for modeling DI&C reliability and CCF in probabilistic risk assessment, thereby supporting risk-informed decision-making and enhancing the reliability and safety of NPPs.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

VIPER project

The VIPER project has so far produced a formal specification of a 32 bit RISC microprocessor, an implementation of that chip in radiation-hard SOS technology, a partial proof of correctness of the implementation which is still being extended, and a large body of supporting software. The time has now come to consider what has been achieved and what directions should be pursued in the future. The most obvious lesson from the VIPER project was the time and effort needed to use formal methods properly. Most of the problems arose in the interfaces between different formalisms, e.g., between the (informal) English description and the HOL spec, between the block-level spec in HOL and the equivalent in ELLA needed by the low-level CAD tools. These interfaces need to be made rigorous or (better) eliminated. VIPER 1A (the latest chip) is designed to operate in pairs, to give protection against breakdowns in service as well as design faults. We have come to regard redundancy and formal design methods as complementary, the one to guard against normal component failures and the other to provide insurance against the risk of the common-cause failures which bedevil reliability predictions. Any future VIPER chips will certainly need improved performance to keep up with increasingly demanding applications. We have a prototype design (not yet specified formally) which includes 32 and 64 bit multiply, instruction pre-fetch, more efficient interface timing, and a new instruction to allow a quick response to peripheral requests. Work is under way to specify this device in MIRANDA, and then to refine the spec into a block-level design by top-down transformations. When the refinement is complete, a relatively simple proof checker should be able to demonstrate its correctness. This paper is presented in viewgraph form.

Kershaw, John↗

Survey of Aging and Monitoring Concerns for Cables and Splices Due to Cable Repair and Replacement

The purpose of this report is to survey aging and monitoring concerns for electrical cable splices in nuclear power plants (NPPs) in long term operation. As portions of existing electrical cable runs in nuclear are replaced over time due to localized events, the total number of splices in NPPs is expected to increase. Relative to cables, the body of knowledge regarding aging of splices and splices in combination with aging cables in nuclear service environments in long-term operations is low. A few reports have considered the aging of cable system components other than cables (Jacobus 1990; Nelson 1998; Villaran and Lofaro 2002), but the nuclear industry has two decades of operating experience since these were published to further enlighten this issue. Herein we discuss electrical cables and splices commonly found in U.S. nuclear power plants, their qualification in safety-related application, and methods for monitoring their health condition. Common environmental stresses that can give rise to cable and splice failure are discussed. The Nuclear Regulatory Commission (NRC) Licensee Event Reports (LER) database was used to identify documented issues of cable and splice failure. The trend in the resultant data over time is considered to see if failures are increasing as plants age. Observations and conclusions of this work include: 1. Cables and splices are highly reliable components. Occurrence rates for events of interest were low and nearly constant over the last 20 years. 2. Common-cause failure for evaluated cable events of interest was observed to primarily be associated with loose connections, which may manifest associated with workmanship issues, thermal cycling, and/or vibration. 3. Replacement of cables is more common than repair, leading to an increase in proportion of new generation cables in the plant over time. 4. Splices on degraded cables have been observed to be problematic. Due to aging NPP infrastructure, including electrical cables, it is expected that such issues will continue to increase. 5. Condition monitoring approaches, while shown to be fruitful for cables, have been shown to be insensitive to degradation of splice sleeves, which are critical to the continued performance of splices. Additional condition monitoring (CM) work is needed to evaluate methods which are sensitive to the degradation of splice components. 6. Extended Material Degradation Assessment (EMDA) knowledge gaps for electrical cables (Bernstein et al. 2014) have not been investigated for splices but may represent similar concerns such as for the accelerated aging process historically used in environmental qualification.

42 ENGINEERING↗

Bayesian And Human Reliability Analysis (hra)-aided Method For The Reliability Analysis Of Software (bahamas)

The purpose of the BAHAMAS code is to provide a simplified process for performing quantitative evaluations of software reliability. The Bayesian and Human Reliability Analysis (HRA)-Aided method for the Reliability Analysis of software (BAHAMAS) was developed specifically to perform quantification under limited data conditions, i.e., when limited testing or operational data are available, such as during early development stages. BAHAMAS essentially examines the quality of a software development life cycle to determine the probability of specific types of software failure. BAHAMAS will have modules to support user input for detailed and simplified analyses. The user interface will also support software common cause failure analysis.

Wang, Congjian (0000000207789927)↗

Comparative Analysis of Static and Dynamic Probabilistic Risk Assessment

This study examines three different methodologies for producing loss-of-mission (LOM) and loss-of-crew (LOC) risks estimates for probabilistic risk assessments (PRA) of crewed spacecraft. The three bottom-up, component-based PRA approaches examined are a traditional static fault tree, a dynamic Monte Carlo simulation, and a fault tree hybrid that incorporates some dynamic elements. These approaches were used to model the reaction control system thruster pod of a generic crewed spacecraft and mission, and a comparative analysis of the methods is presented. The methodologies are assessed in terms of the process of modeling a system, the actionable information produced for the design team, and the overall fidelity of the quantitative risk evaluation generated. The system modeling process is compared in terms of the effort required to generate the initial model, update the model in response to design changes, and support mass-versus-risk trade studies. The results are compared by examining the top-level LOM/LOC estimates and the relative risk driver rankings at the failure mode level. The fidelity of each modeling methodology is discussed in terms of its capability to handle real-world system dynamics such as cold-sparing, changes in mission operations due to loss of redundancy, and common cause failure modes. The paper also discusses the applicability of each methodology to different phases of system development and shows that a single methodology may not be suitable for all of the many purposes of a spacecraft PRA. The fault tree hybrid approach is shown to be best suited to the needs of early assessments during conceptual design phases. As the design begins to mature, the level of detail represented in the risk model must go beyond redundancy and nominal mission operations to include dynamic, time- and state-dependent system responses as well as diverse system capabilities. This is best accomplished using the dynamic simulation approach, since these phenomena are not easily captured by static methods. Ultimately, once the design has been finalized and the goal of the PRA is to provide design validation and requirement verification, more traditional, static fault tree approaches may become as appropriate as the simulation method.

Mattenberger, Christopher J.↗

Conical Seat Shut-Off Valve

A moveable valve for controlling flow of a pressurized working fluid was designed. This valve consists of a hollow, moveable floating piston pressed against a stationary solid seat, and can use the working fluid to seal the valve. This open/closed, novel valve is able to use metal-to-metal seats, without requiring seat sliding action; therefore there are no associated damaging effects. During use, existing standard high-pressure ball valve seats tend to become damaged during rotation of the ball. Additionally, forces acting on the ball and stem create large amounts of friction. The combination of these effects can lead to system failure. In an attempt to reduce damaging effects and seat failures, soft seats in the ball valve have been eliminated; however, the sliding action of the ball across the highly loaded seat still tends to scratch the seat, causing failure. Also, in order to operate, ball valves require the use of large actuators. Positioning the metal-to-metal seats requires more loading, which tends to increase the size of the required actuator, and can also lead to other failures in other areas such as the stem and bearing mechanisms, thus increasing cost and maintenance. This novel non-sliding seat surface valve allows metal-to-metal seats without the damaging effects that can lead to failure, and enables large seating forces without damaging the valve. Additionally, this valve design, even when used with large, high-pressure applications, does not require large conventional valve actuators and the valve stem itself is eliminated. Actuation is achieved with the use of a small, simple solenoid valve. This design also eliminates the need for many seals used with existing ball valve and globe valve designs, which commonly cause failure, too. This, coupled with the elimination of the valve stem and conventional valve actuator, improves valve reliability and seat life. Other mechanical liftoff seats have been designed; however, they have only resulted in increased cost, and incurred other reliability issues. With this novel design, the seat is lifted by simply removing the working fluid pressure that presses it against the seat and no external force is required. By eliminating variables associated with existing ball and globe configurations that can have damaging effects upon a valve, this novel design reduces downtime in rocket engine test schedules and maintenance costs.

Farner, Bruce↗