Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Common Cause Failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

An Application of a Modified Beta Factor Method for the Analysis of Software Common Cause Failures

This paper presents an approach for modeling software common cause failures (CCFs) within digital instrumentation and control (I&C) systems. CCFs consist of a concurrent failure between two or more components due to a shared failure cause and coupling mechanism. This work emphasizes the importance of identifying software-centric attributes related to the coupling mechanisms necessary for simultaneous failures of redundant software components. The groups of components that share coupling mechanisms are called common cause component groups (CCCGs). Most CCF models rely on operational data as the basis for establishing CCCG parameters and predicting CCFs. This work is motivated by two primary concerns: (1) a lack of operational and CCF data for estimating software CCF model parameters; and (2) the need to model single components as part of multiple CCCGs simultaneously. A hybrid approach was developed to account for these concerns by leveraging existing techniques: a modified beta factor model allows single components to be placed within multiple CCCGs, while a second technique provides software-specific model parameters for each CCCG. This hybrid approach provides a means to overcome the limitations of conventional methods while offering support for design decisions under the limited data scenario.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Developing Component-Specific Prior Distributions for Common Cause Failure Alpha Factors

This report presents the development of component-specific common cause failure (CCF) prior distributions for five component categories: pump, valve, strainer, generator, and other. For ease in comparing these component-specific CCF priors with the 2015 generic CCF priors, the same method used to develop the 2015 generic CCF priors was also employed in this report, along with the same set of failure data (1997–2015) featured in INL/EXT-21-43723. For selected CCF templates, this report also evaluates the effects of applying component-specific priors to CCF parameters.

99 GENERAL AND MISCELLANEOUS↗

Common Cause Failure Analysis for Nuclear Power Plant Instrumentation and Control Systems

This presentation is prepared for the IAEA Virtual Consultancy Meeting on Preparation of a Coordinated Research Project on Common Cause Failures in Nuclear Power Plant Instrumentation and Control Systems from February 5 ? 8, 2024. It provides an overview of the INL activities on the common cause failure analysis for instrumentation and control system in nuclear power plants.

99 GENERAL AND MISCELLANEOUS↗

Analyzing Hardware and Software Common Cause Failures in Digital Instrumentation and Control Systems using Dual Error Propagation Method

This paper develops a methodology for quantifying software common cause failures (CCFs) in digital instrumentation and control (I&C) systems of nuclear power plants. To support the transition of analog I&C systems to digital in nuclear power plants, probabilistic risk assessment (PRA) techniques are used. The hardware components of the I&C systems have reliability databases that can be used in the PRA studies. However, the failure data for redundant software components of the systems is sparse. Failure of components constitutes a CCF, wherein two or more components or systems fail due to a single shared cause and coupling mechanism. This paper proposes a quantification approach that can simultaneously model hardware and software components, incorporate the CCFs of software systems in the models, and bridge the gap between the failure quantification of models and the development of CCF parametric databases. We demonstrate the dual error propagation method (DEPM) by developing I&C systems failure models for a representative digital reactor trip system. The DEPM models are built to simulate the control and data flows within the systems and can accommodate failure states. By expanding DEPM to software CCFs, we generated alpha factor parameter estimates for each of the modeled error propagation mechanisms.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Common Cause Failure Evaluation of High Safety-significant Safety-related Digital Instrumentation and Control Systems

Digital instrumentation and control (DI&C) systems in nuclear power plants (NPPs) have many advantages over analog systems but also pose different engineering and technical challenges, such as potential threats due to common cause failures (CCFs). This paper proposes a Platform for Risk Assessment of DI&C (PRADIC) developed by Idaho National Laboratory for dealing with potential software CCFs in DI&C systems of NPPs. The methodology development of PRADIC on the quantitative evaluation of software CCFs in high safety-significant safety-related DI&C systems in NPPs is illustrated in this paper. In PRADIC, qualitative hazard analysis and quantitative reliability and consequence analysis are successively implemented to obtain quantitative risk information, compare with respective risk evaluation acceptance criteria, and provide suggestions for risk reduction and design optimization. A comprehensive case study was also performed and documented in this paper. Results show that PRADIC can effectively identify potential digital-based CCFs, estimate their failure probabilities, and evaluate their impacts to system and plant safety.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Common Cause Failure Evaluation of High Safety Significant Safety-related Digital Instrumentation and Control Systems using IRADIC Technology

Digital instrumentation and control (DI&C) systems in nuclear power plants (NPPs) have many advantages over analog systems but also pose different engineering and technical challenges, such as potential threats due to common cause failures (CCFs). This paper proposes an integrated risk assessment technology for DI&C systems (IRADIC) developed by Idaho National Laboratory for dealing with potential software CCFs in DI&C systems of NPPs. The methodology development of the IRADIC technology on the quantitative evaluation of software CCFs in high safety-significant safety-related DI&C systems in NPPs is illustrated in this paper. In IRADIC, qualitative hazard analysis and quantitative reliability and consequence analysis are successively implemented to obtain quantitative risk information, compare with respective risk evaluation acceptance criteria, and provide suggestions for risk reduction and design optimization. A comprehensive case study was also performed and documented in this paper. Results show that the IRADIC technology can effectively identify potential digital-based CCFs, estimate their failure probabilities, and evaluate their impacts to system and plant safety.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Quantitative evaluation of common cause failures in high safety-significant safety-related digital instrumentation and control systems in nuclear power plants

Digital instrumentation and control (DI&C) systems at nuclear power plants (NPPs) have many advantages over analog systems. They are proven to be more reliable, cheaper, and easier to maintain given obsolescence of analog components. However, they also pose new engineering and technical challenges, such as possibility of common cause failures (CCFs) unique to digital systems. Here this paper proposes a Platform for Risk Assessment of DI&C (PRADIC) that is developed by Idaho National Laboratory (INL). A methodology for evaluation of software CCFs in high safety-significant safety-related DI&C systems of NPPs was developed as part of the framework. The framework integrates three stages of a typical risk assessment—qualitative hazard analysis and quantitative reliability and consequence analyses. The quantified risks compared with respective acceptance criteria provide valuable insights for system architecture alternatives allowing design optimization in terms of risk reduction and cost savings. A comprehensive case study performed to demonstrate the framework's capabilities is documented here in this paper. Results show that the PRADIC is a powerful tool capable to identify potential digital-based CCFs, estimate their probabilities, and evaluate their impacts on system and plant safety.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Investigation of the Use of Dynamic Probabilistic Risk Assessment Methodologies for Identifying Digital I&C System Common Cause Failures

Digital Instrumentation and Control (I&C) systems have a key role in nuclear power plants in the upgrade of aging analog systems. Digital systems improve plant safety and reliability through features such as increased hardware reliability and stability and improved failure detection capability. There is no consensus on which of the current probabilistic risk assessment methods are most suitable for use in the reliability analysis of digital I&C systems. While the traditional event-tree/fault-tree (ET/FT) approach is still used for their reliability modeling, there are concerns regarding this approach in properly accounting for dynamic interactions among system components since potentially significant dependencies among failure events may not be identified and/or their likelihood may not be properly quantified. Dynamic methodologies are expected to provide a much more accurate representation of probabilistic evolution of the I&C systems in time due to their capability to more properly account for complex interactions than the static approach. The applicability of dynamic PRA methodologies for digital I&C system is investigated using the criteria presented in the NUREG/CR-6901, and the comparisons made in NUREG/CR-6901 are updated in light of the latest studies. The Dynamic Event Tree (DET) approach has been identified as one of the top dynamic methods when evaluated against the requirements for the reliability modeling of digital I&C systems. The DET method is a strong candidate for integration into existing PRA studies, as it bears many similarities to the traditional ET approach. In this study, the DET approach has been applied to the Plant Protection System of the APR1400 design, and the results are compared to results from its available traditional ET/FT analysis. Possible approaches to evaluate and quantify the effects of common cause failures on system safety using dynamic methods are also examined.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

An Approach to Automate tools for the Risk Assessment of Digital Instrumentation and Control Systems

Reliable digital instrumentation and control systems (DI&C) are integral for sustaining the continued operation of nuclear power plants. These systems ensure that nuclear reactors operate safely, efficiently, and within regulatory requirements. Yet, the cost of designing and licensing new nuclear DI&C can be prohibitively expensive. Under the U.S. Department of Energy Light Water Reactor Sustainability Program, Idaho National Laboratory has developed a framework for supporting the risk-informed design of DI&C systems by offering methods to support the identification, quantification, and evaluation of risks for various DI&C design architectures. The framework indicates potential software failure modes and provides pathways for quantifying the potential for these software failures, including common cause failures. Using the framework’s systematic approach, challenges for assessing risks within new and existing nuclear DI&C systems can be reduced. Nevertheless, the current framework can be further improved using the convenience of automation. This paper introduces the development of Software for the Hazard Identification and Evaluation of Digital Systems (SHIELDS). SHIELDS is an engineering software package that enables the identification, elimination, and mitigation of potential risks and reduces the burden of deploying reliable DI&C systems. This work introduces plans and techniques to digitize and improve the manual risk assessment modules of the framework. These improvements will save time and increase the repeatability and usability of the framework, making it more accessible to a wider range of users. Ultimately, this introduces SHIELDS and how its modules support efficient development of safe and reliable DI&C systems.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Causal CCF Parameter Estimations 2020

This report documents the quantitative results of the causal common-cause failure (CCF) parameter estimations for the failure cause groups “component,” “design,” “environment,” “human,” and “other,” based on CCF data through 2020 in the U.S. Nuclear Regulatory Commission (NRC) CCF database: https://rads.inl.gov/Pages/CCF.aspx. This report utilizes the same data period (2006–2020) and CCF templates as INL/EXT-21-62940, Revision 1, CCF Parameter Estimations, 2020 Update. The 2015 causal CCF prior distributions for the specific failure cause groups (instead of the 2015 generic CCF prior distributions) were used in this report to estimate the associated causal CCF parameters. All the 2015 causal CCF prior distributions and generic CCF prior distributions were developed in INL/EXT-21-43723, Developing Generic Prior Distributions for Common Cause Failure Alpha Factors and Causal Alpha Factors, using CCF data from 1997 to 2015. These quantitative results were developed to support the causal alpha factor model and should be used as appropriate in probabilistic risk assessment (PRA) studies such as the NRC Significance Determination Process for commercial nuclear power plants in the United States.

99 GENERAL AND MISCELLANEOUS↗

An Integrated Risk Assessment Process of Safety-Related Digital I&C Systems in Nuclear Power Plants

Upgrading the existing analog instrumentation and control (I&C) systems to state-of-the-art digital I&C (DI&C) systems will greatly benefit existing light water reactors. However, the issue of software common cause failure (CCF) remains an obstacle in terms of qualification for digital technologies. Existing analyses of CCFs in I&C systems mainly focus on hardware failures. With the application and upgrading of new DI&C systems, design flaws could cause software CCFs to become a potential threat to plant safety, considering that most redundancy designs use similar digital platforms or software in their operating and application systems. With complex multilayer redundancy designs to meet the single failure criterion, these I&C safety systems are of particular concern in U.S. Nuclear Regulatory Commission licensing procedures. In Fiscal Year 2019, the Risk-Informed Systems Analysis (RISA) Pathway of the U.S. Department of Energy’s Light Water Reactor Sustainability Program initiated a project to develop a risk assessment strategy for delivering a strong technical basis to support effective, licensable, and secure DI&C technologies for digital upgrades and designs. An integrated risk assessment for the DI&C process was proposed for this strategy to identify potential key digital-induced failures, implement reliability analyses of related digital safety I&C systems, and evaluate the unanalyzed sequences introduced by these failures (particularly software CCFs) at the plant level. Here this paper summarizes these RISA efforts in the risk analysis of safety-related DI&C systems at Idaho National Laboratory.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Operating Experience Data Analysis for Digital Instrumentation and Control System Reliability and Risk Assessment in Nuclear Power Plants

The implementation of advanced digital instrumentation and control (DI&C) systems in U.S. nuclear power plants (NPPs) can bring significant advancements in reliability, monitoring, and control capabilities. However, these systems also introduce new challenges, particularly in assessing risks such as common-cause failures (CCFs) and establishing robust reliability estimates for DI&C components. Addressing these challenges is critical for ensuring the safe and efficient operation of NPPs. Recently, Idaho National Laboratory was tasked by the U.S. Nuclear Regulatory Commission (NRC) to conduct a DI&C reliability study using operating experience data from the nuclear industry. The two operating experience data sources for the study are the Institute of Nuclear Power Operations’ Industry Reporting and Information System (IRIS) and the NRC’s Licensee Event Report database which is hosted at Idaho National Laboratory at https://lersearch.inl.gov/LERSearchCriteria.aspx. This report provides a comprehensive examination of DI&C systems, including their architecture, operational advantages, and associated challenges. It reviews existing industry DI&C studies and failure mode taxonomies, along with reliability data from various industries. Through a detailed analysis of these databases, the study provides insights into DI&C system performance. Considerations should be given to incorporate DI&C failure data into the NRC's Integrated Data Collection and Coding System and updating the Reliability and Availability Data System to support ongoing DI&C reliability studies. Recommendations are also provided for modeling DI&C reliability and CCF in probabilistic risk assessment, thereby supporting risk-informed decision-making and enhancing the reliability and safety of NPPs.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Survey of Aging and Monitoring Concerns for Cables and Splices Due to Cable Repair and Replacement

The purpose of this report is to survey aging and monitoring concerns for electrical cable splices in nuclear power plants (NPPs) in long term operation. As portions of existing electrical cable runs in nuclear are replaced over time due to localized events, the total number of splices in NPPs is expected to increase. Relative to cables, the body of knowledge regarding aging of splices and splices in combination with aging cables in nuclear service environments in long-term operations is low. A few reports have considered the aging of cable system components other than cables (Jacobus 1990; Nelson 1998; Villaran and Lofaro 2002), but the nuclear industry has two decades of operating experience since these were published to further enlighten this issue. Herein we discuss electrical cables and splices commonly found in U.S. nuclear power plants, their qualification in safety-related application, and methods for monitoring their health condition. Common environmental stresses that can give rise to cable and splice failure are discussed. The Nuclear Regulatory Commission (NRC) Licensee Event Reports (LER) database was used to identify documented issues of cable and splice failure. The trend in the resultant data over time is considered to see if failures are increasing as plants age. Observations and conclusions of this work include: 1. Cables and splices are highly reliable components. Occurrence rates for events of interest were low and nearly constant over the last 20 years. 2. Common-cause failure for evaluated cable events of interest was observed to primarily be associated with loose connections, which may manifest associated with workmanship issues, thermal cycling, and/or vibration. 3. Replacement of cables is more common than repair, leading to an increase in proportion of new generation cables in the plant over time. 4. Splices on degraded cables have been observed to be problematic. Due to aging NPP infrastructure, including electrical cables, it is expected that such issues will continue to increase. 5. Condition monitoring approaches, while shown to be fruitful for cables, have been shown to be insensitive to degradation of splice sleeves, which are critical to the continued performance of splices. Additional condition monitoring (CM) work is needed to evaluate methods which are sensitive to the degradation of splice components. 6. Extended Material Degradation Assessment (EMDA) knowledge gaps for electrical cables (Bernstein et al. 2014) have not been investigated for splices but may represent similar concerns such as for the accelerated aging process historically used in environmental qualification.

42 ENGINEERING↗

Bayesian And Human Reliability Analysis (hra)-aided Method For The Reliability Analysis Of Software (bahamas)

The purpose of the BAHAMAS code is to provide a simplified process for performing quantitative evaluations of software reliability. The Bayesian and Human Reliability Analysis (HRA)-Aided method for the Reliability Analysis of software (BAHAMAS) was developed specifically to perform quantification under limited data conditions, i.e., when limited testing or operational data are available, such as during early development stages. BAHAMAS essentially examines the quality of a software development life cycle to determine the probability of specific types of software failure. BAHAMAS will have modules to support user input for detailed and simplified analyses. The user interface will also support software common cause failure analysis.

Wang, Congjian (0000000207789927)↗

Failure Mechanism Traceability and Application in Human System Interface of Nuclear Power Plants using RESHA

In recent years, there has been considerable effort to modernize existing and new nuclear power plants with digital instrumentation and control systems (DI&C). However, there has also been considerable concern both by industry and regulatory bodies for the risk and consequence analysis of these systems. Of particular concern are digital common cause failures (CCFs) specifically related to software defects. These “misbehaviors” by the software can occur in both the control and monitoring of a system. While many new methods have been proposed to identify potential software failure modes, such as Systems-theoretic Process Analysis (STPA), Hazard and Consequence Analysis for Digital Systems (HAZCADS), etc., these methods are focused primarily on the control action pathway of a system. In contrast, the information feedback pathway lacks unsafe control actions (UCAs), which are typically related to software basic events; thus, assessment of software basic events in such systems is unclear. In this work, we present the idea of intermediate processors and unsafe information flow (UIF) to help safety analysts trace failure mechanisms in the feedback pathway and how they can be integrated into a fault tree for improved assessment capability. The concepts presented are demonstrated in two comprehensive case studies, a smart sensor integrated platform for unmanned autonomous vehicles and another on a representative advanced human system interface (HSI) for safety critical plant monitoring. The qualitative software basic events are identified, and a fault tree analysis is conducted based on a modified Redundancy-guided Systems-theoretic Hazard Analysis (RESHA) methodology. The case studies demonstrate the use of UIF and intermediate processors in the fault tree to improve traceability of software failures in highly complex digital instrumentation feedback. The improved method can also clarify fault tree construction when multiple component dependencies are present in the system.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Failure Mechanism Traceability and Application in Human System Interface of Nuclear Power Plants using RESHA

In recent years, there has been considerable effort to modernize existing and new nuclear power plants with digital instrumentation and control systems (DI&C). However, there has also been considerable concern both by industry and regulatory bodies for the risk and consequence analysis of these systems. Of particular concern are digital common cause failures (CCFs) specifically related to software defects. These “misbehaviors” by the software can occur in both the control and monitoring of a system. While many new methods have been proposed to identify potential software failure modes, such as Systems-theoretic Process Analysis (STPA), Hazard and Consequence Analysis for Digital Systems (HAZCADS), etc., these methods are focused primarily on the control action pathway of a system. In contrast, the information feedback pathway lacks unsafe control actions (UCAs), which are typically related to software basic events; thus, assessment of software basic events in such systems is unclear. In this work, we present the idea of intermediate processors and unsafe information flow (UIF) to help safety analysts trace failure mechanisms in the feedback pathway and how they can be integrated into a fault tree for improved assessment capability. The concepts presented are demonstrated in two comprehensive case studies, a smart sensor integrated platform for unmanned autonomous vehicles and another on a representative advanced human system interface (HSI) for safety critical plant monitoring. The qualitative software basic events are identified, and a fault tree analysis is conducted based on a modified Redundancy-guided Systems-theoretic Hazard Analysis (RESHA) methodology. The case studies demonstrate the use of UIF and intermediate processors in the fault tree to improve traceability of software failures in highly complex digital instrumentation feedback. The improved method can also clarify fault tree construction when multiple component dependencies are present in the system.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Evaluation of Hardware and Software Bill of Materials (HBOMs/SBOMs) Extraction Methods

Hardware and software bills of materials (HBOMs and SBOMs) provide important visibility into the components, dependencies, and supply chain relationships within programmable digital devices. This visibility is critical for advanced nuclear reactor applications, where use of common or shared hardware components, software libraries, suppliers, or manufacturing processes may create common cause failure (CCF) vulnerabilities despite apparent diversity. This paper evaluates current approaches for obtaining and analyzing HBOMs and SBOMs in support of CCF, diversity and defense-in-depth (D3) assessments, and begins to explore potential methods for artificial intelligence/machine learning-based analysis. The availability of BOM information from advanced reactor manufacturers and vendors, representative hardware and software categories found in advanced reactor systems continues to limit research [13]. This paper compares commonly used BOM formats, including CycloneDX, SPDX, and SWID. It also surveys publicly available tools for generating BOMs from source code, compiled binaries, and hardware-related information, noting limitations in language coverage, system age, and format interoperability. Finally, this paper evaluates methods for correlating BOM data with vulnerability and exploitability information, including VEX, CVE, and CWE resources. The findings indicate that publicly available nuclear-vendor BOMs are limited, making third-party extraction and research into novel analysis techniques necessary.

Cybersecurity↗

An Integrated Framework for Risk Assessment of Safety-related Digital Instrumentation and Control Systems in Nuclear Power Plants: Methodology Refinement and Exploration

This report documents activities performed by Idaho National Laboratory (INL) during Fiscal Year (FY) 2023 for the U.S. Department of Energy (DOE) Light Water Reactor Sustainability (LWRS) Program, Risk Informed Systems Analysis (RISA) Pathway, digital instrumentation and control (DI&C) risk assessment project. In FY 2019, the RISA Pathway initiated a project to develop a risk assessment strategy for delivering a technical basis to support effective, and secure DI&C technologies for digital upgrades/designs. A risk assessment-informed framework was proposed for this strategy, which aims to (1) provide a best-estimate, risk informed capability to quantitatively estimate the safety margin obtained from plant modernization, especially for safety-related DI&C systems, (2) support and supplement existing risk informed DI&C design guides by providing quantitative risk information and evidence, (3) offer a capability of design architecture evaluation of various DI&C systems, (4) assure the long-term safety and reliability of safety-related DI&C systems, and (5) reduce uncertainty in costs and support integration of DI&C systems in the plant. To achieve these technical goals, the LWRS-developed framework provides a means to address relevant technical issues by: (1) defining a risk informed analysis process for DI&C upgrade that integrates hazard analysis, reliability analysis, and consequence analysis, (2) applying risk informed tools to address common cause failures (CCFs) and quantify corresponding failure probabilities for DI&C technologies, particularly software CCFs, (3) evaluating the impact of digital failures at the component level, system level, and plant level, and (4) providing insights and suggestions on designs to manage the risks, thus to support the development and deployment of advanced DI&C technologies in nuclear power plants (NPPs). Adding diversity within a system or components is the primary means to eliminate and mitigate CCFs, but diversity also increases system complexity and may not address all sources of systematic failures. Optimization of diversity and redundancy applications for the safety-critical DI&C systems remains a challenge. To deal with the technical issues in addressing potential software CCFs in safety-related DI&C systems of NPPs and supporting relevant design optimization, the proposed framework provides: (a) A best-estimate, risk informed capability to address new technical digital issues quantitatively, focusing on software CCFs in safety-related DI&C systems of NPPs; (b) A common and a modularized platform for DI&C designers, software developers, cybersecurity analysts, and plant engineers to predict and prevent risk in the early design stage of DI&C systems; (c) Technical bases and risk informed insights to assist users address the risk informed alternatives for evaluation of CCFs in safety-related DI&C systems of NPPs; and (d) A risk informed tool that offers a capability of design architecture evaluation of various DI&C systems to support system design decisions in diversity and redundancy applications. The research and development efforts of this project in FY 2023 are focused on refining current methods on software CCF modeling and estimation and exploring additional innovative approaches to risk assessment of DI&C systems to enable a more comprehensive and complete assessment of various safety-related DI&C design architectures. The primary audience of this report are DI&C designers, engineers, and probabilistic risk assessment (PRA) practitioners. This includes stakeholders, such as the nuclear utilities and regulators who consider the deployment and upgrade of DI&C systems, DI&C software developers and reviewers, and cybersecurity specialists. It should be noted that all the analyses are performed for the demonstration of the methodology, not for the evaluation of an actual digital control system. Results are obtained based on limited design information and testing data.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗