Engineering Papers⌕ Search

Engineering topics

Elks, Carl

Publications and source records attributed to Elks, Carl.

An Integrated Framework for Risk Assessment of High Safety Significant Safety-related Digital Instrumentation and Control Systems in Nuclear Power Plants: Methodology and Demonstration

This report documents the activities performed by Idaho National Laboratory (INL) during Fiscal Year (FY) 2022 for the U.S. Department of Energy (DOE) Light Water Reactor Sustainability (LWRS) Program, Risk Informed Systems Analysis (RISA) Pathway, digital instrumentation and control (DI&C) risk assessment project. In FY 2019, the RISA Pathway initiated a project to develop a risk assessment strategy for delivering a technical basis to support effective and secure DI&C technologies for digital upgrades/designs. A framework was proposed for this strategy, which aims to (1) provide a best-estimate, risk-informed capability to quantitatively and accurately estimate the risk impact of plant modernization, considering the introduction of high safety-significant safety-related (HSSSR) DI&C systems, (2) support and supplement existing risk-informed DI&C design guides by providing quantitative risk information and evidence, (3) offer a capability of design architecture evaluation of various DI&C systems, (4) assure the long-term safety and reliability of HSSSR DI&C systems, and (5) reduce uncertainty in costs and support integration of DI&C systems in the plant. To achieve these technical goals, the framework provides a means to address relevant technical issues by: (1) defining a risk-informed analysis process for DI&C upgrade, that integrates hazard analysis, reliability analysis, and consequence analysis, (2) applying risk-informed tools to address common cause failures (CCFs) and quantify corresponding failure probabilities for DI&C technologies, particularly software CCFs, (3) evaluating the impact of digital failures at the component level, system level, and plant level, and (4) providing insights and suggestions on designs to manage the risks, thus to support the development and deployment of advanced DI&C technologies on nuclear power plant (NPPs).

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Failure Mechanism Traceability and Application in Human System Interface of Nuclear Power Plants using RESHA

In recent years, there has been considerable effort to modernize existing and new nuclear power plants with digital instrumentation and control systems (DI&C). However, there has also been considerable concern both by industry and regulatory bodies for the risk and consequence analysis of these systems. Of particular concern are digital common cause failures (CCFs) specifically related to software defects. These “misbehaviors” by the software can occur in both the control and monitoring of a system. While many new methods have been proposed to identify potential software failure modes, such as Systems-theoretic Process Analysis (STPA), Hazard and Consequence Analysis for Digital Systems (HAZCADS), etc., these methods are focused primarily on the control action pathway of a system. In contrast, the information feedback pathway lacks unsafe control actions (UCAs), which are typically related to software basic events; thus, assessment of software basic events in such systems is unclear. In this work, we present the idea of intermediate processors and unsafe information flow (UIF) to help safety analysts trace failure mechanisms in the feedback pathway and how they can be integrated into a fault tree for improved assessment capability. The concepts presented are demonstrated in two comprehensive case studies, a smart sensor integrated platform for unmanned autonomous vehicles and another on a representative advanced human system interface (HSI) for safety critical plant monitoring. The qualitative software basic events are identified, and a fault tree analysis is conducted based on a modified Redundancy-guided Systems-theoretic Hazard Analysis (RESHA) methodology. The case studies demonstrate the use of UIF and intermediate processors in the fault tree to improve traceability of software failures in highly complex digital instrumentation feedback. The improved method can also clarify fault tree construction when multiple component dependencies are present in the system.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Application of Orthogonal Defect Classification for Software Reliability Analysis

The modernization of existing and new nuclear power plants with digital instrumentation and control systems (DI&C) is a recent and highly trending topic. However, there lacks strong consensus on best-estimate reliability methodologies by both the United States (U.S.) Nuclear Regulatory Commission (NRC) and the industry. This has resulted in hesitation for further modernization projects until a more unified methodology is realized. In this work, we develop an approach called Orthogonal-defect Classification for Assessing Software Reliability (ORCAS) to quantify probabilities of various software failure modes in a DI&C system. The method utilizes accepted industry methodologies for software quality assurance that are also verified by experimental or mathematical formulations. In essence, the approach combines a semantic failure classification model with a reliability growth model to predict (and quantify) the potential failure modes of a DI&C software system. The semantic classification model is used to address the question: How do latent defects in software contribute to different software failure root causes? The use of reliability growth models is then used to address the question: Given the connection between latent defects and software failure root causes, how can we quantify the reliability of the software? A case study was conducted on a representative I&C platform (ChibiOS) running a smart sensor acquisition software developed by Virginia Commonwealth University (VCU). The testing and evidence collection guidance in ORCAS was applied, and defects were uncovered in the software. Qualitative evidence, such as condition coverage, was used to gauge the completeness and trustworthiness of the assessment while quantitative evidence was used to determine the software failure probabilities. The reliability of the software was then estimated and compared to existing operational data of the sensor device. It is demonstrated that by using ORCAS, a semantic reasoning framework can be developed to justify if the software is reliable (or unreliable) while still leveraging the strength of the existing methods.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Failure Mechanism Traceability and Application in Human System Interface of Nuclear Power Plants using RESHA

In recent years, there has been considerable effort to modernize existing and new nuclear power plants with digital instrumentation and control systems (DI&C). However, there has also been considerable concern both by industry and regulatory bodies for the risk and consequence analysis of these systems. Of particular concern are digital common cause failures (CCFs) specifically related to software defects. These “misbehaviors” by the software can occur in both the control and monitoring of a system. While many new methods have been proposed to identify potential software failure modes, such as Systems-theoretic Process Analysis (STPA), Hazard and Consequence Analysis for Digital Systems (HAZCADS), etc., these methods are focused primarily on the control action pathway of a system. In contrast, the information feedback pathway lacks unsafe control actions (UCAs), which are typically related to software basic events; thus, assessment of software basic events in such systems is unclear. In this work, we present the idea of intermediate processors and unsafe information flow (UIF) to help safety analysts trace failure mechanisms in the feedback pathway and how they can be integrated into a fault tree for improved assessment capability. The concepts presented are demonstrated in two comprehensive case studies, a smart sensor integrated platform for unmanned autonomous vehicles and another on a representative advanced human system interface (HSI) for safety critical plant monitoring. The qualitative software basic events are identified, and a fault tree analysis is conducted based on a modified Redundancy-guided Systems-theoretic Hazard Analysis (RESHA) methodology. The case studies demonstrate the use of UIF and intermediate processors in the fault tree to improve traceability of software failures in highly complex digital instrumentation feedback. The improved method can also clarify fault tree construction when multiple component dependencies are present in the system.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Application of Orthogonal Defect Classification for Software Reliability Analysis

The modernization of existing and new nuclear power plants with digital instrumentation and control systems (DI&C) is a recent and highly trending topic. However, there lacks strong consensus on best-estimate reliability methodologies by both the United States (U.S.) Nuclear Regulatory Commission (NRC) and the industry. This has resulted in hesitation for further modernization projects until a more unified methodology is realized. In this work, we develop an approach called Orthogonal-defect Classification for Assessing Software Reliability (ORCAS) to quantify probabilities of various software failure modes in a DI&C system. The method utilizes accepted industry methodologies for software quality assurance that are also verified by experimental or mathematical formulations. In essence, the approach combines a semantic failure classification model with a reliability growth model to predict (and quantify) the potential failure modes of a DI&C software system. The semantic classification model is used to address the question: How do latent defects in software contribute to different software failure root causes? The use of reliability growth models is then used to address the question: Given the connection between latent defects and software failure root causes, how can we quantify the reliability of the software? A case study was conducted on a representative I&C platform (ChibiOS) running a smart sensor acquisition software developed by Virginia Commonwealth University (VCU). The testing and evidence collection guidance in ORCAS was applied, and defects were uncovered in the software. Qualitative evidence, such as condition coverage, was used to gauge the completeness and trustworthiness of the assessment while quantitative evidence was used to determine the software failure probabilities. The reliability of the software was then estimated and compared to existing operational data of the sensor device. It is demonstrated that by using ORCAS, a semantic reasoning framework can be developed to justify software reliability (or unreliability) while still leveraging the strength of the existing methods.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A highly reliable, high performance open avionics architecture for real time Nap-of-the-Earth operations

An Army Fault Tolerant Architecture (AFTA) has been developed to meet real-time fault tolerant processing requirements of future Army applications. AFTA is the enabling technology that will allow the Army to configure existing processors and other hardware to provide high throughput and ultrahigh reliability necessary for TF/TA/NOE flight control and other advanced Army applications. A comprehensive conceptual study of AFTA has been completed that addresses a wide range of issues including requirements, architecture, hardware, software, testability, producibility, analytical models, validation and verification, common mode faults, VHDL, and a fault tolerant data bus. A Brassboard AFTA for demonstration and validation has been fabricated, and two operating systems and a flight-critical Army application have been ported to it. Detailed performance measurements have been made of fault tolerance and operating system overheads while AFTA was executing the flight application in the presence of faults.

Harper, Richard E.↗