Ensuring Interpretability of ML Technologies for Predictive Maintenance on Operating Nuclear Plants
Poster to explain app that we are presenting at the conference
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Poster to explain app that we are presenting at the conference
The domestic nuclear power plant fleet has relied on labor-intensive and time-consuming preventive maintenance programs, thus driving up operation and maintenance costs to achieve high-capacity factors. Artificial intelligence and machine learning can help simplify complex problems, such as diagnosing equipment degradation, to enable more effective decision-making. Benefits will be felt not only within existing analog and digital instrumentation and control, but also work processes, the integration of people with technology, and most importantly, the business case. Together, these hold promise to make nuclear power more efficient and reduce costs associated with operation and maintenance. While the artificial intelligence and machine learning technologies hold significant promise in the nuclear industry, there are challenges or barriers to their adoption. This report outlines the those different machine learning adoption barriers (categorized as historical, technical, economic, regulatory, and user) that the industry must overcome to realize the full benefits of artificial intelligence and machine learning capabilities for long-term economic sustainability. This report also provides solutions for some of these barriers by focusing on improving the explainability of machine learning to encourage trust from the end-user. Trust and explainability are essential to machine learning adoption. This report focuses on research-developed solutions to some of these barriers while analyzing a non-safety-related system, namely the circulating water system. This system frequently experiences waterbox fouling which our models preemptively diagnoses then explains to the operator how those conclusions were reached. This report presents and discusses the inherent trade-off between machine learning performance (in terms of accuracy) and explainability, where highly accurate machine learning methods (such as deep-learning) are the least explainable, and the most explainable methods (such as decision trees) are the least accurate. In addition, explainability of artificial intelligence techniques in terms of transparency and post-hoc metrics are discussed. This report outlines the importance of data novelty and value of new information in evaluating both the explainability and trustworthiness. Novelty detection helps to establish consistency or inconsistency of the new data with respect to the training data. On the other hand, value of information could be a part of the user-centric visualization recommendation system that request additional information to be collected, thereby strengthening the machine learning outcomes. During this project, a copyrighted user-centric visualization that aligns with a human-in-the-loop approach was developed. The user-centric visualization presents different levels of information and can be tailored as per user credentials to gain user confidence. One of the salient features of the user-centric visualization is it presents machine learning methods with explainability metrics. A simplified version of the user-centric visualization was presented to 32 users with varying levels of machine learning expertise. Feedback was solicited to test the hypothesis that the app contained sufficient explainability and that the users would trust the algorithm. Overall, the app was positively received, and the hypothesis was supported. This report discusses the trust-but-verify framework – a potential approach to build user trust artificial intelligence. The framework discusses trust from the human level to artificial intelligence level. The fundamental premise of the trust but verify framework is derived from an observation of nuclear safety culture (i.e., nuclear power plant personnel do not rely on a singular source of data to make a decision). This also ties back to the user-centric visualization that presents different levels of information to achieve both explainability and trustworthiness of artificial intelligence. Even so, the adoption of artificial intelligence and machine learning in the nuclear industry faces additional barriers, namely regulatory and stakeholder readiness. To overcome these challenges, new solutions must gain regulatory approval and cater to stakeholder needs. The Nuclear Regulatory Committee has a 5-year strategic plan which prepares them for reviewing artificial intelligence technologies in licensee submissions. Early and frequent engagement with the regulator is encouraged. Additionally, artificial intelligence solutions should incorporate human-in-the-loop considerations and offer explainability. Stakeholders must prepare by hiring or training staff to adapt to advancing technology in everyday plant tasks.
The domestic nuclear power plant (NPP) fleet has historically relied on labor-intensive and time-consuming predictive maintenance (PdM) programs, thus driving up operation and maintenance (O&M) costs to achieve high-capacity factors. Artificial intelligence (AI) and machine-learning (ML) can help simplify complex problems such as diagnosing equipment degradation to enable more effective decision-making efforts. The benefits of AI will be felt through more efficient plant O&M, improved work processes, and better integration of people and technology. Together, these benefits hold the promise to make nuclear power more sustainable by reducing O&M costs while improving employee engagement. While AI and ML technologies hold significant promise for the nuclear industry, there are challenges or barriers to their adoption. Explainability and trustworthiness of AI are two salient challenges that need to be addressed for wider deployment of these technologies in NPPs. This research focuses specifically on addressing the explainability and trustworthiness of AI technologies to advance the human, technical, and organization (HTO) readiness levels in adopting a risk-informed PdM strategy at commercial NPPs. In addition, this approach can be adapted to enhance the acceptability of AI in other nuclear applications with a few application-specific modifications. The technical approach ensuring wider adoption of AI technologies was developed by Idaho National Laboratory (INL)—in collaboration with Public Service Enterprise Group (PSEG), Nuclear, LLC—by utilizing the circulating water system (CWS) at two PSEG-owned plant sites for demonstration. Focused user studies were performed in collaboration with subject matter experts (SMEs) from PSEG and other nuclear domains to enhance human and organization readiness by building trust in AI-informed technologies. VIsualization for PrEdictive maintenance Recommendation (VIPER)—a Battelle Energy Alliance, LLC, copyrighted software—was developed and expanded to provide a user-centric visualization by incorporating inputs from the collaborating utility, human factors engineering guidelines, and data analysts. The VIPER software enables users, who may be unfamiliar with ML in general, to be interactively engaged by asking technical questions about PdM, work orders, diagnosis results and their confidence levels, the kind of data being used, and the types of ML algorithms employed. This interactive engagement enhances explainability and builds trust. One of the enabling accomplishments was the integration of large language models (LLMs), both text-based and vision-based, in the VIPER software.
Acoustic emission sensors are vital in the nuclear industry for real-time structural health monitoring and early detection of material degradation. By capturing high-frequency stress waves emitted from defects like cracks, corrosion, or fatigue, acoustic emission sensors enable non-invasive monitoring of critical components such as reactor vessels, piping, and containment structures. This technology supports predictive maintenance, enhances safety, and ensures regulatory compliance by providing early warnings of potential failures. It is also instrumental in research, particularly in material testing reactors, where it is used to monitor the behavior of fuels and materials under irradiation, by allowing the detection of cracking or other acoustic signals in real time. This enables the evaluation of performance and accident behavior of advanced fuel concepts.
The fact that light-water reactor operation and maintenance costs are prohibitively expensive and contribute to the premature decommissioning of nuclear power plants is partly due to how the equipment is monitored. In recent years, cloud computing has emerged as a dominant technology, as its low cost, computing and storage adaptability, and ability to host applications across numerous virtual infrastructures potentially make it a cost-effective alternative to onsite storage and diagnostics. In this paper, a technological assessment is carried out on a provisional cloud deployment architecture for a nuclear power plant predictive monitoring system. This cloud-based monitoring system would enable maintenance and diagnostic analysts and other authorized plant users to remotely monitor equipment functionality, thus enabling early fault detection and effective predictive maintenance practices. To provide data processing and storage, sensor device networking, and database management, the Microsoft Azure cloud platform is utilized as part of the proposed cloud architecture; however, this analysis could be extended to other cloud computing service providers as well. The focus of this paper is on application of cloud resources for enabling predictive maintenance, identification of technological hurdles associated with moving to a cloud-computing-based architecture, and potential benefits from moving to a centralized cloud system.
The primary objective of the research presented in this report is to develop scalable technologies that are deployable across plant assets and across the nuclear fleet to achieve risk-informed predictive maintenance (PdM) strategies at commercial nuclear power plants (NPPs). Over the years, the nuclear fleet has relied on labor-intensive and time-consuming preventive maintenance (PM) programs, driving up operation and maintenance (O&M) costs to achieve high capacity factors. A well-constructed risk-informed PdM approach for an identified plant asset has been developed in this research, taking advantage of advancements in data analytics, machine learning (ML), artificial intelligence (AI), physics-informed modeling, and visualization. These technologies would allow commercial NPPs to reliably transition from current labor-intensive PM programs to a technology driven PdM program, eliminating unnecessary O&M costs. The work presented in the report is being developed as part of a collaborative research effort between Idaho National Laboratory and Public Service Enterprise Group Nuclear, LLC. This report (1) reflects the results of work by LWRS Program researchers with PSEG, Nuclear LLC-owned Salem and Hope Creek Nuclear Power Plants; (2) presents utilization of circulating water system (CWS) heterogeneous data and fault modes from both the Salem and Hope Creek nuclear power plant sites to develop salient fault signatures associated with each fault mode; (3) describes the integration of component-level predictive models into a robust system-level model enabled by the federated-transfer learning; (4) describes the development of physics-informed model of circulating water pump and motor; (5) develops a scalable risk and economic model; and (6) outlines the development of a user-centric visualization application. The outcomes presented in this report lays the foundation and provides a much-needed technical basis to focus on explainability and trustworthiness of ML and AI-based technologies, as part of future research.
This report examines the transformative impact of Artificial Intelligence (AI) and Machine Learning (ML) on operations research, private industry, and government sectors, highlighting their applications in automating processes, enhancing decision-making, and optimizing complex systems. AI/ML technologies have revolutionized industries through predictive maintenance, supply chain optimization, and autonomous systems, while also advancing public safety and defense operations. However, challenges such as data integrity, model transparency, and the need for human oversight persist, particularly in high-consequence environments. The report emphasizes the critical role of explainable AI (XAI) and human-computer interaction models like Human-in-the-Loop (HITL) and Human-on-the-Loop (HOTL) in fostering trust and accountability. Balancing automation with ethical responsibility and transparency is essential for the continued successful integration of AI/ML into operational and strategic decision-making frameworks.
Multi-Kernel Support Vector Machine (MK-SVM) is a machine learning classification algorithm that can assist in the development of predictive maintenance strategies for nuclear power plant systems. Predictive maintenance can alleviate maintenance costs and enhance reliability of plant systems. In this work, MK-SVM is utilized for determining the health of the circulating water system (CWS) in a nuclear power plant.
A primary challenge in hydropower industry is the ability to maintain cost-competitiveness, reliability, and security of hydropower assets through evolving power system contexts and aging of the fleet. Maintaining cost-effective and reliable operations under these conditions is expected to require new modernization and maintenance paradigms for changing contexts. Changes in existing practices for O&M will require an understanding of the current state and health of hydropower assets, and the impact of changing paradigms on asset health and reliability. The Hydropower Fleet Intelligence project is developing and evaluating standardized methodologies and analysis tools for data-driven asset reliability and management technologies for hydropower, leading to eventual predictive maintenance planning, repair/replacement decision making, and asset-reliability and cost-optimized operations. A key question is the feasibility of using existing data sets at hydropower facilities to perform assessments of asset reliability. This document uses data from hydropower facilities to assess the potential for using available analytics methods for asset reliability estimates. In addition to reliability assessments, the feasibility of using existing analytics techniques for several other potential applications is discussed. Finally, a case study that a data-driven model is trained to learn nominal operations via vibration data from an asset of a certain plant, and then utilized to identify anomalies on a similar asset from a different plant, highlighting the generic use of proposed Prognostics and Health Management (PHM) approaches.
The current light-water reactor fleet uses time-based maintenance strategies to achieve high-capacity factors. But to make nuclear more competitive in the energy market, these reactors could utilize emerging artificial intelligence (AI) and cloud computing technologies to achieve a cost-effective, predictive-maintenance strategy. This paper presents discussion and results on the application of cloud computing in the nuclear industry. The technical viability of cloud computing was analyzed using data from a boiling-water reactor’s safety relief valve. The models were hosted on three different systems: a local personal computer, Idaho National Laboratory’s high-performance computer system, and Microsoft Azure. The data were loaded and processed, and two types of models were trained in an A/B fashion. Based on the speed at which these actions were completed, it was determined that cloud computing affords adequate computing resources. Additionally, the computing power can scale with the demanded load. To enable cloud computing in the existing fleet, additional sensors, networks, and other requirements must be implemented to ensure a smooth transition from current maintenance strategies. However, the benefit is that the plants no longer need to manage their own servers, software, cybersecurity, and information technology support staff for in-house data analytics purpose. Many of these features can be offloaded to the cloud provider for a potential cost savings. Demonstrating how AI can improve the maintenance and operation of non-safety-related systems seems the likely path forward for implementing AI and cloud computing resources inside nuclear power plants.
The need for an accessible iterative approach for evaluating prospective artificial intelligence (AI)/ML based technologies in the nuclear industry is needed, given the nature of algorithms and rapid advancements. This paper explores existing heuristic design principles for user-centered design and evaluates them based on their relevancy and usefulness for evaluating AI/ ML based technologies. Researchers at the Idaho National Laboratory (INL) have developed a machine learning software application called VIsualization for PrEdictive maintenance Recommendation (VIPER), which is used to help users understand and engage with the tool to learn more about work orders, data used, predictive maintenance, and machine learning (ML) algorithms. Early user research studies used to access VIPER’s technology readiness level have occurred; however, there is room for further improvement of the software through heuristic evaluations along with other methods and user testing. This work describes the applicability of heuristic evaluation methods and cognitive walkthroughs to help ensure human readiness for prospective AI/ ML based applications, using VIPER as a candidate use case. This work supports industry in ensuring that prospective AI/ML based technologies are usable and useful for plant personnel at nuclear power plants, ultimately leading to their safe, reliable, and efficient use.
The need for an accessible iterative approach for evaluating prospective artificial intelligence (AI)/ML based technologies in the nuclear industry is needed, given the nature of algorithms and rapid advancements. This paper explores existing heuristic design principles for user-centered design and evaluates them based on their relevancy and usefulness for evaluating AI/ ML based technologies. Researchers at the Idaho National Laboratory (INL) have developed a machine learning software application called VIsualization for PrEdictive maintenance Recommendation (VIPER), which is used to help users understand and engage with the tool to learn more about work orders, data used, predictive maintenance, and machine learning (ML) algorithms. Early user research studies used to access VIPER?s technology readiness level have occurred; however, there is room for further improvement of the software through heuristic evaluations along with other methods and user testing. This work describes the applicability of heuristic evaluation methods and cognitive walkthroughs to help ensure human readiness for prospective AI/ ML based applications, using VIPER as a candidate use case. This work supports industry in ensuring that prospective AI/ML based technologies are usable and useful for plant personnel at nuclear power plants, ultimately leading to their safe, reliable, and efficient use. PowerPoint for conference that was reviewed in PRS and LRS PRS/CON-25-05379 and INL/CON-25-82946
This report describes the research results to date to identify regulatory implications of DT technologies and their uses. Specifically, this report reviews current regulatory guidance relevant to the application of predictive maintenance DTs, artificial intelligence (AI), and automation. The focus of this review included determination of constraints on the application of DT technology, identification of any regulatory gaps or uncertainties, and clarification of anticipated technical basis information likely to be important for regulatory acceptance of these technologies. The research included review of topical reports and other submissions to the US Nuclear Regulatory Commission (NRC) on technologies relevant to using DTs for predictive maintenance, NRC safety evaluation (SE) reports, and other relevant literature to identify specific regulatory concerns.
This research introduces a comprehensive framework for creating and deploying a digital twin platform for continuous monitoring and predictive maintenance within industrial settings. Through utilizing advanced technologies, including Unreal Engine 5, Unity 3D, the Message Queue Telemetry Transport protocol, Random Forest machine learning algorithms, and Large Language Models (LLMs), we establish a platform that digitally reproduces physical equipment and translates digital controls into real-world actions. This facilitates preventive maintenance approaches and improves operational effectiveness. The digital twin platform gathers sensor data from operational equipment, analyzes it using machine learning, and delivers practical insights to prevent potential malfunctions and enhance equipment performance. Furthermore, the incorporation of a web portal enables efficient monitoring and access to historical data, educational materials, and equipment status information. Preliminary findings indicate that digital twins can transform industrial equipment management and maintenance methodologies.
Nuclear plant sites collect and store large volumes of data collected from various equipment and systems. These datasets typically include plant process parameters, maintenance records, technical logs, online monitoring data, and equipment failure data. The collection of such data affords an opportunity to leverage data-driven machine learning and artificial intelligence technologies to provide diagnostic and prognostic capabilities within the nuclear power industry to reduce operating and maintenance costs. In this way, nuclear energy can become more economically competitive with other energy sources, and premature closures can be avoided. From a maintenance standpoint, savings can be achieved by leveraging machine learning and artificial intelligence technologies to develop data-driven algorithms to better diagnose and predict potential faults within the system. Improved model accuracy can lead to reductions in unnecessary maintenance and more efficient planning of future maintenance, thus lowering the costs associated with parts, labor, and unnecessary planned, forced, or extended outages. From an operations perspective, cost savings can be generated by shifting from route-based monitoring to wireless technologies for online monitoring, and by transitioning from onsite- to cloud-based computing and storage services. Wireless monitoring would reduce the operator manhours required for taking routine measurements, while cloud computing services would generate cost savings by reducing the amount of hardware needing to be purchased and maintained—all while scaling to both computational and storage demands. This report summarizes this project’s effort to shift from costly, labor-intensive preventative maintenance to cheaper predictive maintenance.
The research involves developing scalable technologies for risk-informed predictive analytics to achieve condition-based monitoring and maintenance strategies to reduce overall maintenance costs. The research utilizes data (real-time data, periodic data, and institutional knowledge) related to a particular plant asset from a specific nuclear plant site to develop technologies to scale risk-informed predictive analytic algorithms across different plant assets at the plant site and across the nuclear fleet. The developed algorithms and codes are used to optimize the maintenance strategy and estimate/forecast generation costs based on the state of health of the plant asset. Developed codes specifically include 1. Parameter estimation using plant operation data 2. Federated and Transfer learning model 3. Feature group based Multi-kernel SVM 4. Three state markov model
The current light water reactor fleet uses time-based or failure-based maintenance strategies to achieve high-capacity factors. But to make nuclear more competitive in the energy market, these reactors could utilize emerging technologies in terms of artificial intelligence (AI) and cloud computing to enable a cost-effective, predictive maintenance strategy. This report examines the feasibility of cloud computing for the nuclear industry’s needs in terms of the cloud’s computing capabilities, feasibility, and regulatory concerns. The technical viability of cloud computing was analyzed using one year worth of data from a boiling water reactor’s safety relief valve. Models were hosted on a local desktop, Idaho National Laboratory’s high-performance computer, and Microsoft Azure. Data was loaded, processed, and two types of models were trained in an A/B fashion. Based on the speed at which these actions were completed, it was used to determined that cloud computing has adequate computing resources. Additionally, the computing power can scale with the demanded load. To enable cloud computing in the existing fleet, additional sensors, networks, and other requirements must be implemented to ensure a smooth transition from current maintenance strategies. However, there is a benefit as the plant no longer needs manage their own servers, software, cybersecurity, and IT support staff. Many of these features can be offloaded on to the cloud provider. A comprehensive analysis was completed that showed the current annual cost of operating is more expensive than using cloud computing resources. Lastly, the regulatory framework does not explicitly address AI or autonomous control. Currently, the NRC and other regulatory bodies are evaluating providing guidance to address gaps rather than new regulations to address the use of AI and ML. But since many of the AI applications are focused on non-safety related applications, such as balance-of-plant components, they will likely have little or no regulatory restrictions or necessary approvals. Demonstrating how AI can improve maintenance and operation of these non-safety related systems seems like the likely path forward for implementing AI and cloud computing resources inside nuclear power plants (NPPs).