Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Risk algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Risk-Aware Reinforcement Learning Framework for User-Centric O-RAN

The evolution of Open Radio Access Networks (O-RAN) presents an opportunity to enhance network performance by enabling dynamic orchestration of configuration and optimization parameters (COPs) through online learning methods. However, leveraging this potential requires overcoming the limitations of traditional cell-centric RAN architectures, which lack the necessary flexibility. On the other hand, despite their recent popularity, the practical deployment of online learning frameworks, such as Deep Reinforcement Learning (DRL)-based COP optimization solutions, remains limited due to their risk of deteriorating network performance during the exploration phase. In this article, we propose and analyze a novel risk-aware DRL framework for user-centric RAN (UC-RAN), which offers both the architectural flexibility and COP optimization to exploit this flexibility. We investigate and identify UC-RAN COPs that can be optimized via a soft actor-critic algorithm implementable as an O-RAN application (rApp) to jointly maximize latency satisfaction, reliability satisfaction, area spectral efficiency, and energy efficiency. We use the offline learning on UC-RAN to reliably accelerate DRL training, thus minimizing the risk of DRL deteriorating cellular network performance. Results show that our proposed solution approaches near-optimal performance in just a few hundred iterations with a decrease in risk score by a factor of ten.

6G and beyond↗

Autonomous direct freeform fabrication strategy for multi-axis additive manufacturing

Multi-axis additive manufacturing (M-AM) enables precise material deposition along both planar and curved layers, eliminating the need for support structures through a continuous material deposition approach. In contrast to conventional 2-dimensional planar layers constrained to a fixed building orientation, the deposition on freeform layers demands the specification of guided curves to determine material deposition directions which are no longer to be fixed to a build direction. There are challenges that arise when fabricating components with multiple “build” directions, necessitating the decomposition of geometries and the specification of guided curves for the resulting volumes. Furthermore, multi-axis systems introduce heightened challenges due to an increased degree of freedom in motion. Consequently, the potential risks of collision between the deposited geometry and the motion platform become a notable concern. This research proposes a freeform layering algorithm to address the challenge of seamless transitions between planar and curved layers in the process planning of M-AM. The algorithm computes 3D “printable” layers by leveraging topological information derived from the geometry to be built and integrates collision-free manufacturability considerations into the computational process. These accumulated volumes serve as a “substrate” and support volumes for subsequent deposition, allowing later layers to be built without the need for additional support material. In conclusion, the proposed method successfully devises a freeform layering approach suitable for intricate models that demand substantial support, thus enabling the fabrication of diverse geometries in a manner previously unachievable.

36 MATERIALS SCIENCE↗

Detecting Short Circuits: Post Accident Electric Vehicle Battery Safety Check

Fast and accurate detection of soft short circuits (SCs) in the battery packs of damaged electric vehicles is needed by first responders and mechanics to mitigate the potential risk from battery fires that may occur hours, days, or weeks after an accident. Here, this paper presents an SC-detection algorithm for potentially damaged lithium-ion batteries that works quickly and without a priori knowledge of the battery-pack chemistry, capacity, state of charge, or state of health. The proposed universal SC-detection algorithm is designed to be implemented on an inexpensive handheld device that can connect to and monitor the voltages of all cells in a pack. Transient filtering and linear-quadratic state observation provide estimates of normalized SC current for every cell in the pack. Cells with SC-current estimates outside a sigma-based threshold are detected. Simulations, experiments, and electric vehicle (EV) crash data are used to verify the speed, sensitivity, and accuracy of the method, demonstrating 96% accurate detection of 0.0027 C SCs in under 1 h for 5S cell groups in the lab and no false positives for crashed Volkswagen, Chevrolet, and Tesla vehicles without SCs.

25 - ENERGY STORAGE↗

Adaptive Protection and Validated Models to Enable Deployment of High Penetrations of Solar PV (PV-MOD)

The availability and validation of various PV models in commercial tools differ, with some models not yet thoroughly validated for advanced inverter functionalities and reliable performance under weak system conditions. Many existing models do not fully incorporate new inverter control functions, which can affect system stability. The increasing deployment of solar PV and other inverter-based resources (IBRs), including distributed energy resources (DERs), is influencing the reliable operation of protection schemes in distribution systems and microgrids. Emerging adaptive protection schemes (APS) offer new opportunities for protecting these systems during varying configurations and DER operating conditions, though their demonstration and validation remain limited. Adaptive protection schemes face similar challenges, as they are typically designed for specific configurations. There is a growing need for tools and methodologies to streamline the deployment of adaptive protection for safe and reliable DER integration. The project main objective was to develop and validate high-fidelity generic models of solar PV facilities for stability, protection, EMT, and QSTS analyses. This objective was achieved, and these models can now be integrated into commercial software tools, enabling utilities, vendors, and developers to study high-penetration PV systems more confidently. The project also demonstrated advanced applications of these models, including the design and deployment of adaptive protection schemes in high-penetration field applications and microgrids, supporting grid safety and reliability. Several milestones were reached by the end of the project. A sophisticated inverter test plan was developed, and inverters representative of the North American marketplace were selected. EPRI and NREL tested various inverters, conforming to IEEE standards. Improvements were made to existing generic models of IBR units, IBR plants, and aggregated feeders for various analyses. The first generic electromagnetic transient (EMT) model for a solar PV plant was developed, conforming to IEEE Std 2800™-2022 and validated against laboratory measurements of a 2.2 MVA large-scale battery energy storage system (BESS) inverter. That model was then used to produce reference responses illustrating examples of validated and verified IBR plant models that pass or fail tests for technical minimum capability and performance as specified in the IEEE standard. The developed, tested, and validated generic models can be used for transmission planning, stability assessments, expansion planning, and evaluating potential future IBR interconnection requirements. They can also support interconnection screens and conformity assessments of IBR plants, including solar PV. The project significantly contributed to the ongoing standardization and model-based representation and verification of IBR responses. The project further addressed challenges of common distribution protection schemes with increasing deployment of DER by developing, validating, and demonstrating adaptive protection schemes (APS) that can improve the reliable and safe integration of DER into distribution systems. New APS were designed using improved DER models for three common distribution systems: a radial feeder, a meshed network, and a microgrid. Modeling and hardware-in-the-loop (HIL) testing of the APS were conducted, successfully showing their effectiveness and selectivity. Proof-of-concept field demonstration was achieved for two APS, i.e., one on a radial feeder and another one in a microgrid. Field demonstration could not be achieved for the APS on a meshed network, primarily due apprehension of one utility partner and also due to limited access to the protective algorithms in the network protectors. Guidelines developed from the lessons learned in the project lay out the general process followed in the design, installation, and commissioning of APS for various distribution systems. Distribution utility partners’ apprehension about field demonstration of the new APS were addressed—with varying success—by taking a stepped risk-management approach of modeling of a wide range of sensitivities first, performing in-depth proof-of-concept testing in the laboratory including HIL next, and finally deliberately implementing and commissioning the actual protection equipment and algorithms into parts of—or in parallel operation to—the three real distribution systems. Future work should include pilot projects that further show the acceptable performance of the developed APS before these schemes be rolled out more widely. Inclusion of both utility and original equipment manufacturers (OEMs) in future projects could increase chances of successful field demonstration. Despite challenges in achieving the field demonstration goal of the project for all three APS, the research significantly contributed to the innovation of adaptive protection solutions for scalable and reliable DER integration into distribution systems. This project significantly enhances the understanding of the impact of using appropriate inverter models on distribution and transmission (T&D) systems. By addressing the limitations of existing generic models, the project introduces high-fidelity models for stability, protection, electromagnetic transient (EMT), and quasi-static time series (QSTS) analyses. These models, integrated into commercial software tools, enable utilities, vendors, and developers to confidently study high-penetration PV systems. The project also demonstrates advanced applications, including adaptive protection schemes (APS) for distribution systems and microgrids, ensuring grid safety and reliability. The technical effectiveness and economic feasibility of the methods are evident through the development and validation of sophisticated inverter test plans and the selection of representative inverters. Testing by EPRI and NREL on retail, commercial, and utility-scale inverters, conforming to IEEE standards, underscores the robustness of the models. Improvements to existing generic models for various analyses further enhance their validity and applicability. The project also identifies gaps in common distribution protection schemes and designed new APS using improved DER models, demonstrating their effectiveness through modeling and hardware-in-the-loop (HIL) testing. The project’s benefits to the public are manifold. By advancing the standardization and model-based representation of IBR response, it supports transmission planning, stability assessments, and future IBR interconnection requirements. The generic models can facilitate better communication between transmission planners and developers, supporting expected IBR plant capability and performance. Additionally, the development of APS for radial feeders, meshed networks, and microgrids supports the integration of distributed energy resources (DERs) into distribution systems, enhancing grid reliability and safety. The project’s emphasis on thorough testing and simplicity in design ensures practical and scalable solutions for DER integration.

14 SOLAR ENERGY↗

A scoping study of far-SOL main-wall protection limiters for steady-state operation of compact pilot plant tokamaks

We present a novel method for handling steady-state heat fluxes incident on the main wall of pilot plant-scale magnetic fusion devices, based on the utilization of protection limiters in the far scrape-off layer (SOL). This method helps avoid large plasma-wall gaps, without excessively compromising blanket performance. We present an optimization algorithm for determining the appropriate size and scale of these protection limiters given (1) probability distributions of SOL plasma parameters and (2) assumed risk tolerance. As part of this optimization, we have developed an analytic description of parallel heat fluxes across limiter shadows, and an objective cost function (the ‘Far-SOL Marginal Cost’) to quantify the impact that different main-wall thermal management design choices have on reactor capital cost. Applying the model to a midscale fusion pilot plant concept shows that making use of far-SOL protection limiters can reduce capital costs on the order of $500 M, relative to naively increasing the plasma-wall gap. Our analysis demonstrates that the far-SOL power decay length is the highest-leverage plasma assumption for thermal loading of the first wall, and the primary cost driver for main wall thermal management. The relative cost efficiency of protection limiters increases as assumptions on the far-SOL heat flux become more pessimistic. The concepts described in this paper motivate the further development of far-SOL protection limiters as part of larger efforts to design economical core-edge-wall compatible solutions for a fusion pilot plant.

Design under uncertainty↗

Provable bounds for noise-free expectation values computed from noisy samples

Quantum computing has emerged as a powerful computational paradigm capable of solving problems beyond the reach of classical computers. However, today’s quantum computers are noisy, posing challenges to obtaining accurate results. Here, we explore the impact of noise on quantum computing, focusing on the challenges in sampling bit strings from noisy quantum computers and the implications for optimization and machine learning. We formally quantify the sampling overhead to extract good samples from noisy quantum computers and relate it to the layer fidelity, a metric to determine the performance of noisy quantum processors. Further, we show how this allows us to use the conditional value at risk of noisy samples to determine provable bounds on noise-free expectation values. We discuss how to leverage these bounds for different algorithms and demonstrate our findings through experiments on real quantum computers involving up to 127 qubits. The results show strong alignment with theoretical predictions.

97 MATHEMATICS AND COMPUTING↗

A Large-Scale Analysis to Optimize the Control and V2V Communication Protocols for CDA Agreement-Seeking Cooperation

Cooperative driving automation (CDA) Class C, agreement-seeking cooperation, is an innovative and practical solution that can promote cooperation among general passenger vehicles on the road. However, more comprehensive studies are needed before establishing the standard protocols of agreement-seeking cooperation, such as communication frequency and the duration of cooperation. Here, this article presents an initiative study on the impacts of communication capabilities on agreement-seeking cooperation. Through a large-scale analysis by regulating vehicle-to-vehicle (V2V) communication metrics, this work suggests desirable system parameters that can maximize the benefits of cooperation and ensure reliable operability while avoiding exhaustive communication loads. As the first step, an example agreement-seeking cooperation system is created for a car-following scenario, including decision-making and control algorithms for autonomous vehicles. Then, software-in-the-loop tests explore the performance of the developed system as it encounters various communication risks, such as latency and message packet drops. The system performance metrics are evaluated from various angles, including the time consumed for the agreement-seeking process, cooperation ratio, and the ratio of faulty cooperation. Energy saving from the cooperation is assessed by using simulation software that can run multiple high-fidelity vehicle models simultaneously. Based on the analyses, this article suggests the V2V communication requirements for the reliable operation of CDA agreement-seeking, which can be referred to when developing the standard protocols of agreement-seeking cooperation.

42 ENGINEERING↗

Learning linear optical circuits with coherent states

We analyze the energy and training data requirements for supervised learning of an M-mode linear optical circuit by minimizing an empirical risk defined solely from the action of the circuit on coherent states. When the linear optical circuit acts non-trivially only on k < M unknown modes (i.e. a linear optical k-junta), we provide an energy-efficient, adaptive algorithm that identifies the junta set and learns the circuit. We compare two schemes for allocating a total energy, E, to the learning algorithm. In the first scheme, each of the T random training coherent states has energy E/T. In the second scheme, a single random MT-mode coherent state with energy E is partitioned into T training coherent states. The latter scheme exhibits a polynomial advantage in training data size sufficient for convergence of the empirical risk to the full risk due to concentration of measure on the $(2MT-1)$-sphere. Specifically, generalization bounds for both schemes are proven, which indicate that for ε-approximation of the full risk by the empirical risk with high probability, $O(E^{2/3}M^{2/3}/\epsilon^{2/3})$ training states are sufficient for the first scheme and $O(E^{1/3}M^{1/3}/\epsilon^{2/3})$ training states are sufficient for the second scheme.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Fracture Network Prediction Using Physics-based Machine Learning Algorithms

In recent years, systematic CO2 injection into geological reservoirs across the U.S. has gained traction as a strategy to mitigate greenhouse gas emissions. This approach necessitates precise monitoring to ensure secure containment, minimize risks, and optimize storage management. Our study leverages machine learning (ML) techniques to advance the understanding of CO2 injection processes, focusing on the Illinois Basin. Over a three-year injection period, we analyzed microseismic data, identifying 19 temporal intervals with significant bottom-hole pressure changes. By partitioning microseismic events into these intervals and estimating b-values, we revealed over 100 clusters of events related to fracture initiation or reactivation. Advanced spatial analysis highlighted horizontally-oriented fractures along the NNW-SSE axis. This quantification of fracture networks informs dynamic injection scheduling, work-over strategies, and risk assessments, enhancing carbon capture, utilization, and storage (CCUS) operations. Additionally, our methodology offers valuable insights for oil and gas operations and geothermal development, supporting fracture-based monitoring and risk mitigation.

Kumar, Abhash↗

Simulation Center for Runaway Electron Avoidance and Mitigation (SCREAM SciDAC) (Technical Final Report)

Runaway electrons can severely damage the plasma facing components on ITER during a major disruption and pose a major risk for tokamak fusion. It has been recognized that an adequate disruption mitigation system (DMS) is essential for the safe operation of ITER. The United States is responsible for the design and implementation of the disruption mitigation system on ITER, and in July 2016 the Simulation Center for Runaway Electron Avoidance and Mitigation (SCREAM) was launched by DOE, in a joint Fusion Energy Sciences (FES) and Advanced Scientific Computing Research (ASCR) collaboration. SCREAM was a comprehensive theory and simulation SciDAC center that provided physics guidance in the avoidance and mitigation of runaway electrons, and in tandem with domestic and international experiments, helped establish the qualitative and quantitative bases for safe operational scenarios and viable mitigation techniques. The SCREAM center assembled a national team of experts in runaway electron physics, tokamak disruptions, magnetohydrodynamic (MHD) simulation, and advanced algorithms and computing. The team combined advanced simulation and analysis capability facilitated by direct participation of ASCR SciDAC institutes with theoretical models and code development by FES scientists to focus on the runaway risk for ITER and tokamaks in general. The research scope was focussed on integrated simulations of kinetic runaway electrons, including MHD and fluid models of impurity transport, within a research plan guided by theory. The specific research tasks were (1) establish the fundamental physics of runaway generation, saturation, and dynamical evolution in a tokamak; (2) examine the critical path toward runaway avoidance; and (3) investigate the viability and effectiveness of the leading candidate schemes for runaway mitigation. In all three areas, members of the team carried out scoping studies that established the readiness for rapid and critical advances, especially in the deployment and further development of large-to extreme-scale simulation tools. Our multi-pronged computational approach included (1) relativistic Fokker-Planck solvers with discretization in phase space, (2) self-consistent particle-in-cell techniques, (3) particle-based Monte-Carlo, and (4) MHD-particle hybrid simulations. Cross-check between these different methods provided an additional means for verification and further bolstered the fidelity of our physics prediction. Validation against experimental results brings confidence to the predictive capability for ITER and frequently leads to new ideas for understanding and mitigating the thermal quench driven runaway electron phenomenon.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Noise-aware optimization in nominally identical manufacturing and measuring systems for high-throughput parallel workflows

Device-to-device variability in experimental noise critically impacts reproducibility, especially in automated, high-throughput systems like additive manufacturing farms. While manageable in small labs, such variability can escalate into serious risks at larger scales, such as architectural 3D printing, where noise may cause structural or economic failures. This contribution presents a noise-aware decision-making algorithm that quantifies and models device-specific noise profiles to manage variability adaptively. It uses distributional analysis and pairwise divergence metrics with clustering to choose between single-device and robust multi-device Bayesian optimization strategies. Unlike conventional methods that assume homogeneous devices or enforce generic robustness, the proposed framework explicitly determines whether shared optimization across devices is appropriate based on the degree of inter-device noise heterogeneity. This enables improved performance, reproducibility, and efficiency. An experimental case study involving three nominally identical 3D printers (same brand, model, and close serial numbers) demonstrates reduced redundancy, lower resource usage, and improved reliability, along with improved convergence stability and solution quality through the selection of the appropriate optimization strategy based on the degree of inter-device noise heterogeneity. Overall, this framework establishes a general approach for precision- and resource-aware optimization in scalable, automated experimental platforms, demonstrated here on a representative multi-device 3D printing case study.

Schenk, Christina↗

Risk-informed Graded Approach for Reliability and Performance Assessment of Machine Learning and Artificial Intelligence for Advanced Condition Monitoring Techniques

With the shift away from time-based maintenance and toward condition-based maintenance, and to reduce overall maintenance costs, there has been an upsurge in the usage and development of advanced condition monitoring (ACM) techniques for real-time monitoring of nuclear power plant (NPP) components. ACM is particularly useful in the development of digital twins, which are designed to predict the failure or degradation of plant components. Successful implementation of ACM requires an assessment to inform the development of a risk-informed approach to evaluate the use of ACM to meet Nuclear Regulatory Committee (NRC) regulations for in-service testing (IST) programs. This includes the monitoring and diagnostics of reactor components and systems in current, new, and advanced reactors. A key component in ACM is the usage of machine learning (ML) and artificial intelligence (AI) algorithms that can employ real-time data from instrumentation and sensors to detect and predict reactor component degradations. Such predictive capabilities enable early detection of component degradation so as to help plant personnel plan and execute necessary maintenance. For successful implementation of ML/AI in ACM such that regulatory requirements are met, a risk-informed graded approach is needed to assess the reliability and performance of ML/AI for ACM. The American Society for Mechanical Engineers (ASME) developed their Operations and Maintenance (O&M) Code to provide guidance on safe, reliable O&M of NPPs. The IST section of the O&M Code specifically establishes requirements for IST and examination to gauge operational readiness of components in water-cooled NPPs. This paper presents a state-of-the-art review of how reliability and risk assessment can be integrated with ACM to assess component performance by non-nuclear industries. This is followed by different methodologies and approaches for conducting performance and reliability assessments so as to meet IST requirements for NPP components.

97 - MATHEMATICS AND COMPUTING↗

Risk-informed Graded Approach for Reliability and Performance Assessment for Advanced Condition Monitoring Techniques

With the shift away from time-based maintenance and toward condition-based maintenance, and to reduce overall maintenance costs, there has been an upsurge in the usage and development of advanced condition monitoring (ACM) techniques for real-time monitoring of nuclear power plant (NPP) components. ACM is particularly useful in the development of digital twins, which are designed to predict the failure or degradation of plant components. Successful implementation of ACM requires an assessment to inform the development of a risk-informed approach to evaluate the use of ACM to meet Nuclear Regulatory Committee (NRC) regulations for in-service testing (IST) programs. This includes the monitoring and diagnostics of reactor components and systems in current, new, and advanced reactors. A key component in ACM is the usage of machine learning (ML) and artificial intelligence (AI) algorithms that can employ real-time data from instrumentation and sensors to detect and predict reactor component degradations. Such predictive capabilities enable early detection of component degradation so as to help plant personnel plan and execute necessary maintenance. For successful implementation of ML/AI in ACM such that regulatory requirements are met, a risk-informed graded approach is needed to assess the reliability and performance of ML/AI for ACM. The American Society for Mechanical Engineers (ASME) developed their Operations and Maintenance (O&M) Code to provide guidance on safe, reliable O&M of NPPs. The IST section of the O&M Code specifically establishes requirements for IST and examination to gauge operational readiness of components in water-cooled NPPs. This paper presents a state-of-the-art review of how reliability and risk assessment can be integrated with ACM to assess component performance by non-nuclear industries. This is followed by different methodologies and approaches for conducting performance and reliability assessments so as to meet IST requirements for NPP components.

99 - GENERAL AND MISCELLANEOUS↗

Multiscale drivers of extreme southern California flooding: ENSO, MJO, North Pacific jet, and atmospheric rivers

Extreme rainfall and flooding, driven by a powerful atmospheric river (AR) and a persistent Madden-Julian Oscillation (MJO), hit Southern California in February 2024 during the 2023–2024 El Niño, affecting over 10 million people. ARs are key contributors to extreme rainfall and flooding along the U.S. West Coast. Although the AR-MJO link has been documented, its spatio-temporal variability remains a major forecasting and risk-management challenge. Combining precipitation, stream gauge and demographic data, we quantify the physical drivers and population exposure to this extreme event. Leveraging a Lagrangian MJO precipitation tracking algorithm, we unravel the multiscale interactions responsible for the AR’s development. El Niño favored a large, long-lived MJO that interacted with the North Pacific Jet (NPJ) over more than three weeks. The MJO convective outflow modulated the NPJ by inducing negative potential vorticity advection along the tropopause. The ensuing NPJ extension and acceleration induced explosive cyclogenesis, whose AR-driven moisture transport resulted in extreme rainfall.

Atmospheric dynamics↗

A Sequential Model Predictive and Deep Reinforcement Learning-Based Controller for Distribution System Outage Mitigation under Hurricane Events

This paper proposes a proactive outage mitigation framework for power distribution networks to withstand hurricane-induced disruptions. It leverages Model Predictive Control (MPC) to identify safe lines for proactive switching during hurricanes, minimizing the risk of cascading failures and voltage violations. The switching strategies optimized by MPC are sequentially integrated with a Deep Reinforcement Learning agent using the Advantage Actor-Critic algorithm, enabling dynamic line switching to maximize connected buses and minimize voltage violations in real time. Using a probabilistic hurricane model, the framework predicts line failures and adapts to varying conditions to enhance grid resilience. Simulations on the IEEE 123-bus system demonstrate its effectiveness in maintaining high connectivity and minimizing disruptions. Real-time testing with an RTDS confirms the practicality and reliability of the proposed approach.

Selim, Alaa [University of Connecticut]↗

Root Cause Correlation Analysis of Software Failures via Orthogonal Defect Classification and Natural Language Processing

Systems theoretic process analysis (STPA) is becoming an increasingly popular technique to assess how complex digital software systems can fail. Rather than defining failures by their observable failure events, which may be sparse especially for safety rated nuclear digital instrumentation and control systems (DI&C), failures are defined as postulated unsafe actions under specific contextual conditions. This permits a top-down analysis of system hazards and identifies whether imposed constraints and requirements can sufficiently address undesirable hazards. However, STPA is a qualitative approach at identifying inadequacies in the development process and cannot currently be used to quantify unsafe action likelihoods for probabilistic risk assessment. Therefore, in this work, we examine the root causes of software failure and explore whether a consistent correlation can be linked to specific unsafe action classes. We implement Lbl2Vec, an unsupervised document classification and retrieval algorithm, on a database of 4,096 software defect reports acquired from various open-source software systems. By analyzing sentence structure, embedded labels, and word vectors, we show that certain defect types positively correlate to specific unsafe action classes over others. The correlations developed can be used to estimate the failure probability of safety intended DI&C systems which provides a licensing basis for nuclear plant modernization efforts.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Root Cause Correlation Analysis of Software Failures via Orthogonal Defect Classification and Natural Language Processing

Systems theoretic process analysis (STPA) is becoming an increasingly popular technique to assess how complex digital software systems can fail. Rather than defining failures by their observable failure events, which may be sparse especially for safety rated nuclear digital instrumentation and control systems (DI&C), failures are defined as postulated unsafe actions under specific contextual conditions. This permits a top-down analysis of system hazards and identifies whether imposed constraints and requirements can sufficiently address undesirable hazards. However, STPA is a qualitative approach at identifying inadequacies in the development process and cannot currently be used to quantify unsafe action likelihoods for probabilistic risk assessment. Therefore, in this work, we examine the root causes of software failure and explore whether a consistent correlation can be linked to specific unsafe action classes. We implement Lbl2Vec, an unsupervised document classification and retrieval algorithm, on a database of 4,096 software defect reports acquired from various open-source software systems. By analyzing sentence structure, embedded labels, and word vectors, we show that certain defect types positively correlate to specific unsafe action classes over others. The correlations developed can be used to estimate the failure probability of safety intended DI&C systems which provides a licensing basis for nuclear plant modernization efforts.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE↗