Machine learning selection of differential and integral experiments to optimally reduce nuclear data uncertainties [Slides]
Abstract not provided.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Abstract not provided.
The goal of the EUCLID (Experiments Underpinned by Computational Learning for Improvements in Nuclear Data) project was to reduce compensating errors by utilizing machine learning to both help determine which reactions contain compensating errors as well as optimizing an experiment which can be used to maximally reduce these errors. Compensating errors can adversely impact the predictive power of application simulations, and therefore it’s useful to further constrain nuclear data and reduce these errors. The EUCLID project included building two configurations at the National Criticality Experiments Research Center (NCERC). These two configurations had very different geometries (one was cube-like and one was slab-like). Previous works focus on selection of the target experiment(s), radiation transport capabilities developed in the project, the experiment optimization, and the performance of the experiments. This work will focus only on the safety aspects of performing this experiment, which utilized over 100 kg of weapons-grade plutonium.
Industries around the world are becoming more and more data driven. The nuclear field is no exception with several different applications being proposed. One popular area of research is the use of machine learning in transient detection. This paper seeks to build upon a previous study which made use of the AutoML package TPOT to train traditional machine learning models to classify transient events occurring with a reactor. Synthetic data was once again collected using a GPWR reactor simulator. Data on 12 different events was collected using 15 different initial conditions. Here, a dataset consisting of over 100,000 data points was compiled and used to train 7 different machine learning models using a pre-defined TPOT dictionary with 12 different preprocessing techniques. Three of the trained models were able to produce validation results in the 90s with the expanded dataset. Once the models were trained, it was possible to look into where during the simulation, misclassifications occurred. Using these three models, analysis was done to determine if TPOT could be used to train models that were effective if important features were missing. The results from this were positive with the newly trained models scoring close to the original models. Finally, to conclude this study, the three high performing models were retrained using different random states to see if there was any major variation when different states were used.
Nuclear materials are often demanded to function for extended time in extreme environments, including high radiation fluxes with associated transmutations, high temperature and temperature gradients, mechanical stresses, and corrosive coolants. They also have a wide range of microstructural and chemical makeups, resulting in multifaceted and often out-of-equilibrium interactions. Machine learning (ML) is increasingly being used to tackle these complex time-dependent interactions and aid researchers in developing models and making predictions, sometimes with better accuracy than traditional modeling that focuses on one or two parameters at a time. Conventional practices of acquiring new experimental data in nuclear materials research are often slow and expensive, limiting the opportunity for data-centric ML, but new methods are changing that paradigm. Here we review high-throughput computational and experimental data approaches, especially robotic experimentation and active learning that is based on Gaussian process and Bayesian optimization. We show ML examples in structural materials (e.g., reactor pressure vessel (RPV) alloys and radiation detecting scintillating materials) and highlight new techniques of high-throughput sample preparation and characterizations, and automated radiation/environmental exposures and real-time online diagnostics. Herein, this review suggests that ML models of material constitutive relations in plasticity, damage, and even electronic and optical responses to radiation are likely to become powerful tools as they develop. Finally, we speculate on how the recent trends of using natural language processing (NLP) to aid the collection and analysis of literature data, interpretable artificial intelligence (AI), and the use of streamlined scripting, database, workflow management, and cloud computing platforms that will soon make the utilization of ML techniques as commonplace as the spreadsheet curve-fitting practices of today.
The Ohio State University and Idaho National Laboratory organized the 4 th Big Data for Nuclear Power Plants Workshop in November, 2023 in Columbus, Ohio. Workshop topics were chosen to understand the challenges and gaps that need to be addressed to maximize the impact of data on the nuclear industry, as well as the associated applications and risks. Discussions were focused around six specific application areas: Operation and Maintenance; Machine Learning in Nuclear Materials and Advanced Manufacturing; Cybersecurity; High-Performance Computing and Massive Computation; Big Data and Digital Twins; and Nuclear Non-Proliferation. The opportunities, challenges, and risks identified in the six focus areas explored in this workshop are diverse, but some common themes emerge, such as the importance of data integrity, quality, coverage, privacy, and traceability. Big data and AI/ML tools can be leveraged to reduce costs, optimize human tasking, and reduce human error across various application areas. In order for the nuclear industry to benefit from big data and advanced analytic capabilities, it is essential to address challenges and risks, such as data privacy, model reliability, and computational resource availability. Learning from other industries that have successfully implemented big data and AI/ML technologies, like the aerospace industry, can help the nuclear industry successfully integrate these technologies.
This R&D project, initiated by the DOE Nuclear Physics AI-Machine Learning initiative in 2022, leverages AI to address data processing challenges in high-energy nuclear experiments (RHIC, LHC, and future EIC). Our focus is on developing a demonstrator for real-time processing of high-rate data streams from sPHENIX experiment tracking detectors. The limitations of a 15 kHz maximum trigger rate imposed by the calorimeters can be negated by intelligent use of streaming technology in the tracking system. The approach efficiently identifies low momentum rare heavy flavor events in high-rate p+p collisions (3MHz), using Graph Neural Network (GNN) and High Level Synthesis for Machine Learning (hls4ml). Success at sPHENIX promises immediate benefits, minimizing resources and accelerating the heavy-flavor measurements. The approach is transferable to other fields. For the EIC, we develop a DIS-electron tagger using Artificial Intelligence - Machine Learning (AI-ML) algorithms for real-time identification, showcasing the transformative potential of AI and FPGA technologies in high-energy nuclear and particle experiments real-time data processing pipelines.
In a carbon-constrained world, future uses of nuclear power technologies can contribute to climate change mitigation as the installed electricity generating capacity and range of applications could be much greater and more diverse than with the current plants. To preserve the nuclear industry competitiveness in the global energy market, prognostics and health management (PHM) of plant assets is expected to be important for supporting and sustaining improvements in the economics associated with operating nuclear power plants (NPPs) while maintaining their high availability. Of interest are long-term operation of the legacy fleet to 80 years through subsequent license renewals and economic operation of new builds of either light water reactors or advanced reactor designs. Recent advances in data-driven analysis methods—largely represented by those in artificial intelligence and machine learning—have enhanced applications ranging from robust anomaly detection to automated control and autonomous operation of complex systems. The NPP equipment PHM is one area where the application of these algorithmic advances can significantly improve the ability to perform asset management. This paper provides an updated method-centric review of the full PHM suite in NPPs focusing on data-driven methods and advances since the last major survey article was published in 2015. The main approaches and the state of practice are described, including those for the tasks of data acquisition, condition monitoring, diagnostics, prognostics, and planning and decision-making. Research advances in non-nuclear power applications are also included to assess findings that may be applicable to the nuclear industry, along with the opportunities and challenges when adapting these developments to NPPs. Finally, this paper identifies key research needs in regard to data availability and quality, verification and validation, and uncertainty quantification.
Not Available
Unconstrained physics spaces between two or more nuclear data observables in a library occur when their values can be simultaneously adjusted without violating the uncertainties in either differential information or simulations of relevant integral experiments. Differential data are often too imprecise to fully bound all nuclear data observables of interest for application simulations. Integral data are simulated with combinations of nuclear data so that an error in one observable may be hidden by a counterbalancing error in another. In this manner compensating errors may lurk within nuclear data libraries and these errors have the potential to undermine the predictive power of neutron transport simulations, particularly in situations where there is no conclusive validation experiment that resembles the application of interest. The EUCLID project (Experiments Underpinned by Computational Learning for Improvements in Nuclear Data) developed a preliminary workflow to identify these unconstrained physics spaces by bringing together results from a large collection of integral experiments with their simulated counter-parts as well as differential information that have a one-to-one correspondence to nuclear data. This wealth of information is processed by machine learning tools for subsequent refinement by human experts. Here, we show how the EUCLID work-flow is executed by applying it first to 239 Pu and then to 9 Be nuclear data in ENDF/B-VIII.0.
Unconstrained physics spaces between two or more nuclear data observables in a library occur when their values can be simultaneously adjusted without violating the uncertainties in either differential information or simulations of relevant integral experiments. Differential data are often too imprecise to fully bound all nuclear data observables of interest for application simulations. Integral data are simulated with combinations of nuclear data so that an error in one observable may be hidden by a counterbalancing error in another. In this manner compensating errors may lurk within nuclear data libraries and these errors have the potential to undermine the predictive power of neutron transport simulations, particularly in situations where there is no conclusive validation experiment that resembles the application of interest. The EUCLID project (Experiments Underpinned by Computational Learning for Improvements in Nuclear Data) developed a preliminary workflow to identify these unconstrained physics spaces by bringing together results from a large collection of integral experiments with their simulated counter-parts as well as differential information that have a one-to-one correspondence to nuclear data. This wealth of information is processed by machine learning tools for subsequent refinement by human experts. Here, we show how the EUCLID work-flow is executed by applying it first to 239 Pu and then to 9 Be nuclear data in ENDF/B-VIII.0.
The development of a new fixed-source sensitivity tally capability is currently underway in the MCNP code. In recent research and development efforts that utilize machine learning to both seek problematic nuclear data as well as design experiments optimized to improve the nuclear data, the adjoint-weighted k-eigenvalue sensitivity tally capabilities have been heavily essential. In this paper, the motivation to expand the sensitivity tally capabilities beyond k-eigenvalues toward diverse fixed-source problems along with preliminary results and verification will be discussed.
Current needs of nuclear science and technology include complete, well-documented, and easily verifiable nuclear data. The complete data records require supporting nuclear bibliography, presently stored in dedicated libraries, in addition, to actual data. Additionally, experimental nuclear reaction data (EXFOR) and Nuclear Science References (NSR) databases contain compilations based on primary (journals) and secondary (conference proceedings, theses, preprints, etc.) publications, and data received from authors via private communications. The secondary library materials and private communications often represent a bottleneck for nuclear data verification, compilation, evaluation, and dissemination activities. To address this issue, bibliographic materials were scanned into PDF (Portable Document Format) files and uploaded in a relational database. The traditional scope of nuclear databases that includes meta-data and numbers derived from data in specialized formats was broadened to accommodate the large volumes of original nuclear data publications. The complete PDF publication files were stored in a relational database as Binary Large OBjects (BLOB). This unique collection of nuclear data compilations and supporting publications generate many opportunities for machine learning applications. The Web interfaces for authorized and public access to the EXFOR-NSR nuclear publications database were implemented at the U.S. National Nuclear Data Center, https://www.nndc.bnl.gov/ and IAEA Nuclear Data Section, https://www-nds.iaea.org/ . The current system is complementary to major nuclear libraries and narrowly focused on nuclear data compilation and evaluation procedures. The contents of the PDF database, details of implementation, and Web interface are described. New capabilities for data curation, knowledge preservation, worldwide dissemination, and natural language processing (NLP) applications are given.
Real-time reactor power monitoring is critical for a variety of nuclear applications, spanning safety, security, operations, and maintenance. While machine learning methods have shown promise in monitoring reactor power levels, there is limited research on their efficacy in label-starved environments. The goal of this work is to assess the feasibility of classifying nuclear reactor power level using multisource data in scenarios with limited labels. Data were collected using low-resolution multisensors at four nuclear reactor facilities: two large research reactors and two TRIGA reactors. Within each pair, one reactor dataset served as the source and the other as the target in a transfer learning paradigm. Twenty-three supervised models were trained on labeled sequences of magnetic field and acceleration data from each of the target sites. Self-learning and transfer learning methods were applied to the top performing models to assess their classification performance with increasing amounts of labeled data. While reactor power level classification was achieved with a Matthews Correlation Coefficient of up to 0.739 ± 0.003 and 0.622 ± 0.009 with only 400 sequences per power state for the large research reactor and TRIGA target sites, respectively, self-learning and transfer learning leveraging source site data did not improve target classification performance. These findings suggest that alternative methods, such as higher sensitivity sensors, digital twins, or the use of physics-informed models, are required to enable high-performance classification in machine learning approaches to reactor monitoring with a dearth of target ground truth.
Presentation discusses use of data in radiological emergency response and different methods (applied statistics, spectral clustering, spectral anomaly detection, etc.) that could be used in conjunction with machine learning to optimize response to real-world radiological events.
The Atlas of Neutron Resonances is the most comprehensive compilation of neutron resonances, thermal cross sections, resonance integrals and Maxwellian averaged cross sections generally available. For decades, the Atlas was carefully curated and maintained by Dr. Said Mughabghab who sadly passed on during the summer of 2018 after publishing the 2018 edition of the Atlas . We are continuing the development of this important compendium. To a large extent, the Atlas book is generated from a series of text files given in a single purpose domain-specific format. Therefore, we developed a software API and began the systematic assessment of the Atlas files. With this work past, we are now focusing on new efforts to expand the quality and scope of the Atlas . Current and recently completed projects include a cross comparison of the Atlas bibliography with Nuclear Science References and the EXFOR data library, a better determination of average resonance parameters, and using machine learning to assess the correctness of the spin group assignments of resonances tabulated in the Atlas .
Information from differential nuclear-physics experiments and theory is often too uncertain to accurately define nuclear-physics observables such as cross sections or energy spectra. Integral experimental data, representing the applications of these observables, are often more precise but depend simultaneously on too many of them to unambiguously identify issues in the observable with human expert analysis alone. Here, we explore how we can leverage physics knowledge gained from differential experimental data, nuclear theory, integral experiments, and neutron-transport calculations to better understand nuclear-physics observables in the context of the application area represented by integral experiments. We support this task with machine-learning methods to discern trends in a large amount of convoluted data. Differential and integral information was used in an analysis augmented by the random forest and the Shapley additive explanations metric. We chose as an application area one that is represented by criticality measurements and pulsed-sphere neutron-leakage spectra. We show one representative example ( 241 Pu fission observables) where the combination of differential and integral information allowed to resolve issues in data representing these observables. As a starting point, the machine learning (ML) algorithms highlighted several observables as leading potentially to bias in simulating integral experiments. Differential information, paired with sensitivity to integral quantities, allowed us then to pinpoint one specific observable ( 241 Pu fission cross section) as the main driver of bias. The comparison to integral experiments, on the other hand, allowed us to indicate a likely reliable experiment among several discrepant ones for this observables. In other cases (e.g., 239 Pu observables), we were not able to resolve the confounding introduced by integral experiments but instead highlighted the need for targeted new experiments and theory developments to better constrain the nuclear-physics space for the application area represented by integral experiments. We were able to combine information from differential experimental data, nuclear-physics theory, integral experiments, and neutron-transport simulations of the latter experiments with the help of the random forest algorithm and expert judgment. This combination of knowledge allows to improve our description of nuclear-physics observables as applied to a particular application area.
Distributed multisensor networks record multiple data streams that can be used as inputs to machine learning models designed to classify operations relevant to proliferation at nuclear reactors. The goal of this work is to demonstrate methods to assess the importance of each node (a single multisensor) and region (a group of proximate multisensors) to machine learning model performance in a reactor monitoring scenario. This, in turn, provides insight into model behavior, a critical requirement of data-driven applications in nuclear security. Using data collected at the High Flux Isotope Reactor at Oak Ridge National Laboratory via a network of Merlyn multisensors, two different models were trained to classify the reactor’s operational state: a hidden Markov model (HMM), which is simpler and more transparent, and a feed-forward neural network, which is less inherently interpretable. Traditional wrapper methods for feature importance were extended to identify nodes and regions in the multisensor network with strong positive and negative impacts on the classification problem. These spatial-importance algorithms were evaluated on the two different classifiers. The classification accuracy was then improved relative to baseline models via feature selection from 0.583 to 0.839 and from 0.811 ± 0.005 to 0.884 ± 0.004 for the HMM and feed-forward neural network, respectively. While some differences in node and region importance were observed when using different classifiers and wrapper methods, the nodes near the facility’s cooling tower were consistently identified as important—a conclusion further supported by studies on feature importance in decision trees. Node and region importance methods are model-agnostic, inform feature selection for improved model performance, and can provide insight into opaque classification models in the nuclear security domain.
In FY2020, Savannah River National Laboratory (SRNL) in collaboration with the Discovery Analytics Center (DAC) at Virginia Polytechnic Institute and State University (VT) began developing a demonstration prototype system that uses multiple machine learning and data analytic methods on largescale open data sources to identify new, developing, or undeclared nuclear programs. One of the most challenging aspects of applying machine learning techniques to such a problem is the high likelihood of extremely sparse data from disparate sources. To overcome this challenge, the current work will use a strategic combination of supervised, semi-supervised, and unsupervised learning techniques to ingest and fuse data streams to make a forecast of nuclear activities in a targeted geospatial location. Identifying potential data sources and training supervised learning algorithms is dependent upon the development of a robust foundation of targeted event domains that fundamentally define the nuclear activities of interest. This report documents the definition of a hierarchical structure for both nuclear activity and event domains that will be used to guide the research team in development or use of existing semantic dictionaries that are instrumental to searching, parsing, and categorizing events for the forecasting system’s use.