Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “supervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

SNM Radiation Signature Classification Using Different Semi-Supervised Machine Learning Models

The timely detection of special nuclear material (SNM) transfers between nuclear facilities is an important monitoring objective in nuclear nonproliferation. Persistent monitoring enabled by successful detection and characterization of radiological material movements could greatly enhance the nuclear nonproliferation mission in a range of applications. Supervised machine learning can be used to signal detections when material is present if a model is trained on sufficient volumes of labeled measurements. However, the nuclear monitoring data needed to train robust machine learning models can be costly to label since radiation spectra may require strict scrutiny for characterization. Therefore, this work investigates the application of semi-supervised learning to utilize both labeled and unlabeled data. As a demonstration experiment, radiation measurements from sodium iodide (NaI) detectors are provided by the Multi-Informatics for Nuclear Operating Scenarios (MINOS) venture at Oak Ridge National Laboratory (ORNL) as sample data. Anomalous measurements are identified using a method of statistical hypothesis testing. After background estimation, an energy-dependent spectroscopic analysis is used to characterize an anomaly based on its radiation signatures. In the absence of ground-truth information, a labeling heuristic provides data necessary for training and testing machine learning models. Supervised logistic regression serves as a baseline to compare three semi-supervised machine learning models: co-training, label propagation, and a convolutional neural network (CNN). In each case, the semi-supervised models outperform logistic regression, suggesting that unlabeled data can be valuable when training and demonstrating value in semi-supervised nonproliferation implementations.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Phase behavior of continuous-space systems: A supervised machine learning approach

The phase behavior of complex fluids is a challenging problem for molecular simulations. Supervised machine learning (ML) methods have shown potential for identifying the phase boundaries of lattice models. In this work, we extend these ML methods to continuous-space systems. We propose a convolutional neural network model that utilizes grid-interpolated coordinates of molecules as input data of ML and optimizes the search for phase transitions with different filter sizes. We test the method for the phase diagram of two off-lattice models, namely, the Widom–Rowlinson model and a symmetric freely jointed polymer blend, for which results are available from standard molecular simulations techniques. The ML results show good agreement with results of previous simulation studies with the added advantage that there is no critical slowing down. We find that understanding intermediate structures near a phase transition and including them in the training set is important to obtain the phase boundary near the critical point. The method is quite general and easy to implement and could find wide application to study the phase behavior of complex fluids.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data compression and inference in cosmology with self-supervised machine learning

ABSTRACT The influx of massive amounts of data from current and upcoming cosmological surveys necessitates compression schemes that can efficiently summarize the data with minimal loss of information. We introduce a method that leverages the paradigm of self-supervised machine learning in a novel manner to construct representative summaries of massive data sets using simulation-based augmentations. Deploying the method on hydrodynamical cosmological simulations, we show that it can deliver highly informative summaries, which can be used for a variety of downstream tasks, including precise and accurate parameter inference. We demonstrate how this paradigm can be used to construct summary representations that are insensitive to prescribed systematic effects, such as the influence of baryonic physics. Our results indicate that self-supervised machine learning techniques offer a promising new approach for compression of cosmological data as well as its analysis.

Astronomy & Astrophysics↗

A Weakly Supervised Machine Learning Procedure for Magnet Quench Diagnostics

Voltage taps remain the standard and reliable diagnostic tool for detecting quenches in superconducting magnets. However, they identify a quench only at the time of voltage rise and do not provide information on earlier physical precursors. In this work, we investigate whether acoustic emission data can reveal precursor activity that occurs before conventional voltage detection using machine learning techniques. We introduce an event selection method and a weakly supervised machine learning procedure to learn data-driven criteria for identifying potential acoustic precursors to quenches. Two Convolutional Neural Network (CNN) architectures are trained: one on acoustic sensor events from our selection procedure and one on the Fast Fourier Transforms (FFTs) of these events. Both networks are trained iteratively using confidence-weighted loss functions to associate certain subsets of training data with a precursor label. We evaluate the performance of these models by examining the time distribution of events classified as potential precursors relative to the quench onset. Results indicate that the proposed approach can possibly distinguish acoustic emission events occurring closer to the quench from earlier acoustic activity during ramping, suggesting the potential for flagging quench precursors in acoustic data.

Khan, Maira [Fermilab] (ORCID:0009000891602387)↗

CoMTE: Counterfactual Explanations for Supervised Machine Learning Frameworks on Multivariate Time Series Data

CoMTE is a novel explainability technique that provides counterfactual explanations for supervised machine learning frameworks on multivariate time series data. CoMTE outperforms state-of-the-art explainability methods on several different machine learning frameworks and data sets in comprehensibility and robustness. CoMTE can be used to debug machine learning frameworks and gain a better understanding of the underlying multivariate time series data. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525. SAND2021-1686 O

Leung, VitusJ.↗

Comparison of Supervised and Un-Supervised Machine Learning Algorithms for Threat Detection and Scintillator Performance for Radiation Portal Monitoring

Following the events of September 11, 2001, international border crossing have been equipped with radiation portal monitors (RPMs) to identify illicit radioactive material. Polyvinyl toluene (PVT) scintillators are commonly used due to their low cost and reasonable maintainability, however they offer low spectral resolution. Despite the fact that over twenty years has transpired since this event, radioisotopes are still typically identified by hand-crafted classification algorithms, e.g., total counts or energy windowing, and exhibit relatively poor performance in detecting threats at the low false alarm rates required to support the stream of commerce. While some improvement to performance has been realized via the use of supervised machine learning, these classification algorithms typically utilize simulations in lieu of real data due to the sparsity of data for one or more classes. Accordingly, the performance of these algorithms is somewhat less than optimal when examining experiments or simulations with model mismatch. Consequently, in this work, we examine the application of a number of unsupervised machine learning, anomaly detection based algorithms, to circumvent the inverse crime when analyzing spectroscopy data for RPMs. We also compare anomaly detection results with those obtained via the use of supervised classification detection ML algorithms when model mismatch is introduced between the simulated threat items utilized for training/testing. Finally, we compared the performance of the PVT scintillators to those obtained with higher resolution detectors using both anomaly detection and supervised classification algorithms.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

High-fidelity retrieval from instantaneous line-of-sight returns of nacelle-mounted lidar including supervised machine learning

Abstract. Wind turbine applications that leverage nacelle-mounted Doppler lidar are hampered by several sources of uncertainty in the lidar measurement, affecting both bias and random errors. Two problems encountered especially for nacelle-mounted lidar are solid interference due to intersection of the line of sight with solid objects behind, within, or in front of the measurement volume and spectral noise due primarily to limited photon capture. These two uncertainties, especially that due to solid interference, can be reduced with high-fidelity retrieval techniques (i.e., including both quality assurance/quality control and subsequent parameter estimation). Our work compares three such techniques, including conventional thresholding, advanced filtering, and a novel application of supervised machine learning with ensemble neural networks, based on their ability to reduce uncertainty introduced by the two observed nonideal spectral features while keeping data availability high. The approach leverages data from a field experiment involving a continuous-wave (CW) SpinnerLidar from the Technical University of Denmark (DTU) that provided scans of a wide range of flows both unwaked and waked by a field turbine. Independent measurements from an adjacent meteorological tower within the sampling volume permit experimental validation of the instantaneous velocity uncertainty remaining after retrieval that stems from solid interference and strong spectral noise, which is a validation that has not been performed previously. All three methods perform similarly for non-interfered returns, but the advanced filtering and machine learning techniques perform better when solid interference is present, which allows them to produce overall standard deviations of error between 0.2 and 0.3 m s−1, or a 1 %–22 % improvement versus the conventional thresholding technique, over the rotor height for the unwaked cases. Between the two improved techniques, the advanced filtering produces 3.5 % higher overall data availability, while the machine learning offers a faster runtime (i.e., ∼ 1 s to evaluate) that is therefore more commensurate with the requirements of real-time turbine control. The retrieval techniques are described in terms of application to CW lidar, though they are also relevant to pulsed lidar. Previous work by the authors (Brown and Herges, 2020) explored a novel attempt to quantify uncertainty in the output of a high-fidelity lidar retrieval technique using simulated lidar returns; this article provides true uncertainty quantification versus independent measurement and does so for three techniques rather than one.

47 OTHER INSTRUMENTATION↗

Capturing the Physics of MaNGA Galaxies with Self-supervised Machine Learning

As available data sets grow in size and complexity, advanced visualization tools enabling their exploration and analysis become more important. In modern astronomy, integral field spectroscopic galaxy surveys are a clear example of increasing high dimensionality and complex data sets, which challenges the traditional methods used to extract the physical information they contain. Here, we present the use of a novel self-supervised machine-learning method to visualize the multidimensional information on stellar population and kinematics in the MaNGA survey in a 2D plane. Our framework is insensitive to nonphysical properties such as the size of the integral field unit and is therefore able to order galaxies according to their resolved physical properties. Using the extracted representations, we study how galaxies distribute based on their resolved and global physical properties. We show that even when exclusively using information about the internal structure, galaxies naturally cluster into two well-known categories, rotating main-sequence disks and massive slow rotators, from a purely data-driven perspective, hence confirming distinct assembly channels. Low-mass rotation-dominated quenched galaxies appear as a third cluster only if information about the integrated physical properties is preserved, suggesting a mixture of assembly processes for these galaxies without any particular signature in their internal kinematics that distinguishes them from the two main groups. The framework for data exploration is publicly released with this publication, ready to be used with the MaNGA or other integral field data sets.

79 ASTRONOMY AND ASTROPHYSICS↗

Prediction of Specificity of α-Conotoxins to Subtypes of Human Nicotinic Acetylcholine Receptors with Semi-supervised Machine Learning

Conotoxins are a family of highly toxic neurotoxins composed of cysteine-rich peptides produced by marine cone snails. The most lethal cone snail species to humans is Conus geographus, with fatality rates of up to ∼65% from a single sting, which is caused mostly by the activity of α-conotoxins against human nicotinic acetylcholine receptors (nAChRs). While sequence-based machine learning (ML) classifiers have been trained to identify targets of conotoxins binding voltage-gated ion channels, no ML model has been built to predict the subtype-specific nAChR targets of α-conotoxins. Here, we trained an ML model in a semi-supervised manner to predict the specificity of α-conotoxin binding toward different human nAChR subtypes to overcome the challenge of limited data in subtype-specific nAChR targets of α-conotoxins and the issue that one α-conotoxin can bind multiple nAChR subtypes with high selectivity. We considered additional features of sequences of α-conotoxins in training our ML model, including the secondary structure propensities and electrostatic properties, which resulted in better prediction capability for the ML model. Notably, we identify that most α-conotoxins bind to α3β2, α1γδ, and α7 subtypes of human nAChRs. Our findings from this study provide a framework for predicting targets of various kinds of toxins.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting measures of soil health using the microbiome and supervised machine learning

Soil health encompasses a range of biological, chemical, and physical soil properties that sustain the commercial and ecological value of agroecosystems. Monitoring soil health requires a comprehensive set of diagnostics that can be cost-prohibitive for routine analyses. The soil microbiome provides a rich source of information about soil properties, which can be assayed in a high-throughput, cost-effective way. We evaluated the accuracy of random forest (RF) and support vector machine (SVM) regression and classification models in predicting 12 measures of soil health, tillage status, and soil texture from 16S rRNA gene amplicon data with an operationally relevant sample set. We validated the efficacy of the best performing models against independent datasets and also tested best practices for processing microbiome data for use in machine learning. Soil health metrics could be predicted from microbiome data with the best models achieving a Kappa value of ~0.65, for categorical assessments, and a R2 value of ~0.8, for numerical scores. Biological health ratings were better predicted than chemical or physical ratings. Validation with independent datasets revealed that models had general predictive value for soil properties, including yield. The ecological profiles of several taxa important for model accuracy matched the observed relationships with soil health, including Pyrinomonadaceae, Nitrososphaeraceae, and Candidatus Udeaobacter. Models trained at the highest taxonomic resolution proved most accurate, with losses in accuracy resulting from rarefying, sparsity filtering, and aggregating at higher taxonomic ranks. Furthermore, our study provides the groundwork for developing scalable technology to use microbiome-based diagnostics for the assessment of soil health.

16S rRNA gene↗

A semi-supervised machine learning detector for physics events in tokamak discharges

Databases of physics events have been used in various fusion research applications, including the development of scaling laws and disruption avoidance algorithms, yet they can be time-consuming and tedious to construct. This paper presents a novel application of the label spreading semi-supervised learning algorithm to accelerate this process by detecting distinct events in a large dataset of discharges, given few manually labeled examples. A high detection accuracy (> 85%) for H-L back transitions and initially rotating locked modes is demonstrated on a dataset of hundreds of discharges from DIII-D with manually identified events for which only 3 discharges are initially labeled by the user. Lower yet reasonable performance (~75%) is also demonstrated for the core radiative collapse, an event with a much lower prevalence in the dataset. Additionally, analysis of the performance sensitivity indicates that the same set of algorithmic parameters is optimal for each event. Furthermore, this suggests that the method can be applied to detect a variety of other events not included in this paper, given that the event is well described by a set of 0D signals robustly available on many discharges. Procedures for analysis of new events are demonstrated, showing automatic event detection with increasing fidelity as the user strategically adds manually labeled examples. Detections on Alcator C-Mod and EAST are also shown, demonstrating the potential for this to be used on a multi-tokamak dataset.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A semi-supervised machine learning detector for physics events in tokamak discharges

Databases of physics events have been used in various fusion research applications, including the development of scaling laws and disruption avoidance algorithms, yet they can be time-consuming and tedious to construct. This paper presents a novel application of the label spreading semi-supervised learning algorithm to accelerate this process by detecting distinct events in a large dataset of discharges, given few manually labeled examples. A high detection accuracy (>85%) for H-L back transitions and initially rotating locked modes is demonstrated on a dataset of hundreds of discharges from DIII-D with manually identified events for which only 3 discharges are initially labeled by the user. Lower yet reasonable performance (~75%) is also demonstrated for the core radiative collapse, an event with a much lower prevalence in the dataset. Additionally, analysis of the performance sensitivity indicates that the same set of algorithmic parameters is optimal for each event. This suggests that the method can be applied to detect a variety of other events not included in this paper, given that the event is well described by a set of 0D signals robustly available on many discharges. Procedures for analysis of new events are demonstrated, showing automatic event detection with increasing fidelity as the user strategically adds manually labeled examples. Detections on Alcator C-Mod and EAST are also shown, demonstrating the potential for this to be used on a multi-tokamak dataset.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Battery Charge Curve Prediction via Feature Extraction and Supervised Machine Learning

Real-time onboard state monitoring and estimation of a battery over its lifetime is indispensable for the safe and durable operation of battery-powered devices. In this study, a methodology to predict the entire constant-current cycling curve with limited input information that can be collected in a short period of time is developed. A total of 10 066 charge curves of LiNiO 2 -based batteries at a constant C-rate are collected. With the combination of a feature extraction step and a multiple linear regression step, the method can accurately predict an entire battery charge curve with an error of < 2% using only 10% of the charge curve as the input information. The method is further validated across other battery chemistries (LiCoO 2 -based) using open-access datasets. The prediction error of the charge curves for the LiCoO 2 -based battery is around 2% with only 5% of the charge curve as the input information, indicating the generalization of the developed methodology for predicting battery cycling curves. The developed method paves the way for fast onboard health status monitoring and estimation for batteries during practical applications.

25 ENERGY STORAGE↗

Prediction of the Cu oxidation state from EELS and XAS spectra using supervised machine learning

Abstract Electron energy loss spectroscopy (EELS) and X-ray absorption spectroscopy (XAS) provide detailed information about bonding, distributions and locations of atoms, and their coordination numbers and oxidation states. However, analysis of XAS/EELS data often relies on matching an unknown experimental sample to a series of simulated or experimental standard samples. This limits analysis throughput and the ability to extract quantitative information from a sample. In this work, we have trained a random forest model capable of predicting the oxidation state of copper based on its L-edge spectrum. Our model attains an R 2 score of 0.85 and a root mean square error of 0.24 on simulated data. It has also successfully predicted experimental L-edge EELS spectra taken in this work and XAS spectra extracted from the literature. We further demonstrate the utility of this model by predicting simulated and experimental spectra of mixed valence samples generated by this work. This model can be integrated into a real-time EELS/XAS analysis pipeline on mixtures of copper-containing materials of unknown composition and oxidation state. By expanding the training data, this methodology can be extended to data-driven spectral analysis of a broad range of materials.

36 MATERIALS SCIENCE↗

Detection of Diversion in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors (MRs) pose new challenges for international safeguards. Here, their small size and mass reproducibility make them ideal for deployment in greater numbers and in remote locations, making the job of safeguards inspectors more challenging. Machine learning (ML) is currently being applied to many fields to augment human performance and increase automation; in particular, ML could be used to provide insight for international inspectors to help detect the diversion of nuclear fuel from MR cores. Four ML model types (k-nearest neighbors, decision tree, random forest, and histogram-based gradient boosted ensemble) were trained on integrated flux and critical control drum angle data generated with Serpent 2 for a realistic heat pipe MR design, achieving nearly 100% binary classification accuracy of nominal and diversion core configurations by the end of 1 full power year for three of the four model types. Regression model variants were also trained, using the same input data, for predicting the number of fuel pins diverted. Root-mean-square errors below 5% of the total number of fuel pins were achieved by the 1 full power year mark for all models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Sensor Reduction for Diversion Detection in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors are designed as a smaller, cheaper, and safer alternative to traditional nuclear power plants. Their non-traditional characteristics and prospect of mass production and deployment will likely require new approaches to nuclear safeguards. The primary proliferation concern with microreactors is the diversion of fuel material. Such diversion may produce measurable defects in key physical attributes like neutron flux, which may in turn be detectable using machine learning models. Preliminary work has demonstrated this ability for modeled nominal and diversion scenarios using large quantities of energy integrated neutron flux data. In practice, the number of available sensors for such measurements will be limited and energy integrated flux information will not be available. This work explores the ability of tree-based gradient boosted ensemble models to classify a given microreactor core is nominal or diversion, and determine the number of fuel pins diverted in the case of diversion with reduced numbers of sensors and more realistic detector responses. Classification accuracy of greater than 98% and regression errors as low as 5% of the total number of fuel pins were achieved with as few as 15 sensors, compared to 99% and 4.1% with a maximum of 240 sensors.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗