Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “supervised machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Prediction of the Cu oxidation state from EELS and XAS spectra using supervised machine learning

Abstract Electron energy loss spectroscopy (EELS) and X-ray absorption spectroscopy (XAS) provide detailed information about bonding, distributions and locations of atoms, and their coordination numbers and oxidation states. However, analysis of XAS/EELS data often relies on matching an unknown experimental sample to a series of simulated or experimental standard samples. This limits analysis throughput and the ability to extract quantitative information from a sample. In this work, we have trained a random forest model capable of predicting the oxidation state of copper based on its L-edge spectrum. Our model attains an R 2 score of 0.85 and a root mean square error of 0.24 on simulated data. It has also successfully predicted experimental L-edge EELS spectra taken in this work and XAS spectra extracted from the literature. We further demonstrate the utility of this model by predicting simulated and experimental spectra of mixed valence samples generated by this work. This model can be integrated into a real-time EELS/XAS analysis pipeline on mixtures of copper-containing materials of unknown composition and oxidation state. By expanding the training data, this methodology can be extended to data-driven spectral analysis of a broad range of materials.

36 MATERIALS SCIENCE↗

Detection of Diversion in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors (MRs) pose new challenges for international safeguards. Here, their small size and mass reproducibility make them ideal for deployment in greater numbers and in remote locations, making the job of safeguards inspectors more challenging. Machine learning (ML) is currently being applied to many fields to augment human performance and increase automation; in particular, ML could be used to provide insight for international inspectors to help detect the diversion of nuclear fuel from MR cores. Four ML model types (k-nearest neighbors, decision tree, random forest, and histogram-based gradient boosted ensemble) were trained on integrated flux and critical control drum angle data generated with Serpent 2 for a realistic heat pipe MR design, achieving nearly 100% binary classification accuracy of nominal and diversion core configurations by the end of 1 full power year for three of the four model types. Regression model variants were also trained, using the same input data, for predicting the number of fuel pins diverted. Root-mean-square errors below 5% of the total number of fuel pins were achieved by the 1 full power year mark for all models.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Sensor Reduction for Diversion Detection in a Realistic Heat Pipe Microreactor Using Supervised Machine Learning

Microreactors are designed as a smaller, cheaper, and safer alternative to traditional nuclear power plants. Their non-traditional characteristics and prospect of mass production and deployment will likely require new approaches to nuclear safeguards. The primary proliferation concern with microreactors is the diversion of fuel material. Such diversion may produce measurable defects in key physical attributes like neutron flux, which may in turn be detectable using machine learning models. Preliminary work has demonstrated this ability for modeled nominal and diversion scenarios using large quantities of energy integrated neutron flux data. In practice, the number of available sensors for such measurements will be limited and energy integrated flux information will not be available. This work explores the ability of tree-based gradient boosted ensemble models to classify a given microreactor core is nominal or diversion, and determine the number of fuel pins diverted in the case of diversion with reduced numbers of sensors and more realistic detector responses. Classification accuracy of greater than 98% and regression errors as low as 5% of the total number of fuel pins were achieved with as few as 15 sensors, compared to 99% and 4.1% with a maximum of 240 sensors.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Optimization and supervised machine learning methods for fitting numerical physics models without derivatives

Here, we address the calibration of a computationally expensive nuclear physics model for which derivative information with respect to the fit parameters is not readily available. Of particular interest is the performance of optimization-based training algorithms when dozens, rather than millions or more, of training data are available and when the expense of the model places limitations on the number of concurrent model evaluations that can be performed. As a case study, we consider the Fayans energy density functional model, which has characteristics similar to many model fitting and calibration problems in nuclear physics. We analyze hyperparameter tuning considerations and variability associated with stochastic optimization algorithms and illustrate considerations for tuning in different computational settings.

97 MATHEMATICS AND COMPUTING↗

Prediction of the Cu Oxidation State from EELS and XAS Spectra Using Supervised Machine Learning

Electron energy loss spectroscopy (EELS) and X-ray absorption spectroscopy (XAS) provide detailed information about distributions and locations of atoms, their coordination numbers and oxidation states, and the bonding characteristics [1]. However, analysis of XAS/EELS data often relies on matching the spectra of an unknown experimental sample to a series of simulated or experimental spectra of standard samples. Here, this limits analysis throughput and the ability to extract quantitative information from a sample.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Binary Analysis with Architecture and Code Section Detection using Supervised Machine Learning

When presented with an unknown binary, which may or may not be complete, having the ability to determine information about it is critical to future reverse engineering, particularly in discovering the binary’s intended use and potentially malicious nature. This paper details techniques to both identify the machine architecture of the binary, as well as to locate the important code segments within the file. This identification of unknown binaries makes use of a technique called byte histogram in addition to various machine learning (ML) techniques, which we call “What is it Binary” or WiiBin. Benefits of byte histograms reflect the simplicity of calculation and do not rely on file headers or metadata, allowing for acceptable results when only a small portion of the original file is available. Utilizing WiiBin, we were able to accurately (>80%) determine the architecture of test binaries with as little as a 20% contagious portion of the file present. We were also able to determine the location of code sections within a binary by utilizing the WiiBin framework. Ultimately, the more information that can be gleaned from a binary file, the easier it is to successfully reverse engineer.

99 GENERAL AND MISCELLANEOUS↗

Overview of leakage scenarios in supervised machine learning

Machine learning (ML) provides powerful tools for predictive modeling. ML’s popularity stems from the promise of sample-level prediction with applications across a variety of fields from physics and marketing to healthcare. However, if not properly implemented and evaluated, ML pipelines may contain leakage typically resulting in overoptimistic performance estimates and failure to generalize to new data. This can have severe negative financial and societal implications. Our aim is to expand understanding associated with causes leading to leakage when designing, implementing, and evaluating ML pipelines. Illustrated by concrete examples, we provide a comprehensive overview and discussion of various types of leakage that may arise in ML pipelines.

97 MATHEMATICS AND COMPUTING↗

Integrating Experiments and Well Logs to Predict Caney Shale Static Mechanical Properties During Production with Supervised Machine Learning

Caney shale is one of the emerging oil reservoirs in Oklahoma. Understanding the impact of effective stress on its mechanical properties is critical for predicting hydraulic fracture geometry and overall hydrocarbon production. The objective of our study is to evaluate the impact of effective stress on the dynamic Young’s modulus using ultrasonic velocity measurements for Caney shale samples. A triaxial cell was utilized to measure ultrasonic (P-wave and S-wave) velocities for ten downhole Caney shale samples under various effective stresses to indirectly assess the impact of pore pressure change. The dynamic Young’s moduli estimated from these measurements were integrated with available conventional well logs (excluding sonic logs) and triaxial test results from Benge et al. (2021) to predict the static Young’s modulus using Random Forest (RF) and Extreme Gradient Boosting (XGBoost) models. The results showed that the estimated dynamic Young’s moduli from ultrasonic measurements were higher than the corresponding static Young’s modulus of cores from the same vertical well at similar depths. With increasing effective stress, the dynamic Young’s modulus increased for all samples. The estimated dynamic-to-static correction factor tended to be higher in zones of high neutron porosity (PHIN) and low density compared to other zones. Finally, SHapley Additive exPlanations (SHAP) for RF and XGBoost models identified depth, gamma ray (GR), and PHIN as key features for predicting the static Young’s modulus. This study enhances our understanding of the dynamic and static Young’s moduli for the Caney shale interval, as a function of effective stress and conventional well logs. The findings from this study can improve predictions of production throughout the well's lifespan by offering insights into the mechanical property degradation resulting from pore pressure depletion.

Kholy, Sherif M.↗

Comprehensive Chemical Fingerprinting by Multidimensional GC and Supervised Machine Learning

This project leverages advances in machine learning based data analysis techniques and untargeted omic analytical methods to progress nuclear nonproliferation technologies beyond current capabilities. The developed approaches can be used to identify and detect complex chemical fingerprints of facilities of interest. These techniques have been developed for fields such as metabolomics and genomics but have not been applied to nuclear nonproliferation applications. Adaptation of these techniques for volatile organic compound analysis has far reaching application within the scientific community including environmental chemistry, atmospheric physics, and climate sciences.

97 MATHEMATICS AND COMPUTING↗

Comprehensive Chemical Fingerprinting by Multidimensional GC and Supervised Machine Learning

This project leveraged advances in machine learning based data analysis techniques and untargeted analytical methods for organic analysis to progress nuclear nonproliferation technologies beyond current capabilities. The developed approaches can be used to detect and identify complex chemical fingerprints of facilities of interest. These techniques have been developed for fields such as metabolomics and genomics but have not been applied to nuclear nonproliferation applications. Adaptation of these techniques for volatile organic compound analysis has far reaching application within the scientific community including environmental chemistry, atmospheric physics, and climate sciences.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Unconventional Wells Interference: Supervised Machine Learning for Detecting Fracture Hits

The primary objective of the study was development of a machine learning (ML)-based workflow for fracture hit (“frac hit”) detection and monitoring using shale oil-field data such as drilling surveys, production history (oil and produced water), pressure, and fracking start time and duration records. The ML method takes advantage of long short-term memory (LSTM) and multilayer perceptron (MLP) neural networks to identify the frac hits due to hydraulic communication between the fracking child well(s) and the producing parent well(s) within the same pad (intra-pad interaction) and/or on different pads (inter-pad interaction). It utilizes time series of pressure and production data from within a pad and from adjacent pads. The workflow can capture time variable features of frac hits when the model architecture is deep and wide enough, with enough trainable parameters for deep learning and feature extraction, as demonstrated in this paper by using training and testing subsets of the field data from selected neighboring pads with over a couple of hundred wells. The study was focused on frac-hit interaction among paired wells and demonstrated that the ML model, once trained, can predict the frac-hit probability.

58 GEOSCIENCES↗

Analyzing Natural Language Context in Human-Machine Teaming using Supervised Machine Learning

Building a foundation for trustworthiness and trust verification in multi-asset teaming is the research challenge of Autonomy Teaming and TRAjectories for Complex Trusted Operational Reliability (ATTRACTOR). The Design Reference Mission (DRM) for ATTRACTOR is a search and rescue mission objective governed by a multi-member team consisting of human and machine operators. A crucial component to the effort is the communication between humans and autonomous agents throughout both planning and execution stages of the mission. Intuitive communication methods and modalities are posited as critical enablers for certifying trust and trustworthiness. This paper reports on the data collection and analysis conducted in support of the Human Informed Natural-language GANs Evaluation (HINGE)project to attain explainable and trusted communication between human-machine assets. Two identically curated image description datasets were acquired for HINGE, both consisting of two unique input modalities (typed vs. verbal) and retrieved in two distinct contexts (general vs. specific). The gathered datasets were assessed and compared using Parts-of-Speech (POS)features, sentence similarity metrics, and linguistic analysis. Then, the datasets were modeled and tested separately and in combination with one another using machine learning algorithms. The comparison and testing results reveal a superior dataset, by which a preferred context and input is understood, for generating image representations of missing persons using a Generative Adversarial Network (GAN).

Bryan A Barrows↗