Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Machine learning prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Predicting nepheline precipitation in waste glasses using ternary submixture model and machine learning

Nepheline precipitation in nuclear waste glasses during vitrification can be detrimental due to its negative effect on chemical durability. Developing models to accurately predict nepheline precipitation from compositions is important to increase waste loading since existing models can be overly conservative. In this study, an expanded dataset containing 955 glasses was compiled from literature data, where 355 glasses are for high-level waste (HLW). Previously developed submixture models were refitted using the new dataset, where a misclassification rate of 7.8% was achieved. Nine machine learning (ML) algorithms (e.g., k-nearest neighbor, Gaussian process regression, artificial neural network, support vector machine, decision tree, etc.) were applied to evaluate their ability of predicting nepheline precipitation from compositions. Model accuracy, precision, recall/sensitivity, and F1 score were systemically compared between different ML algorithms and modeling protocols. Good model prediction with an accuracy ~0.9 (misclassification rate of ~10%) was observed with different algorithms under certain protocol. This study evaluated various ML models to predict nepheline precipitations in waste glasses, highlighting the importance of data preparation, modeling protocol, and their effect on model stability and reproducibility. The results provide insights into applying ML to predict glass properties and suggest areas for future research on modeling nepheline precipitations.

Lu, Xiaonan↗

Geometric Measures of Trustworthiness for Machine Learning Predictions

his report details the findings from the research and investigation of Geometric Measures of Trustworthiness for Machine Learning Predictions. We explored the trustworthiness of machine learning (ML) models’ predictions using geometric measures to quantify the similarity of a query point with the training data. Predictive uncertainty in ML can originate from at least three sources: (1) Model uncertainty, which represents the uncertainty in model form (e.g. decision tree, vs neural network) and estimating the model parameters from the training data, (2) Data uncertainty, which represents the natural complexities of the data such as class overlap and inherent noise, and (3) Distributional uncertainty, which represents the mismatch between the training and operational distributions. The proposed measures focus on measuring and explaining the data and distributional uncertainties by measuring the relationships of operational data with the training data.

97 MATHEMATICS AND COMPUTING↗

Human limits in machine learning: prediction of potato yield and disease using soil microbiome data

Abstract Background The preservation of soil health is a critical challenge in the 21st century due to its significant impact on agriculture, human health, and biodiversity. We provide one of the first comprehensive investigations into the predictive potential of machine learning models for understanding the connections between soil and biological phenotypes. We investigate an integrative framework performing accurate machine learning-based prediction of plant performance from biological, chemical, and physical properties of the soil via two models: random forest and Bayesian neural network. Results Prediction improves when we add environmental features, such as soil properties and microbial density, along with microbiome data. Different preprocessing strategies show that human decisions significantly impact predictive performance. We show that the naive total sum scaling normalization that is commonly used in microbiome research is one of the optimal strategies to maximize predictive power. Also, we find that accurately defined labels are more important than normalization, taxonomic level, or model characteristics. ML performance is limited when humans can’t classify samples accurately. Lastly, we provide domain scientists via a full model selection decision tree to identify the human choices that optimize model prediction power. Conclusions Our study highlights the importance of incorporating diverse environmental features and careful data preprocessing in enhancing the predictive power of machine learning models for soil and biological phenotype connections. This approach can significantly contribute to advancing agricultural practices and soil health management.

Aghdam, Rosa↗

Gaussian approximation of dispersion potentials for efficient featurization and machine-learning predictions of metal–organic frameworks

Energy-related descriptors in machine learning are a promising strategy to predict adsorption properties of metal–organic frameworks (MOFs) in the low-pressure regime. Interactions between hosts and guests in these systems are typically expressed as a sum of dispersion and electrostatic potentials. The energy landscape of dispersion potentials plays a crucial role in defining Henry’s constants for simple probe molecules in MOFs. To incorporate more information about this energy landscape, we introduce the Gaussian-approximated Lennard-Jones (GALJ) potential, which fits pairwise Lennard-Jones potentials with multiple Gaussians by varying their heights and widths. The GALJ approach is capable of replicating information that can be obtained from the original LJ potentials and enables efficient development of Gaussian integral (GI) descriptors that account for spatial correlations in the dispersion energy environment. GI descriptors would be computationally inconvenient to compute using the usual direct evaluation of the dispersion potential energy surface. Here, we show that these new GI descriptors lead to improvement in ML predictions of Henry’s constants for a diverse set of adsorbates in MOFs compared to previous approaches to this task.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Dashboard for Visualizing Molecular Property Prediction Machine Learning Results

This is a dashboard for exploring the results of machine learning models for predicting molecular properties from molecular structure. It includes tools for: 1. Modifying molecules to observe the change in predicted properties 2. Exploring the relationship between molecular structure and predicted properties 3. Recommending structurally similar molecules with improved properties 4. Exploring the impact of data subsampling on model performance metrics

Xu, Audrey↗

Machine learning prediction of electron density and temperature from He I line ratios

We propose to utilize machine learning to predict the electron density, ne, and temperature, T e , from He I line intensity ratios. In this approach, training data consist of measured He I line ratios as input and ne and T e measured using other diagnostic(s) as desired output, which is a Langmuir probe in our study. Support vector machine regression analysis is, then, performed with the training data to develop a predictive model for n e and T e , separately. It is confirmed that n e and T e predicted using the developed models agree well with those from the Langmuir probe in the ranges of 0.28 × 10 18 ≤ n e (m -3 ) ≤ 3.8 × 10 18 and 3.2 ≤ T e (eV) ≤ 7.5. The developed models are, further, examined with an evaluation data, which are not included in the training data, and are found to well reproduce absolute values and radial profiles of probe-measured n e and T e .

47 OTHER INSTRUMENTATION↗

Machine Learning Predictions of Simulated Self-Diffusion Coefficients for Bulk and Confined Pure Liquids

Diffusion properties of bulk fluids have been predicted using empirical expressions and machine learning (ML) models, suggesting that predictions of diffusion also should be possible for fluids in confined environments. The ability to quickly and accurately predict diffusion in porous materials would enable new discoveries and spur development in relevant technologies such as separations, catalysis, batteries, and subsurface applications. Here in this work, we apply artificial neural network (ANN) models to predict the simulated self-diffusion coefficients of real liquids in both bulk and pore environments. The training data sets were generated from molecular dynamics (MD) simulations of Lennard-Jones particles representing a diverse set of 14 molecules ranging from ammonia to dodecane over a range of liquid pressures and temperatures. Planar, cylindrical, and hexagonal pore models consisted of walls composed of carbon atoms. Our simple model for these liquids was primarily used to generate ANN training data, but the simulated self-diffusion coefficients of bulk liquids show excellent agreement with experimental diffusion coefficients. ANN models based on simple descriptors accurately reproduced the MD diffusion data for both bulk and confined liquids, including the trend of increased mobility in large pores relative to the corresponding bulk liquid.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Effects of Substance Use and Antisocial Personality on Neuroimaging-Based Machine Learning Prediction of Schizophrenia

Abstract Background and hypothesis Neuroimaging-based machine learning (ML) algorithms have the potential to aid the clinical diagnosis of schizophrenia. However, literature on the effect of prevalent comorbidities such as substance use disorder (SUD) and antisocial personality (ASPD) on these models’ performance has remained unexplored. We investigated whether the presence of SUD or ASPD affects the performance of neuroimaging-based ML models trained to discern patients with schizophrenia (SCH) from controls. Study design We trained an ML model on structural MRI data from public datasets to distinguish between SCH and controls (SCH = 347, controls = 341). We then investigated the model’s performance in two independent samples of individuals undergoing forensic psychiatric examination: sample 1 was used for sensitivity analysis to discern ASPD (N = 52) from SCH (N = 66), and sample 2 was used for specificity analysis to discern ASPD (N = 26) from controls (N = 25). Both samples included individuals with SUD. Study results In sample 1, 94.4% of SCH with comorbid ASPD and SUD were classified as SCH, followed by patients with SCH + SUD (78.8% classified as SCH) and patients with SCH (60.0% classified as SCH). The model failed to discern SCH without comorbidities from ASPD + SUD (AUC = 0.562, 95%CI = 0.400–0.723). In sample 2, the model’s specificity to predict controls was 84.0%. In both samples, about half of the ASPD + SUD were misclassified as SCH. Data-driven functional characterization revealed associations between the classification as SCH and cognition-related brain regions. Conclusion Altogether, ASPD and SUD appear to have effects on ML prediction performance, which potentially results from converging cognition-related brain abnormalities between SCH, ASPD, and SUD.

99 GENERAL AND MISCELLANEOUS↗

Machine learning predictions of near-surface permafrost extent at Teller 27, Teller 47, and the Kougarok 64 Hillslope sites on the Seward Peninsula, Alaska: Supporting Data

Geophysical surveys were conducted at the NGEE Arctic Teller mile marker 27 site, Teller mile marker 47 site, and Kougarok mile marker 64 site during the summers of 2018, 2019, and 2021. Additional data was collected at Teller mile marker 47 during September 2021 and August 2022. These surveys were used to identify locations of near-surface permafrost during the period of maximum seasonal thaw depth for ground truth data used in machine learning predictions of near-surface permafrost extent at each site. This dataset contains CSV files of ground truth observations of near-permafrost presence or absence for each site, where PF = 1 indicates permafrost presence and PF = 0 indicates permafrost absence. The dataset also includes 2 sets of binary rasters (WGS84 UTM zone 3) of permafrost extent for each site using 1) all of the training data and 2) the transferred model. For both sets of rasters, 0 = non-permafrost and 1 = permafrost. Included are 6 *.tif files and 5 *.csv files that include a data dictionary (dd.csv) and file-level metadata (flmd.csv). This dataset is in support of the paper "Machine learning-derived high-resolution maps of near-surface permafrost for three watersheds on the Seward Peninsula, Alaska" that is in review (May 2023). The Next-Generation Experiments: Arctic (NGEE Arctic), a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Machine learning predictions for local electronic properties of disordered correlated electron systems

We present a scalable machine learning (ML) model to predict local electronic properties such as on-site electron number and double occupation for disordered correlated electron systems. Our approach is based on the locality principle, or the nearsightedness nature, of many-electron systems, which means local electronic properties depend mainly on the immediate environment. A ML model is developed to encode this complex dependence of local quantities on the neighborhood. We demonstrate our approach using the square-lattice Anderson-Hubbard model, which is a paradigmatic system for studying the interplay between Mott transition and Anderson localization. We develop a lattice descriptor based on the group-theoretical method to represent the on-site random potentials within a finite region. The resultant feature variables are used as input to a multilayer fully connected neural network, which is trained from data sets of variational Monte Carlo (VMC) simulations on small systems. We show that the ML predictions agree reasonably well with the VMC data. Our work underscores the promising potential of ML methods for multiscale modeling of correlated electron systems.

36 MATERIALS SCIENCE↗

Machine learning prediction and experimental verification of Pt-modified nitride catalysts for ethanol reforming with reduced precious metal loading

Ethanol is the smallest molecule containing C–O, C–C, C–H, and O–H bonds present in biomass-derived oxygenates. The development of inexpensive and selective catalysts for ethanol reforming is important towards the renewable generation of hydrogen from biomass. Transition metal nitrides (TMN) are interesting catalyst support materials that can effectively reduce precious metal loading for the catalysis of ethanol and other oxygenates. Herein theoretical and experimental methods were used to probe platinum-modified molybdenum nitride (Pt/Mo 2 N) surfaces for ethanol reforming. Computations using density-functional theory and machine learning predicted monolayer Pt/Mo 2 N to be highly active and selective for ethanol reforming. Temperature-programmed desorption (TPD) experiments verified that ethanol primarily underwent decomposition on Mo 2 N, and the reaction pathway shifted to reforming on Pt/Mo 2 N surfaces. Additionally, high-resolution electron energy loss spectroscopy (HREELS) results further indicated that while Mo2N decomposed the ethoxy intermediate by cleaving C–C, C–O, and C–H bonds, Pt-modification preserved the C–O bond, resulting in ethanol reforming.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

From atomistic models to machine learning: Predictive design of nanocarbons under extreme conditions

The formation of technologically valuable nanocarbon structures under extreme conditions, such as those produced during high-explosive detonations, remains poorly understood but holds significant potential for the development of controlled synthesis pathways. While detonation shockwaves provide the high-pressure, high-temperature environment required for nanodiamond formation, subsequent cooling and decompression dictate whether the diamond phase is preserved or transformed into other nanocarbon structures. Here, in this study, we employ GPU-accelerated reactive molecular dynamics (ReaxFF) simulations to investigate the graphitization and structural remodeling of detonation nanodiamond under nonlinear quench and pressure-release trajectories. We further investigate how the initial nanodiamond morphology; cuboctahedral, octahedral, or hexagonal prism influences the resulting transformation products. Evolution of nanostructure, allotrope (via simulated x-ray diffraction), carbon hybridization, and ring statistics are tracked during a two-stage quench from 5000 K to 60 GPa. Rapid cooling combined with slow decompression optimizes cubic diamond retention, whereas slow cooling with rapid pressure release promotes surface-to-core graphitization, producing concentric sp 2 -hybridized layers and hollowed inner shells. Octahedral nanodiamonds evolve into carbon nano-onions, initially forming bucky diamonds that progressively transform into fully sp 2 -hybridized structures, while hexagonal prisms preferentially form parallel-stacked graphite layers resembling carbon dots. Transient hexagonal diamond (lonsdaleite) emerges as an interfacial phase, suggesting potential reversibility in the shock-induced graphite-to-diamond transformation pathway transformation route. To extend predictive capabilities, we trained machine learning (ML) regressors on over 10 5 node-hours of molecular dynamics (MD) trajectories. A multilayer perceptron (MLP) model reliably predicts the number of graphitized layers from temperature–pressure trajectories with a coefficient of determination (R 2 ) exceeding 0.90. This high predictive fidelity enables efficient, high-throughput mapping of the synthesis parameter space for optimized graphitization outcomes. Collectively, morphological control combined with optimized quench–decompression conditions promote the selective synthesis of nanocarbon allotropes. This work establishes a data-driven framework for the rational, a priori design of carbon nanomaterials for applications in energy storage, sensing, and biomedicine.

Detonation nanodiamond remodeling↗

Machine learning predictions of irradiation embrittlement in reactor pressure vessel steels

Abstract Irradiation increases the yield stress and embrittles light water reactor (LWR) pressure vessel steels. In this study, we demonstrate some of the potential benefits and risks of using machine learning models to predict irradiation hardening extrapolated to low flux, high fluence, extended life conditions. The machine learning training data included the Irradiation Variable for lower flux irradiations up to an intermediate fluence, plus the Belgian Reactor 2 and Advanced Test Reactor 1 for very high flux irradiations, up to very high fluence. Notably, the machine learning model predictions for the high fluence, intermediate flux Advanced Test Reactor 2 irradiations are superior to extrapolations of existing hardening models. The successful extrapolations showed that machine learning models are capable of capturing key intermediate flux effects at high fluence. Similar approaches, applied to expanded databases, could be used to predict hardening in LWRs under life-extension conditions.

36 MATERIALS SCIENCE↗

Subsurface Characterization and Machine Learning Predictions at Brady Hot Springs: Preprint

Subsurface data analysis, reservoir modeling, and machine learning (ML) techniques have been applied to the Brady Hot Springs (BHS) geothermal field in Nevada, USA to further characterize the subsurface and assist with optimizing reservoir management. Hundreds of reservoir simulations have been conducted in TETRAD-G and CMG STARS to explore different injection and production fluid flow rates and allocations and to develop a training data set for ML. This process included simulating the historical injection and production since 1979 and prediction of future performance through 2040. ML networks were created and trained using TensorFlow based on multilayer perceptron (MLP), long short-term memory (LSTM), and convolutional neural network (CNN) architectures. These networks took as input selected flow rates, injection temperatures, and historical field operation data and produced estimates of future production temperatures. This approach was first successfully tested on a simplified single fracture doublet system, followed by the application to the BHS reservoir. Using an initial BHS dataset with 37 simulated scenarios, the trained and validated network predicted the production temperature for 6 production wells with the mean absolute percentage error of less than 8%. In a complementary analysis effort, the principal component analysis applied to 13 BHS geological parameters revealed that vertical fracture permeability shows the strongest correlation with fault density and fault intersection density. A new BHS reservoir model was developed considering the fault intersection density as proxy for permeability. This new reservoir model helps to explore under-exploited zones in the reservoir. A data gathering plan to obtain additional subsurface data was developed; it includes temperature surveying for three idle injection wells, at which the reservoir simulations indicate high bottom-hole temperatures. The collected data assist with calibrating the reservoir model and may lead to converting these wells to producers to access under-exploited zones in the reservoir. Data gathering activities are planned for the first quarter of 2021.

40 EE - Geothermal Technologies Office (EE-4G)↗

TCR-H: explainable machine learning prediction of T-cell receptor epitope binding on unseen datasets

Artificial-intelligence and machine-learning (AI/ML) approaches to predicting T-cell receptor (TCR)-epitope specificity achieve high performance metrics on test datasets which include sequences that are also part of the training set but fail to generalize to test sets consisting of epitopes and TCRs that are absent from the training set, i.e., are ‘unseen’ during training of the ML model. We present TCR-H, a supervised classification Support Vector Machines model using physicochemical features trained on the largest dataset available to date using only experimentally validated non-binders as negative datapoints. TCR-H exhibits an area under the curve of the receiver-operator characteristic (AUC of ROC) of 0.87 for epitope ‘hard splitting’ (i.e., on test sets with all epitopes unseen during ML training), 0.92 for TCR hard splitting and 0.89 for ‘strict splitting’ in which neither the epitopes nor the TCRs in the test set are seen in the training data. Furthermore, we employ the SHAP (Shapley additive explanations) eXplainable AI (XAI) method for post hoc interrogation to interpret the models trained with different hard splits, shedding light on the key physiochemical features driving model predictions. TCR-H thus represents a significant step towards general applicability and explainability of epitope:TCR specificity prediction.

60 APPLIED LIFE SCIENCES↗

Optimization of simulated high-field side lower hybrid current drive coupling using machine learning predictions of scrape-off layer density

Lower hybrid current drive (LHCD) is a potential source of non-inductive off-axis current drive (CD) for tokamaks. Although LHCD has been successfully deployed on a number of tokamaks, it is highly sensitive to the scrape-off layer (SOL) conditions local to the LHCD launcher. Large gaps between the launcher and plasma core, SOL turbulence, or edge density perturbations due to edge-localized modes can hamper CD or cause large reflected power. These coupling issues in part motivated the installation of an LHCD launcher on the high-field side (HFS) of DIII-D. On the HFS, the SOL is less turbulent and more controllable compared to the low-field side. This quiescence may result in more predictable edge conditions and thus a more predictable CD. Here, in this work, HFS SOL reflectometry measurements are predicted from global plasma parameters using machine learning models. The SOL predictions coupled with the full-wave simulation of the LHCD launcher allow for the prediction of reflected power, directivity, and arcing risk before the discharge. Launcher performance is then optimized using multi-objective Bayesian optimization, finding the shot parameters that result in an optimal SOL density that maximizes CD while minimizing the risk of arcing. The predictions and optimizations of LHCD performance are then accelerated using a surrogate model of the full-wave LHCD simulation.

Bayesian optimization↗

Uncertainty Quantification of Machine Learning Predicted Creep Property of Alumina-Forming Austenitic Alloys

The development of machine learning (ML) approaches in materials science offers the opportunity to exploit existing engineering and developmental alloy datasets, such as Oak Ridge National Laboratory (ORNL)’s consistently measured creep-rupture dataset for alumina-forming austenitic (AFA) alloys, to accelerate their further development. As a first step toward achieving ML insights for improved alloy design, the potential sources of uncertainty and their impacts on ML output are examined. It is observed that the selection of algorithms and features as well as data sampling significantly affects the performance of ML models, either positively or negatively. Further, the performance of various ML models in predicting the creep properties of AFA alloys is compared, with further evaluation by assessment of a small set of new developmental AFA alloys that were not part of the training dataset. The present study demonstrates that uncertainty quantification (UQ) is essential in materials science for evaluating the performance of ML algorithms with specifically selected feature sets and obtaining a comprehensive understanding of their limitations and the resultant capability of effective prediction in complex materials systems.

36 MATERIALS SCIENCE↗

BatteryPro: A Python Toolkit for Battery Data Analysis and Machine Learning Predictions

Analyzing battery test data for research & development can be time-consuming since battery tests often run on the order of months to years, generating large volumes of data. BatteryPro is a comprehensive Python package and software designed to facilitate advanced analysis and performance predictions for battery test data. Developed for battery researchers, it supports data types from widely used battery testing instruments, including MACCOR and Biologic cycling systems. The software provides a variety of tools for extracting and plotting key battery parameters such as time, voltage, capacity, current, and pressure. In addition to its extensive data analysis capabilities, BatteryPro features a dedicated machine learning module that employs a Bayesian Gaussian Mixture Model (GMM) to predict battery performance and degradation. Users can generate synthetic capacity fade data, calculate fade metrics, and leverage predictive models to forecast long-term battery behavior. The software's graphical user interface (GUI) enhances usability, allowing researchers to upload, merge, and analyze multiple data files with full customizability. The GUI also supports machine learning predictions, enabling users to fit models and make predictions based on selected data and parameters. BatteryPro is built using QtDesigner, scikit-learn, matplotlib, and pandas, ensuring a high level of customization, flexibility, and accuracy in battery data analysis. This tool aims to empower researchers with the ability to perform detailed battery analysis and make informed predictions, ultimately advancing the field of battery research.

25 - ENERGY STORAGE↗