Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Machine learning prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Acoustic Energy Release During the Laboratory Seismic Cycle: Insights on Laboratory Earthquake Precursors and Prediction

Abstract Machine learning can predict the timing and magnitude of laboratory earthquakes using statistics of acoustic emissions. The evolution of acoustic energy is critical for lab earthquake prediction; however, the connections between acoustic energy and fault zone processes leading to failure are poorly understood. Here, we document in detail the temporal evolution of acoustic energy during the laboratory seismic cycle. We report on friction experiments for a range of shearing velocities, normal stresses, and granular particle sizes. Acoustic emission data are recorded continuously throughout shear using broadband piezo‐ceramic sensors. The coseismic acoustic energy release scales directly with stress drop and is consistent with concepts of frictional contact mechanics and time‐dependent fault healing. Experiments conducted with larger grains (10.5 μm) show that the temporal evolution of acoustic energy scales directly with fault slip rate. In particular, the acoustic energy is low when the fault is locked and increases to a maximum during coseismic failure. Data from traditional slide‐hold‐slide friction tests confirm that acoustic energy release is closely linked to fault slip rate. Furthermore, variations in the true contact area of fault zone particles play a key role in the generation of acoustic energy. Our data show that acoustic radiation is related primarily to breaking/sliding of frictional contact junctions, which suggests that machine learning‐based laboratory earthquake prediction derives from frictional weakening processes that begin very early in the seismic cycle and well before macroscopic failure.

58 GEOSCIENCES↗

Predicting boron coordination in multicomponent borate and borosilicate glasses using analytical models and machine learning

Accurate prediction of boron coordination in multicomponent glasses is critical in glass science and technology as it strongly affects the properties of borate and borosilicate glasses. We have collected a dataset containing 657 glasses from literature with boron coordination values and developed models using analytical functions based on the well accepted Dell, Xiao and Bray model. Good prediction of boron coordination with a R 2 value higher than 0.8 was obtained. The large variation of boron coordination from experiments, originated from sample preparations and characterizations, led to difficulties in obtaining models with better prediction performance. Various machine learning (ML) algorithms were evaluated and slightly better prediction performance was observed; however, interpretation of the ML models is less straight forward. In conclusion, this study developed various models capable of providing quantitative boron coordination predictions, providing insights into its structural roles in multi-component glasses, and suggesting fruitful areas for future research.

36 MATERIALS SCIENCE↗

Explainable machine learning for hydrogen diffusion in metals and random binary alloys

Hydrogen diffusion in metals and alloys plays an important role in the discovery of new materials for fuel cell and energy storage technology. While analytic models use hand-selected features that have clear physical ties to hydrogen diffusion, they often lack accuracy when making quantitative predictions. Machine learning models are capable of making accurate predictions, but their inner workings are obscured, rendering it unclear which physical features are truly important. To develop interpretable machine learning models to predict the activation energies of hydrogen diffusion in metals and random binary alloys, we create a database for physical and chemical properties of the species and use it to fit six machine learning models. Our models achieve root-mean-squared errors between 98–119 meV on the testing data and accurately predict that elemental Ru has a large activation energy, while elemental Cr and Fe have small activation energies. By analyzing the feature importances of these fitted models, we identify relevant physical properties for predicting hydrogen diffusivity. While metrics for measuring the individual feature importances for machine learning models exist, correlations between the features lead to disagreement between models and limit the conclusions that can be drawn. Instead grouped feature importance, formed by combining the features via their correlations, agree across the six models and reveal that the two groups containing the packing factor and electronic specific heat are particularly significant for predicting hydrogen diffusion in metals and random binary alloys. In conclusion, this framework allows us to interpret machine learning models and enables rapid screening of new materials with the desired rates of hydrogen diffusion.

36 MATERIALS SCIENCE↗

Explainability and extrapolation of machine learning models for predicting the glass transition temperature of polymers

Abstract Machine learning (ML) offers promising tools to develop surrogate models for polymers' structure–property relations. Surrogate models can be built upon existing polymer data and are useful for rapidly predicting the properties of unknown polymers. The accuracy of such ML models appears to depend on the feature space representation of polymers, the range of training data, and learning algorithms. Here, we establish connections between these factors for predicting the glass transition temperature (T g ) of polymers. Our analysis suggests linear models with fewer fitting parameters are as accurate as nonlinear models with many hidden and unexplainable parameters. Also, the performance of a monomer topology‐based ML model is found to be qualitatively identical to that of a physicochemical descriptor‐based ML model. We find that the ML models's performance in the extrapolative region is enhanced as the property range of the training data increases. Moreover, we establish newT g – polymer chemistry correlations via ML. Our work illustrates how ML can advance the fundamental understanding of polymer structure–property correlations and its efficacy for extrapolation problems.

Polymer Science↗

Improving streamflow predictions across CONUS by integrating advanced machine learning models and diverse data

Accurate streamflow prediction is crucial to understand climate impacts on water resources and develop effective adaption strategies. A global long short-term memory (LSTM) model, using data from multiple basins, can enhance streamflow prediction, yet acquiring detailed basin attributes remains a challenge. To overcome this, we introduce the Geo-vision transformer (ViT)-LSTM model, a novel approach that enriches LSTM predictions by integrating basin attributes derived from remote sensing with a ViT architecture. Applied to 531 basins across the Contiguous United States, our method demonstrated superior prediction accuracy in both temporal and spatiotemporal extrapolation scenarios. Geo-ViT-LSTM marks a significant advancement in land surface modeling, providing a more comprehensive and effective tool for better understanding the environment responses to climate change.

Tayal, Kshitij↗

Machine Learning Models for Predicting Molecular UV–Vis Spectra with Quantum Mechanical Properties

Accurate understanding of Ultraviolet–visible (UV–Vis) spectra is critical for highthroughput design of compounds for drug discovery. Experimentally determining UV–Vis spectra can become expensive when dealing with a large quantity of novel molecules. This provides us an opportunity to drive computational advances in molecular property predictions using quantum mechanics and machine learning. In this work, we use both Quantum Mechanically (QM) predicted and measured UV–Vis spectra as input to modify four different machine learning architectures: UVvis-SchNet, UVvis- DTNN, UVvis-Transformer, and UVvis-MPNN. Here we find that the UVvis-MPNN model outperforms the other models when using optimized 3D coordinates and QM predicted spectra as input features. This model has the highest performance for predicting UVVisible spectra with a training RMSE of 0.06 and validation RMSE of 0.08.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Using Ultrasound Image Augmentation and Ensemble Predictions to Prevent Machine-Learning Model Overfitting

Deep learning predictive models have the potential to simplify and automate medical imaging diagnostics by lowering the skill threshold for image interpretation. However, this requires predictive models that are generalized to handle subject variability as seen clinically. Here, we highlight methods to improve test accuracy of an image classifier model for shrapnel identification using tissue phantom image sets. Using a previously developed image classifier neural network—termed ShrapML—blind test accuracy was less than 70% and was variable depending on the training/test data setup, as determined by a leave one subject out (LOSO) holdout methodology. Introduction of affine transformations for image augmentation or MixUp methodologies to generate additional training sets improved model performance and overall accuracy improved to 75%. Further improvements were made by aggregating predictions across five LOSO holdouts. This was done by bagging confidences or predictions from all LOSOs or the top-3 LOSO confidence models for each image prediction. Top-3 LOSO confidence bagging performed best, with test accuracy improved to greater than 85% accuracy for two different blind tissue phantoms. This was confirmed by gradient-weighted class activation mapping to highlight that the image classifier was tracking shrapnel in the image sets. Overall, data augmentation and ensemble prediction approaches were suitable for creating more generalized predictive models for ultrasound image analysis, a critical step for real-time diagnostic deployment.

60 APPLIED LIFE SCIENCES↗

Understanding the role of segmentation on process-structure–property predictions made via machine learning

Here, the present study investigated the effect of porosity surface determination methods on performance of machine learning models used to predict the tensile properties of AlSi10Mg processed by laser powder bed fusion from micro-computed tomography data. Machine learning models applied in this work include support vector machines, neural networks, decision trees, and Bayesian classifiers. The effects of isosurface thresholding and local gradient approaches for porosity segmentation, as well as image filtering schemes, on model precision were evaluated for samples produced under differing levels of global energy density.

36 MATERIALS SCIENCE↗

Predicting Solar Plant Generation with Machine Learning Techniques

Accurately predicting power generation for PV sites is critical for prioritizing relevant operations & maintenance activities, thereby extending the lifetime of a system and improving profit margins. A number of factors influence power generation at PV sites, including local weather, shading and soiling losses, design of modules, DC mismatches, and degradation over time. Other external factors such as curtailment and grid outages can also have a notable impact on power generation. Machine learning techniques can be used to provide more accurate predictions of PV power production by accounting for important weather and climate information neglected by current industry methods. This article will cover the deficiencies of those methods and will show how machine learning can dramatically improve power generation predictions.

14 SOLAR ENERGY↗

Budget Constrained Machine Learning for Early Prediction of Adverse Outcomes for COVID-19 Patients

Background: Machine learning (ML) based risk stratification models of Electronic Health records (EHR) data may help to optimize treatment of COVID-19 patients, but are often limited by their lack of clinical interpretability and cost of laboratory tests. We develop a ML based tool for predicting adverse outcomes based on EHR data to optimize clinical utility under a given cost structure. This cohort study was performed using deidentified EHR data from COVID-19 patients from ProMedica Healthcare in northwest Ohio and southeastern Michigan. Methods: We tested performance of various ML approaches for predicting either increasing ventilatory support or mortality and the set of model features under a budget constraint was optimized via exhaustive search across all combinations of features. Results: The optimal sets of features for predicting ventilation under any budget constraint included demographics and comorbidities (DCM), basic metabolic panel (BMP), D-dimer, lactate dehydrogenase (LDH), erythrocyte sedimentation rate (ESR), CRP, brain natriuretic peptide (BNP), and procalcitonin and for mortality included DCM, BMP, complete blood count, D-dimer, LDH, CRP, BNP, procalcitonin and ferritin. Conclusions: This study presents a quick, accurate and cost-effective method to evaluate risk of deterioration for patients with SARS-CoV-2 infection at the time of clinical evaluation.

Cadena Pico, JoseE.↗

Machine learning and deep learning tools for the automated capture of cancer surveillance data

The National Cancer Institute and the Department of Energy strategic partnership applies advanced computing and predictive machine learning and deep learning models to automate the capture of information from unstructured clinical text for inclusion in cancer registries. Applications include extraction of key data elements from pathology reports, determination of whether a pathology or radiology report is related to cancer, extraction of relevant biomarker information, and identification of recurrence. With the growing complexity of cancer diagnosis and treatment, capturing essential information with purely manual methods is increasingly difficult. These new methods for applying advanced computational capabilities to automate data extraction represent an opportunity to close critical information gaps and create a nimble, flexible platform on which new information sources, such as genomics, can be added. This will ultimately provide a deeper understanding of the drivers of cancer and outcomes in the population and increase the timeliness of reporting. These advances will enable better understanding of how real-world patients are treated and the outcomes associated with those treatments in the context of our complex medical and social environment.

60 APPLIED LIFE SCIENCES↗

Machine-learned impurity level prediction for semiconductors: the example of Cd-based chalcogenides

The ability to predict the likelihood of impurity incorporation and their electronic energy levels in semiconductors is crucial for controlling its conductivity, and thus the semiconductor's performance in solar cells, photodiodes, and optoelectronics. The difficulty and expense of experimental and computational determination of impurity levels makes a data-driven machine learning approach appropriate. In this work, we show that a density functional theory-generated dataset of impurities in Cd-based chalcogenides CdTe, CdSe, and CdS can lead to accurate and generalizable predictive models of defect properties. By converting any semiconductor + impurity system into a set of numerical descriptors, regression models are developed for the impurity formation enthalpy and charge transition levels. These regression models can subsequently predict impurity properties in mixed anion CdX compounds (where X is a combination of Te, Se and S) fairly accurately, proving that although trained only on the end points, they are applicable to intermediate compositions. We make machine-learned predictions of the Fermi-level-dependent formation energies of hundreds of possible impurities in 5 chalcogenide compounds, and we suggest a list of impurities which can shift the equilibrium Fermi level in the semiconductor as determined by the dominant intrinsic defects. Machine learning predictions for the dominating impurities compare well with DFT predictions, revealing the power of machine-learned models in the quick screening of impurities likely to affect the optoelectronic behavior of semiconductors.

36 MATERIALS SCIENCE↗

Machine Learning-based Prediction of Departure from Nucleate Boiling Power for the PSBT Benchmark

Machine Learning (ML) has seen an exponential growth in its applications due to its advanced data driven prediction capabilities. The study presents a data-driven approach as a preliminary attempt to predict the power at which departure from nucleate boiling (DNB) occurs in pressurized water reactors (PWRs) by constructing an advanced ML algorithm that takes outlet pressure, inlet temperature and inlet mass flux as the input features. DNB is a critical heat flux (CHF) phenomenon seen in PWRs. The experimental data from the PWR subchannel and bundle tests (PSBT) benchmark is first used to train an artificial neural network (ANN) to predict the DNB power, which produces a root mean square error (RMSE) of 6.89 kW/m when tested on a blind subset of the PSBT data. Since the PSBT dataset is relatively small to train an accurate ANN, a data augmentation methodology based on generative adversarial networks (GANs) is used to expand the training dataset. By assuming that the real data follows a certain distribution, GANs try to learn that underlying distribution to generate similar synthetic data to augment the database and to improve the predictive capabilities of the ANN. The data generated from GANs are validated using 1-nearest neighbor and kernel maximum mean discrepancy. To further ensure data from GAN is similar to PSBT, the data is tested and filtered out using the sub-channel thermal-hydraulic code CTF. The results indicate that with the addition of 120 data points from GAN the RMSE reduces to 4.84 kW/m showing promising results for future developments.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A machine learning model for predicting the minimum miscibility pressure of CO 2 and crude oil system based on a support vector machine algorithm approach

CO 2 enhanced oil recovery (EOR) is a potential way for carbon capture, utilization and storage (CCUS). Though, the effect of CO 2 injection is greatly influenced by the reservoir conditions. Typically, Minimum miscible pressure (MMP) is selected as one of the key parameters for the screening and evaluation of prospective CO 2 flooding. Conventional slim tube test is both accurate and widely accepted but it is inefficient. Existing empirical formulas for MMPs are easy to be used but have been proved inaccurate and unreliable. Machine learning-based methods have great advantages in predicting MMP. However, only predication accuracy is discussed for most models without the screening of the main control factors and further validation of the model reliability. In this paper, a new prediction model based on support vector machine (SVM) was developed for pure/impure CO 2 and crude oil system. This study was based on 147 sets of MMP data from the literature with full information on reservoir temperature, oil composition and gas composition. The main control factors were screened by several statistical methods. Unlike the conventional prediction models that verified by only prediction accuracy, learning curve and single factor control variable analysis are further validated to obtain the optimum model.

02 PETROLEUM↗

Predicting wind farm operations with machine learning and the P2D‐RANS model: A case study for an AWAKEN site

Abstract The power performance and the wind velocity field of an onshore wind farm are predicted with machine learning models and the pseudo‐2D RANS model, then assessed against SCADA data. The wind farm under investigation is one of the sites involved with the American WAKE experimeNt (AWAKEN). The performed simulations enable predictions of the power capture at the farm and turbine levels while providing insights into the effects on power capture associated with wake interactions that operating upstream turbines induce, as well as the variability caused by atmospheric stability. The machine learning models show improved accuracy compared to the pseudo‐2D RANS model in the predictions of turbine power capture and farm power capture with roughly half the normalized error. The machine learning models also entail lower computational costs upon training. Further, the machine learning models provide predictions of the wind turbulence intensity at the turbine level for different wind and atmospheric conditions with very good accuracy, which is difficult to achieve through RANS modeling. Additionally, farm‐to‐farm interactions are noted, with adverse impacts on power predictions from both models.

17 WIND ENERGY↗

minervachem

Minervachem is a tool for cheminformatics and machine learning in chemistry. It includes both algorithms from existing literature and algorithms which we have developed. Its features center around two main themes: molecular representations and algorithms, both for molecular machine learning. It provides a scikit-learn transformer interface for molecular featurization and an estimator interface for machine learning algorithms. It also provides visualization tools for explaining machine learning predictions. We aim to continue developing this software to improve its performance and usability and expand its capabilities within the realm of molecular machine learning, cheminformatics, and visualization.

Lubbers, Nicholas↗