Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Machine learning prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Subsurface Characterization and Machine Learning Predictions at Brady Hot Springs Results

Geothermal power plants typically show decreasing heat and power production rates over time. Mitigation strategies include optimizing the management of existing wells - increasing or decreasing the fluid flow rates across the wells - and drilling new wells at appropriate locations. The latter is expensive, time-consuming, and subject to many engineering constraints, but the former is a viable mechanism for periodic adjustment of the available fluid allocations. Data and supporting literature from a study describing a new approach combining reservoir modeling and machine learning to produce models that enable strategies for the mitigation of decreased heat and power production rates over time for geothermal power plants. The computational approach used enables translation of sets of potential flow rates for the active wells into reservoir-wide estimates of produced energy and discovery of optimal flow allocations among the studied sets. In our computational experiments, we utilize collections of simulations for a specific reservoir (which capture subsurface characterization and realize history matching) along with machine learning models that predict temperature and pressure timeseries for production wells. We evaluate this approach using an "open-source" reservoir we have constructed that captures many of the characteristics of Brady Hot Springs, a commercially operational geothermal field in Nevada, USA. Selected results from a reservoir model of Brady Hot Springs itself are presented to show successful application to an existing system. In both cases, energy predictions prove to be highly accurate: all observed prediction errors do not exceed 3.68% for temperatures and 4.75% for pressures. In a cumulative energy estimation, we observe prediction errors that are less than 4.04%. A typical reservoir simulation for Brady Hot Springs completes in approximately 4 hours, whereas our machine learning models yield accurate 20-year predictions for temperatures, pressures, and produced energy in 0.9 seconds. This paper aims to demonstrate how the models and techniques from our study can be applied to achieve rapid exploration of controlled parameters and optimization of other geothermal reservoirs. Includes a synthetic, yet realistic, model of a geothermal reservoir, referred to as open-source reservoir (OSR). OSR is a 10-well (4 injection wells and 6 production wells) system that resembles Brady Hot Springs (a commercially operational geothermal field in Nevada, USA) at a high level but has a number of sufficiently modified characteristics (which renders any possible similarity between specific characteristics like temperatures and pressures as purely random). We study OSR through CMG simulations with a wide range of flow allocation scenarios. Includes a dataset with 101 simulated scenarios that cover the period of time between 2020 and 2040 and a link to the published paper about this project, where we focus on the Machine Learning work for predicting OSR's energy production based on the simulation data, as well as a link to the GitHub repository where we have published the code we have developed (please refer to the repository's readme file to see instructions on how to run the code). Additional links are included to associated work led by the USGS to identify geologic factors associated with well productivity in geothermal fields. Below are the high-level steps for applying the same modeling + ML process to other geothermal reservoirs: 1. Develop a geologic model of the geothermal field. The location of faults, upflow zones, aquifers, etc. need to be accounted for as accurately as possible 2. The geologic model needs to be converted to a reservoir model that can be used in a reservoir simulator, such as, for instance, CMG STARS, TETRAD, or FALCON 3. Using native state modeling, the initial temperature and pressure distributions are evaluated, and they become the initial conditions for dynamic reservoir simulations 4....

15 GEOTHERMAL ENERGY↗

Analyzing Machine Learning Predictions of Passive Microwave Brightness Temperature Spectral Difference Over Snow-Covered Terrain in High Mountain Asia

Snow is an important component of the terrestrial freshwater budget in high mountainAsia (HMA) and contributes to the runoff in Himalayan rivers through snowmelt. Despitethe importance of snow in HMA, considerable spatiotemporal uncertainty exists across the different estimates of snow water equivalent for this region. In order to better estimate snow water equivalent, radiative transfer models are often used in conjunction with microwave brightness temperature measurements. In this study, the efficacy of support vector machines (SVMs), a machine learning technique, to predict passive microwave brightness temperature spectral difference (1Tb) as a function of geophysical variables (snow water equivalent, snow depth, snow temperature, and snow density) is explored through a sensitivity analysis. The use of machine learning (as opposed to radiative transfer models) is a relatively new and novel approach for improving snow water equivalent estimates. The Noah-MP land surface model within the NASALand Information System framework is used to simulate the hydrologic cycle over HMA and model geophysical variables that are then used for SVM training. The SVMsserve as a nonlinear map between the geophysical space (modeled in Noah-MP) andthe observation space (1Tb as measured by the radiometer). Advanced MicrowaveScanning Radiometer-Earth Observing System measured passive microwave brightness temperatures over snow-covered locations in the HMA region are used as training data during the SVM training phase. Sensitivity of well-trained SVMs to each Noah-MP modeled state variable is assessed by computing normalized sensitivity coefficients. Sensitivity analysis results generally conform with the known first-order physics. Input states that increase volume scattering of microwave radiation, such as snow density and snow water equivalent, exhibit a plurality of positive normalized sensitivity coefficients. In general, snow temperature was the most sensitive input to the SVM predictions. The sensitivity of each state is location and time dependent. The signs of normalized sensitivity coefficients that indicate physical irrationality are ascribed to significant cross-correlation between Noah-MP simulated states and decreased SVM prediction capability at specific locations due to insufficient training data. SVM prediction pitfalls do exist that serve to highlight the limitations of this particular machine learning algorithm.

high mountain Asia↗

Machine Learning Predicts the Timing and Shear Stress Evolution of Lab Earthquakes Using Active Seismic Monitoring of Fault Zone Processes

Abstract Machine learning (ML) techniques have become increasingly important in seismology and earthquake science. Lab‐based studies have used acoustic emission data to predict time‐to‐failure and stress state, and in a few cases, the same approach has been used for field data. However, the underlying physical mechanisms that allow lab earthquake prediction and seismic forecasting remain poorly resolved. Here, we address this knowledge gap by coupling active‐source seismic data, which probe asperity‐scale processes, with ML methods. We show that elastic waves passing through the lab fault zone contain information that can predict the full spectrum of labquakes from slow slip instabilities to highly aperiodic events. The ML methods utilize systematic changes in P‐wave amplitude and velocity to accurately predict the timing and shear stress during labquakes. The ML predictions improve in accuracy closer to fault failure, demonstrating that the predictive power of the ultrasonic signals improves as the fault approaches failure. Our results demonstrate that the relationship between the ultrasonic parameters and fault slip rate, and in turn, the systematically evolving real area of contact and asperity stiffness allow the gradient boosting algorithm to “learn” about the state of the fault and its proximity to failure. Broadly, our results demonstrate the utility of physics‐informed ML in forecasting the imminence of fault slip at the laboratory scale, which may have important implications for earthquake mechanics in nature.

58 GEOSCIENCES↗

Fast and accurate machine learning prediction of phonon scattering rates and lattice thermal conductivity

Abstract Lattice thermal conductivity is important for many applications, but experimental measurements or first principles calculations including three-phonon and four-phonon scattering are expensive or even unaffordable. Machine learning approaches that can achieve similar accuracy have been a long-standing open question. Despite recent progress, machine learning models using structural information as descriptors fall short of experimental or first principles accuracy. This study presents a machine learning approach that predicts phonon scattering rates and thermal conductivity with experimental and first principles accuracy. The success of our approach is enabled by mitigating computational challenges associated with the high skewness of phonon scattering rates and their complex contributions to the total thermal resistance. Transfer learning between different orders of phonon scattering can further improve the model performance. Our surrogates offer up to two orders of magnitude acceleration compared to first principles calculations and would enable large-scale thermal transport informatics.

36 MATERIALS SCIENCE↗

Machine Learning Prediction of Tritium‐Helium Groundwater Ages in the Central Valley, California, USA

Abstract Groundwater ages provides insight into recharge rates, flow velocities, and vulnerability to contaminants. The ability to predict groundwater ages based on more accessible parameters via Machine Learning (ML) would advance our ability to guide sustainable management of groundwater resources. In this study, ML models were trained and tested on a large data set of tritium concentrations and tritium‐helium groundwater ages from the California Central Valley, a large groundwater basin with complex land use, irrigation, and water management practices. The ML models were trained on 63 features, including location, well construction information, landscape characteristics, and climate variables, water chemistry, and stable isotopes. The Bagging regressor method can accurately classify (F1‐score = 0.91) groundwater samples as either modern or pre‐modern whereas the accuracy of the ML prediction of continuous tritium‐helium groundwater ages is limited and explains only of the variability in this data set. In general, ML groundwater age prediction relies mostly on features related to (a) the source of groundwater recharge, (b) contaminant history, (c) aquifer materials, (d) well construction, and (e) geochemical reactions along flow paths.

54 ENVIRONMENTAL SCIENCES↗

Machine learning prediction of glass transition temperature of conjugated polymers from chemical structure

Predicting the glass transition temperature (T g ) is of critical importance as it governs the thermomechanical performance of conjugated polymers (CPs). Here, we report a predictive modeling framework to predict T g of CPs through the integration of machine learning (ML), molecular dynamics (MD) simulations, and experiments. With 154 T g data collected, an ML model is developed by taking simplified “geometry” of six chemical building blocks as molecular features, where side-chain fraction, isolated rings, fused rings, and bridged rings features are identified as the dominant ones for T g . MD simulations further unravel the fundamental roles of those chemical building blocks in dynamical heterogeneity and local mobility of CPs at a molecular level. The developed ML model is demonstrated for its capability of predicting T g of several new high-performance solar cell materials to a good approximation. The established predictive framework facilitates the design and prediction of T g of complex CPs, paving the way for addressing device stability issues that have hampered the field from developing stable organic electronics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Machine learning predictions of diffusion in bulk and confined ionic liquids using simple descriptors

Ionic liquids have many intriguing properties and widespread applications such as separations and energy storage. However, ionic liquids are complex fluids and predicting their behavior is difficult, particularly in confined environments. We introduce fast and computationally efficient machine learning (ML) models that can predict diffusion coefficients and ionic conductivity of bulk and nanoconfined ionic liquids over a wide temperature range (350–500 K). The ML models are trained on molecular dynamics simulation data for 29 unique ionic liquids as bulk fluids and confined in graphite slit pores. This model is based on simple physical descriptors of the cations and anions such as molecular weight and surface area. Here, we also demonstrate that accurate results can be obtained using only descriptors derived from SMILES (simplified molecular-input line-entry system) codes for the ions with minimal computational effort. This offers a fast and efficient method for estimating diffusion and conductivity of nanoconfined ionic liquids at various temperatures without the need for expensive molecular dynamics simulations.

74 ATOMIC AND MOLECULAR PHYSICS↗

An investigation on machine learning predictive accuracy improvement and uncertainty reduction using VAE-based data augmentation

The confluence of ultrafast computers with large memory, rapid progress in Machine Learning (ML) algorithms, and the availability of large datasets place multiple engineering fields at the threshold of dramatic progress. However, a unique challenge in nuclear engineering is data scarcity because experimentation on nuclear systems is usually more expensive and time-consuming than most other disciplines. One potential way to resolve the data scarcity issue is deep generative learning, which uses certain ML models to learn the underlying distribution of existing data and generate synthetic samples that resemble the real data. In this way, one can significantly expand the dataset to train more accurate predictive ML models. In this study, our objective is to evaluate the effectiveness of data augmentation using variational autoencoder (VAE)-based deep generative models. We investigated whether the data augmentation leads to improved accuracy in the predictions of a deep neural network (DNN) model trained using the augmented data. Additionally, the DNN prediction uncertainties are quantified using Bayesian Neural Networks (BNN) and conformal prediction (CP) to assess the impact on predictive uncertainty reduction. To test the proposed methodology, we used TRACE simulations of steady-state void fraction data based on the NUPEC Boiling Water Reactor Full-size Fine-mesh Bundle Test (BFBT) benchmark. Here, we found that augmenting the training dataset using VAEs has improved the DNN model’s predictive accuracy, improved the prediction confidence intervals, and reduced the prediction uncertainties.

Bayesian neural network↗

Machine learning prediction of enzyme optimum pH

The relationship between pH and enzyme catalytic activity, especially the optimal pH (pH opt ) at which enzymes function, is critical for biotechnological applications. Hence, computational methods to predict pH opt will enhance enzyme discovery and design by facilitating accurate identification of enzymes that function optimally at specific pH levels, and by elucidating sequence-function relationships. Here, in this study, we proposed and evaluated various machine learning methods for predicting pH opt , conducting extensive hyperparameter optimization and training over 11,000 model instances. Our results demonstrate that models utilizing language model embeddings markedly outperform other methods in predicting pHopt. We present EpHod, the best-performing model, to predict pHopt, making it publicly available to researchers. From sequence data, EpHod directly learns structural and biophysical features that relate to pH opt , including proximity of residues to the catalytic centre and the accessibility of solvent molecules. Overall, EpHod presents a promising advancement in pH opt prediction and will potentially speed up the development of enzyme technologies.

97 MATHEMATICS AND COMPUTING↗

A Science Gateway for the Repeatable Analysis of Machine Learning Predicted Gravity Anomalies

In recent years, deep learning has become an increasingly popular alternative for modeling in geoscience applications due to its scalability and efficiency. However, the interpretability, compute, data volume, and hyperparameter tuning requirements of deep learning models make development and monitoring difficult. Furthermore, model explainability and communicating results obtained by these models to users or domain experts is a challenge, as domain experts in geoscience also need to have a deep understanding of how those models function in order to support their scientific works. Here, we describe a science gateway and machine learning pipeline for predicting gravity anomalies from geophysical data. The gateway, built on open-source technologies, provides a holistic view of the pipeline through interactive visualizations aimed at enabling efficient exploratory data analysis. The repeatability, reproducibility, and monitoring capabilities of this overall system allow us to iterate and analyze at scale. Using this pipeline and gateway, we can repeatedly produce accurate high-resolution gravity anomaly datasets. By describing the underlying technologies, implementation, and results, here we provide a foundation for the broader adoption of science gateways into cross-cutting geoscience and machine learning research projects as a means to improve the scientific discovery and collaboration in the geophysics and computational sciences community.

58 GEOSCIENCES↗

Machine Learning Predictions of Transition Probabilities in Atomic Spectra

Forward modeling of optical spectra with absolute radiometric intensities requires knowledge of the individual transition probabilities for every transition in the spectrum. In many cases, these transition probabilities, or Einstein A-coefficients, quickly become practically impossible to obtain through either theoretical or experimental methods. Complicated electronic orbitals with higher order effects will reduce the accuracy of theoretical models. Experimental measurements can be prohibitively expensive and are rarely comprehensive due to physical constraints and sheer volume of required measurements. Due to these limitations, spectral predictions for many element transitions are not attainable. In this work, we investigate the efficacy of using machine learning models, specifically fully connected neural networks (FCNN), to predict Einstein A-coefficients using data from the NIST Atomic Spectra Database. For simple elements where closed form quantum calculations are possible, the data-driven modeling workflow performs well but can still have lower precision than theoretical calculations. For more complicated nuclei, deep learning emerged more comparable to theoretical predictions, such as Hartree–Fock. Unlike experiment or theory, the deep learning approach scales favorably with the number of transitions in a spectrum, especially if the transition probabilities are distributed across a wide range of values. It is also capable of being trained on both theoretical and experimental values simultaneously. In addition, the model performance improves when training on multiple elements prior to testing. The scalability of the machine learning approach makes it a potentially promising technique for estimating transition probabilities in previously inaccessible regions of the spectral and thermal domains on a significantly reduced timeline.

74 ATOMIC AND MOLECULAR PHYSICS↗

The impact of curation errors in the PDBBind Database on machine learning predictions of protein–protein binding affinity

The PDBBind database has been widely utilized for the computational prediction of protein–protein binding affinities. While the accuracy of the PDBBind-curated equilibrium dissociation constants (K D ) has been reported for the protein–ligand subset of the PDBBind database, the curation accuracy has not been reported for the protein–protein subset. Here, we present a detailed manual analysis for the subset of PDBBind records with PubMed Central Open Access primary publications and find that ~19% of these records had K D values that were not supported by their primary publications. The impact of these putative curation errors on the machine learning-based prediction of K D from experimental protein–protein 3D structures was evaluated and correcting the curation errors improved the Pearson correlation coefficient between measured and random forest-predicted log 10 (K D ) values by ~8 percentage points. This finding underscores the importance of dataset accuracy for computational modelling and highlights the need for more stringent curation processes when extracting information from the scientific literature.

59 BASIC BIOLOGICAL SCIENCES↗

Uncertainty bounds for multivariate machine learning predictions on high-strain brittle fracture

Simulation of the crack network evolution on high strain rate impact experiments performed in brittle materials is very compute-intensive. The cost increases even more if multiple simulations are needed to account for the randomness in crack length, location, and orientation, which is inherently found in real-world materials. Constructing a machine learning emulator can make the process faster by orders of magnitude. There has been little work, however, on assessing the error associated with their predictions. Estimating these errors is imperative for meaningful overall uncertainty quantification. In this work, we extend the heteroscedastic uncertainty estimates to bound a multiple output machine learning emulator. Overall, we find that the response prediction is accurate within its predicted errors, but with a somewhat conservative estimate of uncertainty.

36 MATERIALS SCIENCE↗

Physics-Informed Machine-Learning Prediction of Curie Temperatures and Its Promise for Guiding the Discovery of Functional Magnetic Materials

High-performance permanent magnets with a high Curie temperature, containing less critical materials, are integral to zero-carbon energy solutions. We built a machine-learning model trained over available experimentally measured Curie temperature values to predict the T C of multicomponent magnetic materials. We chose two compositions from a pseudo-binary (Zr 1–x Ce x )Fe 2 system, namely, (Zr 0.16 Ce 0.84 )Fe 2 and (Zr 0.94 Ce 0.06 )Fe 2 , to experimentally validate the ability of our model to predict the Curie temperature of novel compounds. We also provided a detailed discussion on the correlation of the Curie temperature with the de Gennes scaling factor in rare-earth intermetallic compounds and its breakdown below a certain rare-earth content. The electronic structure calculations (density of states and Fermi surface) were performed using the density functional theory on selected compounds (Zr 0.16 Ce 0.84 )Fe 2 and (Zr 0.94 Ce 0.06 )Fe 2 to understand the electronic origin of a strong magnetic exchange. We found that the change in the electronic density of states and electron/hole fillings at the Fermi level directly correlate with the Curie temperature. Notably, our model was able to capture these key electronic structure trends, which show that physics-informed machine learning can play a crucial role in designing new high-performance magnets with improved properties for environmentally sustainable applications.

36 MATERIALS SCIENCE↗

Using Computationally-Determined Properties for Machine Learning Prediction of Self-Diffusion Coefficients in Pure Liquids

The ability to predict transport properties of liquids quickly and accurately will greatly improve our understanding of fluid properties both in bulk and complex mixtures, as well as in confined environments. Such information could then be used in the design of materials and processes for applications ranging from energy production and storage to manufacturing processes. As a first step, we consider the use of machine learning (ML) methods to predict the diffusion properties of pure liquids. Recent results have shown that Artificial Neural Networks (ANNs) can effectively predict the diffusion of pure compounds based on the use of experimental properties as the model inputs. In the current study, a similar ANN approach is applied to modeling diffusion of pure liquids using fluid properties obtained exclusively from molecular simulations. A diverse set of 102 pure liquids is considered, ranging from small polar molecules (e.g., water) to large nonpolar molecules (e.g., octane). Self-diffusion coefficients were obtained from classical molecular dynamics (MD) simulations. Since nearly all the molecules are organic compounds, a general set of force field parameters for organic molecules was used. The MD methods are validated by comparing physical and thermodynamic properties with experiment. Computational input features for the ANN include physical properties obtained from the MD simulations as well as molecular properties from quantum calculations of individual molecules. Furthermore, fluid properties describing the local liquid structure were obtained from center of mass radial distribution functions (COM-RDFs). Feature sensitivity analysis revealed that isothermal compressibility, heat of vaporization, and the thermal expansion coefficient were the most impactful properties used as input for the ANN model to predict the MD simulated self-diffusion coefficients. The MD-based ANN successfully predicts the MD self-diffusion coefficients with only a subset (2 to 3) of the available computationally determined input features required. A separate ANN model was developed using literature experimental self-diffusion coefficients as model targets. Although this second ML model was not as successful due to a limited number of data points, a good correlation is still observed between experimental and ML predicted self-diffusion coefficients.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Interpreting machine learning prediction of fire emissions and comparison with FireMIP process-based models

Annual burned areas in the United States have increased 2-fold during the past decades. With more large fires resulting in more emissions of fine particulate matter, an accurate prediction of fire emissions is critical for quantifying the impacts of fires on air quality, human health, and climate. This study aims to construct a machine learning (ML) model with game-theory interpretation to predict monthly fire emissions over the contiguous US (CONUS) and to understand the controlling factors of fire emissions. The optimized ML model is used to diagnose the process-based models in the Fire Modeling Intercomparison Project (FireMIP) to inform future development. Results show promising performance for the ML model, Community Land Model (CLM), and Joint UK Land Environment Simulator-Interactive Fire And Emission Algorithm For Natural Environments (JULES-INFERNO) in reproducing the spatial distributions, seasonality, and interannual variability of fire emissions over the CONUS. Regional analysis shows that only the ML model and CLM simulate the realistic interannual variability of fire emissions for most of the subregions (r >0.95 for ML and r =0.14~0.70 for CLM), except for Mediterranean California, where all the models perform poorly (r =0.74 for ML and r <0.30 for the FireMIP models). Regarding seasonality, most models capture the peak emission in July over the western US. However, all models except for the ML model fail to reproduce the bimodal peaks in July and October over Mediterranean California, which may be explained by the smaller wind speeds of the atmospheric forcing data during Santa Ana wind events and limitations in model parameterizations for capturing the effects of Santa Ana winds on fire activity. Furthermore, most models struggle to capture the spring peak in emissions in the southeastern US, probably due to underrepresentation of human effects and the influences of winter dryness on fires in the models. As for extreme events, both the ML model and CLM successfully reproduce the frequency map of extreme emission occurrence but overestimate the number of months with extremely large fire emissions. Comparing the fire PM 2.5 emissions from the ML model with process-based fire models highlights their strengths and uncertainties for regional analysis and prediction and provides useful insights into future directions for model improvements.

54 ENVIRONMENTAL SCIENCES↗

Machine learning predictions of high-Curie-temperature materials

Technologies that function at room temperature often require magnets with a high Curie temperature, $T$ C , and can be improved with better materials. Discovering magnetic materials with a substantial $T$ C is challenging because of the large number of candidates and the cost of fabricating and testing them. Using the two largest known datasets of experimental Curie temperatures, we develop machine-learning models to make rapid $T$ C predictions solely based on the chemical composition of a material. We train a random-forest model and a k -NN one and predict on an initial dataset of over 2500 materials and then validate the model on a new dataset containing over 3000 entries. The accuracy is compared for multiple compounds' representations (“descriptors”) and regression approaches. A random-forest model provides the most accurate predictions and is not improved by dimensionality reduction or by using more complex descriptors based on atomic properties. Further, a random-forest model trained on a combination of both datasets shows that cobalt-rich and iron-rich materials have the highest Curie temperatures for all binary and ternary compounds. An analysis of the model reveals systematic error that causes the model to over-predict low-$T$ C materials and under-predict high-$T$ C materials. For exhaustive searches to find new high-$T$ C materials, analysis of the learning rate suggests either that much more data is needed or that more efficient descriptors are necessary.

36 MATERIALS SCIENCE↗

Can machine learning predict fuel properties accurately?

High-potential molecules derived from biomass sources may suitably replace or supplement traditional nonrenewable hydrocarbon fuels to reduce pollution and fuel processing cost. Experimental property testing of these bioproducts is usually conducted years after initial bench-scale experiments, due to high experimental costs and/or high volume requirements. However, neglecting to conduct property testing early in the pathway development cycle can lead to investments spent on scaling-up production of bioproducts and biofuels that do not perform as expected. Instead, machine-learning techniques can be used to develop quantitative structure–property relationships for molecules using a relatively large training set of molecular descriptor data. For this study, we compiled measured properties, IR spectra, and molecular descriptors of bio-based molecules from databases and published studies for training models of bioproduct properties. We trained regression models with molecular descriptors and will compare results of different estimators. This study describes the first steps towards a performance prediction tool for bio-based alternative fuels. Keywords: Machine learning, biofuels, jet fuels, fuel properties

Mayer, Morgan A.↗