Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Machine learning prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Machine learning and ligand binding predictions: A review of data, methods, and obstacles

We report that computational predictions of ligand binding is a difficult problem, with more accurate methods being extremely computationally expensive. The use of machine learning for drug binding predictions could possibly leverage the use of biomedical big data in exchange for time-intensive simulations. This paper reviews current trends in the use of machine learning for drug binding predictions, data sources to develop machine learning algorithms, and potential problems that may lead to overfitting and ungeneralizable models. A few popular datasets that can be used to develop virtual high-throughput screening models are characterized using spatial statistics to quantify potential biases. We can see from evaluating some common benchmarks that good performance correlates with models with high-predicted bias scores and models with low bias scores do not have much predictive power. A better understanding of the limits of available data sources and how to fix them will lead to more generalizable models that will lead to novel drug discovery.

59 BASIC BIOLOGICAL SCIENCES↗

A tutorial review of machine learning-based model predictive control methods

Abstract This tutorial review provides a comprehensive overview of machine learning (ML)-based model predictive control (MPC) methods, covering both theoretical and practical aspects. It provides a theoretical analysis of closed-loop stability based on the generalization error of ML models and addresses practical challenges such as data scarcity, data quality, the curse of dimensionality, model uncertainty, computational efficiency, and safety from both modeling and control perspectives. The application of these methods is demonstrated using a nonlinear chemical process example, with open-source code available on GitHub. The paper concludes with a discussion on future research directions in ML-based MPC.

Wu, Zhe [Department of Chemical and Biomolecular E↗

Physics-informed machine learning modeling for predictive control using noisy data

Due to the occurrence of over-fitting at the learning phase, the modeling of chemical processes via artificial neural networks (ANN) by using corrupted data (i.e., noisy data) is an ongoing challenge. Therefore, this work investigates the effect of both Gaussian and non-Gaussian noise on the performance of process-structure based recurrent neural networks (RNN) models, which take the form of partially-connected RNN models in this work, that are used to approximate a class of multi-input-multi-outputs nonlinear systems. Furthermore, two different techniques, specifically Monte Carlo dropout and co-teaching, are utilized in the development of partially-connected RNN models. Here, these two techniques are employed to reduce the over-fitting in ANNs when noisy data is used in the training process and, hence, to improve the open-loop accuracy as well as the closed-loop performance under a Lyapunov-based model predictive controller (MPC). Aspen Plus Dynamics, a well-known high-fidelity process simulator, is used to simulate a large-scale chemical process application in order to demonstrate the anticipated improvements in both open-loop approximation and closed-loop controller performance in the presence of Gaussian and non-Gaussian noise in the data set using physics-informed RNNs.

97 MATHEMATICS AND COMPUTING↗

The Use of Machine Learning Models for Predicting the Dielectric Strength of Gases

Technological advancements in high voltage systems have pushed sulfur hexafluoride (SF6) to its operational limits. Furthermore, this gas has other drawbacks including a high liquefaction temperature and a high global warming potential. Therefore, there has been an urgent need to find alternative gases with high dielectric strength (DS). In this work, density functional theory (DFT) is used to calculate molecular descriptors that are fed into an artificial neural network (ANN) and a random forest (RF). These machine learning (ML) models are then used to predict the DS for hundreds of molecules. A finite element model (FEM) is also used to calculate the electric field profile of multiple simple electrode geometries as the applied voltage to the system is increased. Results indicate that the random forest model has better generalization to unseen data than the neural network. The highest DS value predicted by the RF was 2.16 relative to the experimental DS of SF6. The results also demonstrate how choosing a gas with a higher DS and a geometry with minimal edges and corners can significantly increase the operating voltage of an electrical system. Due to its superior generalization, the RF represents the most promising path toward an accurate DS predictor once sufficient experimental data are available.

Mileski, Matthew [AFIT]↗

Machine Learning-Based Prediction of Distribution Network Voltage and Sensor Allocation

Increasing penetration levels of fast-varying energy resources might negatively affect power system operation. At the same time, sensor deployment throughout distribution networks improves system awareness and enables the development of new and advanced voltage control solutions. Such control techniques rely on accurate prediction in anticipation of voltage violation scenarios. This paper analyzes various approaches to voltage prediction in a distribution system, and it is shown that combining multiple techniques into a single regressor improves its predictive power. Moreover, a two-step regressor is proposed in which initial predictions based on a global regressor are refined by local regressors; in this case, prediction errors decrease significantly. Additionally, a clustering approach is employed to perform sensor allocation so that only the most influential buses are selected for monitoring without diminishing prediction accuracy.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Machine Learning-Based Prediction of Distribution Network Voltage and Sensor Allocation

Increasing penetration of fast-varying energy resources may negatively affect power systems' operation. At the same time, sensor deployment throughout distribution networks improves system awareness and enables the development of new and advanced voltage control solutions. Such control techniques rely on accurate prediction in anticipation to voltage violation scenarios. This paper analyzes various approaches for voltage prediction in a distribution system; it is shown that combining multiple techniques into a single regressor improves its predictive power. Moreover, a two-step regressor is proposed, where initial predictions based on a global regressor are refined by local regressors; in this case, prediction errors decrease significantly. Additionally, a clustering approach is employed for performing sensor allocation, so that only the most influential buses are selected for monitoring without diminishing prediction accuracy.

61 RADIATION PROTECTION AND DOSIMETRY↗

Machine Learning-Based Prediction of Distribution Network Voltage and Sensors Allocation: Preprint

Increasing penetration of fast-varying energy resources may negatively affect power systems' operation. At the same time, sensor deployment throughout distribution networks improves system awareness and enables the development of new and advanced voltage control solutions. Such control techniques rely on accurate prediction in anticipation to voltage violation scenarios. This paper analyzes various approaches for voltage prediction in a distribution system; it is shown that combining multiple techniques into a single regressor improves its predictive power. Moreover, a two-step regressor is proposed, where initial predictions based on a global regressor are refined by local regressors; in this case, prediction errors decrease significantly. Additionally, a clustering approach is employed for performing sensor allocation, so that only the most influential buses are selected for monitoring without diminishing prediction accuracy.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Machine learning models inaccurately predict current and future high-latitude C balances

The high-latitude carbon (C) cycle is a key feedback to the global climate system, yet because of system complexity and data limitations, there is currently disagreement over whether the region is a source or sink of C. Recent advances in big data analytics and computing power have popularized the use of machine learning (ML) algorithms to upscale site measurements of ecosystem processes, and in some cases forecast the response of these processes to climate change. Due to data limitations, however, ML model predictions of these processes are almost never validated with independent datasets. To better understand and characterize the limitations of these methods, we develop an approach to independently evaluate ML upscaling and forecasting. We mimic data-driven upscaling and forecasting efforts by applying ML algorithms to different subsets of regional process-model simulation gridcells, and then test ML performance using the remaining gridcells. In this study, we simulate C fluxes and environmental data across Alaska using ecosys, a process-rich terrestrial ecosystem model, and then apply boosted regression tree ML algorithms to training data configurations that mirror and expand upon existing AmeriFLUX eddy-covariance data availability. We first show that a ML model trained using ecosys outputs from currently-available Alaska AmeriFLUX sites incorrectly predicts that Alaska is presently a modeled net C source. Increased spatial coverage of the training dataset improves ML predictions, halving the bias when 240 modeled sites are used instead of 15. However, even this more accurate ML model incorrectly predicts Alaska C fluxes under 21st century climate change because of changes in atmospheric CO 2 , litter inputs, and vegetation composition that have impacts on C fluxes which cannot be inferred from the training data. Our results provide key insights to future C flux upscaling efforts and expose the potential for inaccurate ML upscaling and forecasting of high-latitude C cycle dynamics.

54 ENVIRONMENTAL SCIENCES↗

MACHINE LEARNING-ENABLED PREDICTION OF TRANSIENT INJECTION MAP IN AUTOMOTIVE INJECTORS WITH UNCERTAINTY QUANTIFICATION

Accurate prediction of injection profiles is a critical aspect of linking injector operation with engine performance and emissions. However, highly resolved injector simulations can take one to two weeks of wall-clock time, which is incompatible with engine design cycles with desired turnaround times of less than a day. Hence, it is important to reduce the time-to-solution of the internal flow simulations by several orders of magnitude to make it compatible with engine simulations. This work demonstrates a data-driven approach for tackling the computational overhead of injector simulations, whereby the transient injection profiles are emulated for a side-oriented, single-hole diesel injector using a Bayesian machine-learning framework. First, an interpretable Bayesian learning strategy was employed to understand the effect of design parameters on the total void fraction field. Then, autoencoders are utilized for efficient dimensionality reduction of the flowfields. Gaussian process models are finally used to predict the spatiotemporal void fraction field at the injector exit for unknown operating conditions. The Gaussian process models produce principled uncertainty estimates associated with the emulated flowfields, which provide the engine designer with valuable information of where the data-driven predictions can be trusted in the design space. The Bayesian flowfield predictions are compared with the corresponding predictions from a deep neural network, which has been transfer-learned from static needle simulations from a previous work by the authors. The emulation framework can predict the void fraction field at the exit of the orifice within a few seconds, thus achieving a speed-up factor of up to 38 x 10(6) over the traditional simulation-based approach of generating transient injection maps.

machine learning↗

Machine learning pipeline to predict defect behavior in metallic alloy systems

The interaction between defect and solute atoms is critical to the thermodynamic and kinetic behavior of metallic alloys under exposure to high-energy radiation, causing irradiation damage in materials. Radiation can generate non-equilibrium concentrations of point defects such as vacancies and interstitials. The excess point defects not only accelerate diffusional processes such as precipitation that cause radiation embrittlement, but also change the pathway of phase transformations, including nucleation processes. Understanding these defect behaviors is complicated by the challenge and complexity of addressing each possible local and discrete distribution of environments and chemical interactions around targeted defects-solute or solute-solute complexes. To resolve the challenge, machine learning regression techniques have emerged as powerful tools that can train and construct an energy model to accurately describe the chemical interactions of solutes and defects. In Fiscal Year 2022, the work focused on the workflow development and demonstration using machine learning regression, density functional theory, cluster expansion, and Monte Carlo simulation to predict the effects of ternary solute elements (e.g., aluminum and molybdenum) and point defects on the Cr-rich $\alpha^{\prime}$ precipitation in multicomponent FeCr model alloys. The computational outcomes include the prediction of the ternary phase diagram, vacancy formation energy for different compositions, and the effect of vacancies on the nucleation of Cr-rich clusters. The simulations predict a pronounced change of Cr solubility in bcc Fe by the addition of Al and the rejection of Al atoms from $\alpha^{\prime}$ precipitates. Additionally, the simulations show the formation of Cr-vacancy clusters as the initial nuclei for stable nucleation and growth of $\alpha^{\prime}$ particles. The results demonstrate important outcomes and applications of using machine learning pipeline to study model or commercial alloys with multicomponent solute species and point defects.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Comparison of permeability predictions on cemented sandstones with physics-based and machine learning approaches

Permeability prediction has been an important problem since the time of Darcy. Most approaches to solve this problem have used either idealized physical models or empirical relations. In recent years, machine learning (ML) has led to more accurate and robust, but less interpretable empirical models. Using 211 core samples collected from 12 wells in the Garn Sandstone from the North Sea, this study compared idealized physical models based on the Carman-Kozeny equation to interpretable ML models. We found that ML models trained on estimates of physical properties are more accurate than physical models. The results show evidence of a threshold of about 10% volume fraction, above which pore-filling cement strongly affects permeability.

03 NATURAL GAS↗

A comparative study of machine learning models for predicting the state of reactive mixing

Mixing phenomena are important mechanisms controlling flow, species transport, and reaction processes in fluids and porous media. Accurate predictions of reactive mixing are critical for many Earth and environmental science problems such as contaminant fate and remediation, macroalgae growth, and plankton biomass evolution. Here, to investigate the evolution of mixing dynamics under different scenarios (e.g., anisotropy, fluctuating velocity fields), a finite-element-based numerical model was built to solve the fast, irreversible bimolecular reaction-diffusion equations to simulate a range of reactive-mixing scenarios. A total of 2,315 simulations were performed using different sets of model input parameters comprising various spatial scales of vortex structures in the velocity field, time-scales associated with velocity oscillations, the perturbation parameter for the vortex-based velocity, anisotropic dispersion contrast (i.e., ratio of longitudinal-to-transverse dispersion), and molecular diffusion. The outputs comprised concentration profiles of reactants and products. The inputs to and outputs from these simulations were concatenated into feature and label matrices, respectively, to train 20 different machine learning (ML) models intended to emulate system behavior. These 20 ML emulators, based on linear methods, Bayesian methods, ensemble learning methods, and multilayer perceptrons (MLPs), were trained to classify the state of mixing and predict three quantities of interest (QoIs) characterizing species production, decay (i.e., average concentration, square of average concentration), and degree of mixing (i.e., variances of species concentration). Unsurprisingly, linear classifiers and regressors failed to reproduce the QoIs; however, ensemble methods (classifiers and regressors) and the MLP model accurately classified the state of reactive mixing and the QoIs. Among ensemble methods, random forest and decision-tree-based AdaBoost faithfully predicted the QoIs. At run time, trained ML emulators produced results times faster than the finite-element simulations. Due to their low computational expense and high accuracy, ensemble and MLP models are excellent emulators for these numerical simulations and great utilities in uncertainty quantification exercises, which can require 1,000s of forward model runs.

97 MATHEMATICS AND COMPUTING↗

Uncertainty quantification of machine learning models to improve streamflow prediction under changing climate and environmental conditions

Machine learning (ML) models, and Long Short-Term Memory (LSTM) networks in particular, have demonstrated remarkable performance in streamflow prediction and are increasingly being used by the hydrological research community. However, most of these applications do not include uncertainty quantification (UQ). ML models are data driven and can suffer from large extrapolation errors when applied to changing climate/environmental conditions. UQ is required to quantify the influence of data noises on model predictions and avoid overconfident projections in extrapolation. In this work, we integrate a novel UQ method, called PI3NN, with LSTM networks for streamflow prediction. PI3NN calculates Prediction Intervals by training 3 Neural Networks. It can precisely quantify the predictive uncertainty caused by the data noise and identify out-of-distribution (OOD) data in a non-stationary condition to avoid overconfident predictions. We apply the PI3NN-LSTM method in the snow-dominant East River Watershed in the western US and in the rain-driven Walker Branch Watershed in the southeastern US. Results indicate that for the prediction data which have similar features as the training data, PI3NN precisely quantifies the predictive uncertainty with the desired confidence level; and for the OOD data where the LSTM network fails to make accurate predictions, PI3NN produces a reasonably large uncertainty indicating that the results are not trustworthy and should avoid overconfidence. PI3NN is computationally efficient, robust in performance, and generalizable to various network structures and data with no distributional assumptions. It can be broadly applied in ML-based hydrological simulations for credible prediction.

54 ENVIRONMENTAL SCIENCES↗

Editorial: Data-driven machine learning for advancing hydrological and hydraulic predictability

The growing influence of machine learning (ML) in every aspect of our lives has led to revolutionary advancements in our understanding, prediction, and decision-making capabilities. One field that stands to benefit greatly from applying these techniques includes hydrology and hydraulics. The ability to predict hydrological and hydraulic phenomena with greater accuracy and reliability is of utmost importance, given the increasing threats posed by climate change and extreme weather/climate events. In this editorial, we explore the significant contributions made by four recent studies that aim to advance hydrological and hydraulic predictability through data-driven ML.

42 ENGINEERING↗

Physics-guided machine learning approaches to predict the ideal stability properties of fusion plasmas

One of the biggest challenges to achieve the goal of producing fusion energy in tokamak devices is the necessity of avoiding disruptions of the plasma current due to instabilities. The disruption event characterization and forecasting (DECAF) framework has been developed in this purpose, integrating physics models of many causal events that can lead to a disruption. Two different machine learning approaches are proposed to improve the ideal magnetohydrodynamic (MHD) no-wall limit component of the kinetic stability model included in DECAF. First, a random forest regressor (RFR), was adopted to reproduce the DCON computed change in plasma potential energy without wall effects for a large database of equilibria from the national spherical torus experiment (NSTX). This tree-based method provides an analysis of the importance of each input feature, giving an insight into the underlying physics phenomena. Secondly, a fully-connected neural network has been trained on sets of calculations with the DCON code, to get an improved closed form equation of the no-wall β limit as a function of the relevant plasma parameters indicated by the RFR. The neural network has been guided by physics theory of ideal MHD in its extension outside the domain of the NSTX experimental data. The estimated value has been incorporated into the DECAF kinetic stability model and tested against a set of experimentally stable and unstable discharges. Moreover, the neural network results were used to simulate a real-time stability assessment using only quantities available in real-time. Finally, the portability of the model was investigated, showing encouraging results by testing the NSTX-trained algorithm on the mega ampere spherical tokamak (MAST).

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗