Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Machine Learning Models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

A Machine Learning Model for Predicting Composition of Catalytic Coprocessing Products from Molecular Beam Mass Spectra

Demand for the development of an automated and integrated refining process for biofuels has increased in recent years due to the lack of generalized process inspection tools. In bio-oil upgrading processes, all process variables are maintained based on the offline specification of intermediates and products. A lack of real-time product specifications in batch-wise monitoring can cause process failure and wasted resources. Therefore, there is a need for a fast and accurate intermediates/product specification tool that can be used for real-time specification to reduce waste and mitigate the risk of process failure. Here, to address this gap, we developed a machine learning (ML) model for predicting speciated bio-oil composition, including paraffin, iso-paraffins, olefins, naphthene, and aromatics. The model is trained using the mass spectra from upgraded products collected in the vapor phase before condensation and predicts the composition of the condensed product. Training ML models using raw mass spectra is challenging due to numerous overlapped peaks originating from different parent compounds. With this in mind, we propose a protocol that (i) transforms raw mass spectra to chemistry-inspired predefined features and (ii) trains decision tree-based models using these features. Our results show that the random forest model was robust against overfitting and had the highest accuracy compared to other models. Moreover, a stochastic ablation method determined the eight most significant features while maximizing the accuracy. Our protocol facilitates real-time compositional analysis of upgraded bio-oils and thus real-time process monitoring. Additionally, this protocol enables the rational design of efficient catalysts and the determination of optimal process conditions.

09 BIOMASS FUELS↗

Applied Machine-Learning Models to Identify Spectral Sub-Types of M Dwarfs from Photometric Surveys

M dwarfs are the most abundant stars in the Solar Neighborhood and they are prime targets for searching for rocky planets in habitable zones. Consequently, a detailed characterization of these stars is in demand. The spectral sub-type is one of the parameters that is used for the characterization and it is traditionally derived from the observed spectra. However, obtaining the spectra of M dwarfs is expensive in terms of observation time and resources due to their intrinsic faintness. We study the performance of four machine-learning (ML) models—K-Nearest Neighbor (KNN), Random Forest (RF), Probabilistic Random Forest (PRF), and Multilayer Perceptron (MLP)—in identifying the spectral sub-types of M dwarfs at a grand scale by deploying broadband photometry in the optical and near-infrared. We trained the ML models by using the spectroscopically identified M dwarfs from the Sloan Digital Sky Survey (SDSS) Data Release (DR) 7, together with their photometric colors that were derived from the SDSS, Two-Micron All-Sky Survey, and Wide-field Infrared Survey Explorer. We found that the RF, PRF, and MLP give a comparable prediction accuracy, 74%, while the KNN provides slightly lower accuracy, 71%. We also found that these models can predict the spectral sub-type of M dwarfs with ~99% accuracy within ±1 sub-type. The five most useful features for the prediction are r - z, r - i, r - J, r - H , and g - z, and hence lacking data in all SDSS bands substantially reduces the prediction accuracy. However, we can achieve an accuracy of over 70% when the r and i magnitudes are available. Since the stars in this study are nearby (d ≲ 1300 pc for 95% of the stars), the dust extinction can reduce the prediction accuracy by only 3%. Finally, we used our optimized RF models to predict the spectral sub-types of M dwarfs from the Catalog of Cool Dwarf Targets for the Transiting Exoplanet Survey Satellite, and we provide the optimized RF models for public use.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine-learning modeling of magnetization dynamics in quasi-equilibrium and driven metallic spin systems

Here, we present a perspective on recent progress in machine-learning (ML) force-field approaches for large-scale Landau–Lifshitz–Gilbert (LLG) simulations of metallic spin systems. Building on a generalization of the Behler–Parrinello (BP) architecture originally developed for quantum molecular dynamics, we develop scalable and transferable ML models that faithfully capture the complex, environment-dependent electron-mediated exchange fields characteristic of itinerant magnets. A central ingredient of this framework is the implementation of symmetry-aware magnetic descriptors based on group-theoretical bispectrum formalisms. Leveraging these ML force fields, LLG simulations faithfully reproduce hallmark non-collinear magnetic orders—such as the 120° and tetrahedral states—on the triangular lattice, and successfully capture the complex spin textures emerging in the mixed-phase states of a square-lattice double-exchange model under thermal quench. We further discuss a generalized potential theory that extends the BP formalism to incorporate both conservative and nonconservative electronic torques, thereby enabling ML models to learn nonequilibrium exchange fields from computationally demanding microscopic approaches such as nonequilibrium Green’s-function techniques. This extension yields quantitatively accurate predictions of voltage-driven domain-wall motion and establishes a foundation for quantum-accurate, multiscale modeling of nonequilibrium spin dynamics and spintronic functionalities.

Descriptors↗

Field Work Proposal ERKJ358: Black-box training for scientific machine learning models (Final Report)

The overarching goal of this project is to develop a scalable black-box training capability for scientific machine learning (SciML) problems that are non-trainable with existing automatic differentiation (AD)-based algorithms. AD assumes that a loss function can be decomposed into a sequence of elementary operations whose derivatives are known. This assumption is violated when the loss function includes a black-box physical model (e.g., a legacy simulator). The current strategy, converting a black-box simulator to an AD-enabled code via differential programming, is inflexible and time-, labor-consuming. Thus, black-box optimization is a main workhorse for training SciML models, e.g., in scientific reinforcement learning, hyper-parameter fine tuning, designing SciML models with adversarial robustness, etc.

97 MATHEMATICS AND COMPUTING↗

Physics-informed machine learning modeling for predictive control using noisy data

Due to the occurrence of over-fitting at the learning phase, the modeling of chemical processes via artificial neural networks (ANN) by using corrupted data (i.e., noisy data) is an ongoing challenge. Therefore, this work investigates the effect of both Gaussian and non-Gaussian noise on the performance of process-structure based recurrent neural networks (RNN) models, which take the form of partially-connected RNN models in this work, that are used to approximate a class of multi-input-multi-outputs nonlinear systems. Furthermore, two different techniques, specifically Monte Carlo dropout and co-teaching, are utilized in the development of partially-connected RNN models. Here, these two techniques are employed to reduce the over-fitting in ANNs when noisy data is used in the training process and, hence, to improve the open-loop accuracy as well as the closed-loop performance under a Lyapunov-based model predictive controller (MPC). Aspen Plus Dynamics, a well-known high-fidelity process simulator, is used to simulate a large-scale chemical process application in order to demonstrate the anticipated improvements in both open-loop approximation and closed-loop controller performance in the presence of Gaussian and non-Gaussian noise in the data set using physics-informed RNNs.

97 MATHEMATICS AND COMPUTING↗

Machine learning modeling and model predictive control of a closed-circuit reverse osmosis system

Closed-circuit reverse osmosis (CCRO) offers a flexible and energy-efficient alternative to conventional reverse osmosis by operating in a semi-batch mode that recycles brine, enabling higher recovery rates and reduced specific energy consumption (SEC). However, developing accurate, system-level dynamic models for CCRO remains challenging due to its nonlinear, multi-phase operation and sensitivity to variable feed water conditions. Traditional modeling approaches, such as NARMAX (nonlinear autoregressive moving average with exogenous inputs), often struggle to generalize across varying inlet feed concentrations, necessitating frequent parameter re-estimation and limiting their utility for real-time control applications. To address these limitations, we developed a long short-term memory (LSTM) neural network model trained on an extensive experimental data set from a CCRO pilot plant. The model accepts three inputs, feed flow rate, recirculation flow rate, and initial feed conductivity, and predicts three key outputs: reject conductivity, feed pump power draw, and recirculation pump power draw. We validated the LSTM model against experimental data, demonstrating its ability to distinguish between different feed conductivities and adapt to variable flow rates. Subsequently, we incorporated the LSTM model within a nonlinear model predictive control (MPC) scheme and conducted closed-loop simulations to optimize the integrated SEC (iSEC). In conclusion, the results project up to a 6% reduction in iSEC by using MPC to optimize performance over the entire experiment duration, without requiring any random excitation for data collection or parameter re-estimation.

Desalination↗

Uncertainty-Based Design: Finite Element and Explainable Machine Learning Modeling of Carbon–Carbon Composites for Ultra-High Temperature Solar Receivers

Design under uncertainty has significantly grown in research developments during the past decade. Additionally, machine learning (ML) and explainable ML (XML) have offered various opportunities to provide reliable predictable models. The current article investigates the use of finite element modeling (FEM), ML and XML predictions, and uncertain-based design of carbon-carbon (C-C) composites for use in ultra-high temperatures. A C-C composite concentrating solar power (CSP) as a microvascular receiver is considered as a case study. These C-C composites are fiber composites with directly integrated carbonized microchannels to form a lightweight, high-absorptivity material that includes an embedded microvascular network of channels. The topology of these microchannels is engineered to optimize heat transfer to a supercritical carbon dioxide (sCO2) heat transfer fluid. The mechanical characterization of C-C composites is highly challenging. Thus, designing every component made of C-C composites for ultra-high temperature applications needs an uncertainty-based analysis. As a part of a comprehensive project on the development of a novel carbonized microvascular C-C composite, this paper explores C-C composite sensitivity analysis, FEM, ML prediction, and XML analysis. The resulting composite can then be carbonized and coated with an oxidation-resistant coating to form a thermally efficient and mechanically robust C-C composite. An ANSYS 3-D-FE model was used to analyze the CSP’s stress/strain. To consider the variability in the mechanical and thermal properties of C-C composites, various mechanical properties are considered as the ANSYS FEM’s input. A synthetic dataset from 730 ANSYS runs was produced to feed into the ML and XML algorithms for uncertainty analysis and prediction. The ML and XML algorithms could accurately predict the CSP stresses/strains.

Daghigh, Vahid (ORCID:0000000298941620)↗

Revealing low-temperature plasma efficacy through a dose-rate assessment by DNA damage detection combined with machine learning models

Low-temperature plasmas have quickly emerged as alternative and unconventional types of radiation that offer great promise for various clinical modalities. As with other types of radiation, the therapeutic efficacy and safety of low-temperature plasmas are ubiquitous concerns, and assessing their dose rates is crucial in clinical settings. Unfortunately, assessing the dose rates by standard dosimetric techniques has been challenging. To overcome this difficulty, we proposed a dose-rate assessment framework that combined the predictive modeling of plasma-induced damage in DNA by machine learning with existing radiation dose-DNA damage correlations. Our results indicated that low-temperature plasmas have a remarkably high dose rate that can be tuned by various process parameters. This attribute is beneficial for inducing radiobiological effects in a more controllable manner.

59 BASIC BIOLOGICAL SCIENCES↗

Streamflow simulation in data-scarce basins using Bayesian and physics-informed machine learning models

Hydrologic predictions at rural watersheds are important but also challenging due to data shortage. Long short-term memory (LSTM) networks are a promising machine learning approach and have demonstrated good performance in streamflow predictions. However, due to its data-hungry nature, most LSTM applications focus on well-monitored catchments with abundant and high-quality observations. In this work, we investigate predictive capabilities of LSTM in poorly monitored watersheds with short observation records. To address three main challenges of LSTM applications in data-scarce locations, i.e., overfitting, uncertainty quantification (UQ), and out-of-distribution prediction, we evaluate different regularization techniques to prevent overfitting, apply a Bayesian LSTM for UQ, and introduce a physics-informed hybrid LSTM to enhance out-of-distribution prediction. Through case studies in two diverse sets of catchments with and without snow influence, we demonstrate that 1) when hydrologic variability in the prediction period is similar to the calibration period, LSTM models can reasonably predict daily streamflow with Nash–Sutcliffe efficiency above 0.8, even with only 2 years of calibration data; 2) when the hydrologic variability in the prediction and calibration periods is dramatically different, LSTM alone does not predict well, but the hybrid model can improve the out-of-distribution prediction with acceptable generalization accuracy; 3) L2 norm penalty and dropout can mitigate overfitting, and Bayesian and hybrid LSTM have no overfitting; and 4) Bayesian LSTM provides useful uncertainty information to improve prediction understanding and credibility. In conclusion, these insights have vital implications for streamflow simulation in watersheds where data quality and availability are a critical issue.

54 ENVIRONMENTAL SCIENCES↗

Detecting Large Explosions With Machine Learning Models Trained on Synthetic Infrasound Data

Explosions produce low-frequency acoustic (infrasound) waves capable of propagating globally, but the spatio-temporal variability of the atmosphere makes detecting events difficult. Machine learning (ML) is well-suited to identify the subtle and nonlinear patterns in explosion infrasound signals, but a previous lack of ground-truth data inhibited training of generalized models. We introduce a physics-based method that propagates infrasound sources through realistic atmospheres to create 28,000 synthetic events, which are used to train ML classifiers. A simple artificial neural network and modern temporal convolutional network discriminate synthetic events from background noise with >90% accuracy and, more importantly, successfully identify the majority of real-world explosion signals recorded during the Humming Road Runner experiment. ML models trained entirely on physics-based synthetics advance explosion detection capabilities and make ML more viable to related fields lacking training data.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Machine learning models for PDE constrained optimization

Partial differential equation (PDE)-constrained optimization problems arise in a variety of scientific and engineering applications, such as topology optimization, electrodynamics, fluid dynamics, and structural dynamics. However, these problems are often challenging and computationally expensive to solve, due to the need to solve the PDEs within the optimization loop. One approach to reducing the computational cost of these methods while providing convergence guarantees is through inexact trust region methods; this method uses lower fidelity solutions of the PDE at early stages of the optimization and adjusts the required accuracy of inexact PDE solvers as the optimization progresses. In this work, we explore the use of machine learning based surrogate models with these inexact trust region methods. We first demonstrate the potential of this approach by using Gaussian processes as the surrogate model and test this on a simple PDE-constrained optimization problem. We then document explorations into improving the computational costs of evolutional deep neural network / neural Galerkin methods, with the eventual goal of using these methods with the inexact trust region algorithms. We are able to speed up these approaches, albeit at the cost of lower accuracy.

97 MATHEMATICS AND COMPUTING↗

Evaluating proxies for the drivers of natural gas productivity using machine-learning models

We report the extensive development of unconventional reservoirs using horizontal drilling and multistage hydraulic fracturing has generated large volumes of reservoir characterization and production data. The analysis of this abundant data using statistical methods and advanced machine-learning (ML) techniques can provide data-driven insights into well performance. Most predictive modeling studies have focused on the impact that different well completion and stimulation strategies have on well production but have not fully exploited the available in situ rock property data to determine its role in reservoir productivity. We have used machine-learning techniques to rank rock mechanical properties, microseismic attributes, and stimulation parameters in the order of their significance for predicting natural gas production from an unconventional reservoir. The data for this study came from a hydraulically fractured well in the Marcellus Shale in Monongalia County, West Virginia. The data classes included measurements aggregated by well completion stage that included (1) gas production, (2) well-log-derived measurements including bulk density, elastic moduli, shear impedance, compressional impedance, brittleness, and gamma measurements, (3) microseismic attributes, (4) long-period long-duration (LPLD) event counts, (5) fracture counts, and (6) stimulation parameters that included the fluid injection volume and average pumping pressure. To identify observable proxies for the drivers of gas production, we evaluated five commonly used ML approaches including multivariate adaptive regression spline, Gaussian mixture model, random forest, gradient boosting, and neural network. We selected five variables including LPLD event count, seismogenic b-value, hydraulic diffusivity, cumulative moment, and fluid volume as the features most likely to impact gas productivity at the stage level in the study area. The data-driven selection of these parameters for their importance in determining gas production can help reservoir engineers design more effective hydraulic-fracture treatments in the Marcellus Shale and other similar unconventional reservoirs. Plain language summary: We use machine-learning methods and data-driven selection of reservoir parameters to rank and better understand their importance in determining gas production, which can help reservoir engineers design more effective hydraulic-fracture treatments in the Marcellus Shale and other similar unconventional reservoirs.

58 GEOSCIENCES↗

Characterize traction–separation relation and interfacial imperfections by data-driven machine learning models

Abstract Interfacial mechanical properties are important in composite materials and their applications, including vehicle structures, soft robotics, and aerospace. Determination of traction–separation (T–S) relations at interfaces in composites can lead to evaluations of structural reliability, mechanical robustness, and failures criteria. Accurate measurements on T–S relations remain challenging, since the interface interaction generally happens at microscale. With the emergence of machine learning (ML), data-driven model becomes an efficient method to predict the interfacial behaviors of composite materials and establish their mechanical models. Here, we combine ML, finite element analysis (FEA), and empirical experiments to develop data-driven models that characterize interfacial mechanical properties precisely. Specifically, eXtreme Gradient Boosting (XGBoost) multi-output regressions and classifier models are harnessed to investigate T–S relations and identify the imperfection locations at interface, respectively. The ML models are trained by macroscale force–displacement curves, which can be obtained from FEA and standard mechanical tests. The results show accurate predictions of T–S relations ( R 2 = 0.988) and identification of imperfection locations with 81% accuracy. Our models are experimentally validated by 3D printed double cantilever beam specimens from different materials. Furthermore, we provide a code package containing trained ML models, allowing other researchers to establish T–S relations for different material interfaces.

97 MATHEMATICS AND COMPUTING↗

MOFSimplify, machine learning models with extracted stability data of three thousand metal–organic frameworks

Abstract We report a workflow and the output of a natural language processing (NLP)-based procedure to mine the extant metal–organic framework (MOF) literature describing structurally characterized MOFs and their solvent removal and thermal stabilities. We obtain over 2,000 solvent removal stability measures from text mining and 3,000 thermal decomposition temperatures from thermogravimetric analysis data. We assess the validity of our NLP methods and the accuracy of our extracted data by comparing to a hand-labeled subset. Machine learning (ML, i.e. artificial neural network) models trained on this data using graph- and pore-geometry-based representations enable prediction of stability on new MOFs with quantified uncertainty. Our web interface, MOFSimplify, provides users access to our curated data and enables them to harness that data for predictions on new MOFs. MOFSimplify also encourages community feedback on existing data and on ML model predictions for community-based active learning for improved MOF stability models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗