Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “empirical machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Deep Nonparametric Estimation of Operators between Infinite Dimensional Spaces

Learning operators between infinitely dimensional spaces is an important learning task arising in machine learning, imaging science, mathematical modeling and simulations, etc. This paper studies the nonparametric estimation of Lipschitz operators using deep neural networks. Non-asymptotic upper bounds are derived for the generalization error of the empirical risk minimizer over a properly chosen network class. Under the assumption that the target operator exhibits a low dimensional structure, our error bounds decay as the training sample size increases, with an attractive fast rate depending on the intrinsic dimension in our estimation. Our assumptions cover most scenarios in real applications and our results give rise to fast rates by exploiting low dimensional structures of data in operator estimation. We also investigate the influence of network structures (e.g., network width, depth, and sparsity) on the generalization error of the neural network estimator and propose a general suggestion on the choice of network structures to maximize the learning efficiency quantitatively.

97 MATHEMATICS AND COMPUTING↗

Machine Learning Applications to Metal-Silicate Equilibria and their Insights into Core Formation

An extensive number of studies have experimentally investigated how elements distribute between metal and silicate phases, to better constrain core-mantle chemical equilibrium. Here, we present a new database compiling all (to our knowledge) experimental data on liquid metal-silicate partitioning from 118 peer-reviewed publications. We applied various machine learning techniques to gain further insights into these partitioning equilibria and their dependencies. We performed a network analysis to investigate the relationship between experiments and partition coefficients, which enables visualizing gaps in the experimental dataset and biases related to varying experimental conditions and analytical setup. In addition, semi-empirical thermodynamic models are commonly used to extrapolate these chemical reactions to the wide range of pressure, temperature and compositional conditions of planetary differentiation. These models are based on linear regressions that assume continuous relationship between partition coefficients and experimental variables. Here, we considered random forest regressions, which are algorithms based on ensembles of decision trees and does not consider continuous effects of each variable. The application of this regression significantly improves the prediction of metal-silicate partitioning for several elements including Ni, Si and Cr. We will show how this new approach improves our understanding of elemental exchange between metal and silicate and their implications for the Earth’s core formation.

siderophile element↗

Using Machine Learning to Predict Future Temperature Outputs in Geothermal Systems

Optimizing the power output, and economic value, of geothermal power plants over decades of operation is a major challenge in renewable energy. Optimizing the output requires the ability to predict the mass flow rates and the output temperatures of production wells based on the inputs of injection wells, as well as the time history of the system. Machine Learning (ML) that incorporates the known physics of geothermal systems is one possible solution to this challenge. In this work, we explore the ability of ML algorithms to predict future temperature outputs based on historical data. Considering the challenges with obtaining an empirical dataset from field data that is large enough to enable reliable ML, we propose an alternate approach: developing a high-fidelity reservoir model and using computational resources to build a dataset that enables ML. As a first step towards achieving this goal, we present preliminary results from applying ML to predict the temperature timeseries of simple modeled geothermal systems. We describe the application of relevant state-of-the-art ML approaches, such as the Long Short-Term Memory (LSTM) networks and Convolutional Neural Networks (CNN), to extract temporal structures in the model data. We assess the accuracy of the forecasts we obtain, compare the selected approaches, and share the lessons learned that would inform the process of training and utilizing ML algorithms for larger and more complex geothermal systems.

GEOTHERMAL ENERGY↗

Data-driven discovery of a formation prediction rule on high-entropy ceramics

The interest in high entropy ceramics (HECs) has increased steadily due to their superior properties. However, the prediction of their formation still poses challenges for the discovery of new systems. Here, we discover a rational rule for designing single-phase high entropy metal diborides (HEBs) using data-driven approach. The machine learning (ML) model is trained on data collected via high-throughput experiments (HTEs). K nearest neighbor (KNN) model shows an experimental validation accuracy of 93.75%. By implementing interpretable ML method, we demonstrate that a mismatch of the bonds between boron and transition metals (δ B-TM ) dominates the formation of HEBs. We propose an empirical rule that HEBs favor forming a single phase when δ B-TM < 3.66; otherwise, multiphase. The rule has a high accuracy of 93.33% for new HEBs predictions. In addition, we contribute 165 high quality HEBs data in total, which can promote the development of materials informatics in HEBs. Furthermore, this data-driven strategy can be expanded to accelerate the search for new HECs, paving a pathway to design novel HECs with superior properties rapidly.

36 MATERIALS SCIENCE↗

Quantum Image Denoising: A Framework via Boltzmann Machines, QUBO, and Quantum Annealing

We investigate a framework for binary image denoising via restricted Boltzmann machines (RBMs) that introduces a denoising objective in quadratic unconstrained binary optimization (QUBO) form and is well-suited for quantum annealing. The denoising objective is attained by balancing the distribution learned by a trained RBM with a penalty term for derivations from the noisy image. We derive the statistically optimal choice of the penalty parameter assuming the target distribution has been well-approximated, and further suggest an empirically supported modification to make the method robust to that idealistic assumption. We also show under additional assumptions that the denoised images attained by our method are, in expectation, strictly closer to the noise-free images than the noisy images are. While we frame the model as an image denoising model, it can be applied to any binary data. As the QUBO formulation is well-suited for implementation on quantum annealers, we test the model on a D-Wave Advantage machine, and also test on data too large for current quantum annealers by approximating QUBO solutions through classical heuristics.

restricted Boltzmann machine↗

Predictive Chemical Kinetic Modeling: Where We Succeed, Where We Struggle, and What Comes Next

Chemical kinetic modeling plays a foundational role in fields ranging from energy to environmental science, pharmaceuticals, and advanced materials. The past two decades have seen remarkable progress, particularly in modeling gas-phase reactions for thermochemical processes, leading to impactful industrial applications such as steam cracking and air quality management. However, new challenges are emerging. The successful development of systematic methodologies for the description of gas-phase kinetics opens the possibility to apply the same approach to the study of more challenging systems. Here, we review recent advances, including ab initio transition state theory-based master equation estimation of elementary rates, automated mechanism generation, machine-learning-assisted kinetics, and uncertainty quantification, and discuss the advances needed to apply the same methodological approach in areas such as heterogeneous catalysis, electrochemistry, liquid-phase and solid-state reactivity, and multiscale model integration. We advocate for the development of targeted tools, especially methods that go beyond empirical tuning toward first-principles-based predictions. We highlight the need for accessible software and AIaugmented workflows to democratize modeling for industry and academia alike. In this perspective, we call attention to not only what has worked but also what remains unsolved, advocating to avoid overemphasizing successes in scientific works at the expense of realism. The next decade should focus on predictive capability, physical accuracy, and community infrastructure (e.g., databases and services) to enable innovation across diverse fields. We argue that kinetic modeling, properly equipped, can accelerate discovery far beyond its traditional domains.

ab initio calculations↗

A deep potential model with long-range electrostatic interactions

Machine learning models for the potential energy of multi-atomic systems, such as the deep potential (DP) model, make molecular simulations with the accuracy of quantum mechanical density functional theory possible at a cost only moderately higher than that of empirical force fields. However, the majority of these models lack explicit long-range interactions and fail to describe properties that derive from the Coulombic tail of the forces. To overcome this limitation, we extend the DP model by approximating the long-range electrostatic interaction between ions (nuclei + core electrons) and valence electrons with that of distributions of spherical Gaussian charges located at ionic and electronic sites. The latter are rigorously defined in terms of the centers of the maximally localized Wannier distributions, whose dependence on the local atomic environment is modeled accurately by a deep neural network. In the DP long-range (DPLR) model, the electrostatic energy of the Gaussian charge system is added to short-range interactions that are represented as in the standard DP model. The resulting potential energy surface is smooth and possesses analytical forces and virial. Missing effects in the standard DP scheme are recovered, improving on accuracy and predictive power. By including long-range electrostatics, DPLR correctly extrapolates to large systems the potential energy surface learned from quantum mechanical calculations on smaller systems. We illustrate the approach with three examples: the potential energy profile of the water dimer, the free energy of interaction of a water molecule with a liquid water slab, and the phonon dispersion curves of the NaCl crystal.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Forecasting Commercial Building Electricity Consumption, Zone Airflow and Zone Temperature: Update - Development of a Generalized Machine Learning Approach

The U.S. power grid is being transformed to make it smarter, more efficient, and cleaner. This transformation is leading to the addition of a significant of energy generated by distributed, variable, and renewable resources. Because of the variable nature of renewable generation, the short- and long-term supply and demand imbalances are less predictable, and conventional approaches to mitigating the imbalances will be less efficient or cost effective. To address this challenge and to support the mission and the vision of the U.S. Department of Energy’s (DOE’s) Office of Energy Efficiency and Renewable Energy (EERE) Building Technologies Office has developed a Grid-Interactive Efficient Building Strategy. The strategy focuses on simultaneously improving building energy efficiency and supporting reliability and resilience of the electric grid more efficiently and at a lower cost. In addition, EERE and DOE’s Office of Electricity created an initiative led by DOE and supported by the national laboratories under the Grid Modernization Lab Consortium structure to enhance grid modernization. The work reported in this document is part of the first set of projects funded under the initiative to design, develop, and validate scalable transactive control technologies for the commercial buildings sector. Transactive controls requires the ability of individual end-use loads to express flexibility as a function of a transactive signal (e.g., price). Empirical grey- and black-box models have been widely used to express flexibility. Although this approach is generally easy to construct and simple to use, it does not capture non-linear behavior that some end-use loads represent. Therefore, Pacific Northwest National Laboratory (PNNL) with support from Western Washington University conducted this research to explore the use of deep machine learning (ML) techniques. The work reported in this document is limited to forecasting whole building electricity consumption, the zone airflow and the zone temperature predictions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Graph-EAM: An Interpretable and Efficient Graph Neural Network Potential Framework

The development of deep learning interatomic potentials has enabled efficient and accurate computations in quantum chemistry and materials science, circumventing computationally expensive ab initio calculations. However, the huge number of learnable parameters in deep learning models and their complex architectures hinder physical interpretability and affect the robustness of the derived potential. In this work, we propose graph-EAM, a lightweight graph neural network (GNN) inspired by the empirical embedded atom method to model the interatomic potential of single-element structures. Four material systems: platinum, niobium, silicon, and amorphous-carbon, for which quantum simulation data sets are publicly available, are examined to demonstrate that graph-EAM can achieve high energy and force prediction accuracy-comparable or better than existing state-of-the-art machine learning models-with much fewer parameters. It is also shown that the explicit inclusion of the angular information via three-body atomic density increases the prediction accuracy. In conclusion, the accuracy and efficiency of potentials obtained from graph-EAM can help accelerate the molecular dynamics simulation.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

In-Situ Calibrated Digital Process Twin Models for Resource Efficient Manufacturing

The chief objective of manufacturing process improvement efforts is to significantly minimize process resources such as time, cost, waste, and consumed energy while improving product quality and process productivity. This paper presents a novel physics-informed optimization approach based on artificial intelligence (AI) to generate digital process twins (DPTs). The utility of the DPT approach is demonstrated in the case of finish machining of aerospace components made from gamma titanium aluminide alloy (γ-TiAl). This particular component has been plagued with persistent quality defects, including surface and sub-surface cracks, which adversely affect resource efficiency. Previous process improvement efforts have been restricted to anecdotal post-mortem investigation and empirical modeling, which fail to address the fundamental issue of how and when cracks occur during cutting. In this work, the integration of in-situ process characterization with modular physics-based models is presented, and machine learning algorithms are used to create a DPT capable of reducing environmental and energy impacts while significantly increasing yield and profitability. Based on the preliminary results presented here, we report an improvement in the overall embodied energy efficiency of over 84%, 93% in process queuing time, 2% in scrap cost, and 93% in queuing cost has been realized for γ-TiAl machining using our novel approach.

42 ENGINEERING↗

Combining Observations and Models: A Review of the CARDAMOM Framework for Data‐Constrained Terrestrial Ecosystem Modeling

The rapid increase in the volume and variety of terrestrial biosphere observations (i.e., remote sensing data and in situ measurements) offers a unique opportunity to derive ecological insights, refine process‐based models, and improve forecasting for decision support. However, despite their potential, ecological observations have primarily been used to benchmark process‐based models, as many past and current models lack the capability to directly integrate observations and their associated uncertainties for parameterization. In contrast, data assimilation frameworks such as the CARbon DAta MOdel fraMework (CARDAMOM) and its suite of process‐based models, known as the Data Assimilation Linked Ecosystem Carbon Model (DALEC), are specifically designed for model‐data fusion. This review, motivated by a recent CARDAMOM community workshop, examines the development and applications of CARDAMOM, with an emphasis on its role in advancing ecosystem process understanding. CARDAMOM employs a Bayesian approach, using a Markov Chain Monte Carlo algorithm to enable data‐driven calibration of DALEC parameters and initial states (i.e., carbon pool sizes) through observation operators. CARDAMOM's unique ability to retrieve localized model process parameters from diverse datasets—ranging from in situ measurements to global satellite observations—makes it a highly flexible tool for analyzing spatially variable ecosystem responses to environmental change. However, assimilating these data also presents challenges, including data quality issues that propagate into model skill, as well as trade‐offs between model complexity, parameter equifinality, and predictive performance. We discuss potential solutions to these challenges, such as reducing parameter equifinality by incorporating new observations. This review also offers community recommendations for incorporating emerging datasets, integrating machine learning techniques, strengthening collaboration with remote sensing, field, and modeling communities, and expanding CARDAMOM's relevance for localized ecosystem monitoring and decision‐making. CARDAMOM enables a deep, mechanistic understanding of terrestrial ecosystem dynamics that cannot be achieved through empirical analyses of observational datasets or weakly constrained models alone.

Bayesian inference↗

Learning curves for drug response prediction in cancer cell lines

Motivated by the size and availability of cell line drug sensitivity data, researchers have been developing machine learning (ML) models for predicting drug response to advance cancer treatment. As drug sensitivity studies continue generating drug response data, a common question is whether the generalization performance of existing prediction models can be further improved with more training data. We utilize empirical learning curves for evaluating and comparing the data scaling properties of two neural networks (NNs) and two gradient boosting decision tree (GBDT) models trained on four cell line drug screening datasets. The learning curves are accurately fitted to a power law model, providing a framework for assessing the data scaling behavior of these models. The curves demonstrate that no single model dominates in terms of prediction performance across all datasets and training sizes, thus suggesting that the actual shape of these curves depends on the unique pair of an ML model and a dataset. The multi-input NN (mNN), in which gene expressions of cancer cells and molecular drug descriptors are input into separate subnetworks, outperforms a single-input NN (sNN), where the cell and drug features are concatenated for the input layer. In contrast, a GBDT with hyperparameter tuning exhibits superior performance as compared with both NNs at the lower range of training set sizes for two of the tested datasets, whereas the mNN consistently performs better at the higher range of training sizes. Moreover, the trajectory of the curves suggests that increasing the sample size is expected to further improve prediction scores of both NNs. These observations demonstrate the benefit of using learning curves to evaluate prediction models, providing a broader perspective on the overall data scaling characteristics. A fitted power law learning curve provides a forward-looking metric for analyzing prediction performance and can serve as a co-design tool to guide experimental biologists and computational scientists in the design of future experiments in prospective research studies.

60 APPLIED LIFE SCIENCES↗

Use of Design of Experiments in Determining Neural Network Architectures for Loss of Control Detection

We describe empirical methods for selecting a neural network architecture to implement belief state inference on generic commercial transport aircraft. We highlight a case study on the planning, execution, and analysis of a set of experiments to determine the configurations of a conditional variational autoencoder (CVAE). Our main contribution is the application of a structured method that can be used for machine learning in many aerospace applications. This method optimizes the structure and training parameters of a neural network for belief state inference, using Design of Experiments (DOE) statistical methodologies. The motivation for this specific DOE analysis was to identify the appropriate hyperparameters for measuring the CVAE reconstruction probability and latent space, such that the measurements can be used to infer qualitative state changes for the aircraft. We demonstrate that this process yields information about a trained neural network’s utility for this specific application, along with a quantifiable range of certainty. We execute 84 experiments using loss-of-control flight maneuver data from the NASA T-2 aircraft, demonstrating that this empirical process allows us to construct cheap and simple models with specific attributes amenable to belief state inference in aerospace applications.

Loss of Control↗

Use of Design of Experiments in Determining Neural Network Architectures for Loss of Control Detection

We describe empirical methods for selecting a neural network architecture to implement belief state inference on generic commercial transport aircraft. We highlight a case study on the planning, execution, and analysis of a set of experiments to determine the configurations of a conditional variational autoencoder (CVAE). Our main contribution is the application of a structured method that can be used for machine learning in many aerospace applications. This method optimizes the structure and training parameters of a neural network for belief state inference, using Design of Experiments (DOE) statistical methodologies. The motivation for this specific DOE analysis was to identify the appropriate hyperparameters for measuring the CVAE reconstruction probability and latent space, such that the measurements can be used to infer qualitative state changes for the aircraft. We demonstrate that this process yields information about a trained neural network’s utility for this specific application, along with a quantifiable range of certainty. We execute 84 experiments using loss-of-control flight maneuver data from the NASA T 2 aircraft, demonstrating that this empirical process allows us to construct cheap and simple models with specific attributes amenable to belief state inference in aerospace applications.

Loss of Control↗

Use of Design of Experiments in Determining Neural Network Architectures for Loss of Control Detection

Abstract—We describe empirical methods for selecting a neural network architecture to implement belief state inference on generic commercial transport aircraft. We highlight a case study on the planning, execution, and analysis of a set of experiments to determine the configurations of a conditional variational autoencoder (CVAE). Our main contribution is the application of a structured method that can be used for machine learning in many aerospace applications. This method optimizes the structure and training parameters of a neural network for belief state inference, using Design of Experiments (DOE) statistical methodologies. The motivation for this specific DOE analysis was to identify the appropriate hyperparameters for measuring the CVAE reconstruction probability and latent space, such that the measurements can be used to infer qualitative state changes for the aircraft. We demonstrate that this process yields information about a trained neural network’s utility for this specific application, along with a quantifiable range of certainty. We execute 84 experiments using loss-of-control flight maneuver data from the NASA T-2 aircraft, demonstrating that this empirical process allows us to construct cheap and simple models with specific attributes amenable to belief state inference in aerospace applications.

neural networks↗

Enhancing high-fidelity neural network potentials through low-fidelity sampling

The efficacy of neural network potentials (NNPs) critically depends on the quality of the configurational datasets used for training. Prior research using empirical potentials has shown that well-selected liquid–solid transitional configurations of a metallic system can be translated to other metallic systems. This study demonstrates that such validated configurations can be relabeled using density functional theory (DFT) calculations, thereby enhancing the development of high-fidelity NNPs. Training strategies and sampling approaches are efficiently assessed using empirical potentials and subsequently relabeled via DFT in a highly parallelized fashion for high-fidelity NNP training. Our results reveal that relying solely on energy and force for NNP training is inadequate to prevent overfitting, highlighting the necessity of incorporating stress terms into the loss functions. To optimize training involving force and stress terms, we propose employing transfer learning to fine-tune the weights, ensuring that the potential surface is smooth for these quantities composed of energy derivatives. This approach markedly improves the accuracy of elastic constants derived from simulations in both empirical potential-based NNPs and relabeled DFT-based NNPs. Overall, this study offers significant insights into leveraging empirical potentials to expedite the development of reliable and robust NNPs at the DFT level.

97 MATHEMATICS AND COMPUTING↗

APOGEE Net: Improving the Derived Spectral Parameters for Young Stars through Deep Learning

Machine learning allows for efficient extraction of physical properties from stellar spectra that have been obtained by large surveys. The viability of machine-learning approaches has been demonstrated for spectra covering a variety of wavelengths and spectral resolutions, but most often for main-sequence (MS) or evolved stars, where reliable synthetic spectra provide labels and data for training. Spectral models of young stellar objects (YSOs) and low-mass MS stars are less well-matched to their empirical counterparts, however, posing barriers to previous approaches to classify spectra of such stars. In this work, we generate labels for YSOs and low-mass MS stars through their photometry. We then use these labels to train a deep convolutional neural network to predict logg, T {sub eff}, and Fe/H for stars with Apache Point Observatory Galactic Evolution Experiment (APOGEE) spectra in the DR14 data set. This “APOGEE Net” has produced reliable predictions of logg for YSOs, with uncertainties of within 0.1 dex and a good agreement with the structure indicated by pre-MS evolutionary tracks, and it correlates well with independently derived stellar radii. These values will be useful for studying pre-MS stellar populations to accurately diagnose membership and ages.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine Learning-Based Process Control for Injection Molding of Recycled Polypropylene

The increased interest in artificial intelligence in manufacturing has driven the adoption of machine learning to optimize processes and improve efficiency. A key challenge in injection molding is the variability of recycled materials, which affects part quality and processing stability. This study presents a novel closed-loop process control approach for injection molding, leveraging machine learning to adaptively predict processing inputs and quality outcomes. The methodology was tested on five blends of recycled polypropylene (rPP), using artificial neural networks (ANNs), linear regression, and polynomial regression to model the relationships between material properties and process parameters. The dataset was split 80/20 into training and testing sets. The ANN model was implemented using TensorFlow and Keras, with six hidden layers of 32 neurons per layer, ReLU activation, and an Adam optimizer. Empirical tuning and early stopping were used to optimize performance and prevent overfitting. Predictions were evaluated based on mean absolute error (MAE), mean squared error (MSE), and percentage error. The results showed that yield stress, ultimate elongation, and part weight were accurately predicted within a 5% error for linear and polynomial regression models and within a 10% error for the ANN. However, modulus predictions were less reliable, with errors of ~11% for ANN and linear regression and ~40% for polynomial regression, reflecting the inherent variability of this property in rPP blends. Predictions of processing inputs had errors ranging from 3% to 25%, depending on the model and response variable. No single modeling approach was consistently superior across all responses, highlighting the complexity of the relationship between material properties, process parameters, and quality metrics. Overall, the work demonstrates that closed-loop process control, powered by machine learning, can effectively predict key quality parameters in injection molding of recycled materials. The proposed approach can improve process stability and material utilization, facilitating increased adoption of sustainable materials.

Krantz, Joshua↗