Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Transfer Learning Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Transformers and Long Short-Term Memory Transfer Learning for GenIV Reactor Temperature Time Series Forecasting

Automated monitoring of the coolant temperature can enable autonomous operation of generation IV reactors (GenIV), thus reducing their operating and maintenance costs. Automation can be accomplished with machine learning (ML) models trained on historical sensor data. However, the performance of ML usually depends on the availability of large amount of training data, which is difficult to obtain for GenIV, as this technology is still under development. We propose the use of transfer learning (TL), which involves utilizing knowledge across different domains, to compensate for this lack of training data. TL can be used to create pre-trained ML models with data from small-scale research facilities, which can then be fine-tuned to monitor GenIV reactors. In this work, we develop pre-trained Transformer and long short-term memory (LSTM) networks by training them on temperature measurements from thermal hydraulic flow loops operating with water and Galinstan fluids at room temperature at Argonne National Laboratory. The pre-trained models are then fine-tuned and re-trained with minimal additional data to perform predictions of the time series of high temperature measurements obtained from the Engineering Test Unit (ETU) at Kairos Power. The performance of the LSTM and Transformer networks is investigated by varying the size of the lookback window and forecast horizon. The results of this study show that LSTM networks have lower prediction errors than Transformers, but LSTM errors increase more rapidly with increasing lookback window size and forecast horizon compared to the Transformer errors.

LSTM↗

Designing a quantum-accurate machine-learning potential to enable large-scale simulations of deuterium under shock

Large-scale molecular dynamics of deuterium under shock can elucidate kinetic processes vital to the target design in inertial confinement fusion and high-energy-density experiments. However, modeling the complex evolution of this material from an insulating molecular state at ambient pressure to an ionized, atomic fluid under strong shock is beyond the capability of simple pair and even bond order potentials. We thus train a quantum-accurate and broadly transferable machine-learning interatomic potential for deuterium using the Chebyshev Interaction Model for Efficient Simulations framework. We show that due to an improved description of the molecular-to-atomic transition, our model is able to better reproduce the ab initio equation of state, radial distribution functions, and principal Hugoniot than bond order potentials. This represents an important step toward large-scale quantum-accurate and nonequilibrium simulations of complicated systems under dynamic changes including phase transitions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A deep learning and finite element approach for exploration of inverse structure–property designs of lightweight hybrid composites

Hybrid composites have important applications, such as high-performance and lightweight materials in aerospace and automotive industries. Hybrid composites utilize the synergy of diverse fillers to achieve desired material properties, but usually have more complicated microstructures. While topology optimization can optimize a particular property, designing hybrid composites for customized mechanical performances, e.g. full-range stress–strain curve, remains challenging. Here, a computational framework that integrated finite element analysis (FEA) and artificial intelligence (AI) methods of Conditional Generative Adversarial Networks (cGAN) deep learning and transfer learning was developed to establish inverse structure–property relationships and design tailor-made hybrid composites. Based on FEA-generated datasets of hybrid fiber-particle–matrix microstructures and their corresponding full-range stress–strain curves, a cGAN architecture was trained to generate tailored microstructures and establish structure–property relationships. Similarity in microstructural features and well-matched stress–strain curves based on the AI-generated composites were achieved. In conclusion, transfer learning was used to expand the pre-trained model for designing different materials systems.

Hybrid composites↗

Bi-fidelity variational auto-encoder for uncertainty quantification

Quantifying the uncertainty of quantities of interest (QoIs) from physical systems is a primary objective in model validation. However, achieving this goal entails balancing the need for computational efficiency with the requirement for numerical accuracy. To address this trade-off, we propose a novel bi-fidelity formulation of variational auto-encoders (BF-VAE) designed to estimate the uncertainty associated with a QoI from low-fidelity (LF) and high-fidelity (HF) samples of the QoI. Here, this model allows for the approximation of the statistics of the HF QoI by leveraging information derived from its LF counterpart. Specifically, we design a bi-fidelity auto-regressive model in the latent space which is integrated within the VAE’s probabilistic encoder–decoder structure. An effective algorithm is proposed to maximize the variational lower bound of the HF log-likelihood in the presence of limited HF data, resulting in the synthesis of HF realizations with a reduced computational cost. Additionally, we introduce the concept of the bi-fidelity information bottleneck (BF-IB) to provide an information-theoretic interpretation of the proposed BF-VAE model. Our numerical results demonstrate that the BF-VAE leads to considerably improved accuracy, as compared to a VAE trained using only HF data, when limited HF data is available.

42 ENGINEERING↗

Elucidating the Role of Hydrogen Bonding in the Optical Spectroscopy of the Solvated Green Fluorescent Protein Chromophore: Using Machine Learning to Establish the Importance of High-Level Electronic Structure

Hydrogen bonding interactions with chromophores in chemical and biological environments play a key role in determining their electronic absorption and relaxation processes, which are manifested in their linear and multidimensional optical spectra. For chromophores in the condensed phase, the large number of atoms needed to simulate the environment has traditionally prohibited the use of high-level excited-state electronic structure methods. By leveraging transfer learning, we show how to construct machine-learned models to accurately predict the high-level excitation energies of a chromophore in solution from only 400 high-level calculations. Here, we show that when the electronic excitations of the green fluorescent protein chromophore in water are treated using EOM-CCSD embedded in a DFT description of the solvent the optical spectrum is correctly captured and that this improvement arises from correctly treating the coupling of the electronic transition to electric fields, which leads to a larger response upon hydrogen bonding between the chromophore and water.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Quantum Hardware-Enabled Molecular Dynamics via Transfer Learning

The ability to perform ab initio molecular dynamics simulations using potential energy surfaces provided by quantum computers would open the door to virtually exact dynamics for a variety of chemical and biochemical systems, with impacts on catalysis and biophysics. Nonetheless, performing molecular dynamics on surfaces produced by quantum hardware has been hampered by the noisy energies typically produced by quantum computers and challenges associated with computing gradients and scaling to large systems interest. A recent set of advances in machine learning, known as transfer learning, provides a new path forward for molecular dynamics simulations on quantum hardware. Transfer learning offers a workaround, where one first trains models on larger, less accurate classical datasets and then refines them on smaller, more accurate quantum datasets. We explore this approach by training machine learning models to predict a molecule's potential energy based on its geometric structure using Behler-Parrinello neural networks. When successfully trained, the model enables energy gradient predictions necessary for dynamic simulations. To reduce the quantum resources needed, the model is initially trained with data derived from classical density functional theory and subsequently refined with a smaller dataset obtained from a variational quantum eigensolver optimization of the unitary coupled cluster ansatz. We show that this approach significantly reduces the size of the needed quantum training dataset while capturing the high accuracies needed within quantum chemistry simulations. The success of this two-step training method opens more opportunities to apply machine learning models on quantum data, a significant stride towards efficient quantum-classical hybrid computational models.

quantum computing↗

Accelerating Hydraulic Fracture Imaging by Deep Transfer Learning

Deep transfer learning has a great success story in computer vision (CV), natural language processing, and many other fields. In this communication, we are going to push forward the deep transfer learning to the hydraulic fracture imaging problem by proposing a two-step approach: 1) train a convolutional neural network (CNN) to reconstruct target geometries by a relatively large amount of approximated field patterns generated from a simplified model and 2) fine-tune the top layers of transferred CNN by a small amount of true field patterns generated through a full model. The advantages include the rapid generation of large amount of data through the simplified model and the high reconstruction accuracy through a careful design of the deep transfer learning. Here, the CNN trained through the deep transfer learning can accurately reconstruct the lateral extent and direction of fractures with unseen conductivity and white Gaussian noise, showing a notable acceleration/accuracy enhancement over the previous CNN trained by mixing a small number of true data into the approximated data as data augmentation.

02 PETROLEUM↗

Dynamics of Aqueous Electrolyte Solutions: Challenges for Simulations

This Perspective article focuses on recent simulation work on the dynamics of aqueous electrolytes. It is well-established that full-charge, nonpolarizable models for water and ions generally predict solution dynamics that are too slow in comparison to experiments. Models with reduced (scaled) charges do better for solution diffusivities and viscosities but encounter issues describing other dynamic phenomena such as nucleation rates of crystals from solution. Polarizable models show promise, especially when appropriately parametrized, but may still miss important physical effects such as charge transfer. First-principles calculations are starting to emerge for these properties that are in principle able to capture polarization, charge transfer, and chemical transformations in solution. Finally, while direct ab initio simulations are still too slow for simulations of large systems over long time scales, machine-learning models trained on appropriate first-principles data show significant promise for accurate and transferable modeling of electrolyte solution dynamics.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Steam Generator Model Design Parameter Sensitivity Study Using Advanced Optimization Tools

This study focuses on design parameter sensitivity studies pertaining to a steam generator (SG) model, using both Python and machine-learning tools. The SG model is a mathematical representation (including fluid flow and heat transfer equations/models/correlations) of a steam-generating unit in a pressurized water reactor (PWR)-type small modular reactor (SMR) system. Design studies involve changing the model’s input design parameters (e.g., temperature, pressure, mass flow rate) to observe the resulting effects on the output of the system (e.g., heat transfer coefficient [HTC], Nusselt number, heat transfer performance). Sensitivity studies analyze the degree to which system output and/or desired parameters (e.g., HTC or heat transfer performance) are sensitive to changes in input parameters. By using machine-learning tools such as the Risk Analysis Virtual Environment (RAVEN) developed at Idaho National Laboratory (INL), detailed design parametric sensitivity studies and model optimization were performed. Six input parameters—namely, the pressure, temperature, and mass flow rate for the inlet of the primary-side (hot fluid) and secondary-side (cold fluid) of the SG—were randomly perturbed via RAVEN’s Monte Carlo Sampler module, using uniform distributions (±1% relative changes). The analysis results give valuable insights into SG system performance and optimization, and provide justification for researching optimized sensor placement to effectively monitor and obtain experimental data.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Steam Generator Model Design Parameter Sensitivity Study Using Advanced Optimization Tools

This study focuses on design parameter sensitivity studies pertaining to a steam generator (SG) model, using both Python and machine-learning tools. The SG model is a mathematical representation (including fluid flow and heat transfer equations/models/correlations) of a steam-generating unit in a pressurized water reactor (PWR)-type small modular reactor (SMR) system. Design studies involve changing the model’s input design parameters (e.g., temperature, pressure, mass flow rate) to observe the resulting effects on the output of the system (e.g., heat transfer coefficient [HTC], Nusselt number, heat transfer performance). Sensitivity studies analyze the degree to which system output and/or desired parameters (e.g., HTC or heat transfer performance) are sensitive to changes in input parameters. By using machine-learning tools such as the Risk Analysis Virtual Environment (RAVEN) developed at Idaho National Laboratory (INL), detailed design parametric sensitivity studies and model optimization were performed. Six input parameters—namely, the pressure, temperature, and mass flow rate for the inlet of the primary-side (hot fluid) and secondary-side (cold fluid) of the SG—were randomly perturbed via RAVEN’s Monte Carlo Sampler module, using uniform distributions (±1% relative changes). The analysis results give valuable insights into SG system performance and optimization, and provide justification for researching optimized sensor placement to effectively monitor and obtain experimental data.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Beyond Point Estimates: Benchmarking Uncertainty Quantification Methods on the AION-1 Astronomical Foundation Model

Foundation models for astronomical surveys offer powerful learned representations that can be transferred to downstream regression tasks such as galaxy property estimation. However, point predictions alone are insufficient for scientific inference; reliable uncertainty quantification (UQ) is essential. We compare seven UQ methods on galaxy property regression using frozen AION-1 foundation-model embeddings, predicting redshift, stellar mass, stellar-population age, gas-phase metallicity, and specific star-formation rate, from Legacy Survey photometry/imaging and DESI spectra, with PROVABGS-derived labels. Distribution-free conformal methods achieve marginal coverage within $\sim$1 pp of the nominal 90% across all properties, while non-conformal baselines (Deep Ensembles, MC~Dropout) fail to calibrate reliably. Among conformal approaches, Conformalized Quantile Regression (CQR) delivers the best coverage in the bin with the poorest model predictions. More importantly, only the Locally Valid and Discriminative (LVD) framework -- particularly when operating on AION-1 embeddings -- also provides finite-sample \emph{local validity}, producing intervals that adapt to each galaxy's local prediction difficulty rather than relying on marginal guarantees alone. These results establish conformal prediction, and LVD in particular, as the preferred UQ framework for uncertainty-aware inference on foundation-model embeddings in astrophysics.

Tame-Narvaez, Karla [Fermilab] (ORCID:000000022249↗

The Rise of Neural Networks for Materials and Chemical Dynamics

Machine learning (ML) is quickly becoming a premier tool for modeling chemical processes and materials. ML-based force fields, trained on large data sets of high-quality electron structure calculations, are particularly attractive due their unique combination of computational efficiency and physical accuracy. This Perspective summarizes some recent advances in the development of neural network-based interatomic potentials. Designing high-quality training data sets is crucial to overall model accuracy. One strategy is active learning, in which new data are automatically collected for atomic configurations that produce large ML uncertainties. Another strategy is to use the highest levels of quantum theory possible. Transfer learning allows training to a data set of mixed fidelity. A model initially trained to a large data set of density functional theory calculations can be significantly improved by retraining to a relatively small data set of expensive coupled cluster theory calculations. These advances are exemplified by applications to molecules and materials.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Explicit solvent machine-learned coarse-grained model of sodium polystyrene sulfonate to capture polymer structure and dynamics

In this article, strongly charged polyelectrolytes (PEs) demonstrate complex solution behavior as a function of chain length, concentrations, and ionic strength. The viscosity behavior is important to understand and is a core quantity for many applications, but aspects remain a challenge. Molecular dynamics simulations using implicit solvent coarse-grained (CG) models successfully reproduce structure, but are often inappropriate for calculating viscosities. To address the need for CG models which reproduce viscoelastic properties of one of the most studied PEs, sodium polystyrene sulfonate (NaPSS), we report our recent efforts in using Bayesian optimization to develop CG models of NaPSS which capture both polymer structure and dynamics in aqueous solutions with explicit solvent. We demonstrate that our explicit solvent CG NaPSS model with the ML-BOP water model [Chan et al. Nat Commun 10, 379 (2019)] quantitatively reproduces NaPSS chain statistics and solution structure. The new explicit solvent CG model is benchmarked against diffusivities from atomistic simulations and experimental specific viscosities for short chains. We also show that our Bayesian-optimized CG model is transferable to larger chain lengths across a range of concentrations. Overall, this work provides a machine-learned model to probe the structural, dynamic, and rheological properties of polyelectrolytes such as NaPSS and aids in the design of novel, strongly charged polymers with tunable structural and viscoelastic properties

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Testing Paired Neural Network Models for Aftershock Identification

Aftershock sequences are a burden to real-time seismic monitoring. Cross-correlation can be used because aftershocks exhibit similar waveforms, but the method is computationally expensive. Deep learning may be an alternative, as it is computationally efficient, but great attention to training and testing is required in order to trust that the model can generalize to new aftershock sequences. This is problematic for aftershock sequences, because large-magnitude earthquakes are unpredictable and are globally widespread. Here, we test several paired neural network (PNN) models trained on a augmented (noise-added) earthquake dataset, to determine whether they can be generalized to process real aftershock sequences. Two aftershock datasets that were originally detected by cross-correlation and subsequently validated by an expert analyst were used. We found that current PNN models struggle to generalize to aftershock sequences. However, we identify approaches to improve training future PNN models and believe that improvements may be achieved by transfer learning.

58 GEOSCIENCES↗

Predicting Kyasanur forest disease in resource-limited settings using event-based surveillance and transfer learning

In recent years, the reports of Kyasanur forest disease (KFD) breaking endemic barriers by spreading to new regions and crossing state boundaries is alarming. Effective disease surveillance and reporting systems are lacking for this emerging zoonosis, hence hindering control and prevention efforts. We compared time-series models using weather data with and without Event-Based Surveillance (EBS) information, i.e., news media reports and internet search trends, to predict monthly KFD cases in humans. We fitted Extreme Gradient Boosting (XGB) and Long Short-Term Memory models at the national and regional levels. We utilized the rich epidemiological data from endemic regions by applying Transfer Learning (TL) techniques to predict KFD cases in new outbreak regions where disease surveillance information was scarce. Overall, the inclusion of EBS data, in addition to the weather data, substantially increased the prediction performance across all models. The XGB method produced the best predictions at the national and regional levels. The TL techniques outperformed baseline models in predicting KFD in new outbreak regions. Novel sources of data and advanced machine-learning approaches, e.g., EBS and TL, show great potential towards increasing disease prediction capabilities in data-scarce scenarios and/or resource-limited settings, for better-informed decisions in the face of emerging zoonotic threats.

60 APPLIED LIFE SCIENCES↗

A data-driven framework for predicting machining stability: employing simulated data, operational modal analysis, and enhanced transfer learning

Chatter, a self-excited vibration phenomenon, presents a significant challenge in machining operations, particularly in high-speed milling, where it can degrade tool life, reduce material removal efficiency, and compromise workpiece quality. Addressing this challenge requires a reliable predictive model that can accommodate the complex dynamics of various machining scenarios. This study introduces a novel, data-driven approach to predicting machining stability, leveraging over 140,000 simulated datasets and employing advanced techniques such as operational modal analysis (OMA), enhanced transfer learning (TL), and receptance coupling substructure analysis (RCSA). By integrating these methodologies, the framework effectively classifies and predicts chatter across diverse operational modes, achieving robust and accurate outcomes. Our model utilizes a Random Forest (RF) classifier trained with the comprehensive dataset, which demonstrates substantial improvements in both predictive accuracy and robustness. Specifically, the RF model achieved an accuracy rate of 85%, an area under the curve (AUC) of 0.90, and an F1 score of 0.88, underscoring its capability to adapt to varying machining configurations. These results highlight the framework’s potential to enhance operational efficiency and machining quality by providing reliable chatter predictions across a broad range of machining parameters. In conclusion, this research thus offers a significant advancement in predictive maintenance for machining processes, enabling more stable and efficient manufacturing operations.

42 ENGINEERING↗

Molecular dipole moment learning via rotationally equivariant derivative kernels in molecular-orbital-based machine learning

This study extends the accurate and transferable molecular-orbital-based machine learning (MOB-ML) approach to modeling the contribution of electron correlation to dipole moments at the cost of Hartree–Fock computations. A MOB pairwise decomposition of the correlation part of the dipole moment is applied, and these pair dipole moments could be further regressed as a universal function of MOs. The dipole MOB features consist of the energy MOB features and their responses to electric fields. An interpretable and rotationally equivariant derivative kernel for Gaussian process regression (GPR) is introduced to learn the dipole moment more efficiently. The proposed problem setup, feature design, and ML algorithm are shown to provide highly accurate models for both dipole moments and energies on water and 14 small molecules. To demonstrate the ability of MOB-ML to function as generalized density-matrix functionals for molecular dipole moments and energies of organic molecules, we further apply the proposed MOB-ML approach to train and test the molecules from the QM9 dataset. The application of local scalable GPR with Gaussian mixture model unsupervised clustering GPR scales up MOB-ML to a large-data regime while retaining the prediction accuracy. In addition, compared with the literature results, MOB-ML provides the best test mean absolute errors of 4.21 mD and 0.045 kcal/mol for dipole moment and energy models, respectively, when training on 110 000 QM9 molecules. The excellent transferability of the resulting QM9 models is also illustrated by the accurate predictions for four different series of peptides.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Comparison of multifidelity machine learning models for potential energy surfaces

Multifidelity modeling is a technique for fusing the information from two or more datasets into one model. It is particularly advantageous when one dataset contains few accurate results and the other contains many less accurate results. Within the context of modeling potential energy surfaces, the low-fidelity dataset can be made up of a large number of inexpensive energy computations that provide adequate coverage of the N-dimensional space spanned by the molecular internal coordinates. The high-fidelity dataset can provide fewer but more accurate electronic energies for the molecule in question. Here, we compare the performance of several neural network-based approaches to multifidelity modeling. We show that the four methods (dual, Δ-learning, weight transfer, and Meng–Karniadakis neural networks) outperform a traditional implementation of a neural network, given the same amount of training data. We also show that the Δ-learning approach is the most practical and tends to provide the most accurate model.

Chemistry↗