Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Transfer Learning Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Quantifying dislocation-type defects in post irradiation examination via transfer learning

The quantitative analysis of dislocation-type defects in irradiated materials is critical to materials characterization in the nuclear energy industry. The conventional approach of an instrument scientist manually identifying any dislocation defects is both time-consuming and subjective, thereby potentially introducing inconsistencies in the quantification. This work approaches dislocation-type defect identification and segmentation using a standard open-source computer vision model, YOLO11, that leverages transfer learning to create a highly effective dislocation defect quantification tool while using only a minimal number of annotated micrographs for training. This model demonstrates the ability to segment both dislocation lines and loops concurrently in micrographs with high pixel noise levels and on two alloys not represented in the training set. Inference of dislocation defects using transmission electron microscopy on three different irradiated alloys relevant to the nuclear energy industry are examined in this work with widely varying pixel noise levels and with completely unrelated composition and dislocation formations for practical post irradiation examination analysis. Code and models are available at https://github.com/idaholab/PANDA.

36 MATERIALS SCIENCE↗

Predicting fault slip via transfer learning

Abstract Data-driven machine-learning for predicting instantaneous and future fault-slip in laboratory experiments has recently progressed markedly, primarily due to large training data sets. In Earth however, earthquake interevent times range from 10’s-100’s of years and geophysical data typically exist for only a portion of an earthquake cycle. Sparse data presents a serious challenge to training machine learning models for predicting fault slip in Earth. Here we describe a transfer learning approach using numerical simulations to train a convolutional encoder-decoder that predicts fault-slip behavior in laboratory experiments. The model learns a mapping between acoustic emission and fault friction histories from numerical simulations, and generalizes to produce accurate predictions of laboratory fault friction. Notably, the predictions improve by further training the model latent space using only a portion of data from a single laboratory earthquake-cycle. The transfer learning results elucidate the potential of using models trained on numerical simulations and fine-tuned with small geophysical data sets for potential applications to faults in Earth.

58 GEOSCIENCES↗

HydraGNN_OPF_GFM_2026 - Ensemble of predictive graph foundation models for power grid applications

This dataset supports research on graph foundation models for optimal power flow (OPF) on electric grids using HydraGNN. It contains heterogeneous graph representations of PGLib-OPF cases spanning systems from 14 to 13,659 buses, together with packed HDF5 datasets for pretraining, feasibility classification, and N-1 contingency analysis. The release includes OPF solution data, downstream fine-tuning datasets, pretrained HeteroSAGE and HeteroHEAT model checkpoints, hyperparameter-optimization summaries across multiple heterogeneous GNN architectures, and aggregated fine-tuning results for sample-efficiency studies. The dataset is designed to enable scalable training, evaluation, and transfer-learning studies for OPF surrogate modeling, including node-level AC-OPF solution prediction, graph-level prediction, feasibility classification, operating-condition generalization, and contingency-response tasks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Quantum AI Based Enhanced Detection of Dementia

Quantum computing has the potential to significantly improve the early detection of Alzheimer's Disease and Related Dementias (ADRD). Quantum-enhanced machine learning can be used to perform an early screening of Alzheimer's disease using brain imaging data based on dataset of MRI scans from both healthy individuals and those diagnosed with Alzheimer's. This study aims to demonstrate the potential of quantum transfer learning to enhance the performance of the classical deep learning model for dementia detection. Using the MRI sagittal images available in the OASIS-2 (64 demented and 72 non-demented subjects between 60 and 96 years), we show how quantum techniques can transform a suboptimal classical model into a more effective solution for dementia detection, highlighting their potential impact on advancing healthcare technology. We begin with a simple classical deep learning model with a significantly smaller number of parameters, which gives suboptimal performance on the problem. Then, we apply different configurations of quantum transfer learning based on the pre-trained weak classifier (Figure 1). We fix the weak classifier's initial convolutional layers at their fixed pre-trained parameters and replace the last set of dense layers with a dressed quantum circuit (DQN), which we train to enhance performance. We performed 4-fold cross-validation for both the classical and the hybrid quantum models and trained them using Pennylane's `default.qubit' simulator and IonQ's Aria-1 simulator (noisy simulation). We showed that with significantly fewer parameters, the quantum transfer learning-based hybrid models showed significant performance enhancement over the base weak classical deep learning model for dementia detection. To classify between a demented and non-demented subject, the accuracy of quantum-based AI methods improved by 6 to 14% compared to classical methods. The sensitivity of the models improved by 4 to 17%. This shows that there are fewer chances of misclassifying demented patients. Figure 2 compares the performance of the hybrid quantum models and their base classical model, and Table 1 summarizes the results. We illustrated that with assistance from quantum machine learning, it is possible to enhance detection for dementia based on brain images. This shows the potential for practical utility of quantum computing in ADRD research.

Bhowmik, Sounak [University of Tennessee, Knoxvill↗

Model Form Error Correction for a Black-Box Thermal Battery Heat Transfer Simulation

Thermal batteries are crucial for supplying power to high-consequence engineering applications such as rockets. Computational simulations have been developed to predict thermal battery behavior, but these simulations often suffer from modeling errors, including model form uncertainty. Addressing this uncertainty can be achieved by quantifying either the model discrepancy in the output or the model form error (MFE) in the governing equation. MFE is particularly valuable as it can be better extrapolated beyond observed outputs, which is essential for predictions involving changes in external system loading, system configuration and geometry, or output quantities. This paper employs a state estimation approach to estimate MFE using experimental data and then utilizes machine learning (ML) to model its relationship with state variables. A nonintrusive technique is used to estimate MFE in a black-box thermal battery heat transfer simulation. The trained machine learning model for MFE is then applied to correct simulation predictions under extrapolated initial conditions and battery configurations. In conclusion, the methodology's performance is evaluated using additional experimental data, demonstrating its effectiveness in improving prediction accuracy.

Batteries↗

Cognitive simulation models for inertial confinement fusion: Combining simulation and experimental data

The design space for inertial confinement fusion (ICF) experiments is vast, and experiments are extremely expensive. Researchers rely heavily on computer simulations to explore the design space in search of high-performing implosions. However, ICF multiphysics codes must make simplifying assumptions, and thus deviate from experimental measurements for complex implosions. For more effective design and investigation, simulations require input from past experimental data to better predict future performance. In this work, we describe a cognitive simulation method for combining simulation and experimental data into a common, predictive model. This method leverages a machine learning technique called “transfer learning,” the process of taking a model trained to solve one task, and partially retraining it on a sparse dataset to solve a different, but related task. In the context of ICF design, neural network models are trained on large simulation databases and partially retrained on experimental data, producing models that are far more accurate than simulations alone. Here, we demonstrate improved model performance for a range of ICF experiments at the National Ignition Facility and predict the outcome of recent experiments with less than 10% error for several key observables. We discuss how the methods might be used to carry out a data-driven experimental campaign to optimize performance, illustrating the key product—models that become increasingly accurate as data are acquired.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Physics-informed neural network with transfer learning (TL-PINN) based on domain similarity measure for prediction of nuclear reactor transients

Nuclear reactor safety and efficiency can be enhanced through the development of accurate and fast methods for prediction of reactor transient (RT) states. Physics informed neural networks (PINNs) leverage deep learning methods to provide an alternative approach to RT modeling. Applications of PINNs in monitoring of RTs for operator support requires near real-time model performance. However, as with all machine learning models, development of a PINN involves time-consuming model training. Here, we show that a transfer learning (TL-PINN) approach achieves significant performance gain, as measured by reduction of the number of iterations for model training. Using point kinetic equations (PKEs) model with six neutron precursor groups, constructed with experimental parameters of the Purdue University Reactor One (PUR-1) research reactor, we generated different RTs with experimentally relevant range of variables. The RTs were characterized using Hausdorff and Fréchet distance. We have demonstrated that pre-training TL-PINN on one RT results in up to two orders of magnitude acceleration in prediction of a different RT. The mean error for conventional PINN and TL-PINN models prediction of neutron densities is smaller than 1%. We have developed a correlation between TL-PINN performance acceleration and similarity measure of RTs, which can be used as a guide for application of TL-PINNs.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Advancing spatiotemporal forecasts of CO 2 plume migration using deep learning networks with transfer learning and interpretation analysis

Accurate and timely forecasts of CO 2 plume distribution throughout the injection and post-injection phases are crucial for detecting plume migration, assessing leakage risks, and supporting operational decisions in geologic carbon storage (GCS). Current convolutional neural network-based approaches primarily focus on spatial information and overlook temporal dependencies in plume distributions, thus limiting their ability to capture dynamic movement effects and provide accurate predictions of plume migration. In this work, we propose two deep learning models, Auto-Encoder (AE)-LSTM and Encoder-Decoder (ED)-ConvLSTM, each uniquely designed to capture both spatial and temporal features. We apply the proposed methods to forecast the dynamic distribution of CO 2 plumes based on 108 reservoir simulations over a 30-year injection and a 30-year post-injection period. The results indicate that the ED-ConvLSTM model outperforms the AE-LSTM model in accurately predicting the spatiotemporal dynamics of CO 2 plume migration, achieving R 2 values above 0.99. To provide a deeper understanding of these model predictions, we employ a gradient-based explanation method on the trained models. This approach provides insights into the influence of input variables on plume migration forecasts and uncovers the underlying prediction mechanisms of the proposed models. Furthermore, we introduce a transfer learning technique, enabling fast and accurate plume migration forecasting in the post-injection phase by leveraging the trained model during the injection phase. This reduces the necessity for extensive data collection or re-training. In conclusion, the methods proposed in our work enhances the performance and interpretability of CO 2 plume migration forecasts, thereby facilitating informed decision-making throughout the entire lifecycle of GCS applications.

58 GEOSCIENCES↗

A transfer learning approach for acoustic emission zonal localization on steel plate-like structure using numerical simulation and unsupervised domain adaptation

The detection and localization of damage in metallic structures using acoustic emission (AE) monitoring and artificial intelligence technology such as deep learning has been widely studied. However, a current challenge of this approach is the difficulty of obtaining sufficient labeled historical AE signals for the training process of deep learning models. This problem can be approached through the implementation of transfer learning. The innovation of this paper lies in the development of a transfer learning approach for AE source localization on a stainless-steel structure when no historical labeled AE signals are available for training. A finite element model is developed to generate numerical AE signals for the training. Unsupervised domain adaptation (UDA) technology is utilized to reduce the distribution difference between the numerical and the realistic AE signals and to derive the localization results of the unlabeled realistic AE signals. Finally, the results suggest that the proposed approach is capable of localizing AE signals with high accuracy in the absence of labeled training data.

42 ENGINEERING↗

Optical information transfer through random unknown diffusers using electronic encoding and diffractive decoding

Free-space optical information transfer through diffusive media is critical in many applications, such as biomedical devices and optical communication, but remains challenging due to random, unknown perturbations in the optical path. We demonstrate an optical diffractive decoder with electronic encoding to accurately transfer the optical information of interest, corresponding to, e.g., any arbitrary input object or message, through unknown random phase diffusers along the optical path. This hybrid electronic-optical model, trained using supervised learning, comprises a convolutional neural network-based electronic encoder and successive passive diffractive layers that are jointly optimized. After their joint training using deep learning, our hybrid model can transfer optical information through unknown phase diffusers, demonstrating generalization to new random diffusers never seen before. The resulting electronic-encoder and optical-decoder model was experimentally validated using a 3D-printed diffractive network that axially spans <70λ, where λ = 0.75 mm is the illumination wavelength in the terahertz spectrum, carrying the desired optical information through random unknown diffusers. The presented framework can be physically scaled to operate at different parts of the electromagnetic spectrum, without retraining its components, and would offer low-power and compact solutions for optical information transfer in free space through unknown random diffusive media.

36 MATERIALS SCIENCE↗

Transfer learning of memory kernels for transferable coarse-graining of polymer dynamics

The present work concerns the transferability of coarse-grained (CG) modeling in reproducing the dynamic properties of the reference atomistic systems across a range of parameters. In particular, we focus on implicit-solvent CG modeling of polymer solutions. The CG model is based on the generalized Langevin equation, where the memory kernel plays the critical role in determining the dynamics in all time scales. Thus, we propose methods for transfer learning of memory kernels. The key ingredient of our methods is Gaussian process regression. By integration with the model order reduction via proper orthogonal decomposition and the active learning technique, the transfer learning can be practically efficient and requires minimum training data. Through two example polymer solution systems, we demonstrate the accuracy and efficiency of the proposed transfer learning methods in the construction of transferable memory kernels. The transferability allows for out-of-sample predictions, even in the extrapolated domain of parameters. Built on the transferable memory kernels, the CG models can reproduce the dynamic properties of polymers in all time scales at different thermodynamic conditions (such as temperature and solvent viscosity) and for different systems with varying concentrations and lengths of polymers.

Ma, Zhan↗

Deep structural clustering for single-cell RNA-seq data jointly through autoencoder and graph neural network

Abstract Single-cell RNA sequencing (scRNA-seq) permits researchers to study the complex mechanisms of cell heterogeneity and diversity. Unsupervised clustering is of central importance for the analysis of the scRNA-seq data, as it can be used to identify putative cell types. However, due to noise impacts, high dimensionality and pervasive dropout events, clustering analysis of scRNA-seq data remains a computational challenge. Here, we propose a new deep structural clustering method for scRNA-seq data, named scDSC, which integrate the structural information into deep clustering of single cells. The proposed scDSC consists of a Zero-Inflated Negative Binomial (ZINB) model-based autoencoder, a graph neural network (GNN) module and a mutual-supervised module. To learn the data representation from the sparse and zero-inflated scRNA-seq data, we add a ZINB model to the basic autoencoder. The GNN module is introduced to capture the structural information among cells. By joining the ZINB-based autoencoder with the GNN module, the model transfers the data representation learned by autoencoder to the corresponding GNN layer. Furthermore, we adopt a mutual supervised strategy to unify these two different deep neural architectures and to guide the clustering task. Extensive experimental results on six real scRNA-seq datasets demonstrate that scDSC outperforms state-of-the-art methods in terms of clustering accuracy and scalability. Our method scDSC is implemented in Python using the Pytorch machine-learning library, and it is freely available at https://github.com/DHUDBlab/scDSC.

Gan, Yanglan↗

Applying Deep Learning for Wildfire Identification: Economical and Accessible Solutions Leveraging Small Datasets

Wildfires significantly impact human health, air quality, visibility, weather, and climate change and cause substantial economic losses. While state and county-operated air quality monitors provide critical insights during wildfires, they are not available in all regions. This highlights the need for affordable, accessible tools that allow the general public to assess air quality impacts. In this study, we apply machine learning with deep neural networks to diagnose air quality rapidly from sky images taken at the Pacific Northwest National Laboratory in Richland, WA, USA. Using a convolutional neural network (CNN) framework, we trained a deep learning model to classify air quality indices based on sky images. By leveraging transfer learning, our approach fine-tunes a pre-trained model on a small dataset of sky images, significantly reducing training time while maintaining high accuracy. Our results demonstrate the potential of deep learning to provide rapid air quality diagnostics during wildfire episodes, offering early warnings to the public and enabling timely mitigation strategies, particularly for vulnerable populations. Additionally, we show that lower respiratory infections pose the highest health risk during acute smoke exposures. Reactive oxygen species (ROS) from wildfire particles further exacerbate health risks by triggering inflammation and other adverse effects.

54 ENVIRONMENTAL SCIENCES↗

Preliminary Transfer Learning Results on Israel Data

In this preliminary report, we use publicly available data recorded in Israel to test and expand upon existing machine learning models for seismic-phase detection and arrival-time measurement. We downloaded 3-years of waveform data from Geofon, and cross referenced the waveforms to Israel bulletin picks (Schardong et al., 2021). The initial results using existing models directly generated ubiquitous false detections and that obscured detections of signals that are clearly visible in the waveforms. However, after applying transfer learning (tuning parameters in the existing ML models using one year of the Israel-network data), the results are encouraging, i.e. ML picks agree within a few tenths of a second with bulletin picks and the number of false detections is greatly reduced. The bulletin picks are a good starting point, but they cannot be considered ground-truth. To test potential improvement in picking using ML we would like to relocate the events using the ML picks to see if the events cluster more tightly at known mine locations. However, in order to constrain event locations, we need ML picks for the whole Israeli-Jordanian network, which requires waveforms that are not publicly available.

58 GEOSCIENCES↗

New Neural Network Cloud Mask Algorithm Based on Radiative Transfer Simulations

Cloud detection and screening constitute critically important first steps required to derive many satellite data products. Traditional threshold-based cloud mask algorithms require a complicated design process and fine tuning for each sensor, and they have difficulties over areas partially covered with snow/ice. Exploiting advances in machine learning techniques and radiative transfer modeling of coupled environmental systems, we have developed a new, threshold-free cloud mask algorithm based on a neural network classifier driven by extensive radiative transfer simulations. Statistical validation results obtained by using collocated CALIOP and MODIS data show that its performance is consistent over different ecosystems and significantly better than the MODIS Cloud Mask (MOD35 C6) during the winter seasons over snow-covered areas in the mid-latitudes. Simulations using a reduced number of satellite channels also show satisfactory results, indicating its flexibility to be configured for different sensors. Comparedto threshold-based methods and previous machine-learning approaches, this new cloud mask (i) does not rely on thresholds, (ii) needs fewer satellite channels, (iii) has superior performance during winter seasons in mid-latitude areas, and (iv) can easily be applied to different sensors.

cloud mask algorithms↗

Machine learning for reactor power monitoring with limited labeled data

Real-time reactor power monitoring is critical for a variety of nuclear applications, spanning safety, security, operations, and maintenance. While machine learning methods have shown promise in monitoring reactor power levels, there is limited research on their efficacy in label-starved environments. The goal of this work is to assess the feasibility of classifying nuclear reactor power level using multisource data in scenarios with limited labels. Data were collected using low-resolution multisensors at four nuclear reactor facilities: two large research reactors and two TRIGA reactors. Within each pair, one reactor dataset served as the source and the other as the target in a transfer learning paradigm. Twenty-three supervised models were trained on labeled sequences of magnetic field and acceleration data from each of the target sites. Self-learning and transfer learning methods were applied to the top performing models to assess their classification performance with increasing amounts of labeled data. While reactor power level classification was achieved with a Matthews Correlation Coefficient of up to 0.739 ± 0.003 and 0.622 ± 0.009 with only 400 sequences per power state for the large research reactor and TRIGA target sites, respectively, self-learning and transfer learning leveraging source site data did not improve target classification performance. These findings suggest that alternative methods, such as higher sensitivity sensors, digital twins, or the use of physics-informed models, are required to enable high-performance classification in machine learning approaches to reactor monitoring with a dearth of target ground truth.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Transfer Learning of High-Fidelity Opacity Spectra in Autoencoders and Surrogate Models

Simulations of high energy density physics are expensive, largely in part for the need to produce nonlocal thermodynamic equilibrium opacities. High-fidelity spectra may reveal new physics in the simulations not seen with low-fidelity spectra, but the cost of these simulations also scales with the level of fidelity of the opacities being used. Neural networks are capable of reproducing these spectra, but neural networks need data to train them, which limits the level of fidelity of the training data. Here this article demonstrates that it is possible to reproduce high-fidelity spectra with median errors in the realm of 3%–4% using as few as 50 samples of high-fidelity Krypton data by performing transfer learning on a neural network trained on many times more low-fidelity data.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗