Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Transfer Learning Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

An adaptive adversarial domain adaptation approach for corn yield prediction

Recently, statistical machine learning and deep learning methods have been widely explored for corn yield prediction. Though successful, machine learning models generated within a specific spatial domain often lose their validity when directly applied to new regions. To address this issue, we designed an unsupervised adaptive domain adversarial neural network (ADANN). Specifically, through domain adversarial training, the ADANN model reduced the impact of domain shift by projecting data from different domains into the same subspace. Also, the ADANN model was designed to be trained in an adaptive way, which guaranteed the model can learn the domain-invariant features and perform accurate yield prediction simultaneously. Informative variables including time-series vegetation indices and sequential weather observations were first collected from multiple data sources and aggregated to the county level. Then, we trained the ADANN model with the extracted features and corresponding reported county-level corn yield from the U.S. Department of Agriculture (USDA). Finally, the trained model was evaluated in four testing years 2016–2019. The U.S. corn belt was used as the study area and counties under study were grouped into two diverse ecological regions. Overall, the experimental results showed that the developed ADANN model had better performance than three other state-of-the-art machine learning models in both local experiments (train and test in the same region) and transfer experiments (train and test in different regions). As the first study using adversarial learning for crop yield prediction, this research demonstrates a novel solution for improving model transferability on crop yield prediction.

59 BASIC BIOLOGICAL SCIENCES↗

High-Accuracy Semiempirical Quantum Models Based on a Minimal Training Set

A great need exists for computationally efficient quantum simulation approaches that can achieve an accuracy similar to high-level theories at a fraction of the computational cost. In this regard, we have leveraged a machine-learned interaction potential based on Chebyshev polynomials to improve density functional tight binding (DFTB) models for organic materials. The benefit of our approach is two-fold: (1) many-body interactions can be corrected for in a systematic and rapidly tunable process, and (2) high-level quantum accuracy for a broad range of compounds can be achieved with ~0.3% of data required for one advanced deep learning potential. Our model exhibits both transferability and extensibility through comparison to quantum chemical results for organic clusters, solid carbon phases, and molecular crystal phase stability rankings. Overall, our efforts thus allow for high-throughput physical and chemical predictions with up to coupled-cluster accuracy for systems that are computationally intractable with standard approaches.

36 MATERIALS SCIENCE↗

BioADAPT-MRC: adversarial learning-based domain adaptation improves biomedical machine reading comprehension task

ABSTRACT Motivation Biomedical machine reading comprehension (biomedical-MRC) aims to comprehend complex biomedical narratives and assist healthcare professionals in retrieving information from them. The high performance of modern neural network-based MRC systems depends on high-quality, large-scale, human-annotated training datasets. In the biomedical domain, a crucial challenge in creating such datasets is the requirement for domain knowledge, inducing the scarcity of labeled data and the need for transfer learning from the labeled general-purpose (source) domain to the biomedical (target) domain. However, there is a discrepancy in marginal distributions between the general-purpose and biomedical domains due to the variances in topics. Therefore, direct-transferring of learned representations from a model trained on a general-purpose domain to the biomedical domain can hurt the model’s performance. Results We present an adversarial learning-based domain adaptation framework for the biomedical machine reading comprehension task (BioADAPT-MRC), a neural network-based method to address the discrepancies in the marginal distributions between the general and biomedical domain datasets. BioADAPT-MRC relaxes the need for generating pseudo labels for training a well-performing biomedical-MRC model. We extensively evaluate the performance of BioADAPT-MRC by comparing it with the best existing methods on three widely used benchmark biomedical-MRC datasets—BioASQ-7b, BioASQ-8b and BioASQ-9b. Our results suggest that without using any synthetic or human-annotated data from the biomedical domain, BioADAPT-MRC can achieve state-of-the-art performance on these datasets. Availability and implementation BioADAPT-MRC is freely available as an open-source project at https://github.com/mmahbub/BioADAPT-MRC. Supplementary information Supplementary data are available at Bioinformatics online.

60 APPLIED LIFE SCIENCES↗

Deep learning-based spatio-temporal fusion for high-fidelity ultra-high-speed X-ray radiography

Full-field ultra-high-speed (UHS) X-ray imaging experiments have been well established to characterize various processes and phenomena. However, the potential of UHS experiments through the joint acquisition of X-ray videos with distinct configurations has not been fully exploited. In this paper, we investigate the use of a deep learning-based spatio-temporal fusion (STF) framework to fuse two complementary sequences of X-ray images and reconstruct the target image sequence with high spatial resolution, high frame rate and high fidelity. We applied a transfer learning strategy to train the model and compared the peak signal-to-noise ratio (PSNR), average absolute difference (AAD) and structural similarity (SSIM) of the proposed framework on two independent X-ray data sets with those obtained from a baseline deep learning model, a Bayesian fusion framework and the bicubic interpolation method. The proposed framework outperformed the other methods with various configurations of the input frame separations and image noise levels. With three subsequent images from the low-resolution (LR) sequence of a four times lower spatial resolution and another two images from the high-resolution (HR) sequence of a 20 times lower frame rate, the proposed approach achieved average PSNRs of 37.57 dB and 35.15 dB, respectively. When coupled with the appropriate combination of high-speed cameras, the proposed approach will enhance the performance and therefore the scientific value of UHS X-ray imaging experiments.

deep learning↗

Automated Fire Detection for Industrial Settings with Pretrained Convolutional Networks

Early fire detection in industrial environments is critical to preventing equipment damage, personal injury, and operational disruptions. Traditional smoke detectors, while effective, often experience delays due to the time required for smoke to reach sensors, allowing fires to spread. Manual fire watch operations and human surveillance of camera feeds are resource-intensive and prone to human error. To address these challenges, this paper explores the application of convolutional neural networks for automated fire detection, specifically in industrial settings. By leveraging 11 different pre-trained machine vision models from TensorFlow and enhancing them with transfer learning on a custom-built industrial fire dataset, we optimized fire detection performance. Here, we analyzed each machine vision model architecture in terms of its depth, width, and input image resolution, considering both resource requirements and detection accuracy. We further explored the option of combining multiple models into an ensemble classifier to evaluate whether the performance improvements could justify the much greater computational complexity and other practical impacts. A cost-benefit analysis is presented to evaluate the trade-offs between performance and computational expense. Our findings identify that EfficientNetV2L, specifically tailored for industrial applications, provides the optimal balance between costs involved in training and using the model versus the overall fire detection performance. Additionally, we present a qualitative analysis of model performance using the technique of gradient-based class activation mapping to provide explainability by visualizing model decisions.

artificial intelligence↗

Enhancing transfer learning in angle-resolved photoemission spectroscopy (ARPES) with spatially-aware representations via graph convolution

A recent application of machine learning has been to spatially-resolved angle-resolved photoemission spectroscopy (ARPES). Here we advance the state-of-the-art by applying representational learning to transform ARPES data into an embedding space of a pre-trained self-supervised learning model, thus enhancing the pipeline that improves the bandstructure classification and domain assignment/segmentation performance compared to a k-means clustering method. In the current iteration, the real-space information is entered into the domain assignment through the graph convolution method, which improves the transfer learning performance of the original self-supervised model. Lastly, an unsupervised automated tool is developed that incorporates these techniques to enable automatic domain assignment.

ARPES↗

Sharing is caring: An extensive analysis of parameter-based transfer learning for the prediction of building thermal dynamics

In recent years deep neural networks have been proposed as a lightweight data-driven model to capture high-dimensional, nonlinear physical processes to predict building thermal responses. However, the need of a large amount of data for the training process of deep neural networks clashes with the potential limited data availability in most existing or new buildings. Transfer learning aims to enhance the performance of a target learner exploiting knowledge from related and similar environments. This study conducted a suite of experiments that leveraged 250 data-driven models based on a synthetic dataset of a building archetype to study the influence of data availability, energy efficiency level, occupancy and climate for the transfer process of thermal dynamics. The performance of the transfer learning process was compared against a classical machine learning approach. Here, the results suggest that building thermal dynamics can be effectively transferred under the same climatic conditions, increasing performance when dealing with different occupancy schedules, efficiency levels and low data availability. Furthermore, the paper compares the performance of both transfer learning and machine learning approaches in an online fashion, to support the implementation in real-world deployment.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Molecular simulation using transfer-learned potentials for the disordered nanoscale structure of nitrogen-doped nanoporous carbons

Machine learning (ML)-based molecular dynamics (MD) simulations of the formation of a class of N-doped nanoporous carbons are performed to assess their disordered partially graphitized nanoscale structure. The study is motivated by the effectiveness of so-called nitrogen assembly carbons (NACs) for catalysis applications. Benchmark simulations for pure-C disordered graphitic systems reveal the importance of reliably capturing the vdW component of the potentials in order to accurately describe the tendency for layering of disordered graphene-like sheets. In our modeling, this is achieved by a transfer learning strategy incorporating features of the energetics from the optB88-vdW DFT functional into potentials initially trained with a less expensive functional, thereby providing a superior description of the pure-C systems. Generation from MD simulations of realistic partially graphitized structures is significantly more challenging for N-doped versus for pure C systems. However, such structures are achieved by a tailored MD simulation protocol mimicking the experimental synthesis process and in particular incorporating an annealing and subsequent quenching stages. Simulated PXRD patterns effectively reproduce the features of experimental observations for NACs, including the appearance of a prominent but broad (002) peak at around 25, and the development of another weaker feature associated with in-layer ordering of mixed C-N graphene-like sheets.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Illuminating the Material World: Autonomous Microscopy to Understand Order, Disorder, and Everything In Between

Artificial intelligence (AI) holds immense promise for revolutionizing microscopy, yet its widespread adoption has been hindered by challenges ranging from user inexperience to limited model transferability and difficulties in operationalizing machine learning. This presentation showcases our approach to developing practical autonomy for materials discovery, aiming to accelerate the integration of AI into everyday microscopy workflows. As shown in Fig. 1, I will focus on three key areas: understanding order-disorder transitions, quantifying point defects, and achieving truly device-scale microscopy. First, I will demonstrate the power of multi-modal knowledge graphs for integrating diverse microscopy data. By combining imaging, spectroscopy, and diffraction data, these graphs provide a holistic view of material behavior, capturing the intricate relationships between different modalities [1,2]. I will present a case study on how these models illuminate the structural and chemical changes associated with irradiation in oxide thin films, revealing critical insights for designing materials for extreme environments like spaceflight and nuclear energy. Specifically, I will show how multi-modal analysis clarifies the evolution of order-disorder transitions under irradiation, a key factor influencing material performance in these applications. Next, I will address the challenge of quantifying point defects in 2D materials. We demonstrate the application of computer vision and transfer learning to accurately identify and classify various defect types, such as vacancies and substitutional atoms, and to quantify their concentrations. This information is crucial for understanding and tailoring the properties of 2D materials for applications in electronics, optoelectronics, and catalysis. For example, I will show how our models can characterize the topological distribution of point defects in MXene transition metal carbides, providing valuable insights for optimizing their performance in energy storage and separation science. Finally, I will discuss our progress toward autonomous device-scale microscopy [3,4]. We are fundamentally redesigning electron microscopes around the principles of machine reasoning, enabling automation beyond basic tasks like sample navigation and data acquisition to include sophisticated experimental design. This approach paves the way for truly reproducible and massively scaled analysis campaigns. I will emphasize the importance of autonomous microscopy platforms for high-throughput materials discovery and characterization, facilitating the rapid screening of materials for a broad range of applications and accelerating the development of next-generation technologies.

36 MATERIALS SCIENCE↗

Data Validation Experiments with a Computer-Generated Imagery Dataset for International Nuclear Safeguards

Computer vision models have great potential as tools for international nuclear safeguards verification activities, but off-the-shelf models require fine-tuning through transfer learning to detect relevant objects. Because open-source examples of safeguards-relevant objects are rare, and to evaluate the potential of synthetic training data for computer vision, we present the Limbo dataset. Limbo includes both real and computer-generated images of uranium hexafluoride containers for training computer vision models. Here, we generated these images iteratively based on results from data validation experiments that are detailed here. The findings from these experiments are applicable both for the safeguards community and the broader community of computer vision research using synthetic data.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Probabilistic locked mode predictor in the presence of a resistive wall and finite island saturation in tokamaks

We present a framework for estimating the probability of locking to an error field in a rotating tokamak plasma. This leverages machine learning methods trained on data from a mode-locking model, including an error field, resistive magnetohydrodynamics modeling of the plasma, a resistive wall, and an external vacuum region, leading to a fifth-order ordinary differential equation (ODE) system. It is an extension of the model without a resistive wall introduced by Akçay et al. [Phys. Plasmas 28, 082106 (2021)]. Tearing mode saturation by a finite island width is also modeled. We vary three pairs of control parameters in our studies: the momentum source plus either the error field, the tearing stability index, or the island saturation term. The order parameters are the time-asymptotic values of the five ODE variables. Normalization of them reduces the system to 2D and facilitates the classification into locked (L) or unlocked (U) states, as illustrated by Akçay et al., [Phys. Plasmas 28, 082106 (2021)]. This classification splits the control space into three regions: L̂, with only L states; Û, with only U states; and a hysteresis (hysteretic) region Ĥ, with both L and U states. In regions L̂ and Û, the cubic equation of torque balance yields one real root. Region Ĥ has three roots, allowing bifurcations between the L and U states. The classification of the ODE solutions into L/U is used to estimate the locking probability, conditional on the pair of the control parameters, using a neural network. We also explore estimating the locking probability for a sparse dataset, using a transfer learning method based on a dense model dataset.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Reducing Southern Ocean Shortwave Radiation Errors in the ERA5 Reanalysis with Machine Learning and 25 Years of Surface Observations

Earth system models struggle to simulate clouds and their radiative effects over the Southern Ocean, partly due to a lack of measurements and targeted cloud microphysics knowledge. We have evaluated biases of downwelling shortwave radiation in the ERA5 climate reanalysis using 25 years (1995–2019) of summertime surface measurements, collected on the Research and Supply Vessel (RSV) Aurora Australis, the Research Vessel (R/V) Investigator, and at Macquarie Island. During October–March daylight hours, the ERA5 simulation of SW down exhibited large errors (mean bias = 54 W m -2 , mean absolute error = 82 W m -2 , root-mean-square error = 132 W m -2 , and R 2 = 0.71). To determine whether we could improve these statistics, we bypassed ERA5’s radiative transfer model for SW down with machine learning–based models using a number of ERA5’s gridscale meteorological variables as predictors. These models were trained and tested with the surface measurements of SW down using a 10-fold shuffle split. An extreme gradient boosting (XGBoost) and a random forest–based model setup had the best performance relative to ERA5, both with a near complete reduction of the mean bias error, a decrease in the mean absolute error and root-mean-square error by 25% ± 3%, and an increase in the R 2 value of 5% ± 1% over the 10 splits. Large improvements occurred at higher latitudes and cyclone cold sectors, where ERA5 performed most poorly. We further interpret our methods using Shapley additive explanations. Our results indicate that data-driven techniques could have an important role in simulating surface radiation fluxes and in improving reanalysis products.

54 ENVIRONMENTAL SCIENCES↗

Graph reinforcement learning for exploring model spaces beyond the standard model

We present a methodology for performing scans of beyond the standard model (BSM) parameter spaces with reinforcement learning. We identify a novel procedure using graph neural networks that is capable of exploring spaces of models without the user specifying a fixed particle content, allowing broad classes of BSM models to be explored—in theory, the technique is applicable to nearly any model space with a prespecified gauge group. We provide a generic procedure by which a suitable graph grammar can be developed for any BSM model that features user-specified symmetry groups and a finite number of different possible particle species, the use of which is applicable to a variety of machine learning tasks over the actions of BSM theories beyond our particular reinforcement learning use case. As a proof of concept, we construct the graph grammar for theories with vectorlike leptons that may or may not be charged under a dark U ( 1 ) group, inspired by portal matter extensions of the sub-GeV vector portal/kinetic mixing simplified dark matter models. We then use this graph grammar to create a reinforcement learning environment tasked with creating models with these vectorlike leptons that are consistent with a list of a variety of precision observables. The reinforcement learning agent succeeds in developing models that can address the observed muon anomalous magnetic moment discrepancy while remaining consistent with flavor violation and electroweak precision observables, including both constructions that have previously been studied as well as new models that have not, to our knowledge, previously been identified. By inspecting the resulting ensembles of models that the agent produces and experimenting with different configurations for our reinforcement learning environment and graph grammar, we also infer various lessons about the development of these environments that can be transferable to reinforcement learning scans of more complicated model spaces and comment on future directions for the development of this technique into a more mature tool. Published by the American Physical Society 2025

Wojcik, George N.↗

Combining synchrotron X-ray diffraction, mechanistic modeling and machine learning for in situ subsurface temperature quantification during laser melting

Laser melting, such as that encountered during additive manufacturing, produces extreme gradients of temperature in both space and time, which in turn influence microstructural development in the material. Qualification and model validation of the process itself and the resulting material necessitate the ability to characterize these temperature fields. However, well established means to directly probe the material temperature below the surface of an alloy while it is being processed are limited. To address this gap in characterization capabilities, a novel means is presented to extract subsurface temperature-distribution metrics, with uncertainty, from in situ synchrotron X-ray diffraction measurements to provide quantitative temperature evolution data during laser melting. Temperature-distribution metrics are determined using Gaussian process regression supervised machine-learning surrogate models trained with a combination of mechanistic modeling (heat transfer and fluid flow) and X-ray diffraction simulation. The trained surrogate model uncertainties are found to range from 5 to 15% depending on the metric and current temperature. The surrogate models are then applied to experimental data to extract temperature metrics from an Inconel 625 nickel superalloy wall specimen during laser melting. The maximum temperatures of the solid phase in the diffraction volume through melting and cooling are found to reach the solidus temperature as expected, with the mean and minimum temperatures found to be several hundred degrees less. The extracted temperature metrics near melting are determined to be more accurate because of the lower relative levels of mechanical elastic strains. However, uncertainties for temperature metrics during cooling are increased due to the effects of thermomechanical stress.

36 MATERIALS SCIENCE↗

Sensor enabled data-driven predictive analytics for modeling and control with high penetration of DERs in distribution systems

The electric power grid is undergoing a tremendous transformation due to the increasing penetration of renewable energy resources beginning with wind and more recently with the distributed energy resources (DERs) such as solar and battery storage. DERs have dramatically changed the role of the distribution systems in the overall power grid, and they are expected to contribute a significant portion of power generation in the future. If current trends for DERs continue, system operation and control will need to change dramatically for improved grid reliability and resiliency. As renewable resources increase in penetration, new and challenging operational, planning, and design problems are expected to emerge. Some of the key challenges that arise in the planning and operation of the future grid are: 1) Quantifying the impact of high DER penetration in distribution systems on bulk grid behavior over multiple time scales. 2) Identifying whether a particular DER configuration/settings have a large impact on the overall grid behavior. These challenges can be addressed in an offline manner using detailed T&D grid models and they can also be addressed in an online manner using sensor measurements. In particular, the advancement and planned growth in sensor technology in power grid over various voltage levels provide us with a unique opportunity to tackle these challenges from a data analytic perspective without needing detailed T&D grid models. A few questions that naturally arise when addressing the challenges from DERs using sensor data are: 1) How can we use limited sensor measurements to monitor & control voltage stability and small signal stability of the bulk system? 2) How can we ensure that the developed data analytic methods are robust to data availability and quality issues? 3) How can we compute the developed analytics in a scalable manner using streaming measurements? In this project, we addressed the aforementioned challenges arising from DERs and answered the questions raised above on how to effectively use the sensor measurements to enhance the reliability and performance of the electric grid. Thus, the overarching goal of this project is to develop effective reduced/representative system models from data that make the computational complexity sufficiently manageable so as to be useful to simulate, analyze, and even control complex non-linear power systems dynamics with large penetrations of DERs. In order to achieve the objective, the project team established a four-fold technical approach 1) Formulated a combined transmission-distribution co-simulation framework for data generation and validation, 2) Derived reduced/representative models of power systems based on data-driven methods for efficient computation and appropriate representation of system behavior, 3) Developed data driven characterization of power system behavior based on transfer operator theory, machine learning and optimization for model estimation, 4) Incorporated a scalable data management and processing architecture using distributed Kafka streaming applications that coordinate input data streams to the developed data analytics. The key accomplishments of the project are: 1) Development of a scalable multi-timescale T&D co-simulation framework (both for steady state and for dynamic co-simulation) using commercial solvers (PSSE and GridLAB-D). The steady-state T&D co-simulation interface is shared with our industry partner (PJM). 2) A structured reduced order dynamic model of distribution systems that can represent partial motor stalling along with a systematic procedure to derive the model parameters. 3) A PMU based online method to monitor, localize and mitigate fault-induced delayed voltage recovery using DER reactive support and load control in distribution systems. 4) Development of linear operator based robust methodologies for dynamic state estimation, uncertainty quantification, system identification and trajectory prediction for power system dynamics. 5) An adaptive damping control for utilizing wind energy resources to provide oscillation damping and system stability. 6) Implementation of Kafka-based framework for efficient processing of streaming data using Linux-based local virtual environment.

DER integration↗

Rapid Adaptation of Chemical Named Entity Recognition Using Few-Shot Learning and LLM Distillation

Named entity recognition (NER) has been widely used in chemical text mining for the automatic identification and extraction of chemical entities. However, existing chemical NER systems primarily focus on scenarios with abundant training data, requiring significant human effort on annotations. This poses challenges for applications in the chemical field, such as catalysis, where many advancements have traditionally relied on trial-and-error investigations and incremental adjustment of variables. This hinders catalysis science and technology progress in addressing emerging energy and environmental crises. In this work, we propose a few-shot NER model that can quickly adapt to extract new types of chemical entities by using only a limited number of annotated examples. Our model employs a metric-learning approach to transfer entity similarity knowledge from high-resource chemical domains (with abundant annotations) to enable effective entity recognition in low-resource specialized domains (limited annotation). We validate the effectiveness of our model on a few-shot chemical NER benchmark built based on six existing chemical NER data sets. Experiments show that the proposed few-shot NER model can achieve reasonable performance with only 5 examples per entity type and shows consistent improvement as the number of examples increases. Furthermore, we demonstrate how the proposed model can be trained with large language model (LLM) annotated data, opening a new pathway for rapid adaptation of NER systems. Furthermore, our approach leverages the knowledge broadness of large language models for chemistry while distilling this knowledge into a lightweight model suitable for efficient and in-house use.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Electronic structure prediction of multi-million atom systems through uncertainty quantification enabled transfer learning

The ground state electron density — obtainable using Kohn-Sham Density Functional Theory (KS-DFT) simulations — contains a wealth of material information, making its prediction via machine learning (ML) models attractive. However, the computational expense of KS-DFT scales cubically with system size which tends to stymie training data generation, making it difficult to develop quantifiably accurate ML models that are applicable across many scales and system configurations. Here, we address this fundamental challenge by employing transfer learning to leverage the multi-scale nature of the training data, while comprehensively sampling system configurations using thermalization. Our ML models are less reliant on heuristics, and being based on Bayesian neural networks, enable uncertainty quantification. We show that our models incur significantly lower data generation costs while allowing confident — and when verifiable, accurate — predictions for a wide variety of bulk systems well beyond training, including systems with defects, different alloy compositions, and at multi-million-atom scales. Moreover, such predictions can be carried out using only modest computational resources.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗