Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “deep transfer learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Contrastive learning for robust representations of neutrino data

In neutrino physics, analyses often depend on large simulated datasets, making it essential for models to generalize effectively to real-world detector data. Contrastive learning, a well-established technique in deep learning, offers a promising solution to this challenge. By applying controlled data augmentations to simulated data, contrastive learning enables the extraction of robust and transferable features. This improves the ability of models trained on simulations to adapt to real experimental data distributions. In this paper, we investigate the application of contrastive learning methods in the context of neutrino physics. Through a combination of empirical evaluations and theoretical insights, we demonstrate how contrastive learning enhances model performance and adaptability. Additionally, we compare it to other domain adaptation techniques, highlighting the unique advantages of contrastive learning for this field. Published by the American Physical Society 2025

Wilkinson, Alex (ORCID:0000000253404506)↗

An Overview of the Usefulness of Machine Learning Techniques on Network Packet Data

Understanding the health and behavior of a computer network allows for better network efficiency and security. We present an overview of various machine learning techniques for classifying network packet data via packet metadata. While some classical machine learning approaches achieve reasonable results, the most accurate classification can be achieved with deep learning. On the four data sets studied herein, a basic deep learning model achieved at or near 100\% classification accuracy. We also propose a method for determining variable importance as a means for potential transfer learning applications to classifying yet unseen network packet data.

97 MATHEMATICS AND COMPUTING↗

Virtual to Physical: Reinforcement Learning to Optimize SNS Particle Accelerator Controls

Complex accelerators must have control systems that can handle dynamic nonlinear environments. This makes traditional control methods unsuitable as they can struggle to adapt to these uncertainties. This provides an ideal environment for reinforcement learning algorithms as they are adaptable and generalizable. We present a reinforcement learning pipeline that can effectively handle the dynamics of a complex accelerator. We test and prove our pipelines capabilities on multiple environments including the Spallation Neutron Source (SNS) and the Beam Test Facility (BTF) at Oakridge National Lab (ORNL). Due to the limited time available to train an online algorithm like reinforcement learning on a real accelerator, we utilize a virtual twin accelerator (VIRAC) developed by ORNL to pretrain the policy and show its ability to converge in the virtual environment. We then test the adaptability of the pretrained RL model by applying it on the real accelerator and comparing the results. Utilizing our Scientific Optimization and Controls Toolkit (SOCT) and open-source standards such as Gymnasium we create and solve for a MEBT orbit correction problem in the SNS and an emittance maximization problem in the BTF. We show how Twin Delayed Deep Deterministic Policy Gradient (TD3) can solve this optimization environment in the virtual accelerator and transfer this policy onto the real accelerator for inference and model retraining. We show how reinforcement learning can be utilized as a control system for complex accelerators and provide a model pipeline for how an implementation performs and can be adapted to new accelerator control problems.

Kasparian, Armen [Thomas Jefferson National Accele↗

Exact constraints and appropriate norms in machine-learned exchange-correlation functionals

Machine learning techniques have received growing attention as an alternative strategy for developing general-purpose density functional approximations, augmenting the historically successful approach of human-designed functionals derived to obey mathematical constraints known for the exact exchange-correlation functional. More recently, efforts have been made to reconcile the two techniques, integrating machine learning and exact-constraint satisfaction. We continue this integrated approach, designing a deep neural network that exploits the exact constraint and appropriate norm philosophy to de-orbitalize the strongly constrained and appropriately normed (SCAN) functional. The deep neural network is trained to replicate the SCAN functional from only electron density and local derivative information, avoiding the use of the orbital-dependent kinetic energy density. The performance and transferability of the machine-learned functional are demonstrated for molecular and periodic systems.

Artificial neural networks↗

Machine-learning force-field models for dynamical simulations of metallic magnets

We review recent advances in machine-learning (ML) force-field methods for Landau–Lifshitz–Gilbert simulations of itinerant electron magnets, focusing on their scalability and transferability. Built on the principle of locality, a deep neural-network model is developed to efficiently and accurately predict electron-mediated forces governing spin dynamics. Symmetry-aware descriptors constructed through a group-theoretical approach ensure rigorous incorporation of both lattice and spin-rotation symmetries. The framework is demonstrated using the prototypical s-d exchange model widely employed in spintronics. ML-enabled large-scale simulations reveal novel nonequilibrium phenomena, including anomalous coarsening of tetrahedral spin order on the triangular lattice and the freezing of phase-separation dynamics in lightly hole-doped, strong-coupling square-lattice systems. These results establish ML force-field frameworks as scalable, accurate, and versatile tools for modeling nonequilibrium spin dynamics in itinerant magnets.

Artificial neural networks↗

Deep Learning-Enabled MS/MS Spectrum Prediction Facilitates Automated Identification Of Novel Psychoactive Substances

The market for illicit drugs has been reshaped by the emergence of more than 1100 new psychoactive substances (NPS) over the past decade, posing a major challenge to the forensic and toxicological laboratories tasked with detecting and identifying them. Tandem mass spectrometry (MS/MS) is the primary method used to screen for NPS within seized materials or biological samples. The most contemporary workflows necessitate labor-intensive and expensive MS/MS reference standards, which may not be available for recently emerged NPS on the illicit market. Here, we present NPS-MS, a deep learning method capable of accurately predicting the MS/MS spectra of known and hypothesized NPS from their chemical structures alone. NPS-MS is trained by transfer learning from a generic MS/MS prediction model on a large data set of MS/MS spectra. We show that this approach enables a more accurate identification of NPS from experimentally acquired MS/MS spectra than any existing method. We demonstrate the application of NPS-MS to identify a novel derivative of phencyclidine (PCP) within an unknown powder seized in Denmark without the use of any reference standards. We anticipate that NPS-MS will allow forensic laboratories to identify more rapidly both known and newly emerging NPS. NPS-MS is available as a web server at https://nps-ms.ca/, which provides MS/MS spectra prediction capabilities for given NPS compounds. Additionally, it offers MS/MS spectra identification against a vast database comprising approximately 8.7 million predicted NPS compounds from DarkNPS and 24.5 million predicted ESI-QToF-MS/MS spectra for these compounds.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fracture Characterization Via AI‐Assisted Analysis of Temperature Logs

Abstract Fractures control fluid flow, mass transport, and heat transfer in a geothermal reservoir. This makes accurate characterization of fracture networks a prerequisite for optimal design and control of a reservoir's exploitation. We develop a deep‐learning procedure to identify fracture locations via interpretation of temporally and spatially continuous downhole temperature measurements. A long short‐term memory fully convolutional network (LSTM‐FCN) is used both to capture long‐term dependencies in sequential temperature data and to distill local features around fractures. A wellbore and fractured‐reservoir thermal model is established to generate temperature data for network training. The trained LSTM‐FCN exhibits a unique ability to detect multiple fractures intersecting a borehole. We use the LSTM‐FCN algorithm to evaluate the effectiveness of different‐stage wellbore temperature measurements on fracture detection in a complex fractured system. Our experiments reveal that the use of various‐stage temperature information as an input feature set improves the robustness of fracture detection to noise interference. This study indicates the practical feasibility of obtaining accurate fracture‐network reconstructions from temperature signals, at reasonable computational cost.

Yang, Xiaoyu↗

Convolutional Neural Networks for Problems in Transport Phenomena: A Theoretical Minimum

Convolutional neural network (CNN), a deep learning algorithm, has gained popularity in technological applications that rely on interpreting images (typically, an image is a 2D field of pixels). Transport phenomena is the science of studying different fields representing mass, momentum, or heat transfer. Some of the common fields are species concentration, fluid velocity, pressure, and temperature. Each of these fields can be expressed as an image(s). Consequently, CNNs can be leveraged to solve specific scientific problems in transport phenomena. Herein, we show that such problems can be grouped into three basic categories: (a) mapping a field to a descriptor (b) mapping a field to another field, and (c) mapping a descriptor to a field. After reviewing the representative transport phenomena literature for each of these categories, we illustrate the necessary steps for constructing appropriate CNN solutions using sessile liquid drops as an exemplar problem. If sufficient training data is available, CNNs can considerably speed up the solution of the corresponding problems. Finally, the present discussion is meant to be minimalistic such that readers can easily identify the transport phenomena problems where CNNs can be useful as well as construct and/or assess such solutions.

97 MATHEMATICS AND COMPUTING↗

Alpert multi-wavelets for functional inverse problems: direct optimization and deep learning

Computational engineering models often contain unknown entities (e.g. parameters, initial and boundary conditions) that require estimation from other measured observable data. Estimating such unknown entities is challenging when they involve spatio-temporal fields because such functional variables often require an infinite-dimensional representation. Here, we address this problem by transforming an unknown functional field using Alpert wavelet bases and truncating the resulting spectrum. Hence the problem reduces to the estimation of few coefficients that can be performed using common optimization methods. We apply this method on a one-dimensional heat transfer problem where we estimate the heat source field varying in both time and space. The observable data is comprised of temperature measured at several thermocouples in the domain. This latter is composed of either copper or stainless steel. The optimization using our method based on wavelets is able to estimate the heat source with an error between 5% and 7%. We analyze the effect of the domain material and number of thermocouples as well as the sensitivity to the initial guess of the heat source. Finally, we estimate the unknown heat source using a different approach based on deep learning techniques where we consider the input and output of a multi-layer perceptron in wavelet form. We find that this deep learning approach is more accurate than the optimization approach with errors below 4%.

97 MATHEMATICS AND COMPUTING↗

Federated learning for 2D synchrotron x-ray diffractometry: a cross-institutional approach for phase quantification of Ti–6Al–4V alloy

High-energy Two dimensional (2D) synchrotron x-ray diffractometry provides important insights into the atomistic structure and phase evolution of materials, yet traditional analysis methods remain complex, knowledge-intensive, and computationally demanding. Deep-learning models offer a powerful alternative for automating their analysis. Institutions that hold these datasets may be unwilling to share their data due to privacy and security policies, as well as the challenges associated with large-scale data transfer. As a result, models trained on local datasets often perform well only on their own data but exhibit bias and poor generalization across different instruments or facilities. To overcome these limitations, we explore federated learning (FL) for 2D synchrotron diffractograms, enabling collaborative model training without exchanging raw data. In this study, 2D synchrotron diffractograms of Ti–6Al–4V alloy collected from two independent facilities are used to train convolutional neural networks for predicting the β-phase volume fraction. Experimental results show that federated global models significantly outperform locally trained models in terms of generalization and achieve accuracy comparable to centralized trained models. These findings demonstrate the potential of FL to enable secure, cross-institutional collaboration and enhance the scalability of deep-learning-based materials characterization.

36 MATERIALS SCIENCE↗

Atomic resolution convergent beam electron diffraction analysis using convolutional neural networks

Two types of convolutional neural network (CNN) models, a discrete classification network and a continuous regression network, were trained to determine local sample thickness from convergent beam diffraction (CBED) patterns of SrTiO 3 collected in a scanning transmission electron microscope (STEM) at atomic column resolution. Acquisition of atomic resolution CBED patterns for this purpose requires careful balancing of CBED feature size in pixels, acquisition speed, and detector dynamic range. The training datasets were derived from multislice simulations, which must be convolved with incoherent source broadening. Sample thicknesses were also determined using quantitative high-angle annular dark-field (HAADF) STEM images acquired simultaneously. The regression CNN performed well on sample thinner than 35 nm, with 70% of the CNN results within 1 nm of HAADF thickness, and 1.0 nm overall root mean square error between the two measurements. The classification CNN was trained for a thicknesses up to 100 nm and yielded 66% of CNN results within one classification increment of 2 nm of HAADF thickness. Our approach depends on methods from computer vision including transfer learning and image augmentation.

36 MATERIALS SCIENCE↗

Subaru Hyper Suprime-Cam revisits the large-scale environmental dependence on galaxy morphology over 360 deg2 at z = 0.3–0.6

This study investigates the role of large-scale environments on the fraction of spiral galaxies at z = 0.3–0.6 sliced to three redshift bins of Δz = 0.1. Here, we sample 276220 massive galaxies in a limited stellar mass of 5 × 1010 solar mass (~M*) over 360 deg2, as obtained from the Second Public Data Release of the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP). By combining projected two-dimensional density information (Shimakawa et al. 2021, MNRAS, 503, 3896) and the CAMIRA cluster catalog (Oguri et al. 2018, PASJ, 70, S20), we investigate the spiral fraction across large-scale overdensities and in the vicinity of red sequence clusters. We adopt transfer learning to reduce the cost of labeling spiral galaxies significantly and then perform stacking analysis across the entire field to overcome the limitations of sample size. Here we employ a morphological classification catalog by the Galaxy Zoo Hubble (Willett et al., 2017, MNRAS, 464, 4176) to train the deep learning model. Based on 74103 sources classified as spirals, we find moderate morphology–density relations on a 10 comoving Mpc scale, thanks to the wide-field coverage of HSC-SSP. Clear deficits of spiral galaxies have also been confirmed, in and around 1136 red sequence clusters. Furthermore, we verify whether there is a large-scale environmental dependence on rest-frame u - r colors of spiral galaxies; such a tendency was not observed in our sample.

Astronomy & Astrophysics↗

Time-lapse seismic inversion for CO 2 saturation with SeisCO2Net: An application to Frio-II site

Seismic monitoring of geological CO 2 storage (GCS) involves highly nonlinear seismic inversion and petrophysical inversion, making it challenging to estimate CO 2 volume efficiently and detect possible early CO 2 leakages. Deep learning (DL) using convolutional neural networks (CNNs) has shown promise in solving highly nonlinear seismic inversion problems. However, direct estimation of CO 2 plume extent/saturation from time-lapse seismic gathers using DL is still underexplored, with no reported field applications to date. The investigation of field data is primarily hindered by scarcity of field data for neural network training. Other obstacles include highly nonlinear seismic-petrophysics inverse relationship, and presence of noise in field seismic data. We introduce SeisCO2Net, a deep CNN that predicts CO 2 saturation maps directly from time-lapse full waveform shot gathers. For training, we use site-specific geological information, fluid flow physics, rock physics, and seismic modeling to generate synthetic datasets that closely resemble the CO 2 storage site. Synthetic tests show promising results, inspiring us to apply SeisCO2Net's trained weights on field data collected at Frio-II GCS site by leveraging transfer learning principles. As reference, we compare SeisCO2Net's predicted CO 2 saturation maps with results obtained from physics-based inversion. Our analyses show both methods display similar CO 2 plume shapes, reasonable CO 2 plume characteristics, and comparable saturation values. Our results suggest pre-training CNNs on physics-informed synthetic datasets and then applying the learned weights to field data is a viable approach to estimating field CO 2 saturation. This method effectively addresses the scarcity of field training data, thus encouraging the feasibility of long-term GCS monitoring.

58 GEOSCIENCES↗

Potential of deep learning methods to enhance satellite-based monitoring of nuclear power plants focusing on remote operation evaluations

The anticipated expansion of the nuclear industry and the deployment of new nuclear reactors (200 + GW of new nuclear capacity by 2050) require the development of monitoring systems that align with safety and security concerns, providing enhanced evaluation capabilities. A remote monitoring system using satellites and deep learning techniques was evaluated for its ability to detect anomalies and capture various features of nuclear reactors independently of the conditions on the ground. Satellite images of current operational and under-construction nuclear power plants were collected from Google Earth Pro as a surrogate database. Subsequently, five datasets were created from the collected images. Transfer learning technique was used for several classification tasks utilizing VGG16, ResNet50V2, Xception, DenseNet121, and MobileNetV2 pre-trained models. In the first task, the capability of the monitoring system to detect abnormal conditions or processes in a nuclear power plant was investigated. In the second task, the ability to capture operational features remotely was examined. As an example, for the purposes of this study, these features included classifying reactors based on type, power range, or onsite condition. Several evaluation metrics were used to compare the performance of the pre-trained models and the overall monitoring system. Here, the evaluation results demonstrated that deep learning techniques and pre-trained models applied to satellite images have the potential to facilitate further and expand capabilities in monitoring systems to assess plant operation details.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Fed-DeepONet: Stochastic Gradient-Based Federated Training of Deep Operator Networks

The Deep Operator Network (DeepONet) framework is a different class of neural network architecture that one trains to learn nonlinear operators, i.e., mappings between infinite-dimensional spaces. Traditionally, DeepONets are trained using a centralized strategy that requires transferring the training data to a centralized location. Such a strategy, however, limits our ability to secure data privacy or use high-performance distributed/parallel computing platforms. To alleviate such limitations, in this paper, we study the federated training of DeepONets for the first time. That is, we develop a framework, which we refer to as Fed-DeepONet, that allows multiple clients to train DeepONets collaboratively under the coordination of a centralized server. To achieve Fed-DeepONets, we propose an efficient stochastic gradient-based algorithm that enables the distributed optimization of the DeepONet parameters by averaging first-order estimates of the DeepONet loss gradient. Then, to accelerate the training convergence of Fed-DeepONets, we propose a moment-enhanced (i.e., adaptive) stochastic gradient-based strategy. Finally, we verify the performance of Fed-DeepONet by learning, for different configurations of the number of clients and fractions of available clients, (i) the solution operator of a gravity pendulum and (ii) the dynamic response of a parametric library of pendulums.

Moya, Christian↗

Tracing and Forecasting Metabolic Indices of Cancer Patients Using Patient-Specific Deep Learning Models

We develop a patient-specific dynamical system model from the time series data of the cancer patient’s metabolic panel taken during the period of cancer treatment and recovery. The model consists of a pair of stacked long short-term memory (LSTM) recurrent neural networks and a fully connected neural network in each unit. It is intended to be used by physicians to trace back and look forward at the patient’s metabolic indices, to identify potential adverse events, and to make short-term predictions. When the model is used in making short-term predictions, the relative error in every index is less than 10% in the L ∞ norm and less than 6.3% in the L 1 norm in the validation process. Once a master model is built, the patient-specific model can be calibrated through transfer learning. As an example, we obtain patient-specific models for four more cancer patients through transfer learning, which all exhibit reduced training time and a comparable level of accuracy. This study demonstrates that this modeling approach is reliable and can deliver clinically acceptable physiological models for tracking and forecasting patients’ metabolic indices.

60 APPLIED LIFE SCIENCES↗