Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “deep transfer learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Potential of deep learning methods to enhance satellite-based monitoring of nuclear power plants focusing on remote operation evaluations

The anticipated expansion of the nuclear industry and the deployment of new nuclear reactors (200 + GW of new nuclear capacity by 2050) require the development of monitoring systems that align with safety and security concerns, providing enhanced evaluation capabilities. A remote monitoring system using satellites and deep learning techniques was evaluated for its ability to detect anomalies and capture various features of nuclear reactors independently of the conditions on the ground. Satellite images of current operational and under-construction nuclear power plants were collected from Google Earth Pro as a surrogate database. Subsequently, five datasets were created from the collected images. Transfer learning technique was used for several classification tasks utilizing VGG16, ResNet50V2, Xception, DenseNet121, and MobileNetV2 pre-trained models. In the first task, the capability of the monitoring system to detect abnormal conditions or processes in a nuclear power plant was investigated. In the second task, the ability to capture operational features remotely was examined. As an example, for the purposes of this study, these features included classifying reactors based on type, power range, or onsite condition. Several evaluation metrics were used to compare the performance of the pre-trained models and the overall monitoring system. Here, the evaluation results demonstrated that deep learning techniques and pre-trained models applied to satellite images have the potential to facilitate further and expand capabilities in monitoring systems to assess plant operation details.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Fed-DeepONet: Stochastic Gradient-Based Federated Training of Deep Operator Networks

The Deep Operator Network (DeepONet) framework is a different class of neural network architecture that one trains to learn nonlinear operators, i.e., mappings between infinite-dimensional spaces. Traditionally, DeepONets are trained using a centralized strategy that requires transferring the training data to a centralized location. Such a strategy, however, limits our ability to secure data privacy or use high-performance distributed/parallel computing platforms. To alleviate such limitations, in this paper, we study the federated training of DeepONets for the first time. That is, we develop a framework, which we refer to as Fed-DeepONet, that allows multiple clients to train DeepONets collaboratively under the coordination of a centralized server. To achieve Fed-DeepONets, we propose an efficient stochastic gradient-based algorithm that enables the distributed optimization of the DeepONet parameters by averaging first-order estimates of the DeepONet loss gradient. Then, to accelerate the training convergence of Fed-DeepONets, we propose a moment-enhanced (i.e., adaptive) stochastic gradient-based strategy. Finally, we verify the performance of Fed-DeepONet by learning, for different configurations of the number of clients and fractions of available clients, (i) the solution operator of a gravity pendulum and (ii) the dynamic response of a parametric library of pendulums.

Moya, Christian↗

Tracing and Forecasting Metabolic Indices of Cancer Patients Using Patient-Specific Deep Learning Models

We develop a patient-specific dynamical system model from the time series data of the cancer patient’s metabolic panel taken during the period of cancer treatment and recovery. The model consists of a pair of stacked long short-term memory (LSTM) recurrent neural networks and a fully connected neural network in each unit. It is intended to be used by physicians to trace back and look forward at the patient’s metabolic indices, to identify potential adverse events, and to make short-term predictions. When the model is used in making short-term predictions, the relative error in every index is less than 10% in the L ∞ norm and less than 6.3% in the L 1 norm in the validation process. Once a master model is built, the patient-specific model can be calibrated through transfer learning. As an example, we obtain patient-specific models for four more cancer patients through transfer learning, which all exhibit reduced training time and a comparable level of accuracy. This study demonstrates that this modeling approach is reliable and can deliver clinically acceptable physiological models for tracking and forecasting patients’ metabolic indices.

60 APPLIED LIFE SCIENCES↗

Deep Learning Approach for High-accuracy Electron Counting of Monolithic Active Pixel Sensor-type Direct Electron Detectors at Increased Electron Dose

Abstract Electron counting can be performed algorithmically for monolithic active pixel sensor direct electron detectors to eliminate readout noise and Landau noise arising from the variability in the amount of deposited energy for each electron. Errors in existing counting algorithms include mistakenly counting a multielectron strike as a single electron event, and inaccurately locating the incident position of the electron due to lateral spread of deposited energy and dark noise. Here, we report a supervised deep learning (DL) approach based on Faster region-based convolutional neural network (R-CNN) to recognize single electron events at varying electron doses and voltages. The DL approach shows high accuracy according to the near-ideal modulation transfer function (MTF) and detector quantum efficiency for sparse images. It predicts, on average, 0.47 pixel deviation from the incident positions for 200 kV electrons versus 0.59 pixel using the conventional counting method. The DL approach also shows better robustness against coincidence loss as the electron dose increases, maintaining the MTF at half Nyquist frequency above 0.83 as the electron density increases to 0.06 e−/pixel. Thus, the DL model extends the advantages of counting analysis to higher dose rates than conventional methods.

Materials Science↗

Deep modelling of plasma and neutral fluctuations from gas puff turbulence imaging

The role of turbulence in setting boundary plasma conditions is presently a key uncertainty in projecting to fusion energy reactors. To robustly diagnose edge turbulence, we develop and demonstrate a technique to translate brightness measurements of HeI line radiation into local plasma fluctuations via a novel integrated deep learning framework that combines neutral transport physics and collisional radiative theory for the $3^3 D - 2^3 P$ transition in atomic helium. The tenets for experimental validity are reviewed, illustrating that this turbulence analysis for ionized gases is transferable to both magnetized and unmagnetized environments with arbitrary geometries. Based upon fast camera data on the Alcator C-Mod tokamak, we present the first 2-dimensional time-dependent experimental measurements of the turbulent electron density, electron temperature, and neutral density revealing shadowing effects in a fusion plasma using a single spectral line.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Accelerating Resonance Searches via Signature-Oriented Pre-training

The search for heavy resonances beyond the Standard Model (BSM) is a key objective at the LHC. While the recent use of advanced deep neural networks for boosted-jet tagging significantly enhances the sensitivity of dedicated searches, it is limited to specific final states, leaving vast potential BSM phase space underexplored. We introduce a novel experimental method, Signature-Oriented Pre-training for Heavy-resonance ObservatioN (Sophon), which leverages deep learning to cover an extensive number of boosted final states. Pre-trained on the comprehensive JetClass-II dataset, the Sophon model learns intricate jet signatures, ensuring the optimal constructions of various jet tagging discriminates and enabling high-performance transfer learning capabilities. We show that the method can not only push widespread model-specific searches to their sensitivity frontier, but also greatly improve model-agnostic approaches, accelerating LHC resonance searches in a broad sense.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Scaling Ensembles of Data-Intensive Quantum Chemical Calculations for Millions of Molecules

Deep learning models are efficient computational tools that can accelerate the inverse design of molecules with desired functional properties by generating predictions at a fraction of the time required by traditional quantum chemical approaches. To ensure that a model maintains accuracy and transferability across broad regions of the chemical space explored during the inverse design, it must be trained on massively large volumes of simulation data. This requires running large-scale ensemble quantum chemical calculations on high-performance computing (HPC) systems for data collection. However, the efficient execution of such large ensemble calculations and the management of large volumes of output data require tools that can judiciously utilize computational resources and manage metadata overhead on the file system. Therefore, we present a high-performance, scalable, ensemble management framework for performing data-intensive quantum chemical electronic structure calculations for organic molecules. This framework provides abstractions to plug different ab initio, first principles, and first principles-based semi-empirical methods and executes them efficiently at large scale on HPC systems. It dynamically distributes tasks to resources and uses tiered storage for managing large collections of files. We employed this framework to process over ten million organic molecules and generate open-source datasets that provide UV-vis absorption spectra by running time-dependent density-functional tight-binding calculations. It is the largest database containing molecular optical spectra that were simulated with quantum chemical methods in a consistent manner.

Mehta, Kshitij↗

B-DeepONet: An enhanced Bayesian DeepONet for solving noisy parametric PDEs using accelerated replica exchange SGLD

Here, the Deep Operator Network (DeepONet) is a neural network architecture used to approximate operators, including the solution operator of parametric PDEs. DeepONets have shown remarkable approximation ability. However, the performance of DeepONets deteriorates when the training data is polluted with noise, a scenario that occurs in practice. To handle noisy data, we propose a Bayesian DeepONet based on replica exchange Langevin diffusion (reLD). Replica exchange uses two particles. The first particle trains a DeepONet to exploit the loss landscape and make predictions. The other particle trains a different DeepONet to explore the loss landscape and escape local minima via swapping. Compared to DeepONets trained with state-of-the-art gradient-based algorithms (e.g., Adam), the proposed Bayesian DeepONet greatly improves the training convergence for noisy scenarios and accurately estimates the uncertainty. To further reduce the high computational cost of the reLD training of DeepONets, we propose (1) an accelerated training framework that exploits the DeepONet's architecture to reduce its computational cost up to 25% without compromising performance and (2) a transfer learning strategy that accelerates training DeepONets for PDEs with different parameter values. Finally, we illustrate the effectiveness of the proposed Bayesian DeepONet using four parametric PDE problems.

97 MATHEMATICS AND COMPUTING↗

Enhanced physics-constrained deep neural networks for modeling vanadium redox flow battery

Numerical simulation has become indispensable in advancing cost-effective process optimization and control of flow batteries. We propose an enhanced version of the physics-constrained deep neural network (PCDNN) approach to provide high-accuracy voltage predictions in the vanadium redox flow batteries (VRFBs). The purpose of the PCDNN approach is to enforce the physics-based zero-dimensional (0D) VRFB model in a neural network to assure model generalization for various battery operation conditions. However, limited by the simplifications of the 0D model, the PCDNN cannot capture sharp voltage changes in the extreme SOC regions. To improve the accuracy of voltage prediction at extreme ranges, we introduce a second (enhanced) DNN to mitigate the prediction errors carried from the 0D model itself and call the resulting approach enhanced PCDNN (ePCDNN). By comparing with experimental data, we demonstrate that the ePCDNN approach can accurately capture the voltage response throughout the charge–discharge cycle, including the tail region of the voltage discharge curve. The loss function for training the ePCDNN is designed to be flexible by adjusting the weights of the physics-constrained DNN and the enhanced DNN. In conclusion, this allows the ePCDNN framework to be transferable to battery systems with variable physical model fidelity.

25 ENERGY STORAGE↗

Analysis-Specific Fast Simulation at the LHC with Deep Learning

Abstract We present a fast-simulation application based on a deep neural network, designed to create large analysis-specific datasets. Taking as an example the generation of W + jet events produced in $$\sqrt{s}=$$ s = 13 TeV proton–proton collisions, we train a neural network to model detector resolution effects as a transfer function acting on an analysis-specific set of relevant features, computed at generation level, i.e., in absence of detector effects. Based on this model, we propose a novel fast-simulation workflow that starts from a large amount of generator-level events to deliver large analysis-specific samples. The adoption of this approach would result in about an order-of-magnitude reduction in computing and storage requirements for the collision simulation workflow. This strategy could help the high energy physics community to face the computing challenges of the future High-Luminosity LHC.

Chen, C.↗

A leaf-level spectral library to support high-throughput plant phenotyping: predictive accuracy and model transfer

Abstract Leaf-level hyperspectral reflectance has become an effective tool for high-throughput phenotyping of plant leaf traits due to its rapid, low-cost, multi-sensing, and non-destructive nature. However, collecting samples for model calibration can still be expensive, and models show poor transferability among different datasets. This study had three specific objectives: first, to assemble a large library of leaf hyperspectral data (n=2460) from maize and sorghum; second, to evaluate two machine-learning approaches to estimate nine leaf properties (chlorophyll, thickness, water content, nitrogen, phosphorus, potassium, calcium, magnesium, and sulfur); and third, to investigate the usefulness of this spectral library for predicting external datasets (n=445) including soybean and camelina using extra-weighted spiking. Internal cross-validation showed satisfactory performance of the spectral library to estimate all nine traits (mean R2=0.688), with partial least-squares regression outperforming deep neural network models. Models calibrated solely using the spectral library showed degraded performance on external datasets (mean R2=0.159 for camelina, 0.337 for soybean). Models improved significantly when a small portion of external samples (n=20) was added to the library via extra-weighted spiking (mean R2=0.574 for camelina, 0.536 for soybean). The leaf-level spectral library greatly benefits plant physiological and biochemical phenotyping, whilst extra-weight spiking improves model transferability and extends its utility.

59 BASIC BIOLOGICAL SCIENCES↗

Deep learning of dynamically responsive chemical Hamiltonians with semiempirical quantum mechanics

Conventional machine-learning (ML) models in computational chemistry learn to directly predict molecular properties using quantum chemistry only for reference data. While these heuristic ML methods show quantum-level accuracy with speeds several orders of magnitude faster than traditional quantum chemistry methods, they suffer from poor extensibility and transferability; i.e., their accuracy degrades on large or new chemical systems. Incorporating quantum chemistry frameworks into the ML models directly solves this problem. Here we take the structure of semiempirical quantum mechanics (SEQM) methods to construct dynamically responsive Hamiltonians. SEQM methods use empirical parameters fitted to experimental properties to construct reduced-order Hamiltonians, facilitating much faster calculations than ab initio methods but with compromised accuracy. By replacing these static parameters with machine-learned dynamic values inferred from the local environment, we greatly improve the accuracy of the SEQM methods. Trained on molecular energies and atomic forces, these dynamically generated Hamiltonian parameters show a strong correlation with atomic hybridization and bonding. Trained with only about 60,000 small organic molecular conformers, the resulting model retains interpretability, extensibility, and transferability when testing on much larger chemical systems and predicting various molecular properties. Overall, this work demonstrates the virtues of incorporating physics-based descriptions with ML to develop models that are simultaneously accurate, transferable, and interpretable.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

EGS Collab Experiment 1: 3D Seismic Velocity Model and Updated Microseismic Catalog Using Transfer-Learning Aided Double-Difference Tomography

This package contains a 3D Seismic velocity model and an updated microseismic catalog associated with a proceedings paper (Chai et al., 2020) published in the 45th Workshop on Geothermal Reservoir Engineering. The 3D_seismic_velocity_model text file contains x (m), y(m), z(m), P-wave velocity (km/s), P-wave velocity quality indicator (1 for well-constrained; 0 for poorly constrained), S-wave velocity (km/s), and S-wave velocity quality indicator (1 for well-constrained; 0 for poorly constrained). The Updated_MEQ_catalog text file contains event origin time, x(m), y(m), z(m), error in x (m), error in y (m), error in z (m), and RMS misfit (millisecond). The 3D_seismic_P-wave_velocity_model animation file shows slices of the 3D P-wave velocity model. The 3D_seismic_S-wave_velocity_model animation file shows slices of the 3D S-wave velocity model. The Interactive_MEQ_locations API file is an interactive visualization of the updated microseismic event locations. The visualization allows users to view the event locations by dragging, rotating, and zooming in. References: Chai, C., Maceira, M., Santos-Villalobos, H. J., Venkatakrishnan, S. V., Schoenball, M., and EGS Collab Team, 2020, Automatic Seismic Phase Picking Using Deep Learning for the EGS Collab Project, in PROCEEDINGS, 45th Workshop on Geothermal Reservoir Engineering, edited, Stanford University, Stanford, California, 45, 1266-1276.

15 GEOTHERMAL ENERGY↗

Estimating cluster masses from SDSS multiband images with transfer learning

ABSTRACT The total masses of galaxy clusters characterize many aspects of astrophysics and the underlying cosmology. It is crucial to obtain reliable and accurate mass estimates for numerous galaxy clusters over a wide range of redshifts and mass scales. We present a transfer-learning approach to estimate cluster masses using the ugriz-band images in the SDSS Data Release 12. The target masses are derived from X-ray or SZ measurements that are only available for a small subset of the clusters. We designed a semisupervised deep learning model consisting of two convolutional neural networks. In the first network, a feature extractor is trained to classify the SDSS photometric bands. The second network takes the previously trained features as inputs to estimate their total masses. The training and testing processes in this work depend purely on real observational data. Our algorithm reaches a mean absolute error (MAE) of 0.232 dex on average and 0.214 dex for the best fold. The performance is comparable to that given by redMaPPer, 0.192 dex. We have further applied a joint integrated gradient and class activation mapping method to interpret such a two-step neural network. The performance of our algorithm is likely to improve as the size of training data set increases. This proof-of-concept experiment demonstrates the potential of deep learning in maximizing the scientific return of the current and future large cluster surveys.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Sparse-Data Deep Learning Strategies for Radiographic Non-Destructive Testing

Radiography is an imaging technique used in a variety of applications, such as medical diagnosis, airport security, and nondestructive testing. We present a deep learning system for extracting information from radiographic images. We perform various prediction tasks using our system, including material classification and regression on the dimensions of a given object that is being radiographed. Our system is designed to address the sparse-data issue for radiographic nondestructive testing applications. It uses a radiographic simulation tool for synthetic data augmentation, and it uses transfer learning with a pre-trained convolutional neural network model. Using this system, our preliminary results indicate that the object geometry regression task saw an improvement of 70% in the R-squared value when using a multi-regime model. In addition, we increase the performance of the object material classification tasks by utilizing data from different imaging systems. In particular, using neutron imaging improved the material classification accuracy by 20% when compared to x-ray imaging.

convolutional neural networks↗

Optimization of the FRIB beam dump: a hybrid genetic algorithm and reinforcement learning approach

The operational envelope of high-power-density systems, such as particle accelerators and advanced nuclear energy systems, is critically constrained by the need to manage extreme thermal loads. To address this, we present a novel hybrid optimization framework combining a genetic algorithm (GA) with a soft actor-critic (SAC) deep reinforcement learning agent. This framework was applied to a practical high-heat-flux problem: redesigning the beam dump at the Facility for Rare Isotope Beams (FRIB) for a power upgrade from 20 kW to 50 kW. The resulting design, validated by three-dimensional conjugate heat transfer simulations, suppresses hazardous hot spots and yields a markedly more uniform temperature distribution. This provides a robust operating margin, increasing the average power-handling capability by 72% relative to the current design, demonstrating the framework’s potential to solve complex thermal management challenges in both accelerator technology and advanced nuclear systems.

Accelerator↗

Physics-based hybrid machine learning for critical heat flux prediction with uncertainty quantification

Critical heat flux (CHF) is a key quantity in nuclear system modeling due to its impact on heat transfer, safety margins, and reactor performance. This study develops and validates an uncertainty-aware hybrid modeling approach that combines machine learning with physics-based models to predict CHF in cases of dryout. The Biasi and Bowring empirical correlations were paired with three ML uncertainty quantification (UQ) techniques: deep neural network (DNN) ensembles, Bayesian neural networks (BNNs), and deep Gaussian processes (DGPs). A pure ML model without a base model was evaluated for comparison. Model performance was assessed under plentiful (7,350 points) and limited (9 points) training data scenarios using parity, uncertainty distributions, and calibration curves. Results show that the Biasi hybrid DNN ensemble achieved the best overall performance, with a mean absolute relative error of 1.846%, and well-calibrated uncertainty estimates. The BNN-based hybrids showed slightly higher error (2.14%) but superior uncertainty calibration. DGP models underperformed, with over 6% error and poor uncertainty calibration. All hybrid models outperformed pure machine learning configurations, demonstrating resistance against data scarcity. These findings indicate that hybrid modeling significantly improves predictive accuracy, interpretability, and resilience to data scarcity. The integration of uncertainty awareness provides actionable confidence in CHF predictions, which is vital for safety-critical decisions in nuclear applications. This hybrid approach offers a viable pathway for deploying ML models in reactor analysis tools while preserving domain knowledge and physical consistency.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Diffractive optical computing in free space

Abstract Structured optical materials create new computing paradigms using photons, with transformative impact on various fields, including machine learning, computer vision, imaging, telecommunications, and sensing. This Perspective sheds light on the potential of free-space optical systems based on engineered surfaces for advancing optical computing. Manipulating light in unprecedented ways, emerging structured surfaces enable all-optical implementation of various mathematical functions and machine learning tasks. Diffractive networks, in particular, bring deep-learning principles into the design and operation of free-space optical systems to create new functionalities. Metasurfaces consisting of deeply subwavelength units are achieving exotic optical responses that provide independent control over different properties of light and can bring major advances in computational throughput and data-transfer bandwidth of free-space optical processors. Unlike integrated photonics-based optoelectronic systems that demand preprocessed inputs, free-space optical processors have direct access to all the optical degrees of freedom that carry information about an input scene/object without needing digital recovery or preprocessing of information. To realize the full potential of free-space optical computing architectures, diffractive surfaces and metasurfaces need to advance symbiotically and co-evolve in their designs, 3D fabrication/integration, cascadability, and computing accuracy to serve the needs of next-generation machine vision, computational imaging, mathematical computing, and telecommunication technologies.

36 MATERIALS SCIENCE↗