Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Transfer Learning Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Contrastive Machine Learning with Gamma Spectroscopy Data Augmentations for Detecting Shielded Radiological Material Transfers

Data analysis techniques can be powerful tools for rapidly analyzing data and extracting information that can be used in a latent space for categorizing observations between classes of data. Machine learning models that exploit learned data relationships can address a variety of nuclear nonproliferation challenges like the detection and tracking of shielded radiological material transfers. The high resource cost of manually labeling radiation spectra is a hindrance to the rapid analysis of data collected from persistent monitoring and to the adoption of supervised machine learning methods that require large volumes of curated training data. Instead, contrastive self-supervised learning on unlabeled spectra can enhance models that are built on limited labeled radiation datasets. This work demonstrates that contrastive machine learning is an effective technique for leveraging unlabeled data in detecting and characterizing nuclear material transfers demonstrated on radiation measurements collected at an Oak Ridge National Laboratory testbed, where sodium iodide detectors measure gamma radiation emitted by material transfers between the High Flux Isotope Reactor and the Radiochemical Engineering Development Center. Label-invariant data augmentations tailored for gamma radiation detection physics are used on unlabeled spectra to contrastively train an encoder, learning a complex, embedded state space with self-supervision. A linear classifier is then trained on a limited set of labeled data to distinguish transfer spectra between byproducts and tracked nuclear material using representations from the contrastively trained encoder. The optimized hyperparameter model achieves a balanced accuracy score of 80.30%. Any given model—that is, a trained encoder and classifier—shows preferential treatment for specific subclasses of transfer types. Regardless of the classifier complexity, a supervised classifier using contrastively trained representations achieves higher accuracy than using spectra when trained and tested on limited labeled data.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

FINETUNA: fine-tuning accelerated molecular simulations

Abstract Progress towards the energy breakthroughs needed to combat climate change can be significantly accelerated through the efficient simulation of atomistic systems. However, simulation techniques based on first principles, such as density functional theory (DFT), are limited in their practical use due to their high computational expense. Machine learning approaches have the potential to approximate DFT in a computationally efficient manner, which could dramatically increase the impact of computational simulations on real-world problems. However, they are limited by their accuracy and the cost of generating labeled data. Here, we present an online active learning framework for accelerating the simulation of atomic systems efficiently and accurately by incorporating prior physical information learned by large-scale pre-trained graph neural network models from the Open Catalyst Project. Accelerating these simulations enables useful data to be generated more cheaply, allowing better models to be trained and more atomistic systems to be screened. We also present a method of comparing local optimization techniques on the basis of both their speed and accuracy. Experiments on 30 benchmark adsorbate-catalyst systems show that our method of transfer learning to incorporate prior information from pre-trained models accelerates simulations by reducing the number of DFT calculations by 91%, while meeting an accuracy threshold of 0.02 eV 93% of the time. Finally, we demonstrate a technique for leveraging the interactive functionality built in to Vienna ab initio Simulation Package (VASP) to efficiently compute single point calculations within our online active learning framework without the significant startup costs. This allows VASP to work in tandem with our framework while requiring 75% fewer self-consistent cycles than conventional single point calculations. The online active learning implementation, and examples using the VASP interactive code, are available in the open source FINETUNA package on Github.

97 MATHEMATICS AND COMPUTING↗

Efficient emulation of relativistic heavy ion collisions with transfer learning

Measurements from the Large Hadron Collider (LHC) and the Relativistic Heavy Ion Collider (RHIC) can be used to study the properties of quark-gluon plasma. Systematic constraints on these properties must combine measurements from different collision systems and methodically account for experimental and theoretical uncertainties. Such studies require a vast number of costly numerical simulations. While computationally inexpensive surrogate models (“emulators”) can be used to efficiently approximate the predictions of heavy ion simulations across a broad range of model parameters, training a reliable emulator remains a computationally expensive task. We use transfer learning to map the parameter dependencies of one model emulator onto another, leveraging similarities between different simulations of heavy ion collisions. By limiting the need for large numbers of simulations to only one of the emulators, this technique reduces the numerical cost of comprehensive uncertainty quantification when studying multiple collision systems and exploring different models.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Trust-Enhancing Probabilistic Transfer Learning for Sparse and Noisy Data Environments

There is an increasing aspiration to utilize machine learning (ML) for various tasks of relevance to national security. ML models have thus far been mostly applied to tasks and domains that, while impactful, have sufficient volume of data. For predictive tasks of national security relevance, ML models of great capacity (ability to approximate nonlinear trends in input-output maps) are often needed to capture the complex underlying physics. However, scientific problems of relevance to national security are often accompanied by various sources of sparse and/or incomplete data, including experiments and simulations, across different regimes of operation, of varying degrees of fidelity, and include noise with different characteristics and/or intensity. State-of-the-art ML models, despite exhibiting superior performance on the task and domain they were trained on, may suffer detrimental loss in performance in such sparse data environments. This report summarizes the results of the Laboratory Directed Research and Development project entitled Trust-Enhancing Probabilistic Transfer Learning for Sparse and Noisy Data Environments. The objective of the project was to develop a new transfer learning (TL) framework that aims to adaptively blend the data across different sources in tackling one task of interest, resulting in enhanced trustworthiness of ML models for mission- and safety-critical systems. The proposed framework determines when it is worth applying TL and how much knowledge is to be transferred, despite uncontrollable uncertainties. The framework accomplishes this by leveraging concepts and techniques from the fields of Bayesian inverse modeling and uncertainty quantification, relying on strong mathematical foundations of probability and measure theories to devise new uncertainty-aware TL workflows.

97 MATHEMATICS AND COMPUTING↗

Ground State Energy Functional with Hartree–Fock Efficiency and Chemical Accuracy

We introduce the deep post Hartree–Fock (DeePHF) method, a machine learning-based scheme for constructing accurate and transferable models for the ground-state energy of electronic structure problems. DeePHF predicts the energy difference between results of highly accurate models such as the coupled cluster method and low accuracy models such as the Hartree–Fock (HF) method, using the ground-state electronic orbitals as the input. It preserves all the symmetries of the original high accuracy model. The added computational cost is less than that of the reference HF or DFT and scales linearly with respect to system size. We examine the performance of DeePHF on organic molecular systems using publicly available data sets and obtain the state-of-art performance, particularly on large data sets.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fine-tuning machine-learned particle-flow reconstruction for new detector geometries in future colliders

We demonstrate transfer learning capabilities in a machine-learned algorithm trained for particle-flow reconstruction in high energy particle colliders. This paper presents a cross-detector fine-tuning study, where we initially pretrain the model on a large full simulation dataset from one detector design, and subsequently fine-tune the model on a sample with a different collider and detector design. Specifically, we use the Compact Linear Collider detector (CLICdet) model for the initial training set and demonstrate successful knowledge transfer to the CLIC-like detector (CLD) proposed for the Future Circular Collider in electron-positron mode. We show that with an order of magnitude less samples from the second dataset, we can achieve the same performance as a costly training from scratch, across particle-level and event-level performance metrics, including jet and missing transverse momentum resolution. Furthermore, we find that the fine-tuned model achieves comparable performance to the traditional rule-based particle-flow approach on event-level metrics after training on 100,000 CLD events, whereas a model trained from scratch requires at least 1 million CLD events to achieve similar reconstruction performance. To our knowledge, this represents the first full-simulation cross-detector transfer learning study for particle-flow reconstruction. These findings offer valuable insights towards building large foundation models that can be fine-tuned across different detector designs and geometries, helping to accelerate the development cycle for new detectors and opening the door to rapid detector design and optimization using machine learning.

43 PARTICLE ACCELERATORS↗

Transfer learning for smart buildings: A critical review of algorithms, applications, and future perspectives

Smart buildings play a crucial role toward decarbonizing society, as globally buildings emit about one-third of greenhouse gases. In the last few years, machine learning has achieved a notable momentum that, if properly harnessed, may unleash its potential for advanced analytics and control of smart buildings, enabling the technique to scale up for supporting the decarbonization of the building sector. In this perspective, transfer learning aims to improve the performance of a target learner exploiting knowledge in related environments. The present work provides a comprehensive overview of transfer learning applications in smart buildings, classifying and analyzing 77 papers according to their applications, algorithms, and adopted metrics. The study identified four main application areas of transfer learning: (1) building load prediction, (2) occupancy detection and activity recognition, (3) building dynamics modeling, and (4) energy systems control. Furthermore, the review highlighted the role of deep learning in transfer learning applications that has been used in more than half of the analyzed studies. The paper also discusses how to integrate transfer learning in a smart building's ecosystem, identifying, for each application area, the research gaps and guidelines for future research directions.

Pinto, G↗

Accurate Prediction of Voltage of Battery Electrode Materials Using Attention-Based Graph Neural Networks

Performing first-principles calculations to discover electrodes’ properties in the large chemical space is a challenging task. While machine learning (ML) has been applied to effectively accelerate those discoveries, most of the applied methods ignore the materials’ spatial information and only use predefined features: based only on chemical compositions. Here, we propose two attention-based graph convolutional neural network techniques to learn the average voltage of electrodes. Our proposed methods, which combine both atomic composition and atomic coordinates in 3D-space, improve the accuracy in voltage prediction significantly when compared to composition-based ML models. The first model directly learns the chemical reaction of electrodes and metal ions to predict their average voltage, whereas the second model combines electrodes’ ML predicted formation energy (E form ) to compute their average voltage. Our E form -based model demonstrates improved accuracy in transferability from our subset of learned Li ions to Na ions. Moreover, we predicted the theoretical voltage of 10 Na x MPO 4 F (M = Ti, Cr, Fe, Cu, Mn, Co, and Ni) fluorophosphate battery frameworks, which are unavailable in the Material Project database. It could be shown that we can expect average voltages higher than 3.1 V from those Na battery frameworks except from the NaTiPO 4 F and TiPO 4 F pair of electrodes, which offer an average voltage of 1.32 V.

25 ENERGY STORAGE↗

Regional-scale soil carbon predictions can be enhanced by transferring global-scale soil–environment relationships

Accurate modelling and mapping soil organic carbon are crucial for supporting soil health restoration and climate change mitigation at both regional and global scales. However, regional soil predictions often suffer from data scarcity and high prediction uncertainty. Utilizing a pre-trained global-to-regional soil carbon predictive model can be a potential solution to address this challenge. Despite its promise, how to construct and apply the global-scale model to enhance regional-scale soil carbon mapping remains largely unexplored. Here, we propose the Global Soil Carbon Pre-trained Model (GSoilCPM), a deep-learning-based domain adaptative model, to enhance regional-scale soil carbon predictions. Based on large amount of environmental covariate data and 106,167 soil samples across the globe, we verify our hypothesis of the effectiveness of this 'global-to-regional' modelling strategy. The pre-trained model can be then transferred and fine-tuned to bridge the regional- and global-scale soil–environment relationships. We applied and validated this modelling strategy in four regional-scale study areas, three in the Northern Hemisphere and one in the Southern Hemisphere, each with distinct environmental background. Compared to traditional modelling approaches as a baseline, four case studies all demonstrated significant improvement in prediction accuracy across diverse environments and varying data availabilities. The average percentage improvement across all regions is 10.93% (absolute values decreased by 1.20 g kg−1 averagely) in MAE and 29.04% (absolute values increased by 0.10 averagely) in CCC. The applicability and future horizons of using GSoilCPM were further discussed. We further reveal that regions with fewer soil samples or lower baseline accuracy benefit more from the pre-trained global model. Our findings highlight the advantages of leveraging the generalized knowledge from global models to enhance specifically localized soil modelling, positioning a potential paradigm shift in digital soil mapping, and far-reaching implications for soil monitoring and land management.

Deep learning↗

Crustal permeability generated through microearthquakes is constrained by seismic moment

We link changes in crustal permeability to informative features of microearthquakes (MEQs) using two field hydraulic stimulation experiments where both MEQs and permeability evolution are recorded simultaneously. The Bidirectional Long Short-Term Memory (Bi-LSTM) model effectively predicts permeability evolution and ultimate permeability increase. Our findings confirm the form of key features linking the MEQs to permeability, offering mechanistically consistent interpretations of this association. Transfer learning correctly predicts permeability evolution of one experiment from a model trained on an alternate dataset and locale, which further reinforces the innate interdependency of permeability-to-seismicity. Models representing permeability evolution on reactivated fractures in both shear and tension suggest scaling relationships in which changes in permeability (Δk) are linearly related to the seismic moment (M) of individual MEQs as Δk ∝ M. This scaling relation rationalizes our observation of the permeability-to-seismicity linkage, contributes to its predictive robustness and accentuates its potential in characterizing crustal permeability evolution using MEQs.

15 GEOTHERMAL ENERGY↗

A unified understanding of minimum lattice thermal conductivity

Here, we propose a first-principles model of minimum lattice thermal conductivity ($κ^{min}_L$) based on a unified theoretical treatment of thermal transport in crystals and glasses. We apply this model to thousands of inorganic compounds and find a universal behavior of $κ^{min}_L$ in crystals in the high-temperature limit: The isotropically averaged $κ^{min}_L$ is independent of structural complexity and bounded within a range from ~0.1 to ~2.6 W/(m K), in striking contrast to the conventional phonon gas model which predicts no lower bound. We unveil the underlying physics by showing that for a given parent compound, $κ^{min}_L$ is bounded from below by a value that is approximately insensitive to disorder, but the relative importance of different heat transport channels (phonon gas versus diffuson) depends strongly on the degree of disorder. Moreover, we propose that the diffuson-dominated $κ^{min}_L$ in complex and disordered compounds might be effectively approximated by the phonon gas model for an ordered compound by averaging out disorder and applying phonon unfolding. With these insights, we further bridge the knowledge gap between our model and the well-known Cahill–Watson–Pohl (CWP) model, rationalizing the successes and limitations of the CWP model in the absence of heat transfer mediated by diffusons. Finally, we construct graph network and random forest machine learning models to extend our predictions to all compounds within the Inorganic Crystal Structure Database (ICSD), which were validated against thermoelectric materials possessing experimentally measured ultralow κ L . Our work offers a unified understanding of $κ^{min}_L$, which can guide the rational engineering of materials to achieve .

42 ENGINEERING↗

Accelerating Resonance Searches via Signature-Oriented Pre-training

The search for heavy resonances beyond the Standard Model (BSM) is a key objective at the LHC. While the recent use of advanced deep neural networks for boosted-jet tagging significantly enhances the sensitivity of dedicated searches, it is limited to specific final states, leaving vast potential BSM phase space underexplored. We introduce a novel experimental method, Signature-Oriented Pre-training for Heavy-resonance ObservatioN (Sophon), which leverages deep learning to cover an extensive number of boosted final states. Pre-trained on the comprehensive JetClass-II dataset, the Sophon model learns intricate jet signatures, ensuring the optimal constructions of various jet tagging discriminates and enabling high-performance transfer learning capabilities. We show that the method can not only push widespread model-specific searches to their sensitivity frontier, but also greatly improve model-agnostic approaches, accelerating LHC resonance searches in a broad sense.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Data Quality Monitoring for the Hadron Calorimeters Using Transfer Learning for Anomaly Detection

The proliferation of sensors brings an immense volume of spatio-temporal (ST) data in many domains, including monitoring, diagnostics, and prognostics applications. Data curation is a time-consuming process for a large volume of data, making it challenging and expensive to deploy data analytics platforms in new environments. Transfer learning (TL) mechanisms promise to mitigate data sparsity and model complexity by utilizing pre-trained models for a new task. Despite the triumph of TL in fields like computer vision and natural language processing, efforts on complex ST models for anomaly detection (AD) applications are limited. In this study, we present the potential of TL within the context of high-dimensional ST AD with a hybrid autoencoder architecture, incorporating convolutional, graph, and recurrent neural networks. Motivated by the need for improved model accuracy and robustness, particularly in scenarios with limited training data on systems with thousands of sensors, this research investigates the transferability of models trained on different sections of the Hadron Calorimeter of the Compact Muon Solenoid experiment at CERN. The key contributions of the study include exploring TL’s potential and limitations within the context of encoder and decoder networks, revealing insights into model initialization and training configurations that enhance performance while substantially reducing trainable parameters and mitigating data contamination effects.

47 OTHER INSTRUMENTATION↗

Machine-Learning Accelerated Studies of Materials with High Performance and Edge Computing

In the studies of materials, experimental measurements often serve as the reference to verify physics theory and modeling; while theory and modeling provide a fundamental understanding of the physics and principles behind. However, the interactions and cross validation between them have long been a challenge even to-date. Not only that inferring a physics model from experimental data is itself a difficult inverse problem, another major challenge is the orders-of-magnitude longer wall-clock time required to carry out high-fidelity computer modeling to match the timescale of experiments. We envisage that by combining high performance computing, data science, and edge computing technology, the current predicament can be alleviated, and a new paradigm of data-driven physics research will open up. For example, we can accelerate computer simulations by first performing the large-scale modeling on high performance computers and train a machine-learned surrogate model. This computationally inexpensive surrogate model can then be transferred to the computing units residing closely to the experimental facilities to perform high-fidelity simulations at a much higher throughout. The model will also be more amenable to analyzing and validating experimental observations in comparable time scales at a much lower computational cost. Further integration of these accelerated computer simulations with an outer machine learning loop can also inform and direct future experiments, while making the inverse problem of physics model inference more tractable. We will demonstrate a proof-of-concept by using a quantum Monte Carlo application, Dynamical Cluster Approximation (DCA++), to machine-learn a surrogate model and accelerate the study of quantum correlated materials.

Li, Ying Wai↗

Transformers and Long Short-Term Memory Transfer Learning for GenIV Reactor Temperature Time Series Forecasting

Automated monitoring of the coolant temperature can enable autonomous operation of generation IV reactors (GenIV), thus reducing their operating and maintenance costs. Automation can be accomplished with machine learning (ML) models trained on historical sensor data. However, the performance of ML usually depends on the availability of large amount of training data, which is difficult to obtain for GenIV, as this technology is still under development. We propose the use of transfer learning (TL), which involves utilizing knowledge across different domains, to compensate for this lack of training data. TL can be used to create pre-trained ML models with data from small-scale research facilities, which can then be fine-tuned to monitor GenIV reactors. In this work, we develop pre-trained Transformer and long short-term memory (LSTM) networks by training them on temperature measurements from thermal hydraulic flow loops operating with water and Galinstan fluids at room temperature at Argonne National Laboratory. The pre-trained models are then fine-tuned and re-trained with minimal additional data to perform predictions of the time series of high temperature measurements obtained from the Engineering Test Unit (ETU) at Kairos Power. The performance of the LSTM and Transformer networks is investigated by varying the size of the lookback window and forecast horizon. The results of this study show that LSTM networks have lower prediction errors than Transformers, but LSTM errors increase more rapidly with increasing lookback window size and forecast horizon compared to the Transformer errors.

LSTM↗

Designing a quantum-accurate machine-learning potential to enable large-scale simulations of deuterium under shock

Large-scale molecular dynamics of deuterium under shock can elucidate kinetic processes vital to the target design in inertial confinement fusion and high-energy-density experiments. However, modeling the complex evolution of this material from an insulating molecular state at ambient pressure to an ionized, atomic fluid under strong shock is beyond the capability of simple pair and even bond order potentials. We thus train a quantum-accurate and broadly transferable machine-learning interatomic potential for deuterium using the Chebyshev Interaction Model for Efficient Simulations framework. We show that due to an improved description of the molecular-to-atomic transition, our model is able to better reproduce the ab initio equation of state, radial distribution functions, and principal Hugoniot than bond order potentials. This represents an important step toward large-scale quantum-accurate and nonequilibrium simulations of complicated systems under dynamic changes including phase transitions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A deep learning and finite element approach for exploration of inverse structure–property designs of lightweight hybrid composites

Hybrid composites have important applications, such as high-performance and lightweight materials in aerospace and automotive industries. Hybrid composites utilize the synergy of diverse fillers to achieve desired material properties, but usually have more complicated microstructures. While topology optimization can optimize a particular property, designing hybrid composites for customized mechanical performances, e.g. full-range stress–strain curve, remains challenging. Here, a computational framework that integrated finite element analysis (FEA) and artificial intelligence (AI) methods of Conditional Generative Adversarial Networks (cGAN) deep learning and transfer learning was developed to establish inverse structure–property relationships and design tailor-made hybrid composites. Based on FEA-generated datasets of hybrid fiber-particle–matrix microstructures and their corresponding full-range stress–strain curves, a cGAN architecture was trained to generate tailored microstructures and establish structure–property relationships. Similarity in microstructural features and well-matched stress–strain curves based on the AI-generated composites were achieved. In conclusion, transfer learning was used to expand the pre-trained model for designing different materials systems.

Hybrid composites↗

Bi-fidelity variational auto-encoder for uncertainty quantification

Quantifying the uncertainty of quantities of interest (QoIs) from physical systems is a primary objective in model validation. However, achieving this goal entails balancing the need for computational efficiency with the requirement for numerical accuracy. To address this trade-off, we propose a novel bi-fidelity formulation of variational auto-encoders (BF-VAE) designed to estimate the uncertainty associated with a QoI from low-fidelity (LF) and high-fidelity (HF) samples of the QoI. Here, this model allows for the approximation of the statistics of the HF QoI by leveraging information derived from its LF counterpart. Specifically, we design a bi-fidelity auto-regressive model in the latent space which is integrated within the VAE’s probabilistic encoder–decoder structure. An effective algorithm is proposed to maximize the variational lower bound of the HF log-likelihood in the presence of limited HF data, resulting in the synthesis of HF realizations with a reduced computational cost. Additionally, we introduce the concept of the bi-fidelity information bottleneck (BF-IB) to provide an information-theoretic interpretation of the proposed BF-VAE model. Our numerical results demonstrate that the BF-VAE leads to considerably improved accuracy, as compared to a VAE trained using only HF data, when limited HF data is available.

42 ENGINEERING↗