Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed training”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

SIDDA: SInkhorn Dynamic Domain Adaptation

Modern neural networks (NNs) often do not generalize well in the presence of a "covariate shift"; that is, in situations where the training and test data distributions differ, but the conditional distribution of classification labels remains unchanged. In such cases, NN generalization can be reduced to a problem of learning more domain-invariant features. Domain adaptation (DA) methods include a range of techniques aimed at achieving this; however, these methods have struggled with the need for extensive hyperparameter tuning, which then incurs significant computational costs. In this work, we introduce SIDDA, an out-of-the-box DA training algorithm built upon the Sinkhorn divergence, that can achieve effective domain alignment with minimal hyperparameter tuning and computational overhead. We demonstrate the efficacy of our method on multiple simulated and real datasets of varying complexity, including simple shapes, handwritten digits, and real astronomical observations. SIDDA is compatible with a variety of NN architectures, and it works particularly well in improving classification accuracy and model calibration when paired with equivariant neural networks (ENNs). We find that SIDDA enhances the generalization capabilities of NNs, achieving up to a ≈40% improvement in classification accuracy on unlabeled target data. We also study the efficacy of DA on ENNs with respect to the varying group orders of the dihedral group DN, and find that the model performance improves as the degree of equivariance increases. Finally, we find that SIDDA enhances model calibration on both source and target data--achieving over an order of magnitude improvement in the ECE and Brier score. SIDDA's versatility, combined with its automated approach to domain alignment, has the potential to advance multi-dataset studies by enabling the development of highly generalizable models.

Pandya, Sneh [Northeastern U.]↗

Measuring the thermal and ionization state of the low- z IGM using likelihood free inference

ABSTRACT We present a new approach to measure the power-law temperature density relationship $T=T_0 (\rho/ \bar{\rho })^{\gamma -1}$ and the UV background photoionization rate $\Gamma _{{{{\rm H\, {\small I}}}}{}}$ of the intergalactic medium (IGM) based on the Voigt profile decomposition of the Ly α forest into a set of discrete absorption lines with Doppler parameter b and the neutral hydrogen column density $N_{\rm H\, {\small I}}$. Previous work demonstrated that the shape of the $b-N_{{{{\rm H\, {\small I}}}}{}}$ distribution is sensitive to the IGM thermal parameters T0 and γ, whereas our new inference algorithm also takes into account the normalization of the distribution, i.e. the line-density dN/dz, and we demonstrate that precise constraints can also be obtained on $\Gamma _{{{{\rm H\, {\small I}}}}{}}$. We use density-estimation likelihood-free inference (DELFI) to emulate the dependence of the $b-N_{{{{\rm H\, {\small I}}}}{}}$ distribution on IGM parameters trained on an ensemble of 624 nyx hydrodynamical simulations at z = 0.1, which we combine with a Gaussian process emulator of the normalization. To demonstrate the efficacy of this approach, we generate hundreds of realizations of realistic mock HST/COS data sets, each comprising 34 quasar sightlines, and forward model the noise and resolution to match the real data. We use this large ensemble of mocks to extensively test our inference and empirically demonstrate that our posterior distributions are robust. Our analysis shows that by applying our new approach to existing Ly α forest spectra at z ≃ 0.1, one can measure the thermal and ionization state of the IGM with very high precision ($\sigma _{\log T_0} \sim 0.08$ dex, σγ ∼ 0.06, and $\sigma _{\log \Gamma _{{{{\rm H\, {\small I}}}}{}}} \sim 0.07$ dex).

79 ASTRONOMY AND ASTROPHYSICS↗

SIDDA: SInkhorn Dynamic Domain Adaptation for image classification with equivariant neural networks

Modern neural networks (NNs) often do not generalize well in the presence of a ‘covariate shift’; that is, in situations where the training and test data distributions differ, but the conditional distribution of classification labels given the data remains unchanged. In such cases, NN generalization can be reduced to a problem of learning more robust, domain-invariant features. Domain adaptation (DA) methods include a broad range of techniques aimed at achieving this; however, these methods have struggled with the need for extensive hyperparameter tuning, which then incurs significant computational costs. In this work, we introduce SInkhorn Dynamic Domain Adaptation (SIDDA), an out-of-the-box DA training algorithm built upon the Sinkhorn divergence, that can achieve effective domain alignment with minimal hyperparameter tuning and computational overhead. We demonstrate the efficacy of our method on multiple simulated and real datasets of varying complexity, including simple shapes, handwritten digits, real astronomical observations, and remote sensing data. These datasets exhibit covariate shifts due to noise, blurring, differences between telescopes, and variations in imaging wavelengths. SIDDA is compatible with a variety of NN architectures, and it works particularly well in improving classification accuracy and model calibration when paired with symmetry-aware equivariant NNs (ENNs). We find that SIDDA consistently enhances the generalization capabilities of NNs, achieving up to a ${\approx}40\%$ improvement in classification accuracy on unlabeled target data, while also providing a more modest performance gain of $\lesssim 1\%$ on labeled source data. We also study the efficacy of DA on ENNs with respect to the varying group orders of the dihedral group DN, and find that the model performance improves as the degree of equivariance increases. Finally, if SIDDA achieves proper domain alignment, it also enhances model calibration on both source and target data, with the most significant gains in the unlabeled target domain—achieving over an order of magnitude improvement in the expected calibration error and Brier score. SIDDA’s versatility across various NN models and datasets, combined with its automated approach to domain alignment, has the potential to significantly advance multi-dataset studies by enabling the development of highly generalizable models.

79 ASTRONOMY AND ASTROPHYSICS↗

RLC4CLR (Reinforcement Learning Controller for Critical Load Restoration Problems)

RLC4CLR demonstrates using a reinforcement learning controller (RLC) to solve a critical load restoration (CLR) problem, which improves the grid resilience after a substation outage event. RLC4CLR consists of two parts. (1) RL environment: This environment encapsulates the CLR problem to be solved and provides interfacing functions to follow the standard OpenAI Gym format. A power system simulator, i.e., OpenDSS, is included to provide the power flow solution. Controller inputs and outputs (RL state and action) as well as the reward are defined in this environment as well. In summary, the RL environment is the problem formulation from which the RL agent can learn. (2) RL training script: The training script enables the RL agent to learn its control policy by interacting with the RL environment. For RL training, an open-sourced RL library, i.e., RLlib, is leveraged which is based on a distributed computing framework (Ray). The training script is designed to be able to be run on both local machine or the NREL HPC system. Other components of RLC4CLR include input data, e.g., grid model (standard IEEE test feeders), and other files used for results analysis.

Zhang, Xiangyu↗

Scale-up Unlearnable Examples Learning with High-performance Computing

Recent advancements in AI models, like ChatGPT, are structured to retain user interactions, which could inadvertently include sensitive healthcare data. In the healthcare field, particularly when radiologists use AI-driven diagnostic tools hosted on online platforms, there is a risk that medical imaging data may be repurposed for future AI training without explicit consent, spotlighting critical privacy and intellectual property concerns around healthcare data usage. Addressing these privacy challenges, a novel approach known as Unlearnable Examples (UEs) has been introduced, aiming to make data unlearnable to deep learning models. A prominent method within this area, called Unlearnable Clustering (UC), has shown improved UE performance with larger batch sizes but was previously limited by computational resources (e.g., a single workstation). To push the boundaries of UE performance with theoretically unlimited resources, we scaled up UC learning across various datasets using Distributed Data Parallel (DDP) training on the Summit supercomputer. Our goal was to examine UE efficacy at high-performance computing (HPC) levels to prevent unauthorized learning and enhance data security, particularly exploring the impact of batch size on UE’s unlearnability. Utilizing the robust computational capabilities of the Summit, extensive experiments were conducted on diverse datasets such as Pets, MedMNist, Flowers, and Flowers102. Our findings reveal that both overly large and overly small batch sizes can lead to performance instability and affect accuracy. However, the relationship between batch size and unlearnability varied across datasets, highlighting the necessity for tailored batch size strategies to achieve optimal data protection. The use of Summit’s high-performance GPUs, along with the efficiency of the DDP framework, facilitated rapid updates of model parameters and consistent training across nodes. Our results underscore the critical role of selecting appropriate batch sizes based on the specific characteristics of each dataset to prevent learning and ensure data security in deep learning applications. The source code is publicly available at https: // github. com/ hrlblab/ UE_ HPC .

Zhu, Yanfan [Vanderbilt University, Nashville, TN,↗

Experimental validation of a high fidelity Monte Carlo neutron transport model of the MIT graphite exponential pile

High-fidelity modeling and simulation were performed for the MIT graphite exponential pile (MGEP) using Monte Carlo neutron transport codes OpenMC and MCNP, and the results were validated by experimental data. The MGEP is being used as the test bed for the design of an autonomous control system for the pile's neutron flux distribution. The main contribution of this work is to generate the training data sets of neutron flux distributions with different locations of control rods that perturb the neutron flux profiles. First, code -to-code cross verification between OpenMC and MCNP was performed to ensure consistency of the numerical modeling within statistical uncertainties. To validate the accuracy of this high-fidelity model, a series of neutron flux measurements were conducted using a Helium-3 (He-3) neutron detector on a mobile platform that is placed inside the pile. Second, the neutron flux profiles were measured in four vertical layers of interest, and compared to the corresponding simulation results. The comparison results shows that the root mean square error is less than 2.5% in the two upper layers, and less than 4.5% in all four measured layers. Here the results validated the accuracy of the modeling and simulation. Finally, the relative change of the neutron flux profiles from moving control rods was analyzed, which identified the layer that has the best sensitivity regarding the control rods movements. Thus, this work identified and provided training data sets of both simulated and experimental neutron flux profiles in the most sensitive layer, paving the path forward to the real-time experimental demonstration of the autonomous control system.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Dynamic Model Agnostic Reliability Evaluation of Machine-Learning Models Integrated in Instrumentation & Control Systems

In recent years, the field of machine learning (ML), specifically neural networks, has grown significantly and has spurred research in its applicability to digital instrumentation and control systems (DI&C). While ML models have shown promise in operational contexts, the trustworthiness of using such algorithms has not been adequately assessed. Failures of ML integrated systems are not well understood, and the lack of comprehensive risk modeling can degrade the trustworthiness in these systems. In recent reports by the National Institute for Standards and Technology (NIST) [1] and the Nuclear Regulatory Commission (NRC) [2], they indicate that trustworthiness in ML is a critical barrier and will play a vital role in the safe, accountable, and secure operation of intelligent systems. Thus, in this work, we demonstrate a dynamic model-agnostic method to quantify the relative reliability of AI/ML predictions by incorporating out-of-distribution (OOD) detection on the training dataset. It is well documented that most ML algorithms excel at interpolation (or near-interpolation) tasks but experience significant performance degradation at extrapolation. The method, referenced as the Laplacian distributed decay for reliability (LADDR), determines the difference between the operational and training datasets which can used to the relative reliability of AI/ML predictions. LADDR is then demonstrated on a feedforward neural network based digital twin used for the prediction of safety significant factors during a loss-of-flow transient. LADDR is used to demonstrate how training data can be used as evidence to support the relative reliability of ML/AI predictions enhancing the overall trustworthiness of the system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Scalable training of graph convolutional neural networks for fast and accurate predictions of HOMO-LUMO gap in molecules

Abstract Graph Convolutional Neural Network (GCNN) is a popular class of deep learning (DL) models in material science to predict material properties from the graph representation of molecular structures. Training an accurate and comprehensive GCNN surrogate for molecular design requires large-scale graph datasets and is usually a time-consuming process. Recent advances in GPUs and distributed computing open a path to reduce the computational cost for GCNN training effectively. However, efficient utilization of high performance computing (HPC) resources for training requires simultaneously optimizing large-scale data management and scalable stochastic batched optimization techniques. In this work, we focus on building GCNN models on HPC systems to predict material properties of millions of molecules. We use HydraGNN, our in-house library for large-scale GCNN training, leveraging distributed data parallelism in PyTorch. We use ADIOS, a high-performance data management framework for efficient storage and reading of large molecular graph data. We perform parallel training on two open-source large-scale graph datasets to build a GCNN predictor for an important quantum property known as the HOMO-LUMO gap. We measure the scalability, accuracy, and convergence of our approach on two DOE supercomputers: the Summit supercomputer at the Oak Ridge Leadership Computing Facility (OLCF) and the Perlmutter system at the National Energy Research Scientific Computing Center (NERSC). We present our experimental results with HydraGNN showing (i) reduction of data loading time up to 4.2 times compared with a conventional method and (ii) linear scaling performance for training up to 1024 GPUs on both Summit and Perlmutter.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Emulating galaxy and peculiar velocity clustering on non-linear scales

We explore the potential of cross-correlating galaxies and peculiar velocities on non-linear scales to enhance cosmological constraints. Leveraging the ABACUSSUMMIT simulation suite and the halo occupation distribution (HOD) formalism, we trained emulator models to describe the non-linear clustering of galaxies and velocities in redshift space. Our analysis demonstrates that combining galaxy and peculiar velocity clustering provides tighter constraints on both HOD and cosmological parameters, particularly on σ8 and w0. We further applied our models to realistic mock catalogues, reproducing the expected density and peculiar velocity errors of type-Ia supernovae, Tully-Fisher and fundamental plane measurements for the combined ZTF and DESI measurements. While systematic biases arise in the HOD parameters, the cosmological constraints remain unbiased, yielding a 3.8% precision measurement on fσ 8 compared to 4.7% when using galaxy clustering alone. We demonstrate that while combining tracers with realistic velocity measurements still yields an improvement, the gains are diminished, highlighting the need for further efforts to reduce velocity measurement uncertainties and correct observational systematics on small scales.

79 ASTRONOMY AND ASTROPHYSICS↗

Model fusion with physics-guided machine learning: Projection-based reduced-order modeling

The unprecedented amount of data generated from experiments, field observations, and large-scale numerical simulations at a wide range of spatiotemporal scales has enabled the rapid advancement of data-driven and especially deep learning models in the field of fluid mechanics. Although these methods are proven successful for many applications, there is a grand challenge of improving their generalizability. This is particularly essential when data-driven models are employed within outer-loop applications like optimization. In this work, we put forth a physics-guided machine learning (PGML) framework that leverages the interpretable physics-based model with a deep learning model. Leveraging a concatenated neural network design from multi-modal data sources, the PGML framework is capable of enhancing the generalizability of data-driven models and effectively protects against or inform about the inaccurate predictions resulting from extrapolation. We apply the PGML framework as a novel model fusion approach combining the physics-based Galerkin projection model and long- to short-term memory (LSTM) network for parametric model order reduction of fluid flows. We demonstrate the improved generalizability of the PGML framework against a purely data-driven approach through the injection of physics features into intermediate LSTM layers. Our quantitative analysis shows that the overall model uncertainty can be reduced through the PGML approach, especially for test data coming from a distribution different than the training data. Moreover, we demonstrate that our approach can be used as an inverse diagnostic tool providing a confidence score associated with models and observations. The proposed framework also allows for multi-fidelity computing by making use of low-fidelity models in the online deployment of quantified data-driven models.

42 ENGINEERING↗

Parameter uncertainties for imperfect surrogate models in the low-noise regime

Abstract Bayesian regression determines model parameters by minimizing the expected loss, an upper bound to the true generalization error. However, this loss ignores model form error, or misspecification, meaning parameter uncertainties are significantly underestimated and vanish in the large data limit. As misspecification is the main source of uncertainty for surrogate models of low-noise calculations, such as those arising in atomistic simulation, predictive uncertainties are systematically underestimated. We analyze the true generalization error of misspecified, near-deterministic surrogate models, a regime of broad relevance in science and engineering. We show that posterior parameter distributions must cover every training point to avoid a divergence in the generalization error and design a compatible ansatz which incurs minimal overhead for linear models. The approach is demonstrated on model problems before application to thousand-dimensional datasets in atomistic machine learning. Our efficient misspecification-aware scheme gives accurate prediction and bounding of test errors in terms of parameter uncertainties, allowing this important source of uncertainty to be incorporated in multi-scale computational workflows.

Swinburne, Thomas D. (ORCID:0000000232554257)↗

Nuclear masses learned from a probabilistic neural network

Machine learning methods and uncertainty quantification have been gaining interest throughout the last several years in low-energy nuclear physics. In particular, Gaussian processes and Bayesian neural networks have increasingly been applied to improve mass model predictions while providing well-quantified uncertainties. In this work, we use the probabilistic Mixture Density Network (MDN) to directly predict the mass excess of the 2016 Atomic Mass Evaluation within the range of measured data, and we extrapolate the inferred models beyond available experimental data. The MDN provides not only mean values but also full posterior distributions both within the training set and extrapolated testing set. We show that the addition of physical information to the feature space increases the accuracy of the match to the training data as well as provides for more physically meaningful extrapolations beyond the the limits of experimental data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Reinforcement Learning via Gaussian Processes with Neural Network Dual Kernels

While deep neural networks (DNNs) and Gaussian Processes (GPs) are both popularly utilized to solve problems in reinforcement learning, both approaches feature undesirable drawbacks for challenging problems. DNNs learn complex non-linear embeddings, but do not naturally quantify uncertainty and are often data-inefficient to train. GPs infer posterior distributions over functions, but popular kernels exhibit limited expressivity on complex and high-dimensional data. Fortunately, recently discovered conjugate and neural tangent kernel functions encode the behavior of overparameterized neural networks in the kernel domain. We demonstrate that these kernels can be efficiently applied to regression and reinforcement learning problems by analyzing a baseline case study.We apply GPs with neural network dual kernels to solve reinforcement learning tasks for the first time. We demonstrate, using the well understood mountain-car problem, that GPs empowered with dual kernels perform at least as well as those using the conventional radial basis function kernel. Finally, we conjecture that by inheriting the probabilistic rigor of GPs and the powerful embedding properties of DNNs, GPs using NN dual kernels will empower future reinforcement learning models on difficult domains.

97 MATHEMATICS AND COMPUTING↗

Performance-Aligned LLMs for Generating Fast HPC Code

Optimizing scientific software is a difficult task because codebases are often large and complex, and performance can depend upon several factors including the algorithm, its implementation, and hardware among others. Causes of poor performance can originate from disparate sources and be difficult to diagnose. Recent years have seen a multitude of work that use large language models (LLMs) to assist in software development tasks. However, these tools are trained to model the distribution of code as text, and are not specifically designed to understand performance aspects of code. In this work, we introduce a reinforcement learning based methodology to align the outputs of code LLMs with performance. This allows us to build upon the current code modeling capabilities of LLMs and extend them to generate better performing code. Here, we demonstrate that our fine-tuned model improves the expected speedup of generated code over base models for a set of benchmark tasks from 0.9 to 1.6 for serial code and 1.9 to 4.5 for OpenMP parallel code.

Computer science↗

A Probabilistic Reasoner Based on Bayes Risk for Damage Detection in Structural Systems

Structural health monitoring (SHM) systems are used to inform operation of structural systems subject to loads and environments that may affect their integrity. SHM systems rely on continuous monitoring of the structure to determine its health state. These systems are often coupled with a model of the deployed structure to determine the consequences of changes in the system by forecasting the response to future states. These models, which may be thought of as digital twins, need to be updated to reflect the latest state of the structural system. This work makes use of an uncertainty-aware machine learning model that enforces distance preservation of the original input space to determine deviations from the training data input space distributions. This workflow enables domain shift detection to determine whether damage is present in the structure. The uncertainty metrics generated by this network are then used in a Bayes risk framework to design an optimal damage detector given cost and risk considerations. The approach is demonstrated on a computational example with simulated damage.

Najera-Flores, David [ATA Engineering, Inc.]↗

Field Emission Mitigation in CEBAF SRF Cavities Using Deep Learning

The Continuous Electron Beam Accelerator Facility (CEBAF) operates hundreds of superconducting radio frequency (SRF) cavities in its two main linear accelerators. Field emission can occur when the cavities are set to high operating RF gradients and is an ongoing operational challenge. This is especially true in newer, higher gradient SRF cavities. Field emission results in damage to accelerator hardware, generates high levels of neutron and gamma radiation, and has deleterious effects on CEBAF operations. So, field emission reduction is imperative for the reliable, high gradient operation of CEBAF that is required by experimenters. Here we explore the use of deep learning architectures via multilayer perceptron to simultaneously model radiation measurements at multiple detectors in response to arbitrary gradient distributions. These models are trained on collected data and could be used to minimize the radiation production through gradient redistribution. This work builds on previous efforts in developing machine learning (ML) models, and is able to produce similar model performance as our previous ML model without requiring knowledge of the field emission onset for each cavity.

Ahammed, K.↗

EM and beam dynamics modeling of CCL with CST Studio

The 800-MeV proton linac at LANSCE consists of a drift-tube linac, which brings the beam to 100 MeV, followed by a coupled-cavity linac (CCL). Each of 44 CCL modules contain multiple tanks, and it is fed by a single 805-MHz klystron. CCL tanks are multi-cell blocks of identical re-entrant sidecoupled cavities, which are followed by drifts with magnetic quadrupole doublets. Bridge couplers – special cavities displaced from the beam axis – electromagnetically couple CCL tanks over such drifts. We have developed 3D CST models of CCL tanks of the LANSCE linac. Their electromagnetic analysis is performed using MicroWave Studio. Beam dynamics is modeled with Particle Studio for bunch trains with realistic beam distributions using the CST calculated RF fields and quadrupole magnetic fields to determine the output beam parameters.

42 ENGINEERING↗

Field Emission Mitigation in CEBAF SRF Cavities Using Deep Learning

The Continuous Electron Beam Accelerator Facility (CEBAF) operates hundreds of superconducting radio frequency (SRF) cavities in its two main linear accelerators. Field emission can occur when the cavities are set to high operating RF gradients and is an ongoing operational challenge. This is especially true in newer, higher gradient SRF cavities. Field emission results in damage to accelerator hardware, generates high levels of neutron and gamma radiation, and has deleterious effects on CEBAF operations. So, field emission reduction is imperative for the reliable, high gradient operation of CEBAF that is required by experimenters. Here we explore the use of deep learning architectures via multilayer perceptron to simultaneously model radiation measurements at multiple detectors in response to arbitrary gradient distributions. These models are trained on collected data and could be used to minimize the radiation production through gradient redistribution. This work builds on previous efforts in developing machine learning (ML) models, and is able to produce similar model performance as our previous ML model without requiring knowledge of the field emission onset for each cavity.

Ahammed, K.↗