Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed training”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Search for HH → bbτ⁺τ⁻ Using Run 3 Scouting Data Analyze b-tagging and tau-tagging Performance with Unified Particle Transformer

B-tagging and tau-tagging performances play an important role in the search for the rare event HH → bbτ⁺τ⁻. A transformer-based neural network, Unified Particle Transformer, is applied for both tagging tasks, and Run 3 proton–proton collision scouting data at center-of-mass energy of 13.6 TeV is used. The scouting data stream accepts events at a much higher rate compared to traditional triggers, but stores only the objects reconstructed in the trigger, no low-level detector information. Therefore, existing taggers trained for the offline event reconstruction cannot be used. Analysis of the SoftMax plots, ROC/AUC curves, confusion matrix, accuracy and losses are used to evaluate model performance. Specifically, the tagging efficiency of the signal and misidentification probability across multiple background processes are compared for varying working points. Different training samples with distinct distributions of jet flavors are utilized and related model performances are analyzed. Interpretability methods, such as Integrated Gradients, may further be applied to study the input features’ influence on the model’s decisions, providing insights into potential improvements.

Chen, Blair [Purdue U., West Lafayette; Fermilab]↗

Enhancing Security and Resiliency in Operational Technology Environments Through Network Slicing and Federated Learning

The growing convergence of Information Technology (IT) and Operational Technology (OT) within Industry 4.0 environments has introduced new demands on industrial network infrastructure. As cyber-physical systems become increasingly interconnected, ensuring the secure, timely, and efficient exchange of critical data is essential. This thesis explores how network slicing, a method of creating isolated virtual network segments, can be applied within OT environments to address challenges such as latency, security, and resource allocation. The first research question addressed in this thesis is: How can OT networks take advantage of NFV and SDN technology to become cyber resilient? This study examines the operational, security, and architectural implications of introducing network slicing into traditionally static OT infrastructures such as Industrial Control Systems (ICS) and SCADA. Through simulated deployments and case studies, the research demonstrates how slicing enables better isolation between critical and non-critical services, thereby improving response time, throughput, and security in sensitive environments. The second question considers: How to dynamically implement network slicing and take advantage of network resources towards integrating decentralized machine learning? In response, this thesis proposes a framework that combines Software-Defined Networking (SDN), Network Function Virtualization (NFV), and Federated Learning (FL) to enable real-time analytics while maintaining data locality. The proposed approach reduces the burden on centralized infrastructure and minimizes privacy risks by supporting on-site training of models across distributed OT nodes, coordinated through dynamically allocated network slices. The third focus explores: How slicing helps to increase the resiliency of OT networks through the orchestration of a dynamic DMZ? To answer this, the thesis presents a method for creating and managing Dynamic Demilitarized Zones (DMZs) using network slicing. This enables flexible and automated isolation of sensitive subsystems during threat scenarios or high-risk operations. Coupled with intelligent orchestration and containerized security services, the dynamic DMZ significantly enhances the system's ability to respond to cyber incidents without halting production. Ultimately, this thesis contributes a comprehensive architecture that blends network slicing with machine learning, secure segmentation, and automation, paving the way for resilient, adaptive, and intelligent OT environments. Performance evaluations across multiple scenarios show improvements in system reliability, threat response time, model accuracy, and resource utilization, providing a strong foundation for future industrial automation systems.

Rodiles Delgado, Brian G↗

Finding MIDDLE Ground: Scalable and Secure Distributed Learning

Edge computing methods allow devices to efficiently train a high-performing, robust, and personalized model for predictive tasks. However, these methods succumb to privacy and scalability concerns such as adversarial data recovery and expensive model communication. Furthermore, edge computing methods unrealistically assume that all devices train an identical model. In practice, edge devices have varying computational and memory constraints which may not allow certain devices to have the space or speed to train a specific model. To overcome these issues, we propose MIDDLE: a model independent distributed learning algorithm which allows heterogeneous edge devices to assist each other’s training while communicating only non-sensitive information. MIDDLE unlocks the ability for edge devices, regardless of computational or memory constraints, to assist each other even with completely different model architectures. Furthermore, MIDDLE does not require model or gradient communication which greatly reduces communication size and time. We prove that MIDDLE attains the optimal convergence rate O(1/sqrt(TM)) of stochastic gradient descent for convex and non-convex smooth optimization (for total iterations T and batch size M). Finally, our experimental results demonstrate that MIDDLE (even in non-IID data settings) attains robust and high-performing models without model or gradient communication.

Bornstein, Marc I.↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

Efficient Distributed Sequence Parallelism for Transformer-Based Image Segmentation

We introduce an efficient distributed sequence parallel approach for training transformer-based deep learning image segmentation models. The neural network models are comprised of a combination of a Vision Transformer encoder with a convolutional decoder to provide image segmentation mappings. The utility of the distributed sequence parallel approach is especially useful in cases where the tokenized embedding representation of image data are too large to fit into standard computing hardware memory. To demonstrate the performance and characteristics of our models trained in sequence parallel fashion compared to standard models, we evaluate our approach using a 3D MRI brain tumor segmentation dataset. We show that training with a sequence parallel approach can match standard sequential model training in terms of convergence. Furthermore, we show that our sequence parallel approach has the capability to support training of models that would not be possible on standard computing resources.

Lyngaas, Isaac↗

Event‐Based Training in Label‐Limited Regimes

Abstract The distribution of attributes assigned using data on independent sensors for a specific source, for example, magnitude, can be richly descriptive for final event characterization and associated uncertainty. Attribute distributions can also provide powerful context for event characterization in the absence of comprehensive annotation. This work develops a way to leverage distributional information across a set of sensors in the absence of comprehensive annotation as a domain‐informed regularization term applied during gradient‐based learning. The regularization term is the basis of event‐based training which I show can be a powerful semi‐supervised learning (SSL) approach. I first use a simple feed forward neural network and a toy data set to outline how data set structure interacts with the assumptions inherent to many semi‐supervised learning approaches. I then demonstrate the effectiveness of event‐based training using a deep convolutional neural network for seismic event classification in Utah, which increases SSL accuracy from 92% to 97% on event classification with a limited number of training labels.

Linville, Lisa M.↗

Integration of Utility Distributed Energy Resource Management System and Aggregators for Evolving Distribution System Operators

With the rapid integration of distributed energy resources (DERs), distribution utilities are faced with new and unprecedented issues. New challenges introduced by high penetration of DERs range from poor observability to overload and reverse power flow problems, under-over-voltages, maloperation of legacy protection systems, and requirements for new planning procedures. Distribution utility personnel are not adequately trained, and legacy control centers are not properly equipped to cope with these issues. Fortunately, distribution energy resource management systems (DERMSs) are emerging software technologies aimed to provide distribution system operators (DSOs) with a specialized set of tools to enable them to overcome the issues caused by DERs and to maximize the benefits of the presence of high penetration of these novel resources. However, as DERMS technology is still emerging, its definition is vague and can refer to very different levels of software hierarchies, spanning from decentralized virtual power plants to DER aggregators and fully centralized enterprise systems (called utility DERMS). Although they are all frequently simply called DERMS, these software technologies have different sets of tools and aim to provide different services to different stakeholders. This paper explores how these different software technologies can complement each other, and how they can provide significant benefits to DSOs in enabling them to successfully manage evolving distribution networks with high penetration of DERs when they are integrated together into the control centers of distribution utilities.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The PAU Survey: Photometric redshifts using transfer learning from simulations

In this paper, we introduce the DEEPZ deep learning photometric redshift (photo-z) code. As a test case, we apply the code to the PAU survey (PAUS) data in the COSMOS field. DEEPZ reduces the σ68 scatter statistic by 50 percent at iAB = 22.5 compared to existing algorithms. This improvement is achieved through various methods, including transfer learning from simulations where the training set consists of simulations as well as observations, which reduces the need for training data. The redshift probability distribution is estimated with a mixture density network (MDN), which produces accurate redshift distributions. Our code includes an autoencoder to reduce noise and extract features from the galaxy SEDs. It also benefits from combining multiple networks, which lowers the photo-z scatter by 10 percent. Furthermore, training with randomly constructed coadded fluxes adds information about individual exposures, reducing the impact of photometric outliers. In addition to opening up the route for higher redshift precision with narrow bands, these machine learning techniques can also be valuable for broad-band surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

An Active Learning-Based Streaming Pipeline for Reduced Data Training of Structure Finding Models in Neutron Diffractometry

Structure determination workloads in neutron diffractometry are computationally expensive and routinely require several hours to many days to determine the structure of a material from its neutron diffraction patterns. The potential for machine learning models trained on simulated neutron scattering patterns to significantly speed up these tasks have been reported recently. However, the amount of simulated data needed to train these models grows exponentially with the number of structural parameters to be predicted and poses a significant computational challenge. To overcome this challenge, we introduce a novel batch-mode active learning (AL) policy that uses uncertainty sampling to simulate training data drawn from a probability distribution that prefers labelled examples about which the model is least certain. We confirm its efficacy in training the same models with ∼ 75% less training data while improving the accuracy. We then discuss the design of an efficient stream-based training workflow that uses this AL policy and present a performance study on two heterogeneous platforms to demonstrate that, compared with a conventional training workflow, the streaming workflow delivers ∼ 20% shorter training time without any loss of accuracy.

Wang, Tianle [Brookhaven National Laboratory (BNL)↗

Deep Image Prior Enabled Full Waveform Inversion (Final Technical Report)

MS Student Naveen Gupta worked on the problem of full waveform inversion (FWI) using neural networks as shown in Figure 1. Our goal was to learn a neural network to represent the subsurface velocity model, which when fed into the FWI module (implemented using a numerical forward model of wave equations) produces amplitude estimates that match with ground-truth observations of amplitude. We used neural networks to solve the inverse problem of estimating velocity distributions for a given seismic amplitude data such that, once trained, our neural network model can generate a distribution of velocity profiles for different random vectors fed as inputs to the neural network model.

97 MATHEMATICS AND COMPUTING↗

Efficient distributed continual learning for steering experiments in real-time

Deep learning has emerged as a powerful method for extracting valuable information from large volumes of data. However, when new training data arrives continuously (i.e., is not fully available from the beginning), incremental training suffers from catastrophic forgetting (i.e., new patterns are reinforced at the expense of previously acquired knowledge). Training from scratch each time new training data becomes available would result in extremely long training times and massive data accumulation. Rehearsal-based continual learning has shown promise for addressing the catastrophic forgetting challenge, but research to date has not addressed performance and scalability. To fill this gap, we propose an approach based on a distributed rehearsal buffer that efficiently complements data-parallel training on multiple GPUs to achieve high accuracy, short runtime, and scalability. It leverages a set of buffers (local to each GPU) and uses several asynchronous techniques for updating these local buffers in an embarrassingly parallel fashion, all while handling the communication overheads necessary to augment input minibatches using unbiased, global sampling. We further propose a generalization of rehearsal buffers to support both classification and generative learning tasks, as well as more advanced rehearsal strategies (notably Dark Experience Replay, leveraging knowledge distillation). We illustrate this approach with a real-life HPC streaming application from the domain of ptychographic image reconstruction. Furthermore, we run extensive experiments on up to 128 GPUs of the ThetaGPU supercomputer to compare our approach with baselines representative of training-from-scratch (the upper bound in terms of accuracy) and incremental training (the lower bound). Results show that rehearsal-based continual learning achieves a top-5 validation accuracy close to the upper bound, while simultaneously exhibiting a runtime close to the lower bound.

Asynchronous data management↗

Novel machine-learning method for spin classification of neutron resonances

The performance of nuclear reactors and other nuclear systems depends on a precise understanding of the neutron interaction cross sections for materials used in these systems. These cross sections exhibit resonant structure whose shape is determined in part by the angular-momentum quantum numbers of the resonances. The correct assignment of the quantum numbers of neutron resonances is, therefore, paramount. In this project, we apply machine learning to automate the quantum number assignments using only the resonances' energies and widths and not relying on detailed transmission or capture measurements. The classifier used for quantum number assignment is trained using stochastically generated resonance sequences whose distributions mimic those of real data. Here we explore the use of several physics-motivated features for training our classifier. These features amount to out-of-distribution tests of a given resonance's widths and resonance-pair spacings. We pay special attention to situations where either capture widths cannot be trusted for classification purposes or where there is insufficient information to classify resonances by the total spin J. We demonstrate the efficacy of our classification approach using simulated and actual 52 Cr resonance data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Time-lapse seismic data inversion for estimating reservoir parameters using deep learning

Geologic carbon sequestration involves the injection of captured carbon dioxide ([Formula: see text]) into subsurface formations for long-term storage. The movement and fate of the injected [Formula: see text] plume is of great concern to regulators because monitoring helps to identify potential leakage zones and determines the possibility of safe long-term storage. To address this concern, we design a deep-learning framework for [Formula: see text] saturation monitoring to determine the geologic controls on the storage of the injected [Formula: see text]. We use different combinations of porosities and permeabilities for a given reservoir to generate saturation and velocity models. We train the deep-learning model with a few time-lapse seismic images and their corresponding changes in saturation values for a particular [Formula: see text] injection site. The deep-learning model learns the mapping from the change in the time-lapse seismic response to the change in [Formula: see text] saturation during the training phase. We then apply the trained model to data sets comprising different time-lapse seismic image slices (corresponding to different time instances) generated using different porosity and permeability distributions that are not part of the training to estimate the [Formula: see text] saturation values along with the plume extent. Our algorithm provides a deep-learning assisted framework for the direct estimation of [Formula: see text] saturation values and plume migration in heterogeneous formations using the time-lapse seismic data. Our method improves the efficiency of time-lapse inversion by streamlining the large number of intermediate steps in the conventional time-lapse inversion workflow. This method also helps to incorporate the geologic uncertainty for a given reservoir by accounting for the statistical distribution of porosity and permeability during the training phase. Tests on different examples verify the effectiveness of our approach.

Geochemistry & Geophysics↗

Solving multiphysics-based inverse problems with learned surrogates and constraints

Abstract Solving multiphysics-based inverse problems for geological carbon storage monitoring can be challenging when multimodal time-lapse data are expensive to collect and costly to simulate numerically. We overcome these challenges by combining computationally cheap learned surrogates with learned constraints. Not only does this combination lead to vastly improved inversions for the important fluid-flow property, permeability, it also provides a natural platform for inverting multimodal data including well measurements and active-source time-lapse seismic data. By adding a learned constraint, we arrive at a computationally feasible inversion approach that remains accurate. This is accomplished by including a trained deep neural network, known as a normalizing flow, which forces the model iterates to remain in-distribution, thereby safeguarding the accuracy of trained Fourier neural operators that act as surrogates for the computationally expensive multiphase flow simulations involving partial differential equation solves. By means of carefully selected experiments, centered around the problem of geological carbon storage, we demonstrate the efficacy of the proposed constrained optimization method on two different data modalities, namely time-lapse well and time-lapse seismic data. While permeability inversions from both these two modalities have their pluses and minuses, their joint inversion benefits from either, yielding valuable superior permeability inversions and CO 2 plume predictions near, and far away, from the monitoring wells.

Yin, Ziyi (ORCID:0000000250248771)↗

An Efficient Distributed Reinforcement Learning for Enhanced Multi-Microgrid Management

Economic dispatch in multi-microgrid (MMG) systems requires coordinating distributed energy resources (DERs) of different microgrids, which leads to a significant increase in the number of states for energy management. In these cases, traditional reinforcement learning (RL) approaches become computationally expensive or output a solution that causes extra-operating costs for the system. This paper proposes an RL approach that employs local learning agents to interact with microgrid environments in a distributed manner and aggregates the outcomes to train the global agent to learn the policy for the MMG system. This distributed exploration and aggregation process provides an effective solution and guides the global agent to learn the dispatch policy efficiently. Case studies are performed on a system with three microgrids with different types of DERs. Results obtained using the proposed RL and comparisons with conventional methods substantiate the effectiveness of the proposed approach in terms of operation costs, computation time, and peak-to-average ratio.

Das, Avijit↗

On the minimum number of radiation field parameters to specify gas cooling and heating functions

Fast and accurate approximations of gas cooling and heating functions are needed for hydrodynamic galaxy simulations. We use machine learning to analyze atomic gas cooling and heating functions in the presence of a generalized incident local radiation field computed by Cloudy. We characterize the radiation field through binned radiation field intensities instead of the photoionization rates used in our previous work. We find a set of 6 energy bins whose intensities exhibit relatively low correlation. We use these bins as features to train machine learning models to predict Cloudy cooling and heating functions at fixed metallicity. We compare the relative SHapley Additive exPlanation (SHAP) value importance of the features. From the SHAP analysis, we identify a feature subset of 3 energy bins (0.5-1, 1-4, and 13-16Ry) with the largest importance and train additional models on this subset. We compare the mean squared errors and distribution of errors on both the entire training data table and a randomly selected 20% test set withheld from model training. The machine learning models trained with 3 and 6 bins, as well as 3 and 4 photoionization rates, have comparable accuracy everywhere, with errors ≳10 times smaller than for the interpolation table of Gnedin and Hollon (2012). We conclude that 3 energy bins (or 3 analogous photoionization rates: molecular hydrogen photodissociation, neutral hydrogen HI, and fully ionized carbon CVI) are sufficient to characterize the dependence of the gas cooling and heating functions on our assumed incident radiation field model.

79 ASTRONOMY AND ASTROPHYSICS↗

On the minimum number of radiation field parameters to specify gas cooling and heating functions

Fast and accurate approximations of gas cooling and heating functions are needed for hydrodynamic galaxy simulations. We use machine learning to analyze atomic gas cooling and heating functions computed by Cloudy in the presence of a generalized incident local radiation field. We characterize the radiation field through binned radiation field intensities instead of the photoionization rates used in our previous work. We find a set of 6 energy bins whose intensities exhibit relatively low correlation. We use these bins as features to train machine learning models to predict Cloudy cooling and heating functions at fixed metallicity. We compare the relative SHapley Additive exPlanation (SHAP) value importance of the features. From the SHAP analysis, we identify a feature subset of 3 energy bins ($0.5-1, 1-4$, and $13-16 \, \mathrm{Ry}$) with the largest importance and train additional models on this subset. We compare the mean squared errors and distribution of errors on both the entire training data table and a randomly selected 20% test set withheld from model training. The machine learning models trained with 3 and 6 bins, as well as 3 and 4 photoionization rates, have comparable accuracy everywhere, with errors $\gtrsim 10$ times smaller than for the interpolation table of Gnedin and Hollon (2012). We conclude that 3 energy bins (or 3 analogous photoionization rates: molecular hydrogen photodissociation, neutral hydrogen HI, and fully ionized carbon CVI) are sufficient to characterize the dependence of the gas cooling and heating functions on our assumed incident radiation field model.

79 ASTRONOMY AND ASTROPHYSICS↗