Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed training”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Performance-Aligned LLMs for Generating Fast HPC Code

Optimizing scientific software is a difficult task because codebases are often large and complex, and performance can depend upon several factors including the algorithm, its implementation, and hardware among others. Causes of poor performance can originate from disparate sources and be difficult to diagnose. Recent years have seen a multitude of work that use large language models (LLMs) to assist in software development tasks. However, these tools are trained to model the distribution of code as text, and are not specifically designed to understand performance aspects of code. In this work, we introduce a reinforcement learning based methodology to align the outputs of code LLMs with performance. This allows us to build upon the current code modeling capabilities of LLMs and extend them to generate better performing code. Here, we demonstrate that our fine-tuned model improves the expected speedup of generated code over base models for a set of benchmark tasks from 0.9 to 1.6 for serial code and 1.9 to 4.5 for OpenMP parallel code.

Computer science↗

A Probabilistic Reasoner Based on Bayes Risk for Damage Detection in Structural Systems

Structural health monitoring (SHM) systems are used to inform operation of structural systems subject to loads and environments that may affect their integrity. SHM systems rely on continuous monitoring of the structure to determine its health state. These systems are often coupled with a model of the deployed structure to determine the consequences of changes in the system by forecasting the response to future states. These models, which may be thought of as digital twins, need to be updated to reflect the latest state of the structural system. This work makes use of an uncertainty-aware machine learning model that enforces distance preservation of the original input space to determine deviations from the training data input space distributions. This workflow enables domain shift detection to determine whether damage is present in the structure. The uncertainty metrics generated by this network are then used in a Bayes risk framework to design an optimal damage detector given cost and risk considerations. The approach is demonstrated on a computational example with simulated damage.

Najera-Flores, David [ATA Engineering, Inc.]↗

Field Emission Mitigation in CEBAF SRF Cavities Using Deep Learning

The Continuous Electron Beam Accelerator Facility (CEBAF) operates hundreds of superconducting radio frequency (SRF) cavities in its two main linear accelerators. Field emission can occur when the cavities are set to high operating RF gradients and is an ongoing operational challenge. This is especially true in newer, higher gradient SRF cavities. Field emission results in damage to accelerator hardware, generates high levels of neutron and gamma radiation, and has deleterious effects on CEBAF operations. So, field emission reduction is imperative for the reliable, high gradient operation of CEBAF that is required by experimenters. Here we explore the use of deep learning architectures via multilayer perceptron to simultaneously model radiation measurements at multiple detectors in response to arbitrary gradient distributions. These models are trained on collected data and could be used to minimize the radiation production through gradient redistribution. This work builds on previous efforts in developing machine learning (ML) models, and is able to produce similar model performance as our previous ML model without requiring knowledge of the field emission onset for each cavity.

Ahammed, K.↗

EM and beam dynamics modeling of CCL with CST Studio

The 800-MeV proton linac at LANSCE consists of a drift-tube linac, which brings the beam to 100 MeV, followed by a coupled-cavity linac (CCL). Each of 44 CCL modules contain multiple tanks, and it is fed by a single 805-MHz klystron. CCL tanks are multi-cell blocks of identical re-entrant sidecoupled cavities, which are followed by drifts with magnetic quadrupole doublets. Bridge couplers – special cavities displaced from the beam axis – electromagnetically couple CCL tanks over such drifts. We have developed 3D CST models of CCL tanks of the LANSCE linac. Their electromagnetic analysis is performed using MicroWave Studio. Beam dynamics is modeled with Particle Studio for bunch trains with realistic beam distributions using the CST calculated RF fields and quadrupole magnetic fields to determine the output beam parameters.

42 ENGINEERING↗

Field Emission Mitigation in CEBAF SRF Cavities Using Deep Learning

The Continuous Electron Beam Accelerator Facility (CEBAF) operates hundreds of superconducting radio frequency (SRF) cavities in its two main linear accelerators. Field emission can occur when the cavities are set to high operating RF gradients and is an ongoing operational challenge. This is especially true in newer, higher gradient SRF cavities. Field emission results in damage to accelerator hardware, generates high levels of neutron and gamma radiation, and has deleterious effects on CEBAF operations. So, field emission reduction is imperative for the reliable, high gradient operation of CEBAF that is required by experimenters. Here we explore the use of deep learning architectures via multilayer perceptron to simultaneously model radiation measurements at multiple detectors in response to arbitrary gradient distributions. These models are trained on collected data and could be used to minimize the radiation production through gradient redistribution. This work builds on previous efforts in developing machine learning (ML) models, and is able to produce similar model performance as our previous ML model without requiring knowledge of the field emission onset for each cavity.

Ahammed, K.↗

Search for HH → bbτ⁺τ⁻ Using Run 3 Scouting Data Analyze b-tagging and tau-tagging Performance with Unified Particle Transformer

B-tagging and tau-tagging performances play an important role in the search for the rare event HH → bbτ⁺τ⁻. A transformer-based neural network, Unified Particle Transformer, is applied for both tagging tasks, and Run 3 proton–proton collision scouting data at center-of-mass energy of 13.6 TeV is used. The scouting data stream accepts events at a much higher rate compared to traditional triggers, but stores only the objects reconstructed in the trigger, no low-level detector information. Therefore, existing taggers trained for the offline event reconstruction cannot be used. Analysis of the SoftMax plots, ROC/AUC curves, confusion matrix, accuracy and losses are used to evaluate model performance. Specifically, the tagging efficiency of the signal and misidentification probability across multiple background processes are compared for varying working points. Different training samples with distinct distributions of jet flavors are utilized and related model performances are analyzed. Interpretability methods, such as Integrated Gradients, may further be applied to study the input features’ influence on the model’s decisions, providing insights into potential improvements.

Chen, Blair [Purdue U., West Lafayette; Fermilab]↗

NASA Tech Briefs, October 2013

Topics include: A Short-Range Distance Sensor with Exceptional Linearity; Miniature Trace Gas Detector Based on Microfabricated Optical Resonators; Commercial Non-Dispersive Infrared Spectroscopy Sensors for Sub-Ambient Carbon Dioxide Detection; Fast, Large-Area, Wide-Bandgap UV Photodetector for Cherenkov Light Detection; Mission Data System Java Edition Version 7; Adaptive Distributed Environment for Procedure Training (ADEPT); LEGEND, a LEO-to-GEO Environment Debris Model; Electronics/Computers; Millimeter-Wave Localizers for Aircraft-to-Aircraft Approach Navigation; Impedance Discontinuity Reduction Between High-Speed Differential Connectors and PCB Interfaces; SpaceCube Version 1.5; High-Pressure Lightweight Thrusters; Non-Magnetic, Tough, Corrosion- and Wear-Resistant Knives From Bulk Metallic Glasses and Composites; Ambient Dried Aerogels; Applications for Gradient Metal Alloys Fabricated Using Additive Manufacturing; Passivation of Flexible YBCO Superconducting Current Lead With Amorphous SiO2 Layer; Propellant-Flow-Actuated Rocket Engine Igniter; Lightweight Liquid Helium Dewar for High-Altitude Balloon Payloads; Method to Increase Performance of Foil Bearings Through Passive Thermal Management; Unibody Composite Pressurized Structure; JWST Integrated Science Instrument Module Alignment Optimization Tool; Radar Range Sidelobe Reduction Using Adaptive Pulse Compression Technique; Digitally Calibrated TR Modules Enabling Real-Time Beamforming SweepSAR Architectures; Electro-Optic Time-to-Space Converter for Optical Detector Jitter Mitigation; Partially Transparent Petaled Mask/Occulter for Visible-Range Spectrum; Educational NASA Computational and Scientific Studies (enCOMPASS); Coarse-Grain Bandwidth Estimation Scheme for Large-Scale Network; Detection of Moving Targets Using Soliton Resonance Effect; High-Efficiency Nested Hall Thrusters for Robotic Solar System Exploration; High-Voltage Clock Driver for Photon-Counting CCD Characterization; Development of the Code RITRACKS; and Enabling Microliquid Chromatography by Microbead Packing of Microchannels.

Source record↗

Enhancing Security and Resiliency in Operational Technology Environments Through Network Slicing and Federated Learning

The growing convergence of Information Technology (IT) and Operational Technology (OT) within Industry 4.0 environments has introduced new demands on industrial network infrastructure. As cyber-physical systems become increasingly interconnected, ensuring the secure, timely, and efficient exchange of critical data is essential. This thesis explores how network slicing, a method of creating isolated virtual network segments, can be applied within OT environments to address challenges such as latency, security, and resource allocation. The first research question addressed in this thesis is: How can OT networks take advantage of NFV and SDN technology to become cyber resilient? This study examines the operational, security, and architectural implications of introducing network slicing into traditionally static OT infrastructures such as Industrial Control Systems (ICS) and SCADA. Through simulated deployments and case studies, the research demonstrates how slicing enables better isolation between critical and non-critical services, thereby improving response time, throughput, and security in sensitive environments. The second question considers: How to dynamically implement network slicing and take advantage of network resources towards integrating decentralized machine learning? In response, this thesis proposes a framework that combines Software-Defined Networking (SDN), Network Function Virtualization (NFV), and Federated Learning (FL) to enable real-time analytics while maintaining data locality. The proposed approach reduces the burden on centralized infrastructure and minimizes privacy risks by supporting on-site training of models across distributed OT nodes, coordinated through dynamically allocated network slices. The third focus explores: How slicing helps to increase the resiliency of OT networks through the orchestration of a dynamic DMZ? To answer this, the thesis presents a method for creating and managing Dynamic Demilitarized Zones (DMZs) using network slicing. This enables flexible and automated isolation of sensitive subsystems during threat scenarios or high-risk operations. Coupled with intelligent orchestration and containerized security services, the dynamic DMZ significantly enhances the system's ability to respond to cyber incidents without halting production. Ultimately, this thesis contributes a comprehensive architecture that blends network slicing with machine learning, secure segmentation, and automation, paving the way for resilient, adaptive, and intelligent OT environments. Performance evaluations across multiple scenarios show improvements in system reliability, threat response time, model accuracy, and resource utilization, providing a strong foundation for future industrial automation systems.

Rodiles Delgado, Brian G↗

Finding MIDDLE Ground: Scalable and Secure Distributed Learning

Edge computing methods allow devices to efficiently train a high-performing, robust, and personalized model for predictive tasks. However, these methods succumb to privacy and scalability concerns such as adversarial data recovery and expensive model communication. Furthermore, edge computing methods unrealistically assume that all devices train an identical model. In practice, edge devices have varying computational and memory constraints which may not allow certain devices to have the space or speed to train a specific model. To overcome these issues, we propose MIDDLE: a model independent distributed learning algorithm which allows heterogeneous edge devices to assist each other’s training while communicating only non-sensitive information. MIDDLE unlocks the ability for edge devices, regardless of computational or memory constraints, to assist each other even with completely different model architectures. Furthermore, MIDDLE does not require model or gradient communication which greatly reduces communication size and time. We prove that MIDDLE attains the optimal convergence rate O(1/sqrt(TM)) of stochastic gradient descent for convex and non-convex smooth optimization (for total iterations T and batch size M). Finally, our experimental results demonstrate that MIDDLE (even in non-IID data settings) attains robust and high-performing models without model or gradient communication.

Bornstein, Marc I.↗

Knowledge Oriented Graph Unified Transformer (KOGUT) v0.1

KOGUT — Knowledge Oriented Graph Unified Transformer KOGUT implements the Relational Graph Transformer (RelGT) architecture for knowledge graph link prediction in biological domains, with a primary focus on microbial growth media prediction. While the original RelGT (arXiv:2505.10960) targets relational tables, time series, and multi-table databases, KOGUT adapts this architecture for heterogeneous biological knowledge graphs, providing first-in-class AI predictive models for microbial cultivation. Key Adaptations Beyond Original RelGT: - Knowledge Graph Focus: Applied to biological KGs with semantic node types (taxa, chemicals, media, phenotypes, environments) versus generic relational database tables, trained on the KG-Microbe knowledge graph (1.3M entities, 2.9M edges, 24 relation types). - Multimodal Node Encoding: Integrates node labels, categories, descriptions, and synonyms from KG metadata through learned embedding layers—adapting relational column features to graph node attributes with textual semantics. - Extended K-Hop Subgraph Strategy: Optimized neighborhood sampling (3-hop default, configurable up to 200 nodes) tuned for sparse biological networks, building on the original local-global attention framework with biological relation preservation. - Biolink Predicate Preservation: Type-specific transformations for 24 biological edge semantics (occurs_in, consumes, produces, has_phenotype, subclass_of) beyond standard relational foreign keys, enabling multi-relation link prediction. - Inductive Learning Support: Enables zero-shot predictions for novel taxa through feature-based embeddings (temperature, oxygen requirements, gram stain, cell shape), extending the original transductive relational benchmark scope to uncultured microorganisms. CheapSOTA Performance Optimizations (This Distribution): - VQ-EMA Centroid Attention: Vector quantization with exponential moving average for improved global context modeling (+5-10% MRR improvement). - HDF5 Precomputed Data Loading: One-time preprocessing of k-hop subgraphs to eliminate redundant graph traversals (2-5× training speedup). - Distributed Data Parallel Training: Multi-GPU support for scaling to larger knowledge graphs (tested on 4× NVIDIA A100 GPUs at NERSC Perlmutter). - Mixed Precision Training: Automatic mixed precision (AMP) for memory efficiency and faster training. Advantages Over Standard Knowledge Graph Embedding Models: Combines RelGT's proven multi-element tokenization (features, type, hop, structure) with graph-native biological representations, enabling interpretable link prediction across heterogeneous entities that standard embedding models (TransE, RotatE, ComplEx) and table-based transformers cannot directly model. Achieves near-perfect performance on microbial growth media prediction (MRR: 0.9966, Precision@1: 0.9932, Hit@10: 1.0000) while maintaining explainability through attention-based reasoning over biological pathways. Training Data: - KG-Microbe merged knowledge graph: 1,379,337 nodes, 2,960,472 edges - 24 biological relation types including taxonomic hierarchies, metabolic interactions, phenotype associations, and environmental relationships - Primary prediction task: Growth media suitability for microbial taxa (biolink:occurs_in, 50K edges) - Multi-relation capability: Predicts links for any of the 24 relation types, including chemical consumption/production, phenotype associations, and taxonomic classification Citation: Original RelGT Architecture: Dwivedi et al., "Relational Graph Transformer", arXiv:2505.10960, 2025 KOGUT Implementation: Knowledge Oriented Graph Unified Transformer for Microbial Growth Media Prediction Developed at Lawrence Berkeley National Laboratory (LBNL) Trained on NERSC Perlmutter supercomputer

Joachimiak, Marcin [Lawrence Berkeley National Lab↗

The location and scope of geographic remote sensing training in the United States

Maps displaying the distribution of graduate departments of geography in the United States and enrollments in remote sensing courses in all geography departments during the past two calendar years were compiled. It was anticipated that the two distributions would show a marked similarity since remote sensing is a relatively new geographic tool requiring specialized training to use as well as equipment not normally found in most geography departments. Thus only the larger graduate departments can afford to devote time and resources to this specialty. A broad correspondence does exist between the graduate departments of geography and the courses in remote ensing. However, the correlation is far from complete and the exceptions are frequent and large enough to cast doubt upon the accuracy of the original hypothesis. Whereas many large departments do offer courses in remote sensing, many smaller colleges and universities do also. A number of possible explanations can be offered for the discrepancies: (1) course titles, (2) the liberal arts orientation of geography departments in many universities, (3) job-oriented skills which many smaller departments have emphasized, and (4) in the tight job market many new graduates of even the larger departments have had to accept position in smaller departments and colleges.

Hawley, A. J.↗

Applications of intelligent computer-aided training

Intelligent computer-aided training (ICAT) systems simulate the behavior of an experienced instructor observing a trainee, responding to help requests, diagnosing and remedying trainee errors, and proposing challenging new training scenarios. This paper presents a generic ICAT architecture that supports the efficient development of ICAT systems for varied tasks. In addition, details of ICAT projects, built with this architecture, that deliver specific training for Space Shuttle crew members, ground support personnel, and flight controllers are presented. Concurrently with the creation of specific ICAT applications, a general-purpose software development environment for ICAT systems is being built. The widespread use of such systems for both ground-based and on-orbit training will serve to preserve task and training expertise, support the training of large numbers of personnel in a distributed manner, and ensure the uniformity and verifiability of training experiences.

Loftin, R. B.↗

A general-purpose development environment for intelligent computer-aided training systems

Space station training will be a major task, requiring the creation of large numbers of simulation-based training systems for crew, flight controllers, and ground-based support personnel. Given the long duration of space station missions and the large number of activities supported by the space station, the extension of space shuttle training methods to space station training may prove to be impractical. The application of artificial intelligence technology to simulation training can provide the ability to deliver individualized training to large numbers of personnel in a distributed workstation environment. The principal objective of this project is the creation of a software development environment which can be used to build intelligent training systems for procedural tasks associated with the operation of the space station. Current NASA Johnson Space Center projects and joint projects with other NASA operational centers will result in specific training systems for existing space shuttle crew, ground support personnel, and flight controller tasks. Concurrently with the creation of these systems, a general-purpose development environment for intelligent computer-aided training systems will be built. Such an environment would permit the rapid production, delivery, and evolution of training systems for space station crew, flight controllers, and other support personnel. The widespread use of such systems will serve to preserve task and training expertise, support the training of many personnel in a distributed manner, and ensure the uniformity and verifiability of training experiences. As a result, significant reductions in training costs can be realized while safety and the probability of mission success can be enhanced.

Savely, Robert T.↗

Neuromorphic learning of continuous-valued mappings from noise-corrupted data. Application to real-time adaptive control

The ability of feed-forward neural network architectures to learn continuous valued mappings in the presence of noise was demonstrated in relation to parameter identification and real-time adaptive control applications. An error function was introduced to help optimize parameter values such as number of training iterations, observation time, sampling rate, and scaling of the control signal. The learning performance depended essentially on the degree of embodiment of the control law in the training data set and on the degree of uniformity of the probability distribution function of the data that are presented to the net during sequence. When a control law was corrupted by noise, the fluctuations of the training data biased the probability distribution function of the training data sequence. Only if the noise contamination is minimized and the degree of embodiment of the control law is maximized, can a neural net develop a good representation of the mapping and be used as a neurocontroller. A multilayer net was trained with back-error-propagation to control a cart-pole system for linear and nonlinear control laws in the presence of data processing noise and measurement noise. The neurocontroller exhibited noise-filtering properties and was found to operate more smoothly than the teacher in the presence of measurement noise.

Troudet, Terry↗

Efficient Distributed Sequence Parallelism for Transformer-Based Image Segmentation

We introduce an efficient distributed sequence parallel approach for training transformer-based deep learning image segmentation models. The neural network models are comprised of a combination of a Vision Transformer encoder with a convolutional decoder to provide image segmentation mappings. The utility of the distributed sequence parallel approach is especially useful in cases where the tokenized embedding representation of image data are too large to fit into standard computing hardware memory. To demonstrate the performance and characteristics of our models trained in sequence parallel fashion compared to standard models, we evaluate our approach using a 3D MRI brain tumor segmentation dataset. We show that training with a sequence parallel approach can match standard sequential model training in terms of convergence. Furthermore, we show that our sequence parallel approach has the capability to support training of models that would not be possible on standard computing resources.

Lyngaas, Isaac↗

Event‐Based Training in Label‐Limited Regimes

Abstract The distribution of attributes assigned using data on independent sensors for a specific source, for example, magnitude, can be richly descriptive for final event characterization and associated uncertainty. Attribute distributions can also provide powerful context for event characterization in the absence of comprehensive annotation. This work develops a way to leverage distributional information across a set of sensors in the absence of comprehensive annotation as a domain‐informed regularization term applied during gradient‐based learning. The regularization term is the basis of event‐based training which I show can be a powerful semi‐supervised learning (SSL) approach. I first use a simple feed forward neural network and a toy data set to outline how data set structure interacts with the assumptions inherent to many semi‐supervised learning approaches. I then demonstrate the effectiveness of event‐based training using a deep convolutional neural network for seismic event classification in Utah, which increases SSL accuracy from 92% to 97% on event classification with a limited number of training labels.

Linville, Lisa M.↗

A-Train Data Depot: Integrating and Visualizing Atmospheric Measurements Along the A-Train Tracks

The succession of US and international satellites that follow each other, seconds to minutes apart, across the local afternoon equator crossing is called the ATrain. The A-Train consists of the following satellites, in order of equator crossing: OCO, EOS Aqua, CloudSat, CALIPSO, PARASOL, and EOS Aura. Flying in such formation increases the number of observations, validates observations, and enables coordination between science observations, resulting in a more complete virtual science platform (Kelly, 2000) The goal of this project is to create the first ever A-Train virtual data portal/center, the A-Train Data Depot, to process, archive, access, visualize, analyze and correlate distributed atmosphere measurements from various A-Train instruments along A-Train tracks. The A-Train Data Depot (ATDD) will enable the free movement of remotely located A-Train data so that they are combined to create a consolidated vertical view of the Earth s Atmosphere along the A-Train tracks. Once the infrastructure of the ATDD is in place, it will be easily evolved to serve data from all A-Train data measurements: one stop shopping. The innovative approach of analyzing and visualizing atmospheric profiles along the platforms track (i.e., time) will be accommodated by reusing the GSFC Atmospheric Composition Data and Information Services Center (ACDISC) visualization and analysis tool, GIOVANNI, existing data reduction tools, on-line archwing for fast data access, and Cooperative Institute for Research in the Atmosphere (CRA) data co-registration tools. Initial measurements utilized include CALIPSO lidar backscatter, CloudSat radar reflectivity, clear air relative humidity, water vapor and temperature from AIRS, and cloud properties and aerosols from both MODIS. This will be followed by associated measurements from MLS, OMI, HIRDLS, and TES. Given the independent nature of instrument/platform development, the ATDD project has been met with many interesting challenges that, once resolved, will provide a much greater understanding of the relative flight dynamics and data co-registration of the suite of A-Train instruments, thus greatly increasing the accuracy of A-Train data analysis. Some of these challenges will be discussed. The project s resulting visualizations and analysis illustrate the importance of managing data so that measurements from various missions can be combined to enhance the understanding of the atmosphere. A-Train data management coordination, as performed here, is extremely significant in facilitating the A-Train science of clouds, precipitation, aerosol and chemistry.

Kempler, Steven↗

Integration of Utility Distributed Energy Resource Management System and Aggregators for Evolving Distribution System Operators

With the rapid integration of distributed energy resources (DERs), distribution utilities are faced with new and unprecedented issues. New challenges introduced by high penetration of DERs range from poor observability to overload and reverse power flow problems, under-over-voltages, maloperation of legacy protection systems, and requirements for new planning procedures. Distribution utility personnel are not adequately trained, and legacy control centers are not properly equipped to cope with these issues. Fortunately, distribution energy resource management systems (DERMSs) are emerging software technologies aimed to provide distribution system operators (DSOs) with a specialized set of tools to enable them to overcome the issues caused by DERs and to maximize the benefits of the presence of high penetration of these novel resources. However, as DERMS technology is still emerging, its definition is vague and can refer to very different levels of software hierarchies, spanning from decentralized virtual power plants to DER aggregators and fully centralized enterprise systems (called utility DERMS). Although they are all frequently simply called DERMS, these software technologies have different sets of tools and aim to provide different services to different stakeholders. This paper explores how these different software technologies can complement each other, and how they can provide significant benefits to DSOs in enabling them to successfully manage evolving distribution networks with high penetration of DERs when they are integrated together into the control centers of distribution utilities.

24 POWER TRANSMISSION AND DISTRIBUTION↗