Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed training”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

HydraGNN v5.0

HydraGNN v5.0 expands the code base into a more portable, scalable, and flexible framework for scientific graph learning, with particular strength in atomistic machine-learning interatomic potentials and large-scale distributed training. The release adds Fully Sharded Data Parallel (FSDP) support alongside existing DDP and DeepSpeed paths, including FSDP-aware checkpointing and optimizer integration, and introduces a configurable multi-precision training workflow supporting FP32, BF16, and FP64 across GPUs and Intel XPUs. For atomistic modeling, HydraGNN v5.0 strengthens its MLIP capabilities through dynamic graph construction at every forward pass, energy-conserving force prediction via automatic differentiation, and per-atom energy loss formulations, while extending EGNN models to properly handle periodic boundary conditions. The release also broadens model expressiveness through graph-level attribute conditioning, adds new multi-task and model-parallel extensions such as MACE support and encoder/decoder branch optimization, and expands application coverage with integrated examples for datasets including OC25, Nabla2-DFT, QCML, Open Polymers 2026, and OPF. In parallel, HydraGNN v5.0 improves production readiness through performance optimizations for large-scale runs, stratified sampling and linear-regression preprocessing utilities, and tested installation scripts for DOE supercomputers including Frontier, Aurora, Perlmutter, and Andes. Overall, the release advances HydraGNN as a robust software platform for scalable graph neural networks across materials science, chemistry, and scientific machine learning workflows

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

Scaling the SciDAC QuantOm Workflow

As part of the Scientific Discovery through Advanced Computing (SciDAC) program, the Quantum Chromodynamics Nuclear Tomography (QuantOM) project aims to analyze data from Deep Inelastic Scattering (DIS) experiments conducted at Jefferson Lab and the upcoming Electron Ion Collider. The DIS data analysis is performed on an event-level by combining the input from theoretical and experimental nuclear physics into a single, composable workflow. The optimization itself (I.e. fitting the experimental data with theoretical predictions) is carried out by a machine / deep learning algorithm. The size of the acquired DIS data as well as the complexity of the workflow itself require that the analysis is performed across multiple GPUs on high performance computing systems, such as Polaris at Argonne National Laboratory. This presentation discusses the novelties and challenges that came along with parallelizing this workflow. Recent results are compared to common distributed training techniques.

Lersch, Daniel↗

Uncertainty quantification and sensitivity analysis of a nuclear thermal propulsion reactor startup sequence

The research presented in this article describes progress in applying stochastic methods, uncertainty quantification, parametric studies, and variance-based sensitivity analysis (also known as Sobol sensitivity analysis) to a full-core model of a nuclear thermal propulsion (NTP) system simulated via the radiation transport code Griffin to simulate neutronics. Our goal is to develop a reduced-order (surrogate) model that can be rapidly sampled with perturbations to multiple input parameters. In this NTP system, reactivity and power feedback affect the rotation of control drums (CDs), which is itself controlled by a hybrid proportional-integral-derivative (PID) controller actuated by the power demand and reactivity feedback from the numerical model. This model uses reactor kinetic feedback (mean generation time [Λ] and effective delayed neutron fraction [ β eff ] from a transient Griffin simulation executed via Griffin’s improved quasi-static solver to provide the kinetic parameters) as inputs to functions that control the CD rotation angle. By investigating numerous stochastic approaches, we developed a dual-purpose surrogate model of the NTP system, using polynomial regression in the Multiphysics Object-Oriented Simulation Environment (MOOSE) Stochastic Tools Module (STM). The trained model can be rapidly sampled while simultaneously perturbing various input parameters, such as coefficients on the PID control or temperature (directly affecting the neutron cross section). The surrogate model delivers accurate (within 5%) results at speeds orders of magnitude faster (minutes, not days of computational time) than the base model. Once the surrogate model has been trained, distributions of the uncertain parameters can be changed at will to investigate the effects of perturbing multiple inputs as well as the effects of these inputs on the model output. For example, coefficients used in the PID control system may vary due to some type of physical interference, or uncertainty may exist in the temperature of the neutron cross sections in various regions of the reactor. A distribution can be placed on these parameters, and operational boundaries can be determined. The goal of this work is to support development of an advanced control system for operating CDs in a functioning NTP system. This work is a scoping study of the MOOSE STM.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

Operating Small Sat Swarms as a Single Entity: Introducing SODA

NASA's decadal survey determined that simultaneous measurements from a 3D volume of space are advantageous for a variety of studies in space physics and Earth science. Therefore, swarm concepts with multiple spacecraft in close proximity are a growing topic of interest in the small satellite community. Among the capabilities needed for swarm missions is a means to maintain operator-specified geometry, alignment, or separation. Swarm stationkeeping poses a planning challenge due to the limited scalability of ground resources. To address scalable control of orbital dynamics, we introduce SODA - Swarm Orbital Dynamics Advisor - a tool that accepts high-level configuration commands and provides the orbital maneuvers needed to achieve the desired type of swarm relative motion. Rather than conventional path planning, SODA's innovation is the use of artificial potential functions to define boundaries and keepout regions. The software architecture includes high fidelity propagation, accommodates manual or automated inputs, displays motion animations, and returns maneuver commands and analytical results. Currently, two swarm types are enabled: in-train distribution and an ellipsoid volume container. Additional swarm types, simulation applications, and orbital destinations are in planning stages.

Conn, Tracie↗

Operating Small Sat Swarms as a Single Entity: Introducing SODA

NASA's decadal survey determined that simultaneous measurements from a 3D volume of space are advantageous for a variety of studies in space physics and Earth science. Therefore, swarm concepts with multiple spacecraft in close proximity are a growing topic of interest in the small satellite community. Among the capabilities needed for swarm missions is a means to maintain operator-specified geometry, alignment, or separation. Swarm stationkeeping poses a planning challenge due to the limited scalability of ground resources. To address scalable control of orbital dynamics, we introduce SODA - Swarm Orbital Dynamics Advisor - a tool that accepts high-level configuration commands and provides the orbital maneuvers needed to achieve the desired type of swarm relative motion. Rather than conventional path planning, SODA's innovation is the use of artificial potential functions to define boundaries and keepout regions. The software architecture includes high fidelity propagation, accommodates manual or automated inputs, displays motion animations, and returns maneuver commands and analytical results. Currently, two swarm types are enabled: in-train distribution and an ellipsoid volume container. Additional swarm types, simulation applications, and orbital destinations are in planning stages.

Conn, Tracie↗

Monte Carlo Tree Search Methods for the Earth-Observing Satellite Scheduling Problem

This work explores on-board planning for the single spacecraft, multiple ground station Earth-observing satellite scheduling problem through artificial neural network function approximation of state–action value estimates generated by Monte Carlo tree search (MCTS). An extensive hyperparameter search is conducted for MCTS on the basis of performance, safety, and downlink opportunity utilization to determine the best hyperparameter combination for data generation. A hyperparameter search is also conducted on neural network architectures. The learned behavior of each network is explored, and each network architecture’s robustness to orbits and epochs outside of the training distributions is investigated. Furthermore, each algorithm is compared with a genetic algorithm, which serves to provide a baseline for optimality. MCTS is shown to compute near-optimal solutions in comparison to the genetic algorithm. The state–action value networks are shown to match or exceed the performance of MCTS in six orders of magnitude less execution time, showing promise for execution on board spacecraft.

Adam P. Herrmann↗

Uncertainty Quantification and Sensitivity Analysis of Non-Nuclear Advanced Controls Testbed Reactor Mockup

The research presented in this report describes our progress in applying stochastic methods and uncertainty quantification, parametric study, and variance-based sensitivity analysis (also known as Sobol sensitivity analysis) to a full-core model of a nuclear thermal propulsion (NTP) system simulated with Griffin, with the goal of developing a reduced order (surrogate) model which can be rapidly sampled while perturbing multiple input parameters. In this NTP system, reactivity and power feedback affect the rotation of control drums, which are controlled by a hybrid proportional, integral and derivative (PID) controller, actuated by the power demand and reactivity feedback from the numerical model. This model uses reactor kinetic feedback (mean generation time and $\beta$ from a transient Griffin simulation executed with the improved quasi-static method to provide the kinetic parameters) as inputs to functions which control the CD rotation angle. Using a number of stochastic method approaches, we developed a dual purpose training-surrogate model of the NTP system using polynomial regression. The trained model can be rapidly sampled while simultaneously perturbing various input parameters of the model, such as coefficients on the PID control, or temperature (directly affect the neutron cross section). The surrogate model delivers accurate results orders-of-magnitude faster (minutes, not days) than the base model. Once the base model has been trained, distributions of the uncertain parameters can be changed at will to investigate the effects of perturbing multiple inputs and their effect on the output. For example, coefficients used in the PID control system may vary due to some physical interference, or there may be uncertainty in the temperature of the neutron cross sections in various regions of the reactor. A distribution can be placed on these parameters and operational boundaries can be determined. The goal of this work is to support development of an advanced control system to operate CDs in a functioning NTP system.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Open Set Recognition for Unknown Waveform Classification

This presentation applies open set recognition to classify unknown waveforms, enabling systems to not only identify known types but also reliably detect when waveforms fall outside the training distribution. This approach enhances robustness by avoiding forced misclassification of novel or anomalous signals.

99 - GENERAL AND MISCELLANEOUS↗

Deep Generative Models that Solve PDEs: Distributed Computing for Training Large Data-Free Models

Recent progress in scientific machine learning (SciML) has opened up the possibility of training novel neural network architectures that solve complex partial differential equations (PDEs). Several (nearly data free) approaches have been recently reported that successfully solve PDEs, with examples including deep feed forward networks, generative networks, and deep encoder-decoder networks. However, practical adoption of these approaches is limited by the difficulty in training these models, especially to make predictions at large output resolutions (≥1024×1024). Here we report on a software framework for data parallel distributed deep learning that resolves the twin challenges of training these large SciML models - training in reasonable time as well as distributing the storage requirements. Our framework provides several out of the box functionality including (a) loss integrity independent of number of processes, (b) synchronized batch normalization, and (c) distributed higher-order optimization methods. We show excellent scalability of this framework on both cloud as well as HPC clusters, and report on the interplay between bandwidth, network topology and bare metal vs cloud. We deploy this approach to train generative models of sizes hitherto not possible, showing that neural PDE solvers can be viably trained for practical applications. We also demonstrate that distributed higher-order optimization methods are 2-3× faster than stochastic gradient-based methods and provide minimal convergence drift with higher batch-size.

PDEs↗

NFPA Distributed Energy Resources Safety Training (DERST) For Emergency Responders

The National Fire Protection Association, with support from the Department of Energy, executed a multi-year initiative to develop, enhance, and disseminate Distributed Energy Resources Safety Training (DERST) tools for U.S. emergency responders. As Distributed Energy Resources (DER)—such as solar photovoltaics, battery energy storage systems (ESS), electric vehicles (EVs), and associated infrastructure—become increasingly prevalent, the NFPA identified a critical need for up-to-date standardized, accessible, and effective safety training tailored for the fire service and related public safety professionals. The project delivered a comprehensive suite of educational resources to improve responders’ abilities to safely manage DER-related incidents. This included: • Revised Modular Training Courses: Updated classroom-based DER safety courses, now modular and accessible nationwide through fire academies and the North American Fire Training Directors (NAFTD) network. • Live Burn Testing & Research: A full-scale controlled burn of a DER-equipped residential structure provided real-world data and insights, forming the basis for updated best practices. • A Gamified Simulation Tool – Firefighters Incident Response Simulation Tool (FIRST): A first-of-its-kind, multiplayer, scenario-based simulation using the Unreal Engine 5.0 to train responders in a realistic virtual, multi-DER incident environment. • Field Familiarization Software Tools & Prop Guide: Digital DER field familiarization evolutions software guide and a prop development manual to support field-based DER training exercises, enhancing responders' hands-on familiarity with DER infrastructure and collaboration on virtual incident responses. • National Dissemination Strategy: Strategic partnerships with NAFTD, Vector Solutions, and others enabled wide-scale distribution, with over 5,000 departments accessing resources and 1,100+ departments adopting the simulator in the first seven months. Also provided a web portal for easy access to all training and simulation programs developed under this grant for the U.S. responder community. Key findings from the project—particularly from the burn test—led to paradigm shifts in fire response tactics. For example, traditional approaches to garage fires may be hazardous if DERs are present, due to explosive off gassing and thermal runaway risks. The new training emphasizes scene assessment, stand-off approaches, thermal imaging verification, and careful post-incident cooling of DER components to prevent reignition. This initiative has had a significant national impact, raising awareness, enhancing preparedness, and supporting safer DER incident response practices. Significant engagement from the media, public safety organizations, and PBS coverage has further amplified the reach and adoption of NFPA’s DER safety training, tools, and simulations.

14 SOLAR ENERGY↗

MDLoader: A Hybrid Model-Driven Data Loader for Distributed Graph Neural Network Training

Scalable data management is essential for processing large scientific dataset on HPC platforms for distributed deep learning. In-memory distributed storage is preferred for its speed, enabling rapid, random, and frequent data access required by stochastic optimizers. Processes use one-sided or collective communication to fetch remote data, with optimal performance depending on (i) dataset characteristics, (ii) training scale, and (iii) interconnection network. Empirical analysis shows collective communication excels with larger mini-batch sizes and/or fewer processes, whereas one-sided communication outperforms at larger scales. We propose MDLoader, a hybrid in-memory data loader for distributed graph neural network training. MDLoader features a model-driven performance estimator that dynamically selects between one-sided and collective communication at the beginning of training using Tree of Parzen Estimators (TPE). Evaluations on NERSC Perlmutter and OLCF Summit show MDLoader outperforms single-backend loaders by up to 2.83 × and predicts the suitable communication method with 96.3% (Perlmutter) and 94.3% (Summit) success rate.

Bae, Jonghyun↗

Cybersecurity Workforce Training for SMR Integration into Distribution Grids: A Competency Framework and Containerized Hands-On Lab for the SMR/DER/Microgrid Boundary

Small modular reactors (SMRs) and microreactors are entering the U.S. distribution grid as synchronous generation on feeders designed for loads and inverter-based distributed energy resources (DERs). No existing cybersecurity training program addresses this intersection of nuclear operations, DER management, and operational technology security. As subcontractor to Iowa State University on the CyDERMS Center, Argonne analyzed the relevant standards and training landscape, translated the resulting gaps into a twelve-objective competency framework across distribution-operator and graduate-analyst role tracks, and built a containerized training lab using a ∼400-bus composite grid model behind a realistically simulated Modbus TCP SCADA stack. The analysis isolates the balance-of-plant / energy-management-system (BOP/EMS) boundary as the critical jurisdictional seam where, as of March 2026, neither NRC nor NERC CIP cleanly claims cybersecurity responsibility for distribution-connected SMRs. The framework maps each objective across NIST CSF 2.0, ISA/IEC 62443, NIST NICE Task–Knowledge–Skill statements, and NRC RG 5.71 awareness-and-training controls. The training lab implements operator-recognition assessment scenarios spanning grid-side disturbances and telemetry-layer anomalies.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Reducing Communication in Graph Neural Network Training

Graph Neural Networks (GNNs) are powerful and flexible neural networks that use the naturally sparse connectivity information of the data. GNNs represent this connectivity as sparse matrices, which have lower arithmetic intensity and thus higher communication costs compared to dense matrices, making GNNs harder to scale to high concurrencies than convolutional or fully-connected neural networks. Here, we introduce a family of parallel algorithms for training GNNs and show that they can asymptotically reduce communication compared to previous parallel GNN training methods. We implement these algorithms, which are based on 1D, 1. 5D, 2D, and 3D sparse-dense matrix multiplication, using torch.distributed on GPU-equipped clusters. Our algorithms optimize communication across the full GNN training pipeline. We train GNNs on over a hundred GPUs on multiple datasets, including a protein network with over a billion edges.

97 MATHEMATICS AND COMPUTING↗

DDStore: Distributed Data Store for Scalable Training of Graph Neural Networks on Large Atomistic Modeling Datasets

Graph neural networks (GNNs) are a class of Deep Learning models used in designing atomistic materials for effective screening of large chemical spaces. To ensure robust prediction, GNN models must be trained on large volumes of atomistic data on leadership class supercomputers. Even with the advent of modern architectures that consist of multiple storage layers that include node-local NVMe devices in addition to device memory for caching large datasets, extreme-scale model training faces I/O challenges at scale.We present DDStore, an in-memory distributed data store designed for GNN training on large-scale graph data. DDStore provides a hierarchical, distributed, data caching technique that combines data chunking, replication, low-latency random access, and high throughput communication. DDStore achieves near-linear scaling for training a GNN model using up to 1000 GPUs on the Summit and Perlmutter supercomputers, and reaches up to a 6.15x reduction in GNN training time compared to state-of-the-art methodologies.

Choi, Jong Youl↗

Fair Concurrent Training of Multiple Models in Federated Learning

Federated learning (FL) enables collaborative learning across multiple clients. In most FL work, all clients train a single learning task. However, the recent proliferation of FL applications may increasingly require multiple FL tasks to be trained simultaneously, sharing clients’ computing resources, which we call Multiple-Model Federated Learning (MMFL). Current MMFL algorithms use naïve average-based client-task allocation schemes that often lead to unfair performance when FL tasks have heterogeneous difficulty levels, as the more difficult tasks may need more client participation to train effectively. Furthermore, in the MMFL setting, we face a further challenge that some clients may prefer training specific tasks to others, and may not even be willing to train other tasks, e.g., due to high computational costs, which may exacerbate unfairness in training outcomes across tasks. We address both challenges by firstly designing FedFairMMFL, a difficulty-aware algorithm that dynamically allocates clients to tasks in each training round, based on the tasks’ current performance levels. We provide guarantees on the resulting task fairness and FedFairMMFL’s convergence rate. We then propose novel auction designs that incentivizes clients to train multiple tasks, so as to fairly distribute clients’ training efforts across the tasks, and extend our convergence guarantees to this setting. Here, we finally evaluate our algorithm with multiple sets of learning tasks on real world datasets, showing that our algorithm improves fairness by improving the final model accuracy and convergence speed of the worst performing tasks, while maintaining the average accuracy across tasks.

Federated learning↗