Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed training”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Machine Learning for Distributed Acoustic Sensing data (MLDAS) v1.0.1

MLDAS is a Python-written package for exploratory data analysis and deep learning training on Distributed Acoustic Sensing data. The machine learning tools are powered by the PyTorch library and designed to work efficiently on large scale datasets using parallel computing. Various SLURM scripts as well as a tutorial have also been made available to allow geophysicists to quickly and easily implement the available tools in their analysis workflow on supercomputer facilities.

Dumont, Vincent↗

Adversarial sampling of unknown and high-dimensional conditional distributions

Many engineering problems require the prediction of realization-to-realization variability or a refined description of modeled quantities. In that case, it is necessary to sample elements from unknown high-dimensional spaces with possibly millions of degrees of freedom. While there exist methods able to sample elements from probability density functions (PDF) with known shapes, several approximations need to be made when the distribution is unknown. In this paper the sampling method, as well as the inference of the underlying distribution, are both handled with a data-driven method known as generative adversarial networks (GAN), which trains two competing neural networks to produce a network that can effectively generate samples from the training set distribution. In practice, it is often necessary to draw samples from conditional distributions. When the conditional variables are continuous, only one (if any) data point corresponding to a particular value of a conditioning variable may be available, which is not sufficient to estimate the conditional distribution. This work handles this problem using an a priori estimation of the conditional moments of a PDF. Herein, two approaches, stochastic estimation, and an external neural network are compared for computing these moments; however, any preferred method can be used. The algorithm is demonstrated in the case of the deconvolution of a filtered turbulent flow field. It is shown that all the versions of the proposed algorithm effectively sample the target conditional distribution with minimal impact on the quality of the samples compared to state-of-the-art methods. Additionally, the procedure can be used as a metric for the diversity of samples generated by a conditional GAN (cGAN) conditioned with continuous variables.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Roadmap and Benchmarking: Privacy in Federated Load Forecasting

Data-driven techniques for energy demand forecasting continue to emerge with promising impacts on distribution grid planning. However, the development of robust and generalizable machine learning models requires that representative high quality training data are available. Distributed energy resources have begun to embed intelligence, gathering large amounts of data on customer demand, behavior, and household devices that are connected to the grid. Though utilities aggregate meter-level demand data for load shaping, demand response, outage management, reliability planning, and billing applications, there lies an inherent privacy concern in sharing consumption data that may identify individual consumer behavioral patterns. Hence, while sharing the data is crucial, the private sensitive customer data must be safeguarded from being exposed or manipulated. In this study, we propose a roadmap for implementing a based privacy preserving framework to support the advancement of data-driven analytics in data-sensitive distributed energy resources environments. The roadmap incorporates federated learning–a distributed training framework, differential privacy–a statistical framework that provides guarantees to safeguard the leakage of sensitive data, secure multiparty computation and homomorphic encryption– techniques for encrypting model gradients and applying secure aggregation on the server. Moreover, we perform baseline experiments on the federated short-term load forecasting (STLF) task using open-source residential load profile datasets, offering insights into the challenges of integrating differential privacy into federated learning.

Abebe, Waqwoya [Oak Ridge National Laboratory (ORN↗

DNTTD (Distributed Non-Negative Tensor Train Decomposition)

The era of exascale computing opens new venues for innovations and discoveries in many scientific, engineering, and commercial fields. However, with the exa flops also come the extra-large high-dimensional data generated by high performance computing. High-dimensional data is presented as multidimensional arrays, aka tensors. The presence of latent (not directly observable) structures in the tensor allows a unique representation and compression of the data by classical tensor factorization techniques. However, the classical tensor methods are not always stable or they can be exponential in their memory requirements, which makes them not suitable for high-dimensional tensors. Tensor train (TT) is a state-of-the-art tensor network introduced for factorization of high-dimensional tensors. TT transforms the initial high-dimensional tensor in a network of three-dimensional tensors that requires only a linear storage. Many real-world data, such as, density, temperature, population, probability, etc., are non-negative and for an easy interpretation, the algorithms preserving non-negativity are preferred. Here, we introduce a distributed non-negative tensor-train and demonstrate its scalability and the compression on synthetic and real world big datasets.

Bhattarai, Manish↗

DP-TwoLevel: two-stage gradient subspace learning for differentially private federated learning

Federated learning (FL) enables collaborative model training across distributed data sources without sharing raw data, but faces fundamental challenges in communication efficiency and privacy. Differentially private (DP) training mitigates information leakage but introduces noise that degrades model performance, especially in high-dimensional settings. We propose DP-TwoLevel, a hierarchical gradient projection method that improves utility under fixed DP constraints by exploiting low-dimensional structure in model updates. Our approach learns a two-level PCA-based representation of gradients and applies DP noise in a reduced-dimensional subspace, thereby lowering the effective noise magnitude while preserving dominant signal components. We evaluate the method across three datasets (MNIST, Fashion-MNIST, CIFAR-10) and three privacy regimes (ϵ∈0.5, 1.0, 2.0). Across nine experimental settings, DP-TwoLevel consistently outperforms DP-FedAvg, achieving an average accuracy improvement of 9.44%, with larger gains observed in lower ϵ(higher-noise) regimes (up to +22.31%). We further analyze scalability across models ranging from 100K to 1.49M parameters and identify a variance-based success criterion: performance remains strong when the projection preserves more than 75% of gradient variance, degrades in a marginal regime (65–75%), and fails below this threshold. Our results demonstrate that structure-aware dimensionality reduction can significantly improve the privacy–utility tradeoff in FL without modifying formal privacy guarantees. We also provide empirical evidence of scaling limitations for global projections and motivate per-layer extensions for larger models.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Stable parallel training of Wasserstein conditional generative adversarial neural networks

In this work, we propose a stable, parallel approach to train Wasserstein conditional generative adversarial neural networks (W-CGANs) under the constraint of a fixed computational budget. Differently from previous distributed GANs training techniques, our approach avoids inter-process communications, reduces the risk of mode collapse and enhances scalability by using multiple generators, each one of them concurrently trained on a single data label. The use of the Wasserstein metric also reduces the risk of cycling by stabilizing the training of each generator. We illustrate the approach on the CIFAR10, CIFAR100, and ImageNet1k datasets, three standard benchmark image datasets, maintaining the original resolution of the images for each dataset. Performance is assessed in terms of scalability and final accuracy within a limited fixed computational time and computational resources. To measure accuracy, we use the inception score, the Fréchet inception distance, and image quality. An improvement in inception score and Fréchet inception distance is shown in comparison to previous results obtained by performing the parallel approach on deep convolutional conditional generative adversarial neural networks as well as an improvement of image quality of the new images created by the GANs approach. Weak scaling is attained on both datasets using up to 2000 NVIDIA V100 GPUs on the OLCF supercomputer Summit.

97 MATHEMATICS AND COMPUTING↗

Gradient-enhanced physics-informed neural networks for forward and inverse PDE problems

Deep learning has been shown to be an effective tool in solving partial differential equations (PDEs) through physics-informed neural networks (PINNs). PINNs embed the PDE residual into the loss function of the neural network, and have been successfully employed to solve diverse forward and inverse PDE problems. However, one disadvantage of the first generation of PINNs is that they usually have limited accuracy even with many training points. Here, we propose a new method, gradient-enhanced physics-informed neural networks (gPINNs), for improving the accuracy of PINNs. gPINNs leverage gradient information of the PDE residual and embed the gradient into the loss function. We tested gPINNs extensively and demonstrated the effectiveness of gPINNs in both forward and inverse PDE problems. Our numerical results show that gPINN performs better than PINN with fewer training points. Additionally, we combined gPINN with the method of residual-based adaptive refinement (RAR), a method for improving the distribution of training points adaptively during training, to further improve the performance of gPINN, especially in PDEs with solutions that have steep gradients.

42 ENGINEERING↗

Network Anomaly Detection in Distributed Edge Computing Infrastructure

As networks continue to grow in complexity and scale, detecting anomalies has become increasingly challenging, particularly in diverse and geographically dispersed environments. Traditional approaches often struggle with managing the computational burden associated with analyzing large-scale network traffic to identify anomalies. This paper introduces a distributed edge computing framework that integrates federated learning with Apache Spark and Kubernetes to address these challenges. We hypothesize that our approach, which enables collaborative model training across distributed nodes, significantly enhances the detection accuracy of network anomalies across different network types. We show that by leveraging distributed computing and containerization technologies, our framework not only improves scalability and fault tolerance but also achieves superior detection performance compared to state-of-the-art methods. Extensive experiments on the UNSW-NB15 and ROAD datasets validate the effectiveness of our approach, demonstrating statistically significant improvements in detection accuracy and training efficiency over baseline models, as confirmed by MannWhitney U and Kolmogorov-Smirnov tests (p<0.05).

Marfo, William [University of Texas at El Paso,Dep↗

MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs

Graph Neural Networks (GNN) are indispensable in learning from graph-structured data, yet their rising computational costs, especially on massively connected graphs, pose significant challenges in terms of execution performance. To tackle this, distributed-memory solutions such as partitioning the graph to concurrently train multiple replicas of GNNs are in practice. However, approaches requiring a partitioned graph usually suffer from communication overhead and load imbalance, even under optimal partitioning and communication strategies due to irregularities in the neighborhood minibatch sampling. This paper proposes practical trade-offs for improving the sampling and communication overheads for representation learn- ing on distributed graphs (using popular GraphSAGE architecture) by developing a parameterized prefetch and eviction scheme on top of the state-of-the-art Amazon DistDGL distributed GNN framework, demonstrating about 15–40% improvement in end-to-end training performance on the NERSC Perlmutter supercomputer for various OGB datasets.

Machine Leanring, high performance comptuing, grap↗

Novel usage of deep learning and high-performance computing in long-baseline neutrino oscillation experiments

Mención Internacional en el título de doctorDeep-learning methods are playing a crucial role in numerous scientific and industrialapplications. Over the past two decades, these techniques have helped in the collection,reconstruction, and analysis of large data samples in particle physics experiments. Themain topic of this PhD research is the study of deep-learning techniques in long-baselineneutrino oscillation experiments. Neutrinos are mysterious light elementary particles,and their investigation is essential to shed light on some of the remaining open questionsin physics. The work presented here describes an algorithm based on a convolutionalneural network developed to provide highly accurate and efficient selections of electronneutrino and muon neutrino interactions in the Deep Underground Neutrino Experiment(DUNE). With this algorithm, the electron neutrino (antineutrino) selection efficiencypeaks at 90% (94%) and exceeds 85% (90%) for reconstructed neutrino energies between2-5 GeV. The selection efficiency for muon neutrino (antineutrino) interactions is foundto have a maximum of 96% (97%) and exceeds 90% (95%) efficiency for reconstructedneutrino energies above 2 GeV. When considering all electron neutrino and antineutrinointeractions as signal (both those appearing from oscillations and those intrinsic tothe beam), a selection purity of 90% is achieved. These event selections are criticalto maximise the sensitivity of the experiment to CP-violating effects, key to furtherunderstand the matter-antimatter asymmetry of the Universe.In high-energy physics experiments, deep learning has also been explored for producingfast simulations and physically-motivated manipulations of simulated images. Some ofthose simulations, such as the light production and detection, are very computationallyexpensive and require novel methods to produce the necessary samples while controllingthe varied underlying physics model parameters. To do so, we invented the model-assistedgenerative adversarial network (MAGAN), first validated on simple generic case studiesand then successfully applied to the DUNE photon-detector simulation.Moreover, we also developed graph neural networks for 3D-voxel classification ofambiguities and optical crosstalk for a different particle physics experiment, most preciselyfor the proposed SuperFGD. This novel 3D-granular plastic-scintillator neutrino detectorwill be used to upgrade the near detector of the T2K neutrino oscillation experiment, and our method reports efficiencies and purities of 94-96% per event in the classificationof particle track voxels.Due to the growth and complexity of deep neural networks, researchers have beeninvestigating techniques to train those networks in a more computationally-efficient way.Many efforts have been made by the community to optimise deep-learning models byparallelising or distributing their training computation across multiple devices. In thisthesis, we study an approach based on data locality for those neural networks that cannotbenefit from scaling their computation due to a significant bottleneck in the data I/O.The research also includes a detailed study on the performance of deep neural networkson hardware accelerator boards.Los métodos de aprendizaje profundo son cada vez más utilizados en numerosas aplicacionescientíficas e industriales hoy en día. Durante las dos últimas décadas, estastécnicas se han empleado en la recolección, reconstrucción y análisis de la gran cantidadde datos generados por experimentos de física de partículas. El tema principal de estatesis doctoral es el uso de estos modelos de aprendizaje profundo en experimentos defísica de neutrinos, en concreto en los experimentos de larga distancia DUNE y T2K. Losneutrinos, partículas fundamentales neutras, de las más ligeras del Universo, pueden serclave para explicar algunas de las cuestiones todavía sin resolver en física fundamental.Entre las diferentes contribuciones que esta tesis ha hecho a su estudio, cabe destacar eldesarrollo de un algoritmo basado en una red de neuronas convolucional para seleccionarcon gran eficiencia y precisión las interacciones de neutrinos electrónicos y muónicos enel Deep Underground Neutrino Experiment (DUNE). La eficiencia de selección obtenidapara neutrinos (antineutrinos) electrónicos alcanza un máximo del 90% (94%) y supera el85% (90%) para neutrinos con energías reconstruidas en el rango 2-5 GeV. La selección deneutrinos (antineutrinos) muónicos tiene una eficiencia máxima del 96% (97%) y excedeel 90% (95%) para neutrinos con energías reconstruidas de más de 2 GeV. Considerandocomo señal todas las interacciones de neutrinos y antineutrinos electrónicos (procedentestanto de oscilaciones como intrínsecos en el haz inicial), se logra una pureza en la seleccióndel 90%. Dichas selecciones de eventos son fundamentales para maximizar la sensibilidaddel experimento a los efectos de violació...

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Scalable Deep-Learning-Accelerated Topology Optimization for Additively Manufactured Materials

Topology optimization (TO) is a popular and powerful computational approach for designing novel structures, materials, and devices. Two computational challenges have limited the applicability of TO to a variety of industrial applications. First, a TO problem often involves a large number of design variables to guarantee sufficient expressive power. Second, many TO problems require a large number of expensive physical model simulations, and those simulations cannot be parallelized. To address these issues, we propose a general scalable deep-learning (DL) based TO framework, referred to as SDL-TO, which utilizes parallel schemes in high performance computing (HPC) to accelerate the TO process for designing additively manufactured (AM) materials. Unlike the existing studies of DL for TO, our framework accelerates TO by learning the iterative history data and simultaneously training on the mapping between the given design and its gradient. The surrogate gradient is learned by utilizing parallel computing on multiple CPUs incorporated with a distributed DL training on multiple GPUs. The learned TO gradient enables a fast online update scheme instead of an expensive update based on the physical simulator or solver. Using a local sampling strategy, we achieve to reduce the intrinsic high dimensionality of the design space and improve the training accuracy and the scalability of the SDL-TO framework. The method is demonstrated by benchmark examples and AM materials design for heat conduction. The proposed SDL-TO framework shows competitive performance compared to the baseline methods but significantly reduces the computational cost by a speed up of around 8.6x over the standard TO implementation.

Bi, Sirui↗

FitCache: A Transparent Drop-In Framework for Multi-Tier Caching to Accelerate Distributed Deep Learning Workloads

Training in Deep learning (DL) remains highly compute- and data-intensive, with I/O becoming a critical bottleneck as models and datasets scale. Recent studies report that data loading can dominate training time, especially on large-scale HPC systems with shared parallel file systems (PFS). Existing caching approaches either rely on single-tier designs or require intrusive modifications to training pipelines, limiting their portability and effectiveness. In this work, we present FitCache, a transparent drop-in framework for multi-tier caching to accelerate distributed DL training by coordinating fast local memory (e.g., DRAM, Persistent Memory (PMem)) and NVMe as hierarchical caches atop PFS. Our design adapts to hardware diversity, i.e., if NVMe is missing, memory transparently acts as a caching tier, ensuring stable performance. FitCache transparently intercepts I/O requests and issues concurrent fetches across all tiers, returning data from the fastest responder without centralized metadata or static redirection paths. FitCache adapts to dynamic workloads and heterogeneous clusters while maintaining POSIX compatibility. Experiments on Frontier (2048 GPUs) and smaller research clusters show that FitCache reduces training time by up to 40% and per-batch I/O latency by up to 71.6% compared to Lustre Orion PFS, offering a drop-in solution for scalable DL training.

Hu, Guangxing [ORNL] (ORCID:0009000283203614)↗

Data-based filtered dissipation rate modelling for multi-modal turbulent combustion: evaluating a priori model generalizability

Manifold-based models offer a computationally efficient alternative to directly transporting the thermochemical state in computational simulations of turbulent reacting flows, projecting the high-dimensional thermochemical state-space onto a low-dimensional manifold. Recent efforts have yielded a manifold-based model applicable to multi-modal combustion, enabling reconstruction of the thermochemical state from solutions to two-dimensional manifold equations in mixture fraction and generalized progress variable that are parameterised by three scalar dissipation rates. In coarse-grained simulations such as Large Eddy Simulation (LES), closure of the multi-modal manifold equations and subfilter variances/covariance requires closure of three filtered scalar dissipation rates. Here, the present work adopts a data-based approach, providing closure for the three filtered scalar dissipation rates via deep neural networks (DNNs). High-fidelity datasets corresponding to an autoigniting n-dodecane jet flame and a bluff body swirl-stabilized confined lifted spray flame of two aviation fuels (Jet-A and C1) with different ignition propensities are leveraged to generate training data that spans a diverse range of thermodynamic conditions and combustion modes, including low- and high-temperature ignition regimes in addition to premixed and nonpremixed behaviour. A final DNN model is trained to enforce inherent physical constraints by learning nonlinear functional transformations of the three filtered scalar dissipation rates. The generalizability of this constrained DNN model is demonstrated a priori via conditional statistics evaluated on the lifted spray flame with C1–a configuration that had not been included in the training data. Excellent DNN agreement with conditional DNS statistics is observed, and integrated gradients are computed to identify the most sensitive input variables. The similarity of the marginal PDFs of the most informative input variables and outputs across configurations are quantified via the Wasserstein metric, demonstrating that data-based models may successfully generalize to unseen parametric conditions so long as the most informative input variables share similar distributions across training and testing datasets.

Data-based modelling↗

Deep-Learning-Based Koopman Modeling for Online Control Synthesis of Nonlinear Power System Transient Dynamics

Power system stability and control have become more challenging due to the increasing uncertainty associated with renewable generation. Here, the performance of conventional control is highly driven by the physics-based offline-developed dynamic models that can deviate from the actual system characteristics under different operating conditions and/or configurations. Data-driven approaches based on online measurements can be a better solution to addressing these issues by capturing real-time operation conditions. This article describes a novel fully data-driven probabilistic framework to derive a linear representation of postcontingency grid dynamics and online prescribe control based on the derived model to enhance transient stability. The complex nonlinear power system dynamics is approximated by a linear model by using multiple neural network modules that infer distributions of the observations and introducing a Koopman layer to sample possible Koopman linear models from the inferred distributions. The trained model features linearity that can be easily incorporated into the existing linear control design paradigm and ease the controller design process. The effectiveness of Koopman-based control designs is validated through comparative case studies, which demonstrate increased prediction accuracy and control performance when applied to a power system with heterogeneous generator dynamics.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Recent Progress in Developing Nb 3 Sn Conductors With Increased Specific Heat

A possible approach to reducing Nb 3 Sn magnet training is to increase the energy margin of Nb 3 Sn conductors by enhancing their specific heat (C p ). For this study, we have been developing Nb 3 Sn conductors with increased C p by incorporating substances with high C p at 2-10 K based on a conductor design that is compatible with standard Nb 3 Sn strand production. In the past couple of years our efforts have been mainly focused on improving strand design (e.g., position of high- C p filaments, thickness of the Cu tube for the high-C p filaments, ratio of the Cu powder to the high-C p substance, filament spacing, etc.) in order to obtain good strand drawability and to reduce degradation after rolling, which is needed for production of Rutherford cables. We also tried a new high-C p substance, Gd 2 O 2 S, and verified that it has much higher C p over the whole magnetic field range than the Gd 2 O 3 we used before. With that development work we can now produce high-C p strands with good drawability and low levels of degradation after rolling. This paper reports our findings and the current status of the development of high-C p Nb 3 Sn conductors.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Potential of the Julia Programming Language for High Energy Physics Computing

Research in high energy physics (HEP) requires huge amounts of computing and storage, putting strong constraints on the code speed and resource usage. To meet these requirements, a compiled high-performance language is typically used; while for physicists, who focus on the application when developing the code, better research productivity pleads for a high-level programming language. A popular approach consists of combining Python, used for the high-level interface, and C++, used for the computing intensive part of the code. A more convenient and efficient approach would be to use a language that provides both high-level programming and high-performance. The Julia programming language, developed at MIT especially to allow the use of a single language in research activities, has followed this path. In this paper the applicability of using the Julia language for HEP research is explored, covering the different aspects that are important for HEP code development: runtime performance, handling of large projects, interface with legacy code, distributed computing, training, and ease of programming. The study shows that the HEP community would benefit from a large scale adoption of this programming language. The HEP-specific foundation libraries that would need to be consolidated are identified.

97 MATHEMATICS AND COMPUTING↗

Model-parallel Fourier neural operators as learned surrogates for large-scale parametric PDEs

Fourier neural operators (FNOs) are a recently introduced neural network architecture for learning solution operators of partial differential equations (PDEs), which have been shown to perform significantly better than comparable deep learning approaches. Once trained, FNOs can achieve speed-ups of multiple orders of magnitude over conventional numerical PDE solvers. However, due to the high dimensionality of their input data and network weights, FNOs have so far only been applied to two-dimensional or small three-dimensional problems. To remove this limited problem-size barrier, we propose a model-parallel version of FNOs based on domain-decomposition of both the input data and network weights. Here, we demonstrate that our model-parallel FNO is able to predict time-varying PDE solutions of over 2.6 billion variables on Perlmutter using up to 512 A100 GPUs and show an example of training a distributed FNO on the Azure cloud for simulating multiphase CO 2 dynamics in the Earth’s subsurface.

58 GEOSCIENCES↗

D2NO: Efficient handling of heterogeneous input function spaces with distributed deep neural operators

Neural operators have been applied in various scientific fields, such as solving parametric partial differential equations, dynamical systems with control, and inverse problems. However, challenges arise when dealing with input functions that exhibit heterogeneous properties, requiring multiple sensors to handle functions with minimal regularity. To address this issue, discretization-invariant neural operators have been used, allowing the sampling of diverse input functions with different sensor locations. However, existing frameworks still require an equal number of sensors for all functions. We propose a novel distributed approach to further relax the discretization requirements and solve the heterogeneous dataset challenges. Our method involves partitioning the input function space and processing individual input functions using independent and separate neural networks. A centralized neural network is used to handle shared information across all output functions. This distributed methodology reduces the number of gradient descent back-propagation steps, improving efficiency while maintaining accuracy. Here, we demonstrate that the corresponding neural network is a universal approximator of continuous nonlinear operators and present three numerical examples to validate its performance.

97 MATHEMATICS AND COMPUTING↗