Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scalable deep learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

232 records · Page 13

Streamlining Ocean Dynamics Modeling with Fourier Neural Operators: A Multiobjective Hyperparameter and Architecture Optimization Approach

Training an effective deep learning model to learn ocean processes involves careful choices of various hyperparameters. We leverage DeepHyper’s advanced search algorithms for multiobjective optimization, streamlining the development of neural networks tailored for ocean modeling. The focus is on optimizing Fourier neural operators (FNOs), a data-driven model capable of simulating complex ocean behaviors. Selecting the correct model and tuning the hyperparameters are challenging tasks, requiring much effort to ensure model accuracy. DeepHyper allows efficient exploration of hyperparameters associated with data preprocessing, FNO architecture-related hyperparameters, and various model training strategies. We aim to obtain an optimal set of hyperparameters leading to the most performant model. Moreover, on top of the commonly used mean squared error for model training, we propose adopting the negative anomaly correlation coefficient as the additional loss term to improve model performance and investigate the potential trade-off between the two terms. The numerical experiments show that the optimal set of hyperparameters enhanced model performance in single timestepping forecasting and greatly exceeded the baseline configuration in the autoregressive rollout for long-horizon forecasting up to 30 days. Utilizing DeepHyper, we demonstrate an approach to enhance the use of FNO in ocean dynamics forecasting, offering a scalable solution with improved precision.

97 MATHEMATICS AND COMPUTING↗

FitCache: A Transparent Drop-In Framework for Multi-Tier Caching to Accelerate Distributed Deep Learning Workloads

Training in Deep learning (DL) remains highly compute- and data-intensive, with I/O becoming a critical bottleneck as models and datasets scale. Recent studies report that data loading can dominate training time, especially on large-scale HPC systems with shared parallel file systems (PFS). Existing caching approaches either rely on single-tier designs or require intrusive modifications to training pipelines, limiting their portability and effectiveness. In this work, we present FitCache, a transparent drop-in framework for multi-tier caching to accelerate distributed DL training by coordinating fast local memory (e.g., DRAM, Persistent Memory (PMem)) and NVMe as hierarchical caches atop PFS. Our design adapts to hardware diversity, i.e., if NVMe is missing, memory transparently acts as a caching tier, ensuring stable performance. FitCache transparently intercepts I/O requests and issues concurrent fetches across all tiers, returning data from the fastest responder without centralized metadata or static redirection paths. FitCache adapts to dynamic workloads and heterogeneous clusters while maintaining POSIX compatibility. Experiments on Frontier (2048 GPUs) and smaller research clusters show that FitCache reduces training time by up to 40% and per-batch I/O latency by up to 71.6% compared to Lustre Orion PFS, offering a drop-in solution for scalable DL training.

Hu, Guangxing [ORNL] (ORCID:0009000283203614)↗

Towards intelligent emergency control for large-scale power systems: Convergence of learning, physics, computing and control

Here, this paper has delved into the pressing need for intelligent emergency control in large-scale power systems, which are experiencing significant transformations and are operating closer to their limits with more uncertainties. Learning-based control methods are promising and have shown effectiveness for intelligent power system control. However, when they are applied to large-scale power systems, there are multifaceted challenges such as scalability, adaptiveness, and security posed by the complex power system landscape, which demand comprehensive solutions. The paper first proposes and instantiates a convergence framework for integrating power systems physics, machine learning, advanced computing, and grid control to realize intelligent grid control at a large scale. Our developed methods and platform based on the convergence framework have been applied to a large (more than 3000 buses) Texas power system, and tested with 56 000 scenarios. Our work achieved a 26% reduction in load shedding on average and outperformed existing rule-based control in 99.7% of the test scenarios. The results demonstrated the potential of the proposed convergence framework and DRL-based intelligent control for the future grid.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Catalyzing deep decarbonization with federated battery diagnosis and prognosis for better data management in energy storage systems

Industrial data analytics methods play a central role in improving energy storage performance and efficiency, impacting the future of electrified transportation and renewable electricity generation. However, significant challenges hinder the large-scale deployment of batteries. Conventional methods rely on centralized collection and processing of fleet-level data, leading to database size issues and privacy concerns due to potential data breaches. To enable scalable deployment of battery management systems, this article proposes a federated battery diagnosis and prognosis model, which distributes the processing of battery standard current-voltage-time-usage data in a privacy-preserving manner. Instead of transferring the raw data, this approach communicates only the locally processed parameters, thus reducing communication load and preserving data confidentiality. The federated model offers a paradigm shift in battery health management through privacy-preserving distributed methods for battery data processing and lifetime prediction, ensuring the reliable and sustainable deployment of lithium-ion batteries in a rapidly evolving world.

asset health management↗

Evaluating probabilistic deep learning methods for uncertainty quantification of temperature downscaling

Deep learning (DL) has emerged as a promising tool for downscaling coarse-resolution climate data to high-resolution outputs, enabling improved regional climate predictions. A critical aspect of DL-based downscaling is the incorporation of uncertainty quantification (UQ), which enhances the interpretability and reliability of predictions—key factors for climate risk assessment and decision-making. This study develops a DL model to downscale 2 m temperature across the contiguous United States using reanalysis datasets. We systematically evaluate three epistemic UQ methods—deep ensembles (DEns), Monte Carlo dropout (MCD), and Flipout—based on their probabilistic accuracy, downscaling performance, sensitivity to geographical features, and computational efficiency. Results indicate that MCD generally outperforms Flipout and DEns in terms of calibration and downscaling accuracy. However, DEns demonstrate lower calibration errors in coastal regions, indicating its higher confidence within these areas. Flipout, in contrast, is more sensitive to elevation gradients and exhibits higher calibration errors in mountainous regions. Hence, the choice of UQ method for this task depends on the specific requirements of the application. For applications that prioritize overall calibration, downscaling accuracy, and computational efficiency, MCD is a strong candidate. These findings highlight the importance of selecting UQ methods based on application-specific requirements, such as geographical context and computational constraints. By addressing the trade-offs between UQ methods, this study provides actionable insights for improving the reliability, scalability, and utility of DL-based downscaling in climate science.

Environmental sciences↗

Conditioned quantum-assisted deep generative surrogate for particle-calorimeter interactions

Particle collisions at accelerators like the Large Hadron Collider (LHC), recorded by experiments such as ATLAS and CMS, enable precise standard model measurements and searches for new phenomena. Simulating these collisions significantly influences experiment design and analysis but incurs immense computational costs, projected at millions of CPU-years annually during the high luminosity LHC (HL-LHC) phase. Currently, simulating a single event with Geant4 consumes around 1000 CPU seconds, with calorimeter simulations especially demanding. To address this, we propose a conditioned quantum-assisted generative model, integrating a conditioned variational autoencoder (VAE) and a conditioned restricted Boltzmann machine (RBM). Our RBM architecture is tailored for D-Wave’s Pegasus-structured advantage quantum annealer for sampling, leveraging the flux bias for conditioning. This approach combines classical RBMs as universal approximators for discrete distributions with quantum annealing’s speed and scalability. We also introduce an adaptive method for efficiently estimating effective inverse temperature, and validate our framework on Dataset 2 of CaloChallenge.

97 MATHEMATICS AND COMPUTING↗

Interpretable AI forecasting for numerical relativity waveforms of quasicircular, spinning, nonprecessing binary black hole mergers

We present a deep-learning artificial intelligence model (AI) that is capable of learning and forecasting the late-inspiral, merger and ringdown of numerical relativity waveforms that describe quasicircular, spinning, nonprecessing binary black hole mergers. We used the NRHybSur3dq8 surrogate model to produce train, validation and test sets of ℓ = |m| = 2 waveforms that cover the parameter space of binary black hole mergers with mass ratios q ≤ 8 and individual spins |s$^{z}_ {(1,2)}$| ≤ 0.8. These waveforms cover the time range t ∊ [-5000 M, 130 M], where t = 0M marks the merger event, defined as the maximum value of the waveform amplitude. We harnessed the ThetaGPU supercomputer at the Argonne Leadership Computing Facility to train our AI model using a training set of 1.5 million waveforms. We used 16 NVIDIA DGX A100 nodes, each consisting of 8 NVIDIA A100 Tensor Core GPUs and 2 AMD Rome CPUs, to fully train our model within 3.5 h. Our findings show that artificial intelligence can accurately forecast the dynamical evolution of numerical relativity waveforms in the time range t ∊ [-100 M, 130 M]. Sampling a test set of 190,000 waveforms, we find that the average overlap between target and predicted waveforms is ≳99% over the entire parameter space under consideration. We also combined scientific visualization and accelerated computing to identify what components of our model take in knowledge from the early and late-time waveform evolution to accurately forecast the latter part of numerical relativity waveforms. This work aims to accelerate the creation of scalable, computationally efficient and interpretable artificial intelligence models for gravitational wave astrophysics

79 ASTRONOMY AND ASTROPHYSICS↗

Physics-informed latent neural operator for real-time predictions of time-dependent parametric PDEs

Deep operator network (DeepONet) has shown significant promise as surrogate models for systems governed by partial differential equations (PDEs), enabling accurate mappings between infinite-dimensional function spaces. However, when applied to systems with high-dimensional input-output mappings arising from large numbers of spatial and temporal collocation points, these models often require heavily overparameterized networks, leading to long training times. Latent DeepONet addresses some of these challenges by introducing a two-step approach: first learning a reduced latent space using a separate model, followed by operator learning within this latent space. While efficient, this method is inherently data-driven and lacks mechanisms for incorporating physical laws, limiting its robustness and generalizability in data-scarce settings. Here, in this work, we propose PI-Latent-NO, a physics-informed latent neural operator framework that integrates governing physics directly into the learning process. Our architecture features two coupled DeepONets trained end-to-end: a Latent-DeepONet that learns a low-dimensional representation of the solution, and a Reconstruction-DeepONet that maps this latent representation back to the physical space. By embedding PDE constraints into the training via automatic differentiation, our method eliminates the need for labeled training data and ensures physics-consistent predictions. The proposed framework is both memory and compute-efficient, exhibiting near-constant scaling with problem size and demonstrating significant speedups over traditional physics-informed operator models. We validate our approach on a range of parametric PDEs, showcasing its accuracy, scalability, and suitability for real-time prediction in complex physical systems.

Latent representations↗

Learning high-dimensional parametric maps via reduced basis adaptive residual networks

We propose a scalable framework for the learning of high-dimensional parametric maps via adaptively constructed residual network (ResNet) maps between reduced bases of the inputs and outputs. When just few training data are available, it is beneficial to have a compact parametrization in order to ameliorate the ill-posedness of the neural network training problem. By linearly restricting high-dimensional maps to informed reduced bases of the inputs, one can compress high-dimensional maps in a constructive way that can be used to detect appropriate basis ranks, equipped with rigorous error estimates. A scalable neural network learning framework is thus to learn the nonlinear compressed reduced basis mapping. Unlike the reduced basis construction, however, neural network constructions are not guaranteed to reduce errors by adding representation power, making it difficult to achieve good practical performance. Inspired by recent approximation theory that connects ResNets to sequential minimizing flows, we present an adaptive ResNet construction algorithm. This algorithm allows for depth-wise enrichment of the neural network approximation, in a manner that can achieve good practical performance by first training a shallow network and then adapting. We prove universal approximation of the associated neural network class for $L^2_v$ functions on compact sets. Our overall framework allows for constructive means to detect appropriate breadth and depth, and related compact parametrizations of neural networks, significantly reducing the need for architectural hyperparameter tuning. Numerical experiments for parametric PDE problems and a 3D CFD wing design optimization parametric map demonstrate that the proposed methodology can achieve remarkably high accuracy for limited training data, and outperformed other neural network strategies we compared against.

42 ENGINEERING↗

Machine learning predictions for local electronic properties of disordered correlated electron systems

We present a scalable machine learning (ML) model to predict local electronic properties such as on-site electron number and double occupation for disordered correlated electron systems. Our approach is based on the locality principle, or the nearsightedness nature, of many-electron systems, which means local electronic properties depend mainly on the immediate environment. A ML model is developed to encode this complex dependence of local quantities on the neighborhood. We demonstrate our approach using the square-lattice Anderson-Hubbard model, which is a paradigmatic system for studying the interplay between Mott transition and Anderson localization. We develop a lattice descriptor based on the group-theoretical method to represent the on-site random potentials within a finite region. The resultant feature variables are used as input to a multilayer fully connected neural network, which is trained from data sets of variational Monte Carlo (VMC) simulations on small systems. We show that the ML predictions agree reasonably well with the VMC data. Our work underscores the promising potential of ML methods for multiscale modeling of correlated electron systems.

36 MATERIALS SCIENCE↗

Quantum computing in power systems

Electric power systems provide the backbone of modern industrial societies. Enabling scalable grid analytics is the keystone to successfully operating large transmission and distribution systems. However, today's power systems are suffering from ever-increasing computational burdens in sustaining the expanding communities and deep integration of renewable energy resources, as well as managing huge volumes of data accordingly. These unprecedented challenges call for transformative analytics to support the resilient operations of power systems. Recently, the explosive growth of quantum computing techniques has ignited new hopes of revolutionizing power system computations. Quantum computing harnesses quantum mechanisms to solve traditionally intractable computational problems, which may lead to ultra-scalable and efficient power grid analytics. This paper reviews the newly emerging application of quantum computing techniques in power systems. We present a comprehensive overview of existing quantum-engineered power analytics from different operation perspectives, including static analysis, transient analysis, stochastic analysis, optimization, stability, and control. We thoroughly discuss the related quantum algorithms, their benefits and limitations, hardware implementations, and recommended practices. We also review the quantum networking techniques to ensure secure communication of power systems in the quantum era. Finally, we discuss challenges and future research directions. This paper will hopefully stimulate increasing attention to the development of quantum-engineered smart grids.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Machine Learning-enabled Scalable Performance Prediction of Scientific Codes

Hardware architectures become increasingly complex as the compute capabilities grow to exascale. Here, we present the Analytical Memory Model with Pipelines (AMMP) of the Performance Prediction Toolkit (PPT). PPT-AMMP takes high-level source code and hardware architecture parameters as input and predicts runtime of that code on the target hardware platform, which is defined in the input parameters. PPT-AMMP transforms the code to an (architecture-independent) intermediate representation, then (i) analyzes the basic block structure of the code, (ii) processes architecture-independent virtual memory access patterns that it uses to build memory reuse distance distribution models for each basic block, and (iii) runs detailed basic-block level simulations to determine hardware pipeline usage. PPT-AMMP uses machine learning and regression techniques to build the prediction models based on small instances of the input code, then integrates into a higher-order discrete-event simulation model of PPT running on Simian PDES engine. We validate PPT-AMMP on four standard computational physics benchmarks and present a use case of hardware parameter sensitivity analysis to identify bottleneck hardware resources on different code inputs. We further extend PPT-AMMP to predict the performance of a scientific application code, namely, the radiation transport mini-app SNAP. To this end, we analyze multi-variate regression models that accurately predict the reuse profiles and the basic block counts. We validate predicted SNAP runtimes against actual measured times.

97 MATHEMATICS AND COMPUTING↗

Scalable Risk Assessment of Rare Events in Power Systems With Uncertain Wind Generation and Loads

Risk assessment of rare events has become increasingly important in power system planning and operation with the increasing integration of renewable energy and the presence of system uncertainties. However, quantifying the risk posed by rare events via the traditional method, i.e., Monte Carlo sampling (MCS), incurs substantial computational expense stemming from the vast ensemble of power flow simulations. To accelerate the assessment, this paper proposes a Deep Neural Network (DNN)-kernelized vector-valued Gaussian Process (VVGP) approach with excellent computational efficiency while maintaining high accuracy. Consequently, serving as a surrogate model for the power flow solver, the DNN-kernelized VVGP enables significantly faster but accurate risk assessment compared to the power flow solver. The developed surrogate model evaluates low-order N - k events that contain more than 90% instances by adeptly capturing the topological features while the high-order N - k events are assessed via a power flow solver, thereby striking a balance between computational efficiency and uncertainty quantification accuracy. Moreover, the model incorporates a Support Vector Machine (SVM) classifier to resample concerning low-probability tail events to counteract the biases potentially introduced during the DNN-kernelized VVGP evaluations. Simulations conducted on the modified IEEE 24-bus, 118-bus, and European 1354-bus systems demonstrate that the proposed method maintains the accuracy benchmark set by MCS while significantly reducing computational demands in large-scale power systems as compared to other state-of-the-art methods.

17 WIND ENERGY↗

Grid-Connected Modular Soft-Switching Solid State Transformers (M-S4T)

The objective of this project is to develop and verify the concept of a flexible and modular soft-switching solid-state transformer (M-S4T) for direct grid-connected applications. The ability to directly connect power electronics converters to the medium voltage grid (4 kV – 13 kV), and to potentially replace the passive and bulky, but ubiquitous 60 hertz service transformer in the 25 kVA to 100 kVA range, with a more flexible and controllable device, has been regarded as the ‘holy grail’ in grid control. However, this has proven to be extremely difficult. This project has developed the solutions to several key challenges of the direct grid-connected power electronics and realized a 7.2 kV M-S4T prototype. First, a protection method to protect the M-S4T from the high voltages (110 kV for the 13 kV system) that occur on the grid due to transients and lightning strikes have been developed and experimentally verified. Second, the realization and the operation of the M-S4T based on high-voltage SiC devices (>3.3 kV) and a medium-frequency medium-voltage low-leakage transformer in a single-stage solid-state transformer with zero-voltage switching, low dv/dt, and low electromagnetic interference has been successfully demonstrated up to 7.5 kV peak. Third, an oil-cooling system and stable communication and distributed control system for converter module voltage sharing have been developed and experimentally verified. The developed M-S4T has realized a modular universal high-performance power conversion system. This conversion system is scalable to different voltage and power levels and adaptable to four-quadrant bidirectional operation. Moreover, the use of passive cooling techniques meets the equipment life requirements, and the lightning protection scheme fulfills the basic insulation level specifications for direct grid connection. Such power conversion system opens up near-term opportunities, including energy storage, solar PV, or electric vehicle charging with significant cost and footprint savings. In the longer term, the possibility of replacing the utility distribution transformer with an M-S4T will be transformative for future distribution grids with a compact footprint and full controllability to enable high renewable energy and storage penetration. In addition to the main project, this report expands on the Plus-Up projected including as part of the main award. This project developed and demonstrated the technology for autonomous collaborative inverters that can be connected in an ad hoc manner to the grid. The aim of the project was to: (1) evaluate the existing techniques for grid-connected inverters and find their limitations; (2) develop detailed requirements for grid-connected inverters in the modern grid with millions of active nodes; (3) design a unified control strategy that brings more autonomy and intelligence to grid-connected inverters, and addresses parts of the issues with the existing techniques. The proposed technique, called UniCon, enables inverters to 1) connect/disconnect to/from the grid in an ad hoc manner; (2) work based on local sensing. Slow communication could be used for a more optimized behavior; (3) work automatically in both grid-forming/grid-following mode; (4) handle large disturbances, e.g., big load step and fault, in an oscillation-free manner; (5) work collaboratively with other inverters in steady-state and during transients. UniCon can be implemented in the middle-level control; hence it is agnostic to the vendor and to the implementation of the inner voltage/current and protection loops. Furthermore, a new synchronization scheme, based on deep learning, was developed that can extract the grid voltage phase and amplitude in a stable manner. The method is cheap to implement can improve the dynamic performance of the grid-connected inverters during fast transients, e.g., fault. The proposed control scheme was validated by (1) MATLAB/Simulink; (2) hardware-in-the-loop results, and; (3) experimental results using three inverters that form a microgrid in a down-scaled feeder. Lastly, both the M-S4T and UniCon have achieved promising tangible paths to markets. In the case of the M-S4T, the underlying technology — the Soft Switching Solid State Transformer (S4T) developed at the Georgia Tech Center for Distributed Energy (GT-CDE) has been licensed by GridBlock from the Georgia Tech Research Corporation, and GridBlock has been working with manufacturing partner Jabil (one of the largest US-based contract manufacturers) and system integrator Power Secure (largest deployer of microgrids in the US with 4.7 GW under management), to meet the strong initial demand. Similarly, GridBlock has an exclusive license to the UniCon technology, developed under this award by GT-CDE. The UniCon provides an intermediate control layer that enables the implementation of the higher-level ‘transactive’ control commands for the system. The architecture of the system - slow communications with the cloud for system optimization and setpoints, and the use of locally measured quantities for real-time control, provide a very robust and secure way of implementing a real-time must-run grid that is also secure and stable. This is a brand-new functionality that is critical for the future grid and key to GridBlock’s business model.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Self-supervised and multi-fidelity learning for extended predictive soil spectroscopy

Infrared spectroscopy is a cost-effective, non-destructive, and environmentally benign technology that is increasingly recognized as an important solution for meeting the global demand for soil data. While both near-infrared (NIR) and mid-infrared (MIR) diffuse reflectance spectroscopy enable rapid estimation of soil properties, they present a significant trade-off: NIR offers superior scalability and lower operational costs, whereas MIR provides higher analytical fidelity by capturing fundamental molecular vibrations. In this study, we propose a self-supervised, multi-fidelity learning framework designed to bridge this gap. Our approach leverages large-scale MIR spectral libraries to learn a compact, transferable latent representation, into which NIR spectra are subsequently aligned for downstream prediction. The workflow consists of pretraining a latent model on a large MIR library, adapting the representation using a smaller paired NIR–MIR dataset, and evaluating generalization on an independent external test set. Across a range of chemical and physical soil properties, we found that MIR-derived embeddings improved prediction accuracy relative to baseline models that used raw MIR inputs. Predictions derived from the spectrum conversion (NIR to MIR) task did not match the performance of the original MIR spectra but were similar or superior to predictive performance of NIR-only models, suggesting the unified spectral latent space can effectively leverage the larger and more diverse MIR dataset for prediction of soil properties not well represented in current NIR libraries.

54 ENVIRONMENTAL SCIENCES↗

Scalable, Dual-Mode Occupancy Sensing for Commercial Venues

The ability to establish accurate occupancy is a key enabler for improved energy efficiency in human-occupied indoor spaces. The accurate counts can be used to adjust, in real-time, air exchange, heating, and cooling, to closely match the environmental needs of the occupants rather than to use a single setting based on maximum room occupancy. This project focuses on the use of overhead fisheye cameras and artificial intelligence (AI) algorithms as the technology to realize accurate people counting in large commercial spaces. We developed Computational Occupancy Sensing SYstem ("COSSY''), which supports large-scale deployment (thousands of square feet) and straightforward installation (using conventional power over ethernet), while providing high accuracy and low cost. COSSY has been validated at Boston University and by a third party, demonstrating people-counting accuracy of at least 89.5% in large spaces (2,000 sqft). A cost/benefit analysis performed for Boston commercial market indicates that a full payback of COSSY deployment costs due to energy savings would occur within 3-6 months from installation. Applications of COSSY reach far beyond saving energy. The precise people counting in real time could also be useful for emergency response (e.g., fire), in spatial analytics (e.g., office space management) and in some health applications (e.g., social distance assessment during pandemics).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗