Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “deep learning (DL)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗

TNet: A Model-Constrained Tikhonov Network Approach for Inverse Problems

Deep learning (DL), in particular deep neural networks, by default is purely data-driven and in general does not require physics. This is the strength of DL but also one of its key limitations when applied to science and engineering problems in which underlying physical properties—such as stability, conservation, and positivity—and accuracy are required. DL methods in their original forms are often not capable of respecting the underlying mathematical models or achieving desired accuracy even in big-data regimes. On the other hand, many data-driven science and engineering problems, such as inverse problems, typically have limited experimental or observational data, and DL would overfit the data in this case. Leveraging information encoded in the underlying mathematical models, we argue, not only compensates for missing information in low data regimes but also provides opportunities to equip DL methods with the underlying physics, hence promoting better generalization. This paper develops a model-constrained DL approach and its variant TNet—a Tikhonov neural network—which are capable of learning not only information hidden in the training data but also in the underlying mathematical models to solve inverse problems governed by partial differential equations in low data regimes. We provide the constructions and some theoretical results for the proposed approaches for both linear and nonlinear inverse problems. Since TNet is designed to learn inverse solutions with Tikhonov regularization, it is interpretable: in fact it recovers Tikhonov solutions for linear cases while potentially approximating Tikhonov solutions for nonlinear inverse problems. We also prove that data randomization can enhance not only the smoothness of the networks but also their generalizations. Comprehensive numerical results confirm the theoretical findings and show that with even as little as 1 training data sample for one-dimensional (1D) deconvolution, 5 for an inverse 2D heat conductivity problem, 100 for inverse initial conditions for a time-dependent 2D Burgers’s equation, and 50 for inverse initial conditions for 2D Navier–Stokes equations, TNet solutions can be as accurate as Tikhonov solutions while being several orders of magnitude faster. Furthermore, this is possible owing to the model-constrained term, replications, and randomization.

97 MATHEMATICS AND COMPUTING↗

Deep inverse photonic design: A tutorial

Deep learning (DL) has been used to design deep neural networks (DNNs) which have recently been applied to solving inverse problems in artificial electromagnetic materials (AEMs). Although inverse problems are often ill-posed, and therefore are difficult to solve, deep inverse models (DIMs) have achieved impressive results often surpassing capabilities possible with other approaches. We overview the process of deep inverse learning applied to AEM problems, including the building of data sets, design of a forward model, and comparison of inverse approaches including limitations. We conclude by detailing some important outstanding issues of deep inverse design of AEMs, and present an outlook of this exciting field.

42 ENGINEERING↗

The Application of Artificial Intelligence and Deep Learning to Visually Identify Micrometeoroid and Orbital Debris Impacts

Recent advancements in Artificial Intelligence (AI) have made Machine Learning (ML) techniques readily available for practical applications while using a fraction of time that was previously required. In particular, the use of Deep Learning (DL) algorithms has advanced the field of image and pattern recognition. With the use of Deep Learning algorithms, Micrometeoroid and Orbital Debris(MMOD) penetrations can be identified with high accuracy and give possibilities to new understandings of hypervelocity impacts.

Deep Learning↗

Deep Learning Based Superconducting Radio-Frequency Cavity Fault Classification at Jefferson Laboratory

This work investigates the efficacy of deep learning (DL) for classifying C100 superconducting radio-frequency (SRF) cavity faults in the Continuous Electron Beam Accelerator Facility (CEBAF) at Jefferson Lab. CEBAF is a large, high-power continuous wave recirculating linac that utilizes 418 SRF cavities to accelerate electrons up to 12 GeV. Recent upgrades to CEBAF include installation of 11 new cryomodules (88 cavities) equipped with a low-level RF system that records RF time-series data from each cavity at the onset of an RF failure. Typically, subject matter experts (SME) analyze this data to determine the fault type and identify the cavity of origin. This information is subsequently utilized to identify failure trends and to implement corrective measures on the offending cavity. Manual inspection of large-scale, time-series data, generated by frequent system failures is tedious and time consuming, and thereby motivates the use of machine learning (ML) to automate the task. This study extends work on a previously developed system based on traditional ML methods (Tennant and Carpenter and Powers and Shabalina Solopova and Vidyaratne and Iftekharuddin, Phys. Rev. Accel. Beams, 2020, 23, 114601), and investigates the effectiveness of deep learning approaches. The transition to a DL model is driven by the goal of developing a system with sufficiently fast inference that it could be used to predict a fault event and take actionable information before the onset (on the order of a few hundred milliseconds). Because features are learned, rather than explicitly computed, DL offers a potential advantage over traditional ML. Specifically, two seminal DL architecture types are explored: deep recurrent neural networks (RNN) and deep convolutional neural networks (CNN). We provide a detailed analysis on the performance of individual models using an RF waveform dataset built from past operational runs of CEBAF. In particular, the performance of RNN models incorporating long short-term memory (LSTM) are analyzed along with the CNN performance. Furthermore, comparing these DL models with a state-of-the-art fault ML model shows that DL architectures obtain similar performance for cavity identification, do not perform quite as well for fault classification, but provide an advantage in inference speed.

97 MATHEMATICS AND COMPUTING↗

Improving and Assessing the Quality of Uncertainty Quantification in Deep Learning

Deep learning (DL) models have enjoyed increased attention in recent years because of their powerful predictive capabilities. While many successes have been achieved, standard deep learning methods suffer from a lack of uncertainty quantification (UQ). While the development of methods for producing UQ from DL models is an active area of current research, little attention has been given to the quality of the UQ produced by such methods. In order to deploy DL models to high-consequence applications, high-quality UQ is necessary. This report details the research and development conducted as part of a Laboratory Directed Research and Development (LDRD) project at Sandia National Laboratories. The focus of this project is to develop a framework of methods and metrics for the principled assessment of UQ quality in DL models. This report presents an overview of UQ quality assessment in traditional statistical modeling and describes why this approach is difficult to apply in DL contexts. An assessment on relatively simple simulated data is presented to demonstrate that UQ quality can differ greatly between DL models trained on the same data. A method for simulating image data that can then be used for UQ quality assessment is described. A general method for simulating realistic data for the purpose of assessing a model’s UQ quality is also presented. A Bayesian uncertainty framework for understanding uncertainty and existing metrics is described. Research that came out of collaborations with two university partners are discussed along with a software toolkit that is currently being developed to implement the UQ quality assessment framework as well as serve as a general guide to incorporating UQ into DL applications.

97 MATHEMATICS AND COMPUTING↗

Deep learning inversion of gravity data for detection of CO 2 plumes in overlying aquifers

In this work, we developed an effective U-Net based deep learning (DL) model for inversion of surface gravity data on a rectangular grid to predict 2-D high-resolution subsurface CO 2 distribution along a vertical cross-section due to CO 2 leakage through a wellbore within a deep CO 2 storage reservoir. We used synthetic data to model two types of CO 2 leakage scenarios: one CO 2 plume in a shallow aquifer (single plume case), and two plumes present at different depths (double plume case). The 3-D synthetic plume samples were created by sampling among predetermined CO 2 plume depths, saturations, and volumes. The corresponding surface gravity data on a rectangular grid were generated by a 3-D forward model. The U-Net model detected 72% of single-plume samples, and one or both plumes in 75% of double-plume samples. Most of the undetected single plumes have small gravity field strengths below the typical noise level of 5 μGal. This model generated reproducible, reliable predictions with acceptable errors and demonstrated improved spatial resolution over the conventional least-squares inversion. In contrast to the conventional least-squares inversion, which often overestimates the size of its target and underestimates its density, this U-Net model accurately delineated the boundary of a target. Furthermore, this DL inversion detected deep, small, or low saturation CO 2 plumes that are often more difficult to resolve with conventional gravity inversion methods. We note the limitations of this feasibility study, including the use of synthetic data with regular CO 2 plume shapes, and the prediction of a 2-D plume cross-section rather than the full 3-D plume, as well, we recognize the lower detection fraction for double-plume scenarios. Nevertheless, this study demonstrates that DL gravity inversion is a promising and potentially superior method to conventional least-squares inversion. Our U-Net based deep learning inversion approach may be adapted for inversion of other types of geophysical data. DL inversion can facilitate near real-time monitoring of geologic carbon sequestration to provide site operators with prompt information about subsurface CO 2 distribution for risk management and mitigation.

58 GEOSCIENCES↗

Recent advances and applications of deep learning methods in materials science

Deep learning (DL) is one of the fastest-growing topics in materials data science, with rapidly emerging applications spanning atomistic, image-based, spectral, and textual data modalities. DL allows analysis of unstructured data and automated identification of features. The recent development of large materials databases has fueled the application of DL methods in atomistic prediction in particular. In contrast, advances in image and spectral data have largely leveraged synthetic data enabled by high-quality forward models as well as by generative unsupervised DL methods. In this article, we present a high-level overview of deep learning methods followed by a detailed discussion of recent developments of deep learning in atomistic simulation, materials imaging, spectral analysis, and natural language processing. For each modality we discuss applications involving both theoretical and experimental data, typical modeling approaches with their strengths and limitations, and relevant publicly available software and datasets. We conclude the review with a discussion of recent cross-cutting work related to uncertainty quantification in this field and a brief perspective on limitations, challenges, and potential growth areas for DL methods in materials science.

36 MATERIALS SCIENCE↗

Analyzing inference workloads for spatiotemporal modeling

Ensuring power grid resiliency, forecasting climate conditions, and optimization of transportation infrastructure are some of the many application areas where data is collected in both space and time. Spatiotemporal modeling is about modeling those patterns for forecasting future trends and carrying out critical decision-making by leveraging machine learning/deep learning. Once trained offline, field deployment of trained models for near real-time inference could be challenging because performance can vary significantly depending on the environment, available compute resources and tolerance to ambiguity in results. Users deploying spatiotemporal models for solving complex problems can benefit from analytical studies considering a plethora of system adaptations to understand the associated performance-quality trade-offs. To facilitate the co-design of next-generation hardware architectures for field deployment of trained models, it is critical to characterize the workloads of these deep learning (DL) applications during inference and assess their computational patterns at different levels of the execution stack. In this paper, we develop several variants of deep learning applications that use spatiotemporal data from dynamical systems. We study the associated computational patterns for inference workloads at different levels, considering relevant models (Long short-term Memory, Convolutional Neural Network and Spatio-Temporal Graph Convolution Network), DL frameworks (Tensorflow and PyTorch), precision (FP16, FP32, AMP, INT16 and INT8), inference runtime (ONNX and AI Template), post-training quantization (TensorRT) and platforms (Nvidia DGX A100 and Sambanova SN10 RDU). Overall, our findings indicate that although there is potential in mixed-precision models and post-training quantization for spatiotemporal modeling, extracting efficiency from contemporary GPU systems might be challenging. Instead, co-designing custom accelerators by leveraging optimized High Level Synthesis frameworks (such as SODA High-Level Synthesizer for customized FPGA/ASIC targets) can make workload-specific adjustments to enhance the efficiency.

97 MATHEMATICS AND COMPUTING↗

Surrogate Modeling of Nonlinear Dynamic Systems: A Comparative Study

Surrogate models play a vital role in overcoming the computational challenge in designing and analyzing nonlinear dynamic systems, especially in the presence of uncertainty. This paper presents a comparative study of different surrogate modeling techniques for nonlinear dynamic systems. Four surrogate modeling methods, namely, Gaussian process (GP) regression, a long short-term memory (LSTM) network, a convolutional neural network (CNN) with LSTM (CNN-LSTM), and a CNN with bidirectional LSTM (CNN-BLSTM), are studied and compared. All these model types can predict the future behavior of dynamic systems over long periods based on training data from relatively short periods. The multi-dimensional inputs of surrogate models are organized in a nonlinear autoregressive exogenous model (NARX) scheme to enable recursive prediction over long periods, where current predictions replace inputs from the previous time window. Three numerical examples, including one mathematical example and two nonlinear engineering analysis models, are used to compare the performance of the four surrogate modeling techniques. The results show that the GP-NARX surrogate model tends to have more stable performance than the other three deep learning (DL)-based methods for the three particular examples studied. The tuning effort of GP-NARX is also much lower than its deep learning-based counterparts.

42 ENGINEERING↗

Data optimization for large batch distributed training of deep neural networks

Distributed training in deep learning (DL) is common practice as data and models grow. The current practice for distributed training of deep neural networks faces the challenges of communication bottlenecks when operating at scale, and model accuracy deterioration with an increase in global batch size. Present solutions focus on improving message exchange efficiency as well as implementing techniques to tweak batch sizes and models in the training process. The loss of training accuracy typically happens because the loss function gets trapped in a local minima. We observe that the loss landscape minimization is shaped by both the model and training data and propose a data optimization approach that utilizes machine learning to implicitly smooth out the loss landscape resulting in fewer local minima. Our approach filters out data points which are less important to feature learning, enabling us to speed up the training of models on larger batch sizes to improved accuracy.

Gahlot, Shubhankar↗

A Public Data Set of Auto-Generated Geotagged PV Site Equipment, Generated via Deep Learning

In this research, we present a data set over 100 photovoltaic (PV) sites in TX, which have been automatically geotagged via a fully autonomous deep learning (DL) pipeline. Specifically, locations of inverters, tracker/fixed tilt rows, batteries, and substations are labeled algorithmically. To ensure high data quality, all systems have been reviewed manually and any deep learning errors have been corrected. This public data set, as well as the open-sourced pipeline used to generate it, is valuable for site planning, modelling, and insurance purposes. Given time and resources, we hope to extend the data set to additional states/regions in the US.

14 SOLAR ENERGY↗

Improving deep learning model performance under parametric constraints for materials informatics applications

Abstract Modern machine learning (ML) and deep learning (DL) techniques using high-dimensional data representations have helped accelerate the materials discovery process by efficiently detecting hidden patterns in existing datasets and linking input representations to output properties for a better understanding of the scientific phenomenon. While a deep neural network comprised of fully connected layers has been widely used for materials property prediction, simply creating a deeper model with a large number of layers often faces with vanishing gradient problem, causing a degradation in the performance, thereby limiting usage. In this paper, we study and propose architectural principles to address the question of improving the performance of model training and inference under fixed parametric constraints. Here, we present a general deep-learning framework based on branched residual learning (BRNet) with fully connected layers that can work with any numerical vector-based representation as input to build accurate models to predict materials properties. We perform model training for materials properties using numerical vectors representing different composition-based attributes of the respective materials and compare the performance of the proposed models against traditional ML and existing DL architectures. We find that the proposed models are significantly more accurate than the ML/DL models for all data sizes by using different composition-based attributes as input. Further, branched learning requires fewer parameters and results in faster model training due to better convergence during the training phase than existing neural networks, thereby efficiently building accurate models for predicting materials properties.

36 MATERIALS SCIENCE↗

Knowledge-informed deep learning for hydrological model calibration: an application to Coal Creek Watershed in Colorado

Abstract. Deep learning (DL)-assisted inverse mapping has shown promise in hydrological model calibration by directly estimating parameters from observations. However, the increasing computational demand for running the state-of-the-art hydrological model limits sufficient ensemble runs for its calibration. In this work, we present a novel knowledge-informed deep learning method that can efficiently conduct the calibration using a few hundred realizations. The method involves two steps. First, we determine decisive model parameters from a complete parameter set based on the mutual information (MI) between model responses and each parameter computed by a limited number of realizations (∼50). Second, we perform more ensemble runs (e.g., several hundred) to generate the training sets for the inverse mapping, which selects informative model responses for estimating each parameter using MI-based parameter sensitivity. We applied this new DL-based method to calibrate a process-based integrated hydrological model, the Advanced Terrestrial Simulator (ATS), at Coal Creek Watershed, CO. The calibration is performed against observed stream discharge (Q) and remotely sensed evapotranspiration (ET) from the water year 2017 to 2019. Preliminary MI analysis on 50 realizations resulted in a down-selection of 7 out of 14 ATS model parameters. Then, we performed a complete MI analysis on 396 realizations and constructed the inverse mapping from informative responses to each of the selected parameters using a deep neural network. Compared with calibration using observations covering all time steps, the new inverse mapping improves parameter estimations, thus enhancing the performance of ATS forward model runs. The Nash–Sutcliffe efficiency (NSE) of streamflow predictions increases from 0.53 to 0.8 when calibrating against Q alone. Using ET observations, on the other hand, does not show much improvement on the performance of ATS modeling mainly due to both the uncertainty of the remotely sensed product and the insufficient coverage of the model ET ensemble in capturing the observation. By using observed Q only, we further performed a multiyear analysis and show that Q is best simulated (NSE > 0.8) by including in the calibration the dry-year flow dynamics that show more sensitivity to subsurface characteristics than the other wet years. Moreover, when continuing the forward runs till the end of 2021, the calibrated models show similar simulation performances during this evaluation period as the calibration period, demonstrating the ability of the estimated parameters in capturing climate sensitivity. Our success highlights the importance of leveraging data-driven knowledge in DL-assisted hydrological model calibration.

54 ENVIRONMENTAL SCIENCES↗

Editorial: Functional microcircuits in the brain and in artificial intelligent systems

Fundamental principles underlying higher-order cognitive functions remain elusive, but recent breakthroughs in neurophysiology and deep learning offer new perspectives. First, experimental studies have uncovered neural circuit motifs consisting of various neuron types; see Brain Initiative Cell Census Network (https://www.nature.com/collections/cicghheddj). For example, inhibitory neuron types expressing exclusive genes have specific targets and distinct functions (Pfeffer et al., 2013). Furthermore, diverse neuron types in cortex and their connectomes were identified in cortical columns (Jiang et al., 2015); see also Barth et al. (2016) for a debate on neuron types. Second, artificial neural networks were originally inspired by structures of the brain (McCulloch and Pitts, 1943) and could be trained to perform complex functions similar to human perception/cognition by deep learning (DL) (Lecun et al., 2015).

59 BASIC BIOLOGICAL SCIENCES↗

Decentralized Distributed Proximal Policy Optimization (DD-PPO) for High Performance Computing Scheduling on Multi-User Systems

Resource allocation in High Performance Computing (HPC) environments presents a complex and multifaceted challenge for job scheduling algorithms. Beyond the efficient allocation of system resources, schedulers must account for and optimize multiple performance metrics, including job wait time and system throughput. Traditional heuristic-based scheduling algorithms increasingly struggle and lack the efficiency needed to meet the demands and address the complexity and scale of modern HPC systems. Consequently, recent research efforts have focused on leveraging advancements in Artificial Intelligence (AI) and Deep Learning (DL), particularly Reinforcement Learning (RL), to develop more adaptable and intelligent scheduling strategies. Previous RL-based scheduling approaches have explored a range of algorithms, from Deep Q-Networks (DQN) to Proximal Policy Optimization (PPO), and more recently, hybrid methods that integrate Graph Neural Networks (GNNs) with RL techniques. However, a common limitation across these methods is their reliance on relatively small datasets, with few methods being evaluated using large-scale, multi-million-job trace datasets representative of real-world HPC workloads. Moreover, existing RL schedulers face scalability issues due to centralized policy updates, which hinder training efficiency and performance when applied to large datasets. This study introduces a novel RL-based scheduler utilizing Decentralized Distributed Proximal Policy Optimization (DD-PPO) algorithm, which supports large-scale distributed training across multiple workers without requiring parameter synchronization at every step. By eliminating reliance on centralized updates to a shared policy, the DD-PPO scheduler enhances scalability, training efficiency, and sample utilization. Experimental validation using a large real-world dataset containing over 11.5 million job traces collected from petascale HPC systems over six years assesses the influence of dataset scale on training effectiveness and compares DD-PPO performance to traditional and advanced scheduling approaches. The experimental results demonstrate improved scheduling performance in comparison to both heuristic-based schedulers and existing RL-based scheduling algorithms.

AI↗

Improving deep learning performance for predicting large-scale geological ${{CO}_{2}}$ sequestration modeling through feature coarsening

Physics-based reservoir simulation for fluid flow in porous media is a numerical simulation method to predict the temporal-spatial patterns of state variables (e.g. pressure p) in porous media, and usually requires prohibitively high computational expense due to its non-linearity and the large number of degrees of freedom (DoF). This work describes a deep learning (DL) workflow to predict the pressure evolution as fluid flows in large-scale 3-dimensional(3D) heterogeneous porous media. In particular, we develop an efficient feature coarsening technique to extract the most representative information and perform the training and prediction of DL at the coarse scale, and further recover the resolution at the fine scale by spatial interpolation. We validate the DL approach to predict pressure field against physics-based simulation data for a field-scale 3D geologic CO 2 sequestration reservoir model. We evaluate the impact of feature coarsening on DL performance, and observe that the feature coarsening not only decreases the training time by >74% and reduces the memory consumption by >75%, but also maintains temporal error 0.63% on average. Besides, the DL workflow provides predictive efficiency with 1406 times speedup compared to physics-based numerical simulation. The key findings from this research significantly improve the training and prediction efficiency of deep learning model to deal with large-scale heterogeneous reservoir models, and thus it can also be further applied to accelerate workflows of history matching and reservoir optimization for close-loop reservoir management.

58 GEOSCIENCES↗

Simultaneously improving accuracy and computational cost under parametric constraints in materials property prediction tasks

Abstract Modern data mining techniques using machine learning (ML) and deep learning (DL) algorithms have been shown to excel in the regression-based task of materials property prediction using various materials representations. In an attempt to improve the predictive performance of the deep neural network model, researchers have tried to add more layers as well as develop new architectural components to create sophisticated and deep neural network models that can aid in the training process and improve the predictive ability of the final model. However, usually, these modifications require a lot of computational resources, thereby further increasing the already large model training time, which is often not feasible, thereby limiting usage for most researchers. In this paper, we study and propose a deep neural network framework for regression-based problems comprising of fully connected layers that can work with any numerical vector-based materials representations as model input. We present a novel deep regression neural network, iBRNet, with branched skip connections and multiple schedulers, which can reduce the number of parameters used to construct the model, improve the accuracy, and decrease the training time of the predictive model. We perform the model training using composition-based numerical vectors representing the elemental fractions of the respective materials and compare their performance against other traditional ML and several known DL architectures. Using multiple datasets with varying data sizes for training and testing, We show that the proposed iBRNet models outperform the state-of-the-art ML and DL models for all data sizes. We also show that the branched structure and usage of multiple schedulers lead to fewer parameters and faster model training time with better convergence than other neural networks. Scientific contribution: The combination of multiple callback functions in deep neural networks minimizes training time and maximizes accuracy in a controlled computational environment with parametric constraints for the task of materials property prediction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗