Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “deep learning (DL)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

DEEP CELLULAR RECURRENT NEURAL ARCHITECTURE FOR EFFICIENT MULTIDIMENSIONAL TIME-SERIES DATA PROCESSING

Efficient processing of time series data is a fundamental yet challenging problem in pattern recognition. Though recent developments in machine learning and deep learning have enabled remarkable improvements in processing large scale datasets in many application domains, most are designed and regulated to handle inputs that are static in time. Many real-world data, such as in biomedical, surveillance and security, financial, manufacturing and engineering applications, are rarely static in time, and demand models able to recognize patterns in both space and time. Current machine learning (ML) and deep learning (DL) models adapted for time series processing tend to grow in complexity and size to accommodate the additional dimensionality of time. Specifically, the biologically inspired learning based models known as artificial neural networks that have shown extraordinary success in pattern recognition, tend to grow prohibitively large and cumbersome in the presence of large scale multi-dimensional time series biomedical data such as EEG. Consequently, this work aims to develop representative ML and DL models for robust and efficient large scale time series processing. First, we design a novel ML pipeline with efficient feature engineering to process a large scale multi-channel scalp EEG dataset for automated detection of epileptic seizures. With the use of a sophisticated yet computationally efficient time-frequency analysis technique known as harmonic wavelet packet transform and an efficient self-similarity computation based on fractal dimension, we achieve state-of-the-art performance for automated seizure detection in EEG data. Subsequently, we investigate the development of a novel efficient deep recurrent learning model for large scale time series processing. For this, we first study the functionality and training of a biologically inspired neural network architecture known as cellular simultaneous recurrent neural network (CSRN). We obtain a generalization of this network for multiple topological image processing tasks and investigate the learning efficacy of the complex cellular architecture using several state-of-the?art training methods. Finally, we develop a novel deep cellular recurrent neural network (CDRNN) architecture based on the biologically inspired distributed processing used in CSRN for processing time series data. The proposed DCRNN leverages the cellular recurrent architecture to promote extensive weight sharing and efficient, individualized, synchronous processing of multi-source time series data. Experiments on a large scale multi-channel scalp EEG, and a machine fault detection dataset show that the proposed DCRNN offers state-of-the-art recognition performance while using substantially fewer trainable recurrent units.

Vidyaratne, Lasitha S.↗

Learning dynamical systems from data: An introduction to physics-guided deep learning

Modeling complex physical dynamics is a fundamental task in science and engineering. Traditional physics-based models are first-principled, explainable, and sample-efficient. However, they often rely on strong modeling assumptions and expensive numerical integration, requiring significant computational resources and domain expertise. While deep learning (DL) provides efficient alternatives for modeling complex dynamics, they require a large amount of labeled training data. Furthermore, its predictions may disobey the governing physical laws and are difficult to interpret. Physics-guided DL aims to integrate first-principled physical knowledge into data-driven methods. It has the best of both worlds and is well equipped to better solve scientific problems. Recently, this field has gained great progress and has drawn considerable interest across discipline Here, we introduce the framework of physics-guided DL with a special emphasis on learning dynamical systems. We describe the learning pipeline and categorize state-of-the-art methods under this framework. We also offer our perspectives on the open challenges and emerging opportunities.

97 MATHEMATICS AND COMPUTING↗

Simurgh: A Framework for Cad-Driven Deep Learning Based X-Ray CT Reconstruction

High-resolution X-ray computed tomography (XCT) is an important technique for the inspection of additively manufactured (AM) parts. While XCT is typically used off-line to inspect a subset of manufactured parts, significantly accelerating measurement speed while retaining accuracy would enable use of XCT for in-line inspection to rapidly identify defects in each part as it is manufactured. Here, we propose a deep learning (DL) based approach that uses computer aided design (CAD) models of the AM parts and physics-based information to rapidly produce high-quality reconstructions from sparse XCT measurements without high quality ground truth data. Our approach uses a generative adversarial neural network (GAN) to produced realistic training data from the CAD-based simulations and a deep neural network that is trained using data from the first stage to produce accurate 3D reconstructions. Using experimental XCT data of metal parts, we demonstrate enhanced defect detection capabilities while dramatically reducing the scan time.

Ziabari, Amir↗

Counter Data Paucity through Adversarial Invariance Encoding: A Case Study on Modeling Battery Thermal Runaway

Lithium-ion batteries, widely used for their durability and high energy storage, face the risk of internal short circuits leading to catastrophic thermal runaway events. These events, triggered by external stimuli like mechanical loads, pose safety concerns in applications such as electric vehicles. Detecting and understanding thermal runaway events is crucial, but physics-driven models struggle to explain the non-linear evolution of battery temperature during these events, considering factors like material composition and state-of-charge. Due to the rarity of these events and the cost of data collection, we propose a deep learning (DL) model to predict battery temperature responses during thermal runaway. The challenge lies in the scarcity of data, making traditional DL models prone to overfitting and learning low-quality representations of the complex process.Our approach introduces a novel few-shot architecture that incorporates an adversarially governed invariant encoding process. This architecture aims to distill "invariant" relationships by addressing distributional shifts in data across various battery properties, facilitating the detection of thermal runaway events. Specifically, our results demonstrate that deep learning models conditioned on these "invariant" representations outperform state-of-the-art baselines, achieving a remarkable 96.8% performance improvement in terms of the popular metric MAPE. This framework presents a promising direction for enhancing battery safety modeling, particularly in the context of rare and complex events like thermal runaway. Our code and code and dataset used for the paper are public1.

Tabassum, Anika [ORNL] (ORCID:0000000254600955)↗

Enabling deeper learning on big data for materials informatics applications

Abstract The application of machine learning (ML) techniques in materials science has attracted significant attention in recent years, due to their impressive ability to efficiently extract data-driven linkages from various input materials representations to their output properties. While the application of traditional ML techniques has become quite ubiquitous, there have been limited applications of more advanced deep learning (DL) techniques, primarily because big materials datasets are relatively rare. Given the demonstrated potential and advantages of DL and the increasing availability of big materials datasets, it is attractive to go for deeper neural networks in a bid to boost model performance, but in reality, it leads to performance degradation due to the vanishing gradient problem. In this paper, we address the question of how to enable deeper learning for cases where big materials data is available. Here, we present a general deep learning framework based on Individual Residual learning (IRNet) composed of very deep neural networks that can work with any vector-based materials representation as input to build accurate property prediction models. We find that the proposed IRNet models can not only successfully alleviate the vanishing gradient problem and enable deeper learning, but also lead to significantly (up to 47%) better model accuracy as compared to plain deep neural networks and traditional ML techniques for a given input materials representation in the presence of big data.

36 MATERIALS SCIENCE↗

Intercomparison of Deep Learning Model Architectures for Atmospheric River Prediction

With a rapid surge in the application of machine learning (ML) for a diverse range of tasks in climate science, the present study addresses a challenge for climate scientists when selecting the optimal ML or deep learning (DL) architecture for a given application. In particular, a DL intercomparison study was performed with a focus on forecasting the position of atmospheric rivers (ARs) on short-range time scales (up to 5-day lead times). AR predictions from multiple DL architectures, including various types of convolutional autoencoders and a vision transformer (ViT), were compared against ECMWF ERA5 reanalysis and hindcasts from a global climate model. DL models with similar trainable parameters were trained on ERA5 reanalysis data and AR positions derived from a thresholding algorithm to ensure a fair comparison among the DL models. Each model’s performance and accuracy in forecasting AR location and key input fields within a 5-day window were assessed using metrics of root-mean-square error, anomaly correlation, and mean intersection over union. The ViT architecture outperformed other autoencoder models in most of the metrics. Incorporating additional meteorological fields only yielded slight improvements in forecasting certain fields at longer lead times. The results also suggest that a smaller number of input time steps or smaller number of autoregressive steps can achieve better prediction skills, while also improving the overall computational efficiency. This research offers valuable insights into the strengths and weaknesses of different DL techniques for AR forecasting, hopefully guiding the development of improved models for forecasting this phenomenon.

54 ENVIRONMENTAL SCIENCES↗

Towards multi-fidelity deep learning of wind turbine wakes

We report engineering wake models that accurately predict wake in a computationally efficient manner are very important for tasks such as layout optimization and control of wind farms. In this paper, we explore an application of deep learning (DL) to learn the wake model from hierarchies of physics-based approaches ranging from analytical models to an approximate form of the Reynolds-averaged Navier-Stokes equations. We first illustrate the application of principal component analysis to obtain a lower-dimensional representation that allows a computationally tractable training and deployment of DL models. Then, the DL model is trained to learn the mapping from input parameter space to the principal components, which are then used to reconstruct the three-dimensional flow field. Additionally, we investigate a composite framework consisting of two neural networks to learn the correlation between low- and high-fidelity data with Gauss and curl models treated as proxies for low- and high-fidelity models, respectively. The prediction from both DL models matches well with the high-fidelity data with a maximum relative percentage error for the kinetic energy flux of <1%. This work opens up possibilities for data-efficient construction of surrogate models for wake prediction that can be used to study the influence of wind speed and yaw angles on wind farm power production.

17 WIND ENERGY↗

A robust estimator of mutual information for deep learning interpretability

Abstract We develop the use of mutual information (MI), a well-established metric in information theory, to interpret the inner workings of deep learning (DL) models. To accurately estimate MI from a finite number of samples, we present GMM-MI (pronounced ‘Jimmie’), an algorithm based on Gaussian mixture models that can be applied to both discrete and continuous settings. GMM-MI is computationally efficient, robust to the choice of hyperparameters and provides the uncertainty on the MI estimate due to the finite sample size. We extensively validate GMM-MI on toy data for which the ground truth MI is known, comparing its performance against established MI estimators. We then demonstrate the use of our MI estimator in the context of representation learning, working with synthetic data and physical datasets describing highly non-linear processes. We train DL models to encode high-dimensional data within a meaningful compressed (latent) representation, and use GMM-MI to quantify both the level of disentanglement between the latent variables, and their association with relevant physical quantities, thus unlocking the interpretability of the latent representation. We make GMM-MI publicly available in this GitHub repository.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Deep Learning-based Surrogate Model for Efficient Reservoir Simulation in Large-scale Geological Carbon Storage: Application in IBDP Dataset

This project introduces an advanced deep learning (DL)-based surrogate modeling approach to enhance the efficiency and accuracy of large-scale geological carbon storage (GCS) simulations. Using the Illinois Basin Decatur Project (IBDP) dataset as training data, the study employs a residual U-Net architecture to predict critical state variables such as pressure and CO₂ saturation, as well as CO₂ plume migration. By incorporating key geological parameters (e.g., porosity, permeability, and rock facies) and physics-informed inputs like the diffusive time of flight and time step, the DL model effectively reduces computational complexity while maintaining robust physical constraints. Compared to traditional simulators like Eclipse, the DL model achieves remarkable accuracy, with a root mean square error (RMSE) of 1.57 psi for pressure and 0.007 for saturation, and dramatically reduces computational time from hours to just 69.9 seconds for 50-step simulations. These results demonstrate the potential of innovative DL methodologies to improve the predictivity and operational efficiency of GCS simulations, providing a reliable foundation for decision-making in CCS operations. Supported by the SMART initiative, this project underscores the success of leveraging computational innovations to advance CCS technologies.

advanced deep learning↗

Deep Learning Classification of Cheatgrass Invasion in the Western United States Using Biophysical and Remote Sensing Data

Cheatgrass (Bromus tectorum) invasion is driving an emerging cycle of increased fire frequency and irreversible loss of wildlife habitat in the western US. Yet, detailed spatial information about its occurrence is still lacking for much of its presumably invaded range. Deep learning (DL) has demonstrated success for remote sensing applications but is less tested on more challenging tasks like identifying biological invasions using sub-pixel phenomena. We compare two DL architectures and the more conventional Random Forest and Logistic Regression methods to improve upon a previous effort to map cheatgrass occurrence at >2% canopy cover. High-dimensional sets of biophysical, MODIS, and Landsat-7 ETM+ predictor variables are also compared to evaluate different multi-modal data strategies. All model configurations improved results relative to the case study and accuracy generally improved by combining data from both sensors with biophysical data. Cheatgrass occurrence is mapped at 30 m ground sample distance (GSD) with an estimated 78.1% accuracy, compared to 250-m GSD and 71% map accuracy in the case study. Furthermore, DL is shown to be competitive with well-established machine learning methods in a limited data regime, suggesting it can be an effective tool for mapping biological invasions and more broadly for multi-modal remote sensing applications.

54 ENVIRONMENTAL SCIENCES↗

Deep Neural Network Algorithm for CMC Microstructure Characterization and Variability Quantification

Microstructure characterization and variability quantification are crucial for understanding ceramic matrix composites (CMCs) mechanical behavior and deformation mechanisms across length scales. Traditionally, analyses of the micrographs obtained from microscopy are labor-intensive. However, with the vast improvement in computer vision (CV) and deep learning (DL), an automated algorithm can be designed to extract essential microstructure variability from micrographs which can then be used to construct a statistically representative volume element (SRVE). The DL-based algorithm spans the taxonomy of microstructure analyses, including semantic segmentation of microstructure constituents, secondary phases, matrix/fiber interface, and defects, and quantifying the microstructure variability in terms of probability distributions. In this work, C/SiNC and SiC/SiNC CMCs microstructures are semantically segmented through a deep convolutional neural network, followed by variability quantification through the implementation of a fully connected regression layer, hence forming a deep regression network. The deep regression network operates in a feedforward regime, in which the neuron output signal traverses through the network in a unidirectional manner. The weight tensor associated with each layer is updated through a backpropagation stochastic gradient descent approach. The input gray-scale image obtained through in-house scanning electron microscope and confocal microscope micrographs is augmented through affine transformations to increase the training set size, which is then processed through four strided convolutional layers. This compresses the image resolution by half at each layer while increasing the image depth by applying different filters (image encoding). The class activation maps (CAMs) corresponding to the applied filters highlight the key architectural features and assist with the semantic segmentation of the microstructure.

Hamza, Mohamed H.↗

Structural characterization of an intrinsically disordered protein complex using integrated small-angle neutron scattering and computing

Characterizing structural ensembles of intrinsically disordered proteins (IDPs) and intrinsically disordered regions (IDRs) of proteins is essential for studying structure–function relationships. Due to the different neutron scattering lengths of hydrogen and deuterium, selective labeling and contrast matching in small-angle neutron scattering (SANS) becomes an effective tool to study dynamic structures of disordered systems. However, experimental timescales typically capture measurements averaged over multiple conformations, leaving complex SANS data for disentanglement. We hereby demonstrate an integrated method to elucidate the structural ensemble of a complex formed by two IDRs. We use data from both full contrast and contrast matching with residue-specific deuterium labeling SANS experiments, microsecond all-atom molecular dynamics (MD) simulations with four molecular mechanics force fields, and an autoencoder-based deep learning (DL) algorithm. From our combined approach, we show that selective deuteration provides additional information that helps characterize structural ensembles. We find that among the four force fields, a99SB-disp and CHARMM36m show the strongest agreement with SANS and NMR experiments. In addition, our DL algorithm not only complements conventional structural analysis methods but also successfully differentiates NMR and MD structures which are indistinguishable on the free energy surface. Finally, we present an ensemble that describes experimental SANS and NMR data better than MD ensembles generated by one single force field and reveal three clusters of distinct conformations. Our results demonstrate a new integrated approach for characterizing structural ensembles of IDPs.

59 BASIC BIOLOGICAL SCIENCES↗

Predicting future well performance for environmental remediation design using deep learning

Here in this study, we developed a deep learning (DL) framework with a multi-channel three-dimensional convolutional neural network (MC3D-CNN) to predict well performance and thereby assist future environmental remediation design. Such prediction of extraction well performance at designated locations is critical for configuring pump-and-treat (P&T) well network design and operation, setting reasonable target closure dates for overall remedying, and estimating remedy costs. The framework is developed with operational and monitoring data routinely collected during P&T remedy operations, including well extraction and injection rates as well as in situ contaminant concentrations. Traditionally, the collected data were rarely used for purposes other than assessing past well performance and the accuracy of the conceptual site model. However, recent advances in data-driven computational approaches enable better use of the large datasets to inform future well performance, enhance site characterization, and improve remediation planning. In this study, we established a DL framework to integrate transient three-dimensional contaminant plumes and multiple aquifer properties (e.g., hydraulic conductivity and hydrostratigraphic maps) to identify characteristic patterns controlling and representing extraction well mass recovery, aiming at providing future mass recovery estimates for existing wells and candidate wells at any proposed locations. We evaluated our framework by using a realistic synthetic dataset generated from a well-calibrated flow and transport model used in the 200 West Area of the U.S. Department of Energy’s Hanford Site in southeastern Washington state. The multi-channel feature in our framework allows integration of various types and temporal densities of training datasets for DL model development. Overall, we found that the trained DL model achieved an accuracy of over 90% in ranking extraction well performance in validation datasets, and over 80% in predicting high-performance-ranking well locations. This data-informed approach provides a flexible tool to support adaptive site management, streamline decision-making, and potentially reduce remediation time and costs. Our DL framework can be used as a filtering tool to improve the current P&T network optimization design by reducing the number of candidate well locations.

54 ENVIRONMENTAL SCIENCES↗

Reducing Urban Traffic Congestion Using Deep Learning and Model Predictive Control

This article proposes a deep learning (DL)-based control algorithm—DL velocity-based model predictive control (VMPC)—for reducing traffic congestion with slowly time-varying traffic signal controls. This control algorithm consists of system identification using DL and traffic signal control using VMPC. For the training process of DL, we established a modeling error entropy loss as the criteria inspired by the theory of stochastic distribution control (SDC) originated by the fourth author. Herein, simulation results show that the proposed algorithm can reduce traffic congestion with a slowly varying traffic signal control input. Results of an ablation study demonstrate that this algorithm compares favorably to other model-based controllers in terms of prediction error, signal varying speed, and control effectiveness.

97 MATHEMATICS AND COMPUTING↗

A Multiscale Deep Learning Model for Soil Moisture Integrating Satellite and In Situ Data

Deep learning (DL) models trained on hydrologic observations can perform extraordinarily well, but they can inherit deficiencies of the training data, such as limited coverage of in situ data or low resolution/accuracy of satellite data. In this work, we propose a novel multiscale DL scheme learning simultaneously from satellite and in situ data to predict 9 km daily soil moisture (5 cm depth). Based on spatial cross-validation over sites in the conterminous United States, the multiscale scheme obtained a median correlation of 0.901 and root-mean-square error of 0.034 m 3 /m 3 . It outperformed the Soil Moisture Active Passive satellite mission's 9 km product, DL models trained on in situ data alone, and land surface models. Our 9 km product showed better accuracy than previous 1 km satellite downscaling products, highlighting limited impacts of improving resolution. Not only is our product useful for planning against floods, droughts, and pests, our scheme is generically applicable to geoscientific domains with data on multiple scales, breaking the confines of individual data sets.

54 ENVIRONMENTAL SCIENCES↗

Deep Learning Systems for Increased Safeguards Surveillance Review Productivity

Nuclear safeguards inspectors expend significant time and maintain intense focus in reviewing video surveillance for safeguards relevant events. To increase efficiency and reduce the time burden of safeguards inspectors performing surveillance data review, this paper presents a novel deep learning (DL) systems concept to integrate generalized DL models into the safeguards surveillance review workflow. The Agency is investigating several DL algorithms for object recognition, localization, tracking, and flagging relevant activities. The project team is working closely with nuclear safeguards inspectors to identify review use cases (based on specific safeguards objectives) and collect their associated requirements. We focused on CANDU and LWR Nuclear Power Plants (NPPs) and their associated dry storage areas as these present a particularly heavy burden on the inspector surveillance review process due to the number of these facilities under safeguards worldwide. Initial DL algorithm results on safeguards data are promising. Using a convolutional neural network, the team attained a mean average precision (mAP) of 92.9% identifying and localizing spent fuel (SF) casks from a 475 surveillance image dataset. Further, the team had initial success in training a recurrent neural network to identify reactor area activities in video clips, successfully indicating when SF casks enter or exit a pool. We discuss how such DL algorithms would be integrated into the Next Generation Surveillance Review (NGSR) software application. Another issue impacting review productivity is the long time inspectors may have to wait when running these algorithms in NGSR. We present a concept to pre-process remotely collected surveillance data with DL models as the data arrives to IAEA headquarters so that results are already available when starting a new review in NGSR. The proposed DL system concept shows a pathway and workflow for increasing an inspector’s surveillance review productivity by quickly and accurately identifying declared and undeclared safeguards relevant objects and activities in large quantities of surveillance imagery data.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Evaluating probabilistic deep learning methods for uncertainty quantification of temperature downscaling

Deep learning (DL) has emerged as a promising tool for downscaling coarse-resolution climate data to high-resolution outputs, enabling improved regional climate predictions. A critical aspect of DL-based downscaling is the incorporation of uncertainty quantification (UQ), which enhances the interpretability and reliability of predictions—key factors for climate risk assessment and decision-making. This study develops a DL model to downscale 2 m temperature across the contiguous United States using reanalysis datasets. We systematically evaluate three epistemic UQ methods—deep ensembles (DEns), Monte Carlo dropout (MCD), and Flipout—based on their probabilistic accuracy, downscaling performance, sensitivity to geographical features, and computational efficiency. Results indicate that MCD generally outperforms Flipout and DEns in terms of calibration and downscaling accuracy. However, DEns demonstrate lower calibration errors in coastal regions, indicating its higher confidence within these areas. Flipout, in contrast, is more sensitive to elevation gradients and exhibits higher calibration errors in mountainous regions. Hence, the choice of UQ method for this task depends on the specific requirements of the application. For applications that prioritize overall calibration, downscaling accuracy, and computational efficiency, MCD is a strong candidate. These findings highlight the importance of selecting UQ methods based on application-specific requirements, such as geographical context and computational constraints. By addressing the trade-offs between UQ methods, this study provides actionable insights for improving the reliability, scalability, and utility of DL-based downscaling in climate science.

Environmental sciences↗

Advancing the Limits of InSAR to Detect Crustal Displacement from Low-Magnitude Earthquakes through Deep Learning

Detecting surface deformation associated with low-magnitude (M w ≤ 5) seismicity using interferometric synthetic aperture radar (InSAR) is challenging due to the subtlety of the signal and the often challenging imaging environments. However, low-magnitude earthquakes are potential precursors to larger seismic events, and thus characterizing the crustal displacement associated with them is crucial for regional seismic hazard assessment. We combine InSAR time-series techniques with a Deep Learning (DL) autoencoder denoiser to detect the magnitude and extent of crustal deformation from the M w = 3.4 Gallina, New Mexico earthquake that occurred on 30 July 2020. Although InSAR alone cannot detect event-related deformation from such a low-magnitude seismic event, application of the DL method reveals maximum displacements as small as (±2.5 mm) in the vicinity of both the fault and earthquake epicenter without prior knowledge of the fault system. This finding improves small-scale displacement discernment with InSAR by an order of magnitude relative to previous studies. We additionally estimate best-fitting fault parameters associated with the observed deformation. The application of the DL technique unlocks the potential for low-magnitude earthquake studies, providing new insights into local fault geometries and potential risks from higher-magnitude earthquakes. This technique also permits low-magnitude event monitoring in areas where seismic networks are sparse, allowing for the possibility of global fault deformation monitoring.

58 GEOSCIENCES↗