Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “deep learning (DL)”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Simultaneously improving accuracy and computational cost under parametric constraints in materials property prediction tasks

Abstract Modern data mining techniques using machine learning (ML) and deep learning (DL) algorithms have been shown to excel in the regression-based task of materials property prediction using various materials representations. In an attempt to improve the predictive performance of the deep neural network model, researchers have tried to add more layers as well as develop new architectural components to create sophisticated and deep neural network models that can aid in the training process and improve the predictive ability of the final model. However, usually, these modifications require a lot of computational resources, thereby further increasing the already large model training time, which is often not feasible, thereby limiting usage for most researchers. In this paper, we study and propose a deep neural network framework for regression-based problems comprising of fully connected layers that can work with any numerical vector-based materials representations as model input. We present a novel deep regression neural network, iBRNet, with branched skip connections and multiple schedulers, which can reduce the number of parameters used to construct the model, improve the accuracy, and decrease the training time of the predictive model. We perform the model training using composition-based numerical vectors representing the elemental fractions of the respective materials and compare their performance against other traditional ML and several known DL architectures. Using multiple datasets with varying data sizes for training and testing, We show that the proposed iBRNet models outperform the state-of-the-art ML and DL models for all data sizes. We also show that the branched structure and usage of multiple schedulers lead to fewer parameters and faster model training time with better convergence than other neural networks. Scientific contribution: The combination of multiple callback functions in deep neural networks minimizes training time and maximizes accuracy in a controlled computational environment with parametric constraints for the task of materials property prediction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Cross-property deep transfer learning framework for enhanced predictive analytics on small materials data

Abstract Artificial intelligence (AI) and machine learning (ML) have been increasingly used in materials science to build predictive models and accelerate discovery. For selected properties, availability of large databases has also facilitated application of deep learning (DL) and transfer learning (TL). However, unavailability of large datasets for a majority of properties prohibits widespread application of DL/TL. We present a cross-property deep-transfer-learning framework that leverages models trained on large datasets to build models on small datasets of different properties. We test the proposed framework on 39 computational and two experimental datasets and find that the TL models with only elemental fractions as input outperform ML/DL models trained from scratch even when they are allowed to use physical attributes as input, for 27/39 (≈ 69%) computational and both the experimental datasets. We believe that the proposed framework can be widely useful to tackle the small data challenge in applying AI/ML in materials science.

36 MATERIALS SCIENCE↗

Deep learning applications in visual data for benign and malignant hematologic conditions: a systematic review and visual glossary

Deep learning (DL) is a subdomain of artificial intelligence algorithms capable of automatically evaluating subtle graphical features to make highly accurate predictions, which was recently popularized in multiple imaging-related tasks. Because of its capabilities to analyze medical imaging such as radiology scans and digitized pathology specimens, DL has significant clinical potential as a diagnostic or prognostic tool. Coupled with rapidly increasing quantities of digital medical data, numerous novel research questions and clinical applications of DL within medicine have already been explored. Similarly, DL research and applications within hematology are rapidly emerging, although these are still largely in their infancy. Given the exponential rise of DL research for hematologic conditions, it is essential for the practising hematologist to be familiar with the broad concepts and pitfalls related to these new computational techniques. This narrative review provides a visual glossary for key deep learning principles, as well as a systematic review of published investigations within malignant and non-malignant hematologic conditions, organized by the different phases of clinical care. In order to assist the unfamiliar reader, this review highlights key portions of current literature and summarizes important considerations for the critical understanding of deep learning development and implementations in clinical practice.

60 APPLIED LIFE SCIENCES↗

Deep learning to estimate permeability using geophysical data

Time-lapse electrical resistivity tomography (ERT) is a popular geophysical method to estimate three-dimensional (3D) permeability fields from electrical potential difference measurements. Traditional inversion and data assimilation methods are used to ingest this ERT data into hydrogeophysical models to estimate permeability. Due to ill-posedness and the curse of dimensionality, existing inversion strategies provide poor estimates and low resolution of the 3D permeability field. Recent advances in deep learning provide us with powerful algorithms to overcome this challenge. This paper presents a deep learning (DL) framework to estimate the 3D subsurface permeability from time-lapse ERT data. To test the feasibility of the proposed framework, we train DL-enabled inverse models on simulation data. Each measurement in both synthetic and field data is standardized by removing the mean and scaling the time-series to unit variance. This pre-processing step is necessary to bring simulation data closer to field observations. Subsurface process models based on hydrogeophysics are used to generate this synthetic data. Training performed on limited simulation data resulted in the DL model over-fitting. An advanced data augmentation based on mixup is implemented to generate additional training samples to overcome this issue. This mixup technique creates weakly labeled (low-fidelity) samples from strongly labeled (high-fidelity) data. The weakly labeled training data is then used to develop DL-enabled inverse models and reduce over-fitting. As both time-lapse ERT (1133048 features/realization) and 3D permeability (585453 features/realization) data samples are from a high-dimensional space, principal component analysis (PCA) is employed to reduce dimensionality. Encoded ERT and encoded permeability are generated using the trained PCA estimators. A deep neural network is then trained to map the encoded ERT to encoded permeability. This mixup training and unsupervised learning allowed us to build a fast and reasonably accurate DL-based inverse model under limited simulation data. Results show that proposed weak supervised learning can capture salient spatial features in the 3D permeability field. Quantitatively, the average mean squared error (in terms of the natural log) on the strongly labeled training, validation, and test datasets is less than 0.5. The R 2 -score (global metric) is greater than 0.75, and the percent error in each cell (local metric) is less than 10%. Finally, an added benefit in terms of computational cost is that the proposed DL-based inverse model is at least O(10 4 ) times faster than running a forward model once it is trained. Data generation, DL model training, and hyperparameter tuning to identify optimal neural network architectures utilized high-performance computing resources while the DL inference is performed on a standard laptop. Approximately, O(10 5 ) processor hours are used for generating data and DL tuning and training. We acknowledge that the data generation and DL model development are expensive. But once a DL model is trained, it can be re-used for inversion rapidly for the given system, with set physics and domain. Note that traditional inversion may require multiple forward model simulations (e.g., in the order of 10 to 1000), which are very expensive. This computational savings ≈ O(10 5 ) – O(10 7 )) makes the proposed DL-based inverse model attractive for subsurface imaging and real-time ERT monitoring applications due to fast and yet reasonably accurate estimations of permeability field.

58 GEOSCIENCES↗

Surrogate Model Based Optimization for Finding Robust Deep Learning Model Architectures

Deep Learning (DL) models are increasingly used throughout the sciences. However, their performance and usefulness depend greatly on their architecture which is defined by hyperparameters such as the number of nodes, layers, the learning rate, etc. Tuning these hyperparameters is time-consuming because evaluating their performance requires a lengthy training step. Stochastic optimizers used in training lead to performance variability and potentially prediction reliability issues. In this talk, we will describe an automated optimization method based on surrogate models and active learning strategies for tuning DL model architectures. We take into account the prediction variability with the goal to identify architectures that make reliable and robust predictions. We demonstrate our developments on an application arising in particle physics.

deep learning↗

UNNT: A novel Utility for comparing Neural Net and Tree-based models

The use of deep learning (DL) is steadily gaining traction in scientific challenges such as cancer research. Advances in enhanced data generation, machine learning algorithms, and compute infrastructure have led to an acceleration in the use of deep learning in various domains of cancer research such as drug response problems. In our study, we explored tree-based models to improve the accuracy of a single drug response model and demonstrate that tree-based models such as XGBoost (eXtreme Gradient Boosting) have advantages over deep learning models, such as a convolutional neural network (CNN), for single drug response problems. However, comparing models is not a trivial task. To make training and comparing CNNs and XGBoost more accessible to users, we developed an open-source library called UNNT (A novel Utility for comparing Neural Net and Tree-based models). The case studies, in this manuscript, focus on cancer drug response datasets however the application can be used on datasets from other domains, such as chemistry.

59 BASIC BIOLOGICAL SCIENCES↗

Elastic Resource Management for Deep Learning Applications in a Container Cluster

The increasing demand for learning from massive datasets is restructuring our economy. Effective learning, however, involves nontrivial computing resources. Most businesses utilize commercial infrastructure providers (e.g., AWS) to host their computing clusters in the cloud, where various jobs compete for available resources. While cloud resource management is a fruitful research field that has made many advances in production, such as Kubernetes and YARN, few efforts have been invested to further optimize the system performance, especially for deep learning (DL) training jobs in a container cluster. This work introduces FlowCon, a system that is able to monitor the individual evaluation functions of DL jobs at runtime, and thus to make placement decisions on resource allocations elastically. Here, we present a detailed design and implementation of FlowCon and conduct intensive experiments over various DL models. The results demonstrate that FlowCon significantly improves DL job completion time and resource utilization efficiency, compared to default systems. According to the results, FlowCon is able to improve the completion time by up to 68.8% and meanwhile, reduce the makespan by 18.0%, in the presence of various DL job workloads.

97 MATHEMATICS AND COMPUTING↗

A Framework for Deep Learning Emulation of Numerical Models With a Case Study in Satellite Remote Sensing

Numerical models based on physics represent the state of the art in Earth system modeling and comprise our best tools for generating insights and predictions. Despite rapid growth in computational power, the perceived need for higher model resolutions overwhelms the latest generation computers, reducing the ability of modelers to generate simulations for understanding parameter sensitivities and characterizing variability and uncertainty. Thus, surrogate models are often developed to capture the essential attributes of the full-blown numerical models. Recent successes of machine learning methods, especially deep learning (DL), across many disciplines offer the possibility that complex nonlinear connectionist representations may be able to capture the underlying complex structures and nonlinear processes in Earth systems. A difficult test for DL-based emulation, which refers to function approximation of numerical models, is to understand whether they can be comparable to traditional forms of surrogate models in terms of computational efficiency while simultaneously reproducing model results in a credible manner. A DL emulation that passes this test may be expected to perform even better than simple models with respect to capturing complex processes and spatiotemporal dependencies. Here, we examine, with a case study in satellite-based remote sensing, the hypothesis that DL approaches can credibly represent the simulations from a surrogate model with comparable computational efficiency. Our results are encouraging in that the DL emulation reproduces the results with acceptable accuracy and often even faster performance. We discuss the broader implications of our results in light of the pace of improvements in high-performance implementations of DL and the growing desire for higher resolution simulations in the Earth sciences.

Bayesian Deep Learning↗

Graph interpolating activation improves both natural and robust accuracies in data-efficient deep learning

Improving the accuracy and robustness of deep neural nets (DNNs) and adapting them to small training data are primary tasks in deep learning (DL) research. In this paper, we replace the output activation function of DNNs, typically the data-agnostic softmax function, with a graph Laplacian-based high-dimensional interpolating function which, in the continuum limit, converges to the solution of a Laplace–Beltrami equation on a high-dimensional manifold. Furthermore, we propose end-to-end training and testing algorithms for this new architecture. The proposed DNN with graph interpolating activation integrates the advantages of both deep learning and manifold learning. Compared to the conventional DNNs with the softmax function as output activation, the new framework demonstrates the following major advantages: First, it is better applicable to data-efficient learning in which we train high capacity DNNs without using a large number of training data. Second, it remarkably improves both natural accuracy on the clean images and robust accuracy on the adversarial images crafted by both white-box and black-box adversarial attacks. Third, it is a natural choice for semi-supervised learning. This paper is a significant extension of our earlier work published in NeurIPS, 2018. For reproducibility, the code is available at https://github.com/BaoWangMath/DNN-DataDependentActivation .

Mathematics↗

Deep learning-based spatio-temporal estimate of greenhouse gas emissions using satellite data

Accurate estimation of greenhouse gases (GHGs) emissions is very important for developing mitigation strategies to climate change by controlling and reducing GHG emissions. This project aims to develop multiple deep learning approaches to estimate anthropogenic greenhouse gas emissions using multiple types of satellite data. NO2 concentration is chosen as an example of GHGs to evaluate the proposed approach. Two sentinel satellites (sentinel-2 and sentinel-5P) provide multiscale observations of GHGs from 10-60m resolution (sentinel-2) to ~kilometer scale resolution (sentinel-5P). Among multiple deep learning (DL) architectures evaluated, two best DL models demonstrate that key features of spatio-temporal satellite data and additional information (e.g., observation times and/or coordinates of ground stations) can be extracted using convolutional neural networks and feed forward neural networks, respectively. In particular, irregular time series data from different NO 2 observation stations limit the flexibility of long short-term memory architecture, requiring zero-padding to fill in missing data. However, deep neural operator (DNO) architecture can stack time-series data as input, providing the flexibility of input structure without zero-padding. As a result, the DNO outperformed other deep learning architectures to account for time-varying features. Overall, temporal patterns with smooth seasonal variations were predicted very well, while frequent fluctuation patterns were not predicted well. In addition, uncertainty quantification using conformal inference method is performed to account for prediction ranges. Overall, this research will lead to a new groundwork for estimating greenhouse gas concentrations using multiple satellite data to enhance our capability of tracking the cause of climate change and developing mitigation strategies.

54 ENVIRONMENTAL SCIENCES↗

Machine and Deep Learning: Artificial Intelligence Application in Biotic and Abiotic Stress Management in Plants

Biotic and abiotic stresses significantly affect plant fitness, resulting in a serious loss in food production. Biotic and abiotic stresses predominantly affect metabolite biosynthesis, gene and protein expression, and genome variations. However, light doses of stress result in the production of positive attributes in crops, like tolerance to stress and biosynthesis of metabolites, called hormesis. Advancement in artificial intelligence (AI) has enabled the development of high-throughput gadgets such as high-resolution imagery sensors and robotic aerial vehicles, i.e., satellites and unmanned aerial vehicles (UAV), to overcome biotic and abiotic stresses. These High throughput (HTP) gadgets produce accurate but big amounts of data. Significant datasets such as transportable array for remotely sensed agriculture and phenotyping reference platform (TERRA-REF) have been developed to forecast abiotic stresses and early detection of biotic stresses. For accurately measuring the model plant stress, tools like Deep Learning (DL) and Machine Learning (ML) have enabled early detection of desirable traits in a large population of breeding material and mitigate plant stresses. In this review, advanced applications of ML and DL in plant biotic and abiotic stress management have been summarized.

59 BASIC BIOLOGICAL SCIENCES↗

Deep Learning enabled spectral energy conversion for in situ exposure measurements

A detector-specific deep learning (DL) approach is presented for spectra-to-exposure conversion using large-format sodium iodide (NaI(Tl)) detectors deployed for in situ environmental radiation measurements in emergency response scenarios. Accurate determination of exposure from NaI spectra is challenging due to poor energy resolution, partial energy absorption, and the strong sensitivity of traditionally deployed analytical conversion methods to calibrated source geometry and pre-deployment assumptions. Here, to address these limitations, a multi-layer perceptron model was trained on a hybrid in situ /Monte Carlo dataset constructed to span a broad range of photon energies, spatial extents, and realistic deployment variability, representative of general in situ emergency response conditions. The DL model was evaluated against commonly fielded analytical approaches under matched simulation conditions, including a single-factor method, a G-function method, and a modeled pressurized ion chamber (PIC) baseline. This study was intentionally computational in scope to enable controlled, like-for-like comparisons between conversion techniques while minimizing confounding real-world variability. Comparison to the modeled PIC provides contextual benchmarking and is not intended as a field inter-comparison with deployed instruments. Across the evaluated 20 keV to 3 MeV energy range, the DL approach consistently exhibited higher accuracy and reduced variance relative to the analytical methods against a deterministically calculated exposure. This may indicate improved robustness to spectral complexity without reliance on source-, geometric-, or spectral region-specific optimization. While results do not represent real-world validation, the presented work demonstrates that deep learning may effectively learn the nonlinear detector response-to-exposure relationship for asymmetric NaI(Tl) detectors and offers a promising pathway for improving in situ exposure estimation using spectroscopic systems already integrated into initial real-time emergency response operations.

61 RADIATION PROTECTION AND DOSIMETRY↗

ML-based Dimension Reduction Strategies

Deep learning (DL)--based surrogate models have achieved success in various applications in carbon capture and storage (CCS). However, the model training on high-dimensional spaces is computationally expensive and impractical for large-scale and complex geological models, because the models usually contain hundreds of thousands to millions of grid cells, each with a set of parameters. Furthermore, the high cost of generating training data with sufficient variation is another limitation of model training on high-dimensional spaces, which may result in overfitting and reduce the model efficiency and prediction performance. We proposed the workflow incorporating dimension reduction methods and deep learning models, which aim to extract the latent variables of input parameters and output state variables, and then build the mapping function at the latent spaces. The proposed workflow can significantly reduce the computational complexity in solving both forward and inverse problems compared to models trained on high-dimensional spaces. Dimensionality reduction models showed great potential in workflows for fast reservoir simulation, history matching, prior model generation, visualization, and more, ultimately enhancing DL model performance in related SMART Work Packages.

Hosseini, Seyyed↗

UIR-Net: Object Detection in Infrared Imaging of Thermomechanical Processes in Automotive Manufacturing

Thermomechanical processes (TMPs) such as resistance spot welding (RSW) and hot stamping are widely used in automotive manufacturing. Recent advancement in sensing technology has led to an increasing adoption of thermographic cameras to capture the infrared (IR) radiation of a metal part (or component of a part) during its thermomechanical processing or immediately after the process when the part is still hot. Detecting the object(s) of interest from raw IR images is an essential step in analyzing these data. Deep learning (DL) has been a recent success for object detection (OD), but the application of DL-based OD for industrial IR images in manufacturing is largely lagging behind. The major contribution of this work, which is also the distinction from previous OD studies, is the capability of building the OD model with unlabeled IR images, i.e., imaging data without accurate information indicating the object position. Here, the architecture of Unsupervised IR Image Net (UIR-Net) is designed to accommodate the unique characteristics of IR images from TMPs in manufacturing. This study presents a novel method for OD in unlabeled IR images from TMPs. The proposed method, called UIR-Net, consists of two components: label generation and DL model construction. Two case studies from automotive manufacturing, RSW and hot stamping, are reported to demonstrate the feasibility and effectiveness of the proposed method.

42 ENGINEERING↗

Unbundling Smart Meter Services Through Spatio-Temporal Decomposition Agents in DER-rich Environment

Smart meters and the advanced metering infrastructure (AMI) facilitate distribution system operators (DSOs) to gather information on energy consumption at the customer level. With the increasing penetration of building-level intermittent distributed energy resources (DERs) behind the meter, DER information is not available to DSOs. At the same time, smart meter enables users to participate in grid, with real-time information. Information for behind the meter is needed by user to coordinate building level assets for maximum benefits. The concept of unbundled smart meter (USM) needs agents to decompose smart meter measurements to provide service to DSO as well as customers. In this paper, we propose a Spatio-Temporal Decomposition Agent (STDA) for USM based on Artificial Intelligence (AI). STDA can help users optimize their energy usage, help DSO to utilize building assets for the grid operation. The energy usage strategy developed by STDA is suitable for different users, and can be customized by deep learning (DL) models according to the different energy consumption habits of each user. The power prediction performance results of various DL models and evaluation using a set of data from a Hawaii utility is presented. Furthermore, STDA integration with Home Energy management Systems (HEMS) to manage resources is presented and validated. STDA pre-processes the measurements before model training, and provides the spatio-temporal decomposed forecasting.

42 ENGINEERING↗

Time-lapse seismic inversion for CO 2 saturation with SeisCO2Net: An application to Frio-II site

Seismic monitoring of geological CO 2 storage (GCS) involves highly nonlinear seismic inversion and petrophysical inversion, making it challenging to estimate CO 2 volume efficiently and detect possible early CO 2 leakages. Deep learning (DL) using convolutional neural networks (CNNs) has shown promise in solving highly nonlinear seismic inversion problems. However, direct estimation of CO 2 plume extent/saturation from time-lapse seismic gathers using DL is still underexplored, with no reported field applications to date. The investigation of field data is primarily hindered by scarcity of field data for neural network training. Other obstacles include highly nonlinear seismic-petrophysics inverse relationship, and presence of noise in field seismic data. We introduce SeisCO2Net, a deep CNN that predicts CO 2 saturation maps directly from time-lapse full waveform shot gathers. For training, we use site-specific geological information, fluid flow physics, rock physics, and seismic modeling to generate synthetic datasets that closely resemble the CO 2 storage site. Synthetic tests show promising results, inspiring us to apply SeisCO2Net's trained weights on field data collected at Frio-II GCS site by leveraging transfer learning principles. As reference, we compare SeisCO2Net's predicted CO 2 saturation maps with results obtained from physics-based inversion. Our analyses show both methods display similar CO 2 plume shapes, reasonable CO 2 plume characteristics, and comparable saturation values. Our results suggest pre-training CNNs on physics-informed synthetic datasets and then applying the learned weights to field data is a viable approach to estimating field CO 2 saturation. This method effectively addresses the scarcity of field training data, thus encouraging the feasibility of long-term GCS monitoring.

58 GEOSCIENCES↗

Effect of image resolution on automated classification of chest X-rays

Deep learning (DL) models have received much attention lately for their ability to achieve expert-level performance on the accurate automated analysis of chest X-rays (CXRs). Recently available public CXR datasets include high resolution images, but state-of-the-art models are trained on reduced size images due to limitations on graphics processing unit memory and training time. As computing hardware continues to advance, it has become feasible to train deep convolutional neural networks on high-resolution images without sacrificing detail by downscaling. This study examines the effect of increased resolution on CXR classification performance. We used the publicly available MIMIC-CXR-JPG dataset, comprising 377,110 high resolution CXR images for this study. We applied image downscaling from native resolution to 2048 × 2048 pixels, 1024 × 1024 pixels, 512 × 512 pixels, and 256 × 256 pixels and then we used the DenseNet121 and EfficientNet-B4 DL models to evaluate clinical task performance using these four downscaled image resolutions. We find that while some clinical findings are more reliably labeled using high resolutions, many other findings are actually labeled better using downscaled inputs. We qualitatively verify that tasks requiring a large receptive field are better suited to downscaled low resolution input images, by inspecting effective receptive fields and class activation maps of trained models. Lastly, we show that stacking an ensemble across resolutions outperforms each individual learner at all input resolutions while providing interpretable scale weights, indicating that diverse information is extracted across resolutions.

47 OTHER INSTRUMENTATION↗

Deep learning multiphysics network for imaging CO 2 saturation and estimating uncertainty in geological carbon storage

Multiphysics inversion exploits different types of geophysical data that often complement each other and aims to improve overall imaging resolution and reduce uncertainties in geophysical interpretation. Despite the advantages, traditional multiphysics inversion is challenging because it requires a large amount of computational time and intensive human interactions for preprocessing data and finding trade-off parameters. These issues make it nearly impossible for traditional multiphysics inversion to be applied as a real-time monitoring tool for geological carbon storage. In this paper, we present a deep learning (DL) multiphysics network for imaging CO 2 saturation in real time. The multiphysics network consists of three encoders for analysing seismic, electromagnetic and gravity data and shares one decoder for combining imaging capabilities of the different geophysical data for better predicting CO 2 saturation. The network is trained on pairs of CO 2 label models and multiphysics data so that it can directly image CO 2 saturation. Here we use the bootstrap aggregating method to enhance the imaging accuracy and estimate uncertainties associated with CO 2 saturation images. Using realistic CO 2 label models and multiphysics data derived from the Kimberlina CO 2 storage model, we evaluate the performance of the deep learning multiphysics network and compare its imaging results to those from the deep learning single-physics networks. Our modelling experiments show that the deep learning multiphysics network for seismic, electromagnetic, and gravity data not only improves the imaging accuracy but also reduces uncertainties associated with CO 2 saturation images. Our results also suggest that the deep learning multiphysics network for the non-seismic data (i.e., electromagnetic and gravity) can be used as an effective low-cost monitoring tool in between regular seismic monitoring.

58 GEOSCIENCES↗