Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “deep neural network”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Improving Deep Neural Networks’ Training for Image Classification With Nonlinear Conjugate Gradient-Style Adaptive Momentum

Momentum is crucial in stochastic gradient-based optimization algorithms for accelerating or improving training deep neural networks (DNNs). In deep learning practice, the momentum is usually weighted by a well-calibrated constant. However, tuning the hyperparameter for momentum can be a significant computational burden. In this article, we propose a novel adaptive momentum for improving DNNs training; this adaptive momentum, with no momentum-related hyperparame- ter required, is motivated by the nonlinear conjugate gradient (NCG) method. Stochastic gradient descent (SGD) with this new adaptive momentum eliminates the need for the momentum hyperparameter calibration, allows using a significantly larger learning rate, accelerates DNN training, and improves the final accuracy and robustness of the trained DNNs. For example, SGD with this adaptive momentum reduces classification errors for training ResNet110 for CIFAR10 and CIFAR100 from 5.25% to 4.64% and 23.75% to 20.03%, respectively. Furthermore, SGD, with the new adaptive momentum, also benefits adversarial training and, hence, improves the adversarial robustness of the trained DNNs.

97 MATHEMATICS AND COMPUTING↗

Estimating Lossy Compressibility of Scientific Data Using Deep Neural Networks

Simulation based scientific applications generate increasingly large amounts of data on high-performance computing (HPC) systems. To allow data to be stored and analyzed efficiently, data compression is often utilized to reduce the volume and velocity of data. However, a question often raised by domain scientists is the level of compression that can be expected so that they can make more informed decisions, balancing between accuracy and performance. In this letter, we propose a deep neural network based approach for estimating the compressibility of scientific data. To train the neural network, we build both general features as well as compressor-specific features so that the characteristics of both data and lossy compressors are captured in training. Our approach is demonstrated to outperform a prior analytical model as well as a sampling based approach in the case of a biased estimation, i.e., for SZ. However, for the unbiased estimation (i.e., ZFP), the sampling based approach yields the best accuracy, despite the high overhead involved in sampling the target dataset.

97 MATHEMATICS AND COMPUTING↗

Modeling Liquid Water by Climbing up Jacob’s Ladder in Density Functional Theory Facilitated by Using Deep Neural Network Potentials

Within the framework of Kohn–Sham density functional theory (DFT), the ability to provide good predictions of water properties by employing a strongly constrained and appropriately normed (SCAN) functional has been extensively demonstrated in recent years. Here, we further advance the modeling of water by building a more accurate model on the fourth rung of Jacob’s ladder with the hybrid functional, SCAN0. In particular, we carry out both classical and Feynman path-integral molecular dynamics calculations of water with the SCAN0 functional and the isobaric–isothermal ensemble. To generate the equilibrated structure of water, a deep neural network potential is trained from the atomic potential energy surface based on ab initio data obtained from SCAN0 DFT calculations. For the electronic properties of water, a separate deep neural network potential is trained by using the Deep Wannier method based on the maximally localized Wannier functions of the equilibrated trajectory at the SCAN0 level. The structural, dynamic, and electric properties of water were analyzed. The hydrogen-bond structures, density, infrared spectra, diffusion coefficients, and dielectric constants of water, in the electronic ground state, are computed by using a large simulation box and long simulation time. For the properties involving electronic excitations, we apply the GW approximation within many-body perturbation theory to calculate the quasiparticle density of states and bandgap of water. Compared to the SCAN functional, mixing exact exchange mitigates the self-interaction error in the meta-generalized-gradient approximation and further softens liquid water toward the experimental direction. For most of the water properties, the SCAN0 functional shows a systematic improvement over the SCAN functional. However, some important discrepancies remain. The H-bond network predicted by the SCAN0 functional is still slightly overstructured compared to the experimental results.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Identifying Genomic Islands with Deep Neural Networks

Background Horizontal gene transfer is the main source of adaptability for bacteria, through which genes are obtained from different sources including bacteria, archaea, viruses, and eukaryotes. This process promotes the rapid spread of genetic information across lineages, typically in the form of clusters of genes referred to as genomic islands (GIs). Different types of GIs exist, and are often classified by the content of their cargo genes or their means of integration and mobility. While various computational methods have been devised to detect different types of GIs, no single method is capable of detecting all types. Results We propose a method, which we call Shutter Island, that uses a deep learning model (Inception V3, widely used in computer vision) to detect genomic islands. The intrinsic value of deep learning methods lies in their ability to generalize. Via a technique called transfer learning, the model is pre-trained on a large generic dataset and then re-trained on images that we generate to represent genomic fragments. We demonstrate that this image-based approach generalizes better than the existing tools. Conclusions We used a deep neural network and an image-based approach to detect the most out of the correct GI predictions made by other tools, in addition to making novel GI predictions. The fact that the deep neural network was re-trained on only a limited number of GI datasets and then successfully generalized indicates that this approach could be applied to other problems in the field where data is still lacking or hard to curate.

Computer Vision↗

Power System Event Identification Based on Deep Neural Network With Information Loading

Online power system event identification and classification are crucial to enhancing the reliability of transmission systems. In this study, we develop a deep neural network (DNN) based approach to identify and classify power system events by leveraging real-world measurements from hundreds of phasor measurement units (PMUs) and labels from thousands of events. Two innovative designs are embedded into the baseline model built on convolutional neural networks (CNNs) to improve the event classification accuracy. First, we propose a graph signal processing based PMU sorting algorithm to improve the learning efficiency of CNNs. Second, we deploy information loading based regularization to strike the right balance between memorization and generalization for the DNN. Numerical results based on real-world dataset from the Eastern Interconnection of the U.S power transmission grid show that the combination of PMU based sorting and the information loading based regularization techniques help the proposed DNN approach achieve highly accurate event identification and classification results.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Estimating Watershed Subsurface Permeability From Stream Discharge Data Using Deep Neural Networks

Subsurface permeability is a key parameter in watershed models that controls the contribution from the subsurface flow to stream flows. Since the permeability is difficult and expensive to measure directly at the spatial extent and resolution required by fully distributed watershed models, estimation through inverse modeling has had a long history in subsurface hydrology. The wide availability of stream surface flow data, compared to groundwater monitoring data, provides a new data source to infer soil and geologic properties using integrated surface and subsurface hydrologic models. As most of the existing methods have shown difficulty in dealing with highly nonlinear inverse problems, we explore the use of deep neural networks for inversion owing to their successes in mapping complex, highly nonlinear relationships. We train various deep neural network (DNN) models with different architectures to predict subsurface permeability from stream discharge hydrograph at the watershed outlet. The training data are obtained from ensemble simulations of hydrographs corresponding to an permeability ensemble using a fully-distributed, integrated surface-subsurface hydrologic model. The trained model is then applied to estimate the permeability of the real watershed using its observed hydrograph at the outlet. Our study demonstrates that the permeabilities of the soil and geologic facies that make significant contributions to the outlet discharge can be more accurately estimated from the discharge data. Their estimations are also more robust with observation errors. Compared to the traditional ensemble smoother method, DNNs show stronger performance in capturing the nonlinear relationship between permeability and stream hydrograph to accurately estimate permeability. Our study sheds new light on the value of the emerging deep learning methods in assisting integrated watershed modeling by improving parameter estimation, which will eventually reduce the uncertainty in predictive watershed models.

54 ENVIRONMENTAL SCIENCES↗

Understanding the surface wave characteristics using 2D particle-in-cell simulation and deep neural network

Here, the characteristics of the surface waves along the interface between a plasma and a dielectric material have been investigated using kinetic particle-in-cell simulations. A microwave source of GHz frequency has been used to trigger the surface wave in the system. The outcome indicates that the surface wave gets excited along the interface of plasma and the dielectric tube and appears as light and dark patterns in the electric field profiles. The dependency of radiation pressure on the dielectric permittivity and supplied input frequency has been investigated. Further, we assessed the capabilities of neural networks to predict the radiation pressure for a given system. The proposed deep neural network model is aimed at developing accurate and efficient data-driven plasma surface wave devices.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Machine Learned Hückel Theory: Interfacing Physics and Deep Neural Networks

The Hückel Hamiltonian is an incredibly simple tight-binding model known for its ability to capture qualitative physics phenomena arising from electron interactions in molecules and materials. Part of its simplicity arises from using only two types of empirically fit physics-motivated parameters: the first describes the orbital energies on each atom and the second describes electronic interactions and bonding between atoms. By replacing these empirical parameters with machine-learned dynamic values, we vastly increase the accuracy of the extended Hückel model. The dynamic values are generated with a deep neural network, which is trained to reproduce orbital energies and densities derived from density functional theory. The resulting model retains interpretability, while the deep neural network parameterization is smooth and accurate and reproduces insightful features of the original empirical parameterization. Altogether, this work shows the promise of utilizing machine learning to formulate simple, accurate, and dynamically parameterized physics models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Physics-Informed Deep Neural Networks for Learning Parameters and Constitutive Relationships in Subsurface Flow Problems

In this paper, we present a physics informed deep neural network (DNN) method for estimating parameters and unknown physics (constitutive relationships) in partial differential equation (PDE) models. We use PDEs in addition to measurements to train DNNs to approximate unknown parameters and constitutive relationships as well as states. The proposed approach increases the accuracy of DNN approximations of partially known functions when a limited number of measurements is available and allows for training DNNs when no direct measurements of the functions of interest are available. We employ physics informed DNNs to estimate the unknown space-dependent diffusion coefficient in a linear diffusion equation and an unknown constitutive relationship in a non-linear diffusion equation. For the parameter estimation problem, we assume that partial measurements of the coefficient and states are available and demonstrate that under these conditions, the proposed method is more accurate than state-of-the-art methods. For the non-linear diffusion PDE model with a fully unknown constitutive relationship (i.e., no measurements of constitutive relationship are available), the physics informed DNN method can accurately estimate the non-linear constitutive relationship based on state measurements only. Finally, we demonstrate that the proposed method remains accurate in the presence of measurement noise.

42 ENGINEERING↗

Phase retrieval for refraction-enhanced x-ray radiography using a deep neural network

X-ray refraction-enhanced radiography (RER) or phase contrast imaging is widely used to study internal discontinuities within materials. The resulting radiograph captures both the decrease in intensity caused by material absorption along the x-ray path, as well as the phase shift, which is highly sensitive to gradients in density. A significant challenge lies in effectively analyzing the radiographs to decouple the intensity and phase information and accurately ascertain the density profile. Conventional algorithms often yield ambiguous and unrealistic results due to difficulties in including physical constraints and other relevant information. We have developed an algorithm that uses a deep neural network to address these issues and applied it to extract the detailed density profile from an experimental RER. To generalize the applicability of our algorithm, we have developed a technique that quantitatively evaluates the complexity of the phase retrieval process based on the characteristics of the sample and the configuration of the experiment. Accordingly, this evaluation aids in the selection of the neural network architecture for each specific case. Beyond RER, the model has potential applications for other diagnostics where phase retrieval analysis is required.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Enhanced physics-constrained deep neural networks for modeling vanadium redox flow battery

Numerical simulation has become indispensable in advancing cost-effective process optimization and control of flow batteries. We propose an enhanced version of the physics-constrained deep neural network (PCDNN) approach to provide high-accuracy voltage predictions in the vanadium redox flow batteries (VRFBs). The purpose of the PCDNN approach is to enforce the physics-based zero-dimensional (0D) VRFB model in a neural network to assure model generalization for various battery operation conditions. However, limited by the simplifications of the 0D model, the PCDNN cannot capture sharp voltage changes in the extreme SOC regions. To improve the accuracy of voltage prediction at extreme ranges, we introduce a second (enhanced) DNN to mitigate the prediction errors carried from the 0D model itself and call the resulting approach enhanced PCDNN (ePCDNN). By comparing with experimental data, we demonstrate that the ePCDNN approach can accurately capture the voltage response throughout the charge–discharge cycle, including the tail region of the voltage discharge curve. The loss function for training the ePCDNN is designed to be flexible by adjusting the weights of the physics-constrained DNN and the enhanced DNN. In conclusion, this allows the ePCDNN framework to be transferable to battery systems with variable physical model fidelity.

25 ENERGY STORAGE↗

Quantum Perturbation Theory Using Tensor Cores and a Deep Neural Network

In this work, time-independent quantum response calculations are performed using Tensor cores. This is achieved by mapping density matrix perturbation theory onto the computational structure of a deep neural network. The main computational cost of each deep layer is dominated by tensor contractions, i.e., dense matrix–matrix multiplications, in mixed-precision arithmetics, which achieves close to peak performance. Quantum response calculations are demonstrated and analyzed using self-consistent charge density-functional tight-binding theory as well as coupled-perturbed Hartree–Fock theory. For linear response calculations, a novel parameter-free convergence criterion is presented that is well-suited for numerically noisy low-precision floating point operations and we demonstrate a peak performance of almost 200 Tflops using the Tensor cores of two Nvidia A100 GPUs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A detailed study of interpretability of deep neural network based top taggers

Abstract Recent developments in the methods of explainable artificial intelligence (XAI) allow researchers to explore the inner workings of deep neural networks (DNNs), revealing crucial information about input–output relationships and realizing how data connects with machine learning models. In this paper we explore interpretability of DNN models designed to identify jets coming from top quark decay in high energy proton–proton collisions at the Large Hadron Collider. We review a subset of existing top tagger models and explore different quantitative methods to identify which features play the most important roles in identifying the top jets. We also investigate how and why feature importance varies across different XAI metrics, how correlations among features impact their explainability, and how latent space representations encode information as well as correlate with physically meaningful quantities. Our studies uncover some major pitfalls of existing XAI methods and illustrate how they can be overcome to obtain consistent and meaningful interpretation of these models. We additionally illustrate the activity of hidden layers as neural activation pattern diagrams and demonstrate how they can be used to understand how DNNs relay information across the layers and how this understanding can help to make such models significantly simpler by allowing effective model reoptimization and hyperparameter tuning. These studies not only facilitate a methodological approach to interpreting models but also unveil new insights about what these models learn. Incorporating these observations into augmented model design, we propose the particle flow interaction network model and demonstrate how interpretability-inspired model augmentation can improve top tagging performance.

97 MATHEMATICS AND COMPUTING↗

Development of interatomic potential for Al–Tb alloys using a deep neural network learning method

An interatomic potential for the Al–Tb alloy around the composition of Al 90 Tb 10 is developed using the deep neural network (DNN) learning method. The atomic configurations and the corresponding total potential energies and forces on each atom obtained from ab initio molecular dynamics (AIMD) simulations are collected to train a DNN model to construct the interatomic potential for the Al–Tb alloy. Here we show that the obtained DNN model can well reproduce the energies and forces calculated by AIMD simulations. Molecular dynamics (MD) simulations using the DNN interatomic potential also accurately describe the structural properties of the Al 90 Tb 10 liquid, such as partial pair correlation functions (PPCFs) and bond angle distributions, in comparison with the results from AIMD simulations. Furthermore, the developed DNN interatomic potential predicts the formation energies of the crystalline phases of the Al–Tb system with an accuracy comparable to ab initio calculations. The structure factors of the Al 90 Tb 10 metallic liquid and glass obtained by MD simulations using the developed DNN interatomic potential are also in good agreement with the experimental X-ray diffraction data. The development of short-range order (SRO) in the Al 90 Tb 10 liquid and the undercooled liquid is also analyzed and three dominant SROs, i.e., Al-centered distorted icosahedron (DISICO) and Tb-centered ‘3661’ and ‘15551’ clusters, respectively, are identified.

36 MATERIALS SCIENCE↗

Learning viscoelasticity models from indirect data using deep neural networks

In this study, we propose a novel approach to model viscoelasticity materials, where rate-dependent and non-linear constitutive relationships are approximated with deep neural networks. We assume that inputs and outputs of the neural networks are not directly observable, and therefore common training techniques with input–output pairs for the neural networks are inapplicable. To that end, we develop a novel computational approach to both calibrate parametric and learn neural-network-based constitutive relations of viscoelasticity materials from indirect displacement data in the context of multiple-physics systems. We show that limited displacement data holds sufficient information to quantify the viscoelasticity behavior. We formulate the inverse computation – modeling viscoelasticity properties from observed displacement data – as a PDE-constrained optimization problem and minimize the error functional using a gradient-based optimization method. The gradients are computed by a combination of automatic differentiation and implicit function differentiation rules. The effectiveness of our method is demonstrated through numerous benchmark problems in geomechanics and porous media transport.

97 MATHEMATICS AND COMPUTING↗

Factorized visual representations in the primate visual system and deep neural networks

Object classification has been proposed as a principal objective of the primate ventral visual stream and has been used as an optimization target for deep neural network models (DNNs) of the visual system. However, visual brain areas represent many different types of information, and optimizing for classification of object identity alone does not constrain how other information may be encoded in visual representations. Information about different scene parameters may be discarded altogether (‘invariance’), represented in non-interfering subspaces of population activity (‘factorization’) or encoded in an entangled fashion. In this work, we provide evidence that factorization is a normative principle of biological visual representations. In the monkey ventral visual hierarchy, we found that factorization of object pose and background information from object identity increased in higher-level regions and strongly contributed to improving object identity decoding performance. We then conducted a large-scale analysis of factorization of individual scene parameters – lighting, background, camera viewpoint, and object pose – in a diverse library of DNN models of the visual system. Models which best matched neural, fMRI, and behavioral data from both monkeys and humans across 12 datasets tended to be those which factorized scene parameters most strongly. Notably, invariance to these parameters was not as consistently associated with matches to neural and behavioral data, suggesting that maintaining non-class information in factorized activity subspaces is often preferred to dropping it altogether. Thus, we propose that factorization of visual scene information is a widely used strategy in brains and DNN models thereof.

59 BASIC BIOLOGICAL SCIENCES↗

A Deep Neural Network for Accurate and Robust Prediction of the Glass Transition Temperature of Polyhydroxyalkanoate Homo- and Copolymers

The purpose of this study was to develop a data-driven machine learning model to predict the performance properties of polyhydroxyalkanoates (PHAs), a group of biosourced polyesters featuring excellent performance, to guide future design and synthesis experiments. A deep neural network (DNN) machine learning model was built for predicting the glass transition temperature, Tg, of PHA homo- and copolymers. Molecular fingerprints were used to capture the structural and atomic information of PHA monomers. The other input variables included the molecular weight, the polydispersity index, and the percentage of each monomer in the homo- and copolymers. The results indicate that the DNN model achieves high accuracy in estimation of the glass transition temperature of PHAs. In addition, the symmetry of the DNN model is ensured by incorporating symmetry data in the training process. The DNN model achieved better performance than the support vector machine (SVD), a nonlinear ML model and least absolute shrinkage and selection operator (LASSO), a sparse linear regression model. The relative importance of factors affecting the DNN model prediction were analyzed. Sensitivity of the DNN model, including strategies to deal with missing data, were also investigated. Compared with commonly used machine learning models incorporating quantitative structure–property (QSPR) relationships, it does not require an explicit descriptor selection step but shows a comparable performance. The machine learning model framework can be readily extended to predict other properties.

quantitative structure–property relationship (QSPR↗