Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “deep learning method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Load Profile Inpainting for Missing Load Data Restoration and Baseline Estimation

This paper introduces a Generative Adversarial Nets (GAN) based, Load Profile Inpainting Network (Load-PIN) for restoring missing load data segments and estimating the baseline for a demand response event. The inputs are time series load data before and after the inpainting period together with explanatory variables (e.g., weather data). Here, we propose a Generator structure consisting of a coarse network and a fine-tuning network. The coarse network provides an initial estimation of the data segment in the inpainting period. The fine-tuning network consists of self-attention blocks and gated convolution layers for adjusting the initial estimations. Loss functions are specially designed for the fine-tuning and the discriminator networks to enhance both the point-to-point accuracy and realisticness of the results. We test the Load-PIN on three real-world data sets for two applications: patching missing data and deriving baselines of conservation voltage reduction (CVR) events. We benchmark the performance of Load-PIN with five existing deep-learning methods. Our simulation results show that, compared with the state-of-the-art methods, Load-PIN can handle varying-length missing data events and achieve 15-30% accuracy improvement.

14 SOLAR ENERGY↗

Development of interatomic potential for Al–Tb alloys using a deep neural network learning method

An interatomic potential for the Al–Tb alloy around the composition of Al 90 Tb 10 is developed using the deep neural network (DNN) learning method. The atomic configurations and the corresponding total potential energies and forces on each atom obtained from ab initio molecular dynamics (AIMD) simulations are collected to train a DNN model to construct the interatomic potential for the Al–Tb alloy. Here we show that the obtained DNN model can well reproduce the energies and forces calculated by AIMD simulations. Molecular dynamics (MD) simulations using the DNN interatomic potential also accurately describe the structural properties of the Al 90 Tb 10 liquid, such as partial pair correlation functions (PPCFs) and bond angle distributions, in comparison with the results from AIMD simulations. Furthermore, the developed DNN interatomic potential predicts the formation energies of the crystalline phases of the Al–Tb system with an accuracy comparable to ab initio calculations. The structure factors of the Al 90 Tb 10 metallic liquid and glass obtained by MD simulations using the developed DNN interatomic potential are also in good agreement with the experimental X-ray diffraction data. The development of short-range order (SRO) in the Al 90 Tb 10 liquid and the undercooled liquid is also analyzed and three dominant SROs, i.e., Al-centered distorted icosahedron (DISICO) and Tb-centered ‘3661’ and ‘15551’ clusters, respectively, are identified.

36 MATERIALS SCIENCE↗

Machine learning the spectral function of a hole in a quantum antiferromagnet

Understanding charge motion in a background of interacting quantum spins is a fundamental problem in quantum many-body physics. The most extensively studied model for this problem is the so-called t-t'-t''-J model, where the determination of the parameter t' in the context of cuprate superconductors is challenging. Here we present a theoretical study of the spectral functions of a mobile hole in the t-t'-t''-J model using two machine-learning techniques: K-nearest neighbor regression (KNN) and a feed-forward neural network (FFNN). We employ the self-consistent Born approximation to generate a dataset of about 1.3 x 10 5 spectral functions. Here we show that, for the forward problem, both methods allow for the accurate and efficient prediction of spectral functions, allowing, e.g., rapid searches through parameter space. Furthermore, we find that for the inverse problem (inferring Hamiltonian parameters from spectra), the FFNN can, but the KNN cannot, accurately predict the model parameters using merely the density of states. Our results suggest that it may be possible to use deep-learning methods to predict materials parameters from experimentally measured spectral functions.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Mitigation of spatial nonstationarity with vision transformers

Spatial nonstationarity, the location variance of features’ statistical distributions, is ubiquitous in many natural settings. For example, in geological reservoirs rock matrix porosity varies vertically due to geomechanical compaction trends, in mineral deposits grades vary due to sedimentation and concentration processes, in hydrology rainfall varies due to the atmosphere and topography interactions, and in metallurgy crystalline structures vary due to differential cooling. Conventional geostatistical modeling workflows rely on the assumption of stationarity to be able to model spatial features for geostatistical inference. Nevertheless, this is often not a realistic assumption when dealing with nonstationary spatial data and this has motivated a variety of nonstationary spatial modeling workflows such as trend and residual decomposition, cosimulation with secondary features, and spatial segmentation and independent modeling over stationary subdomains. The advent of deep learning technologies has enabled new workflows for modeling spatial relationships. However, there is a paucity of demonstrated best practice and general guidance on mitigation of spatial nonstationarity with deep learning in the geospatial context. We demonstrate the impact of two common types of geostatistical spatial nonstationarity on deep learning model prediction performance and propose the mitigation of such impacts using self-attention (vision transformer) models. We demonstrate the utility of vision transformers for the mitigation of nonstationarity with relative errors as low as 10%, exceeding the performance of alternative deep learning methods such as convolutional neural networks. We establish best practice by demonstrating the ability of self-attention networks for modeling large-scale spatial relationships in the presence of commonly observed geospatial nonstationarity.

58 GEOSCIENCES↗

DNCON2_Inter: predicting interchain contacts for homodimeric and homomultimeric protein complexes using multiple sequence alignments of monomers and deep learning

Deep learning methods that achieved great success in predicting intrachain residue-residue contacts have been applied to predict interchain contacts between proteins. However, these methods require multiple sequence alignments (MSAs) of a pair of interacting proteins (dimers) as input, which are often difficult to obtain because there are not many known protein complexes available to generate MSAs of sufficient depth for a pair of proteins. In recognizing that multiple sequence alignments of a monomer that forms homomultimers contain the co-evolutionary signals of both intrachain and interchain residue pairs in contact, we applied DNCON2 (a deep learning-based protein intrachain residue-residue contact predictor) to predict both intrachain and interchain contacts for homomultimers using multiple sequence alignment (MSA) and other co-evolutionary features of a single monomer followed by discrimination of interchain and intrachain contacts according to the tertiary structure of the monomer. We name this tool DNCON2_Inter. Allowing true-positive predictions within two residue shifts, the best average precision was obtained for the Top-L/10 predictions of 22.9% for homodimers and 17.0% for higher-order homomultimers. In some instances, especially where interchain contact densities are high, DNCON2_Inter predicted interchain contacts with 100% precision. We also developed Con_Complex, a complex structure reconstruction tool that uses predicted contacts to produce the structure of the complex. Using Con_Complex, we show that the predicted contacts can be used to accurately construct the structure of some complexes. Our experiment demonstrates that monomeric multiple sequence alignments can be used with deep learning to predict interchain contacts of homomeric proteins.

59 BASIC BIOLOGICAL SCIENCES↗

Exact enforcement of temporal continuity in sequential physics-informed neural networks

The use of deep learning methods in scientific computing represents a potential paradigm shift in engineering problem solving. One of the most prominent developments is Physics-Informed Neural Networks (PINNs), in which neural networks are trained to satisfy partial differential equations (PDEs). While this method shows promise, the standard version has been shown to struggle in accurately predicting the dynamic behavior of time-dependent problems. To address this challenge, methods have been proposed that decompose the time domain into multiple segments, employing a distinct neural network in each segment and directly incorporating continuity between them in the loss function of the minimization problem. In this work we introduce a method to exactly enforce continuity between successive time segments via a solution ansatz. This hard constrained sequential PINN (HCS-PINN) method is simple to implement and eliminates the need for any loss terms associated with temporal continuity. The method is tested for a number of benchmark problems involving both linear and non-linear PDEs. Examples include various first order time dependent problems in which traditional PINNs struggle, namely advection, Allen–Cahn, and Korteweg–de Vries equations. Furthermore, second and third order time-dependent problems are demonstrated via wave and Jerky dynamics examples, respectively. Notably, the Jerky dynamics problem is chaotic, making the problem especially sensitive to temporal accuracy. Finally, the numerical experiments conducted with the proposed method demonstrated superior convergence and accuracy over both traditional PINNs and the soft-constrained counterparts.

42 ENGINEERING↗

Probabilistic neural networks for improved analyses with phenomenological R -matrix

Here we present a method for measurement analyses based on probabilistic deep neural networks that provide several advantages over conventional analyses with phenomenological models. These include predicting physical quantities directly from data, the rapid generation of statistically robust uncertainties, and the ability to bypass some parameters that may induce ambiguities and complications in data analysis. As deep learning methods make predictions through “black boxes,” the uncertainty quantification is typically challenging. We use a probabilistic framework that provides thorough uncertainty quantification and is straightforward to follow in practice. With the network architecture based on the Transformer, we demonstrate the current method for predicting nuclear resonance parameters from scattering data using the phenomenological R-matrix model.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A Framework for Deep Learning Emulation of Numerical Models With a Case Study in Satellite Remote Sensing

Numerical models based on physics represent the state of the art in Earth system modeling and comprise our best tools for generating insights and predictions. Despite rapid growth in computational power, the perceived need for higher model resolutions overwhelms the latest generation computers, reducing the ability of modelers to generate simulations for understanding parameter sensitivities and characterizing variability and uncertainty. Thus, surrogate models are often developed to capture the essential attributes of the full-blown numerical models. Recent successes of machine learning methods, especially deep learning (DL), across many disciplines offer the possibility that complex nonlinear connectionist representations may be able to capture the underlying complex structures and nonlinear processes in Earth systems. A difficult test for DL-based emulation, which refers to function approximation of numerical models, is to understand whether they can be comparable to traditional forms of surrogate models in terms of computational efficiency while simultaneously reproducing model results in a credible manner. A DL emulation that passes this test may be expected to perform even better than simple models with respect to capturing complex processes and spatiotemporal dependencies. Here, we examine, with a case study in satellite-based remote sensing, the hypothesis that DL approaches can credibly represent the simulations from a surrogate model with comparable computational efficiency. Our results are encouraging in that the DL emulation reproduces the results with acceptable accuracy and often even faster performance. We discuss the broader implications of our results in light of the pace of improvements in high-performance implementations of DL and the growing desire for higher resolution simulations in the Earth sciences.

Bayesian Deep Learning↗

HDBind: encoding of molecular structure with hyperdimensional binary representations

Traditional methods for identifying “hit” molecules from a large collection of potential drug-like candidates rely on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have a significant limitation in that they require exceptional computing capabilities for even relatively small collections of molecules. Increasingly large and complex state-of-the-art deep learning approaches have gained popularity with the promise to improve the productivity of drug design, notorious for its numerous failures. However, as deep learning models increase in their size and complexity, their acceleration at the hardware level becomes more challenging. Hyperdimensional Computing (HDC) has recently gained attention in the computer hardware community due to its algorithmic simplicity relative to deep learning approaches. The HDC learning paradigm, which represents data with high-dimension binary vectors, allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas (computer vision, bioinformatics, mass spectrometery, remote sensing, edge devices, etc.). To the best of our knowledge, our work is the first to consider HDC for the task of fast and efficient screening of modern drug-like compound libraries. We also propose the first HDC graph-based encoding methods for molecular data, demonstrating consistent and substantial improvement over previous work. We compare our approaches to alternative approaches on the well-studied MoleculeNet dataset and the recently proposed LIT-PCBA dataset derived from high quality PubChem assays. We demonstrate our methods on multiple target hardware platforms, including Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs), showing at least an order of magnitude improvement in energy efficiency versus even our smallest neural network baseline model with a single hidden layer. Our work thus motivates further investigation into molecular representation learning to develop ultra-efficient pre-screening tools. We make our code publicly available at https://github.com/LLNL/hdbind.

59 BASIC BIOLOGICAL SCIENCES↗

Regional Medium-Term Hourly Electricity Demand Forecasting Based on LSTM

This paper aims to forecast high-resolution (hourly) aggregated load for a certain region in the medium term (a few days to over a year). One region is defined as some places with similar climate characteristics because the climate influences people's daily lifestyles and hence the electric usage. We decom- pose the electric usage records into two parts: base load and seasonal load. Considering both temperature and time factors, different deep learning methods are adopted to characterize them. The first goal of our approach is to predict the peak load which is critical for power system planning. Furthermore, our proposed forecast method can provide the depiction of the hourly load profile to provide customized load curves for high- level real-time applications. The proposed method is tested on real-world historical data collected by CAISO, BPA, and PACW. The experimental results show that trained by three years of data, our method could reduce the prediction error for one-year lead hourly load below 5% MAPE, and predict the occurrence of the peak load for next year in CAISO with an error within three days. Furthermore, as a byproduct, an interesting observation on the impact of COVID-19 on human life was made and discussed based on these case studies.

deep learning↗

Regional Medium-Term Hourly Electricity Demand Forecasting Based on LSTM

This paper aims to forecast high-resolution (hourly) aggregated load for a certain region in the medium term (a few days to over a year). One region is defined as some places with similar climate characteristics because the climate influences people's daily lifestyles and hence the electric usage. We decompose the electric usage records into two parts: base load and seasonal load. Considering both temperature and time factors, different deep-learning methods are adopted to characterize them. The first goal of our approach is to predict the peak load which is critical for power system planning. Furthermore, our proposed forecast method can provide the depiction of the hourly load profile to provide customized load curves for high-level real-time applications. The proposed method is tested on real-world historical data collected by CAISO, BPA, and PACW. The experimental results show that trained by three years of data, our method could reduce the prediction error for a one-year lead hourly load below $5\%$ MAPE, and predict the occurrence of the peak load for next year in CAISO with an error within three days. Furthermore, as a byproduct, an interesting observation on the impact of COVID-19 on human life was made and discussed based on these case studies.

deep learning↗

Regional Medium-Term Hourly Electricity Demand Forecasting Based on LSTM: Preprint

This paper aims to forecast high-resolution (hourly) aggregated load for a certain region in the medium term (a few days to over a year). One region is defined as some places with similar climate characteristics because the climate influences people's daily lifestyles and hence the electric usage. We decompose the electric usage records into two parts: base load and seasonal load. Considering both temperature and time factors, different deep-learning methods are adopted to characterize them. The first goal of our approach is to predict the peak load which is critical for power system planning. Furthermore, our proposed forecast method can provide the depiction of the hourly load profile to provide customized load curves for high-level real-time applications. The proposed method is tested on real-world historical data collected by CAISO, BPA, and PACW. The experimental results show that trained by three years of data, our method could reduce the prediction error for a one-year lead hourly load below 5% MAPE, and predict the occurrence of the peak load for next year in CAISO with an error within three days. Furthermore, as a byproduct, an interesting observation on the impact of COVID-19 on human life was made and discussed based on these case studies.

deep learning↗

Approximating Nash Equilibrium in Day-ahead Electricity Market Bidding with Multi-agent Deep Reinforcement Learning

In this paper, a day-ahead electricity market bidding problem with multiple strategic generation company (GEN-CO) bidders is studied. The problem is formulated as a Markov game model, where GENCO bidders interact with each other todevelop their optimal day-ahead bidding strategies. Considering unobservable information in the problem, a model-free and data-driven approach, known as multi-agent deep deterministic policy gradient (MADDPG), is applied for approximating the Nash equilibrium (NE) in the above Markov game. The MADDPG algorithm has the advantage of generalization due to the automatic feature extraction ability of the deep neural networks. The algorithm is tested on an IEEE 30-bus system with three competitive GENCO bidders in both an uncongested caseand a congested case. Comparisons with a truthful bidding strategy and state-of-the-art deep reinforcement learning methods including deep Q network and deep deterministic policy gradient (DDPG) demonstrate that the applied MADDPG algorithm can find a superior bidding strategy for all the market participants with increased profit gains. In addition, the comparison with a conventional model-based method shows that the MADDPG algorithm has higher computational efficiency, which is feasible for real-world applications.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Three-Dimensional Analysis of Particle Distribution on Filter Layers inside N95 Respirators by Deep Learning

The global COVID-19 pandemic has changed many aspects of daily lives. Wearing personal protective equipment, especially respirators (face masks), has become common for both the public and medical professionals, proving to be effective in preventing spread of the virus. Nevertheless, a detailed understanding of respirator filtration-layer internal structures and their physical configurations is lacking. Here, we report three-dimensional (3D) internal analysis of N95 filtration layers via X-ray tomography. Using deep learning methods, we uncover how the distribution and diameters of fibers within these layers directly affect contaminant particle filtration. The average porosity of the filter layers is found to be 89.1%. Contaminants are more efficiently captured by denser fiber regions, with fibers <1.8 μm in diameter being particularly effective, presumably because of the stronger electric field gradient on smaller diameter fibers. This study provides critical information for further development of N95-type respirators that combine high efficiency with good breathability.

42 ENGINEERING↗

Precision Polishing of Ablator Capsules via in situ Process Monitoring and Machine Learning–Based Optimization

In inertial confinement fusion (ICF) experiments seeking output gains of unity and beyond, the quality of the ablator capsule is paramount for minimizing the hydrodynamic mix that quenches the central hot spot. Defects in the form of foreign particles or missing mass on the surface and within the wall of the capsule are primary offenders. High-density carbon capsules made for ICF experiments at the National Ignition Facility are precision polished to achieve surface smoothness on the order of a few nanometers as well as to minimize isolated defects in the form of pits. Given the critical role of this process, we are developing smart manufacturing techniques with the goal of elevating the efficiency of this process. Our approach is to use MEMS (micro-electromechanical systems)–based sensors to capture the fine vibration signals generated during the polishing process and combine them with synchronized visual feedback as needed. Beyond using these sensors for process monitoring, we use specific deep learning methods to analyze the data and extract correlations with both the process parameters and the final performance of the polishing run. Here, in this work, we describe the multiple fronts we have explored in this regard and the results we have gotten so far. This approach promises to have the potential to ultimately provide real-time feedback that can be used to ensure the progress of the run as well as a means for faster optimization.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Latent-space time evolution of non-intrusive reduced-order models using Gaussian process emulation

Non-intrusive reduced-order models (ROMs) have recently generated considerable interest for constructing computationally efficient counterparts of nonlinear dynamical systems emerging from various domain sciences. They provide a low-dimensional emulation framework for systems that may be intrinsically high-dimensional. This is accomplished by utilizing a construction algorithm that is purely data-driven. It is no surprise, therefore, that the algorithmic advances of machine learning have led to non-intrusive ROMs with greater accuracy and computational gains. However, in bypassing the utilization of an equation-based evolution, it is often seen that the interpretability of the ROM framework suffers. This becomes more problematic when black-box deep learning methods are used which are notorious for lacking robustness outside the physical regime of the observed data. In this article, we propose the use of a novel latent-space interpolation algorithm based on Gaussian process regression. Notably, this reduced-order evolution of the system is parameterized by control parameters to allow for interpolation in space. The use of this procedure also allows for a continuous interpretation of time which allows for temporal interpolation. The latter aspect provides information, with quantified uncertainty, about full-state evolution at a finer resolution than that utilized for training the ROMs. This research assesses the viability of this algorithm for an advection-dominated system given by the inviscid shallow water equations.

97 MATHEMATICS AND COMPUTING↗

LSTM-Based Data Integration to Improve Snow Water Equivalent Prediction and Diagnose Error Sources

Accurate prediction of snow water equivalent (SWE) can be valuable for water resource managers. Recently, deep learning methods such as long short-term memory (LSTM) have exhibited high accuracy in simulating hydrologic variables and can integrate lagged observations to improve prediction, but their benefits were not clear for SWE simulations. Here we tested an LSTM network with data integration (DI) for SWE in the western United States to integrate 30-day-lagged or 7-day-lagged observations of either SWE or satellite-observed snow cover fraction (SCF) to improve future predictions. SCF proved beneficial only for shallow-snow sites during snowmelt, while lagged SWE integration significantly improved prediction accuracy for both shallow- and deep-snow sites. The median Nash–Sutcliffe model efficiency coefficient (NSE) in temporal testing improved from 0.92 to 0.97 with 30-day-lagged SWE integration, and root-mean-square error (RMSE) and the difference between estimated and observed peak SWE values d max were reduced by 41% and 57%, respectively. DI effectively mitigated accumulated model and forcing errors that would otherwise be persistent. Moreover, by applying DI to different observations (30-day-lagged, 7-day-lagged), we revealed the spatial distribution of errors with different persistent lengths. For example, integrating 30-day-lagged SWE was ineffective for ephemeral snow sites in the southwestern United States, but significantly reduced monthly-scale biases for regions with stable seasonal snowpack such as high-elevation sites in California. These biases are likely attributable to large interannual variability in snowfall or site-specific snow redistribution patterns that can accumulate to impactful levels over time for nonephemeral sites. These results set up benchmark levels and provide guidance for future model improvement strategies.

54 ENVIRONMENTAL SCIENCES↗

Reduced volume of diabetic pancreatic islets in rodents detected by synchrotron X-ray phase-contrast microtomography and deep learning network

The pancreatic islet is a highly structured micro-organ that produces insulin in response to rising blood glucose. Here we develop a label-free and automatic imaging approach to visualize the islets in situ in diabetic rodents by the synchrotron radiation X-ray phase-contrast microtomography (SRμCT) at the ID17 station of the European Synchrotron Radiation Facility. The large-size images (3.2 mm × 15.97 mm) were acquired in the pancreas in STZ-treated mice and diabetic GK rats. Each pancreas was dissected by 3000 reconstructed images. The image datasets were further analysed by a self-developed deep learning method, AA-Net. All islets in the pancreas were segmented and visualized by the three-dimension (3D) reconstruction. After quantifying the volumes of the islets, we found that the number of larger islets (=>1500 μm 3 ) was reduced by 2-fold (wt 1004 ± 94 vs GK 419 ± 122, P < 0.001) in chronically developed diabetic GK rat, while in STZ-treated diabetic mouse the large islets were decreased by half (189 ± 33 vs 90 ± 29, P < 0.001) compared to the untreated mice. Our study provides a label-free tool for detecting and quantifying pancreatic islets in situ. It implies the possibility of monitoring the state of pancreatic islets in vivo diabetes without labelling.

59 BASIC BIOLOGICAL SCIENCES↗