Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “deep learning methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A Framework for Deep Learning Emulation of Numerical Models With a Case Study in Satellite Remote Sensing

Numerical models based on physics represent the state of the art in Earth system modeling and comprise our best tools for generating insights and predictions. Despite rapid growth in computational power, the perceived need for higher model resolutions overwhelms the latest generation computers, reducing the ability of modelers to generate simulations for understanding parameter sensitivities and characterizing variability and uncertainty. Thus, surrogate models are often developed to capture the essential attributes of the full-blown numerical models. Recent successes of machine learning methods, especially deep learning (DL), across many disciplines offer the possibility that complex nonlinear connectionist representations may be able to capture the underlying complex structures and nonlinear processes in Earth systems. A difficult test for DL-based emulation, which refers to function approximation of numerical models, is to understand whether they can be comparable to traditional forms of surrogate models in terms of computational efficiency while simultaneously reproducing model results in a credible manner. A DL emulation that passes this test may be expected to perform even better than simple models with respect to capturing complex processes and spatiotemporal dependencies. Here, we examine, with a case study in satellite-based remote sensing, the hypothesis that DL approaches can credibly represent the simulations from a surrogate model with comparable computational efficiency. Our results are encouraging in that the DL emulation reproduces the results with acceptable accuracy and often even faster performance. We discuss the broader implications of our results in light of the pace of improvements in high-performance implementations of DL and the growing desire for higher resolution simulations in the Earth sciences.

Bayesian Deep Learning↗

HDBind: encoding of molecular structure with hyperdimensional binary representations

Traditional methods for identifying “hit” molecules from a large collection of potential drug-like candidates rely on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have a significant limitation in that they require exceptional computing capabilities for even relatively small collections of molecules. Increasingly large and complex state-of-the-art deep learning approaches have gained popularity with the promise to improve the productivity of drug design, notorious for its numerous failures. However, as deep learning models increase in their size and complexity, their acceleration at the hardware level becomes more challenging. Hyperdimensional Computing (HDC) has recently gained attention in the computer hardware community due to its algorithmic simplicity relative to deep learning approaches. The HDC learning paradigm, which represents data with high-dimension binary vectors, allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas (computer vision, bioinformatics, mass spectrometery, remote sensing, edge devices, etc.). To the best of our knowledge, our work is the first to consider HDC for the task of fast and efficient screening of modern drug-like compound libraries. We also propose the first HDC graph-based encoding methods for molecular data, demonstrating consistent and substantial improvement over previous work. We compare our approaches to alternative approaches on the well-studied MoleculeNet dataset and the recently proposed LIT-PCBA dataset derived from high quality PubChem assays. We demonstrate our methods on multiple target hardware platforms, including Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs), showing at least an order of magnitude improvement in energy efficiency versus even our smallest neural network baseline model with a single hidden layer. Our work thus motivates further investigation into molecular representation learning to develop ultra-efficient pre-screening tools. We make our code publicly available at https://github.com/LLNL/hdbind.

59 BASIC BIOLOGICAL SCIENCES↗

Regional Medium-Term Hourly Electricity Demand Forecasting Based on LSTM

This paper aims to forecast high-resolution (hourly) aggregated load for a certain region in the medium term (a few days to over a year). One region is defined as some places with similar climate characteristics because the climate influences people's daily lifestyles and hence the electric usage. We decom- pose the electric usage records into two parts: base load and seasonal load. Considering both temperature and time factors, different deep learning methods are adopted to characterize them. The first goal of our approach is to predict the peak load which is critical for power system planning. Furthermore, our proposed forecast method can provide the depiction of the hourly load profile to provide customized load curves for high- level real-time applications. The proposed method is tested on real-world historical data collected by CAISO, BPA, and PACW. The experimental results show that trained by three years of data, our method could reduce the prediction error for one-year lead hourly load below 5% MAPE, and predict the occurrence of the peak load for next year in CAISO with an error within three days. Furthermore, as a byproduct, an interesting observation on the impact of COVID-19 on human life was made and discussed based on these case studies.

deep learning↗

Regional Medium-Term Hourly Electricity Demand Forecasting Based on LSTM

This paper aims to forecast high-resolution (hourly) aggregated load for a certain region in the medium term (a few days to over a year). One region is defined as some places with similar climate characteristics because the climate influences people's daily lifestyles and hence the electric usage. We decompose the electric usage records into two parts: base load and seasonal load. Considering both temperature and time factors, different deep-learning methods are adopted to characterize them. The first goal of our approach is to predict the peak load which is critical for power system planning. Furthermore, our proposed forecast method can provide the depiction of the hourly load profile to provide customized load curves for high-level real-time applications. The proposed method is tested on real-world historical data collected by CAISO, BPA, and PACW. The experimental results show that trained by three years of data, our method could reduce the prediction error for a one-year lead hourly load below $5\%$ MAPE, and predict the occurrence of the peak load for next year in CAISO with an error within three days. Furthermore, as a byproduct, an interesting observation on the impact of COVID-19 on human life was made and discussed based on these case studies.

deep learning↗

Regional Medium-Term Hourly Electricity Demand Forecasting Based on LSTM: Preprint

This paper aims to forecast high-resolution (hourly) aggregated load for a certain region in the medium term (a few days to over a year). One region is defined as some places with similar climate characteristics because the climate influences people's daily lifestyles and hence the electric usage. We decompose the electric usage records into two parts: base load and seasonal load. Considering both temperature and time factors, different deep-learning methods are adopted to characterize them. The first goal of our approach is to predict the peak load which is critical for power system planning. Furthermore, our proposed forecast method can provide the depiction of the hourly load profile to provide customized load curves for high-level real-time applications. The proposed method is tested on real-world historical data collected by CAISO, BPA, and PACW. The experimental results show that trained by three years of data, our method could reduce the prediction error for a one-year lead hourly load below 5% MAPE, and predict the occurrence of the peak load for next year in CAISO with an error within three days. Furthermore, as a byproduct, an interesting observation on the impact of COVID-19 on human life was made and discussed based on these case studies.

deep learning↗

Approximating Nash Equilibrium in Day-ahead Electricity Market Bidding with Multi-agent Deep Reinforcement Learning

In this paper, a day-ahead electricity market bidding problem with multiple strategic generation company (GEN-CO) bidders is studied. The problem is formulated as a Markov game model, where GENCO bidders interact with each other todevelop their optimal day-ahead bidding strategies. Considering unobservable information in the problem, a model-free and data-driven approach, known as multi-agent deep deterministic policy gradient (MADDPG), is applied for approximating the Nash equilibrium (NE) in the above Markov game. The MADDPG algorithm has the advantage of generalization due to the automatic feature extraction ability of the deep neural networks. The algorithm is tested on an IEEE 30-bus system with three competitive GENCO bidders in both an uncongested caseand a congested case. Comparisons with a truthful bidding strategy and state-of-the-art deep reinforcement learning methods including deep Q network and deep deterministic policy gradient (DDPG) demonstrate that the applied MADDPG algorithm can find a superior bidding strategy for all the market participants with increased profit gains. In addition, the comparison with a conventional model-based method shows that the MADDPG algorithm has higher computational efficiency, which is feasible for real-world applications.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Precision Polishing of Ablator Capsules via in situ Process Monitoring and Machine Learning–Based Optimization

In inertial confinement fusion (ICF) experiments seeking output gains of unity and beyond, the quality of the ablator capsule is paramount for minimizing the hydrodynamic mix that quenches the central hot spot. Defects in the form of foreign particles or missing mass on the surface and within the wall of the capsule are primary offenders. High-density carbon capsules made for ICF experiments at the National Ignition Facility are precision polished to achieve surface smoothness on the order of a few nanometers as well as to minimize isolated defects in the form of pits. Given the critical role of this process, we are developing smart manufacturing techniques with the goal of elevating the efficiency of this process. Our approach is to use MEMS (micro-electromechanical systems)–based sensors to capture the fine vibration signals generated during the polishing process and combine them with synchronized visual feedback as needed. Beyond using these sensors for process monitoring, we use specific deep learning methods to analyze the data and extract correlations with both the process parameters and the final performance of the polishing run. Here, in this work, we describe the multiple fronts we have explored in this regard and the results we have gotten so far. This approach promises to have the potential to ultimately provide real-time feedback that can be used to ensure the progress of the run as well as a means for faster optimization.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Latent-space time evolution of non-intrusive reduced-order models using Gaussian process emulation

Non-intrusive reduced-order models (ROMs) have recently generated considerable interest for constructing computationally efficient counterparts of nonlinear dynamical systems emerging from various domain sciences. They provide a low-dimensional emulation framework for systems that may be intrinsically high-dimensional. This is accomplished by utilizing a construction algorithm that is purely data-driven. It is no surprise, therefore, that the algorithmic advances of machine learning have led to non-intrusive ROMs with greater accuracy and computational gains. However, in bypassing the utilization of an equation-based evolution, it is often seen that the interpretability of the ROM framework suffers. This becomes more problematic when black-box deep learning methods are used which are notorious for lacking robustness outside the physical regime of the observed data. In this article, we propose the use of a novel latent-space interpolation algorithm based on Gaussian process regression. Notably, this reduced-order evolution of the system is parameterized by control parameters to allow for interpolation in space. The use of this procedure also allows for a continuous interpretation of time which allows for temporal interpolation. The latter aspect provides information, with quantified uncertainty, about full-state evolution at a finer resolution than that utilized for training the ROMs. This research assesses the viability of this algorithm for an advection-dominated system given by the inviscid shallow water equations.

97 MATHEMATICS AND COMPUTING↗

LSTM-Based Data Integration to Improve Snow Water Equivalent Prediction and Diagnose Error Sources

Accurate prediction of snow water equivalent (SWE) can be valuable for water resource managers. Recently, deep learning methods such as long short-term memory (LSTM) have exhibited high accuracy in simulating hydrologic variables and can integrate lagged observations to improve prediction, but their benefits were not clear for SWE simulations. Here we tested an LSTM network with data integration (DI) for SWE in the western United States to integrate 30-day-lagged or 7-day-lagged observations of either SWE or satellite-observed snow cover fraction (SCF) to improve future predictions. SCF proved beneficial only for shallow-snow sites during snowmelt, while lagged SWE integration significantly improved prediction accuracy for both shallow- and deep-snow sites. The median Nash–Sutcliffe model efficiency coefficient (NSE) in temporal testing improved from 0.92 to 0.97 with 30-day-lagged SWE integration, and root-mean-square error (RMSE) and the difference between estimated and observed peak SWE values d max were reduced by 41% and 57%, respectively. DI effectively mitigated accumulated model and forcing errors that would otherwise be persistent. Moreover, by applying DI to different observations (30-day-lagged, 7-day-lagged), we revealed the spatial distribution of errors with different persistent lengths. For example, integrating 30-day-lagged SWE was ineffective for ephemeral snow sites in the southwestern United States, but significantly reduced monthly-scale biases for regions with stable seasonal snowpack such as high-elevation sites in California. These biases are likely attributable to large interannual variability in snowfall or site-specific snow redistribution patterns that can accumulate to impactful levels over time for nonephemeral sites. These results set up benchmark levels and provide guidance for future model improvement strategies.

54 ENVIRONMENTAL SCIENCES↗

Reduced volume of diabetic pancreatic islets in rodents detected by synchrotron X-ray phase-contrast microtomography and deep learning network

The pancreatic islet is a highly structured micro-organ that produces insulin in response to rising blood glucose. Here we develop a label-free and automatic imaging approach to visualize the islets in situ in diabetic rodents by the synchrotron radiation X-ray phase-contrast microtomography (SRμCT) at the ID17 station of the European Synchrotron Radiation Facility. The large-size images (3.2 mm × 15.97 mm) were acquired in the pancreas in STZ-treated mice and diabetic GK rats. Each pancreas was dissected by 3000 reconstructed images. The image datasets were further analysed by a self-developed deep learning method, AA-Net. All islets in the pancreas were segmented and visualized by the three-dimension (3D) reconstruction. After quantifying the volumes of the islets, we found that the number of larger islets (=>1500 μm 3 ) was reduced by 2-fold (wt 1004 ± 94 vs GK 419 ± 122, P < 0.001) in chronically developed diabetic GK rat, while in STZ-treated diabetic mouse the large islets were decreased by half (189 ± 33 vs 90 ± 29, P < 0.001) compared to the untreated mice. Our study provides a label-free tool for detecting and quantifying pancreatic islets in situ. It implies the possibility of monitoring the state of pancreatic islets in vivo diabetes without labelling.

59 BASIC BIOLOGICAL SCIENCES↗

HD-Bind: Encoding of Molecular Structure with Low Precision, Hyperdimensional Binary Representations

Publicly available collections of drug-like molecules have grown to comprise tens of billions of compounds due to advances in combinatorial chemistry. Traditional methods for identifying "hit" molecules from a large collection of potential drug-like candidates have relied on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have the major drawback that they require exceptional computing capabilities for even relatively small collections of molecules. Hyperdimensional Computing (HDC) is a recently-proposed learning paradigm that represents data with high-dimension binary vectors; this allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas. We consider existing HDC approaches for molecular property classification and introduce two novel encodings of a commonly-used molecular representation, the extended connectivity fingerprint (ECFP). We show that HDC-based inference methods are as much as 91 times more efficient than traditional machine learning methods, and achieve an acceleration of nearly nine orders of magnitude compared to molecular docking. Our results show that HDC accelerated methods retain competitive accuracy on a number of well-studied tasks such as molecular property predictions using the MoleculeNet dataset, and bind/no-bind activity classification using the DUD-E and LIT-PCBA datasets. Our work thus motivates further investigation into molecular representation learning to develop ultraefficient pre-screening tools.

Jones, William↗

An end-to-end deep learning solution for automated LiDAR tree detection in the urban environment

Cataloging and classifying trees in the urban environment is a crucial step in urban and environmental planning; however, manual collection and maintenance of this data is expensive and time-consuming. Although algorithmic approaches that rely on remote sensing data have been developed for tree detection in forests, they generally struggle in the more varied urban environment. This work proposes a novel end-to-end deep learning method for the detection of trees in the urban environment from remote sensing data. Specifically, we develop and train a novel PointNet-based neural network architecture to predict tree locations directly from LiDAR data augmented with multi-spectral imagery. We compare this model to a number of high-performing baselines on a large and varied dataset in the Southern California region, and find that our method outperforms all baselines in terms of tree detection ability (75.5% F-score) and positional accuracy (2.28 meter root mean squared error), while being highly efficient. We then analyze and compare the sources of errors, and how these reveal the strengths and weaknesses of each approach. Our results highlight the importance of fusing spectral and structural information for remote sensing tasks in complex urban environments.

54 ENVIRONMENTAL SCIENCES↗

Low Activity Tritium Detection in CCDs Using Deep Learning Techniques

Here, this study explores the use of charge-coupled devices (CCDs) for detecting low-energy beta particles from tritium decay - a critical signal for nuclear safety, nuclear nonproliferation, and environmental monitoring. We employ a dual approach utilizing both measured CCD data and detailed Geant4 simulations. Our analysis compares classical techniques with advanced deep learning methods, including convolutional neural networks (CNNs), autoencoders trained exclusively on tritium data, and preliminary studies on boosted decision trees (BDTs). The CNN, trained on mixed signal/background datasets, demonstrates superior classification performance, while the autoencoder shows the potential of unsupervised, background-agnostic strategies when background characteristics are poorly defined. These results highlight the excellent sensitivity achievable thanks to the background rejection made possible by information-rich CCD data, paving the way for improved portable tritium monitoring.

Autoencoder↗

Knowledge-informed deep learning for hydrological model calibration: an application to Coal Creek Watershed in Colorado

Abstract. Deep learning (DL)-assisted inverse mapping has shown promise in hydrological model calibration by directly estimating parameters from observations. However, the increasing computational demand for running the state-of-the-art hydrological model limits sufficient ensemble runs for its calibration. In this work, we present a novel knowledge-informed deep learning method that can efficiently conduct the calibration using a few hundred realizations. The method involves two steps. First, we determine decisive model parameters from a complete parameter set based on the mutual information (MI) between model responses and each parameter computed by a limited number of realizations (∼50). Second, we perform more ensemble runs (e.g., several hundred) to generate the training sets for the inverse mapping, which selects informative model responses for estimating each parameter using MI-based parameter sensitivity. We applied this new DL-based method to calibrate a process-based integrated hydrological model, the Advanced Terrestrial Simulator (ATS), at Coal Creek Watershed, CO. The calibration is performed against observed stream discharge (Q) and remotely sensed evapotranspiration (ET) from the water year 2017 to 2019. Preliminary MI analysis on 50 realizations resulted in a down-selection of 7 out of 14 ATS model parameters. Then, we performed a complete MI analysis on 396 realizations and constructed the inverse mapping from informative responses to each of the selected parameters using a deep neural network. Compared with calibration using observations covering all time steps, the new inverse mapping improves parameter estimations, thus enhancing the performance of ATS forward model runs. The Nash–Sutcliffe efficiency (NSE) of streamflow predictions increases from 0.53 to 0.8 when calibrating against Q alone. Using ET observations, on the other hand, does not show much improvement on the performance of ATS modeling mainly due to both the uncertainty of the remotely sensed product and the insufficient coverage of the model ET ensemble in capturing the observation. By using observed Q only, we further performed a multiyear analysis and show that Q is best simulated (NSE > 0.8) by including in the calibration the dry-year flow dynamics that show more sensitivity to subsurface characteristics than the other wet years. Moreover, when continuing the forward runs till the end of 2021, the calibrated models show similar simulation performances during this evaluation period as the calibration period, demonstrating the ability of the estimated parameters in capturing climate sensitivity. Our success highlights the importance of leveraging data-driven knowledge in DL-assisted hydrological model calibration.

54 ENVIRONMENTAL SCIENCES↗

Robust High-Throughput Phenotyping with Deep Segmentation Enabled by a Web-Based Annotator

The abilities of plant biologists and breeders to characterize the genetic basis of physiological traits are limited by their abilities to obtain quantitative data representing precise details of trait variation, and particularly to collect this data at a high-throughput scale with low cost. Although deep learning methods have demonstrated unprecedented potential to automate plant phenotyping, these methods commonly rely on large training sets that can be time-consuming to generate. Intelligent algorithms have therefore been proposed to enhance the productivity of these annotations and reduce human efforts. We propose a high-throughput phenotyping system which features a Graphical User Interface (GUI) and a novel interactive segmentation algorithm: Semantic-Guided Interactive Object Segmentation (SGIOS). By providing a user-friendly interface and intelligent assistance with annotation, this system offers potential to streamline and accelerate the generation of training sets, reducing the effort required by the user. Our evaluation shows that our proposed SGIOS model requires fewer user inputs compared to the state-of-art models for interactive segmentation. As a case study of the use of the GUI applied for genetic discovery in plants, we present an example of results from a preliminary genome-wide association study (GWAS) of in planta regeneration in Populus trichocarpa (poplar). We further demonstrate that the inclusion of a semantic prior map with SGIOS can accelerate the training process for future GWAS, using a sample of a dataset extracted from a poplar GWAS of in vitro regeneration. The capabilities of our phenotyping system surpass those of unassisted humans to rapidly and precisely phenotype our traits of interest. The scalability of this system enables large-scale phenomic screens that would otherwise be time-prohibitive, thereby providing increased power for GWAS, mutant screens, and other studies relying on large sample sizes to characterize the genetic basis of trait variation. Our user-friendly system can be used by researchers lacking a computational background, thus helping to democratize the use of deep segmentation as a tool for plant phenotyping.

54 ENVIRONMENTAL SCIENCES↗

The stellar parameters and elemental abundances from low-resolution spectra – I. 1.2 million giants from LAMOST DR8

As a typical data-driven method, deep learning becomes a natural choice for analysing astronomical data. In this study, we built a deep convolutional neural network (NN) to estimate basic stellar parameters $T\rm {_{eff}}$, log g , metallicity ([M/H] and [Fe/H]) and [α/M] along with nine individual elemental abundances ([C/Fe], [N/Fe], [O/Fe], [Mg/Fe], [Al/Fe], [Si/Fe], [Ca/Fe], [Mn/Fe], and [Ni/Fe]). The NN is trained using common stars between the APOGEE survey and the LAMOST survey. We used low-resolution spectra from LAMOST survey as input, and measurements from APOGEE as labels. For stellar spectra with the signal-to-noise ratio in g band larger than 10 in the test set, the mean absolute error (MAE) is 29 K for $T\rm {_{eff}}$, 0.07 dex for log g , 0.03 dex for both [Fe/H] and [M/H], and 0.02 dex for [α/M]. The MAE of most elements is between 0.02 and 0.04 dex. The trained NN was applied to 1210 145 giants, including sub-giants, from LAMOST DR8 within the range of stellar parameters 3500 K < $T\rm {_{eff}}$ < 5500 K, 0.0 dex < log g < 4.0 dex, −2.5 dex < [Fe/H] < 0.5 dex. The distribution of our results in the chemical spaces is highly consistent with APOGEE labels and stellar parameters show consistency with external high-resolution measurements from GALAH. The results in this study allow us to further studies based on LAMOST data and deepen our understanding of the accretion and evolution history of the Milky Way. The electronic version of the value added catalog is available at http://www.lamost.org/dr8/v1.1/doc/vac.

79 ASTRONOMY AND ASTROPHYSICS↗

De Novo Design of High‐Affinity Miniprotein Binders Targeting Francisella Tularensis Virulence Factor

Abstract Francisella tularensis poses considerable public health risk due to its high infectivity and potential for bioterrorism. Francisella‐like lipoprotein (Flpp3), a key virulence factor unique to Francisella, plays critical roles in infection and immune evasion, making it a promising target for therapeutic development. However, the lack of well‐defined binding pockets and structural information on native interactions has hindered structure‐guided ligand discovery against Flpp3. Here, we used a combination of physics‐based and deep‐learning methods to design high‐affinity miniprotein binders targeting two distinct sites on Flpp3. We identified four binders for site I with binding affinities ranging between 24–110 nM. For the second site, an initial binder showed a dissociation constant ( K D ) of 81 nM, and subsequent site saturation mutagenesis yielded variants with sub‐nanomolar affinities. Circular dichroism confirmed the topology of designed miniproteins. The X‐ray crystal structure of Flpp3 in complex with a site I binder is nearly identical to the design model (Cα root‐mean‐square deviation (RMSD): 0.9 Å). These designed miniproteins provide research tools to explore the roles of Flpp3 in tularemia and should enable the development of new therapeutic candidates.

Gokce‐Alpkilic, Gizem [Molecular Engineering and S↗

Understanding and Leveraging the I/O Patterns of Emerging Machine Learning Analytics

The scientific community is currently experiencing unprecedented amounts of data generated by cutting-edge science facilities. Soon facilities will be producing up to 1 PB/s which will force scientist to use more autonomous techniques to learn from the data. The adoption of machine learning methods, like deep learning techniques, in large-scale workflows comes with a shift in the workflow’s computational and I/O patterns. These changes often include iterative processes and model architecture searches, in which datasets are analyzed multiple times in different formats with different model configurations in order to find accurate, reliable and efficient learning models. This shift in behavior brings changes in I/O patterns at the application level as well at the system level. These changes also bring new challenges for the HPC I/O teams, since these patterns contain more complex I/O workloads. In this paper we discuss the I/O patterns experienced by emerging analytical codes that rely on machine learning algorithms and highlight the challenges in designing efficient I/O transfers for such workflows. We comment on how to leverage the data access patterns in order to fetch in a more efficient way the required input data in the format and order given by the needs of the application and how to optimize the data path between collaborative processes. We will motivate our work and show performance gains with a study case of medical applications.

Gainaru, Ana↗