Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “deep generative models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

On the Stochastic Stability of Deep Markov Models

Deep Markov models (DMM) are generative models which are scalable and expressive generalization of Markov models for representation, learning, and inference problems. DMMs using deep neural networks to parametrize the transition of Markov probability distributions have recently been shown to provide more expressiveness in modeling sequential data and dynamical system responses. However, the fundamental stochastic stability guarantees of such models have not been thoroughly investigated. In this paper, we present a rigorous analytical method to prove the necessary and sufficient conditions of DMM's stochastic stability. This task is achieved by spectral analysis of the efficiently computed Jacobians of probabilistic maps modeled by deep neural networks. We make theoretical connections between the eigenvalues of neural network's weights and the different activation function types used on the stability and overall dynamic behavior of DMMs with Gaussian distributions. We empirically substantiate our theoretical results on stochastic stability and eigenvalue spectra via several numerical experiments. Formal stability guarantees of DMMs can substantially improve their robustness and trustworthiness, necessary for reliable use in safety-critical real-world applications.

Drgona, Jan↗

Atmospheric Structure Prediction for Infrasound Propagation Modeling Using Deep Learning

Abstract Infrasound is generated by a variety of natural and anthropogenic sources. Infrasonic waves travel through the dynamic atmosphere, which can change on the order of minutes to hours. Infrasound propagation largely depends on the wind and temperature structure of the atmosphere. Numerical weather prediction models are available to provide atmospheric specifications, but uncertainties in these models exist and they are computationally expensive to run. Machine learning has proven useful in predicting tropospheric weather using Long Short‐Term Memory (LSTM) networks. An LSTM network is utilized to make atmospheric specification predictions up to ∼30 km for three different training and testing scenarios: (a) the model is trained and tested using only radiosonde data from the Albuquerque, NM, USA station, (b) the model is trained on radiosonde stations across the contiguous US, excluding the Albuquerque, NM, USA station, which was reserved for testing, and (c) the model is trained and tested on radiosonde stations across the contiguous US. Long Short‐Term Memory predictions are compared to a state‐of‐the‐art reanalysis model and show cases where the LSTM outperforms, performs equally as well, or underperforms in comparison to the state‐of‐the‐art. Regional and temporal trends in model performance across the US are also discussed. Results suggest that the LSTM model is a viable tool for predicting atmospheric specifications for infrasound propagation modeling.

54 ENVIRONMENTAL SCIENCES↗

Tropical Cirrus in Global Storm-Resolving Models: 1. Role of Deep Convection

Pervasive cirrus clouds in the upper troposphere and tropical tropopause layer (TTL) influence the climate by altering the top-of-atmosphere radiation balance and stratospheric water vapor budget. These cirrus are often associated with deep convection, which global climate models must parameterize and struggle to accurately simulate. By comparing high-resolution global storm-resolving models from the Dynamics of the Atmospheric general circulation Modeled On Non-hydrostatic Domains (DYAMOND) intercomparison that explicitly simulate deep convection to satellite observations, we assess how well these models simulate deep convection, convectively generated cirrus, and deep convective injection of water into the TTL over representative tropical land and ocean regions. The DYAMOND models simulate deep convective precipitation, organization, and cloud structure fairly well over land and ocean regions, but with clear intermodel differences. All models produce frequent overshooting convection whose strongest updrafts humidify the TTL and are its main source of frozen water. Intermodel differences in cloud properties and convective injection exceed differences between land and ocean regions in each model. We argue that, with further improvements, global storm-resolving models can better represent tropical cirrus and deep convection in present and future climates than coarser-resolution climate models. To realize this potential, they must use available observations to perfect their ice microphysics and dynamical flow solvers.

54 ENVIRONMENTAL SCIENCES↗

Modeling gas migration through clay-based buffer material using coupled multiphase fluid flow and geomechanics with stress-dependent gas permeability

A model for gas migration through clay-based buffer material is developed for modeling gas generation and migration associated with deep geologic nuclear waste disposal. The model is based on a multiphase fluid flow and geomechanics simulator that is adapted to consider enhanced gas flow when gas pressure is high enough to approach the confining stress magnitude. A key feature in the model is a direct coupling between gas permeability and stress, through a non-linear stress-dependent permeability function. The model was first tested and calibrated by modelling two different laboratory gas migration tests on Wyoming (MX-80) bentonite samples. The calibrated model was then applied to model gas migration through a bentonite buffer of a large-scale gas injection test (Lasgit) conducted at the Äspö Hard Rock Laboratory in Sweden. Observed preferential gas migration along interfaces (between compacted blocks and along the canister surface) required explicit representation of such interfaces in the model. The model with the stress-dependent gas permeability accurately captured observed experimental responses in terms of gas breakthrough time, peak gas pressure, and cumulative gas flow rates. The calibrated model was finally applied to simulate migration of hydrogen gas generated within a breached nuclear waste canister over 10,000 years, involving migration of much larger gas volumes. For the considered gas generation rate and host rock properties, the generated gas could migrate through the bentonite buffer and released into the surrounding host rock at a maximum gas pressure somewhat higher than the initial total stress, though a significant amount of hydrogen remained within the buffer. This modelling sets the stage for further detailed analysis of the impact of hydrogen gas generation on the long-term performance of nuclear waste repositories.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

On the Stochastic Stability of Deep Markov Models

Deep Markov models (DMM) are generative models which are scalable and expressive generalization of Markov models for representation, learning, and inference problems. However, the fundamental stochastic stability guarantees of such models have not been thoroughly investigated. In this paper, we present a novel stability analysis method and provide sufficient conditions of DMM's stochastic stability. The proposed stability analysis is based on the contraction of probabilistic maps modeled by deep neural networks. We make connections between the spectral properties of neural network's weights and different types of used activation function on the stability and overall dynamic behavior of DMMs with Gaussian distributions. Based on the theory, we propose a few practical methods for designing constrained DMMs with guaranteed stability. We empirically substantiate our theoretical results via intuitive numerical experiments using the proposed stability constraints.

Drgona, Jan↗

Wavelet and Deep-Learning-Based Approach for Generation System Problematic Parameters Identification and Calibration

Accurate models of generation systems are critical for maintaining reliable and secure grid operations. In this paper, a novel and systematic approach is proposed to identify and calibrate the generation system problematic parameters using continuous wavelet transform (CWT) and advanced deep-learning technology. The phasor measurement unit (PMU) data are used through “event playback” to check whether the parameter calibration is required, and if yes, a group of suspicious parameters will be identified as the primary problematic parameter candidates (PPCs). These primary PPCs are randomly perturbed to generate the event playback simulation data, which are used by the CWT and convolutional neural networks (CNNs) to further narrow down the primary PPCs into a smaller set of candidates. Then, the identified candidates are perturbed again to generate massive event playback simulation data for training a parameter calibration neural network. Here, we designed a multi-output neural network structure to find the mappings between the perturbed parameters and the simulation data using both CNN and long short-term memory (LSTM) models. Finally, the well-trained and tested CNN-LSTM model is used to estimate the accurate value of the suspicious parameters with actual PMU measurements. The proposed CNN-LSTM network can accurately and reliably estimate the generation-system problematic parameters, and has better performance when compared to other machine-learning methods, such as the multilayer perceptron network and the conditional variational autoencoder method. The accuracy and effectiveness of the proposed approach have been validated through simulation and real-world data.

24 POWER TRANSMISSION AND DISTRIBUTION↗

ORNL_AISD_DL-HLgap

This dataset provides supplementary molecular dataset of Deep Learning Workflow for the Inverse Design of Molecules with Specific Optoelectronic Properties. The dataset comprises three main directories such as GDB-9_dataset, Low_HL_Gap_dataset, and High_HL_Gap_dataset which individually has csv files, smiles_txt files, pdb files and xyz files containing information of molecular structures, properties and coordinates generated from deep learning workflow using generative model, surrogate model and DFTB calculation results. GDB-9_dataset contains the molecular data extracted from the original GDB-9 dataset with additional data of DFTB HL gap, surrogate HL gap and molecular property analysis. (the number of atoms, aromaticity and double bond equivalent) Low_HL_Gap_dataset and High_HL_Gap_dataset contains series of dataset for different generations with further split to train and test dataset that were obtained from the iterative workflow described in the manuscript. Additional directory Chemiscope_visualization in Low_HL_Gap_dataset directory contains compressed json files to visualize molecules using chemiscope.org page or application to help readers examine generated molecules.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

LENS: Learning Enabled Network Synthesis

RTRC and UMD have developed novel machine learning based methods under the ARPA-E DIFFERENTIATE program for rapid acceleration of hypothesis generation in complex architecture design spaces involving both discrete choices of component inclusion and interconnection and continuous parametric decisions. The project named Learning Enabled Network Synthesis (LENS) further demonstrated the developed methods on challenging electrical power converter design problems by identifying the most suitable circuit topologies and simultaneously selecting the most appropriate components to achieve optimized design of power converter with improved performances. We demonstrated that LENS could enable exploration of very large design space of circuit topologies and components by addressing the limitations of conventional design process in non-linear, high switching speed, multi-dimensional power converter design and optimization. The key innovation developed in LENS is the seamless integration of statistical learning and logical reasoning techniques and building on the individual strengths of these techniques for rapid hypothesis discovery. The main component of LENS comprises of: 1) Graph Reasoning Engine (GRE) to enforce composition rules that rapidly reject all discrete architectures that are composed incorrectly and generates an adaptive database of feasible designs which can be used by ML modules, 2) Graph Generative Learning module which is a deep neural network based generative model for graph architectures which can enable design space exploration beyond the dataset generated by the GRE, 3) Graph Reduced Order Model (ROM) for graph domains for accelerating computation of output metrics, and 4) Active learning and Rule Discovery module for sample efficient learning and extracting logical rules from the learned ML models which will be integrated in the GRE to enhance the filtering effectiveness. LENS approach can be applied to any design domains where designs can be represented as multi-attribute graphs. The LENS team integrated the various technical innovations listed above into an optimization pipeline and exercised the optimization pipeline on the converter design problem. The LENS project demonstrated that the developed AI/ML technologies can be used to generate novel converter circuits >45x faster than experts on chosen use-cases. This can enable faster design space exploration and identification of new designs which are not considered by experts due to the increasing design space complexity. This has significant potential impact on the public and energy needs of the country. It is currently estimated that 30% of all electrical powers generated passes through power converters. The future estimate is that 80% of all power generated would be passing through converters. LENS fills a critical gap in this space since by accelerating the design process the designers would be able to generate more efficient converters which can lead to significant energy savings for the country.

42 ENGINEERING↗

Integrating particle flavor into deep learning models for hadronization

Hadronization models used in event generators are physics-inspired functions with many tunable parameters. Since we do not understand hadronization from first principles, there have been multiple proposals to improve the accuracy of hadronization models by utilizing more flexible parametrizations based on neural networks. These recent proposals have focused on the kinematic properties of hadrons, but a full model must also include particle flavor. In this paper, we show how to build a deep learning-based hadronization model that includes both kinematic (continuous) and flavor (discrete) degrees of freedom. Our approach is based on generative adversarial networks and we show the performance within the context of the cluster hadronization model within the erwig event generator.

Chan, Jay↗

NRAP-Open-IAM: Generic Aquifer Component Development and Testing

The Generic Aquifer Model calculates the concentrations of dissolved salt and dissolved CO 2 surrounding a leaking legacy well. The Generic Aquifer model can also estimate the size of an “impact plume” where concentration changes exceed user-specified thresholds. The model is a component of NRAP-Open-IAM, an open-source Integrated Assessment Model (IAM) developed by the National Risk Assessment Partnership (NRAP) to perform risk assessment for geologic CO 2 storage. The input parameters were selected to cover a wide range of groundwater aquifers and leakage rates. The generic aquifer model was developed using a generative adversarial deep learning network, trained using a large synthetic dataset of STOMP multiphase flow simulations. The deep learning model predictions of dissolved salt and dissolved CO 2 in the aquifer compare well to the original STOMP simulation results. The extent of aquifer impacted by leaking CO 2 or brine is calculated using a user-defined mass fraction threshold. The aquifer impact volumes calculated based on STOMP simulation results compare well to those calculated based on the deep learning model. In a provided python script, gridded observation results from the generic aquifer component of NRAP-Open-IAM are converted to HDF5 format files for monitoring design with the DREAM code.

54 ENVIRONMENTAL SCIENCES↗

Joint Modeling of Quasar Variability and Accretion Disk Reprocessing Using Latent Stochastic Differential Equations

Quasars are bright active galactic nuclei powered by the accretion of matter around supermassive black holes at the center of galaxies. Their stochastic brightness variability depends on the physical properties of the accretion disk and black hole. The upcoming Rubin Observatory Legacy Survey of Space and Time (LSST) is expected to observe tens of millions of quasars, so there is a need for efficient techniques like machine learning that can handle the large volume of data. Quasar variability is believed to be driven by an X-ray corona, which is reprocessed by the accretion disk and emitted as UV/optical variability. We are the first to introduce an auto-differentiable simulation of the accretion disk and reprocessing. We use the simulation as a direct component of our neural network to jointly model the driving variability and reprocessing, trained with supervised learning on simulated LSST-like 10 yr quasar light curves. We encode the light curves using a transformer encoder, and the driving variability is reconstructed using latent stochastic differential equations, a physically motivated generative deep learning method that can model continuous-time stochastic dynamics. By embedding the physical processes of the driving signal and reprocessing into our network, we achieve a model that is more robust and interpretable. We demonstrate that our model outperforms a Gaussian process regression baseline and can infer accretion disk parameters and time delays between wave bands, even for out-of-distribution driving signals. Our approach provides a powerful framework that can be adapted to solve other inverse problems in multivariate time series.

Fagin, Joshua [City Univ. of New York (CUNY), NY (↗

Hypothesis testing via AI: Generating physically interpretable models of scientific data with machine learning (Full Technical Report)

Deep learning has demonstrated an exceptional ability to solve complex tasks (an engineering success); however, it has done so at the expense of the ability to generate new knowledge (a scientific failure). We propose an alternative framework—entitled Deep Symbolic Regression (DSR)—in which artificial neural networks (NNs) rapidly generate hypotheses about physical relationships among inputs. This framework bypasses the need to interpret an NN altogether, while still leveraging the representational power of deep learning. The resulting models are tractable mathematical expressions, which are inherently and readily human interpretable and can provide insights into underlying physical phenomena. Further, we fold this methodology into the scientific process by allowing the scientist to directly integrate a priori knowledge and beliefs to accelerate learning. We demonstrate this methodology on symbolic regression—the problem of rediscovering underlying expressions describing a dataset—and achieve state-of-the-art performance across a wide variety of symbolic regression problems. Further, we generalize our DSR framework to apply to the more general class of symbolic optimization problems, in which one seeks to optimize a sequence of symbols or “tokens” under a black-box reward function. Examples of other symbolic optimization problems include neural architecture search and computational antibody design. Our generalized tool, Deep Symbolic Optimization (DSO), has been demonstrated on the task of learning symbolic control policies for reinforcement learning environments, and has been adopted as an enabling capability for computational antibody design.

97 MATHEMATICS AND COMPUTING↗

Machine Learned Hückel Theory: Interfacing Physics and Deep Neural Networks

The Hückel Hamiltonian is an incredibly simple tight-binding model known for its ability to capture qualitative physics phenomena arising from electron interactions in molecules and materials. Part of its simplicity arises from using only two types of empirically fit physics-motivated parameters: the first describes the orbital energies on each atom and the second describes electronic interactions and bonding between atoms. By replacing these empirical parameters with machine-learned dynamic values, we vastly increase the accuracy of the extended Hückel model. The dynamic values are generated with a deep neural network, which is trained to reproduce orbital energies and densities derived from density functional theory. The resulting model retains interpretability, while the deep neural network parameterization is smooth and accurate and reproduces insightful features of the original empirical parameterization. Altogether, this work shows the promise of utilizing machine learning to formulate simple, accurate, and dynamically parameterized physics models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Learning generative neural networks with physics knowledge

Deep generative neural networks have enabled modeling complex distributions, but incorporating physics knowledge into the neural networks is still challenging and is at the core of current physics-based machine learning research. To this end, we propose a physics generative neural network (PhysGNN), a new class of generative neural networks for learning unknown distributions in a physical system described by partial differential equations (PDE). PhysGNN couples PDE systems with generative neural networks. It is a fully differentiable model that allows back-propagation of gradients through both numerical PDE solvers and generative neural networks, and is trained by minimizing the discrete Wasserstein distance between generated and observed probability distributions of the PDE outputs using the stochastic gradient descent method. Moreover, PhysGNN does not require adversarial training like standard generative neural networks, which offers better stability than adversarial training. We show that PhysGNN can learn complex distributions in stochastic inverse problems, where conventional methods such as maximum likelihood estimation and momentum matching methods may be inapplicable when little knowledge is known about the form of unknown distributions or the physical model is too complex. Furthermore, our method allows physics-based generative neural network training for learning complex distributions in the context of differential equations.

97 MATHEMATICS AND COMPUTING↗

An open database of computed bulk ternary transition metal dichalcogenides

Abstract We present a dataset of structural relaxations of bulk ternary transition metal dichalcogenides (TMDs) computed via plane-wave density functional theory (DFT). We examined combinations of up to two chalcogenides with seven transition metals from groups 4–6 in octahedral (1T) or trigonal prismatic (2H) coordination. The full dataset consists of 672 unique stoichiometries, with a total of 50,337 individual configurations generated during structural relaxation. Our motivations for building this dataset are (1) to develop a training set for the generation of machine and deep learning models and (2) to obtain structural minima over a range of stoichiometries to support future electronic analyses. We provide the dataset as individual VASP xml files as well as all configurations encountered during relaxations collated into an ASE database with the corresponding total energy and atomic forces. In this report, we discuss the dataset in more detail and highlight interesting structural and electronic features of the relaxed structures.

36 MATERIALS SCIENCE↗

De novo design of small beta barrel proteins

Small beta barrel proteins are attractive targets for computational design because of their considerable functional diversity despite their very small size (<70 amino acids). However, there are considerable challenges to designing such structures, and there has been little success thus far. Because of the small size, the hydrophobic core stabilizing the fold is necessarily very small, and the conformational strain of barrel closure can oppose folding; also intermolecular aggregation through free beta strand edges can compete with proper monomer folding. Here, we explore the de novo design of small beta barrel topologies using both Rosetta energy–based methods and deep learning approaches to design four small beta barrel folds: Src homology 3 (SH3) and oligonucleotide/oligosaccharide-binding (OB) topologies found in nature and five and six up-and-down-stranded barrels rarely if ever seen in nature. Both approaches yielded successful designs with high thermal stability and experimentally determined structures with less than 2.4 Å rmsd from the designed models. Using deep learning for backbone generation and Rosetta for sequence design yielded higher design success rates and increased structural diversity than Rosetta alone. The ability to design a large and structurally diverse set of small beta barrel proteins greatly increases the protein shape space available for designing binders to protein targets of interest.

59 BASIC BIOLOGICAL SCIENCES↗