Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “deep generative models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Language models for the prediction of SARS-CoV-2 inhibitors

The COVID-19 pandemic highlights the need for computational tools to automate and accelerate drug design for novel protein targets. We leverage deep learning language models to generate and score drug candidates based on predicted protein binding affinity. We pre-trained a deep learning language model (BERT) on ∼9.6 billion molecules and achieved peak performance of 603 petaflops in mixed precision. Our work reduces pre-training time from days to hours, compared to previous efforts with this architecture, while also increasing the dataset size by nearly an order of magnitude. For scoring, we fine-tuned the language model using an assembled set of thousands of protein targets with binding affinity data and searched for inhibitors of specific protein targets, SARS-CoV-2 Mpro and PLpro. We utilized a genetic algorithm approach for finding optimal candidates using the generation and scoring capabilities of the language model. Our generalizable models accelerate the identification of inhibitors for emerging therapeutic targets.

Blanchard, Andrew E.↗

Leveraging generative AI for urban digital twins: a scoping review on the autonomous generation of urban data, scenarios, designs, and 3D city models for smart city advancement

The digital transformation of modern cities by integrating advanced information, communication, and computing technologies has marked the epoch of data-driven smart city applications for efficient and sustainable urban management. Despite their effectiveness, these applications often rely on massive amounts of high-dimensional and multi-domain data for monitoring and characterizing different urban sub-systems, presenting challenges in application areas that are limited by data quality and availability, as well as costly efforts for generating urban scenarios and design alternatives. As an emerging research area in deep learning, Generative Artificial Intelligence (GenAI) models have demonstrated their unique values in content generation. This paper aims to explore the innovative integration of GenAI techniques and urban digital twins to address challenges in the planning and management of built environments with focuses on various urban sub-systems, such as transportation, energy, water, and building and infrastructure. The survey starts with the introduction of cutting-edge generative AI models, such as the Generative Adversarial Networks (GAN), Variational Autoencoders (VAEs), Generative Pre-trained Transformer (GPT), followed by a scoping review of the existing urban science applications that leverage the intelligent and autonomous capability of these techniques to facilitate the research, operations, and management of critical urban subsystems, as well as the holistic planning and design of the built environment. Based on the review, we discuss potential opportunities and technical strategies that integrate GenAI models into the next-generation urban digital twins for more intelligent, scalable, and automated smart city development and management.

3D city modeling↗

Quantum-assisted associative adversarial network: applying quantum annealing in deep learning

Abstract Generative models have the capacity to model and generate new examples from a dataset and have an increasingly diverse set of applications driven by commercial and academic interest. In this work, we present an algorithm for learning a latent variable generative model via generative adversarial learning where the canonical uniform noise input is replaced by samples from a graphical model. This graphical model is learned by a Boltzmann machine which learns low-dimensional feature representation of data extracted by the discriminator. A quantum processor can be used to sample from the model to train the Boltzmann machine. This novel hybrid quantum-classical algorithm joins a growing family of algorithms that use a quantum processor sampling subroutine in deep learning, and provides a scalable framework to test the advantages of quantum-assisted learning. For the latent space model, fully connected, symmetric bipartite and Chimera graph topologies are compared on a reduced stochastically binarized MNIST dataset, for both classical and quantum sampling methods. The quantum-assisted associative adversarial network successfully learns a generative model of the MNIST dataset for all topologies. Evaluated using the Fréchet inception distance and inception score, the quantum and classical versions of the algorithm are found to have equivalent performance for learning an implicit generative model of the MNIST dataset. Classical sampling is used to demonstrate the algorithm on the LSUN bedrooms dataset, indicating scalability to larger and color datasets. Though the quantum processor used here is a quantum annealer, the algorithm is general enough such that any quantum processor, such as gate model quantum computers, may be substituted as a sampler.

Wilson, Max (ORCID:0000000207983391)↗

Large language models generate functional protein sequences across diverse families

Deep-learning language models have shown promise in various biotechnological applications, including protein design and engineering. Here, in this paper, we describe ProGen, a language model that can generate protein sequences with a predictable function across large protein families, akin to generating grammatically and semantically correct natural language sentences on diverse topics. The model was trained on 280 million protein sequences from >19,000 families and is augmented with control tags specifying protein properties. ProGen can be further fine-tuned to curated sequences and tags to improve controllable generation performance of proteins from families with sufficient homologous samples. Artificial proteins fine-tuned to five distinct lysozyme families showed similar catalytic efficiencies as natural lysozymes, with sequence identity to natural proteins as low as 31.4%. ProGen is readily adapted to diverse protein families, as we demonstrate with chorismate mutase and malate dehydrogenase.

59 BASIC BIOLOGICAL SCIENCES↗

Machine Intelligence to Detect, Characterise, and Defend against Influence Operations in the Information Environment

Social media has enabled a new era of manipulation in the information and cognitive domains. Deceptive content—misleading, falsified, and fabricated—is routinely created and spread in the modern social media environment with the intent to create confusion and widen political and social divides, and exploit the societal conflict exacerbated by these divides in the real-world (aka physical domain). Such disinformation campaigns demonstrate a threat to the integrity of economic, political, cultural, public health, and national security institutions around the world. In this work we overview our artificial intelligence (AI) capabilities to detect, describe, and defend against information operations on Twitter as an example social platform to understand the influence of misleading and falsified content diffusion and better enable those charged with defending against such manipulation to enable responsive parties to counter it. We first present novel linguistically-informed deep learning (DL) models for misinformation and disinformation detection, and present an in-depth linguistic analysis of psycho-linguistic markers across broad deception categories. We then demonstrate how our models perform in the multilingual and multimodal setting and categorize falsified and misleading content based on the intent to deceive. We also provide a large-scale analysis to describe user behavior and spread patterns while engaging with deceptive content and report novel findings about the immediate diffusion of deceptive content by characterizing the vulnerable sub-populations and their demographics, and explicitly measuring speed and scale of deception spread to uncover who shares deceptive content, how quickly, how much, and how evenly. In addition, we measure audience reactions to misinformation and disinformation at scale, distinguishing the reactions of users identified as bots versus humans. Finally, we take advantage of deep translation and generation models to create unique solutions for real-time defense against digital deception and discuss how to apply causal inference to prescribe and intervene into strategic communications jointly across information, cognitive, and physical domains.

artificial intelligence, deep learning, neural lan↗

Iterative sampling of expensive simulations for faster deep surrogate training

Deep neural network (DNN) surrogates of expensive physics simulations are enabling a rapid change in the way that common experimental design and analysis tasks are approached. Surrogate models allow simulations to be performed in parallel and separately from downstream tasks, thereby enabling analyses that would be impossible with the simulation in-the-loop; surrogates based on DNNs can effectively emulate diverse non-scalar data of the types collected in fusion and laboratory-astrophysics experiments. The challenge is in training the surrogate model, for which large ensembles of physics simulations must be run, preferably without wasting computational effort on uninteresting simulations. Here, in this paper, we present an iterative sampling scheme that can preferentially propose simulations in interesting regions of parameter space without neglecting unexplored regions, allowing high-quality and wide-ranging surrogate models to be trained using 2–3 times fewer simulations compare to space-filling designs. Our approach uses an explicit importance function defined on the simulation output space, balanced against a measure of simulation density which serves as a proxy for surrogate accuracy. It is easy to implement and can be tuned to find interesting simulations early in the study, allowing surrogates to be trained quickly and refined as new simulations become available; this represents an important step towards the routine generation of deep surrogate models quickly enough to be truly relevant to experimental work.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Analytic Modeling of a Deep Shielding Problem

Previous generations of scientists would make tremendous efforts to simplify non-tractable problems and generate simpler models that preserved the fundamental physics. This process involved applying assumptions and simplifications to reduce the complexity of the problem until it reached a solvable form. Each assumption and simplification was chosen and applied with the intent to preserve the essential physics of the problem, since, if the core physics of the problem were eliminated, the simplified model served no purpose. Moreover, if done correctly, solutions to the reduced model would serve as useful approximations to the original problem. In a sense, solving the simple models laid the ground-work for and provided insight into the more complex problem. Today, however, the affordability of high performance computing has essentially replaced the process for analyzing complex problems. Rather than \building up" a problem by understanding smaller, simpler models, a user generally relies on powerful computational tools to directly arrive at solutions to complex problems. As computational resources grow, users continue trying to simulate new, more complex, or more detailed problems, resulting in continual stress on both the code and computational resources. When these resources are limited, the user will have to make concessions by simplifying the problem while trying to preserve important details. In the context of MCNP, simplifications typically come as reductions in geometry, or by using variance reduction techniques. Both approaches can influence the physics of the problem, leading to potentially inaccurate or non-physical results. Errors can also be introduced as a result of faulty input into a computational tool: something as simple as transposing numbers in a tally input can result in incorrect answers.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Analytic Modeling of a Deep Shielding Problem

Previous generations of scientists would make tremendous efforts to simplify non-tractable problems and generate simpler models that preserved the fundamental physics. This process involved applying assumptions and simplifications to reduce the complexity of the problem until it reached a solvable form. Each assumption and simplification was chosen and applied with the intent to preserve the essential physics of the problem, since, if the core physics of the problem were eliminated, the simplified model served no purpose. Moreover, if done correctly, solutions to the reduced model would serve as useful approximations to the original problem. In a sense, solving the simple models laid the ground-work for and provided insight into the more complex problem. Today, however, the affordability of high performance computing has essentially replaced the process for analyzing complex problems. Rather than "building up" a problem by understanding smaller, simpler models, a user generally relies on powerful computational tools to directly arrive at solutions to complex problems. As computational resources grow, users continue trying to simulate new, more complex, or more detailed problems, resulting in continual stress on both the code and computational resources. When these resources are limited, the user will have to make concessions by simplifying the problem while trying to preserve important details. In the context of the Monte Carlo N-Particle radiation transport simulation tool, simplifications typically come as reductions in geometry, or by using variance reduction techniques. Both approaches can influence the physics of the problem, leading to potentially inaccurate or non-physical results. Errors can also be introduced as a result of faulty input into a computational tool: something as simple as transposing numbers in a tally input can result in incorrect answers. In this paradigm, reduced complexity computational and analytical models still have an important purpose. The explicit form of an analytic solution is arguably the best way to understand the qualitative properties of simple models. In contrast to "building up" a complex problem through understanding simpler problems, results from detailed computational scenarios can be better explained by "building down" the complex model through simple models rooted in the fundamental or essential phenomenology. Simplified analytic and computational models can be used to 1) increase a user's confidence in the computational solution of a complex model, 2) confirm there are no user input errors, and 3) ensure essential assumptions of the simulation tool are preserved. This process of using analytic models to develop a more valuable analysis of simulation results is named the results assessment methodology. The utility of the results assessment methodology and a complimentary sensitivity analysis is exemplified through the analysis of the neutron flux in a dry used fuel storage cask. This application was chosen due to current scientific interest in used nuclear fuel storage.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Analytic Modeling of a Deep Shielding Problem

Previous generations of scientists would make tremendous efforts to simplify non-tractable problems and generate simpler models that preserved the fundamental physics. This process involved applying assumptions and simplifications to reduce the complexity of the problem until it reached a solvable form. Each assumption and simplification was chosen and applied with the intent to preserve the essential physics of the problem, since, if the core physics of the problem were eliminated, the simplified model served no purpose. Moreover, if done correctly, solutions to the reduced model would serve as useful approximations to the original problem. In a sense, solving the simple models laid the ground-work for and provided insight into the more complex problem. Today, however, the affordability of high performance computing has essentially replaced the process for analyzing complex problems. Rather than "building up" a problem by understanding smaller, simpler models, a user generally relies on powerful computational tools to directly arrive at solutions to complex problems. As computational resources grow, users continue trying to simulate new, more complex, or more detailed problems, resulting in continual stress on both the code and computational resources. When these resources are limited, the user will have to make concessions by simplifying the problem while trying to preserve important details. In the context of the Monte Carlo N-Particle radiation transport simulation tool, simplifications typically come as reductions in geometry, or by using variance reduction techniques. Both approaches can influence the physics of the problem, leading to potentially inaccurate or non-physical results. Errors can also be introduced as a result of faulty input into a computational tool: something as simple as transposing numbers in a tally input can result in incorrect answers. In this paradigm, reduced complexity computational and analytical models still have an important purpose. The explicit form of an analytic solution is arguably the best way to understand the qualitative properties of simple models. In contrast to "building up" a complex problem through understanding simpler problems, results from detailed computational scenarios can be better explained by "building down" the complex model through simple models rooted in the fundamental or essential phenomenology. Simplified analytic and computational models can be used to 1) increase a user's confidence in the computational solution of a complex model, 2) con firm there are no user input errors, and 3) ensure essential assumptions of the simulation tool are preserved. This process of using analytic models to develop a more valuable analysis of simulation results is named the results assessment methodology. The utility of the results assessment methodology and a complimentary sensitivity analysis is exempli fied through the analysis of the neutron flux in a dry used fuel storage cask. This application was chosen due to current scientific interest in used nuclear fuel storage.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Learning effective physical laws for generating cosmological hydrodynamics with Lagrangian deep learning

The goal of generative models is to learn the intricate relations between the data to create new simulated data, but current approaches fail in very high dimensions. When the true data-generating process is based on physical processes, these impose symmetries and constraints, and the generative model can be created by learning an effective description of the underlying physics, which enables scaling of the generative model to very high dimensions. In this work, we propose Lagrangian deep learning (LDL) for this purpose, applying it to learn outputs of cosmological hydrodynamical simulations. The model uses layers of Lagrangian displacements of particles describing the observables to learn the effective physical laws. The displacements are modeled as the gradient of an effective potential, which explicitly satisfies the translational and rotational invariance. The total number of learned parameters is only of order 10, and they can be viewed as effective theory parameters. Further, we combine N-body solver fast particle mesh (FastPM) with LDL and apply it to a wide range of cosmological outputs, from the dark matter to the stellar maps, gas density, and temperature. The computational cost of LDL is nearly four orders of magnitude lower than that of the full hydrodynamical simulations, yet it outperforms them at the same resolution. We achieve this with only of order 10 layers from the initial conditions to the final output, in contrast to typical cosmological simulations with thousands of time steps. This opens up the possibility of analyzing cosmological observations entirely within this framework, without the need for large dark-matter simulations.

79 ASTRONOMY AND ASTROPHYSICS↗

Connectivity-informed drainage network generation using deep convolution generative adversarial networks

Abstract Stochastic network modeling is often limited by high computational costs to generate a large number of networks enough for meaningful statistical evaluation. In this study, Deep Convolutional Generative Adversarial Networks (DCGANs) were applied to quickly reproduce drainage networks from the already generated network samples without repetitive long modeling of the stochastic network model, Gibb’s model. In particular, we developed a novel connectivity-informed method that converts the drainage network images to the directional information of flow on each node of the drainage network, and then transforms it into multiple binary layers where the connectivity constraints between nodes in the drainage network are stored. DCGANs trained with three different types of training samples were compared; (1) original drainage network images, (2) their corresponding directional information only, and (3) the connectivity-informed directional information. A comparison of generated images demonstrated that the novel connectivity-informed method outperformed the other two methods by training DCGANs more effectively and better reproducing accurate drainage networks due to its compact representation of the network complexity and connectivity. This work highlights that DCGANs can be applicable for high contrast images common in earth and material sciences where the network, fractures, and other high contrast features are important.

42 ENGINEERING↗

Enhancing molecular design efficiency: Uniting language models and generative networks with genetic algorithms

This study examines the effectiveness of generative models in drug discovery, material science, and polymer science, aiming to overcome constraints associated with traditional inverse design methods relying on heuristic rules. Generative models generate synthetic data resembling real data, enabling deep learning model training without extensive labeled datasets. They prove valuable in creating virtual libraries of molecules for material science and facilitating drug discovery by generating molecules with specific properties. While generative adversarial networks (GANs) are explored for these purposes, mode collapse restricts their efficacy, limiting novel structure variability. To address this, we introduce a masked language model (LM) inspired by natural language processing. Although LMs alone can have inherent limitations, we propose a hybrid architecture combining LMs and GANs to efficiently generate new molecules, demonstrating superior performance over standalone masked LMs, particularly for smaller population sizes. This hybrid LM-GAN architecture enhances efficiency in optimizing properties and generating novel samples.

97 MATHEMATICS AND COMPUTING↗

On the Stochastic Stability of Deep Markov Models

Deep Markov models (DMM) are generative models which are scalable and expressive generalization of Markov models for representation, learning, and inference problems. DMMs using deep neural networks to parametrize the transition of Markov probability distributions have recently been shown to provide more expressiveness in modeling sequential data and dynamical system responses. However, the fundamental stochastic stability guarantees of such models have not been thoroughly investigated. In this paper, we present a rigorous analytical method to prove the necessary and sufficient conditions of DMM's stochastic stability. This task is achieved by spectral analysis of the efficiently computed Jacobians of probabilistic maps modeled by deep neural networks. We make theoretical connections between the eigenvalues of neural network's weights and the different activation function types used on the stability and overall dynamic behavior of DMMs with Gaussian distributions. We empirically substantiate our theoretical results on stochastic stability and eigenvalue spectra via several numerical experiments. Formal stability guarantees of DMMs can substantially improve their robustness and trustworthiness, necessary for reliable use in safety-critical real-world applications.

Drgona, Jan↗

Atmospheric Structure Prediction for Infrasound Propagation Modeling Using Deep Learning

Abstract Infrasound is generated by a variety of natural and anthropogenic sources. Infrasonic waves travel through the dynamic atmosphere, which can change on the order of minutes to hours. Infrasound propagation largely depends on the wind and temperature structure of the atmosphere. Numerical weather prediction models are available to provide atmospheric specifications, but uncertainties in these models exist and they are computationally expensive to run. Machine learning has proven useful in predicting tropospheric weather using Long Short‐Term Memory (LSTM) networks. An LSTM network is utilized to make atmospheric specification predictions up to ∼30 km for three different training and testing scenarios: (a) the model is trained and tested using only radiosonde data from the Albuquerque, NM, USA station, (b) the model is trained on radiosonde stations across the contiguous US, excluding the Albuquerque, NM, USA station, which was reserved for testing, and (c) the model is trained and tested on radiosonde stations across the contiguous US. Long Short‐Term Memory predictions are compared to a state‐of‐the‐art reanalysis model and show cases where the LSTM outperforms, performs equally as well, or underperforms in comparison to the state‐of‐the‐art. Regional and temporal trends in model performance across the US are also discussed. Results suggest that the LSTM model is a viable tool for predicting atmospheric specifications for infrasound propagation modeling.

54 ENVIRONMENTAL SCIENCES↗

Tropical Cirrus in Global Storm-Resolving Models: 1. Role of Deep Convection

Pervasive cirrus clouds in the upper troposphere and tropical tropopause layer (TTL) influence the climate by altering the top-of-atmosphere radiation balance and stratospheric water vapor budget. These cirrus are often associated with deep convection, which global climate models must parameterize and struggle to accurately simulate. By comparing high-resolution global storm-resolving models from the Dynamics of the Atmospheric general circulation Modeled On Non-hydrostatic Domains (DYAMOND) intercomparison that explicitly simulate deep convection to satellite observations, we assess how well these models simulate deep convection, convectively generated cirrus, and deep convective injection of water into the TTL over representative tropical land and ocean regions. The DYAMOND models simulate deep convective precipitation, organization, and cloud structure fairly well over land and ocean regions, but with clear intermodel differences. All models produce frequent overshooting convection whose strongest updrafts humidify the TTL and are its main source of frozen water. Intermodel differences in cloud properties and convective injection exceed differences between land and ocean regions in each model. We argue that, with further improvements, global storm-resolving models can better represent tropical cirrus and deep convection in present and future climates than coarser-resolution climate models. To realize this potential, they must use available observations to perfect their ice microphysics and dynamical flow solvers.

54 ENVIRONMENTAL SCIENCES↗

Modeling gas migration through clay-based buffer material using coupled multiphase fluid flow and geomechanics with stress-dependent gas permeability

A model for gas migration through clay-based buffer material is developed for modeling gas generation and migration associated with deep geologic nuclear waste disposal. The model is based on a multiphase fluid flow and geomechanics simulator that is adapted to consider enhanced gas flow when gas pressure is high enough to approach the confining stress magnitude. A key feature in the model is a direct coupling between gas permeability and stress, through a non-linear stress-dependent permeability function. The model was first tested and calibrated by modelling two different laboratory gas migration tests on Wyoming (MX-80) bentonite samples. The calibrated model was then applied to model gas migration through a bentonite buffer of a large-scale gas injection test (Lasgit) conducted at the Äspö Hard Rock Laboratory in Sweden. Observed preferential gas migration along interfaces (between compacted blocks and along the canister surface) required explicit representation of such interfaces in the model. The model with the stress-dependent gas permeability accurately captured observed experimental responses in terms of gas breakthrough time, peak gas pressure, and cumulative gas flow rates. The calibrated model was finally applied to simulate migration of hydrogen gas generated within a breached nuclear waste canister over 10,000 years, involving migration of much larger gas volumes. For the considered gas generation rate and host rock properties, the generated gas could migrate through the bentonite buffer and released into the surrounding host rock at a maximum gas pressure somewhat higher than the initial total stress, though a significant amount of hydrogen remained within the buffer. This modelling sets the stage for further detailed analysis of the impact of hydrogen gas generation on the long-term performance of nuclear waste repositories.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

On the Stochastic Stability of Deep Markov Models

Deep Markov models (DMM) are generative models which are scalable and expressive generalization of Markov models for representation, learning, and inference problems. However, the fundamental stochastic stability guarantees of such models have not been thoroughly investigated. In this paper, we present a novel stability analysis method and provide sufficient conditions of DMM's stochastic stability. The proposed stability analysis is based on the contraction of probabilistic maps modeled by deep neural networks. We make connections between the spectral properties of neural network's weights and different types of used activation function on the stability and overall dynamic behavior of DMMs with Gaussian distributions. Based on the theory, we propose a few practical methods for designing constrained DMMs with guaranteed stability. We empirically substantiate our theoretical results via intuitive numerical experiments using the proposed stability constraints.

Drgona, Jan↗