Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Generative learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Future Building Archetypes for Los Angeles (2100 Projection)

This dataset (Data.zip) includes empirical and machine learning-generated building information for the Los Angeles urban region. The MAv1_LA.csv file provides the baseline 2015 building data while Final_IECC_LO_2100_GAN.csv represents generative adversarial network-projected urban morphologies for the year 2100. Building archetypes were created for both datasets (Basecase_LA_Archetype.csv and LA_Simulation_2100_GAN_Archetype.csv) using footprint area as the key aggregation variable. More details about the dataset are provided in the attached readme file (README_LA_Archetype_MAv1.txt)

AutoBEM

Identifying Meteorological Influences on Marine Low Cloud Mesoscale Morphology Using Satellite Classifications

Marine low cloud mesoscale morphology in the southeastern Pacific Ocean is analyzed using a large dataset of machine-learning generated classifications spanning three years. Meteorological variables and cloud properties are composited 10by mesoscale cloud type, showing distinct meteorological regimes of marine low cloud organization from the tropics to the midlatitudes. The presentation of mesoscale cellular convection, with respect to geographic distribution, boundary layer structure, and large-scale environmental conditions, agrees with prior knowledge. Two tropical and subtropical cumuliform boundary layer regimes, suppressed cumulus and clustered cumulus, are studied in detail. The patterns in precipitation, circulation, column water vapor, and cloudiness are consistent with the representation of marine shallow mesoscale convective 15 self-aggregation by large eddy simulations of the boundary layer. Although they occur under similar large-scale conditions, the suppressed and clustered low cloud types are found to be well-separated by variables associated with low-level mesoscale circulation, with surface wind divergence being the clearest discriminator between them, whether reanalysis or satellite observations are used. Clustered regimes are associated with surface convergence and suppressed regimes are associated with surface divergence.

Johannes Mohrmann

A Gigaparsec-scale Hydrodynamic Volume Reconstructed with Deep Learning

The next generation of spectroscopic surveys will map the large-scale structure of the Universe at high redshifts (2 ≤ z ≤ 5) using millions of quasar spectra, enabling major advances in constraining both the standard cosmological model and its extensions. Robust cosmological analyses of these data sets require numerical simulations that both cover gigaparsec volumes and resolve features on ∼10 kpc scales and smaller. However, running such large-volume, high-resolution hydrodynamic simulations is computationally prohibitive. We present a generative deep learning model that enhances a low-resolution, gigaparsec-scale (960 h −1 Mpc) hydrodynamic simulation using a smaller (80 h −1 Mpc) high-resolution input hydrodynamic simulation as training data. The resulting enhanced simulation reproduces the line-of-sight power spectrum to within ∼10% and the three-dimensional power spectrum at the ∼20% level at intermediate to small scales (k ≲ 2 h Mpc −1 ). Our method shows strong promise for producing realistic simulations for cosmological analyses with current surveys such as the Dark Energy Spectroscopic Instrument and upcoming next-generation experiments, but further improvements are needed to accurately recover the large-scale modes. We publicly release the enhanced hydrodynamic simulation, along with a halo catalog from a companion N-body dark matter simulation to support the calibration of data analysis pipelines for these large-scale surveys.

Convolutional neural networks

Innovating Training through Immersive Environments: Generation Y, Exploratory Learning, and Serious Games

Over the next decade, those entering Service and Joint Staff positions within the military will come from a different generation than the current leadership. They will come from Generation Y and have differing preferences for learning. Immersive learning environments like serious games and virtual world initiatives can complement traditional training methods to provide a better overall training program for staffs. Generation Y members desire learning methods which are relevant and interactive, regardless of whether they are delivered over the internet or in person. This paper focuses on a project undertaken to assess alternative training methods to teach special operations staffs. It provides a summary of the needs analysis used to consider alternatives and to better posture the Department of Defense for future training development.

Gendron, Gerald

A Public Data Set of Auto-Generated Geotagged PV Site Equipment, Generated via Deep Learning

In this research, we present a data set over 100 photovoltaic (PV) sites in TX, which have been automatically geotagged via a fully autonomous deep learning (DL) pipeline. Specifically, locations of inverters, tracker/fixed tilt rows, batteries, and substations are labeled algorithmically. To ensure high data quality, all systems have been reviewed manually and any deep learning errors have been corrected. This public data set, as well as the open-sourced pipeline used to generate it, is valuable for site planning, modelling, and insurance purposes. Given time and resources, we hope to extend the data set to additional states/regions in the US.

14 SOLAR ENERGY

Data Generation for Machine Learning Interatomic Potentials and Beyond

The field of data-driven chemistry is undergoing an evolution, driven by innovations in machine learning models for predicting molecular properties and behavior. Recent strides in ML-based interatomic potentials have paved the way for accurate modeling of diverse chemical and structural properties at the atomic level. The key determinant defining MLIP reliability remains the quality of the training data. A paramount challenge lies in constructing training sets that capture specific domains in the vast chemical and structural space. This Review navigates the intricate landscape of essential components and integrity of training data that ensure the extensibility and transferability of the resulting models. We delve into the details of active learning, discussing its various facets and implementations. We outline different types of uncertainty quantification applied to atomistic data acquisition and the correlations between estimated uncertainty and true error. The role of atomistic data samplers in generating diverse and informative structures is highlighted. Furthermore, we discuss data acquisition via modified and surrogate potential energy surfaces as an innovative approach to diversify training data. The Review also provides a list of publicly available data sets that cover essential domains of chemical space.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Using Machine-Learning to Dynamically Generate Operationally Acceptable Strategic Reroute Options

The newly developed Trajectory Option Set (TOS), a preference-weighted set of alternative routes submitted by flight operators, is a capability in the U.S. traffic flow management system that enables automated trajectory negotiation between flight operators and Air Navigation Service Providers. The objective of this paper is to describe and demonstrate an approach for automatically generating pre-departure and airborne TOSs that have a high probability of operational acceptance. The approach uses hierarchical clustering of historical route data to identify route candidates. The probability of operational acceptance is then estimated using predictors trained on historical flight plan amendment data using supervised machine learning algorithms, allowing the routes with highest probability of operational acceptance to be selected for the TOS. Features used describe historical route usage, difference in flight time and downstream demand to capacity imbalance. A random forest was found to be the best performing algorithm for learning operational acceptability, with a model accuracy of 0.96. The approach is demonstrated for an historical pre-departure flight from Dallas/Fort Worth International Airport to Newark Liberty International Airport.

Evans, Antony

Using Machine-Learning to Dynamically Generate Operationally Acceptable Strategic Reroute Options

The newly developed Trajectory Option Set (TOS), a preference-weighted set of alternative routes submitted by flight operators, is a capability in the U.S. traffic flow management system that enables automated trajectory negotiation between flight operators and Air Navigation Service Providers. The objective of this paper is to describe and demonstrate an approach for automatically generating pre-departure and airborne TOSs that have a high probability of operational acceptance. The approach uses hierarchical clustering of historical route data to identify route candidates. The probability of operational acceptance is then estimated using predictors trained on historical flight plan amendment data using supervised machine learning algorithms, allowing the routes with highest probability of operational acceptance to be selected for the TOS. Features used describe historical route usage, difference in flight time and downstream demand to capacity imbalance. A random forest was found to be the best performing algorithm for learning operational acceptability, with a model accuracy of 0.96. The approach is demonstrated for an historical pre-departure flight from Dallas/Fort Worth International Airport to Newark Liberty International Airport.

Evans, Antony

SDYN-GANs: Adversarial learning methods for multistep generative models for general order stochastic dynamics

We introduce adversarial learning methods for data-driven generative modeling of dynamics of nth-order stochastic systems. Our approach builds on Generative Adversarial Networks (GANs) with generative model classes based on stable m-step stochastic numerical integrators. From observations of trajectory samples, we introduce methods for learning long-time predictors and stable representations of the dynamics. Our approaches use discriminators based on Maximum Mean Discrepancy (MMD), training protocols using both conditional and marginal distributions, and methods for learning dynamic responses over different time-scales. We show how our approaches can be used for modeling physical systems to learn force-laws, damping coefficients, and noise-related parameters. Our adversarial learning approaches provide methods for obtaining stable generative models for dynamic tasks including long-time prediction and developing simulations for stochastic systems.

• Artificial intelligence (AI) / machine learning

Learning from examples - Generation and evaluation of decision trees for software resource analysis

A general solution method for the automatic generation of decision (or classification) trees is investigated. The approach is to provide insights through in-depth empirical characterization and evaluation of decision trees for software resource data analysis. The trees identify classes of objects (software modules) that had high development effort. Sixteen software systems ranging from 3,000 to 112,000 source lines were selected for analysis from a NASA production environment. The collection and analysis of 74 attributes (or metrics), for over 4,700 objects, captured information about the development effort, faults, changes, design style, and implementation style. A total of 9,600 decision trees were automatically generated and evaluated. The trees correctly identified 79.3 percent of the software modules that had high development effort or faults, and the trees generated from the best parameter combinations correctly identified 88.4 percent of the modules on the average.

Selby, Richard W.

Learning turbulent flows with generative models for super resolution and sparse flow reconstruction

Neural operators are promising surrogates for dynamical systems but when trained with standard L 2 losses they tend to oversmooth fine-scale turbulent structures. Here, we show that combining operator learning with generative modeling overcomes this limitation. We consider three practical turbulent-flow challenges where conventional neural operators fail: spatio-temporal super-resolution, forecasting, and sparse flow reconstruction. For Schlieren jet super-resolution, an adversarially trained neural operator (adv-NO) reduces the energy-spectrum error by 15 × while preserving sharp gradients at neural operator-like inference cost. For 3D homogeneous isotropic turbulence, adv-NO trained on only 160 timesteps from a single trajectory forecasts accurately for five eddy-turnover times and offers 114 × wall-clock speed-up at inference than the baseline diffusion-based forecasters, enabling near-real-time rollouts. For reconstructing cylinder wake flows from highly sparse Particle Tracking Velocimetry-like inputs, a conditional generative model infers full 3D velocity and pressure fields with correct phase alignment and statistics. These advances enable accurate reconstruction and forecasting at low compute cost, bringing near-real-time analysis and control within reach in experimental and computational fluid mechanics.

Fluid dynamics

Reweighting configurations generated by transferable, machine learned models for protein sidechain backmapping

Multiscale modeling requires the linking of models at different levels of detail, with the goal of gaining accelerations from lower fidelity models while recovering fine details from higher resolution models. Communication across resolutions is particularly important in modeling soft matter, where tight couplings exist between molecular-level details and mesoscale structures. While multiscale modeling of biomolecules has become a critical component in exploring their structure and self-assembly, backmapping from coarse-grained to fine-grained, or atomistic, representations presents a challenge, despite recent advances through machine learning. A major hurdle, especially for strategies utilizing machine learning, is that backmappings can only approximately recover the atomistic ensemble of interest. We demonstrate conditions for which backmapped configurations may be reweighted to exactly recover the desired atomistic ensemble. By training separate decoding models for each sidechain type, we develop an algorithm based on normalizing flows and geometric algebra attention to autoregressively propose backmapped configurations for any protein sequence. Critical for reweighting with modern protein force fields, our trained models include all hydrogen atoms in the backmapping and make probabilities associated with atomistic configurations directly accessible. We also demonstrate, however, that reweighting is extremely challenging despite state-of-the-art performance on recently developed metrics and generation of configurations with low energies in atomistic protein force fields. Through detailed analysis of configurational weights, we show that machine-learned backmappings must not only generate configurations with reasonable energies, but also correctly assign relative probabilities under the generative model. These are broadly important considerations in generative modeling of atomistic molecular configurations.

Monroe, Jacob I. [Univ. of Arkansas, Fayetteville,

Large-Scale Alfvenic Impulses on the Sun: How They Are Generated and What We Learn From Them

NASA GSFC The Sun's atmosphere hosts a wide variety of magnetosonic disturbances. These wave modes are detected, almost exclusively, by examining images of the Sun's magnetic atmosphere and looking for propagating distortions. Although none of the Sun's plasma parameters are measured directly, we derive a great deal of information from these observations. In fact, by modeling these propagating disturbances, we may be able to derive the most accurate estimates plasma parameters. From observations absorption, refraction, reflection, and coupling of numerous wave modes, we advance our knowledge of the Sun's magnetic field, temperature, density, and current. The Sun's continuous oscillation, coronal mass ejections, flares, and other dynamic phenomena can produce wave disturbances which are observable from near-Earth space. Several of these disturbances have been traced from the inner corona out into the heliosphere. From the generation of these disturbances, we are able to learn about the phenomena which create them as well as the media through which they re-propagating. The presentation will include a discussion of the generation of Alfvenic disturbances on the Sun, ways we observe these disturbances, and how recent advances in modeling and analysis have brought us closer to determining solar in situ parameters.

Thompson, Barbara

Generative AI models for learning flow maps of stochastic dynamical systems in bounded domains

Simulating stochastic differential equations (SDEs) in bounded domains, presents significant computational challenges due to particle exit phenomena, which requires accurate modeling of interior stochastic dynamics and boundary interactions. Despite the success of machine learning-based methods in learning SDEs, existing learning methods are not applicable to SDEs in bounded domains because they cannot accurately capture the particle exit dynamics. We present a unified hybrid data-driven approach that combines a conditional diffusion model with an exit prediction neural network to capture both interior stochastic dynamics and boundary exit phenomena. Our ML model consists of two major components: a neural network that learns exit probabilities using binary cross-entropy loss with rigorous convergence guarantees, and a training-free diffusion model that generates state transitions for non-exiting particles using closed-form score functions. The two components are integrated through a probabilistic sampling algorithm that determines particle exit at each time step and generates appropriate state transitions. Here, the performance of the proposed approach is demonstrated via three test cases: a one-dimensional simplified problem for theoretical verification, a two-dimensional advection-diffusion problem in a bounded domain, and a three-dimensional problem of interest to magnetically confined fusion plasmas.

Bounded domains

A robust synthetic data generation framework for machine learning in high-resolution transmission electron microscopy (HRTEM)

Machine learning techniques are attractive options for developing highly-accurate analysis tools for nanomaterials characterization, including high-resolution transmission electron microscopy (HRTEM). However, successfully implementing such machine learning tools can be difficult due to the challenges in procuring sufficiently large, high-quality training datasets from experiments. In this work, we introduce Construction Zone, a Python package for rapid generation of complex nanoscale atomic structures which enables fast, systematic sampling of realistic nanomaterial structures and can be used as a random structure generator for large, diverse synthetic datasets. Using Construction Zone, we develop an end-to-end machine learning workflow for training neural network models to analyze experimental atomic resolution HRTEM images on the task of nanoparticle image segmentation purely with simulated databases. Further, we study the data curation process to understand how various aspects of the curated simulated data—including simulation fidelity, the distribution of atomic structures, and the distribution of imaging conditions—affect model performance across three benchmark experimental HRTEM image datasets. Using our workflow, we are able to achieve state-of-the-art segmentation performance on these experimental benchmarks and, further, we discuss robust strategies for consistently achieving high performance with machine learning in experimental settings using purely synthetic data. Construction Zone and its documentation are available at https://github.com/lerandc/construction_zone.

36 MATERIALS SCIENCE

Geometry-complete diffusion for 3D molecule generation and optimization

Abstract Generative deep learning methods have recently been proposed for generating 3D molecules using equivariant graph neural networks (GNNs) within a denoising diffusion framework. However, such methods are unable to learn important geometric properties of 3D molecules, as they adopt molecule-agnostic and non-geometric GNNs as their 3D graph denoising networks, which notably hinders their ability to generate valid large 3D molecules. In this work, we address these gaps by introducing the Geometry-Complete Diffusion Model (GCDM) for 3D molecule generation, which outperforms existing 3D molecular diffusion models by significant margins across conditional and unconditional settings for the QM9 dataset and the larger GEOM-Drugs dataset, respectively. Importantly, we demonstrate that GCDM’s generative denoising process enables the model to generate a significant proportion of valid and energetically-stable large molecules at the scale of GEOM-Drugs, whereas previous methods fail to do so with the features they learn. Additionally, we show that extensions of GCDM can not only effectively design 3D molecules for specific protein pockets but can be repurposed to consistently optimize the geometry and chemical composition of existing 3D molecules for molecular stability and property specificity, demonstrating new versatility of molecular diffusion models. Code and data are freely available on GitHub .

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Hierarchical screening for Li-based solid electrolytes using fast, interpretable machine-learned potentials

Li-based solid-state electrolyte materials enable safer, all-solid-state batteries but the computational search for candidates with favorable stability and Li-ion conductivity is challenging due to the size of the search space and the cost of evaluating transport properties with ab initio methods. The prohibitive cost of high-throughput screening with DFT has lead to the development of surrogate models using geometric analysis, empirical potentials, and descriptors for ionic transport. Here, I will discuss a hierarchical screening approach for identifying promising materials using a combination of density functional theory, bond-valence methods, and machine learning potentials generated with the Ultra-Fast Force Fields (UF3) framework. We show how the inexpensive bond-valence method can be used to guide the generation of training samples for machine learning, in addition to filtering candidates. Finally, we apply the hierarchical workflow to screen for ionic conductivity across a database of Li-containing compounds.

Materials discovery