Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Generative models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Hierarchical transfer learning: an agile and equitable strategy for machine-learning interatomic models

Machine-learned interatomic models are growing in popularity due to their ability to afford near quantum-accurate predictions for complex phenomena with orders-of-magnitude greater computational efficiency. However, these models struggle when applied to systems of many element types due to the approximately exponential increase in number of parameters that must be determined. To mitigate this challenge, we present a new hierarchical transfer learning approach that allows the fitting problem to be decomposed into smaller independent and reusable parameter blocks that enable development of explicitly chemically extensible ML-IAM. Application of this strategy is demonstrated for C and N mixtures under conditions ranging from nominally ambient to ~10,000 K and 200 GPa for compositions from 0 to 100% N. Ultimately, this strategy makes model generation for chemically complex systems more tractable and efficient, facilitates comprehensive model validation, and makes ML-IAM development for problems of this nature more accessible to users with limited access to extreme computing infrastructure.

Lindsey, Rebecca K. [Univ. of Michigan, Ann Arbor,↗

ATAT: Astronomical Transformer for time series and Tabular data

Context. The advent of next-generation survey instruments, such as theVera C. RubinObservatory and its Legacy Survey of Space and Time (LSST), is opening a window for new research in time-domain astronomy. The Extended LSST Astronomical Time-Series Classification Challenge (ELAsTiCC) was created to test the capacity of brokers to deal with a simulated LSST stream. Aims. Our aim is to develop a next-generation model for the classification of variable astronomical objects. We describe ATAT, the Astronomical Transformer for time series And Tabular data, a classification model conceived by the ALeRCE alert broker to classify light curves from next-generation alert streams. ATAT was tested in production during the first round of the ELAsTiCC campaigns. Methods. ATAT consists of two transformer models that encode light curves and features using novel time modulation and quantile feature tokenizer mechanisms, respectively. ATAT was trained on different combinations of light curves, metadata, and features calculated over the light curves. We compare ATAT against the current ALeRCE classifier, a balanced hierarchical random forest (BHRF) trained on human-engineered features derived from light curves and metadata. Results. When trained on light curves and metadata, ATAT achieves a macro F1 score of 82.9 ± 0.4 in 20 classes, outperforming the BHRF model trained on 429 features, which achieves a macro F1 score of 79.4 ± 0.1. Conclusions. The use of transformer multimodal architectures, combining light curves and tabular data, opens new possibilities for classifying alerts from a new generation of large etendue telescopes, such as theVera C. RubinObservatory, in real-world brokering scenarios.

Astronomy & Astrophysics↗

Thresholding Analysis and Feature Extraction from 3D Ground Penetrating Radar Data for Noninvasive Assessment of Peanut Yield

This study explores the efficacy of utilizing a novel ground penetrating radar (GPR) acquisition platform and data analysis methods to quantify peanut yield for breeding selection, agronomic research, and producer management and harvest applications. Sixty plots comprising different peanut market types were scanned with a multichannel, air-launched GPR antenna. Image thresholding analysis was performed on 3D GPR data from four of the channels to extract features that were correlated to peanut yield with the objective of developing a noninvasive high-throughput peanut phenotyping and yield-monitoring methodology. Plot-level GPR data were summarized using mean, standard deviation, sum, and the number of nonzero values (counts) below or above different percentile threshold values. Best results were obtained for data below the percentile threshold for mean, standard deviation and sum. Data both below and above the percentile threshold generated good correlations for count. Correlating individual GPR features to yield generated correlations of up to 39% explained variability, while combining GPR features in multiple linear regression models generated up to 51% explained variability. The correlations increased when regression models were developed separately for each peanut type. This research demonstrates that a systematic search of thresholding range, analysis window size, and data summary statistics is necessary for successful application of this type of analysis. The results also establish that thresholding analysis of GPR data is an appropriate methodology for noninvasive assessment of peanut yield, which could be further developed for high-throughput phenotyping and yield-monitoring, adding a new sensor and new capabilities to the growing set of digital agriculture technologies.

54 ENVIRONMENTAL SCIENCES↗

Stochastic representation and conditioning of process-based geological model by deep generative and recognition networks

Accurate and realistic geological modeling is the core of oil and gas development and production. In recent years, process-based methods are developed to produce highly realistic geological models by simulating the physical processes that reproduce the sedimentary events and develop the geometry. However, the complex dynamic processes are extremely expensive to simulate, making process-based models difficult to be conditioned to field data. In this work, we propose a comprehensive generative adversarial network framework as a machine-learning-assisted approach for mimicking the outputs of process-based geological models with fast generation. The main objective of our work is to obtain a continuous parametrization of the highly realistic process-based geological models which enables us to calibrate the models and condition the models to data. Numerical results are presented to illustrate the capability of our proposed methodology.

58 GEOSCIENCES↗

dGen (Distributed Generation Market Demand) Model Data: Alpha Release

Open sourced data needed to run the basic alpha release version of the dGen model. Includes a pre-generated agent file of 100,000 agents in pickle file format along with the base schema and table data in parquet format that are needed to create a postgreSQL database for the model to interact with.

14 SOLAR ENERGY↗

The impact of coupled air–sea interaction on extreme East Asian summer monsoon simulation in CMIP5 models

Abstract In this study, the relationship between the ability to simulate air–sea interactions over the western North Pacific (WNP), and to reproduce the extreme East Asian summer monsoon (EASM), were investigated by comparing the performances of several global climate models (GCMs). High ranked in air–sea interaction simulation (HRA) and low ranked in air–sea interaction simulation (LRA) models were selected, according to their performance in simulating relations between sea surface temperature (SST) and precipitation over the WNP, from the ensemble of models that participated in the third and fifth phases of the Coupled Model Intercomparison Project (CMIP3, CMIP5). Compared with CMIP3 models, CMIP5 models exhibited improved simulations of the distinctive air–sea interaction over the WNP, namely, the strong atmospheric forcing on the ocean. Among CMIP5 models, HRA models, which reproduced intrinsic negative correlations between precipitation and SST over the WNP, could simulate the extreme EASM better than LRA models. In particular, HRA models generated a more realistic spatial distribution of the extreme EASM compared with LRA models. The defects of the LRA models resulted from distorted synoptic fields, including underestimated geopotential height and overestimated low‐level wind over the WNP, inducing unrealistic moisture supply and convection due to the exaggerated SST forcing. In contrast, reasonable air–sea interactions represented in HRA models lead to realistic synoptic fields over the WNP, and proper simulation of the extreme EASM.

Kim, Taehyung↗

Feasibility of Formulating Ecosystem Biogeochemical Models From Established Physical Rules

Abstract To improve the predictive capability of ecosystem biogeochemical models (EBMs), we discuss the feasibility of formulating biogeochemical processes using physical rules that have underpinned the many successes in computational physics and chemistry. We argue that the currently popular empirically based approaches, such as multiplicative empirical response functions and the law of the minimum, will not lead to EBM formulations that can be continuously refined to incorporate improved mechanistic understanding and empirical observations of biogeochemical processes. Instead, we propose that EBM parameterizations, as a lossy data compression problem, can be better formulated using established physical rules widely used in computational physics and chemistry, and different biogeochemical processes can be more robustly integrated within a reactive‐transport framework. Through several examples, we demonstrate how mathematical representations derived from physical rules can improve understanding of relevant biogeochemical processes and enable more effective communication between modelers, observationalists, and experimentalists regarding essential questions, such as what measurements are needed to meaningfully inform models and how can models generate new process‐level hypotheses to test in empirical studies. Finally, while empirical models with more parameters are often less robust, physical rules‐based models can be more robust and show lower predictive equifinality, stemming from their enhanced consistency in representations of processes, interactions and spatial scaling.

54 ENVIRONMENTAL SCIENCES↗

Physics-informed semantic inpainting: Application to geostatistical modeling

A fundamental problem in geostatistical modeling is to infer the heterogeneous geological field based on limited measurements and some prior spatial statistics. Semantic inpainting, a technique for image processing using deep generative models, has been recently applied for this purpose, demonstrating its effectiveness in dealing with complex spatial patterns. However, the original semantic inpainting framework incorporates only information from direct measurements, while in geostatistics indirect measurements are often plentiful. In this work, to overcome this limitation, we propose a physics-informed semantic inpainting framework, employing the Wasserstein Generative Adversarial Network with Gradient Penalty (WGAN-GP) and jointly incorporating the direct and indirect measurements by exploiting the underlying physical laws. Our simulation results for a high-dimensional problem with 512 dimensions show that in the new method, the physical conservation laws are satisfied and contribute in enhancing the inpainting performance compared to using only the direct measurements.

54 ENVIRONMENTAL SCIENCES↗

Modeling household-level party composition behavior for multiparty activities: a random parameter nested logit modeling approach

This study presents findings of a household-level party composition model for multiparty activities. It exploits data from a comprehensive Household Travel Survey conducted by Chicago Metropolitan Agency of Planning. The study estimates a random parameter nested logit model to capture households’ unobserved preference heterogeneity and non-proportional substitution patterns in terms of activity party composition for multiparty activities. A wide variety of household demographics, activity attributes and residential neighborhood characteristics are examined in this paper. The magnitude of the impacts of the determinants are tested in this study by analyzing the elasticity of the variables, which suggests that household demographics and attributes of the multiparty activities have significant effects on the household-level activity party composition. Residential neighborhood characteristics, although somewhat less impactful, still play a meaningful role. This model will be implemented within the POLARIS transportation systems simulator to improve the activity generation modeling workflow, and the prediction accuracy of various activity-travel components.

activity party composition↗

Proposal for New Plant Controller and Electrical Controller

This is the first revision of this memo capturing in more detail some items discussed at the last WECC MVS meeting in January, 2022 during the REMWG part of the meeting. The items here are presented for further enhancements in the electrical controls model and plant controller model (specifically for hybrid-plants or plants with multiple aggregated inverter-based generation models). These will need to be discussed and refined, and then implemented for benchmark testing and final approval.

42 ENGINEERING↗

Implementation of INCL nuclear model in GENIE Generator

The Liège Intranuclear Cascade (INCL) model is a nuclear-physics model that simulates hadron (baryon, anti-baryon and meson) reactions on nuclei, for incident energies ranging from a few tens of MeV to 10-20 GeV. The INCL model has been well validated by hadron scattering data. In my work, I implement an interface in GENIE to use the INCL nuclear model in the simulations of both the initial state of the target nucleus and the Final State Interaction in neutrino-nucleus interaction. It has a consistent treatment of nuclear models in both neutrino interaction and hadron rescattering. A full event record including neutrino vertex and each vertex of hadron rescattering has been accomplished. Several processes, e.g. cluster production, Delta transportation and de-excitation will also be included as benefits of the implementation of the INCL model in GENIE. I will show some initial simulation results showcasing the new GENIE features and discuss plans for making them available for use in experimental analyses.

Liu, Liang [Fermilab] (ORCID:000000026753925X)↗

Near-Infrared Spectroscopy can Predict Anatomical Abundance in Corn Stover

Feedstock heterogeneity is a key challenge impacting the deconstruction and conversion of herbaceous lignocellulosic biomass to biobased fuels, chemicals, and materials. Upstream processing to homogenize biomass feedstock streams into their anatomical components via air classification allows for a more tailored approach to subsequent mechanical and chemical processing. Here, we show that differing corn stover anatomical tissues respond differently to pretreatment and enzymatic hydrolysis and therefore, a one-size-fits-all approach to chemical processing biomass is inappropriate. To inform on-line downstream processing, a robust and high-throughput analytical technique is needed to quantitatively characterize the separated biomass. Predictive correlation of near-infrared spectra to biomass chemical composition is such a technique. Here, we demonstrate the capability of models developed using an “off-the-shelf,” industrially relevant spectrometer with limited spectral range to make strong predictions of both cell wall chemical composition and the relative abundance of anatomical components of the corn stover, the latter for the first time ever. Gaussian process regression (GPR) yields stronger correlations (average R 2 v = 88% for chemical composition and 95% for anatomical relative abundance) than the more commonly used partial least squares (PLS) regression (average R 2 v = 84% for chemical composition and 92% for anatomical relative abundance). In nearly all cases, both GPR and PLS outperform models generated using neural networks. These results highlight the potential for coupling NIRS with predictive models based on GPR due to the potential to yield more robust correlations.

09 BIOMASS FUELS↗

Debiasing with Diffusion: Probabilistic Reconstruction of Dark Matter Fields from Galaxies with CAMELS

Abstract Galaxies are biased tracers of the underlying cosmic web, which is dominated by dark matter (DM) components that cannot be directly observed. Galaxy formation simulations can be used to study the relationship between DM density fields and galaxy distributions. However, this relationship can be sensitive to assumptions in cosmology and astrophysical processes embedded in galaxy formation models, which remain uncertain in many aspects. In this work, we develop a diffusion generative model to reconstruct DM fields from galaxies. The diffusion model is trained on the CAMELS simulation suite that contains thousands of state-of-the-art galaxy formation simulations with varying cosmological parameters and subgrid astrophysics. We demonstrate that the diffusion model can predict the unbiased posterior distribution of the underlying DM fields from the given stellar density fields while being able to marginalize over uncertainties in cosmological and astrophysical models. Interestingly, the model generalizes to simulation volumes ≈500 times larger than those it was trained on and across different galaxy formation models. The code for reproducing these results can be found athttps://github.com/victoriaono/variational-diffusion-cdm✎.

Astronomy & Astrophysics↗

ERA5-Land Data for LASSO-CACTI Overview Paper

The European Centre for Medium-Range Weather Forecasts (ECMWF) generated a soil reanalysis dataset for the land component of the fifth generation of European ReAnalysis (ERA5), referred to as ERA5-Land. This is a model-generated dataset, with the original version available for the period 1950 to present. The version archived in this DOE ARM product is a subset of the data is for the period of the CACTI field campaign plus several preceding months, specifically from August 1, 2018 through March 22, 2019 with hourly intervals. The ARM copy is also a sub-region of the original global product; the ARM copy is for -60 to -5 °N by -105 to -30 °W. Only variables necessary to drive the WRF-Hydro model are included, which are the 2-m temperature and specific humidity, 10-m wind components, surface pressure, rain rate, and downward surface short and longwave radiation. These data have been obtained from the Copernicus Data Store.

10m wind u-component↗

Ephemeral Learning - Augmenting Triggers with Online-Trained Normalizing Flows

The large data rates at the LHC require an online trigger system to select relevant collisions. Rather than compressing individual events, we propose to compress an entire data set at once. We use a normalizing flow as a deep generative model to learn the probability density of the data online. The events are then represented by the generative neural network and can be inspected offline for anomalies or used for other analysis purposes. We demonstrate our new approach for a toy model and a correlation-enhanced bump hunt.

97 MATHEMATICS AND COMPUTING↗

Application of the metabolic modeling pipeline in KBase to categorize reactions, predict essential genes, and predict pathways in an isolate genome

The DOE Systems Biology Knowledgebase (KBase) platform offers a range of powerful tools for the reconstruction, refinement, and analysis of genome-scale metabolic models built from microbial isolate genomes. In this chapter, we describe and demonstrate these tools in action with an analysis of isoprene production in the Bacillus subtilis DSM genome. Two different methods are applied to build initial metabolic models for the DSM genome, then the models are gapfilled in three different growth conditions. Next, flux balance analysis (FBA) and flux variability analysis (FVA) techniques are applied to both study the growth of these models in minimal media and classify reactions within each model based on essentiality and functionality. The models are applied with the FBA method to predict essential genes, which are then compared to an updated list of essential genes obtained for B. subtilis 168, a very similar strain to the DSM isolate. The models are also applied to simulate Biolog growth conditions, and these results are compared with Biolog data collected for B. subtilis 168. Finally, the DSM metabolic models are applied to explore the pathways and genes responsible for producing isoprene in this strain. These studies demonstrate the accuracy and utility of models generated from the KBase pipelines, as well as exploring the tools available for analyzing these models.

DOE knowledgebase↗

Unsupervised probabilistic models for sequential Electronic Health Records

We develop an unsupervised probabilistic model for heterogeneous Electronic Health Record (EHR) data. Utilizing a mixture model formulation, our approach directly models sequences of arbitrary length, such as medications and laboratory results. This allows for subgrouping and incorporation of the dynamics underlying heterogeneous data types. The model consists of a layered set of latent variables that encode underlying structure in the data. These variables represent subject subgroups at the top layer, and unobserved states for sequences in the second layer. We train this model on episodic data from subjects receiving medical care in the Kaiser Permanente Northern California integrated healthcare delivery system. The resulting properties of the trained model generate novel insight from these complex and multifaceted data. In addition, we show how the model can be used to analyze sequences that contribute to assessment of mortality likelihood.

59 BASIC BIOLOGICAL SCIENCES↗

Application of the Metabolic Modeling Pipeline in KBase to Categorize Reactions, Predict Essential Genes, and Predict Pathways in an Isolate Genome

The DOE Systems Biology Knowledgebase (KBase) platform offers a range of powerful tools for the reconstruction, refinement, and analysis of genome-scale metabolic models built from microbial isolate genomes. In this chapter, we describe and demonstrate these tools in action with an analysis of isoprene production in the Bacillus subtilis DSM genome. Two different methods are applied to build initial metabolic models for the DSM genome, then the models are gapfilled in three different growth conditions. Next, flux balance analysis (FBA) and flux variability analysis (FVA) techniques are applied to both study the growth of these models in minimal media and classify reactions within each model based on essentiality and functionality. The models are applied with the FBA method to predict essential genes, which are then compared to an updated list of essential genes obtained for B. subtilis 168, a very similar strain to the DSM isolate. The models are also applied to simulate Biolog growth conditions, and these results are compared with Biolog data collected for B. subtilis 168. Finally, the DSM metabolic models are applied to explore the pathways and genes responsible for producing isoprene in this strain. These studies demonstrate the accuracy and utility of models generated from the KBase pipelines, as well as exploring the tools available for analyzing these models.

Allen, Benjamin↗