Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “deep generative models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Model Intercomparison of the Impacts of Varying Cloud Droplet–Nucleating Aerosols on the Life Cycle and Microphysics of Isolated Deep Convection

The microphysical impacts of aerosol particles on scattered isolated deep convective cells near Houston, Texas, on 19 June 2013, are examined using multiple cloud-resolving model (CRM) simulations initialized with vertical profiles of low and high concentrations of cloud droplet–nucleating aerosols. These simulations formed part of the Model Intercomparison Project (MIP) conducted by the Deep Convective Working Group of the Aerosol, Cloud, Precipitation and Climate (ACPC) initiative. Each CRM generated a field of convective cells representing those observed during the case study with varying degrees of accuracy. The Tracking and Object-Based Analysis of Clouds (tobac) cell-tracking algorithm was applied to each MIP CRM simulation to track relatively long-lived convective cells (20–60 min). Most of the CRMs produced similar aerosol loading impacts on the warm phase of tracked cell properties with reduced autoconversion and accretion growth of rain, increased cloud water, reduced rainfall, and reduced near-surface evaporation of rain. The sign of aerosol impacts on the warm-phase properties of the convective cells was also quite consistent over cell lifetimes with the greatest magnitude of influence in the first half of the life cycle in most CRMs. In contrast, the ice-phase response to aerosol loading was highly variable among CRMs and included increases or decreases in ice amounts at inconsistent stages of the cell life cycle and midlevel versus upper-level changes in ice. This intermodel variability in ice is indicative both of the complex indirect interactions between aerosols and ice-phase processes in deep convection and their associated parameterizations.

Aerosol-cloud interaction↗

Enhancing Convection Parameterization for Next Generation E3SM (Final Technical Report)

This technical report summarizes the achievements during the project period. The general goal of the project is to enhance the representation of atmospheric deep convection in the DOE global climate model E3SM. By implementing a stochastic convection scheme into E3SM, we improved the simulation of precipitation intensity and frequency. As a result, it has a profound impact on the simulation of other fields in the model, including aerosol scavenging by precipitation and aerosol radiative effects. We also implemented a convective cloud microphysics scheme into E3SMv2 to represent more accurately the interaction between convection and grid-scale clouds. An extensive analysis of model simulations as part of the Next Generation Development (NGD) convection scheme assessment project found that including the convective microphysics scheme improved the simulation of Madden-Julian oscillations (MJO) in the model. To make convective parameterization schemes more scale-aware as the GCM resolution increases, we analyzed the output from cloud-resolving model (CRM) simulations of both organized and unorganized convection by coarse-graining the CRM data to different GCM resolutions. Using a novel approach to modify the convective closure, we demonstrated that the convection scheme used in E3SM (ZM scheme) can be made scale-aware.

54 ENVIRONMENTAL SCIENCES↗

Entanglement engineering of optomechanical systems by reinforcement learning

Entanglement is fundamental to quantum information science and technology, yet controlling and manipulating entanglement—so-called entanglement engineering—for arbitrary quantum systems remains a formidable challenge. There are two difficulties: the fragility of quantum entanglement and its experimental characterization. We develop a model-free deep reinforcement-learning (RL) approach to entanglement engineering, in which feedback control together with weak continuous measurement and partial state observation is exploited to generate and maintain desired entanglement. We employ quantum optomechanical systems with linear or nonlinear photon–phonon interactions to demonstrate the workings of our machine-learning-based entanglement engineering protocol. In particular, the RL agent sequentially interacts with one or multiple parallel quantum optomechanical environments, collects trajectories, and updates the policy to maximize the accumulated reward to create and stabilize quantum entanglement over an arbitrary amount of time. The machine-learning-based model-free control principle is applicable to the entanglement engineering of experimental quantum systems in general.

97 MATHEMATICS AND COMPUTING↗

Understanding Relationships Between Satellite, Model, and Ground-Based Surface Temperature Characterizations From Overcast to Clear Conditions in Support of Satellite Remote Sensing of Clouds and Radiation

Accurate and consistent global estimates of cloud coverage and their properties are fundamental to long-term Earth radiation budget (ERB) monitoring efforts like the Clouds and the Earth’s Radiant Energy System (CERES) project. Cloud detection algorithms often apply thresholding approaches to identify where clouds occur by comparing satellite-measured radiances with those that are expected under cloud-free conditions. In addition, once a cloud is detected, the derivation of cloud optical and microphysical properties also requires knowledge of the background radiances below the cloud. In the infrared, knowledge of the surface emissivity and the expected skin temperature under both cloudy and cloud-free conditions is needed. These traits are generally well known over the oceans. Over land, however, comparisons between satellite-derived land surface temperature (LST) with that characterized in numerical weather analyses reveal large differences in many parts of the world, often exceeding 5 K, which can lead to significant satellite cloud detection and cloud property retrieval errors. Furthermore, clouds have a dramatic influence on the LST, and therefore characterization of that model parameter also depends on the capability of the model to accurately resolve clouds. Thus, the LST characterized in models is, at times, a poor approximation for what would otherwise be observed, thereby impeding accurate satellite cloud retrievals. As a result, we seek to develop a more robust method for estimating the LST required for satellite cloud characterizations. This effort is accomplished through a combination of surface emission/air temperature relationship studies in all-sky conditions using ground measurement stations, along with deep neural network (DNN) estimates of expected LST under overcast and cloud-free conditions. We demonstrate that substituting DNN-predicted LST for that generated by numerical models can mitigate model-inherent diurnal dependencies and reduce overall bias and uncertainty relative to satellite/ground observations by 0.5–4 K and 0.5–2 K, respectively. It is expected that this work will lead to improved satellite cloud retrievals that enhance ERB monitoring efforts.

B Scarino↗

hippynn Python Package

hippynn is a python package for defining, training, and applying neural networks to atomistic systems. In particular, it focuses on Hierarchical Interacting Particle Neural Networks (HIP-NNs), a deep learning architecture for atomistic systems. HIP-NNs take input data describing the properties of atomistic systems (often generated using ab-initio quantum mechanics) to learn fast and accurate models. A trained HIP-NN can predict potential energy surfaces, atomic charges, and more. hippynn uses PyTorch for portable high-performance code, including both CPU and GPU support. hippynn allows for extensive customization, including user-defined models and loss functions, to facilitate future research into extensions of HIP-NN as well as other atomistic deep learning models.

Lubbers, Nicolas↗

Encoding Frequency Constraints in Preventive Unit Commitment Using Deep Learning With Region-of-Interest Active Sampling

With the increasing penetration of renewable energy, frequency response and its security are of significant concerns for reliable power system operations. Frequency-constrained unit commitment (FCUC) is proposed to address this challenge. Despite existing efforts in modeling frequency characteristics in unit commitment (UC), current strategies can only handle oversimplified low-order frequency response models and do not consider wide-range operating conditions. This paper presents a generic data-driven framework for FCUC under high renewable penetration. Here, deep neural networks (DNNs) are trained to predict the frequency response using real data or high-fidelity simulation data. Next, the DNN is reformulated as a set of mixed-integer linear constraints to be incorporated into the ordinary UC formulation. In the data generation phase, all possible power injections are considered, and a region-of-interest active sampling is proposed to include power injection samples with frequency nadirs closer to the UFLC threshold, which enhances the accuracy of frequency constraints in FCUC. The proposed FCUC is investigated on the IEEE 39-bus system. Then, a full-order dynamic model simulation using PSS/E verifies the effectiveness of FCUC in frequency-secure generator commitments.

42 ENGINEERING↗

Discovering Ca II absorption lines with a neural network

Quasar absorption line analysis is critical for studying gas and dust components and their physical and chemical properties as well as the evolution and formation of galaxies in the early universe. Calcium II (Ca II ) absorbers, which are one of the dustiest absorbers and are located at lower redshifts than most other absorbers, are especially valuable when studying physical processes and conditions in recent galaxies. However, the number of known quasar Ca II absorbers is relatively low due to the difficulty of detecting them with traditional methods. In this work, we developed an accurate and quick approach to search for Ca II absorption lines using deep learning. In our deep learning model, a convolutional neural network, tuned using simulated data, is used for the classification task. The simulated training data are generated by inserting artificial Ca II absorption lines into original quasar spectra from the Sloan Digital Sky Survey (SDSS), while an existing Ca II catalogue is adopted as the test set. The resulting model achieves an accuracy of 96 per cent on the real data in the test set. Our solution runs thousands of times faster than traditional methods, taking a fraction of a second to analyse thousands of quasars, while traditional methods may take days to weeks. The trained neural network is applied to quasar spectra from SDSS’s DR7 and DR12 and discovered 399 new quasar Ca II absorbers. In addition, we confirmed 409 known quasar Ca II absorbers identified previously by other research groups through traditional methods.

79 ASTRONOMY AND ASTROPHYSICS↗

Next-generation yeast-two-hybrid analysis with Y2H-SCORES identifies novel interactors of the MLA immune receptor

Protein-protein interaction networks are one of the most effective representations of cellular behavior. In order to build these models, high-throughput techniques are required. Next-generation interaction screening (NGIS) protocols that combine yeast two-hybrid (Y2H) with deep sequencing are promising approaches to generate interactome networks in any organism. However, challenges remain to mining reliable information from these screens and thus, limit its broader implementation. Here, we present a computational framework, designated Y2H-SCORES, for analyzing high-throughput Y2H screens. Y2H-SCORES considers key aspects of NGIS experimental design and important characteristics of the resulting data that distinguish it from RNA-seq expression datasets. Three quantitative ranking scores were implemented to identify interacting partners, comprising: 1) significant enrichment under selection for positive interactions, 2) degree of interaction specificity among multi-bait comparisons, and 3) selection of in-frame interactors. Using simulation and an empirical dataset, we provide a quantitative assessment to predict interacting partners under a wide range of experimental scenarios, facilitating independent confirmation by one-to-one bait-prey tests. Simulation of Y2H-NGIS enabled us to identify conditions that maximize detection of true interactors, which can be achieved with protocols such as prey library normalization, maintenance of larger culture volumes and replication of experimental treatments. Y2H-SCORES can be implemented in different yeast-based interaction screenings, with an equivalent or superior performance than existing methods. Proof-of-concept was demonstrated by discovery and validation of novel interactions between the barley nucleotide-binding leucine-rich repeat (NLR) immune receptor MLA6, and fourteen proteins, including those that function in signaling, transcriptional regulation, and intracellular trafficking.

59 BASIC BIOLOGICAL SCIENCES↗

Predictability and feedbacks of the ocean-soil-plant-atmosphere water cycle: deep learning water conductance in Earth System Model

This white paper responds to Focal Area 2. We seek to build predictive models of leaf and surface conductance of water by implementing deep learning (DL) data assimilation techniques. These new models would then be implemented in existing Land Surface Models (LSMs) and Earth System models (ESMs), generating novel water cycle feedbacks. In doing so, we would improve predictability of expected changes in land precipitation, soil moisture, and vegetation dynamics in the long-term, and the role of land cover on the impacts and feedbacks of extreme weather events in the short-term

54 ENVIRONMENTAL SCIENCES↗

Synthetic Infrasound Data for Machine Learning Detectors

Synthetic data is a powerful tool to generate large amounts of training data for machine learning models. The methods outlined in this report will be used to retrain the deep learning classifier for increased accuracy. Synthetic data will be useful to address the natural class imbalance between the different categories in the original ML work. Additionally, these tools will be applied for a variety of signal analysis methods that would use signals with a known signal-to-noise ratio for validation and testing.

58 GEOSCIENCES↗

Electrical and Seismic Response of Saline Permafrost Soil during Freeze - Thaw Transition: Supporting Data

These data were collected during a repeat freeze-thaw cycle on a saline permafrost portion of a core to understand mechanical and geophysical property changes over time. The core used for the experiment was collected from the NGEE-Arctic Site 0 transect site located in the Barrow Environmental Observatory (BEO) in Barrow, Alaska. This region features polygonal ground, and the core was collected in the center of a flat centered polygon. The saline permafrost portion of the core was at the bottom 260-280 cm below ground surface. Data files (*.csv) include: Electrical signal change during freeze-thaw-refreeze cycle at 1 Hz; Spectral data - Complex resistivity data during Freeze-thaw of the saline permafrost core; Seismic data; and CT cross sections (*.jpg) of the core sample from approximately the same location: (A, D) before the test (thawed), (B, E) after the test (frozen), and (C, F) after deep freeze (frozen).The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

The contribution of galaxy clusters to the soft X-ray background

The present paper examines models for the expected contribution to the X-ray background (XRB) from clusters of galaxies in a flat universe with a spectrum of fluctuations corresponding to the standard cold dark matter picture and for a power-law spectrum P(k) varies as k sup n with n = -2. A self-similar scaling model for clusters proposed by Kaiser (1986) is presented, along with several variants. Photon statistics from synthesized maps of deep Einstein fields show a granularity similar to that deduced by Hamilton and Helfand (1987) from the Einstein data but overproduces by a large factor the number of discrete sources observed as clusters, and the fluctuation limit derived from the recent correlation analyses of the background. Two alternative models are designed to be consistent with the local luminosity function data; for these, expected deep images for the Rosat satellite are generated. Clusters of galaxies should be a significant source of background at energies around 2 keV, and it is proposed that clusters of galaxies contribute up to 30 percent of the unresolved component.

Blanchard, A.↗

ML-based Dimension Reduction Strategies

Deep learning (DL)--based surrogate models have achieved success in various applications in carbon capture and storage (CCS). However, the model training on high-dimensional spaces is computationally expensive and impractical for large-scale and complex geological models, because the models usually contain hundreds of thousands to millions of grid cells, each with a set of parameters. Furthermore, the high cost of generating training data with sufficient variation is another limitation of model training on high-dimensional spaces, which may result in overfitting and reduce the model efficiency and prediction performance. We proposed the workflow incorporating dimension reduction methods and deep learning models, which aim to extract the latent variables of input parameters and output state variables, and then build the mapping function at the latent spaces. The proposed workflow can significantly reduce the computational complexity in solving both forward and inverse problems compared to models trained on high-dimensional spaces. Dimensionality reduction models showed great potential in workflows for fast reservoir simulation, history matching, prior model generation, visualization, and more, ultimately enhancing DL model performance in related SMART Work Packages.

Hosseini, Seyyed↗

End-to-end deep learning pipeline for real-time Bragg peak segmentation: from training to large-scale deployment

X-ray crystallography reconstruction, which transforms discrete X-ray diffraction patterns into three-dimensional molecular structures, relies critically on accurate Bragg peak finding for structure determination. As X-ray free electron laser (XFEL) facilities advance toward MHz data rates (1 million images per second), traditional peak finding algorithms that require manual parameter tuning or exhaustive grid searches across multiple experiments become increasingly impractical. While deep learning approaches offer promising solutions, their deployment in high-throughput environments presents significant challenges in automated dataset labeling, model scalability, edge deployment efficiency, and distributed inference capabilities. We present an end-to-end deep learning pipeline with three key components: (1) a data engine that combines traditional algorithms with our peak matching algorithm to generate high-quality training data at scale, (2) a modular architecture that scales from a few million to hundreds of million parameters, enabling us to train large expert-level models offline while deploying smaller, distilled models at the edge, and (3) a decoupled producer-consumer architecture that separates specialized data source layer from model inference, enabling flexible deployment across diverse computing environments. Using this integrated approach, our pipeline achieves accuracy comparable to traditional methods tuned by human experts while eliminating the need for experiment-specific parameter tuning. Although current throughput requires optimization for MHz facilities, our system's scalable architecture and demonstrated model compression capabilities provide a foundation for future high-throughput XFEL deployments.

Wang, Cong↗

Modeling Spatial Asymmetries in Teleconnected Extreme Temperatures

Abstract Combining strengths from deep learning and extreme value theory can help describe complex relationships between variables where extreme events have significant impacts (e.g., environmental or financial applications). Neural networks learn complicated nonlinear relationships from large datasets under limited parametric assumptions. By definition, the number of occurrences of extreme events is small, which limits the ability of the data-hungry, nonparametric neural network to describe rare events. Inspired by recent extreme cold winter weather events in North America caused by atmospheric blocking, we examine several probabilistic generative models for the entire multivariate probability distribution of daily boreal winter surface air temperature. We propose metrics to measure spatial asymmetries, such as long-range anticorrelated patterns that commonly appear in temperature fields during blocking events. Compared to vine copulas, the statistical standard for multivariate copula modeling, deep learning methods show improved ability to reproduce complicated asymmetries in the spatial distribution of ERA5 temperature reanalysis, including the spatial extent of in-sample extreme events.

Krock, Mitchell L.↗

Composition of Orientale basin deposits and implications for the lunar basin-forming process

The geologic, spectral and geochemical characteristics of the lunar Orientale basin are discussed, and a model is defined for the generation of Orientale basin deposits. The data indicate that the basin ejecta is composed mainly of anorthositic deposits and no crustal material. The crater was originally 500-600 km across and 50-60 km deep, the latter being too shallow to reach the projected 100 km crustal depth. A proportional growth model is judged acceptable for the Orientale basin. Finally, it is concluded that neither the Apollo 14 nor 16 missions obtained Orientale ejecta material, which could in any case be unidentifiable until more thorough samplings are made of a large portion of the lunar surface.

Spudis, P. D.↗

Measurements and Model Improvement: Insight into NWP Model Error Using Doppler Lidar and Other WFIP2 Measurement Systems

Abstract Doppler-lidar wind-profile measurements at three sites were used to evaluate NWP model errors from two versions of NOAA’s 3-km-grid HRRR model, to see whether updates in the latest version 4 reduced errors when compared against the original version 1. Nested (750-m grid) versions of each were also tested to see how grid spacing affected forecast skill. The measurements were part of the field phase of the Second Wind Forecasting Improvement Project (WFIP2), an 18-month deployment into central Oregon–Washington, a major wind-energy-producing region. This study focuses on errors in simulating marine intrusions, a summertime, 600–800-m-deep, regional sea-breeze flow found to generate large errors. HRRR errors proved to be complex and site dependent. The most prominent error resulted from a premature drop in modeled marine-intrusion wind speeds after local midnight, when lidar-measured winds of greater than 8 m s −1 persisted through the next morning. These large negative errors were offset at low levels by positive errors due to excessive mixing, complicating the interpretation of model “improvement,” such that the updates to the full-scale versions produced mixed results, sometimes enhancing but sometimes degrading model skill. Nesting consistently improved model performance, with version 1’s nest producing the smallest errors overall. HRRR’s ability to represent the stages of sea-breeze forcing was evaluated using radiation budget, surface-energy balance, and near-surface temperature measurements available during WFIP2. The significant site-to-site differences in model error and the complex nature of these errors mean that field-measurement campaigns having dense arrays of profiling sensors are necessary to properly diagnose and characterize model errors, as part of a systematic approach to NWP model improvement. Significance Statement Dramatic increases in NWP model skill will be required over the coming decades. This paper describes the role of major deployments of accurate profiling sensors in achieving that goal and presents an example from the Second Wind Forecast Improvement Program (WFIP2). Wind-profile data from scanning Doppler lidars were used to evaluate two versions of HRRR, the original and an updated version, and nested versions of each. This study focuses on the ability of updated HRRR versions to improve upon predicting a regional sea-breeze flow, which was found to generate large errors by the original HRRR. Updates to the full-scale HRRR versions produced mixed results, but the finer-mesh versions consistently reduced model errors.

Meteorology & Atmospheric Sciences↗

RNA language models predict mutations that improve RNA function

Structured RNA lies at the heart of many central biological processes, from gene expression to catalysis. RNA structure prediction is not yet possible due to a lack of high-quality reference data associated with organismal phenotypes that could inform RNA function. We present GARNET (Gtdb Acquired RNa with Environmental Temperatures), a new database for RNA structural and functional analysis anchored to the Genome Taxonomy Database (GTDB). GARNET links RNA sequences to experimental and predicted optimal growth temperatures of GTDB reference organisms. Using GARNET, we develop sequence- and structure-aware RNA generative models, with overlapping triplet tokenization providing optimal encoding for a GPT-like model. Leveraging hyperthermophilic RNAs in GARNET and these RNA generative models, we identify mutations in ribosomal RNA that confer increased thermostability to the Escherichia coli ribosome. The GTDB-derived data and deep learning models presented here provide a foundation for understanding the connections between RNA sequence, structure, and function.

59 BASIC BIOLOGICAL SCIENCES↗