Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “deep generative models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Synthetic data-driven deep learning for label-free autonomous atomic force microscopy

Atomic force microscopy (AFM) is a widely used tool for nanoscale characterization across materials science, energy research, and biology. However, its adoption in high-throughput materials discovery and statistically driven studies remains limited by a strong dependence on expert operator input and by the scarcity of annotated experimental AFM datasets needed to enable data-driven automation. Here, we introduce SimuScan, a synthetic-data–driven framework that enables reliable AFM feature identification, segmentation, and targeted imaging without requiring large manually labeled experimental datasets. SimuScan generates tunable, high-fidelity synthetic AFM images of defined morphologies while incorporating realistic experimental artifacts, including tip–sample convolution, noise, flattening distortions, and surface debris. These datasets are shown to support scalable, label-free training of modern deep learning models for AFM analysis. When integrated into data-driven AFM workflows, SimuScan-trained models can locate and analyze nanoscale structures across large datasets and guide targeted follow-up imaging. We validate this approach on nanostructured surfaces, DNA assemblies, and bacterial cells, demonstrating robust generalization across diverse sample types with minimal operator intervention. More broadly, this work establishes a general strategy for generating explicitly conditioned, task-relevant synthetic data to improve the reliability of downstream models in autonomous microscopy.

Millan-Solsona, Ruben [Oak Ridge National Laborato↗

Data and Code for Understanding Generative AI Content with Embedding Models

This repository contains code for the experiments in the paper "Understanding Generative AI Content with Embedding Models". Constructing high-quality features is critical to any quantitative data analysis. While feature engineering was historically addressed by carefully hand-crafting data representations based on domain expertise, deep neural networks (DNNs) now offer a radically different approach. DNNs implicitly engineer features by transforming their input data into hidden feature vectors called embeddings. For embedding vectors produced by foundation models -- which are trained to be useful across many contexts -- we demonstrate that simple and well-studied dimensionality-reduction techniques such as Principal Component Analysis uncover inherent heterogeneity in input data concordant with human-understandable explanations. Of the many applications for this framework, we find empirical evidence that there is intrinsic separability between real samples and those generated by artificial intelligence (AI).

Vargas, Max [Pacific Northwest National Laboratory↗

How well are hazards associated with derechos reproduced in regional climate simulations?

Abstract. A 15-member ensemble of convection-permitting regional simulations of the fast-moving and destructive derecho of 29–30 June 2012 that impacted the northeastern urban corridor of the USA is presented. This event generated 1100 reports of damaging winds, generated significant wind gusts over an extensive area of up to 500 000 km2, caused several fatalities, and resulted in widespread loss of electrical power. Extreme events such as this are increasingly being used within pseudo-global-warming experiments to examine the sensitivity of historical, societally important events to global climate non-stationarity and how they may evolve as a result of changing thermodynamic and dynamic contexts. As such it is important to examine the fidelity with which such events are described in hindcast experiments. The regional simulations presented herein are performed using the Weather Research and Forecasting (WRF) model. The resulting ensemble is used to explore simulation fidelity relative to observations for wind gust magnitudes, spatial scales of convection (as is manifest in high composite reflectivity, cREF), and both rainfall and hail production as a function of model configuration (microphysics parameterization, lateral boundary conditions (LBCs), start date, use of nudging, compiler choice, damping, and number of vertical levels). We also examine the degree to which each ensemble member differs with respect to key mesoscale drivers of convective systems (e.g., convective available potential energy and vertical wind shear) and critical manifestations of deep convection, e.g., vertical velocities, cold-pool generation, and how those properties relate to the correct characterization of the associated atmospheric hazards (wind gusts and hail). Use of a double-moment, seven-class scheme with number concentrations for all species (including hail and graupel) results in the greatest fidelity of model-simulated wind gusts and convective structure to the observations of this event. All ensemble members, however, fail to capture the intensity of the event in terms of the spatial extent of convection and the production of high near-surface wind gusts. We further show very high sensitivity to the LBCs employed and specifically that simulation fidelity is higher for simulations nested within ERA-Interim compared to ERA5. Excess convective available potential energy (CAPE) in all ensemble members after the derecho passage leads to excess production of convective cells, wind gusts, cREF > 40 dBZ, and precipitation during a frontal passage on the subsequent day. This event proved very challenging to forecast in real time and to reproduce in the 15-member hindcast simulation ensemble presented here. Future work could examine if simulations with other initial and lateral boundary conditions can achieve greater fidelity.

Shepherd, Tristan (ORCID:0000000186276419)↗

Artificial Intelligence Transforming Post-Translational Modification Research

Post-Translational Modifications (PTMs) are covalent changes to amino acids that occur after protein synthesis, including covalent modifications on side chains and peptide backbones. Many PTMs profoundly impact cellular and molecular functions and structures, and their significance extends to evolutionary studies as well. In light of these implications, we have explored how artificial intelligence (AI) can be utilized in researching PTMs. Initially, rationales for adopting AI and its advantages in understanding the functions of PTMs are discussed. Then, various deep learning architectures and programs, including recent applications of language models, for predicting PTM sites on proteins and the regulatory functions of these PTMs are compared. Finally, our high-throughput PTM-data-generation pipeline, which formats data suitably for AI training and predictions is described. We hope this review illuminates areas where future AI models on PTMs can be improved, thereby contributing to the field of PTM bioengineering.

59 BASIC BIOLOGICAL SCIENCES↗

Classification of Clouds and Deep Convection from GEOS-5 Using Satellite Observations

With the increased resolution of global atmospheric models and the push toward global cloud resolving models, the resemblance of model output to satellite observations has become strikingly similar. As we progress with our adaptation of the Goddard Earth Observing System Model, Version 5 (GEOS-5) as a high resolution cloud system resolving model, evaluation of cloud properties and deep convection require in-depth analysis beyond a visual comparison. Outgoing long-wave radiation (OLR) provides a sufficient comparison with infrared (IR) satellite imagery to isolate areas of deep convection. We have adopted a binning technique to generate a series of histograms for OLR which classify the presence and fraction of clear sky versus deep convection in the tropics that can be compared with a similar analyses of IR imagery from composite Geostationary Operational Environmental Satellite (GOES) observations. We will present initial results that have been used to evaluate the amount of deep convective parameterization required within the model as we move toward cloud system resolving resolutions of 10- to 1-km globally.

Putman, William↗

Machine learning–assisted prediction of heat fluxes through thermally anisotropic building envelopes

Thermally anisotropic building envelope (TABE) is a novel active building envelope that can save energy use to maintain thermal comfort in buildings by redirecting heat and coolness from building envelopes to thermal loops. Finite element models (FEMs) can be used to compute the heat fluxes through TABEs, but the high computational cost of finite element simulations has prevented parametric studies and design optimizations. This paper proposes a domain knowledge–informed, finite element–based machine learning framework to reduce the computation cost for the energy management of buildings installed with TABE that uses a ground thermal loop. First, the training heat flux data set was generated by FEM simulations with different thermal loop schedules. Then, both shallow learning models (i.e., multivariate linear regression and eXtreme Gradient Boost, or XGBoost) and a deep learning model (i.e., deep neural network, or DNN) were trained to predict the heat fluxes. Domain knowledge was used for data preprocessing and feature selection. Finally, the suitability of the selected machine learning model was tested under different thermal loop schedules. Herein, the case study results showed that: (1) XGBoost can be as accurate as DNN (coefficient of determination equal to 0.81) with much less training time; (2) the annual energy cost savings for different thermal loop schedules obtained by the XGBoost-predicted and FEM-calculated heat fluxes are consistent, having a difference of only 4%; and (3) XGBoost can reduce the computation time for the annual energy analysis of the case study building with a given thermal loop schedule from around 12 h by using FEM to less than 1 min.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Nanoindentation mapping defects filtration for heterogeneous materials using generative adversarial networks

Advanced composite materials with multiple phases and heterogeneous microstructure necessitate spatial mapping characterization of elastic modulus to develop constitutive relations and overall mechanical response. Such modulus mapping can be obtained using the nanoindentation technique, where the indenter tip raster over the selected microstructure region. Typically, a surface preparation procedure is done in the specimens to ensure proper contact between the indenter tip and sample surface. However, a near-perfect surface finish is unachievable in heterogeneous materials, primarily with ceramic reinforcements, due to the differential material removal rate during polishing. Thus, the nanoindenter records localized erroneous measurements due to differences in surface roughness and corresponding force response. This study establishes a novel deep learning-based strategy to rectify incorrect experimental spatial measurements acquire during nanoindentation modulus mapping. Here, the integrated bicubic interpolation and generative adversarial networks (GANs) model was trained using 14 ceramic and 18 metallic data sets, each comprising 65,536 measurements. The developed algorithm was validated against experimental measurements on four unknown specimens. The standard deviation in measured elastic modulus reduces by ~50% in ceramics and ~72% in metallic samples. This computational framework proposes a novel approach to reducing uncertainty in materials’ properties using state-of-the-art computer vision techniques.

36 MATERIALS SCIENCE↗

Model metamers reveal divergent invariances between biological and artificial neural networks

Deep neural network models of sensory systems are often proposed to learn representational transformations with invariances like those in the brain. To reveal these invariances, we generated ‘model metamers’, stimuli whose activations within a model stage are matched to those of a natural stimulus. Metamers for state-of-the-art supervised and unsupervised neural network models of vision and audition were often completely unrecognizable to humans when generated from late model stages, suggesting differences between model and human invariances. Targeted model changes improved human recognizability of model metamers but did not eliminate the overall human–model discrepancy. The human recognizability of a model’s metamers was well predicted by their recognizability by other models, suggesting that models contain idiosyncratic invariances in addition to those required by the task. Metamer recognizability dissociated from both traditional brain-based benchmarks and adversarial vulnerability, revealing a distinct failure mode of existing sensory models and providing a complementary benchmark for model assessment.

59 BASIC BIOLOGICAL SCIENCES↗

Interpretable Data-Driven Probabilistic Power System Load Margin Assessment with Uncertain Renewable Energy and Loads

The increasing uncertainties caused by the high-penetration of stochastic renewable generation resources poses a significant threat to the power system voltage stability. To address this issue, this paper proposes a probabilistic deep kernel learning enabled surrogate model to extract the hidden relationship between uncertain sources, i.e., wind power and loads, and load margin for probabilistic load margin assessment (PLMA). Unlike other deep learning approaches, a kernel SHAP provides the sensitivity analysis as well as interpretability of the inputs to outputs influences. This allows identifying the critical factors that affect load margin so that corrective control can be initiated for stability enhancement. Numerical results carried out on the IEEE 118-bus power system demonstrate the accuracy and efficiency of the proposed data-driven PLMA scheme.

deep kernel learning↗

Super-resolution and segmentation deep learning for breast cancer histopathology image analysis

Traditionally, a high-performance microscope with a large numerical aperture is required to acquire high-resolution images. However, the images’ size is typically tremendous. Therefore, they are not conveniently managed and transferred across a computer network or stored in a limited computer storage system. As a result, image compression is commonly used to reduce image size resulting in poor image resolution. Here, we demonstrate custom convolution neural networks (CNNs) for both super-resolution image enhancement from low-resolution images and characterization of both cells and nuclei from hematoxylin and eosin (H&E) stained breast cancer histopathological images by using a combination of generator and discriminator networks so-called super-resolution generative adversarial network-based on aggregated residual transformation (SRGAN-ResNeXt) to facilitate cancer diagnosis in low resource settings. The results provide high enhancement in image quality where the peak signal-to-noise ratio and structural similarity of our network results are over 30 dB and 0.93, respectively. The derived performance is superior to the results obtained from both the bicubic interpolation and the well-known SRGAN deep-learning methods. In addition, another custom CNN is used to perform image segmentation from the generated high-resolution breast cancer images derived with our model with an average Intersection over Union of 0.869 and an average dice similarity coefficient of 0.893 for the H&E image segmentation results. Finally, we propose the jointly trained SRGAN-ResNeXt and Inception U-net Models, which applied the weights from the individually trained SRGAN-ResNeXt and inception U-net models as the pre-trained weights for transfer learning. The jointly trained model’s results are progressively improved and promising. We anticipate these custom CNNs can help resolve the inaccessibility of advanced microscopes or whole slide imaging (WSI) systems to acquire high-resolution images from low-performance microscopes located in remote-constraint settings.

60 APPLIED LIFE SCIENCES↗

Open Data and Deep Semantic Segmentation for Automated Extraction of Building Footprints

Advances in machine learning and computer vision, combined with increased access to unstructured data (e.g., images and text), have created an opportunity for automated extraction of building characteristics, cost-effectively, and at scale. These characteristics are relevant to a variety of urban and energy applications, yet are time consuming and costly to acquire with today’s manual methods. Several recent research studies have shown that in comparison to more traditional methods that are based on features engineering approach, an end-to-end learning approach based on deep learning algorithms significantly improved the accuracy of automatic building footprint extraction from remote sensing images. However, these studies used limited benchmark datasets that have been carefully curated and labeled. How the accuracy of these deep learning-based approach holds when using less curated training data has not received enough attention. The aim of this work is to leverage the openly available data to automatically generate a larger training dataset with more variability in term of regions and type of cities, which can be used to build more accurate deep learning models. In contrast to most benchmark datasets, the gathered data have not been manually curated. Thus, the training dataset is not perfectly clean in terms of remote sensing images exactly matching the ground truth building’s foot-print. A workflow that includes data pre-processing, deep learning semantic segmentation modeling, and results post-processing is introduced and applied to a dataset that include remote sensing images from 15 cities and five counties from various region of the USA, which include 8,607,677 buildings. The accuracy of the proposed approach was measured on an out of sample testing dataset corresponding to 364,000 buildings from three USA cities. The results favorably compared to those obtained from Microsoft’s recently released US building footprint dataset.

97 MATHEMATICS AND COMPUTING↗

The star formation history of the Large Magellanic Cloud

Deep photometric observations of stars in three fields of the LMC are presented, and these data are interpreted using synthetic CMDs and LFs generated from overshoot models. The field CMDs and LFs with a star formation rate that experienced a large increase (4 +/- 0.5) x 10 exp 9 yr ago is successfully modeled. The precise age of this 'burst' depends sensitively on the characteristics of the models. Classical (i.e., nonovershoot) models yield a burst age about 2 x 10 exp 9 yr younger than the value obtained. An initial mass function with slope of 2.35 (the Salpeter value) and a mean field star metallicity of Fe/H of about -0.7 are consistent with the photometric data and LFs. It is suggested that the star formation rate in the LMC was globally quite low during at least the first half of its lifetime, and that a major event triggered a substantial and relatively sudden increase in the star formation rate throughout the entire LMC which persisted for several 10 exp 9 yr and even up to the present epoch in some parts of that galaxy.

Bertelli, Gianpaolo↗

Discrete fracture network model benchmarks developed and applied in a DECOVALEX-2023 repository performance assessment study

This study presents newly developed benchmarks for modeling flow and transport within discrete fracture networks (DFNs) and useful methods for analyzing the results. The new benchmarks are designed to test modeling approaches for use in probabilistic performance assessment models of deep geologic repositories in fractured rock. The benchmarks simulate flow and transport through a 1 km 3 block of fractured rock. The first simulates migration of a short pulse of tracer through a simple network of four intersecting fractures. The second adds 1089 stochastically generated fractures. The third changes the pulse to a continuous point source. Evaluation of model performance relies on moment analysis and comparison of the results of different models. The expected nondimensional first moment of the conservative tracer for each benchmark is 1. The benchmarks were simulated by teams from Canada, Czechia, Germany, Korea, Sweden, Taiwan, and the United States as part of a DECOVALEX-2023 study (decovalex.org). The teams used various approaches, including explicit DFN modeling, DFN upscaling to an equivalent continuous porous medium (ECPM), and a combination of both methods. Transport mechanisms are modeled using either the advection-dispersion equation or particle tracking. Results demonstrate strong agreement among the models in breakthrough behavior up to the 75th percentile. Significant deviations in first moments and well-clustered outputs led to the identification of inaccuracies in several models. Such findings exemplify the benefit of exercising these benchmarks and using the presented methods to test DFN flow and transport models.

Benchmark↗

Probing ExoMiner for Effectiveness against False Alarms in Kepler Data

We present a study on the effectiveness of ExoMiner against False Alarms in Kepler data. ExoMiner is a deep learning model that was used to validate around 370 Kepler Objects of Interest. We follow the analysis conducted in Coughlin et al (2017) “DR25 Robovetter Completeness and Effectiveness” for Robovetter, a rule-based model used to vet TCEs for this data release and automatically generate the Q1-Q17 DR 25 KOI Table. The ExoMiner model is trained on observed transit data from Kepler Q1-Q17 DR25 and evaluated on Kepler inverted and scrambled data. The results provide a more comprehensive insight into the capacities and limitations of ExoMiner, especially the vetting of not-transit-like signals and, more generally, the use of deep learning models to model transit photometry data for vetting and validation purposes.

exoplanet↗

Irreducible Tests for Space Mission Sequencing Software

As missions extend further into space, the modeling and simulation of their every action and instruction becomes critical. The greater the distance between Earth and the spacecraft, the smaller the window for communication becomes. Therefore, through modeling and simulating the planned operations, the most efficient sequence of commands can be sent to the spacecraft. The Space Mission Sequencing Software is being developed as the next generation of sequencing software to ensure the most efficient communication to interplanetary and deep space mission spacecraft. Aside from efficiency, the software also checks to make sure that communication during a specified time is even possible, meaning that there is not a planet or moon preventing reception of a signal from Earth or that two opposing commands are being given simultaneously. In this way, the software not only models the proposed instructions to the spacecraft, but also validates the commands as well.To ensure that all spacecraft communications are sequenced properly, a timeline is used to structure the data. The created timelines are immutable and once data is as-signed to a timeline, it shall never be deleted nor renamed. This is to prevent the need for storing and filing the timelines for use by other programs. Several types of timelines can be created to accommodate different types of communications (activities, measurements, commands, states, events). Each of these timeline types requires specific parameters and all have options for additional parameters if needed. With so many combinations of parameters available, the robustness and stability of the software is a necessity. Therefore a baseline must be established to ensure the full functionality of the software and it is here where the irreducible tests come into use.

sequencing test↗

The glass-ceiling convective regime and the origin and diversity of coronae on Venus

Venus and Earth are rocky planets of roughly the same size and bulk density, yet their surface volcanic and tectonic features appear substantially different. On Venus, the coexistence of large volcanic highlands—interpreted as the surface expression of long-lived mantle plumes—alongside coronae, smaller features thought to be caused by transient thermal diapirs, remains enigmatic. Using two-dimensional numerical models of mantle convection with sharp and broad mineral phase transitions for pyrolite, we show that both scales of upwellings can be generated in a stagnant lid planet with an interior temperature 250 to 400 K warmer than Earth’s. The smaller plumes originate from a ~600 km deep internal layer that exists as a consequence of the different sequence of mineral phase transitions that occur in warmer mantles less processed and differentiated by partial melting and volcanism. Future models that include melting will provide further tests of our hypothesis.

Science & Technology - Other Topics↗

Severity factor kinetic model as a strategic parameter of hydrothermal processing (steam explosion and liquid hot water) for biomass fractionation under biorefinery concept

Hydrothermal processes are an attractive clean technology and cost-effective engineering platform for bio-refineries based in the conversion of biomass to biofuels and high-value bioproducts under the basis of sustainability and circular bioeconomy. The deep and detailed knowledge of the structural changes by the severity of biomasses hydrothermal fractionation is scientifically and technological needed in order to improve processes effectiveness, reactors designs, and industrial application of the multi-scale target compounds obtained by stream explosion and liquid hot water systems. The concept of the severity factor [log 10 (R 0 )] established > 30 years ago, continue to be a useful index that can provide a simple description of the relationship between the operational conditions for biomass fractionation in second generation of biorefineries. Furthermore, this review develops a deep explanation of the hydrothermal severity factor based in lignocellulosic biomass fractionation with emphasis in research advances, pretreatment operations and the applications of severity factor kinetic model.

09 BIOMASS FUELS↗