Engineering PapersSearch

SEARCH · Engineering Papers

Results for “synthetic data generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A Convolution Neural Network for Voltage Event Classification at a Photovoltaic Inverter

This paper presents a convolutional neural network (CNN) developed to identify voltage events in photovoltaic (PV) inverters. The CNN is trained on synthetic data generated using the IEEE 13-bus distribution feeder model and evaluated on field measured data collected from Energy Northwest’s Horn Rapids Solar, Storage, and Training (HRSST) facility. The study focuses on two common voltage events: faults and voltage sags. The CNN is configured to analyze voltage and current waveforms from three-phase PV systems, demonstrating excellent accuracy during training. Field data from the HRSST facility is employed to assess its real-world performance, where the CNN achieves perfect identification of faults and voltage sags in a sample of nine events. This work highlights the potential of the proposed method to enhance PV protection schemes, providing a robust foundation for improved voltage event detection and grid reliability.

Cornachione, Matthew A.

Predictive Modeling and Uncertainty Quantification in Condition Monitoring of Active Components: A Reactor Coolant Pump Use Case

This work develops data-driven models for onset of thermal barrier leakage in reactor coolant pumps. It incorporates uncertainty quantification to enhance the reliability and robustness of pre- dictions. Using synthetic data generated by the Generic Pressurized Water Reactor simulator, realistic degradation scenarios were simulated across lifecycle stages—beginning, middle, and end of life. Key variables, including differential pressure, flow rate, vibration, and temperatures, were analyzed using machine learning framework. The fully connected neural network models demonstrated exceptional performance, achieving R2 scores exceeding 0.99 and root mean square errors as low as around 8.23 × 10-2 gallon per minute (gpm) for the three stages of the lifecy- cle. UQ analysis further validated the model’s robustness, with narrow uncertainty bounds during steady-state operations and appropriately wider bounds during transitional phases, reflecting the physical behavior of the system. This work addresses important gaps in real-time condition moni- toring and regulatory compliance by integrating advanced condition monitoring technologies with UQ into IST programs. The ability to detect thermal barrier leakage early and quantify prediction reliability supports optimizing maintenance strategies while ensuring nuclear power plants’ safe and reliable operation.

99 - GENERAL AND MISCELLANEOUS

A Data-Driven Method for Synthetic Extreme Weather Generation and Solar Impact Assessment: Preprint

High-resolution, high-fidelity weather datasets are essential for testing and evaluating the resilience of power systems, particularly under extreme weather conditions. However, existing extreme weather datasets are typically derived from historical events that are localized and may lack the spatial and temporal resolution or scenario diversity needed to test largescale power systems. In this work, we propose a synthetic extreme weather simulation approach capable of generating targeted extreme events, such as hurricanes, using publicly available data sources. Preliminary results demonstrate the impact of a simulated Category 1 hurricane on renewable generation and critical infrastructure in California. The work aims to provide a flexible approach for creating multiple types of extreme weather scenarios across different regions, enabling comprehensive system stress testing, training, and resilience assessment.

24 POWER TRANSMISSION AND DISTRIBUTION

An ultra-fast method for generating synthetic down-scattered neutron data for inertial confinement fusion implosions

In inertial confinement fusion experiments at the National Ignition Facility, asymmetries are probed by a variety of neutron diagnostics, including neutron imaging systems, real-time neutron activation diagnostics (RTNADs), and neutron spectrometers. It is often useful to generate synthetic data based on these diagnostics to validate and tune models. However, current methods of doing so using Monte Carlo particle tracing are time-consuming. In this paper, an ultra-fast method is presented for generating synthetic neutron images, RTNAD data, and spectrometry data using line integrals and 3D convolutions. While it does not contain as much physics as particle tracing codes, it is thousands of times faster and produces nearly identical data. This enables analysis techniques that depend on generating large amounts of synthetic data, which will prove very useful for the study of asymmetries going forward.

Deuterium

An investigation on machine learning predictive accuracy improvement and uncertainty reduction using VAE-based data augmentation

The confluence of ultrafast computers with large memory, rapid progress in Machine Learning (ML) algorithms, and the availability of large datasets place multiple engineering fields at the threshold of dramatic progress. However, a unique challenge in nuclear engineering is data scarcity because experimentation on nuclear systems is usually more expensive and time-consuming than most other disciplines. One potential way to resolve the data scarcity issue is deep generative learning, which uses certain ML models to learn the underlying distribution of existing data and generate synthetic samples that resemble the real data. In this way, one can significantly expand the dataset to train more accurate predictive ML models. In this study, our objective is to evaluate the effectiveness of data augmentation using variational autoencoder (VAE)-based deep generative models. We investigated whether the data augmentation leads to improved accuracy in the predictions of a deep neural network (DNN) model trained using the augmented data. Additionally, the DNN prediction uncertainties are quantified using Bayesian Neural Networks (BNN) and conformal prediction (CP) to assess the impact on predictive uncertainty reduction. To test the proposed methodology, we used TRACE simulations of steady-state void fraction data based on the NUPEC Boiling Water Reactor Full-size Fine-mesh Bundle Test (BFBT) benchmark. Here, we found that augmenting the training dataset using VAEs has improved the DNN model’s predictive accuracy, improved the prediction confidence intervals, and reduced the prediction uncertainties.

Bayesian neural network

Monte Carlo toolkit for designing and validating step-range-filter spectrometer designs

Here, we present a Monte Carlo toolkit for validating step range filter (SRF) spectrometer designs. Geant4 is used to transport charged particles through the SRF filters to generate synthetic SRF data that include realistic CR-39 effects. Synthetic SRF spectra generated by this method inherently account for instrument response and allow for the quantification of SRF performance before shots. The usefulness of this toolkit is demonstrated through its application to a number of problems. A new broadband SRF for the ∼10 MeV wide 3He3He proton spectrum is validated, and an analysis method for analyzing 3He3He-p SRF data that accounts for instrument response is put forth. In addition, an SRF design for the compact recoil-proton spectrometer (CRS) on the Z-machine is validated. Finally, a new calibration technique for the DD-p SRF is proposed and validated.

Johnson, T. M. (ORCID:0000000193032949)

Simulation Aspects in the Study of Rectification of Satellite Scanner Data

Complete sensor/platform modelling is derived and used for the generation of synthetic data and for rectification studies of satellite scanner data. All satellite position and sensor attitude parameters are recovered. Rectification accuracy improves marginally when using more than 25 control points, and is highly sensitive to errors in image point identification.

Mikhail, E. M.

A Visual Analytic Platform for Interactive Validation of Human Mobility Simulations

Human mobility insights guide domain experts in an array of decisions, including critical infrastructure design, disaster response, epidemic modeling, national security, and policy making. Due to the inherent noise and privacy concerns in real-world individual-level mobility data, it is often preferred to leverage simulators that generate synthetic mobility data instead. However, it is critical to inspect and validate the output of such simulators to ensure the synthetic data is aligned with the characteristics of the population and the area of interest known to domain experts. While there exist many quantitative approaches for validating synthetic data, we argue it is also important to also validate such data qualitatively to capture aspects that are known to domain experts but difficult to quantify. In this work, we demonstrate a visual analytic platform that empowers domain experts to interact with their simulation outputs along spatial and temporal dimensions. By augmenting automated techniques and human skills, our visual analytic platform is a step towards interactive capabilities for model steering and quality control of mobility simulators.

Monadjemi, Shayan

Evaluating the Efficacy of Conditional Variational Autoencoders in Generating Synthetic Single Nuclei RNA-Seq Data for Space Biology Research

Astronauts are subject to unique stressors during spaceflight, leading to changes in their cellular function. However, neither astronauts nor model organisms respond the same to spaceflight, and research implicates a contribution of omics components in differential responses. Understanding how gene expression affects astronaut health is critical for the success of long-term space missions, prompting interest in developing personalized predictive models leveraging artificial intelligence (AI) and machine learning (ML) techniques. Developing such models requires extensive data, which is challenging to obtain and share. This study explores the use of conditional variational autoencoders (CVAEs) to synthetically generate single-nuclei RNA-seq (snRNA-seq) data. CVAEs build on standard variational autoencoders (VAEs) by conditioning data generation on covariates like sample identity and mission parameters, enhancing the relevance of generated data for specific contexts. For our work, we built two CVAEs with varying degrees of sparsity to optimize both interpretability and generative power. We train and validate models on existing snRNA-seq data collected from the brain tissue of mice subjected to spaceflight conditions and their ground control counterparts. We evaluate model performance using statistical tests and visualizations to compare synthetic data to real data. We aim to demonstrate that these prototype CVAE architectures could be used in future space biology work and that this is a method worth further exploring.

Sarah Golts

Synthetic Scientific Image Generation with VAE, GAN, and Diffusion Model Architectures

Generative AI (genAI) has emerged as a powerful tool for synthesizing diverse and complex image data, offering new possibilities for scientific imaging applications. This review presents a comprehensive comparative analysis of leading generative architectures, ranging from Variational Autoencoders (VAEs) to Generative Adversarial Networks (GANs) on through to Diffusion Models, in the context of scientific image synthesis. We examine each model's foundational principles, recent architectural advancements, and practical trade-offs. Our evaluation, conducted on domain-specific datasets including microCT scans of rocks and composite fibers, as well as high-resolution images of plant roots, integrates both quantitative metrics (SSIM, LPIPS, FID, CLIPScore) and expert-driven qualitative assessments. Results show that GANs, particularly StyleGAN, produce images with high perceptual quality and structural coherence. Diffusion-based models for inpainting and image variation, such as DALL-E 2, delivered high realism and semantic alignment but generally struggled in balancing visual fidelity with scientific accuracy. Importantly, our findings reveal limitations of standard quantitative metrics in capturing scientific relevance, underscoring the need for domain-expert validation. We conclude by discussing key challenges such as model interpretability, computational cost, and verification protocols, and discuss future directions where generative AI can drive innovation in data augmentation, simulation, and hypothesis generation in scientific research.

Generative Adversarial Networks

Estimates of Ground Temperature and Atmospheric Moisture from CERES Observations

A method is developed to retrieve surface ground temperature (T(sub g)) and atmospheric moisture using clear sky fluxes (CSF) from CERES-TRMM observations. In general, the clear sky outgoing longwave radiation (CLR) is sensitive to upper level moisture (q(sub l)) over wet regions and (T(sub g)) over dry regions The clear sky window flux from 800 to 1200/cm (RadWn) is sensitive to low level moisture (q(sub t)) and T(sub g). Combining these two measurements (CLR and RadWn), Tg and q(sub h) can be estimated over land, while q(sub h) and q(sub l) can be estimated over the oceans. The approach capitalizes on the availability of satellite estimates of CLR and RadWn and other auxiliary satellite data. The basic methodology employs off-line forward radiative transfer calculations to generate synthetic CSF data from two different global 4-dimensional data assimilation products. Simple linear regression is used to relate discrepancies in CSF to discrepancies in T(sub g), q(sub h) and q(sub l). The slopes of the regression lines define sensitivity parameters that can be exploited to help interpret mismatches between satellite observations and model-based estimates of CSF. For illustration, we analyze the discrepancies in the CSF between an early implementation of the Goddard Earth Observing System Data Assimilation System (GEOS-DAS) and a recent operational version of the European Center for Medium-Range Weather Prediction data assimilation system. In particular, our analysis of synthetic total and window region SCF differences (computed from two different assimilated data sets) shows that simple linear regression employing Delta(T(sub g)) and broad layer Delta(q(sub l) from .500 hPa to surface and Delta(q(sub h)) from 200 to .300 hPa provides a good approximation to the full radiative transfer calculations. typically explaining more than 90% of the 6-hourly variance in the flux differences. These simple regression relations can be inverted to "retrieve" the errors in the geophysical parameters. Uncertainties (normalized by standard deviation) in the monthly mean retrieved parameters range from 7% for Delta(T(sub g)) to about 20% for Delta(q(sub l)). Our initial application of the methodology employed an early CERES-TRMM data set (CLR and Radwn) to assess the quality of the GEOS2 data. The results showed that over the tropical and subtropical oceans GEOS2 is, in general, too wet in the upper troposphere (mean bias of 0.99 mm) and too dry in the lower troposphere (mean bias of -4.7 min). We note that these errors, as well as a cold bias in the T(sub g). have largely been corrected in the current version of GEOS-2 with the introduction of a land surface model, a moist turbulence scheme and the assimilation of SSM/I total precipitable water.

Wu, Man Li C.

Estimates of Ground Temperature and Atmospheric Moisture from CERES Observations

A method is developed to retrieve surface ground temperature (Tg) and atmospheric moisture using clear sky fluxes (CSF) from CERES-TRMM observations. In general, the clear sky outgoing long-wave radiation (CLR) is sensitive to upper level moisture (q(sub h)) over wet regions and Tg over dry regions The clear sky window flux from 800 to 1200 /cm (RadWn) is sensitive to low level moisture (q(sub j)) and Tg. Combining these two measurements (CLR and RadWn), Tg and q(sub h) can be estimated over land, while q(sub h) and q(sub t) can be estimated over the oceans. The approach capitalizes on the availability of satellite estimates of CLR and RadWn and other auxiliary satellite data. The basic methodology employs off-line forward radiative transfer calculations to generate synthetic CSF data from two different global 4-dimensional data assimilation products. Simple linear regression is used to relate discrepancies in CSF to discrepancies in Tg, q(sub h) and q(sub t). The slopes of the regression lines define sensitivity parameters that can be exploited to help interpret mismatches between satellite observations and model-based estimates of CSF. For illustration, we analyze the discrepancies in the CSF between an early implementation of the Goddard Earth Observing System Data Assimilation System (GEOS-DAS) and a recent operational version of the European Center for Medium-Range Weather Prediction data assimilation system. In particular, our analysis of synthetic total and window region SCF differences (computed from two different assimilated data sets) shows that simple linear regression employing (Delta)Tg and broad layer (Delta)q(sub l) from 500 hPa to surface and (Delta)q(sub h) from 200 to 500 hPa provides a good approximation to the full radiative transfer calculations, typically explaining more than 90% of the 6-hourly variance in the flux differences. These simple regression relations can be inverted to "retrieve" the errors in the geophysical parameters. Uncertainties (normalized by standard deviation) in the monthly mean retrieved parameters range from 7% for (Delta)T to about 20% for (Delta)q(sub t). Our initial application of the methodology employed an early CERES-TRMM data set (CLR and Radwn) to assess the quality of the GEOS2 data. The results showed that over the tropical and subtropical oceans GEOS2 is, in general, too wet in the upper troposphere (mean bias of 0.99 mm) and too dry in the lower troposphere (mean bias of -4.7 mm). We note that these errors, as well as a cold bias in the Tg, have largely been corrected in the current version of GEOS-2 with the introduction of a land surface model, a moist turbulence scheme and the assimilation of SSTM/I total precipitable water.

Wu, Man Li C.

Denoising diffusion probabilistic models for generative alloy design

Inverse material design is an extremely challenging optimization task made difficult by, in part, the highly nonlinear relationship linking performance with composition. Quantitative approaches have improved significantly owing to advances in high throughput experimentation and computational thermodynamics. However, existing physics-based tools are mostly forward models; input a chemistry and obtain a prediction. More recently the materials community has leveraged advances in the machine learning community to establish novel inverse design frameworks. Very recently denoising diffusion probabilistic models have been shown to be extremely powerful generators producing synthetic data of various modalities e.g. images, text, audio, tables, etc.. In this work a novel framework for alloy design and optimization is proposed leveraging these class of models. Five key generative tasks are demonstrated (1) unconditional generation (2) composition conditioned generation (3) property conditioned generation (4) multi-feedstock conditioned generation and (5) generative optimization. These methods were tested on three case studies: high entropy alloy design, superalloy binder jet additive manufacturing, and in-situ dual-feedstock wire-arc additive manufacturing. Results indicate that the established models are extremely flexible, expressive, and robust. The architecture’s flexibility and training procedure empower the model to learn complex intra-compositional and composition-property relationships. Furthermore, the probabilistic nature of these models makes them well suited for addressing solution non-uniqueness and tackling uncertainty quantification tasks. While the fidelity and quantity of the underlying training data is paramount, we envision that future alloy design frameworks will make extensive use of these kinds of machine learning models as “search” tools bolstering the utility of experimental and computational approaches.

36 MATERIALS SCIENCE

Data Science and Urban Air Mobility: Challenges and Opportunities

Aviation is broadly a combination of aircraft, airspace and airports. The data science life cycle comprises of five steps - capture, maintain, process, analyze and communicate. The presentation introduces the legacy of conventional aviation research in the context of the data science life cycle to motivate the challenges with Urban Air Mobility, a field that is quite nascent. A summary of recent research will be presented to highlight the innovative ways to address the challenges. Examples provided will include the generation of synthetic data, encounter models from simulations, and leveraging novel and diverse data sets from traditional transportation and non-aviation sources, to analyze problems of operation in urban airspace. Finally, opportunities will be identified for further exploration, niche development and filling the gaps in the field of data science for UAM.

Urban Air Mobility

Data Science and Urban Air Mobility: Challenges and Opportunities

Aviation is broadly a combination of aircraft, airspace and airports. The data science life cycle comprises of five steps - capture, maintain, process, analyze and communicate. The presentation introduces the legacy of conventional aviation research in the context of the data science life cycle to motivate the challenges with Urban Air Mobility, a field that is quite nascent. A summary of recent research will be presented to highlight the innovative ways to address the challenges. Examples provided will include the generation of synthetic data, encounter models from simulations, and leveraging novel and diverse data sets from traditional transportation and non-aviation sources, to analyze problems of operation in urban airspace. Finally, opportunities will be identified for further exploration, niche development and filling the gaps in the field of data science for UAM.

Urban Air Mobility

Data Science Challenges for Urban Air Mobility

Aviation is a combination of aircraft, airspace and airports. The data science life cycle comprises of five steps - capture, maintain, process, analyze and communicate. The presentation introduces the legacy of conventional aviation research in the context of the data science life cycle to motivate the challenges with Urban Air Mobility, a field that is quite nascent. A summary of recent research will be presented to highlight the innovative ways to address the challenges. Examples provided will include the generation of synthetic data, encounter models from simulations, and leveraging novel and diverse data sets from traditional transportation and non-aviation sources, to analyze problems of operation in urban airspace. Finally, opportunities will be identified for further exploration, niche development and filling the gaps in the field of data science for UAM.

Data Science

Data Science Challenges for Urban Air Mobility

Aviation is a combination of aircraft, airspace and airports. The data science life cycle comprises of five steps - capture, maintain, process, analyze and communicate. The presentation introduces the legacy of conventional aviation research in the context of the data science life cycle to motivate the challenges with Urban Air Mobility, a field that is quite nascent. A summary of recent research will be presented to highlight the innovative ways to address the challenges. Examples provided will include the generation of synthetic data, encounter models from simulations, and leveraging novel and diverse data sets from traditional transportation and non-aviation sources, to analyze problems of operation in urban airspace. Finally, opportunities will be identified for further exploration, niche development and filling the gaps in the field of data science for UAM.

Data Science