Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT↗

A novel closed-form inversion of the convection–diffusion equation for rapid convection, diffusion, and source profile estimation

To simplify and routinize particle transport analysis in fusion devices, a novel closed form linear inversion of the 1-D convection diffusion equation to estimate diffusion and convection profiles D(r ⃗ ), v(r ⃗ ) and source distribution s(r ⃗ ), of a single species from measured data is derived and demonstrated on synthetic data. Profile estimates of D(r ⃗ ), v(r ⃗ ), s(r ⃗ ) and their uncertainties are given as a matrix expression constructed directly from the incoming density data of the transported species in space and time, as well as physics assumptions such as particle conservation and experimental geometry. The derived matrix expression can be applied to a pumped or non-pumped recycling species, or a non-recycling species that is effectively “pumped” by plasma-facing surfaces.

Hinson, Edward [ORNL] (ORCID:000000019713140X)↗

Anomaly Detection for Online Monitoring of Thermocouple Sensors in the Advanced Test Reactor

This study explores data-driven anomaly detection methods to analyze sensor fail- ures in the Advanced Gas Reactor (AGR) nuclear fuel irradiation experiments. Specifically, we examine failures of thermocouples (TCs), which are critical for mon- itoring and controlling in-reactor temperatures during operation. Failures were pri- marily observed during abrupt power transitions and manifested as sensor drop-outs, drifts, or unexplained behavior. We applied three time-series analysis techniques— rolling mean smoothing, matrix profile, and vector auto-regression (VAR)—to de- tect anomalies in TC data prior to failure events. The rolling mean method effec- tively highlighted deviations aligned with reported failures, while the matrix profile provided partial early warning but sometimes flagged normal fluctuations during power-down periods. VAR shows potential in capturing multivariate dependencies but requires further calibration. A rare case of TC drift was also documented, which did not result in failure, underscoring the challenge of building predictive models with sparse positive examples. Our findings demonstrate that traditional statistical tools can aid anomaly detection but have limited predictive power without richer training data. We propose future directions including synthetic data generation, real- time surrogate modeling, and multi-modal feature integration. This work provides a foundation for applying robust anomaly detection frameworks to mission-critical sensor systems in experimental settings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

An efficient method to propagate model uncertainty when inverting seismic data for time domain seismic moment tensors

SUMMARY We present a computationally efficient method to approximately propagate uncertainty when linearly inverting seismic data for point source, time variable moment tensor components. The method is based on the assumption that the data residual, given by the difference between the observed seismic data and the data predicated by a linear inversion, contains the effects of both data and model uncertainty. Our method uses a distribution of data residuals, added directly to the data, in a pseudo-Monte Carlo scheme. Using the assumption that the data residual is a stochastic process, we use the well-known Karhunen–Loève (KL) theorem to construct a distribution of data residuals, where the required basis functions are constructed using Fourier series. The Fourier series are scaled by a product of a random variable and the real-valued spectral amplitudes of the original data residual’s spectrum. Thus, the Fourier series and spectral amplitudes are eigenfunction-eigenvalue pairs used in the KL-based construction of data residual distribution. Using tests with synthetic data, we show that our method compares closely with a Finite Difference Monte Carlo (FDMC) method that we presented previously. More importantly, the method presented here is computationally several orders of magnitude faster than our previous FDMC method, and requires no a priori assumptions of model and/or data uncertainty.

Poppeliers, Christian (ORCID:0000000159526849)↗

Deep learning inversion of gravity data for detection of CO 2 plumes in overlying aquifers

In this work, we developed an effective U-Net based deep learning (DL) model for inversion of surface gravity data on a rectangular grid to predict 2-D high-resolution subsurface CO 2 distribution along a vertical cross-section due to CO 2 leakage through a wellbore within a deep CO 2 storage reservoir. We used synthetic data to model two types of CO 2 leakage scenarios: one CO 2 plume in a shallow aquifer (single plume case), and two plumes present at different depths (double plume case). The 3-D synthetic plume samples were created by sampling among predetermined CO 2 plume depths, saturations, and volumes. The corresponding surface gravity data on a rectangular grid were generated by a 3-D forward model. The U-Net model detected 72% of single-plume samples, and one or both plumes in 75% of double-plume samples. Most of the undetected single plumes have small gravity field strengths below the typical noise level of 5 μGal. This model generated reproducible, reliable predictions with acceptable errors and demonstrated improved spatial resolution over the conventional least-squares inversion. In contrast to the conventional least-squares inversion, which often overestimates the size of its target and underestimates its density, this U-Net model accurately delineated the boundary of a target. Furthermore, this DL inversion detected deep, small, or low saturation CO 2 plumes that are often more difficult to resolve with conventional gravity inversion methods. We note the limitations of this feasibility study, including the use of synthetic data with regular CO 2 plume shapes, and the prediction of a 2-D plume cross-section rather than the full 3-D plume, as well, we recognize the lower detection fraction for double-plume scenarios. Nevertheless, this study demonstrates that DL gravity inversion is a promising and potentially superior method to conventional least-squares inversion. Our U-Net based deep learning inversion approach may be adapted for inversion of other types of geophysical data. DL inversion can facilitate near real-time monitoring of geologic carbon sequestration to provide site operators with prompt information about subsurface CO 2 distribution for risk management and mitigation.

58 GEOSCIENCES↗

A system identification approach for non-intrusive reduced order modeling of radiation-induced photocurrents

In this study, development of compact photocurrent models is currently dominated by analytical techniques that rely on physical assumptions to render the governing equations solvable in a closed form. Violation of these assumptions can reduce the accuracy of the models and/or limit their scope. In this paper we show that system identification of nonlinear state-space systems can serve as an alternative numerical basis for non-intrusive reduced order modeling of photocurrent effects. To that end we develop a compact gray box photocurrent model (GBPM) by using a state-space representation with a low-dimensional latent state equation that mimics a mathematical model for the response of an idealized class of devices to ionizing radiation. In so doing we obtain a model that learns the dynamics of a quantity of interest directly from its measurements without requiring snapshots of the internal device state or its discretized model, and can be inferred from very small data sets. To demonstrate the approach we train the GBPM using a small experimental data set for a Z5236 Zener diode and a small synthetic data set obtained by simulating a synthetic pn-junction device. We then compare the GBPMs with black box models trained on the same data and show that performance of the latter is limited by the size of the data set, while the former are able to achieve excellent performance in both the reproductive and the predictive regimes.

97 MATHEMATICS AND COMPUTING↗

Evaluation of a preliminary regional Earth model through comparison of synthetic and observed waveform data

In this report, we document the process related to developing a regional geologic model of a 605 x 1334 km area centered around Utah and encompassing surrounding states. This model is developed to test the effect that composition of a model has on the generation of synthetic data with the intent of using this information to improve upon full waveform moment tensor inversions. We compare observed data from three seismic events and five stations to the synthetic data generated by a preliminary model derived from a geologic framework model (GFM) developed by the USGS. The synthetic data and observed data comparisons indicate that our preliminary model performs well at smaller offset distances in the northern and central sections of the model. However, the southern stations consistently display synthetic data P- and S-wave arrival times that do not match the observed data arrival times, indicating that the velocity structure of the southern part of the model especially is inaccurate.

58 GEOSCIENCES↗

Gaussian processes for inferring parton distributions

The extraction of parton distribution functions (PDFs) from experimental or lattice QCD data is an ill-posed inverse problem, where regularization strongly impacts both systematic uncertainties and the reliability of the results. We study a framework based on Gaussian Process Regression (GPR) to reconstruct PDFs from lattice QCD matrix elements. Within a Bayesian framework, Gaussian processes serve as flexible priors that encode uncertainties, correlations, and constraints without imposing rigid functional forms. We investigate a wide range of kernel choices, mean functions, and hyperparameter treatments. We quantify information gained from the data using the Kullback-Leibler divergence. Synthetic data tests demonstrate the consistency and robustness of the method. Our study establishes GPR as a systematic and non-parametric approach to PDF reconstruction, offering controlled uncertainty estimates and reduced model bias in lattice QCD analyses.

hadronic spectroscopy↗

Monitoring the Morphology of M87* in 2009-2017 with the Event Horizon Telescope

The Event Horizon Telescope (EHT) has recently delivered the first resolved images of M87*, the supermassive black hole in the center of the M87 galaxy. These images were produced using 230 GHz observations performed in 2017 April. Additional observations are required to investigate the persistence of the primary image feature—a ring with azimuthal brightness asymmetry—and to quantify the image variability on event horizon scales. To address this need, we analyze M87* data collected with prototype EHT arrays in 2009, 2011, 2012, and 2013. While these observations do not contain enough information to produce images, they are sufficient to constrain simple geometric models. We develop a modeling approach based on the framework utilized for the 2017 EHT data analysis and validate our procedures using synthetic data. Applying the same approach to the observational data sets, we find the M87* morphology in 2009–2017 to be consistent with a persistent asymmetric ring of ∼40 μas diameter. The position angle of the peak intensity varies in time. In particular, we find a significant difference between the position angle measured in 2013 and 2017. These variations are in broad agreement with predictions of a subset of general relativistic magnetohydrodynamic simulations. We show that quantifying the variability across multiple observational epochs has the potential to constrain the physical properties of the source, such as the accretion state or the black hole spin.

79 ASTRONOMY AND ASTROPHYSICS↗

Photovoltaic Analysis and Response Support (PARS) Platform for Solar Situational Awareness and Resiliency Services

The project's primary objective is to develop a digital-twin based Photovoltaic (PV) Analysis and Response Support (PARS) platform, which aims to provide real-time situational awareness and optimal response plans. This platform is designed to enhance the performance of hybrid PV systems, making them competitive with or even superior to conventional generation resources. The PARS platform enabled the project team to develop and evaluate an extensive suite of grid support functionalities for the hybrid PV systems to enhance grid performance, across key areas including visibility, dispatchability, security, resilience, and reliability. Given the global push toward achieving 100% clean energy by 2035, there is a significant increase in the integration of inverter-based resources (IBRs) throughout the energy grid. Effectively managing the inherent variability and uncertainty associated with IBRs is crucial for ensuring cost-effectiveness, reliability, and security in both the main grid and islanded microgrids. Constrained to a limited array of IEEE test systems or standard feeder models, traditional IBR modeling struggles to assimilate new field data, accurately reflect system dynamics, and adapt to the evolving energy landscape. In our project, we embraced a Digital Twin (DT) strategy for crafting the PARS platform. A digital twin acts as a precise virtual counterpart of a physical system, built on historical data and continuously honed with real-time insights. This enables the high-fidelity DT to accurately mirror current system operations and forecast future scenarios. Consequently, the PARS platform becomes an ideal environment for testing and refining monitoring, control, power, and energy management algorithms designed to boost hybrid PV system performance. The defining feature of the PARS platform, distinguishing it from other advanced simulation tools, is its exceptional adaptability. This is achieved by employing actual network topologies and utilizing real-time field data for fine-tuning and calibration, ensuring a close emulation of real-world conditions. The project deliverables include: 1) High-fidelity IBR models and tools for real-time parameterization, utilizing real-time field measurements to refine IBR models for enhanced accuracy and performance; 2) Grid-forming and Grid-following capabilities to deliver resilience services, including blackstart, voltage and frequency support, cold-load pick-up, power reserves, and three-phase load balancing across grid-connected and microgrid settings; 3) Machine learning-based forecasting tools and methods for generating synthetic data and topologies, creating diverse and realistic simulation environments for evaluating varied operational scenarios; 4) Advanced microgrid power and energy management algorithms for optimizing the integration and operation of PV, storage, and demand response resources within both feeder and community scales. The power grid data sets are provided by four utility companies in North Carolina and the New York Power Administration. Acting as industry advisors, our industry partners communicated stakeholder needs and regulatory standards to the research teams, aiding technology transfer by incorporating the developed methodologies into their daily operations. This collaboration ensures that the PARS platform, functioning as a power system digital twin, enhances our understanding of IBR dynamic behaviors and enables the development and evaluation of IBR control functions that match or exceed the capabilities of conventional synchronous generators.

14 SOLAR ENERGY↗

Synthetic High Impedance Fault Data through Deep Convolutional Generated Adversarial Network

High impedance faults (HIFs) have always been significant challenge in the power grids. Researchers have developed some advanced protective methods to detect the HIFs. To test and validate these methods, large amounts of HIF data are required. This paper presents a synthetic HIF data generating method using the deep convolutional generated adversarial network (DCGAN). The DCGAN includes a generator module to create synthetic HIF waveform from random noises; and a discriminator module to identify the flaws of those synthetic data, which ultimately help improve the quality of the synthetic data created by the generator. To test the fidelity of the generated synthetic HIF data, two different HIF-detection methods have been applied. Extensive simulation results have validated the effectiveness of using the DCGAN to create synthetic HIF data.

Yang, Kun↗

SRBench++: Principled Benchmarking of Symbolic Regression With Domain-Expert Interpretation

Symbolic regression searches for analytic expressions that accurately describe studied phenomena. The main promise of this approach is that it may return an interpretable model that can be insightful to users, while maintaining high accuracy. The current standard for benchmarking these algorithms is SRBench, which evaluates methods on hundreds of datasets that are a mix of real-world and simulated processes spanning multiple domains. At present, the ability of SRBench to evaluate interpretability is limited to measuring the size of expressions on real-world data, and the exactness of model forms on synthetic data. In practice, model size is only one of many factors used by subject experts to determine how interpretable a model truly is. Furthermore, SRBench does not characterize algorithm performance on specific, challenging sub-tasks of regression such as feature selection and evasion of local minima. In this work, we propose and evaluate an approach to benchmarking SR algorithms that addresses these limitations of SRBench by 1) incorporating expert evaluations of interpretability on a domain-specific task, and 2) evaluating algorithms over distinct properties of data science tasks. We evaluate 12 modern symbolic regression algorithms on these benchmarks and present an in-depth analysis of the results, discuss current challenges of symbolic regression algorithms and highlight possible improvements for the benchmark itself.

97 MATHEMATICS AND COMPUTING↗

Pooling Data Improves Multimodel IDF Estimates over Median-Based IDF Estimates: Analysis over the Susquehanna and Florida

Traditional multimodel methods for estimating future changes in precipitation intensity, duration, and frequency (IDF) curves rely on mean or median of models’ IDF estimates. Such multimodel estimates are impaired by large estimation uncertainty, shadowing their efficacy in planning efforts. Here, assuming that each climate model is one representation of the underlying data generating process, i.e., the Earth system, we propose a novel extension of current methods through pooling model data: (i) evaluate performance of climate models in simulating the spatial and temporal variability of the observed annual maximum precipitation (AMP), (ii) bias-correct and pool historical and future AMP data of reasonably performing models, and (iii) compute IDF estimates in a nonstationary framework from pooled historical and future model data. Pooling enhances fitting of the extreme value distribution to the data and assumes that data from reasonably performing models represent samples from the “true” underlying data generating distribution. Through Monte Carlo simulations with synthetic data, we show that return periods derived from pooled data have smaller biases and lesser uncertainty than those derived from ensembles of individual model data. We apply this method to NA-CORDEX models to estimate changes in 24-h precipitation intensity–frequency (PIF) estimates over the Susquehanna watershed and Florida peninsula. Our approach identifies significant future changes at more stations compared to median-based PIF estimates. The analysis suggests that almost all stations over the Susquehanna and at least two-thirds of the stations over the Florida peninsula will observe significant increases in 24-h precipitation for 2–100-yr return periods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

NREL Stratus - Enabling Workflows to Fuse Data Streams, Modeling, Simulation, and Machine Learning

Integrating cloud services into advanced computing facilities provides significant new capabilities over focusing solely on traditional high performance computing (HPC) workloads. This brings complementary capabilities as well as enabling new focused roles for HPC. They are especially potent for workflows that fuse data streams, modeling and simulation ('modsim') and machine learning. A key challenge to adopting a hybrid edge-cloud-HPC model is to align optimal capability, data, and user intent on the right resources for each step in a workflow.?The NREL Stratus service provides a basis for this: Stratus layers capabilities needed to make?cloud services accessible to a lab-based scientific community on commercial offerings, and; currently supports upwards of 200 projects ranging from IOT integration to traditional modeling and simulation. This provides a real-world inventory of scientific workflow elements. A growing knowledge base enables placing these elements appropriately between the edge, cloud, and traditional HPC. This paper outlines a vision via reference architecture and the application of that architecture in a typical workflow highlighting multiple components: sensor data intake, cleaning and transforming (edge/cloud suitable); generation of synthetic data through modsim, computationally heavy ML training and hyperparameter optimization (HPC suitable), and; inference and deployment (cloud ideal). Every step in such a workflow involves a cost-benefit analysis regarding the data movement, computational efficiency, availability, latency, and resource capabilities. The reference architecture and examples outlined allow for understanding new opportunities in the context of emerging workflows that combine IOT, cloud, and HPC to bolster scientific productivity.

AI↗

28 NREL Stratus - Enabling Workflows to Fuse Data Streams, Modeling, Simulation, and Machine Learning: Preprint

Integrating cloud services into advanced computing facilities provides significant new capabilities over focusing solely on traditional high performance computing (HPC) workloads. This brings complementary capabilities as well as enabling new focused roles for HPC. They are especially potent for workflows that fuse data streams, modeling and simulation ('modsim') and machine learning. A key challenge to adopting a hybrid edge-cloud-HPC model is to align optimal capability, data, and user intent on the right resources for each step in a workflow.?The NREL Stratus service provides a basis for this: Stratus layers capabilities needed to make?cloud services accessible to a lab-based scientific community on commercial offerings, and; currently supports upwards of 200 projects ranging from IOT integration to traditional modeling and simulation. This provides a real-world inventory of scientific workflow elements. A growing knowledge base enables placing these elements appropriately between the edge, cloud, and traditional HPC. This paper outlines a vision via reference architecture and the application of that architecture in a typical workflow highlighting multiple components: sensor data intake, cleaning and transforming (edge/cloud suitable); generation of synthetic data through modsim, computationally heavy ML training and hyperparameter optimization (HPC suitable), and; inference and deployment (cloud ideal). Every step in such a workflow involves a cost-benefit analysis regarding the data movement, computational efficiency, availability, latency, and resource capabilities. The reference architecture and examples outlined allow for understanding new opportunities in the context of emerging workflows that combine IOT, cloud, and HPC to bolster scientific productivity.

AI↗

Generating Synthetic Time Series Photovoltaic Data with Real-World Physical Challenges and Noise for Use in Algorithm Test and Validation

The PV Fleet Data Initiative and other projects seek the develop algorithms for automated analysis of PV time series data for extraction of statistical information and other parameters of the data such as degradation rates, soiling loss information, tracker performance, clipping or curtailment, system availability and other valuable information. While there is a vast body of PV data available for application of said extraction algorithms it is difficult to validate these algorithms because the true parameters to be extracted are not known. There has been a wide use of synthetic data in the literature for algorithm validation but this synthetic data is typically very bounded by the problem or topic at hand. The PV Fleet Data Initiative project has demonstrated that real time series PV data almost always includes a host of data quality and physical problems that, in reality, any automated PV abstraction algorithm must handle appropriately. For this reason, this work describes the development of a complex synthetic PV times series data set that includes data quality and physical problems that have been experienced in real world PV data. The various quality and physical problems are documented in the synthetic data so that users can test the validity of various PV extraction algorithms as well as develop new algorithms to solve problems this data set can support.

14 SOLAR ENERGY↗

TwinMe4AD: WGAN-based Digital Twins for Anomaly Detection

SAND2024-08373O TwinMe4AD is a Python-based software tool designed for anomaly detection using digital twins that closely mimic real, wearable healthcare datasets. The tool is invaluable for scenarios where collecting data is either expensive or impractical, serving as a privacy-preserving solution. Sensitive information is protected by training deep learning models on synthetic data derived from real datasets. One of TwinMe4AD's key features is its anomaly detection capability, which is based on fourth-order moments of parameters. This versatile approach can be applied across a range of datasets, from univariate to multivariate, making it compatible with various types of data. It also generates synthetic twins using Wasserstein Generative Adversarial Networks (WGANs), allowing users to create a small cohort of a population similar to that of a village population. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Poorey, Kunal↗

On the use of NMR distance measurements for assessing surface site homogeneity

The past few decades have seen tremendous growth in the area of single-site heterogeneous catalysis, which aims to combine the best aspects of homogeneous and heterogeneous catalysis, namely molecular-level site control and ease of separation/recycling. Despite this, we still do not have a means of assessing site homogeneity and whether the produced catalyst is indeed a “single-site”. Recent developments have enabled the use of NMR-based distance measurements to determine the conformations and configurations of surface sites, leading to the question whether such measurements can be used to distinguish materials containing either single or multiple surface sites with otherwise indistinguishable NMR properties. Here, we describe a Monte Carlo-based multi-structure search algorithm and its application to the determination of multi-site structures from supported metal complexes. The sensitivity of REDOR data to the existence of multiple sites is assessed using synthetic data and prior literature examples are revisited to determine whether the single-site approximation was indeed appropriate. We lastly apply this new methodology to differentiate the configurations of zirconocene complexes grafted onto alumina supports that were thermally treated at different temperatures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗