Engineering PapersSearch

SEARCH · Engineering Papers

Results for “synthetic data generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT

Constrained GAN-Generated X-Ray CT Data For Self-Supervised And Foundation-Model Segmentation Of Concrete Microstructures

Three-dimensional characterization of materials using X-ray computed tomography (XCT) is challenging due to the complexity of internal structures, noise, and variations in resolution. Traditional computer vision models often struggle to accurately segment these images, particularly in domain-specific applications like materials science. While supervised deep learning approaches have been developed to address the limitations of conventional algorithms, they typically require large amounts of labeled training data and often fail to generalize across different datasets. Self-supervised, few-and zero-shot learning methods have gained prominence in natural image processing and segmentation tasks, but their application to scientific imaging remains limited due to the unique structural complexity, noise, and textural artifacts present in materials science data. In this work, we investigate how domain adaptation, leveraging physics-based and GAN-generated synthetic data, impacts segmentation performance. We introduce a modified Contrastive Unpaired Translation (CUT) model designed to generate realistic labeled data, which can be used for training, pre-training, and fine-tuning segmentation models for real XCT microstructure data. We evaluate the performance of two segmentation approaches: a self-supervised network (SSL-ALPNet) and a foundation model (Segment Anything Model), assessing their improvements when pre-trained and/or fine-tuned on the synthesized data. Our results demonstrate that leveraging synthetic data significantly enhances segmentation performance, particularly in challenging materials science applications.

Ziabari, Amir [ORNL] (ORCID:000000034776457X)

Cloverleaf Data Artifacts for ArtIMis LDRD

This report summarizes the use of the open-source CloverLeaf/CloverLeaf3D mini-apps to generate synthetic data sets to train foundation models for the ArtIMis LDRD DI. These data artifacts are intended to be used by LANL collaborators and shared externally with our university and institutional partners. Note that CloverLeaf/CloverLeaf3D is not a LANL simulation code.

97 MATHEMATICS AND COMPUTING

AI-Driven Crack Detection for Remanufacturing Cylinder Heads Using Deep Learning and Engineering-Informed Data Augmentation

Detecting cracks in cylinder heads traditionally relies on manual inspection, which is time-consuming and susceptible to human error. As an alternative, automated object detection utilizing computer vision and machine learning models has been explored. However, these methods often face challenges due to a lack of sufficiently annotated training data, limited image diversity, and the inherently small size of cracks. Addressing these constraints, this paper introduces a novel automated crack-detection method that enhances data availability through a synthetic data generation technique. Unlike general data augmentation practices, our method involves copying cracks from one location to another, guided by both random and informed engineering decisions about likely crack formations due to cyclic thermomechanical loads. The innovative aspect of our approach lies in the integration of domain-specific engineering knowledge into the synthetic generation process, which substantially improves detection accuracy. We evaluate our method’s effectiveness using two metrics: the F2 score, which emphasizes recall to prioritize detecting all potential cracks, and mean average precision (MAP), a standard measure in object detection. Experimental results demonstrate that, without engineering insights, our method increases the F2 score from 0.40 to 0.65, while maintaining a stable MAP. Incorporating detailed engineering knowledge further enhances the F2 score to 0.70 and improves MAP to 0.57, representing increases of 63% and 43%, respectively. These results confirm that our approach not only mitigates the limitations of traditional data augmentation but also significantly advances the reliability and precision of crack detection in industrial settings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

CAFE AU LAIT: Compute-Aware Federated Augmented Low-Rank AI Training

Federated finetuning is crucial for unlocking the knowledge embedded in pretrained Large Language Models (LLMs) when data are geographically distributed across clients. Unlike finetuning with data from a single institution, federated finetuning allows collaboration across multiple institutions, enabling the utilization of diverse and decentralized datasets while preserving data privacy. Given the high computing costs of LLM training and the emphasis on energy efficiency in Federated Learning (FL), Low-Rank Adaptation (LoRA) has emerged as a widely adopted algorithm due to its significantly reduced number of trainable parameters. However, this assumes that all data silos have the necessary computing resources to compute local updates of LLMs. Nevertheless, in practice, the computing resources across clients are highly heterogeneous: while some may have access to hundreds of GPUs, others might have limited or no GPU access. Recently, federated finetuning using synthetic data has been proposed, allowing clients to participate in a collaborative training run without training LLMs locally. However, our experimental results reveal a performance gap between models trained using synthetic data and those trained using local updates. Motivated by the observed heterogeneity in computing resources and the performance gap, we propose a novel two-stage algorithm that leverages the storage and computing capabilities of a strong server. In the first stage, under the coordination of the strong server, clients with limited computing resources collaborate to generate synthetic data, which is transferred to and stored on the strong server. In the second stage, the strong server uses this synthetic data on behalf of the resource-constrained clients to perform federated LoRA finetuning alongside clients with sufficient computing resources. This approach ensures that all clients can participate in the finetuning process. Experimental results demonstrate that incorporating local updates from even a small fraction of clients improves performance compared to using synthetic data for all clients. Furthermore, we incorporate the Gaussian mechanism in both stages to guarantee client-level differential privacy.

Wang, Jiayi [ORNL]

Discovery of Activities via Statistical Clustering of Fixation Patterns

Human behavior often consists of a series of distinct activities, each characterized by a unique pattern of interaction with the visual environment. This is true even in a restricted domain, such as a pilot flying an airplane; in this case, activities with distinct visual signatures might be things like communicating, navigating, monitoring, etc. We propose a novel analysis method for gaze-tracking data, to perform blind discovery of these hypothetical activities. We compare, not individual fixations, but groups of fixations aggregated over a fixed time interval (Tau). We assume that the environment has been divided into a finite set of discrete areas-of-interest (AOIs). For a given time interval, we compute the proportion of time spent fixating each AOI, resulting in an N-dimensional vector, where N is the number of AOIs. These proportions can be converted to integer counts by multiplying by Tau divided by the average fixation duration, a parameter that we fix at 283 milliseconds. We compare different intervals by computing the chi-squared statistic. The p-value associated with the statistic is the likelihood of observing the data under the hypothesis that the data in the two intervals were generated by a single process with a single set of probabilities governing the fixation of each AOI. We cluster the intervals, first by merging adjacent intervals that are sufficiently similar, optionally shifting the boundary between non-merged intervals to maximize the difference. Then we compare and cluster non-adjacent intervals. The method is evaluated using synthetic data generated by a hand-crafted set of activities. While the method generally finds more activities than put into the simulation, we have obtained agreement as high as 80 percent between the inferred activity labels and ground truth.

Eye Movements

Sample Analysis at Mars Instrument Simulator

The Sample Analysis at Mars Instrument Simulator (SAMSIM) is a numerical model dedicated to plan and validate operations of the Sample Analysis at Mars (SAM) instrument on the surface of Mars. The SAM instrument suite, currently operating on the Mars Science Laboratory (MSL), is an analytical laboratory designed to investigate the chemical and isotopic composition of the atmosphere and volatiles extracted from solid samples. SAMSIM was developed using Matlab and Simulink libraries of MathWorks Inc. to provide MSL mission planners with accurate predictions of the instrument electrical, thermal, mechanical, and fluid responses to scripted commands. This tool is a first example of a multi-purpose, full-scale numerical modeling of a flight instrument with the purpose of supplementing or even eliminating entirely the need for a hardware engineer model during instrument development and operation. SAMSIM simulates the complex interactions that occur between the instrument Command and Data Handling unit (C&DH) and all subsystems during the execution of experiment sequences. A typical SAM experiment takes many hours to complete and involves hundreds of components. During the simulation, the electrical, mechanical, thermal, and gas dynamics states of each hardware component are accurately modeled and propagated within the simulation environment at faster than real time. This allows the simulation, in just a few minutes, of experiment sequences that takes many hours to execute on the real instrument. The SAMSIM model is divided into five distinct but interacting modules: software, mechanical, thermal, gas flow, and electrical modules. The software module simulates the instrument C&DH by executing a customized version of the instrument flight software in a Matlab environment. The inputs and outputs to this synthetic C&DH are mapped to virtual sensors and command lines that mimic in their structure and connectivity the layout of the instrument harnesses. This module executes, and thus validates, complex command scripts prior to their up-linking to the SAM instrument. As an output, this module generates synthetic data and message logs at a rate that is similar to the actual instrument.

Benna, Mehdi

Predicting critical heat flux with uncertainty quantification and domain generalization using conditional variational autoencoders and deep neural networks

Deep generative models (DGMs) can generate synthetic data samples that closely resemble the original dataset, addressing data scarcity. In this work, we developed a conditional variational autoencoder (CVAE) to augment critical heat flux (CHF) data used for the 2006 Groeneveld lookup table. To compare with traditional methods, a fine-tuned deep neural network (DNN) regression model was evaluated on the same dataset. Both models achieved small mean absolute relative errors, with the CVAE showing more favorable results. Uncertainty quantification (UQ) was performed using repeated CVAE sampling and DNN ensembling. The DNN ensemble improved performance over the baseline, while the CVAE maintained consistent results with less variability and higher confidence. Both models achieved small errors inside and outside the training domain, with slightly larger errors outside. Altogether, the CVAE performed better than the DNN in predicting CHF and exhibited better uncertainty behavior.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Knowledge-guided graph machine learning for spatially distributed prediction of daily discharge and nitrogen export dynamics

Spatially distributed prediction of streamflow and nitrogen export dynamics is essential for precision management of agricultural watersheds. While temporal deep learning models such as Long Short-Term Memory (LSTM) have shown strong performance at basin scales, their ability to generalize spatially is limited by insufficient representation of spatial dependencies and flow paths, particularly under data-scarce conditions. To address this gap, we propose HydroGraphNet, a knowledge-guided graph machine learning framework that integrates process-based knowledge and explicit spatial learning into temporal modeling. This framework incorporates directed graph topology to encode watershed connectivity and upstream inflows, with mass balance constraints to improve physical consistency. To enhance generalization in sparsely monitored regions, HydroGraphNet is pretrained on synthetic data generated by the SWAT+ (Soil and Water Assessment Tool Plus) model. We evaluated HydroGraphNet in the Upper Sangamon River Basin (44 HUC-12 subwatersheds, 2001–2020) against two LSTM baselines: a lumped basin-level model and a distributed variant. When benchmarked on SWAT+ simulations in pretraining, HydroGraphNet improved test NSEs by 8.9% (discharge) and 13.7% (NO₃–N load) in temporal extrapolation, and by 27.1% and 34.7% in spatial extrapolation, relative to the Lumped LSTM baseline. After fine-tuning with USGS monitoring data, the model achieved mean test NSE (KGE) scores of 0.768 (0.861) for discharge and 0.626 (0.664) for NO₃–N load, substantially outperforming baselines. Attribution analysis further highlighted the importance of upstream inflow representation and graph-based spatial learning in capturing cross-subwatershed dependencies. The model also reproduced seasonal hydrological and biogeochemical patterns consistent with known processes, demonstrating its robustness and process fidelity for spatially distributed prediction. Altogether, HydroGraphNet advances the integration of physical knowledge and spatially explicit learning in hydrological modeling, offering a generalizable framework for distributed modeling to support spatially targeted water quality management in data-scarce watersheds.

54 ENVIRONMENTAL SCIENCES

A three-point velocity estimation method for two-dimensional coarse-grained imaging data

Time delay and velocity estimation methods have been widely studied subjects in the context of signal processing, with applications in many different fields of physics. The velocity of waves or coherent fluctuation structures is commonly estimated as the distance between two measurement points divided by the time lag that maximizes the cross correlation function between the measured signals, but this is demonstrated to result in erroneous estimates for two spatial dimensions. We present an improved method to accurately estimate both components of the velocity vector, relying on three non-aligned measurement points. We introduce a stochastic process describing the fluctuations as a superposition of uncorrelated pulses moving in two dimensions. Using this model, we show that the three-point velocity estimation method, using time delays calculated through cross correlations, yields the exact velocity components when all pulses have the same velocity. The two- and three-point methods are tested on synthetic data generated from realizations of such processes for which the underlying velocity components are known. The results reveal the superiority of the three-point technique. Finally, we demonstrate the applicability of the velocity estimation on gas puff imaging data of strongly intermittent plasma fluctuations due to the radial motion of coherent, blob-like structures at the boundary of the Alcator C-Mod tokamak.

Materials Science

Anomaly Detection for Online Monitoring of Thermocouple Sensors in the Advanced Test Reactor

This study explores data-driven anomaly detection methods to analyze sensor fail- ures in the Advanced Gas Reactor (AGR) nuclear fuel irradiation experiments. Specifically, we examine failures of thermocouples (TCs), which are critical for mon- itoring and controlling in-reactor temperatures during operation. Failures were pri- marily observed during abrupt power transitions and manifested as sensor drop-outs, drifts, or unexplained behavior. We applied three time-series analysis techniques— rolling mean smoothing, matrix profile, and vector auto-regression (VAR)—to de- tect anomalies in TC data prior to failure events. The rolling mean method effec- tively highlighted deviations aligned with reported failures, while the matrix profile provided partial early warning but sometimes flagged normal fluctuations during power-down periods. VAR shows potential in capturing multivariate dependencies but requires further calibration. A rare case of TC drift was also documented, which did not result in failure, underscoring the challenge of building predictive models with sparse positive examples. Our findings demonstrate that traditional statistical tools can aid anomaly detection but have limited predictive power without richer training data. We propose future directions including synthetic data generation, real- time surrogate modeling, and multi-modal feature integration. This work provides a foundation for applying robust anomaly detection frameworks to mission-critical sensor systems in experimental settings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

BATMODS-lite [SWR-25-108]

Battery Analysis and Training Models for Optimization and Design Studies (BATMODS) is a Python package with an API for pre-built battery models. The original purpose of the package was to quickly generate synthetic data for machine learning models to train with. However, the models are generally useful for any battery simulations or analysis. BATMODS-lite includes the following: 1) A library and API for pre-built battery models 2) Kinetic/transport properties for common battery materials

Randall, Corey [National Laboratory of the Rockies

Spectral Information from Gapped Data. a Comparison of Technique

The fast Fourier transformations (FFT) is used to estimate power spectra of continuous signals evenly sampled on discrete domains. The problem of finding power spectra on unevenly sampled domains, in particular a regularly spaced domain with gaps is discussed. The analysis of the ACRIM solar bolometric intensity data, obtained with a 3/5 on and 2/5 off duty cycle of approximately 100 minutes, would benefit from the techniques. The comparative effectiveness of three different analysis techniques applied to synthetic data generated on gapped domain is reported.

Kuhn, J. R.

Multi-Parent Clustering Algorithms from Stochastic Grammar Data Models

We introduce a statistical data model and an associated optimization-based clustering algorithm which allows data vectors to belong to zero, one or several "parent" clusters. For each data vector the algorithm makes a discrete decision among these alternatives. Thus, a recursive version of this algorithm would place data clusters in a Directed Acyclic Graph rather than a tree. We test the algorithm with synthetic data generated according to the statistical data model. We also illustrate the algorithm using real data from large-scale gene expression assays.

Mjoisness, Eric

Global Precipitation Measurement: GPM Microwave Imager (GMI) Algorithm Development Approach

This slide presentation reviews the approach to the development of the Global Precipitation Measurement algorithm. This presentation includes information about the responsibilities for the development of the algorithm, and the calibration. Also included is information about the orbit, and the sun angle. The test of the algorithm code will be done with synthetic data generated from the Precipitation Processing System (PPS).

Stocker, Erich Franz

Inversion of Multiangular Polarimetric Measurements Over Open and Coastal Ocean Waters: A Joint Retrieval Algorithm for Aerosol and Water-Leaving Radiance Properties

Ocean color remote sensing is a challenging task over coastal waters due to the complex optical properties of aerosols and hydrosols. In order to conduct accurate atmospheric correction, we previously implemented a joint retrieval algorithm, hereafter referred to as the Multi-Angular Polarimetric Ocean coLor (MAPOL) algorithm, to obtain the aerosol and water-leaving signal simultaneously. The MAPOL algorithm has been validated with synthetic data generated by a vector radiative transfer model, and good retrieval performance has been demonstrated in terms of both aerosol and ocean water optical properties (Gao et al., 2018). In this work we applied the algorithm to airborne polarimetric measurements from the Research Scanning Polarimeter (RSP) over both open and coastal ocean waters acquired in two field campaigns: the Ship-Aircraft Bio-Optical Research (SABOR) in 2014 and the North Atlantic Aerosols and Marine Ecosystems Study (NAAMES) in 2015 and 2016. Two different yet related bio-optical models are designed for ocean water properties. One model aligns with traditional open ocean water bio-optical models that parameterize the ocean optical properties in terms of the concentration of chlorophyll a. The other is a generalized bio-optical model for coastal waters that includes seven free parameters to describe the absorption and scattering by phytoplankton, colored dissolved organic matter, and nonalgal particles. The retrieval errors of both aerosol optical depth and the water-leaving radiance are evaluated. Through the comparisons with ocean color data products from both in situ measurements and the Moderate Resolution Imaging Spectroradiometer (MODIS), and the aerosol product from both the High Spectral Resolution Lidar (HSRL) and the Aerosol Robotic Network (AERONET), the MAPOL algorithm demonstrates both flexibility and accuracy in retrieving aerosol and water-leaving radiance properties under various aerosol and ocean water conditions.

Gao, Meng

Distributed Target Tracking With Optimal Data Migration

The paper presents an Extended Kalman Filter based framework for airborne target tracking using dynamic information fusion from multi-modal sensors with geodiversity. First, the algorithm execution location is determined using an optimal data migration strategy, next the sensors information is dynamically fused at each estimation instance using validity flag for each sensor reading, finally the target estimation is updated based on the fused innovation vector. The approach is applied to synthetic data generated from the radar and camera models located on the ground for the simulated target flight in Reflection simulation environment.

Distributed sensing