Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

GOOML Big Kahuna Forecast Modeling and Genetic Optimization Files

This submission includes example files associated with the Geothermal Operational Optimization using Machine Learning (GOOML) Big Kahuna fictional power plant, which uses synthetic data to model a fictional power plant. A forecast was produced using the GOOML data model framework and fictional input data, and a genetic optimization is included which determines optimal flash plant parameters. The inputs and outputs associated with the forecast and genetic optimization are included. The input and output files consist of data, configuration files, and plots. A link to the Physics-Guided Neural Networks (phygnn) GitHub repository is also included, which augments a traditional neural network loss function with a generic loss term that can be used to guide the neural network to learn physical or theoretical constraints. phygnn is used by the GOOML framework to help integrate its machine learning models into the relevant physics and engineering applications. Note that the data included in this submission are intended to provide a demonstration of GOOML's capabilities. Additional files that have not been released to the public are needed for users to run these models and reproduce these results. Units can be found in the readme data resource.

15 GEOTHERMAL ENERGY↗

Analysis of urban area land cover using SEASAT Synthetic Aperture Radar data

Digitally processed SEASAT synthetic aperture raar (SAR) imagery of the Denver, Colorado urban area was examined to explore the potential of SAR data for mapping urban land cover and the compatability of SAR derived land cover classes with the United States Geological Survey classification system. The imagery is examined at three different scales to determine the effect of image enlargement on accuracy and level of detail extractable. At each scale the value of employing a simplistic preprocessing smoothing algorithm to improve image interpretation is addressed. A visual interpretation approach and an automated machine/visual approach are employed to evaluate the feasibility of producing a semiautomated land cover classification from SAR data. Confusion matrices of omission and commission errors are employed to define classification accuracies for each interpretation approach and image scale.

Henderson, F. M.↗

Potentials for change detection using Seasat synthetic aperture radar data

The use of synthetic aperture radar (SAR) images for detecting change on the earth's surface is highly dependent on target orientation, azimuth angle, and sensor depression angle. SAR data can be used for change detection when consistency is maintained in radar wavelength, polarization, azimuth directions, and off-nadir depression angle. The interaction of these parameters and the imaged surface for change detection are shown in examples drawn from (1) Los Angeles, CA, (2) southern Florida, (3) Imperial Valley, CA, (4) a desert region west of Tucson, AZ, and (5) western Kansas. SAR imagery is used to emphasize the geometric form, and roughness, of the earth's surface. As changes in the roughness of the surface occur over time, temporal SAR images will indicate those differences. Several guidelines for change detection studies using imaging radar are derived from the examples.

Bryan, M. L.↗

Artificial neural network approach for multiphase segmentation of battery electrode nano-CT images

The segmentation of tomographic images of the battery electrode is a crucial processing step, which will have an additional impact on the results of material characterization and electrochemical simulation. However, manually labeling X-ray CT images (XCT) is time-consuming, and these XCT images are generally difficult to segment with histographical methods. We propose a deep learning approach with an asymmetrical depth encode-decoder convolutional neural network (CNN) for real-world battery material datasets. This network achieves high accuracy while requiring small amounts of labeled data and predicts a volume of billions voxel within few minutes. While applying supervised machine learning for segmenting real-world data, the ground truth is often absent. The results of segmentation are usually qualitatively justified by visual judgement. We try to unravel this fuzzy definition of segmentation quality by identifying the uncertainty due to the human bias diluted in the training data. Further CNN trainings using synthetic data show quantitative impact of such uncertainty on the determination of material’s properties. Nano-XCT datasets of various battery materials have been successfully segmented by training this neural network from scratch. We will also show that applying the transfer learning, which consists of reusing a well-trained network, can improve the accuracy of a similar dataset.

25 ENERGY STORAGE↗

Hiperclust

This software leverages transfer learning to analyze atom probe tomography (APT) data. It is trained on synthetic data and then applies this knowledge to predict the optimal number of clusters for a given APT dataset. Initially, the software used preliminary clustering to estimate the general structure of the data. Based on this, it provides suggestions for key parameters like minimum cluster size and minimum number of points. These parameters are critical for algorithms like HDBSCAN, ensuring accurate cluster formation without the need for trial-and-error testing. The software runs on High-Performance computing (HPC) systems, enabling fast, scalable analysis of large APT datasets, ultimately saving time and improving the reliability of clustering outcomes.

Tang, Yalei [Idaho National Laboratory (INL), Idah↗

Cloverleaf Data Artifacts for ArtIMis LDRD

This report summarizes the use of the open-source CloverLeaf/CloverLeaf3D mini-apps to generate synthetic data sets to train foundation models for the ArtIMis LDRD DI. These data artifacts are intended to be used by LANL collaborators and shared externally with our university and institutional partners. Note that CloverLeaf/CloverLeaf3D is not a LANL simulation code.

97 MATHEMATICS AND COMPUTING↗

Computational Characterization and Model Verification for 3D Microstructure Reconstruction of Additively-Manufactured Materials

The goal of this study is to characterize and validate the texture and grain topology of additively-manufactured anisotropic three-dimensional (3D) polycrystalline microstructures. The special focus is on developing methodologies to compare the grain shapes and orientations of two-dimensional (2D) and 3D microstructure representations using the same metric. To generate statistical data, synthetic microstructures are reconstructed from experimental data using Markov random field (MRF). The statistical similarity between the experimental and synthetic microstructures is verified by comparing their grain topologies. A universal measure to compare 2D and 3D grains is portrayed through the concept of image moments that are invariant to shape transformations. The graphical plots developed based on moment invariants to compare the 2D and 3D grains are used to verify the synthetic model

Materials Characterization↗

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT↗

A novel closed-form inversion of the convection–diffusion equation for rapid convection, diffusion, and source profile estimation

To simplify and routinize particle transport analysis in fusion devices, a novel closed form linear inversion of the 1-D convection diffusion equation to estimate diffusion and convection profiles D(r ⃗ ), v(r ⃗ ) and source distribution s(r ⃗ ), of a single species from measured data is derived and demonstrated on synthetic data. Profile estimates of D(r ⃗ ), v(r ⃗ ), s(r ⃗ ) and their uncertainties are given as a matrix expression constructed directly from the incoming density data of the transported species in space and time, as well as physics assumptions such as particle conservation and experimental geometry. The derived matrix expression can be applied to a pumped or non-pumped recycling species, or a non-recycling species that is effectively “pumped” by plasma-facing surfaces.

Hinson, Edward [ORNL] (ORCID:000000019713140X)↗

Anomaly Detection for Online Monitoring of Thermocouple Sensors in the Advanced Test Reactor

This study explores data-driven anomaly detection methods to analyze sensor fail- ures in the Advanced Gas Reactor (AGR) nuclear fuel irradiation experiments. Specifically, we examine failures of thermocouples (TCs), which are critical for mon- itoring and controlling in-reactor temperatures during operation. Failures were pri- marily observed during abrupt power transitions and manifested as sensor drop-outs, drifts, or unexplained behavior. We applied three time-series analysis techniques— rolling mean smoothing, matrix profile, and vector auto-regression (VAR)—to de- tect anomalies in TC data prior to failure events. The rolling mean method effec- tively highlighted deviations aligned with reported failures, while the matrix profile provided partial early warning but sometimes flagged normal fluctuations during power-down periods. VAR shows potential in capturing multivariate dependencies but requires further calibration. A rare case of TC drift was also documented, which did not result in failure, underscoring the challenge of building predictive models with sparse positive examples. Our findings demonstrate that traditional statistical tools can aid anomaly detection but have limited predictive power without richer training data. We propose future directions including synthetic data generation, real- time surrogate modeling, and multi-modal feature integration. This work provides a foundation for applying robust anomaly detection frameworks to mission-critical sensor systems in experimental settings.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

An efficient method to propagate model uncertainty when inverting seismic data for time domain seismic moment tensors

SUMMARY We present a computationally efficient method to approximately propagate uncertainty when linearly inverting seismic data for point source, time variable moment tensor components. The method is based on the assumption that the data residual, given by the difference between the observed seismic data and the data predicated by a linear inversion, contains the effects of both data and model uncertainty. Our method uses a distribution of data residuals, added directly to the data, in a pseudo-Monte Carlo scheme. Using the assumption that the data residual is a stochastic process, we use the well-known Karhunen–Loève (KL) theorem to construct a distribution of data residuals, where the required basis functions are constructed using Fourier series. The Fourier series are scaled by a product of a random variable and the real-valued spectral amplitudes of the original data residual’s spectrum. Thus, the Fourier series and spectral amplitudes are eigenfunction-eigenvalue pairs used in the KL-based construction of data residual distribution. Using tests with synthetic data, we show that our method compares closely with a Finite Difference Monte Carlo (FDMC) method that we presented previously. More importantly, the method presented here is computationally several orders of magnitude faster than our previous FDMC method, and requires no a priori assumptions of model and/or data uncertainty.

Poppeliers, Christian (ORCID:0000000159526849)↗

A modified objective mapping technique for scatterometer wind data

A method for generating high-resolution wind maps from scatterometer data was developed and tested on synthetic data for the northeast Pacific Ocean. It is shown that, unlike the wind fields generated by current GCMs, the wind maps constructed by this method retain the high spatial resolution of the scatterometer wherever adequate measurements exist. For the NASA scatterometer, this method would produce every 12 hours a wind map with spatial resolution that preserves the small-scale features of the original data over about half the mapped region. Over the rest of the region, maps with somewhat lower resolution and accuracy will be obtained.

Kelly, Kathryn A.↗

On the Isothermality of Solar Plasmas

Recent measurements have shown that the quiet unstructured solar corona observed at the solar limb is close to isothermal, at a temperature that does not appear to change over wide areas or with time. Some in dividual active loop structures have also been found to be nearly iso thermal both along their axis and across their cross-section. Even a complex active region observed at the solar limb has been found to be composed of three distinct isothermal plasmas. If confirmed, these r esults would pose formidable challenges to the current theoretical understanding of the thermal structure and heating of the solar corona. For example, no current theoretical model can explain the excess dens ities and lifetimes of many observed loops if the loops are in fact i sothermal. All of these measurements are based on the so-called emiss ion measure (EM) diagnostic technique that is applied to a set of opt ically thin lines under the assumption of isothermal plasma. It provi des simultaneous measurement of both the temperature and EM. However, no study has ever been carried out to quantify the uncertainties in the technique and to rigorously assess its ability to discriminate bet ween isothermal and multithermal plasmas. Such a study is the topic o f the present work. We define a formal measure of the uncertainty in the EM diagnostic technique that can easily be applied to real data. We here apply it to synthetic data based on a variety of assumed plas ma thermal distributions, and develop a method to quantitatively asse ss the degree of multithermality of a plasma.

Landi, E.↗

Comparison of Computational and Experimental Microphone Array Results for an 18%-Scale Aircraft Model

An 18%-scale, semi-span model is used as a platform for examining the efficacy of microphone array processing using synthetic data from numerical simulations. Two hybrid RANS/LES codes coupled with Ffowcs Williams-Hawkings solvers are used to calculate 97 microphone signals at the locations of an array employed in the NASA LaRC 14x22 tunnel. Conventional, DAMAS, and CLEAN-SC array processing is applied in an identical fashion to the experimental and computational results for three different configurations involving deploying and retracting the main landing gear and a part span flap. Despite the short time records of the numerical signals, the beamform maps are able to isolate the noise sources, and the appearance of the DAMAS synthetic array maps is generally better than those from the experimental data. The experimental CLEAN-SC maps are similar in quality to those from the simulations indicating that CLEAN-SC may have less sensitivity to background noise. The spectrum obtained from DAMAS processing of synthetic array data is nearly identical to the spectrum of the center microphone of the array, indicating that for this problem array processing of synthetic data does not improve spectral comparisons with experiment. However, the beamform maps do provide an additional means of comparison that can reveal differences that cannot be ascertained from spectra alone.

Lockard, David P.↗

Deep learning inversion of gravity data for detection of CO 2 plumes in overlying aquifers

In this work, we developed an effective U-Net based deep learning (DL) model for inversion of surface gravity data on a rectangular grid to predict 2-D high-resolution subsurface CO 2 distribution along a vertical cross-section due to CO 2 leakage through a wellbore within a deep CO 2 storage reservoir. We used synthetic data to model two types of CO 2 leakage scenarios: one CO 2 plume in a shallow aquifer (single plume case), and two plumes present at different depths (double plume case). The 3-D synthetic plume samples were created by sampling among predetermined CO 2 plume depths, saturations, and volumes. The corresponding surface gravity data on a rectangular grid were generated by a 3-D forward model. The U-Net model detected 72% of single-plume samples, and one or both plumes in 75% of double-plume samples. Most of the undetected single plumes have small gravity field strengths below the typical noise level of 5 μGal. This model generated reproducible, reliable predictions with acceptable errors and demonstrated improved spatial resolution over the conventional least-squares inversion. In contrast to the conventional least-squares inversion, which often overestimates the size of its target and underestimates its density, this U-Net model accurately delineated the boundary of a target. Furthermore, this DL inversion detected deep, small, or low saturation CO 2 plumes that are often more difficult to resolve with conventional gravity inversion methods. We note the limitations of this feasibility study, including the use of synthetic data with regular CO 2 plume shapes, and the prediction of a 2-D plume cross-section rather than the full 3-D plume, as well, we recognize the lower detection fraction for double-plume scenarios. Nevertheless, this study demonstrates that DL gravity inversion is a promising and potentially superior method to conventional least-squares inversion. Our U-Net based deep learning inversion approach may be adapted for inversion of other types of geophysical data. DL inversion can facilitate near real-time monitoring of geologic carbon sequestration to provide site operators with prompt information about subsurface CO 2 distribution for risk management and mitigation.

58 GEOSCIENCES↗

A system identification approach for non-intrusive reduced order modeling of radiation-induced photocurrents

In this study, development of compact photocurrent models is currently dominated by analytical techniques that rely on physical assumptions to render the governing equations solvable in a closed form. Violation of these assumptions can reduce the accuracy of the models and/or limit their scope. In this paper we show that system identification of nonlinear state-space systems can serve as an alternative numerical basis for non-intrusive reduced order modeling of photocurrent effects. To that end we develop a compact gray box photocurrent model (GBPM) by using a state-space representation with a low-dimensional latent state equation that mimics a mathematical model for the response of an idealized class of devices to ionizing radiation. In so doing we obtain a model that learns the dynamics of a quantity of interest directly from its measurements without requiring snapshots of the internal device state or its discretized model, and can be inferred from very small data sets. To demonstrate the approach we train the GBPM using a small experimental data set for a Z5236 Zener diode and a small synthetic data set obtained by simulating a synthetic pn-junction device. We then compare the GBPMs with black box models trained on the same data and show that performance of the latter is limited by the size of the data set, while the former are able to achieve excellent performance in both the reproductive and the predictive regimes.

97 MATHEMATICS AND COMPUTING↗

Evaluation of a preliminary regional Earth model through comparison of synthetic and observed waveform data

In this report, we document the process related to developing a regional geologic model of a 605 x 1334 km area centered around Utah and encompassing surrounding states. This model is developed to test the effect that composition of a model has on the generation of synthetic data with the intent of using this information to improve upon full waveform moment tensor inversions. We compare observed data from three seismic events and five stations to the synthetic data generated by a preliminary model derived from a geologic framework model (GFM) developed by the USGS. The synthetic data and observed data comparisons indicate that our preliminary model performs well at smaller offset distances in the northern and central sections of the model. However, the southern stations consistently display synthetic data P- and S-wave arrival times that do not match the observed data arrival times, indicating that the velocity structure of the southern part of the model especially is inaccurate.

58 GEOSCIENCES↗

Gaussian processes for inferring parton distributions

The extraction of parton distribution functions (PDFs) from experimental or lattice QCD data is an ill-posed inverse problem, where regularization strongly impacts both systematic uncertainties and the reliability of the results. We study a framework based on Gaussian Process Regression (GPR) to reconstruct PDFs from lattice QCD matrix elements. Within a Bayesian framework, Gaussian processes serve as flexible priors that encode uncertainties, correlations, and constraints without imposing rigid functional forms. We investigate a wide range of kernel choices, mean functions, and hyperparameter treatments. We quantify information gained from the data using the Kullback-Leibler divergence. Synthetic data tests demonstrate the consistency and robustness of the method. Our study establishes GPR as a systematic and non-parametric approach to PDF reconstruction, offering controlled uncertainty estimates and reduced model bias in lattice QCD analyses.

hadronic spectroscopy↗