Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Software for Simulating Remote Sensing Systems

The Application Research Toolbox (ART) is a collection of computer programs that implement algorithms and mathematical models for simulating remote sensing systems. The ART is intended to be especially useful for performing design-tradeoff studies and statistical analyses to support the rational development of design requirements for multispectral imaging systems. Among other things, the ART affords a capability to synthesize coarser-spatial-resolution image-data sets from finer-spatial-resolution data sets and multispectral-image-data products from hyperspectral-image-data products. The ART also provides for synthesis of image-degradation effects, including point-spread functions, misregistration of spectral images, and noise. The ART can utilize real or synthetic data sets, along with sensor specifications, to create simulated data sets. In one example of a typical application, simulated data pertaining to an existing multispectral sensor system are used to verify the data collected by the system in operation. In the case of a proposed sensor system, the simulated data can be used to conduct trade studies and statistical analyses to ensure that the sensor system will satisfy the requirements of potential scientific, academic, and commercial user communities.

Zanoni, Vicki↗

Software for Simulating Remote Sensing Systems

The Application Research Toolbox (ART) is a collection of computer programs that implement algorithms and mathematical models for simulating remote sensing systems. The ART is intended to be especially useful for performing design-tradeoff studies and statistical analyses to support the rational development of design requirements for multispectral imaging systems. Among other things, the ART affords a capability to synthesize coarser-spatial-resolution image-data sets from finer-spatial-resolution data sets and multispectral-image-data products from hyperspectral-image-data products. The ART also provides for synthesis of image-degradation effects, including point-spread functions, misregistration of spectral images, and noise. The ART can utilize real or synthetic data sets, along with sensor specifications, to create simulated data sets. In one example of a typical application, simulated data pertaining to an existing multispectral sensor system are used to verify the data collected by the system in operation. In the case of a proposed sensor system, the simulated data can be used to conduct trade studies and statistical analyses to ensure that the sensor system will satisfy the requirements of potential scientific, academic, and commercial user communities.

Vicki Zanoni↗

Simulating Remote Sensing Systems

The Application Research Toolbox (ART) is a collection of computer programs that implement algorithms and mathematical models for simulating remote sensing systems. The ART is intended to be especially useful for performing design-tradeoff studies and statistical analyses to support the rational development of design requirements for multispectral imaging systems. Among other things, the ART affords a capability to synthesize coarser-spatial-resolution image-data products. The ART also provides for simulations of image-degradation effects, including point-spread functions, misregistration of spectral images, and noise. The ART can utilize real or synthetic data sets, along with sensor specifications, to create simulated data sets. In one example of a particular application, simulated imagery of a coarse resolution system was created using high-resolution imagery from another system in order to perform a radiometric cross-comparison. In the case of a proposed sensor system, the simulated data can be used to conduct trade studies and statistical analyses to ensure that the sensor system will satisfy the requirements of potential scientific, academic, and commercial user communities.

Zanoni, Vicki↗

Flow over an espresso cup: inferring 3-D velocity and pressure fields from tomographic background oriented Schlieren via physics-informed neural networks

Tomographic background oriented Schlieren (Tomo-BOS) imaging measures density or temperature fields in three dimensions using multiple camera BOS projections, and is particularly useful for instantaneous flow visualizations of complex fluid dynamics problems. We propose a new method based on physics-informed neural networks (PINNs) to infer the full continuous three-dimensional (3-D) velocity and pressure fields from snapshots of 3-D temperature fields obtained by Tomo-BOS imaging. The PINNs seamlessly integrate the underlying physics of the observed fluid flow and the visualization data, hence enabling the inference of latent quantities using limited experimental data. In this hidden fluid mechanics paradigm, we train the neural network by minimizing a loss function composed of a data mismatch term and residual terms associated with the coupled Navier–Stokes and heat transfer equations. We first quantify the accuracy of the proposed method based on a two-dimensional synthetic data set for buoyancy-driven flow, and subsequently apply it to the Tomo-BOS data set, where we are able to infer the instantaneous velocity and pressure fields of the flow over an espresso cup based only on the temperature field provided by the Tomo-BOS imaging. Moreover, we conduct an independent PIV experiment to validate the PINN inference for the unsteady velocity field at a centre plane. To explain the observed flow physics, we also perform systematic PINN simulations at different Reynolds and Richardson numbers and quantify the variations in velocity and pressure fields. Furthermore, the results in this paper indicate that the proposed deep learning technique can become a promising direction in experimental fluid mechanics.

97 MATHEMATICS AND COMPUTING↗

HumoNet: A Framework for Realistic Modeling and Simulation of Human Mobility Network

Understanding, analyzing, and predicting human mobility and dynamics are valuable to solving pressing problems, developing effective plans, and prescribing timely remedies. As a computational approach, realistic human mobility simulations allow us to understand, analyze, and predict complex systems, including human societies. Accurate simulations rely on (1) the model that captures interactions and behaviors of myriad entities in our society and (2) the mapping of model instances to real-world entities. Taking this into account, this paper introduces the Human Mobility Network simulation framework (HumoNet), an integrated patterns of life (POL) simulation framework that leverages real-world data layers including transportation networks, points of interest, populations, popularity, and human trajectories. HumoNet is a data informed model in which agents are equipped with activities, locomotion, and planning capabilities. To simulate realistic kinematic maneuvers of individuals in transportation networks, HumoNet harnesses a microscopic traffic simulator that provides interaction among vehicles and traffic objects. In this paper, we describe the framework, outline our methodologies, and discuss the data processing and challenges of each data layer. Through experiments, we demonstrate that our simulations capture key features of human mobility by comparing them to the literature and real data using standard measures of human mobility (i.e., the radius of gyration, number of locations visited, level of exploration) and metrics scoring (i.e., Jensen-Shannon divergence). We envision that the synthetic data produced by HumoNet will serve as a benchmark for analyzing epidemics, deploying EV charging networks, and validating AI/ML tasks such as location prediction.

Kim, Joon-Seok↗

Making Invisible Visible: Data-Driven Seismic Inversion With Spatio-Temporally Constrained Data Augmentation

Deep learning and data-driven approaches have shown great potential in scientific domains. The promise of data-driven techniques relies on the availability of a large volume of high-quality training datasets. Due to the high cost of obtaining data through expensive physical experiments, instruments, and simulations, data augmentation techniques for scientific applications have emerged as a new direction for obtaining scientific data recently. However, existing data augmentation techniques originating from computer vision yield physically unacceptable data samples that are not helpful for the domain problems that we are interested in. In this article, we develop new data augmentation techniques based on convolutional neural networks. Specifically, our generative models leverage different physics knowledge (such as governing equations, observable perception, and physics phenomena) to improve the quality of the synthetic data. To validate the effectiveness of our data augmentation techniques, we apply them to solve a subsurface seismic full-waveform inversion using simulated CO 2 leakage data. Our interest is to invert for subsurface velocity models associated with very small CO 2 leakage. We validate the performance of our methods using comprehensive numerical tests. Here via comparison and analysis, we show that data-driven seismic imaging can be significantly enhanced by using our data augmentation techniques. Particularly, the imaging quality has been improved by 15% in test scenarios of general-sized leakage and 17% in small-sized leakage when using an augmented training set obtained with our techniques.

58 GEOSCIENCES↗

Leveraging Inequality-Constrained Data for Enhanced Liquidus Temperature Prediction in Nuclear Waste Glass Melts

Inequality-constrained data are frequently discarded in engineering, leading to significant information loss in data-scarce domains like glass characterization in nuclear waste vitrification. This paper presents a nonparametric censored-data regression framework based on an l1-norm optimization criterion that leverages slack variables to integrate left-, right-, and interval-constrained observations into training without distributional assumptions. Validated on synthetic data and a Physics-Informed Neural Network (PINN) for predicting liquidus temperature (TL), the method improved R2 from 0.60 to 0.89 and reduced Mean Absolute Error (MAE) by 48% (51.46 to 26.89?rC) on deterministic values. The traditional models failed to satisfy any inequality constraints while the proposed l1-norm PINN satisfies 81.25% of the constraints. The proposed framework effectively extracts actionable information from previously unusable data to enhance predictive accuracy, reduce epistemic uncertainty, and ensure physical consistency in complex industrial applications.

Garcia-Morado, Erick↗

The effects of earth model uncertainty on the inversion of seismic data for seismic source functions

SUMMARY We use Monte Carlo simulations to explore the effects of earth model uncertainty on the estimation of the seismic source time functions that correspond to the six independent components of the point source seismic moment tensor. Specifically, we invert synthetic data using Green’s functions estimated from a suite of earth models that contain stochastic density and seismic wave-speed heterogeneities. We find that the primary effect of earth model uncertainty on the data is that the amplitude of the first-arriving seismic energy is reduced, and that this amplitude reduction is proportional to the magnitude of the stochastic heterogeneities. Also, we find that the amplitude of the estimated seismic source functions can be under- or overestimated, depending on the stochastic earth model used to create the data. This effect is totally unpredictable, meaning that uncertainty in the earth model can lead to unpredictable biases in the amplitude of the estimated seismic source functions.

58 GEOSCIENCES↗

Uncertainty-Aware Machine Learning for Small-Angle X-ray Scattering Analysis in Autonomous Experimentation

Small-angle X-ray scattering (SAXS) is a powerful high-throughput characterization tool for probing nanoscale structure in native sample environments, providing real-time morphological information such as nanoparticle size and shape during synthesis. However, automated SAXS data analysis for extracting meaningful structural parameters is non-trivial and remains a bottleneck in closed-loop experimentation towards autonomous materials discovery, which demands fast, reliable, and uncertainty-aware data analysis. Here, we develop a machine-learning approach for automated SAXS analysis tailored to closed-loop nanoparticle synthesis. A Random Forest (RF) regression model is trained on 100,000 synthetic SAXS curves generated from polydisperse spherical nanoparticles with realistic background contributions. Using normalized one-dimensional SAXS intensity profiles as input, the RF model directly predicts nanoparticle radius, size polydispersity, and background parameters, while the ensemble standard deviation across trees provides built-in uncertainty quantification (UQ). On synthetic data, we show that combining fit-quality metrics (R 2 , MAE) with thresholds on prediction uncertainty reliably identifies accurate parameter estimates without access to ground truth. We then apply the trained model to 365 experimental SAXS profiles of citrate-reduced gold nanoparticles synthesized using an automated droplet-flow microreactor with in situ SAXS at a synchrotron beamline, classifying the results into high- and low-confidence subsets based on UQ metrics. Finally, we integrate RF-based SAXS analysis into a simulated closed-loop optimization campaign using Gaussian process Bayesian optimization to minimize nanoparticle polydispersity, benchmarking against conventional automated Levenberg–Marquardt fitting. The RF-guided campaign exhibits substantially faster convergence and lower relative opportunity cost (∼0.07 vs ∼0.3), demonstrating that uncertainty-aware machine-learning SAXS analysis significantly enhances the efficiency and robustness of autonomous nanomaterials synthesis workflows.

Bayesian optimization↗

A robust estimator of mutual information for deep learning interpretability

Abstract We develop the use of mutual information (MI), a well-established metric in information theory, to interpret the inner workings of deep learning (DL) models. To accurately estimate MI from a finite number of samples, we present GMM-MI (pronounced ‘Jimmie’), an algorithm based on Gaussian mixture models that can be applied to both discrete and continuous settings. GMM-MI is computationally efficient, robust to the choice of hyperparameters and provides the uncertainty on the MI estimate due to the finite sample size. We extensively validate GMM-MI on toy data for which the ground truth MI is known, comparing its performance against established MI estimators. We then demonstrate the use of our MI estimator in the context of representation learning, working with synthetic data and physical datasets describing highly non-linear processes. We train DL models to encode high-dimensional data within a meaningful compressed (latent) representation, and use GMM-MI to quantify both the level of disentanglement between the latent variables, and their association with relevant physical quantities, thus unlocking the interpretability of the latent representation. We make GMM-MI publicly available in this GitHub repository.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Dense Seismic Array Study of a Legacy Underground Nuclear Test at the Nevada National Security Site

The complex postdetonation geologic structures that form after an underground nuclear explosion are difficult to constrain because increased heterogeneity around the damage zone affects seismic waves that propagate through the explosion site. Generally, a vertical rubble-filled structure known as a chimney is formed after an underground nuclear explosion that is composed of debris that falls into the subsurface cavity generated by the explosion. Compared with chimneys that collapse fully, leaving a surface crater, partially collapsed chimneys can have remnant subsurface cavities left in place above collapsed rubble. The 1964 nuclear test HADDOCK, conducted at the Nevada test site (now the Nevada National Security Site), formed a partially collapsed chimney with no surface crater. Understanding the subsurface structure of these features has significant national security applications, such as aiding the study of suspected underground nuclear explosions under a treaty verification. In this study, we investigated the subsurface architecture of the HADDOCK legacy nuclear test using hybrid 2D–3D active source seismic reflection and refraction data. The seismic data were acquired using 275 survey shots from the Seismic Hammer (a 13,000 kg weight drop) and 65 survey shots from a smaller accelerated weight drop, both recorded by ~1000 three-component 5 Hz geophones. First-arrival, P-wave tomographic modeling shows a low-velocity anomaly at ~200 m depth, likely an air-filled cavity caused by partial collapse of the rock column into the temporary postdetonation cavity. A high-velocity anomaly between 20 and 60 m depth represents spall-related compaction of the shallow alluvium. Hints of low velocities are also present near the burial depth (~364 m). The reflection seismic data show a prominent subhorizontal reflector at ~300 m depth, a short-curved reflector at ~200 m, and a high-amplitude reflector at ~50 m depth. Comparisons of the reflection sections to synthetic data and borehole stratigraphy suggest that these features correspond to the alluvium–tuff contact, the partial collapse cavity, and the spalled layer, respectively.

58 GEOSCIENCES↗

Efficient bulk-loading of gridfiles

This paper considers the problem of bulk-loading large data sets for the gridfile multiattribute indexing technique. We propose a rectilinear partitioning algorithm that heuristically seeks to minimize the size of the gridfile needed to ensure no bucket overflows. Empirical studies on both synthetic data sets and on data sets drawn from computational fluid dynamics applications demonstrate that our algorithm is very efficient, and is able to handle large data sets. In addition, we present an algorithm for bulk-loading data sets too large to fit in main memory. Utilizing a sort of the entire data set it creates a gridfile without incurring any overflows.

Leutenegger, Scott T.↗

2021 Smoky Mountains Conference Data Challenge Synthetic-to-Real Domain Adaptation for Autonomous Driving Dataset

The dataset is comprised of both real and synthetic images from a vehicle's forward-facing camera. Each camera image is accompanied by a corresponding pixel-level semantic segmentation image (all files are .png files). In total, the dataset contains 5600 images in the training/validation set and 1400 images in the testing set. The training dataset contains mostly synthetic RGB images collected with a wide range of weather and lighting conditions using the CARLA simulator [1]. In addition, the training data also includes a small pre-selected subset of data from the Cityscapes training dataset – which is comprised of RGB-segmentation image pairs from driving scenarios in various European cities [2]. The testing data is split into three sets. The first set contains synthetic CARLA images with weather/lighting conditions that were not present in the training set. The second set is a subset of the Cityscapes testing dataset. Finally, the third set is an unknown testing set which will not be revealed to the participants until after the submission deadline. [1] Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., and Koltun, V. (2017, October). CARLA: An open urban driving simulator. In Conference on robot learning (pp. 1-16). PMLR. [2] Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., ... and Schiele, B. (2016). The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3213-3223).

99 GENERAL AND MISCELLANEOUS↗

KGML-ag: a modeling framework of knowledge-guided machine learning to simulate agroecosystems: a case study of estimating N<sub>2</sub>O emission using data from mesocosm experiments

Abstract. Agricultural nitrous oxide (N2O) emission accounts for a non-trivial fraction of global greenhouse gas (GHG) budget. To date, estimating N2O fluxes from cropland remains a challenging task because the related microbial processes (e.g., nitrification and denitrification) are controlled by complex interactions among climate, soil, plant and human activities. Existing approaches such as process-based (PB) models have well-known limitations due to insufficient representations of the processes or uncertainties of model parameters, and due to leverage recent advances in machine learning (ML) a new method is needed to unlock the “black box” to overcome its limitations such as low interpretability, out-of-sample failure and massive data demand. In this study, we developed a first-of-its-kind knowledge-guided machine learning model for agroecosystems (KGML-ag) by incorporating biogeophysical and chemical domain knowledge from an advanced PB model, ecosys, and tested it by comparing simulating daily N2O fluxes with real observed data from mesocosm experiments. The gated recurrent unit (GRU) was used as the basis to build the model structure. To optimize the model performance, we have investigated a range of ideas, including (1) using initial values of intermediate variables (IMVs) instead of time series as model input to reduce data demand; (2) building hierarchical structures to explicitly estimate IMVs for further N2O prediction; (3) using multi-task learning to balance the simultaneous training on multiple variables; and (4) pre-training with millions of synthetic data generated from ecosys and fine-tuning with mesocosm observations. Six other pure ML models were developed using the same mesocosm data to serve as the benchmark for the KGML-ag model. Results show that KGML-ag did an excellent job in reproducing the mesocosm N2O fluxes (overall r2=0.81, and RMSE=3.6 mgNm-2d-1 from cross validation). Importantly, KGML-ag always outperforms the PB model and ML models in predicting N2O fluxes, especially for complex temporal dynamics and emission peaks. Besides, KGML-ag goes beyond the pure ML models by providing more interpretable predictions as well as pinpointing desired new knowledge and data to further empower the current KGML-ag. We believe the KGML-ag development in this study will stimulate a new body of research on interpretable ML for biogeochemistry and other related geoscience processes.

54 ENVIRONMENTAL SCIENCES↗

Hybrid data-driven and model-informed online tool wear detection in milling machines

Precision machining tool wear is responsible for low product throughput and quality. Monitoring the tool wear online is vital to prevent degradation in machining quality. However, direct real-time tool wear measurement is not practical. This paper presents residual-based anomaly detection models, combining a hybrid model comprised of a physics-based model and a data-driven model (a decision tree or a neural network) to predict signals of interest (e.g., power or forces) under nominal conditions, followed by Page’s cumulative sum test for detecting tool wear on-line using the computer numerical control machine measurements. The most informative features are ranked using dynamic programming and its approximation variants from real-time measurements and machine settings, such as the width of cut, depth of cut, feed rate and spindle speed, that serve as inputs to the predictive models. The baseline nominal model is incrementally updated with experimental data via a gradient boosted adaptation model to generate the residuals that account for discrepancies between the actual machine data under normal conditions and the baseline nominal model predictions. The hybrid model is validated against 20 Mazak milling machine experimental tests and one Haas run-to-failure experiment. The proposed anomaly detector is applied to synthetic data from simulations of the physics-based model at different operating conditions, measurement noise levels, and tool wear levels, and the methods were able to achieve an overall 92% accuracy in data with 1% noise. The anomaly detection methods based on hybrid model reduced the false alarms of either the data-driven or physical-based models alone, and are found to be capable of good online detection of tool wear.

42 ENGINEERING↗

Hybrid data-driven and model-informed online tool wear detection in milling machines

Precision machining tool wear is responsible for low product throughput and quality. Monitoring the tool wear online is vital to prevent degradation in machining quality. However, direct real-time tool wear measurement is not practical. This paper presents residual-based anomaly detection models, combining a hybrid model comprised of a physics-based model and a data-driven model (a decision tree or a neural network) to predict signals of interest (e.g., power or forces) under nominal conditions, followed by Page’s cumulative sum test for detecting tool wear on-line using the computer numerical control machine measurements. The most informative features are ranked using dynamic programming and its approximation variants from real-time measurements and machine settings, such as the width of cut, depth of cut, feed rate and spindle speed, that serve as inputs to the predictive models. The baseline nominal model is incrementally updated with experimental data via a gradient boosted adaptation model to generate the residuals that account for discrepancies between the actual machine data under normal conditions and the baseline nominal model predictions. The hybrid model is validated against 20 Mazak milling machine experimental tests and one Haas run-to-failure experiment. The proposed anomaly detector is applied to synthetic data from simulations of the physics-based model at different operating conditions, measurement noise levels, and tool wear levels, and the methods were able to achieve an overall 92% accuracy in data with 1% noise. The anomaly detection methods based on hybrid model reduced the false alarms of either the data-driven or physical-based models alone, and are found to be capable of good online detection of tool wear.

Online anomaly detection↗

Distributed Target Tracking With Optimal Data Migration

The paper presents an Extended Kalman Filter based framework for airborne target tracking using dynamic information fusion from multi-modal sensors with geodiversity. First, the algorithm execution location is determined using an optimal data migration strategy, next the sensors information is dynamically fused at each estimation instance using validity flag for each sensor reading, finally the target estimation is updated based on the fused innovation vector. The approach is applied to synthetic data generated from the radar and camera models located on the ground for the simulated target flight in Reflection simulation environment.

Distributed sensing↗

DisasterAWARE – A Global Alerting Platform for Flood Events

The rising number of flooding events combined with increased urbanization is contributing to significant economic losses due to damages to structures and infrastructures. From a risk reduction and resilience perspective, it is not only essential to forecast flood risk and potential impacts, but also to disseminate the information to stakeholders on the ground for rapid implementation of mitigation and response measures. This paper provides (i) an introduction to DisasterAWARE®, a global alerting system, that is used to disseminate flood risk information to stakeholders across the globe, and (ii) a discussion of the models implemented using earth observation data (Synthetic Aperture Radar and optical imagery) for near real-time assessment of flood severity and potential flood impacts to infrastructures. While the models are still in their nascent stage, a case study implementation of the models for the 2020 flooding event in Africa is presented to showcase the model integration with DisasterAWARE®.

54 ENVIRONMENTAL SCIENCES↗