Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

CALPHAD Uncertainty Quantification and TDBX

CALPHAD uncertainty quantification (UQ) is the foundation of materials design with quantified confidence. We report a framework and software packages to enable CALPHAD UQ assessment and calculation using commercial CALPHAD software (Thermo-Calc). This Bayesian inference framework is coupled with a Markov chain Monte Carlo algorithm to establish uncertainty traces with a given thermodynamic database file (TDB) and corresponding experimental data points. This general framework is demonstrated with the Ni–Cr binary system. The algorithm is firstly validated on synthetic data with known ground truth. Then it is applied to real experimental data to generate posterior traces. We develop a file format named TDBX, which provides a single source of truth by combining the original TDB content and the traces for each assessed Gibbs energy parameter. CALPHAD UQ calculations are performed based on the TDBX file, from which uncertainties for phase boundaries, enthalpy curves, and solidification range are collected as examples of basic design parameters. This TDBX file with corresponding scripts are made open-source. Finally, the combination of CALPHAD UQ assessments and calculations connected by TDBX supports uncertainty-assisted modeling, enabling the integrated application of modern design with uncertainty methodologies to computational materials design.

36 MATERIALS SCIENCE↗

Energy landscapes from cryo-EM snapshots: a benchmarking study

Abstract Biomolecules undergo continuous conformational motions, a subset of which are functionally relevant. Understanding, and ultimately controlling biomolecular function are predicated on the ability to map continuous conformational motions, and identify the functionally relevant conformational trajectories. For equilibrium and near-equilibrium processes, function proceeds along minimum-energy pathways on one or more energy landscapes, because higher-energy conformations are only weakly occupied. With the growing interest in identifying functional trajectories, the need for reliable mapping of energy landscapes has become paramount. In response, various data-analytical tools for determining structural variability are emerging. A key question concerns the veracity with which each data-analytical tool can extract functionally relevant conformational trajectories from a collection of single-particle cryo-EM snapshots. Using synthetic data as an independently known ground truth, we benchmark the ability of four leading algorithms to determine biomolecular energy landscapes and identify the functionally relevant conformational paths on these landscapes. Such benchmarking is essential for systematic progress toward atomic-level movies of continuous biomolecular function.

59 BASIC BIOLOGICAL SCIENCES↗

Physical discovery in representation learning via conditioning on prior knowledge

Recent advances in electron, scanning probe, optical, and chemical imaging and spectroscopy yield bespoke data sets containing the information of structure and functionality of complex systems. In many cases, the resulting data sets are underpinned by low-dimensional simple representations encoding the factors of variability within the data. The representation learning methods seek to discover these factors of variability, ideally further connecting them with relevant physical mechanisms. However, generally, the task of identifying the latent variables corresponding to actual physical mechanisms is extremely complex. Here, we present an empirical study of an approach based on conditioning the data on the known (continuous) physical parameters and systematically compare it with the previously introduced approach based on the invariant variational autoencoders. The conditional variational autoencoder (cVAE) approach does not rely on the existence of the invariant transforms and hence allows for much greater flexibility and applicability. Interestingly, cVAE allows for limited extrapolation outside of the original domain of the conditional variable. However, this extrapolation is limited compared to the cases when true physical mechanisms are known, and the physical factor of variability can be disentangled in full. We further show that introducing the known conditioning results in the simplification of the latent distribution if the conditioning vector is correlated with the factor of variability in the data, thus allowing us to separate relevant physical factors. We initially demonstrate this approach using 1D and 2D examples on a synthetic data set and then extend it to the analysis of experimental data on ferroelectric domain dynamics visualized via piezoresponse force microscopy.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Decoding the shift-invariant data: applications for band-excitation scanning probe microscopy *

A shift-invariant variational autoencoder (shift-VAE) is developed as an unsupervised method for the analysis of spectral data in the presence of shifts along the parameter axis, disentangling the physically-relevant shifts from other latent variables. Using synthetic data sets, we show that the shift-VAE latent variables closely match the ground truth parameters. The shift VAE is extended towards the analysis of band-excitation piezoresponse force microscopy data, disentangling the resonance frequency shifts from the peak shape parameters in a model-free unsupervised manner. The extensions of this approach towards denoising of data and model-free dimensionality reduction in imaging and spectroscopic data are further demonstrated. This approach is universal and can also be extended to analysis of x-ray diffraction, photoluminescence, Raman spectra, and other data sets.

36 MATERIALS SCIENCE↗

Validation of non-negative matrix factorization for rapid assessment of large sets of atomic pair distribution function data

The use of the non-negative matrix factorization (NMF) technique is validated for automatically extracting physically relevant components from atomic pair distribution function (PDF) data from time-series data such as in situ experiments. The use of two matrix-factorization techniques, principal component analysis and NMF, on PDF data is compared in the context of a chemical synthesis reaction taking place in a synchrotron beam, applying the approach to synthetic data where the correct composition is known and on measured PDFs from previously published experimental data. The NMF approach yields mathematical components that are very close to the PDFs of the chemical components of the system and a time evolution of the weights that closely follows the ground truth. Lastly, it is discussed how this would appear in a streaming context if the analysis were being carried out at the beamline as the experiment progressed.

36 MATERIALS SCIENCE↗

Distributed Optimization for Nonrigid Nano-Tomography

Resolution level and reconstruction quality in nano-computed tomography (nano-CT) are in part limited by the stability of microscopes, because the magnitude of mechanical vibrations during scanning becomes comparable to the imaging resolution, and the ability of the samples to resist radiation induced deformations during data acquisition. In such cases, there is no incentive in recovering the sample state at different time steps like in time-resolved reconstruction methods, but instead the goal is to retrieve a single reconstruction at the highest possible spatial resolution and without any imaging artifacts. Here we propose a distributed optimization solver for tomographic imaging of samples at the nanoscale. Our approach solves the tomography problem jointly with projection data alignment, nonrigid sample deformation correction, and regularization. Projection data consistency is regulated by dense optical flow estimated by Farneback's algorithm, leading to sharp sample reconstructions with less artifacts. Synthetic data tests show robustness of the method to Poisson and low-frequency background noise. We accelerated the solver on multi-GPU systems and validated the method on three nano-imaging experimental data sets.

97 MATHEMATICS AND COMPUTING↗

A Bayesian approach to time-domain photonic Doppler velocimetry analysis

Photonic Doppler velocimetry (PDV) is an established technique for measuring the velocities of fast-moving surfaces in high-energy-density experiments. In the standard approach to PDV analysis, the short-time Fourier transform (STFT) is used to generate a spectrogram from which the velocity history of the target is inferred. The user chooses the form, duration, and separation of the window function. Here, in this study, we present a Bayesian approach to infer the velocity directly from the PDV oscilloscope trace, without using the spectrogram for analysis. This is clearly a difficult inference problem due to the highly periodic nature of the data, but we find that with carefully chosen prior distributions for the model parameters, we can accurately recover the injected velocity from synthetic data. We validate this method using PDV data collected at the STAR two-stage light gas gun at Sandia National Laboratories, recovering shock-front velocities in quartz that are consistent with those inferred using the STFT-based approach and are interpolated across regions of low signal-to-noise data. Although this method does not rely on the same user choices as the STFT, we caution that it can be prone to misspecification if the chosen model is not sufficient to capture the velocity behavior. Analysis using posterior predictive checks can be used to establish whether a better model is required, although more complex models come with additional computational cost, often taking more than several hours to converge when sampling the Bayesian posterior. We, therefore, recommend it be viewed as a complementary method to that of the STFT-based approach.

Allison, James R. [First Light Fusion Ltd., Yarnto↗

Compactly‐Supported Nonstationary Kernels for Computing Exact Gaussian Processes on Big Data

The Gaussian process (GP) is a widely used method for analyzing large-scale data sets, including spatio-temporal measurements of nonlinear processes that are now commonplace in the environmental sciences. Traditional implementations of GPs involve stationary kernels (also termed covariance functions) that limit their flexibility, and exact methods for inference that prevent application to data sets with more than about 10,000 points. Modern approaches to address stationarity assumptions generally fail to accommodate large data sets, while all attempts to address scalability focus on approximating the Gaussian likelihood, which can involve subjectivity and lead to inaccuracies. In this work, we explicitly derive an alternative kernel that can discover and encode both sparsity and nonstationarity. We embed the kernel within a fully Bayesian GP model and leverage high-performance computing resources to enable the analysis of massive data sets. We demonstrate the favorable performance of our novel kernel relative to existing exact and approximate GP methods across a variety of synthetic data examples. Furthermore, we conduct space–time prediction based on more than 1 million measurements of daily maximum temperature and verify that our results outperform state-of-the-art methods in the Earth sciences. More broadly, having access to exact GPs that use ultra-scalable, sparsity-discovering, nonstationary kernels allows GP methods to truly compete with a wide variety of machine learning methods.

Gaussian processes↗

Predictive Modeling and Uncertainty Quantification in Condition Monitoring of Active Components: A Reactor Coolant Pump Use Case

This work develops data-driven models for onset of thermal barrier leakage in reactor coolant pumps. It incorporates uncertainty quantification to enhance the reliability and robustness of pre- dictions. Using synthetic data generated by the Generic Pressurized Water Reactor simulator, realistic degradation scenarios were simulated across lifecycle stages—beginning, middle, and end of life. Key variables, including differential pressure, flow rate, vibration, and temperatures, were analyzed using machine learning framework. The fully connected neural network models demonstrated exceptional performance, achieving R2 scores exceeding 0.99 and root mean square errors as low as around 8.23 × 10-2 gallon per minute (gpm) for the three stages of the lifecy- cle. UQ analysis further validated the model’s robustness, with narrow uncertainty bounds during steady-state operations and appropriately wider bounds during transitional phases, reflecting the physical behavior of the system. This work addresses important gaps in real-time condition moni- toring and regulatory compliance by integrating advanced condition monitoring technologies with UQ into IST programs. The ability to detect thermal barrier leakage early and quantify prediction reliability supports optimizing maintenance strategies while ensuring nuclear power plants’ safe and reliable operation.

99 - GENERAL AND MISCELLANEOUS↗

Benchmarking Soft Sensors for Remote Monitoring of On-Site Wastewater Treatment Plants

On-site wastewater treatment plants (OSTs) are usually unattended, so failures often remain undetected and lead to prolonged periods of reduced performance. To stabilize the performance of unattended plants, soft sensors could expose faults and failures to the operator. In a previous study, we developed soft sensors and showed that soft sensors with data from unmaintained physical sensors can be as accurate as soft sensors with data from maintained ones. The monitored variables were pH and dissolved oxygen (DO), and soft sensors were used to predict nitrification performance. In the present study, we use synthetic data and monitor three plants to test these soft sensors. We find that a long solids retention time and a moderate aeration rate improve the pH soft-sensor accuracy and that the aeration regime is the main operational parameter affecting the accuracy of the DO soft sensor. We demonstrate that integrated design of monitoring and control is necessary to achieve robustness when extrapolating from one OST to another in the absence of plant-specific fine-tuning. Additionally, we provide a unique labeled dataset for further feature and data-driven soft-sensor development. Our benchmarking results indicate that it is feasible to monitor OSTs with unmaintained sensors and without plant-specific tuning of the developed soft sensors. This is expected to drastically reduce monitoring costs for OST-based sanitation systems.

54 ENVIRONMENTAL SCIENCES↗

Conformal Hierarchical Simulation-Based Inference with Local Validity

Trustworthy and interpretable uncertainty quantification is a long-standing challenge in artificial intelligence. Simulation-based inference (SBI) comprises a broad swath of approaches for estimating latent parameters with uncertainties. Although flexible neural density estimators in SBI can be remark- ably expressive capturing highly structured, high-dimensional posteriors their credible regions can be badly mis-calibrated and are often only accompanied by heuristic coverage checks. We present the first SBI framework that delivers finite-sample local valid coverage guarantees that hold in the neighborhood of each observation. Our framework can couple any off-the-shelf hierarchical SBI engine with a confor- mal Bayesian post-processing step that operates on the posterior predictive density. A kernel-weighted conformity score adapts the conformal quantile to the local geometry of the data, yielding prediction sets that are simultaneously (i) marginally calibrated, (ii) locally valid, and (iii) hierarchical, handling global and observation-specific parameters in a single pass. Through experiments on synthetic data and benchmarks from neuroscience and physics, we show that our approach attains 1 − α coverage, where prior SBI methods under- or over-cover. Our approach also maintains a competitive, credible set size with minimal computational overhead. Finally, our approach can be used to make predictions on real data and give valid credible regions modulo weight-initialization-based model mis-specification.

Trivedi, Shubhendu [Fermilab]↗

Flow over an espresso cup: inferring 3-D velocity and pressure fields from tomographic background oriented Schlieren via physics-informed neural networks

Tomographic background oriented Schlieren (Tomo-BOS) imaging measures density or temperature fields in three dimensions using multiple camera BOS projections, and is particularly useful for instantaneous flow visualizations of complex fluid dynamics problems. We propose a new method based on physics-informed neural networks (PINNs) to infer the full continuous three-dimensional (3-D) velocity and pressure fields from snapshots of 3-D temperature fields obtained by Tomo-BOS imaging. The PINNs seamlessly integrate the underlying physics of the observed fluid flow and the visualization data, hence enabling the inference of latent quantities using limited experimental data. In this hidden fluid mechanics paradigm, we train the neural network by minimizing a loss function composed of a data mismatch term and residual terms associated with the coupled Navier–Stokes and heat transfer equations. We first quantify the accuracy of the proposed method based on a two-dimensional synthetic data set for buoyancy-driven flow, and subsequently apply it to the Tomo-BOS data set, where we are able to infer the instantaneous velocity and pressure fields of the flow over an espresso cup based only on the temperature field provided by the Tomo-BOS imaging. Moreover, we conduct an independent PIV experiment to validate the PINN inference for the unsteady velocity field at a centre plane. To explain the observed flow physics, we also perform systematic PINN simulations at different Reynolds and Richardson numbers and quantify the variations in velocity and pressure fields. Furthermore, the results in this paper indicate that the proposed deep learning technique can become a promising direction in experimental fluid mechanics.

97 MATHEMATICS AND COMPUTING↗

HumoNet: A Framework for Realistic Modeling and Simulation of Human Mobility Network

Understanding, analyzing, and predicting human mobility and dynamics are valuable to solving pressing problems, developing effective plans, and prescribing timely remedies. As a computational approach, realistic human mobility simulations allow us to understand, analyze, and predict complex systems, including human societies. Accurate simulations rely on (1) the model that captures interactions and behaviors of myriad entities in our society and (2) the mapping of model instances to real-world entities. Taking this into account, this paper introduces the Human Mobility Network simulation framework (HumoNet), an integrated patterns of life (POL) simulation framework that leverages real-world data layers including transportation networks, points of interest, populations, popularity, and human trajectories. HumoNet is a data informed model in which agents are equipped with activities, locomotion, and planning capabilities. To simulate realistic kinematic maneuvers of individuals in transportation networks, HumoNet harnesses a microscopic traffic simulator that provides interaction among vehicles and traffic objects. In this paper, we describe the framework, outline our methodologies, and discuss the data processing and challenges of each data layer. Through experiments, we demonstrate that our simulations capture key features of human mobility by comparing them to the literature and real data using standard measures of human mobility (i.e., the radius of gyration, number of locations visited, level of exploration) and metrics scoring (i.e., Jensen-Shannon divergence). We envision that the synthetic data produced by HumoNet will serve as a benchmark for analyzing epidemics, deploying EV charging networks, and validating AI/ML tasks such as location prediction.

Kim, Joon-Seok↗

Making Invisible Visible: Data-Driven Seismic Inversion With Spatio-Temporally Constrained Data Augmentation

Deep learning and data-driven approaches have shown great potential in scientific domains. The promise of data-driven techniques relies on the availability of a large volume of high-quality training datasets. Due to the high cost of obtaining data through expensive physical experiments, instruments, and simulations, data augmentation techniques for scientific applications have emerged as a new direction for obtaining scientific data recently. However, existing data augmentation techniques originating from computer vision yield physically unacceptable data samples that are not helpful for the domain problems that we are interested in. In this article, we develop new data augmentation techniques based on convolutional neural networks. Specifically, our generative models leverage different physics knowledge (such as governing equations, observable perception, and physics phenomena) to improve the quality of the synthetic data. To validate the effectiveness of our data augmentation techniques, we apply them to solve a subsurface seismic full-waveform inversion using simulated CO 2 leakage data. Our interest is to invert for subsurface velocity models associated with very small CO 2 leakage. We validate the performance of our methods using comprehensive numerical tests. Here via comparison and analysis, we show that data-driven seismic imaging can be significantly enhanced by using our data augmentation techniques. Particularly, the imaging quality has been improved by 15% in test scenarios of general-sized leakage and 17% in small-sized leakage when using an augmented training set obtained with our techniques.

58 GEOSCIENCES↗

Leveraging Inequality-Constrained Data for Enhanced Liquidus Temperature Prediction in Nuclear Waste Glass Melts

Inequality-constrained data are frequently discarded in engineering, leading to significant information loss in data-scarce domains like glass characterization in nuclear waste vitrification. This paper presents a nonparametric censored-data regression framework based on an l1-norm optimization criterion that leverages slack variables to integrate left-, right-, and interval-constrained observations into training without distributional assumptions. Validated on synthetic data and a Physics-Informed Neural Network (PINN) for predicting liquidus temperature (TL), the method improved R2 from 0.60 to 0.89 and reduced Mean Absolute Error (MAE) by 48% (51.46 to 26.89?rC) on deterministic values. The traditional models failed to satisfy any inequality constraints while the proposed l1-norm PINN satisfies 81.25% of the constraints. The proposed framework effectively extracts actionable information from previously unusable data to enhance predictive accuracy, reduce epistemic uncertainty, and ensure physical consistency in complex industrial applications.

Garcia-Morado, Erick↗

The effects of earth model uncertainty on the inversion of seismic data for seismic source functions

SUMMARY We use Monte Carlo simulations to explore the effects of earth model uncertainty on the estimation of the seismic source time functions that correspond to the six independent components of the point source seismic moment tensor. Specifically, we invert synthetic data using Green’s functions estimated from a suite of earth models that contain stochastic density and seismic wave-speed heterogeneities. We find that the primary effect of earth model uncertainty on the data is that the amplitude of the first-arriving seismic energy is reduced, and that this amplitude reduction is proportional to the magnitude of the stochastic heterogeneities. Also, we find that the amplitude of the estimated seismic source functions can be under- or overestimated, depending on the stochastic earth model used to create the data. This effect is totally unpredictable, meaning that uncertainty in the earth model can lead to unpredictable biases in the amplitude of the estimated seismic source functions.

58 GEOSCIENCES↗

Uncertainty-Aware Machine Learning for Small-Angle X-ray Scattering Analysis in Autonomous Experimentation

Small-angle X-ray scattering (SAXS) is a powerful high-throughput characterization tool for probing nanoscale structure in native sample environments, providing real-time morphological information such as nanoparticle size and shape during synthesis. However, automated SAXS data analysis for extracting meaningful structural parameters is non-trivial and remains a bottleneck in closed-loop experimentation towards autonomous materials discovery, which demands fast, reliable, and uncertainty-aware data analysis. Here, we develop a machine-learning approach for automated SAXS analysis tailored to closed-loop nanoparticle synthesis. A Random Forest (RF) regression model is trained on 100,000 synthetic SAXS curves generated from polydisperse spherical nanoparticles with realistic background contributions. Using normalized one-dimensional SAXS intensity profiles as input, the RF model directly predicts nanoparticle radius, size polydispersity, and background parameters, while the ensemble standard deviation across trees provides built-in uncertainty quantification (UQ). On synthetic data, we show that combining fit-quality metrics (R 2 , MAE) with thresholds on prediction uncertainty reliably identifies accurate parameter estimates without access to ground truth. We then apply the trained model to 365 experimental SAXS profiles of citrate-reduced gold nanoparticles synthesized using an automated droplet-flow microreactor with in situ SAXS at a synchrotron beamline, classifying the results into high- and low-confidence subsets based on UQ metrics. Finally, we integrate RF-based SAXS analysis into a simulated closed-loop optimization campaign using Gaussian process Bayesian optimization to minimize nanoparticle polydispersity, benchmarking against conventional automated Levenberg–Marquardt fitting. The RF-guided campaign exhibits substantially faster convergence and lower relative opportunity cost (∼0.07 vs ∼0.3), demonstrating that uncertainty-aware machine-learning SAXS analysis significantly enhances the efficiency and robustness of autonomous nanomaterials synthesis workflows.

Bayesian optimization↗

A robust estimator of mutual information for deep learning interpretability

Abstract We develop the use of mutual information (MI), a well-established metric in information theory, to interpret the inner workings of deep learning (DL) models. To accurately estimate MI from a finite number of samples, we present GMM-MI (pronounced ‘Jimmie’), an algorithm based on Gaussian mixture models that can be applied to both discrete and continuous settings. GMM-MI is computationally efficient, robust to the choice of hyperparameters and provides the uncertainty on the MI estimate due to the finite sample size. We extensively validate GMM-MI on toy data for which the ground truth MI is known, comparing its performance against established MI estimators. We then demonstrate the use of our MI estimator in the context of representation learning, working with synthetic data and physical datasets describing highly non-linear processes. We train DL models to encode high-dimensional data within a meaningful compressed (latent) representation, and use GMM-MI to quantify both the level of disentanglement between the latent variables, and their association with relevant physical quantities, thus unlocking the interpretability of the latent representation. We make GMM-MI publicly available in this GitHub repository.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗