Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Calibration verification for stochastic agent-based disease spread models

Accurate disease spread modeling is crucial for identifying the severity of outbreaks and planning effective mitigation efforts. To be reliable when applied to new outbreaks, model calibration techniques must be robust. However, current methods frequently forgo calibration verification (a stand-alone process evaluating the calibration procedure) and instead use overall model validation (a process comparing calibrated model results to data) to check calibration processes, which may conceal errors in calibration. In this work, we develop a stochastic agent-based disease spread model to act as a testing environment as we test two calibration methods using simulation-based calibration, which is a synthetic data calibration verification method. The first calibration method is a Bayesian inference approach using an empirically-constructed likelihood and Markov chain Monte Carlo (MCMC) sampling, while the second method is a likelihood-free approach using approximate Bayesian computation (ABC). Simulation-based calibration suggests that there are challenges with the empirical likelihood calculation used in the first calibration method in this context. These issues are alleviated in the ABC approach. Despite these challenges, we note that the first calibration method performs well in a synthetic data model validation test similar to those common in disease spread modeling literature. We conclude that stand-alone calibration verification using synthetic data may benefit epidemiological researchers in identifying model calibration challenges that may be difficult to identify with other commonly used model validation techniques.

60 APPLIED LIFE SCIENCES↗

Data-Informed Synthetic Networks of Water Distribution Systems for Resilience Analysis in Puerto Rico

The increasing potential of infrastructure disruptions calls for high-quality infrastructure models to be used in resilience analysis and decision making. Unfortunately, many utilities and communities do not have access to accurate and detailed models due to a lack of data and resources. Furthermore, security restrictions on sharing infrastructure models present roadblocks to research, analysis, and decision making. Recent advances in the development of synthetic water distribution models provide a potential solution to this problem. There is an opportunity to improve these methods by leveraging incomplete pipe datasets to aid synthetic network generation. To address this gap, we developed a methodology for synthetic network generation that incorporates partial pipe data using a modification of the minimum cost flow algorithm for network generation and pipe sizing. This methodology demonstrates how partial pipe data can be leveraged to improve site-specific synthetic network generation. For the study area of Mayagüez, Puerto Rico, a synthetic model generated using 50% of real pipe data matches the pressure of the validation system with an average error of 23.5 m of head, which improves upon the average error of 31.6 m of head produced by a synthetic model generated using no data of the real pipes. Additionally, synthetic networks are shown to replicate the pressure response under a disruption scenario of the validation network, suggesting potential use in resilience analysis.

resilience analysis↗

Constrained GAN-Generated X-Ray CT Data For Self-Supervised And Foundation-Model Segmentation Of Concrete Microstructures

Three-dimensional characterization of materials using X-ray computed tomography (XCT) is challenging due to the complexity of internal structures, noise, and variations in resolution. Traditional computer vision models often struggle to accurately segment these images, particularly in domain-specific applications like materials science. While supervised deep learning approaches have been developed to address the limitations of conventional algorithms, they typically require large amounts of labeled training data and often fail to generalize across different datasets. Self-supervised, few-and zero-shot learning methods have gained prominence in natural image processing and segmentation tasks, but their application to scientific imaging remains limited due to the unique structural complexity, noise, and textural artifacts present in materials science data. In this work, we investigate how domain adaptation, leveraging physics-based and GAN-generated synthetic data, impacts segmentation performance. We introduce a modified Contrastive Unpaired Translation (CUT) model designed to generate realistic labeled data, which can be used for training, pre-training, and fine-tuning segmentation models for real XCT microstructure data. We evaluate the performance of two segmentation approaches: a self-supervised network (SSL-ALPNet) and a foundation model (Segment Anything Model), assessing their improvements when pre-trained and/or fine-tuned on the synthesized data. Our results demonstrate that leveraging synthetic data significantly enhances segmentation performance, particularly in challenging materials science applications.

Ziabari, Amir [ORNL] (ORCID:000000034776457X)↗

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally corrected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a surrogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [Illinois U., Chicago]↗

Elucidating and predicting the dynamic evolution of water and land systems due to natural and energy-related forcings

Focal Area(s): 3. Insight gleaned from complex data (both observed and simulated) using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge-guided AI; & 1. Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). Science Challenge: Interactions between water, land, and energy systems are complex and occur on a variety of scales, ranging from local to basinal to regional. Accurately predicting the behavior of ground water and surface water systems for 5-10 years and beyond requires an understanding of the current system and the ability to model both the natural system at scale and human-induced forcings related to energy and other activities. Artificial intelligence and machine learning (AI/ML) combined with modern compilation and integration efforts for U.S. groundwater and surface water systems present potential solutions to bolstering detailed physics-based models of these systems. Big data tied with ML and physics-based modeling can drive breakthroughs in understanding the earth system, but research is often impeded by data access (e.g., privacy issues), quality, formats, gaps, multi-source, multi-scale, integration, and spatiotemporal challenges. Effective integration of real data and simulated (synthetic) data that fill gaps is critical. Overcoming these complex data and model integration challenges will enable a transformational approach to acquiring enhanced understanding of environmental systems.

54 ENVIRONMENTAL SCIENCES↗

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally cor- rected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a sur- rogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [U. Illinois, Chicago]↗

Encoding nonlinear and unsteady aerodynamics of limit cycle oscillations using nonlinear sparse Bayesian learning

This article investigates the applicability of a recently proposed, nonlinear sparse Bayesian learning (NSBL) algorithm to identify and estimate the complex aerodynamics of limit cycle oscillations. NSBL provides a semi-analytical framework for determining the data-optimal sparse model nested within a (potentially) over-parameterized model. This is particularly relevant to nonlinear dynamical systems where modelling approaches involve the use of physics-based and data-driven components. In such cases, the data-driven components, where analytical descriptions of the physical processes are not readily available, are often prone to overfitting, meaning that the empirical aspects of these models will often involve the calibration of an unnecessarily large number of parameters. While an overparameterized model may fit the observed data well, such models may be inadequate for making predictions in regimes that are different from those wherein the data were recorded. In view of this, it is desirable to not only calibrate the model parameters, but also identify the optimal compromise between data fit and model complexity. In this article, we exhibit the optimal model discovery for an aeroelastic system wherein the structural dynamics are well-known and described by a differential equation model, coupled with a semi-empirical aerodynamic model for laminar separation flutter, resulting in low-amplitude limit cycle oscillations (LCO). To illustrate the performance of the algorithm, in this article, we use synthetic data and demonstrate the ability of the algorithm to correctly rediscover the optimal model and model parameters, given a known data-generating model. The synthetic data are generated from a forward simulation of a known differential equation model with parameters selected so as to mimic the dynamics observed in wind-tunnel experiments. Subsequently, we demonstrate the performance of the algorithm for model selection using noisy LCO data from wind tunnel experiments. As there is no ground truth available for the experimental data case, we provide a comparison between NSBL and Bayesian model selection to validate the results, and demonstrate the use of NSBL as an efficient alternative to traditional methods.

97 MATHEMATICS AND COMPUTING↗

Detecting Large Explosions With Machine Learning Models Trained on Synthetic Infrasound Data

Explosions produce low-frequency acoustic (infrasound) waves capable of propagating globally, but the spatio-temporal variability of the atmosphere makes detecting events difficult. Machine learning (ML) is well-suited to identify the subtle and nonlinear patterns in explosion infrasound signals, but a previous lack of ground-truth data inhibited training of generalized models. We introduce a physics-based method that propagates infrasound sources through realistic atmospheres to create 28,000 synthetic events, which are used to train ML classifiers. A simple artificial neural network and modern temporal convolutional network discriminate synthetic events from background noise with >90% accuracy and, more importantly, successfully identify the majority of real-world explosion signals recorded during the Humming Road Runner experiment. ML models trained entirely on physics-based synthetics advance explosion detection capabilities and make ML more viable to related fields lacking training data.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Practical guide to understanding goodness-of-fit metrics used in chemical state modeling of x-ray photoelectron spectroscopy data by synthetic line shapes using nylon as an example

Chemical state analysis of a sample surface through fitting bell-shaped curves to x-ray photoelectron spectroscopic polymer data is reviewed using nylon to introduce and discuss aspects of data analysis. Different strategies for modeling chemistry in nylon spectra are presented and in so doing, a case is made to include in published science the design logic and implementation in terms of line shapes and optimization parameter constraints between components in a peak model. Imperfections in line shape relative to the true shape for photoemission lines, when compensated for using constraints to optimization parameters, are shown to provide chemical state information about a sample that justify, for peak models constructed with these limitations, metrics for goodness-of-fit different from those expected for pulse-counted data.

Materials Science↗

Seismic Spatial Gradients and Machine Learning-Based Classifiers for Explosion Monitoring (LDRD 218327)

This final report summarizes the work completed under the Laboratory Directed Research and Development (LDRD) project “Seismic Spatial Gradients as a Machine Learning-Based Classifier for Explosion Monitoring.” The overarching goal of the project was to explore the efficacy of using machine learning-based classification algorithms where the input data are the spatial gradient of the seismic wavefield collected at a single point on the Earth’s surface. The methods that I describe here are in direct contrast to conventional methods of seismic discrimination which typically rely on a spatially extended network of instruments and physics-based wavefield attributes such as, for example, the ratio between $\textit{P}$ and $\textit{S}$ waves. Rather, we use the spatial gradient of the seismic wavefield observed at a single point on the Earth’s surface and data processing approaches inspired by the machine learning community. We tested two algorithms, a neural network and a modified version of principal component analysis termed Spectrally Filtered Principal Component Analysis (SFPCA). To test these algorithms, we first conducted a series of numerical tests using synthetic data and then conducted a small-scale controlled field experiment. The tests using synthetic data showed that both algorithms had high success rates on gradiometric data, even when simulated noise was added to the signal. Furthermore, we found that using seismic spatial gradients increased the performance of our discrimination algorithms when compared to using just the traditional translational motion seismic data. The tests with field data also showed a high degree of discriminative success.

58 GEOSCIENCES↗

Low-pass spectral analysis of time-resolved serial femtosecond crystallography data

Low-pass spectral analysis (LPSA) is a recently developed dynamics retrieval algorithm showing excellent retrieval properties when applied to model data affected by extreme incompleteness and stochastic weighting. In this work, we apply LPSA to an experimental time-resolved serial femtosecond crystallography (TR-SFX) dataset from the membrane protein bacteriorhodopsin (bR) and analyze its parametric sensitivity. While most dynamical modes are contaminated by nonphysical high-frequency features, we identify two dominant modes, which are little affected by spurious frequencies. The dynamics retrieved using these modes shows an isomerization signal compatible with previous findings. We employ synthetic data with increasing timing uncertainty, increasing incompleteness level, pixel-dependent incompleteness, and photon counting errors to investigate the root cause of the high-frequency contamination of our TR-SFX modes. By testing a range of methods, we show that timing errors comparable to the dynamical periods to be retrieved produce a smearing of dynamical features, hampering dynamics retrieval, but with no introduction of spurious components in the solution, when convergence criteria are met. Using model data, we are able to attribute the high-frequency contamination of low-order dynamical modes to the high levels of noise present in the data. Finally, we propose a method to handle missing observations that produces a substantial dynamics retrieval improvement from synthetic data with a significant static component. Reprocessing of the bR TR-SFX data using the improved method yields dynamical movies with strong isomerization signals compatible with previous findings.

59 BASIC BIOLOGICAL SCIENCES↗

Maven: a multimodal foundation model for supernova science

Abstract A common setting in astronomy is the availability of a small number of high-quality observations, and larger amounts of either lower-quality observations or synthetic data from simplified models. Time-domain astrophysics is a canonical example of this imbalance, with the number of supernovae observed photometrically outpacing the number observed spectroscopically by multiple orders of magnitude. At the same time, no data-driven models exist to understand these photometric and spectroscopic observables in a common context. Contrastive learning objectives, which have grown in popularity for aligning distinct data modalities in a shared embedding space, provide a potential solution to extract information from these modalities. We present Maven, the first foundation model for supernova science. To construct Maven, we first pre-train our model to align photometry and spectroscopy from 0.5 M synthetic supernovae using a contrastive objective. We then fine-tune the model on 4702 observed supernovae from the Zwicky transient facility. Maven reaches state-of-the-art performance on both classification and redshift estimation, despite the embeddings not being explicitly optimized for these tasks. Through ablation studies, we show that pre-training with synthetic data improves overall performance. In the upcoming era of the Vera C. Rubin observatory, Maven will serve as a valuable tool for leveraging large, unlabeled and multimodal time-domain datasets.

Zhang, Gemma (ORCID:0000000280198082)↗

Suppressing simulation bias in multi-modal data using transfer learning

Abstract Many problems in science and engineering require making predictions based on few observations. To build a robust predictive model, these sparse data may need to be augmented with simulated data, especially when the design space is multi-dimensional. Simulations, however, often suffer from an inherent bias. Estimation of this bias may be poorly constrained not only because of data sparsity, but also because traditional predictive models fit only one type of observed outputs, such as scalars or images, instead of all available output data modalities, which might have been acquired and simulated at great cost. To break this limitation and open up the path for multi-modal calibration, we propose to combine a novel, transfer learning technique for suppressing the bias with recent developments in deep learning, which allow building predictive models with multi-modal outputs. First, we train an initial neural network model on simulated data to learn important correlations between different output modalities and between simulation inputs and outputs. Then, the model is partially retrained, or transfer learned, to fit the experiments; a method that has never been implemented in this type of architecture. Using fewer than 10 inertial confinement fusion experiments for training, transfer learning systematically improves the simulation predictions while a simple output calibration, which we design as a baseline, makes the predictions worse. We also offer extensive cross-validation with real and carefully designed synthetic data. The method described in this paper can be applied to a wide range of problems that require transferring knowledge from simulations to the domain of experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Deep Learning At Depth: Estimating subsurface parameters from geophysical monitoring data

Geophysical imaging techniques are a non-invasive way to image the subsurface and understand both subsurface solid (rock/soil) and fluid property distributions and their evolution in time. Inversions of the geophysical data, such as Electrical Resistance Tomography (ERT) data, are solved to estimate the subsurface property distributions, such as conductivity, and many inversion techniques smooth out sharp gradients in rock or fluid property distributions. Sharp gradients in subsurface properties tend to be present in situations with complex subsurface structures, which are common in many subsurface applications. We have successfully demonstrated that it is possible to inform, or constrain, inversions with neural networks trained on synthetic data with complex subsurface structures. Initial results suggest this process may be optimizable to yield property distributions that better represent the true property distributions than the same inversion process without the neural network constraint. Future work would optimize the neural network performance for this application and then apply the synthetic-data trained neural network to real data to understand the utility and performance of this technique for real data sets.

47 OTHER INSTRUMENTATION↗

Application of machine learning techniques for fast MeV x-ray spectra unfolding from filter stack spectrometer data

Recovery of MeV x-ray spectra from detector signals is difficult because the response matrix inversion is ill-conditioned and current methods are too slow for high-repetition-rate experiments. In this work, we make use of neural networks to unfold MeV x-ray spectra from measurements obtained with a filter stack spectrometer at rates of near 40 Hz. The neural network was trained on synthetic data and tested on both synthetic and experimental data, the latter obtained in two separate experiments performed at the Omega EP laser facility. We show here that this unfolding method has good performance on synthetic data and that it is a promising option for experimental data of up to 40 MeV. The accuracy on experimental data is verified by using a simple forward model to compare against measured values.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Bayesian Optimization of Non-Invariant Systems with Constraints Developed for Application to the ECR Ion Source VENUS

In this work, we consider the optimization of non-invariant systems with both safety and control constraints. We present a new approach based on Bayesian optimization for the dynamic, safe and controlled optimization of such systems. Although there are other possible use cases, we focus on the application to the electron cyclotron resonance ion source VENUS. From experimental data, we have observed that VENUS behaves to first order as a non-invariant dynamic system with moving areas of instability. Our novel approach aims at providing a tool that can maintain system optimization in a safe way. This is accomplished by making sure the objective function, the beam current in the case of VENUS, does not fall under an operational minimum, while simultaneously requiring the optimization to avoid areas where VENUS is unstable. We compare the result of our approach on synthetic data modeled to mimic the behavior of VENUS with two methods from the literature, a standard Bayesian optimizer and a safe Bayesian optimizer, both adapted to deal with dynamic systems. A cross Student T-test is conducted to show the significance of the improvement given by the new method we introduce here, regarding the two preexisting methods we compared to. The results of the tests conducted on synthetic data show that the proposed method succeeds at maintaining the system optimized and obeys the predefined constraints better than the literature methods explored.

Bayesian optimization↗

Rapid failure mode classification and quantification in batteries: A deep learning modeling framework

Unique, rapid identification and quantification of the dominant aging modes in lithium-ion batteries (LiBs) with early and non-specialized test data is a significant scientific challenge. Leveraging synthetic-data, deep-learning (DL) techniques have great potential to enable fast and robust classification and quantification of battery aging modes that produce different patterns of cell aging. This study, for the first time, presents a synthetic–data-based DL modeling framework for rapid and automatic classification and quantification of battery-aging modes and resultant aging with experimental validation. Availing synthetic dQ.dV -1 curves for ~26000 initial conditions and aging modes, the framework classified the dominant aging modes, for cells undergoing fast charge, in fewer than 100 cycles. Upon classification, the framework quantified the evolution of the aging modes, which were often nonuniform with cycling, for 22 gr/NMC532 pouch cells tested up to 600 cycles at different charging rates (1C–9C).

25 ENERGY STORAGE↗

Estimating uncertainty: A Bayesian approach to modelling photosynthesis in C3 leaves

The Farquhar-von Caemmerer-Berry (FvCB) model is extensively used to model pho-tosynthesis from gas exchange measurements. Since its publication, many methods have been developed to measure, or more accurately estimate, parameters of this model. Here, we have created a tool that uses Bayesian statistics to fit photosyn-thetic parameters using concurrent gas exchange and chlorophyll fluorescence mea-surements whilst evaluating the reliability of the parameter estimation. We have tested this tool on synthetic data and experimental data from rice leaves. Our results indicate that reliable parameter estimation can be achieved whilst only keeping one parameter, Km, that is, Michaelis constant for CO2 by Rubisco, prefixed. Additionally, we show that including detailed low CO2 measurements at low light levels increases reliability and suggests this as a new standard measurement protocol. By providing an estimated distribution of parameter values, the tool can be used to evaluate the quality of data from gas exchange and chlorophyll fluorescence measurement proto-cols. Compared to earlier model fitting methods, the use of a Bayesian statistics-based tool minimizes human interaction during fitting, reducing the subjectivity which is essential to most existing tools. A user friendly, interactive Bayesian tool script is provided.

Bayesian statistics, leaf photosynthesis, mesophyl↗