Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Hybrid Flush and Synthetic Air Data Filter for Entry Vehicle Atmospheric State Estimation

A hybrid flush/synthetic air data sensing filter utilizing Kalman-Schmidt and Rach-Tung-Striebel smoothers is developed to obtain entry vehicle atmosphere estimates. The filter/smoother blends information from pressure sensors distributed on the heatshield with measurements of the vehicle aerodynamic forces and moments computed from mass properties and inertial measurement unit data, and prior estimates of the atmosphere. The filter produces estimates of the atmospheric conditions along the entry trajectory, and systematic error estimates to reconcile differences between the pressure and aerodynamic data sources. The filter is applied to data acquired during the Mars Science Laboratory and Mars 2020 entry, descent, and landing at Gale crater and at Jezero crater, respectively. The results show that the hybrid filter produces estimates of the freestream flight condition with lower uncertainty than either the flush or synthetic air data algorithms. The filter accomplishes this result by incorporating additional data and computing estimates of systematic error parameters in the pressure data and the aerodynamic model to further reduce the uncertainties.

Christopher D. Karlgaard↗

Calibration verification for stochastic agent-based disease spread models

Accurate disease spread modeling is crucial for identifying the severity of outbreaks and planning effective mitigation efforts. To be reliable when applied to new outbreaks, model calibration techniques must be robust. However, current methods frequently forgo calibration verification (a stand-alone process evaluating the calibration procedure) and instead use overall model validation (a process comparing calibrated model results to data) to check calibration processes, which may conceal errors in calibration. In this work, we develop a stochastic agent-based disease spread model to act as a testing environment as we test two calibration methods using simulation-based calibration, which is a synthetic data calibration verification method. The first calibration method is a Bayesian inference approach using an empirically-constructed likelihood and Markov chain Monte Carlo (MCMC) sampling, while the second method is a likelihood-free approach using approximate Bayesian computation (ABC). Simulation-based calibration suggests that there are challenges with the empirical likelihood calculation used in the first calibration method in this context. These issues are alleviated in the ABC approach. Despite these challenges, we note that the first calibration method performs well in a synthetic data model validation test similar to those common in disease spread modeling literature. We conclude that stand-alone calibration verification using synthetic data may benefit epidemiological researchers in identifying model calibration challenges that may be difficult to identify with other commonly used model validation techniques.

60 APPLIED LIFE SCIENCES↗

Fan Beam Emission Tomography for Estimating Scalar Properties in Laminar Flames

A new method of estimating temperatures and gas species concentrations (CO2 and H2O) in a laminar flame is reported. The path-integrated, spectral radiation intensities emitted from a laminar flame at multiple wavelengths and view angles are calculated using a narrow band radiation model. Synthetic data, in the form of radial profiles of temperature and gas concentrations, are used in these calculations. The calculations mimic measurements that would theoretically be obtained using a mid-infrared spectrometer with a scanner. The path integrated spectral radiation intensities are deconvoluted using a maximum likelihood estimation method in conjunction with an iterative scheme. The deconvolution algorithm accounts for the self-absorption of radiation by the intervening gases, and provides the local temperature and gas species concentrations. The deconvoluted temperatures and gas concentrations are compared with the synthetic data used for calculating the spectral radiation intensities. The deconvoluted temperatures and gas species concentrations are within 0.5 % of the synthetic data. The deconvolution algorithm is expected to provide combustion researchers with an easy method of obtaining the radial profiles of major gas species concentrations and temperatures in laminar flames non-intrusively using a mid-infrared spectrometer with a scanner.

Lim, Jongmook↗

The two-way time synchronization system via a satellite voice channel

A newly developed two-way time synchronization system is described in this paper. The system uses one voice channel at a SCPC satellite digital communication earth station, whose bandwidth is only 45 kHz, thus saving satellite resources greatly. The system is composed of one master station and one or several, up to sixty-two, secondary stations. The master and secondary stations are equipped with the same equipment, including a set of timing equipment, a synthetic data terminal for time synchronizing, and a interface unit between the data terminal and the satellite earth station. The synthetic data terminal for time synchronization also has an IRIG-B code generator and a translator. The data terminal of master station is the key part of whole system. The system synchronization process is full automatic, which is controlled by the master station. Employing an autoscanning technique and conversational mode, the system accomplishes the following tasks: linking up liaison with each secondary station in turn, establishing a coarse time synchronization, calibrating date (years, months, days) and time of day (hours, minutes, seconds), precisely measuring the time difference between local station and the opposite station, exchanging measurement data, statistically processing the data, rejecting error terms, printing the data, calculating the clock difference and correcting the phase, thus realizing real-time synchronization from one point to multiple points. We also designed an adaptive phase circuit to eliminate the phase ambiguity of the PSK demodulator. The experiments have shown that the time synchronization accuracy is better than 2 mu S. The system has been put into regular operation.

Heng-Qiu, Zheng↗

Data-Informed Synthetic Networks of Water Distribution Systems for Resilience Analysis in Puerto Rico

The increasing potential of infrastructure disruptions calls for high-quality infrastructure models to be used in resilience analysis and decision making. Unfortunately, many utilities and communities do not have access to accurate and detailed models due to a lack of data and resources. Furthermore, security restrictions on sharing infrastructure models present roadblocks to research, analysis, and decision making. Recent advances in the development of synthetic water distribution models provide a potential solution to this problem. There is an opportunity to improve these methods by leveraging incomplete pipe datasets to aid synthetic network generation. To address this gap, we developed a methodology for synthetic network generation that incorporates partial pipe data using a modification of the minimum cost flow algorithm for network generation and pipe sizing. This methodology demonstrates how partial pipe data can be leveraged to improve site-specific synthetic network generation. For the study area of Mayagüez, Puerto Rico, a synthetic model generated using 50% of real pipe data matches the pressure of the validation system with an average error of 23.5 m of head, which improves upon the average error of 31.6 m of head produced by a synthetic model generated using no data of the real pipes. Additionally, synthetic networks are shown to replicate the pressure response under a disruption scenario of the validation network, suggesting potential use in resilience analysis.

resilience analysis↗

Synthetic Hyperspectral Data for Global Water Quality Algorithm Development

Eutrophication and increasing prevalence of potentially toxic algal blooms (cyanoHABs) among global inland water bodies have become a major ecological concern and require direct attention. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local aquatic processes, to spatially resolved global products. Planned aquatic biogeochemistry remote sensing data products from hyperspectral imagers such as NASA’s Surface Biology and Geology (SBG) mission and relevant aquatic sensor sensitivity precursor airborne imaging spectrometer data provide unprecedented radiometric resolution and sensor sensitivity for characterizing complex aquatic ecosystems. However, scarcity of high-quality freshwater in-situ optical data hinders our capability to develop and validate robust retrieval algorithms. A state-of-the-art synthetic dataset of paired top-of-atmosphere, bottom-of-atmosphere, and optical and biogeophysical data was developed through radiative transfer modeling to simulate natural freshwater ecosystems. A synthetic or precursor dataset for SBG is being used to train robust machine learning models to derive water quality products pertinent to SBG mission objectives. The dataset is also used to show the potential of performing vigorous aquatic sensitivity studies and explored pathways for how best to optimize hyperspectral data for machine learning development. A processing pipeline and resultant global synthetic/precursor dataset for inland waters is presented to establish the innovation for water quality studies of inland waters globally. Optical Society of America Imaging and Applied Optics Congress, Hyperspectral Imaging and Sounding of the Environment (OSA HISE) Meeting, 19-23 July 2021, Virtual Meeting, https://www.osa.org/enus/meetings/osa_meetings/optical_sensors_and_sensing_congress/program/hyperspectral_imaging_and_sounding_of_the_environm/

Synthetic↗

Constrained GAN-Generated X-Ray CT Data For Self-Supervised And Foundation-Model Segmentation Of Concrete Microstructures

Three-dimensional characterization of materials using X-ray computed tomography (XCT) is challenging due to the complexity of internal structures, noise, and variations in resolution. Traditional computer vision models often struggle to accurately segment these images, particularly in domain-specific applications like materials science. While supervised deep learning approaches have been developed to address the limitations of conventional algorithms, they typically require large amounts of labeled training data and often fail to generalize across different datasets. Self-supervised, few-and zero-shot learning methods have gained prominence in natural image processing and segmentation tasks, but their application to scientific imaging remains limited due to the unique structural complexity, noise, and textural artifacts present in materials science data. In this work, we investigate how domain adaptation, leveraging physics-based and GAN-generated synthetic data, impacts segmentation performance. We introduce a modified Contrastive Unpaired Translation (CUT) model designed to generate realistic labeled data, which can be used for training, pre-training, and fine-tuning segmentation models for real XCT microstructure data. We evaluate the performance of two segmentation approaches: a self-supervised network (SSL-ALPNet) and a foundation model (Segment Anything Model), assessing their improvements when pre-trained and/or fine-tuned on the synthesized data. Our results demonstrate that leveraging synthetic data significantly enhances segmentation performance, particularly in challenging materials science applications.

Ziabari, Amir [ORNL] (ORCID:000000034776457X)↗

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally corrected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a surrogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [Illinois U., Chicago]↗

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally cor- rected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a sur- rogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [U. Illinois, Chicago]↗

Inversion of limb radiance measurements - An operational algorithm

The limb radiance inversion radiometer (LRIR) and limb infrared monitor of the stratosphere (LIMS) experiments aboard the Nimbus 6 and 7 spacecraft have made observations of infrared emission by CO2, O3, H2O, HNO3, and NO2 at the earth's limb. This paper describes a method by which such measurements can be inverted to give vertical distributions of temperature and mixing ratios as functions of pressure. The simple and efficient approach was successfully applied to the LRIR data and subsequently in the initial assessment of the LIMS data. Inversion of synthetic data indicates the size of the errors to be expected as a result of the assumptions and instrumental errors. Retrievals of measured LIMS radiances are shown as examples and compared to in situ observations. The differences are comparable to those obtained with the more complex retrieval scheme used to process the LIMS archival products. Some problems are noted.

Bailey, P. L.↗

Encoding nonlinear and unsteady aerodynamics of limit cycle oscillations using nonlinear sparse Bayesian learning

This article investigates the applicability of a recently proposed, nonlinear sparse Bayesian learning (NSBL) algorithm to identify and estimate the complex aerodynamics of limit cycle oscillations. NSBL provides a semi-analytical framework for determining the data-optimal sparse model nested within a (potentially) over-parameterized model. This is particularly relevant to nonlinear dynamical systems where modelling approaches involve the use of physics-based and data-driven components. In such cases, the data-driven components, where analytical descriptions of the physical processes are not readily available, are often prone to overfitting, meaning that the empirical aspects of these models will often involve the calibration of an unnecessarily large number of parameters. While an overparameterized model may fit the observed data well, such models may be inadequate for making predictions in regimes that are different from those wherein the data were recorded. In view of this, it is desirable to not only calibrate the model parameters, but also identify the optimal compromise between data fit and model complexity. In this article, we exhibit the optimal model discovery for an aeroelastic system wherein the structural dynamics are well-known and described by a differential equation model, coupled with a semi-empirical aerodynamic model for laminar separation flutter, resulting in low-amplitude limit cycle oscillations (LCO). To illustrate the performance of the algorithm, in this article, we use synthetic data and demonstrate the ability of the algorithm to correctly rediscover the optimal model and model parameters, given a known data-generating model. The synthetic data are generated from a forward simulation of a known differential equation model with parameters selected so as to mimic the dynamics observed in wind-tunnel experiments. Subsequently, we demonstrate the performance of the algorithm for model selection using noisy LCO data from wind tunnel experiments. As there is no ground truth available for the experimental data case, we provide a comparison between NSBL and Bayesian model selection to validate the results, and demonstrate the use of NSBL as an efficient alternative to traditional methods.

97 MATHEMATICS AND COMPUTING↗

Helioseismic Measurements of Convective Power in Solar Cycle 24

Constraining the parameters under which convection in the solar interior operates has important implications for describing how energy is transported by plasma motions, and various physical models have been employed to provide some theoretical estimates on the expected power spectrum. Past attempts to measure the convective power distributed among large spatial scales have, however, found differing and incompatible values. Here, we present measurements of the convective power spectrum in the upper convection zone for Carrington rotations in Solar Cycle 24 obtained from the helioseismic signal corresponding to East-West flows. We also perform calibration on synthetic data using the global acoustic GALE code to make assessments of the flow velocities at various length scales without the need for performing inversions. This allows us to derive the convective power spectrum from the flow maps and to compare with the results from the raw travel times. These results are compared against predictions of convective power in global models of convection produced by the EULAG code. Finally, we show how the steps in our analysis procedure (for example data segmentation, filtering, etc.) affect our estimates of the convective power by comparing with the synthetic data from the GALE code.

Heliophysics↗

Detecting Large Explosions With Machine Learning Models Trained on Synthetic Infrasound Data

Explosions produce low-frequency acoustic (infrasound) waves capable of propagating globally, but the spatio-temporal variability of the atmosphere makes detecting events difficult. Machine learning (ML) is well-suited to identify the subtle and nonlinear patterns in explosion infrasound signals, but a previous lack of ground-truth data inhibited training of generalized models. We introduce a physics-based method that propagates infrasound sources through realistic atmospheres to create 28,000 synthetic events, which are used to train ML classifiers. A simple artificial neural network and modern temporal convolutional network discriminate synthetic events from background noise with >90% accuracy and, more importantly, successfully identify the majority of real-world explosion signals recorded during the Humming Road Runner experiment. ML models trained entirely on physics-based synthetics advance explosion detection capabilities and make ML more viable to related fields lacking training data.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Augmented Reality Data Generation for Training Deep Learning Neural Network

One of the major challenges in deep learning is retrieving sufficiently large labeled training datasets, which can become expensive and time consuming to collect. A unique approach to training segmentation is to use Deep Neural Network (DNN) models with a minimal amount of initial labeled training samples. The procedure involves creating synthetic data and using image registration to calculate affine transformations to apply to the synthetic data. The method takes a small dataset and generates a highquality augmented reality synthetic dataset with strong variance while maintaining consistency with real cases. Results illustrate segmentation improvements in various target features and increased average target confidence.

Torres, Gil↗

Practical guide to understanding goodness-of-fit metrics used in chemical state modeling of x-ray photoelectron spectroscopy data by synthetic line shapes using nylon as an example

Chemical state analysis of a sample surface through fitting bell-shaped curves to x-ray photoelectron spectroscopic polymer data is reviewed using nylon to introduce and discuss aspects of data analysis. Different strategies for modeling chemistry in nylon spectra are presented and in so doing, a case is made to include in published science the design logic and implementation in terms of line shapes and optimization parameter constraints between components in a peak model. Imperfections in line shape relative to the true shape for photoemission lines, when compensated for using constraints to optimization parameters, are shown to provide chemical state information about a sample that justify, for peak models constructed with these limitations, metrics for goodness-of-fit different from those expected for pulse-counted data.

Materials Science↗

Estimating the Single-Trial Characteristics of Event-Related Responses: Evaluation of the MCERP Algorithm

Single-trial event-related responses collected during the course of an experiment are typically averaged before analysis resulting in a rather crude picture of event-related brain dynamics. It has been quite clear for some time that these responses exhibit trial-to-trial variability: however, the computational techniques necessary to deal with such responses in noisy conditions have not been available. To this end we have developed the multiple-component, event-related potential model (mcERP), which assumes that the each event-related response consists of a sum of multiple evoked components each described by a stereotypical waveshape. These waveshapes are allowed to vary in amplitude and onset latency from trial to trial, which allows us to capture, to first-order, the trial-dependent variations in event-related brain dynamics. We have constructed many sets of synthetic data designed to simulate intracortical recordings from a 15 channel, linear-array multielectrode implanted acutely in V1 of an awake-behaving macaque undergoing visual stimulation with a red light flash. This synthetic data was used to characterize the performance of the mcERP algorithm. First we quantified the degree to which such trial-to-trial variability aids in the identification of multiple components, and we demonstrate that amplitude variability is a more important factor in component separation than latency variability. Second, we quantified the behavior of the algorithm under two distinct signal-to-noise ratio (SNR) conditions: Gaussian noise independently present in each channel, and highly correlated (1/f distributed), far-field noise presented identically in each channel of the array. The mcERP algorithm was found to be robust to noise accurately identifying all component waveshapes and their associated single-trial characteristics down to SNR levels of -20dB for Gaussian noise and -7dB for 1/f far-field noise. Comparisons of the performance of this algorithm with factor analysis (FA) and independent component analysis (ICA) will be described by Knuth et al. (SFN abstracts, 2002). In addition, the advantages of application of mcERP to real data will be described by Shah et al, (these abstracts, 2002: SFN abstracts, 2002).

Knuth, K. H.↗

Low-pass spectral analysis of time-resolved serial femtosecond crystallography data

Low-pass spectral analysis (LPSA) is a recently developed dynamics retrieval algorithm showing excellent retrieval properties when applied to model data affected by extreme incompleteness and stochastic weighting. In this work, we apply LPSA to an experimental time-resolved serial femtosecond crystallography (TR-SFX) dataset from the membrane protein bacteriorhodopsin (bR) and analyze its parametric sensitivity. While most dynamical modes are contaminated by nonphysical high-frequency features, we identify two dominant modes, which are little affected by spurious frequencies. The dynamics retrieved using these modes shows an isomerization signal compatible with previous findings. We employ synthetic data with increasing timing uncertainty, increasing incompleteness level, pixel-dependent incompleteness, and photon counting errors to investigate the root cause of the high-frequency contamination of our TR-SFX modes. By testing a range of methods, we show that timing errors comparable to the dynamical periods to be retrieved produce a smearing of dynamical features, hampering dynamics retrieval, but with no introduction of spurious components in the solution, when convergence criteria are met. Using model data, we are able to attribute the high-frequency contamination of low-order dynamical modes to the high levels of noise present in the data. Finally, we propose a method to handle missing observations that produces a substantial dynamics retrieval improvement from synthetic data with a significant static component. Reprocessing of the bR TR-SFX data using the improved method yields dynamical movies with strong isomerization signals compatible with previous findings.

59 BASIC BIOLOGICAL SCIENCES↗

Maven: a multimodal foundation model for supernova science

Abstract A common setting in astronomy is the availability of a small number of high-quality observations, and larger amounts of either lower-quality observations or synthetic data from simplified models. Time-domain astrophysics is a canonical example of this imbalance, with the number of supernovae observed photometrically outpacing the number observed spectroscopically by multiple orders of magnitude. At the same time, no data-driven models exist to understand these photometric and spectroscopic observables in a common context. Contrastive learning objectives, which have grown in popularity for aligning distinct data modalities in a shared embedding space, provide a potential solution to extract information from these modalities. We present Maven, the first foundation model for supernova science. To construct Maven, we first pre-train our model to align photometry and spectroscopy from 0.5 M synthetic supernovae using a contrastive objective. We then fine-tune the model on 4702 observed supernovae from the Zwicky transient facility. Maven reaches state-of-the-art performance on both classification and redshift estimation, despite the embeddings not being explicitly optimized for these tasks. Through ablation studies, we show that pre-training with synthetic data improves overall performance. In the upcoming era of the Vera C. Rubin observatory, Maven will serve as a valuable tool for leveraging large, unlabeled and multimodal time-domain datasets.

Zhang, Gemma (ORCID:0000000280198082)↗