Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “synthetic data generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A Physics-Based Digital Twin for Wave Elevation and Seabed Moment Estimation of Offshore Monopiles: Preprint

In this work, we present a proof of concept of a physics-based digital twin for a monopile structure (with overhead inertia) subjected to wave loading. The digital twin is formulated using reduced-order models derived from first principles and combined with a Kalman filter for state estimation. The proposed framework estimates the monopile top motion, the wave elevation, and the section forces and moments along the pile using primarily acceleration measurements at the monopile top. Key innovations include the use of a hydrodynamic shape function to represent distributed wave loading in a compact and computationally efficient manner, and the introduction of a shaping filter to augment the state-space with wave kinematics. Synthetic measurement data are generated using OpenFAST and used as a reference to assess the performance of the digital twin. Results demonstrate that the wave elevation can be accurately reconstructed without direct sea-state measurements as long as the wave regime is inertia-dominated. Under the ideal tested conditions, the total hydrodynamic force and sea-bed bending moment are estimated with relative errors on the order of 1% and correlation coefficients exceeding 96%. Future work will evaluate the estimator's performance under operational uncertainties and more complex loading conditions.

17 WIND ENERGY↗

Developing Data-Driven Synthetic Infrastructure Models for Resilience Analysis

Research on infrastructure resilience has produced promising methods to simulate and optimize complex networks to improve performance. However, restrictions on sharing infrastructure models and the steep cost of developing and maintaining infrastructure models presents a roadblock to adoption. To overcome this limitation, this research focuses on methods to create data-driven infrastructure models that will help improve infrastructure resilience and security. The analysis couples incomplete utility data, geospatial data, machine learning, and synthetic network generation methods to rapidly develop and update infrastructure models. The methods are validated using realistic utility models and site-specific data, with a focus on Puerto Rico due to its unique infrastructure challenges and available data. This research highlights promising opportunities for the use of synthetic network generation and machine learning to create infrastructure models when very little data is available. Results demonstrate that hybrid methods, which combine sparse utility data with synthetic models, can enhance model accuracy, and machine learning can predict model attributes using training data from other models. However, the complexity of infrastructure systems means that even minor changes in network connectivity can significantly impact simulation results. Resilience analysis using synthetic infrastructure models shows that while some system behaviors are preserved, the magnitude of disruptions may not be accurately represented, indicating the need for more research and validation before using synthetic models for critical infrastructure investment decisions. The framework outlined in this report represents a significant advance to infrastructure model development and could be applied to additional domains and sites. Future research will continue to streamline and validate methods to help reduce roadblocks to resilience analysis.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A platform to measure isentropes from proton-heated warm dense matter on short pulse laser facilities

We describe the development of an experimental platform that measures the release isentrope of materials heated isochorically to temperatures of a few electron volts, using short-pulse laser-produced protons to heat the sample and long-pulse laser-produced x rays to perform streaked x-ray radiography. The density profiles derived from the radiography data are integrated to generate pressure–density isentropes, independent of prior knowledge of the equation of state of the sample material. In order to understand the sensitivities of isentrope extraction from radiography data, we analyze synthetic radiographs generated by a radiation hydrodynamics code. Noise reduction and high spatial resolution are critical for isentrope reconstruction, as demonstrated by the analysis of a proof-of-principle shot day on the OMEGA-EP facility. In conclusion, the data demonstrate the feasibility of the platform for characterizing isentropes, and we discuss the necessary improvements to enhance precision in differentiating between equation-of-state models.

Equations of state↗

An Investigation Into Possible Systematic Effects on Neutron Star Radius Estimates using NICER-like Synthetic Data

Neutron star cores contain the densest matter in the observable universe. The state of this matter is of interest in numerous fields, but laboratory experiments cannot explore this matter. Although the composition of the matter would be of great interest, macroscopic observables such as the neutron star mass-radius relation depend primarily on the equation of state (EOS). As a result, precise and reliable radius measurements would be valuable in constraining the EOS. However, most attempts at radius measurements are susceptible to systematic errors, meaning the inferred radius can be significantly biased even though the fit to the data appears to be statistically good. Previous studies suggested that radii inferred using X-ray data provided by NASA’s NICER mission may be more immune from such systematic errors. This is in part because, compared with previous measurements that obtained averaged spectra and fluxes, NICER observes millisecond pulsars by timing the photons so precisely, it is possible to obtain the spectrum as a function of rotational phase and see variations such as heated regions on the star rotate into and out of view hundreds of times per second. This extra information seems promising to break degeneracies and mitigate systematic errors, but a more in-depth study is necessary. We report the first steps of that study, in which we generate NICER-like synthetic data and determine the quality of fit and bias in radius obtained when we fit the data using a model different from the model used to generate the data.

Isiah Holt↗

Aeroacoustic Computations of a Generic Low Boom Concept in Landing Configuration: Part 3 - Aerodynamic Validation and Noise Source Identification

Wind tunnel test data were used to validate the predicted aerodynamic behavior of a 15%-scale version of a generic, low-boom aircraft. The test was conducted in the NASA Langley 14- by 22-Foot Subsonic Tunnel to determine the low-speed aerodynamic characteristics of the model. Measured steady surface pressures and global forces were used to validate predicted aerodynamic results obtained from high-fidelity simulations of the model as installed in the tunnel. Very good agreement between predicted and measured aerodynamic trends was demonstrated, providing the impetus to proceed with companion airframe noise simulations that were conducted in a free-air setting. Computed near-field flow variables acquired on a permeable data surface were used to generate synthetic pressure records for two 800-element phased microphone arrays positioned overhead and to the side of the model. The array data were beamformed to generate noise source localization maps for the clean model and several landing configurations. Primary and secondary airframe sources were identified and their relative strengths were determined. Far-field integrated noise spectra for the full aircraft, as well as individual components, were obtained from the source maps via integration of tailored regions. The analysis showed that noise produced by the landing gear was the dominant contributor to the far-field acoustic signature of the model, followed by flap noise. The effects of permeable data surface end caps, spatial resolution, array orientation, angle of attack, component interaction, velocity scaling, and numerical precision on synthetic far-field spectra were also evaluated.

low boom↗

Validation of the DESI 2024 Lyα forest BAO analysis using synthetic datasets

The first year of data from the Dark Energy Spectroscopic Instrument (DESI) contains the largest set of Lyman-α (Lyα) forest spectra ever observed. This data, collected in the DESI Data Release 1 (DR1) sample, has been used to measure the Baryon Acoustic Oscillation (BAO) feature at redshift z = 2.33. In this work, we use a set of 150 synthetic realizations of DESI DR1 to validate the DESI 2024 Lyα forest BAO measurement presented in [1]. The synthetic data sets are based on Gaussian random fields using the log-normal approximation. We produce realistic synthetic DESI spectra that include all major contaminants affecting the Lyα forest. The synthetic data sets span a redshift range 1.8 < z < 3.8, and are analyzed using the same framework and pipeline used for the DESI 2024 Lyα forest BAO measurement. To measure BAO, we use both the Lyα auto-correlation and its cross-correlation with quasar positions. We use the mean of correlation functions from the set of DESI DR1 realizations to show that our model is able to recover unbiased measurements of the BAO position. We also fit each mock individually and study the population of BAO fits in order to validate BAO uncertainties and test our method for estimating the covariance matrix of the Lyα forest correlation functions. Finally, we discuss the implications of our results and identify the needs for the next generation of Lyα forest synthetic data sets, with the top priority being to simulate the effect of BAO broadening due to non-linear evolution.

79 ASTRONOMY AND ASTROPHYSICS↗

A novel methodology for gamma-ray spectra dataset procurement over varying standoff distances and source activities

The adoption of machine learning approaches for gamma-ray spectroscopy has received considerable attention in the literature. Many studies have investigated the deployment of various algorithm architectures to a specific task. However, little attention has been afforded to the development of the datasets leveraged to train the models. Such training datasets typically span a set of environmental or detector parameters to encompass a problem space of interest to a user. Variations in these measurement parameters will also induce fluctuations in the detector response, including expected pile-up and ground scatter effects. Fundamental to this work is the understanding that 1) the underlying spectral shape varies as the measurement parameters change and 2) the statistical uncertainties associated with two spectra impact their level of similarity. While previous studies attribute some arbitrary discretization to the measurement parameters for the generation of their synthetic training data, this work introduces a principled methodology for efficient spectral-based discretization of a problem space. A signal-to-noise ratio (SNR) respective spectral comparison measure and a Gaussian Process Regression (GPR) model are used to predict the spectral similarity across a range of measurement parameters. This innovative approach effectively showcased its capability by dividing a problem space, ranging from 5 cm to 100 cm standoff distances and 5 μCi–100 μCi of 137 Cs, into three unique combinations of measurement parameters. The findings from this work will aid in creating more robust datasets, which incorporate many possible measurement scenarios, reduce the number of required experimental test set measurements, and possibly enable experimental training data collection for gamma-ray spectroscopy.

data science↗

Interpolation of computed gamma-ray detector response functions

Gamma-ray spectra measured by traditional detectors contain features that result from a combination of the effects of detector materials/geometry, the incident gamma-ray energy, and the angle of entry. The features, such as the full-energy photopeak, Compton continuum, annihilation peak, and escape peaks, are governed by simple relationships depending on incident energy and have been known for a long time. Monte Carlo computer simulations of gamma rays interacting with a detector will show these features, and with a resolution function applied, the results should look similar to real measurements. The traditional approach to creating a detector response function requires many separate simulations of monoenergetic gamma rays striking the detector. This paper presents a new approach to developing computed detector response functions. The new approach involves a much smaller number of monoenergetic gamma-ray simulations and uses interpolation to quickly generate the responses of gamma rays that were not simulated. During the interpolation process, the underlying physics equations are used to accurately compute the response of a given energy gamma ray from the small set of simulations. Such work enables accelerated generation of synthetic radiation detector data.

Detector response↗

Explainable machine learning for incipient anomaly detection in compact molten salt heat exchanger with overlapping feature distributions

High-temperature molten salt-cooled reactors (MSCRs) are a promising next-generation nuclear technology option, offering efficient power conversion and inherent safety features. However, the reliability of these systems depends on the robust operation of heat exchangers (HXs), which are susceptible to failure due to temperature gradients and channel plugging caused by fluid freezing. Conventional monitoring methods, relying on inlet and outlet measurements, lack the spatial resolution needed to detect early-stage faults. We propose a novel design of a compact salt-to-salt matrix-type HX design consisting of interleaved arrays of parallel tubes, with integrated synthetic fiber optic distributed temperature sensing (DTS) to enable localized detection of incipient faults. To evaluate performance of this design, we generate high-fidelity synthetic data using heat transfer computational modeling to simulate channel plugging, and introduce sensor noise for realistic modeling of measurements. The dataset comprises of 97% normal operation and 3% anomaly cases, with each anomaly class representing 1% of the data. These early anomalies result in overlapping temperature profiles between normal and faulty channels, producing a non-separable dataset that challenges traditional classification techniques. We benchmark eight supervised machine learning (ML) models and demonstrate that XGBoost achieves the highest performance. To improve transparency, we develop an explainability framework combining Shapley values and partially ordered sets (POSETs) to quantify and structurally analyze feature importance. This approach identifies both dominant predictors and ambiguous feature relationships, enhancing trust and interpretability. Our results highlight the potential of combining DTS and explainable ML with intelligent feature selection to improve predictive maintenance and ensure operational resilience in advanced nuclear systems.

Prantikos, Konstantinos [Argonne National Laborato↗

Prediction of Hydrological Drought: What Can We Learn From Continental-Scale Offline Simulations?

Land surface model experiments are used to quantify, across the coterminous United States, the contributions (isolated and combined) of soil moisture and snowpack initialization to the skill of seasonal streamflow forecasts at multiple leads and for different start dates. Forecasted streamflows are compared to naturalized streamflow observations where available and to synthetic (model-generated) streamflow data elsewhere. We find that snow initialization has a major impact on skill in the mountainous western U.S. and in a portion of the northern Great Plains; a mid-winter (January 1) initialization of snow in these areas leads to significant skill in the spring melting season. Soil moisture initialization also contributes to skill, and although the maximum contributions are not as large as those seen for snow initialization, the soil moisture contributions extend across a much broader geographical area. Soil moisture initialization can contribute to skill at long leads (up to 5 or 6 months), particularly for forecasts issued during winter.

Koster, Randal↗

Comparison of Boeing 777 Landing Gear Noise Simulations with Flight Test Data

Acoustic phased microphone array measurements of aircraft flyover noise acquired during the 2005 Quiet Technology Demonstrator II test were used to assess the accuracy of high-fidelity, full-scale simulations of landing gear noise produced by a large civilian aircraft. The simulations, conducted with the lattice Boltzmann solver PowerFLOW®, used a highly accurate digital model of a Boeing 777-300ER aircraft with the nose and main landing gear components replicating the full-scale geometries. The simulations were performed for aircraft parameters that matched those recorded during the flyover test conditions. For benchmarking purposes, several aircraft configurations were simulated: a) nose landing gear deployed with main landing gear and wing high-lift devices stowed, b) nose and main landing gear deployed with wing high-lift devices stowed and c) nose and main landing gear with wing high-lift devices deployed. To facilitate direct comparison with measured data, the simulated data sets were used to generate synthetic pressure records at the same array microphone locations as those used during the flight test. Broadly self-consistent beamforming techniques and procedures were used to process the synthetic pressure records and the measured data. Integration of select regions of the beamform maps containing the nose or main landing gear yielded good agreement between predicted and measured integrated far-field spectra for forward directivity angles where airframe noise is more prominent.

airframe noise↗

PDV Inspection and Analysis Demonstration: 2024 PDV Workshop

This document walks a user through a demonstration of working with PDV digitizer data using python. This demonstration and included suggested exercises will be used at the 2024 PDV workshop hands-on session as an example and skill-development training session. The tutorial allows the user to generate synthetic but realistic PDV waveform data and visualize/inspect the results using spectrograms and waveform viewing tools.

97 MATHEMATICS AND COMPUTING↗

Aeroacoustic Computations of a Transonic Truss-Braced Wing Aircraft: Part 2 – Acoustic Signature and Noise Source Identification

High-fidelity, time-dependent simulations of a Boeing-designed, transonic, truss-braced-wing aircraft in cruise (clean) and landing configurations are leveraged to generate synthetic microphone-phased-array data for airframe noise prediction and assessment. These data sets are used to compute source localization (beamform) maps to determine the location and strength of primary and secondary airframe noise sources associated with this unique configuration. The synthetic phased-array implementation mimics the setup of a flight test. As this study is ongoing, preliminary integrated far-field spectra for the cruise configuration obtained at multiple spatial resolutions revealed significant tonal content that lacked convergence with increased resolution. The origin of several of these tones and their unusual convergence behavior was traced to the larger-than-normal trailing-edge thickness of the “as-tested” cruise model being simulated. Reducing the trailing-edge thickness to more realistic values eliminated most of the tones at low to moderate frequencies and improved spectrum convergence significantly. Applying lessons learned from the cruise simulations, several modifications to the geometry of the landing configuration were made and are described in this work. Results from permeable and solid Ffowcs-Williams and Hawkings surfaces at two different spatial resolutions (coarse and medium) are used to illustrate the major noise sources and determine convergence of the CLEAN integrated noise levels for the entire aircraft as well as major subcomponents. We demonstrate that the low-frequency content of the far-field spectrum is dominated by noise generated from the main landing gear, while the medium- and high-frequency content is dominated by the wing-leading-edge Krueger flaps. Since analysis of the acoustic maps for the landing configuration revealed several clusters of multiple sources along the wing leading edge, “high resolution” processing of the array data was used to distinguish more accurately the locations of sources.

Transonic Truss-Braced Wing↗

Improving Sim-to-Real Transfer in Vision-Based Robot Navigation Via Instance-Level GAN-Based Data Augmentation

Achieving robust vision-based robotic tasks requires large amounts of data, which are often difficult to obtain in real-world scenarios. Simulators and synthetic data offer a cost-effective alternative, but the visual gap between simulation and reality hinders the performance of models when deployed in real-world environments. In this paper, we present a data augmentation pipeline that integrates a foundation model (Segment Anything Model) with an unsupervised image-to-image translation model (CycleGAN) for instance-level domain transfer from simulation to reality. This pipeline enables the generation of realistic labeled data from synthetic images for training supervised machine learning models in vision-based navigation tasks. We evaluate our approach on real-world data for ego-vehicle pose estimation, a critical autonomous navigation task involving the prediction of cross-track position and heading angle relative to road center line markings. The results of our tests show that our GAN-based data augmentation pipeline significantly outperforms models trained solely on simulation data or on data processed with standard image augmentation methods for sim-to-real transfer, enhancing model robustness and generalizability in real-world scenarios. Our method provides a scalable and flexible data augmentation tool for leveraging large synthetic datasets to enhance vision-based robotic navigation tasks.

artificial intelligence↗

Seismicity-constrained fault detection and characterization with a multitask machine learning model

Geological fault detection and characterization are crucial for understanding subsurface dynamics across scales. While methods for fault delineation based on either seismicity location analysis or seismic image reflector discontinuity are well-established, a systematic approach that integrates both data types remains absent. We develop a novel machine learning model that unifies seismic reflector images and seismicity location information to automatically identify geological faults and characterize their geometrical properties. The model encodes a seismic image and a seismicity location image separately, and fuses the encoded features with a spatial-channel attention fusion module to improve the learning of important features in both inputs. We design an automated strategy to generate high-quality synthetic training data and labels. To improve the realism of the seismicity location image, we include random seismicity noise and missing seismicity location associated with some of the faults. We validate the model’s efficacy and accuracy using synthetic data examples and two field data examples. Moreover, we show that fine-tuning the trained model with a small, domain-specific dataset enhances its fidelity for field data applications. The results demonstrate that integrating seismicity location and seismic images into a unified framework allows the end-to-end neural network to achieve higher fidelity and accuracy in delineating subsurface faults and their geometrical properties compared with image-only fault detection methods. Our approach offers an adaptive data-driven tool for geological fault characterization and seismic hazard mitigation, bridging the gap between seismicity location and image-based fault detection methods.

58 GEOSCIENCES↗

Data volume reduction for imaging radar polarimetry

Two data reduction algorithms developed using the scattering and phase matrix approaches are described. In the scattering matrix approach, the scattering matrices of four consecutive along-track pixels are averaged and in the phase matrix approach, the phase matrices of four consecutive along-track pixels are averaged. The basic procedures necessary to generate a synthetic polarization image from original data sets are discussed. The two algorithms are evaluated in terms of data volume reduction and the number of errors introduced in the synthesized images. It is observed that the reduced data set produced by the scattering matrix algorithm is smaller than that generated by the phase matrix algorithm; however, greater errors are introduced into the data set by the scattering matrix algorithm than the phase algorithm. Flowcharts for the scattering and phase matrix approaches and for synthesis of uncompressible data are presented.

Dubois, Pascale C.↗

Electron-impact excitation data for W 2+ in support of tungsten spectroscopy and re-deposition measurements for magnetically-confined plasmas

Abstract To better understand plasma wall interactions involving tungsten, accurate atomic structure and electron-impact driven collisional processes for near-neutral ion stages of tungsten are required. Complementing existing work on neutral and singly ionised tungsten, atomic structure and collisional calculations for W 2+ electron-impact excitation have been completed. These excitation calculations are an important component of S/XB coefficients for near-neutral charge states, which may be used to spectroscopically infer re-deposition of tungsten at the plasma-solid boundary of fusion relevant devices. With W 2+ in particular having emission lines that can be observed at ultraviolet (UV) wavelengths, while higher charge states of tungsten are unlikely to have lines possible to observe outside of the vacuum UV range. The atomic structure was generated using the General-purpose Relativistic Atomic Structure Package (GRASP 0 ), implementing the Multi-configuration Dirac Fock approach. This structure was the basis for a subsequent Dirac R -matrix electron-impact excitation calculation to provide Maxwellian averaged rate coefficients. A synthetic spectrum was generated from this data using a collisional-radiative model to predict the strongest W III spectral lines and these lines were compared to emission from the Compact Toroidal Hybrid (CTH) plasma device. Several of the strongest W III lines are observed in CTH and agree well with the modelled line wavelengths and intensities, a table of these lines is provided that could be observed in other devices.

McCann, M. (ORCID:0000000215321240)↗