Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “synthetic data generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Generation of topographic terrain models utilizing synthetic aperture radar and surface level data

Topographical terrain models are generated by digitally delineating the boundary of the region under investigation from the data obtained from an airborne synthetic aperture radar image and surface elevation data concurrently acquired either from an airborne instrument or at ground level. A set of coregistered boundary maps thus generated are then digitally combined in three dimensional space with the acquired surface elevation data by means of image processing software stored in a digital computer. The method is particularly applicable for generating terrain models of flooded regions covered entirely or in part by foliage.

Imhoff, Marc L.↗

Data-Informed Synthetic Networks of Water Distribution Systems for Resilience Analysis in Puerto Rico

The increasing potential of infrastructure disruptions calls for high-quality infrastructure models to be used in resilience analysis and decision making. Unfortunately, many utilities and communities do not have access to accurate and detailed models due to a lack of data and resources. Furthermore, security restrictions on sharing infrastructure models present roadblocks to research, analysis, and decision making. Recent advances in the development of synthetic water distribution models provide a potential solution to this problem. There is an opportunity to improve these methods by leveraging incomplete pipe datasets to aid synthetic network generation. To address this gap, we developed a methodology for synthetic network generation that incorporates partial pipe data using a modification of the minimum cost flow algorithm for network generation and pipe sizing. This methodology demonstrates how partial pipe data can be leveraged to improve site-specific synthetic network generation. For the study area of Mayagüez, Puerto Rico, a synthetic model generated using 50% of real pipe data matches the pressure of the validation system with an average error of 23.5 m of head, which improves upon the average error of 31.6 m of head produced by a synthetic model generated using no data of the real pipes. Additionally, synthetic networks are shown to replicate the pressure response under a disruption scenario of the validation network, suggesting potential use in resilience analysis.

resilience analysis↗

GenAI-Based Digital Twins Aided Data Augmentation Increases Accuracy in Real-Time Cokurtosis-Based Anomaly Detection of Wearable Data

Early detection of potential infectious disease outbreaks is crucial for developing effective interventions. In this study, we introduce advanced anomaly detection methods tailored for health datasets collected from wearables, offering insights at both individual and population levels. Leveraging real-world physiological data from wearables, including heart rate and activity, we developed a framework for the early detection of infection in individuals. Despite the availability of data from recent pandemics, substantial gaps remain in data collection, hindering method development. To bridge this gap, we utilized Wasserstein Generative Adversarial Networks (WGANs) to generate realistic synthetic wearable data, augmenting our dataset for training. Subsequently, we use these augmented datasets to implement a cokurtosis-based technique for anomaly detection in multivariate time-series data. Our approach includes a comprehensive assessment of uncertainties in synthetic data compared to the actual data upon which it was modeled, as well as the uncertainty associated with fine-tuning anomaly detection thresholds in physiological measurements. Through our work, we present an enhanced method for early anomaly detection in multivariate datasets, with promising applications in healthcare and beyond. This framework could revolutionize early detection strategies and significantly impact public health response efforts in future pandemics.

Data-Driven Digital Twins↗

Unsupervised Process Anomaly Detection and Identification Using the Leave-One-Variable-Out Approach

Automated anomaly detection and identification can signal equipment issues and pinpoint causes in large-scale industrial systems. For systems with limited failure history, unsupervised machine learning methods can be utilized as they do not require past failures. This study introduces the leave-one-variable-out (LOVO) model, which masks one variable at a time to predict the others, learning underlying process correlations. Detection performance was assessed with synthetic and experimental data, while identification performance used only synthetic data due to its ability to generate labeled anomaly types. For detection using synthetic data, the LOVO model generally outperformed comparative models; while using experimental data, the comparative methods outperformed the LOVO model. However, the comparative methods required selecting a latent size, and these conclusions pertain to using the optimal size. In practice, it would not be feasible to always select the optimal value, and incorrect selections impacted performance. In contrast, the LOVO model does not require a latent space. For identification using synthetic data, the LOVO model was slightly outperformed in interpretability and repeatability but still demonstrated impressive results. These outcomes suggest that the LOVO model is an effective model and may be more easily implemented without the challenging tuning process of selecting a latent size.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Data volume reduction for imaging radar polarimetry

Two alternative methods are presented for digital reduction of synthetic aperture multipolarized radar data using scattering matrices, or using Stokes matrices, of four consecutive along-track pixels to produce averaged data for generating a synthetic polarization image.

Zebker, Howard A.↗

Data volume reduction for imaging radar polarimetry

Two alternative methods are disclosed for digital reduction of synthetic aperture multipolarized radar data using scattering matrices, or using Stokes matrices, of four consecutive along-track pixels to produce averaged data for generating a synthetic polarization image.

Zebker, Howard A.↗

Generating synthetic signaling networks for in silico modeling studies

Predictive models of signaling pathways have proven to be difficult to develop. Reasons include the uncertainty in the number of species, the complexity in species’ interactions, and the sparseness and uncertainty in experimental data. Traditional approaches to developing mechanistic models rely on collecting experimental data and fitting a single model to that data. This approach works for simple systems but has proven unreliable for complex systems such as biological signaling networks. For example, uncertainty and sparseness of the data often result in overfitted models that have little predictive value beyond recapitulating the experimental data itself. Thus, there is a need to develop new approaches to create predictive mechanistic models of complex systems. However, to determine the effectiveness of any new algorithm, a baseline model is needed to test its performance. To meet this need, we developed a method for generating artificial synthetic networks that are reasonably realistic and thus can be treated as ground truth models. These synthetic models can then be used to generate synthetic data for developing and testing algorithms designed to recover the underlying network topology and associated parameters. Here, we describe a simple approach for generating synthetic signaling networks that can be used for this purpose.

42 ENGINEERING↗

BatteryPro: A Python Toolkit for Battery Data Analysis and Machine Learning Predictions

Analyzing battery test data for research & development can be time-consuming since battery tests often run on the order of months to years, generating large volumes of data. BatteryPro is a comprehensive Python package and software designed to facilitate advanced analysis and performance predictions for battery test data. Developed for battery researchers, it supports data types from widely used battery testing instruments, including MACCOR and Biologic cycling systems. The software provides a variety of tools for extracting and plotting key battery parameters such as time, voltage, capacity, current, and pressure. In addition to its extensive data analysis capabilities, BatteryPro features a dedicated machine learning module that employs a Bayesian Gaussian Mixture Model (GMM) to predict battery performance and degradation. Users can generate synthetic capacity fade data, calculate fade metrics, and leverage predictive models to forecast long-term battery behavior. The software's graphical user interface (GUI) enhances usability, allowing researchers to upload, merge, and analyze multiple data files with full customizability. The GUI also supports machine learning predictions, enabling users to fit models and make predictions based on selected data and parameters. BatteryPro is built using QtDesigner, scikit-learn, matplotlib, and pandas, ensuring a high level of customization, flexibility, and accuracy in battery data analysis. This tool aims to empower researchers with the ability to perform detailed battery analysis and make informed predictions, ultimately advancing the field of battery research.

25 - ENERGY STORAGE↗

TwinMe4AD: WGAN-based Digital Twins for Anomaly Detection

SAND2024-08373O TwinMe4AD is a Python-based software tool designed for anomaly detection using digital twins that closely mimic real, wearable healthcare datasets. The tool is invaluable for scenarios where collecting data is either expensive or impractical, serving as a privacy-preserving solution. Sensitive information is protected by training deep learning models on synthetic data derived from real datasets. One of TwinMe4AD's key features is its anomaly detection capability, which is based on fourth-order moments of parameters. This versatile approach can be applied across a range of datasets, from univariate to multivariate, making it compatible with various types of data. It also generates synthetic twins using Wasserstein Generative Adversarial Networks (WGANs), allowing users to create a small cohort of a population similar to that of a village population. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Poorey, Kunal↗

Applying Machine‐Learning Methods to Laser Acceleration of Protons: Lessons Learned From Synthetic Data

ABSTRACT In this study, we consider three different machine‐learning methods—a three‐hidden‐layer neural network, support vector regression, and Gaussian process regression—and compare how well they can learn from a synthetic data set for proton acceleration in the Target Normal Sheath Acceleration regime. The synthetic data set was generated from a previously published theoretical model by Fuchs et al. 2005 that we modified. Once trained, these machine‐learning methods can assist with efforts to maximize the peak proton energy, or with the more general problem of configuring the laser system to produce a proton energy spectrum with desired characteristics. In our study, we focus on both the accuracy of the machine‐learning methods and the performance on one GPU including memory consumption. Although it is arguably the least sophisticated machine‐learning model we considered, support vector regression performed very well in our tests.

Desai, Ronak↗

Digital simulation of the SIR-C sensor electronics

In this paper software for simulation of the response of the SIR-C sensor to a point target is described. Synthetic SAR data is generated by passing successive chirps through a simulation of the transmitter electronics, propagation path and receiver electronics. This result is then processed with a digital correlator to yield the point target response of the system. This allows an accurate assessment of the effect of the radar design on the final image product.

Klein, Jeffrey D.↗

Black Hole Mergers as Probes of Structure Formation

Intense structure formation and reionization occur at high redshift, yet there is currently little observational information about this very important epoch. Observations of gravitational waves from massive black hole (MBH) mergers can provide us with important clues about the formation of structures in the early universe. Past efforts have been limited to calculating merger rates using different models in which many assumptions are made about the specific values of physical parameters of the mergers, resulting in merger rate estimates that span a very wide range (0.1 - 104 mergers/year). Here we develop a semi-analytical, phenomenological model of MBH mergers that includes plausible combinations of several physical parameters, which we then turn around to determine how well observations with the Laser Interferometer Space Antenna (LISA) will be able to enhance our understanding of the universe during the critical z 5 - 30 structure formation era. We do this by generating synthetic LISA observable data (total BH mass, BH mass ratio, redshift, merger rates), which are then analyzed using a Markov Chain Monte Carlo method. This allows us to constrain the physical parameters of the mergers. We find that our methodology works well at estimating merger parameters, consistently giving results within 1- of the input parameter values. We also discover that the number of merger events is a key discriminant among models. This helps our method be robust against observational uncertainties. Our approach, which at this stage constitutes a proof of principle, can be readily extended to physical models and to more general problems in cosmology and gravitational wave astrophysics.

Alicea-Munoz, E.↗

Using Black Hole Mergers to Explore Structure Formation

Observations of gravitational waves from massive black hole mergers will open a new window into the era of structure formation in the early universe. Past efforts have concentrated on calculating merger rates using different physical assumptions, resulting in merger rate estimates that span a wide range (0.1 - 1 0A4 mergers/year). We develop a semi-analytical, phenomenological model of massive black hole mergers that includes plausible combinations of several physical parameters, which we then turn around to determine how well observations with the Laser Interferometer Space Antenna (LISA) will be able to enhance our understanding of the universe during the critical z approx. 5 - 30 epoch. Our approach involves generating synthetic LISA observable data (total BH masses, BH mass ratios, redshifts, merger rates), which are then analyzed using a Markov Chain Monte Carlo method, thus finding constraints for the physical parameters of the mergers. We find that our method works well at estimating merger parameters and that the number of merger events is a key discriminant among models, therefore making our method robust against observational uncertainties. Our approach can also be extended to more physically-driven models and more general problems in cosmology.

Alicea-Munoz, E.↗

Using Black Hole Mergers to Explore Structure Formation

Observations of gravitational waves from massive black hole mergers will open a new window into the era of structure formation in the early universe. Past efforts have concentrated on calculating merger rates using different physical assumptions, resulting in merger rate estimates that span a wide range (0.1 - 10(exp 4) mergers/year). We develop a semi-analytical, phenomenological model of massive black hole mergers that includes plausible combinations of several physical parameters, which we then turn around to determine how well observations with the Laser Interferometer Space Antenna (LISA) will be able to enhance our understanding of the universe during the critical z approximately equal to 5-30 epoch. Our approach involves generating synthetic LISA observable data (total BH masses, BH mass ratios, redshifts, merger rates), which are then analyzed using a Markov Chain Monte Carlo method, thus finding constraints for the physical parameters of the mergers. We find that our method works well at estimating merger parameters and that the number of merger events is a key discriminant among models, therefore making our method robust against observational uncertainties. Our approach can also be extended to more physically-driven models and more general problems in cosmology. This work is supported in part by the Cooperative Education Program at NASA/GSFC.

Alicea-Munoz, E.↗

Black Hole Mergers as Probes of Structure Formation

Observations of gravitational waves from massive black hole (MBH) mergers can provide us with important clues about the era of structure formation in the early universe. Previous research in this field has been limited to calculating merger rates of MBHs using different models where many assumptions are made about the specific values of physical parameters of the mergers, resulting in merger rate estimates that span 5 to 6 orders of magnitude. We develop a semi-analytical, phenomenological model that includes plausible combinations of several physical parameters involved in the mergers. which we then turn around to determine how well LISA observations will be able to enhance our understanding of the universe during the critical z approximately equal to 5-30 structure formation era. We do this by generating synthetic LISA observable data (masses, redshifts, merger rates), which are then analyzed using a Markov Chain Monte Carlo (MCMC) method. This allows us to constrain the physical parameters of the mergers.

Alicea-Munoz, Emily↗

Mapping of CO2 at High Spatiotemporal Resolution using Satellite Observations: Global distributions from OCO-2

Satellite observations of CO2 offer new opportunities to improve our understanding of the global carbon cycle. Using such observations to infer global maps of atmospheric CO2 and their associated uncertainties can provide key information about the distribution and dynamic behavior of CO2, through comparison to atmospheric CO2 distributions predicted from biospheric, oceanic, or fossil fuel flux emissions estimates coupled with atmospheric transport models. Ideally, these maps should be at temporal resolutions that are short enough to represent and capture the synoptic dynamics of atmospheric CO2. This study presents a geostatistical method that accomplishes this goal. The method can extract information about the spatial covariance structure of the CO2 field from the available CO2 retrievals, yields full coverage (Level 3) maps at high spatial resolutions, and provides estimates of the uncertainties associated with these maps. The method does not require information about CO2 fluxes or atmospheric transport, such that the Level 3 maps are informed entirely by available retrievals. The approach is assessed by investigating its performance using synthetic OCO-2 data generated from the PCTM/ GEOS-4/CASA-GFED model, for time periods ranging from 1 to 16 days and a target spatial resolution of 1deg latitude x 1.25deg longitude. Results show that global CO2 fields from OCO-2 observations can be predicted well at surprisingly high temporal resolutions. Even one-day Level 3 maps reproduce the large-scale features of the atmospheric CO2 distribution, and yield realistic uncertainty bounds. Temporal resolutions of two to four days result in the best performance for a wide range of investigated scenarios, providing maps at an order of magnitude higher temporal resolution relative to the monthly or seasonal Level 3 maps typically reported in the literature.

Hammerling, Dorit M.↗

A Physics-Based Digital Twin for Wave Elevation and Seabed Moment Estimation of Offshore Monopiles: Preprint

In this work, we present a proof of concept of a physics-based digital twin for a monopile structure (with overhead inertia) subjected to wave loading. The digital twin is formulated using reduced-order models derived from first principles and combined with a Kalman filter for state estimation. The proposed framework estimates the monopile top motion, the wave elevation, and the section forces and moments along the pile using primarily acceleration measurements at the monopile top. Key innovations include the use of a hydrodynamic shape function to represent distributed wave loading in a compact and computationally efficient manner, and the introduction of a shaping filter to augment the state-space with wave kinematics. Synthetic measurement data are generated using OpenFAST and used as a reference to assess the performance of the digital twin. Results demonstrate that the wave elevation can be accurately reconstructed without direct sea-state measurements as long as the wave regime is inertia-dominated. Under the ideal tested conditions, the total hydrodynamic force and sea-bed bending moment are estimated with relative errors on the order of 1% and correlation coefficients exceeding 96%. Future work will evaluate the estimator's performance under operational uncertainties and more complex loading conditions.

17 WIND ENERGY↗

An Investigation Into Possible Systematic Effects on Neutron Star Radius Estimates using NICER-like Synthetic Data

Neutron star cores contain the densest matter in the observable universe. The state of this matter is of interest in numerous fields, but laboratory experiments cannot explore this matter. Although the composition of the matter would be of great interest, macroscopic observables such as the neutron star mass-radius relation depend primarily on the equation of state (EOS). As a result, precise and reliable radius measurements would be valuable in constraining the EOS. However, most attempts at radius measurements are susceptible to systematic errors, meaning the inferred radius can be significantly biased even though the fit to the data appears to be statistically good. Previous studies suggested that radii inferred using X-ray data provided by NASA’s NICER mission may be more immune from such systematic errors. This is in part because, compared with previous measurements that obtained averaged spectra and fluxes, NICER observes millisecond pulsars by timing the photons so precisely, it is possible to obtain the spectrum as a function of rotational phase and see variations such as heated regions on the star rotate into and out of view hundreds of times per second. This extra information seems promising to break degeneracies and mitigate systematic errors, but a more in-depth study is necessary. We report the first steps of that study, in which we generate NICER-like synthetic data and determine the quality of fit and bias in radius obtained when we fit the data using a model different from the model used to generate the data.

Isiah Holt↗