Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Reliable Measures of Spread in High Dimensional Latent Spaces

Understanding geometric properties of the latent spaces of natural language processing models allows the manipulation of these properties for improved performance on downstream tasks. One such property is the amount of data spread in a model’s latent space, or how fully the available latent space is being used. We demonstrate that the commonly used measures of data spread, average cosine similarity and a partition function min/max ratio I (V), do not provide reliable metrics to compare the use of latent space across data distributions. We propose and examine six alternative measures of data spread, all of which improve over these current metrics when applied to seven synthetic data distributions. Of our proposed measures, we recommend one principal component-based measure and one entropy-based measure that provide reliable, relative measures of spread and can be used to compare models of different sizes and dimensionalities.

97 MATHEMATICS AND COMPUTING↗

Distributed non-negative matrix factorization with determination of the number of latent features

The holistic analysis and understanding of the latent (that is, not directly observable) variables and patterns buried in large datasets is crucial for data-driven science, decision making and emergency response. Such exploratory analyses require devising unsupervised learning methods for data mining and extraction of the latent features, and non-negative matrix factorization (NMF) is one of the prominent such methods. NMF is based on compute-intense non-convex constrained minimization, which, for large datasets requires fast and distributed algorithms. However, current parallel implementations of NMF fail to estimate the number of latent features. In practice, identifying these features is both difficult and significant for pattern recognition and latent feature analysis, especially for large dense matrices. Here, we introduce a distributed NMF algorithm coupled with distributed custom clustering followed by a stability analysis on dense data, which we call DnMFk, to determine the number of latent variables. The results on synthetic data and the classical Swimmer data set demonstrate the accuracy of model determination while scaling nearly linearly across multiple processors for large data. Further, we employ DnMFk to determine the number of hidden features from a terabyte matrix.

97 MATHEMATICS AND COMPUTING↗

Airborne hyperspectral imaging of cover crops through radiative transfer process-guided machine learning

Cover cropping between cash crop growing seasons is a multifunctional conservation practice. Timely and accurate monitoring of cover crop traits, notably aboveground biomass and nutrient content, is beneficial to agricultural stakeholders to improve management and understand outcomes. Currently, there is a scarcity of spatially and temporally resolved information for assessing cover crop growth. Remote sensing has a high potential to fill this need, but conventional empirical regression operated with coarse-resolution multispectral data has large uncertainties. Therefore, this study utilized airborne hyperspectral imaging techniques and developed new process-guided machine learning approaches (PGML) for cover crop monitoring. Specifically, we deployed an airborne hyperspectral system covering visible to shortwave-infrared wavelengths (400–2400 nm) to acquire high spatial (0.5 m) and spectral (3–5 nm) resolution reflectance over 23 cover crop fields across Central Illinois in March and April of 2021. Airborne hyperspectral surface reflectance with high spectral and spatial resolution can be well matched with field data to quantify cover crop traits. Furthermore, the PGML models were pre-trained by synthetic data from soil-vegetation radiative transfer modeling (one million records), and then fine-tuned with field data of cover crop biomass and nutrient content. Results show that airborne hyperspectral data with PGML can achieve high accuracy to predict cover crop aboveground biomass (R 2 = 0.72, relative RMSE = 15.16%) and nitrogen content (R 2 = 0.69, relative RMSE = 16.59%) through leave-one-field-out cross-validation. Unlike the pure data-driven approach (e.g., partial least-squares regression), PGML incorporated radiative transfer knowledge and obtained higher predictive performance with fewer field data. Meanwhile, with field data for model fine-tuning, PGML predicted biomass more accurately than the inversion of radiative transfer models. Here we also found that the red edge has a high contribution in quantifying aboveground biomass and nitrogen content, followed by green and shortwave spectra. This study demonstrated the first attempt of utilizing hyperspectral remote sensing to accurately quantify cover crop traits. We highlight the strength of PGML in exploiting sensing data to quantify ecosystem variables to advance agroecosystem monitoring for sustainable agricultural management.

60 APPLIED LIFE SCIENCES↗

Subspace-Driven Learning for Anomaly Detection in Process Transients

Nuclear power plant (NPP) monitoring and diagnostic centers are actively investigating and implementing automated anomaly detection algorithms to help plants catch anomalies sooner, thereby preventing or reducing the duration of unexpected shutdowns. Current machine learning-based anomaly detection methods are expected to be highly effective during stable, full-power operations because NPPs typically operate as baseload power generators, meaning there are extensive operating data available from plant equipment. However, it is expected that anomaly detection methods will face significant challenges during transient conditions (i.e., when power output falls below full power) because plants only occasionally operate at these lower power levels, generating sparse transient operational data, and resulting in false alarms or missed detections. Here, to address this issue, transfer learning is used, which for this problem leverages knowledge (in the form of learned features) from stable, full-power operations to improve detection accuracy during transient conditions, even with limited data. In this effort, a novel subspace approach is developed to transfer a subset of the data features from full power operation to transients. This approach is validated through experiments using synthetic data and was found to outperform two baseline transfer learning approaches in anomaly detection performance across a range of amounts of transient data used in the training process.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Direct structural retrieval from gas-phase ultrafast diffraction data using a genetic algorithm

Ultrafast scattering techniques such as ultrafast electron diffraction and ultrafast x-ray diffraction have been utilized to elucidate the structural dynamics, reaction intermediates, and final products in molecular reactions following photoexcitation. The time-dependent structures are typically not directly retrieved from the experimental data, but they rely on comparison with calculations. The genetic algorithm (GA), a global optimization strategy, can be used to retrieve the molecular structures directly from diffraction patterns without any theoretical input. However, the robustness of the GA with respect to real experimental conditions such as a limited momentum transfer range, noise, and artifacts has not been studied in detail. In this work, we characterize the performance of the GA with simulated data that mimic realistic experimental conditions. We have developed and implemented a variant of the GA specific to diffraction measurements which performs better in the presence of imperfect data compared to the standard implementation of the GA. We demonstrate this method with both synthetic data and experimental ultrafast electron diffraction data on the UV-induced photodissociation of trifluoroiodomethane (C⁢F 3⁡ I) molecules.

74 ATOMIC AND MOLECULAR PHYSICS↗

Analytical Modeling of Exoplanet Transit Spectroscopy with Dimensional Analysis and Symbolic Regression

Abstract The physical characteristics and atmospheric chemical composition of newly discovered exoplanets are often inferred from their transit spectra, which are obtained from complex numerical models of radiative transfer. Alternatively, simple analytical expressions provide insightful physical intuition into the relevant atmospheric processes. The deep-learning revolution has opened the door for deriving such analytical results directly with a computer algorithm fitting to the data. As a proof of concept, we successfully demonstrate the use of symbolic regression on synthetic data for the transit radii of generic hot-Jupiter exoplanets to derive a corresponding analytical formula. As a preprocessing step, we use dimensional analysis to identify the relevant dimensionless combinations of variables and reduce the number of independent inputs, which improves the performance of the symbolic regression. The dimensional analysis also allowed us to mathematically derive and properly parameterize the most general family of degeneracies among the input atmospheric parameters that affect the characterization of an exoplanet atmosphere through transit spectroscopy.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine Learning-based Prediction of Departure from Nucleate Boiling Power for the PSBT Benchmark

Machine Learning (ML) has seen an exponential growth in its applications due to its advanced data driven prediction capabilities. The study presents a data-driven approach as a preliminary attempt to predict the power at which departure from nucleate boiling (DNB) occurs in pressurized water reactors (PWRs) by constructing an advanced ML algorithm that takes outlet pressure, inlet temperature and inlet mass flux as the input features. DNB is a critical heat flux (CHF) phenomenon seen in PWRs. The experimental data from the PWR subchannel and bundle tests (PSBT) benchmark is first used to train an artificial neural network (ANN) to predict the DNB power, which produces a root mean square error (RMSE) of 6.89 kW/m when tested on a blind subset of the PSBT data. Since the PSBT dataset is relatively small to train an accurate ANN, a data augmentation methodology based on generative adversarial networks (GANs) is used to expand the training dataset. By assuming that the real data follows a certain distribution, GANs try to learn that underlying distribution to generate similar synthetic data to augment the database and to improve the predictive capabilities of the ANN. The data generated from GANs are validated using 1-nearest neighbor and kernel maximum mean discrepancy. To further ensure data from GAN is similar to PSBT, the data is tested and filtered out using the sub-channel thermal-hydraulic code CTF. The results indicate that with the addition of 120 data points from GAN the RMSE reduces to 4.84 kW/m showing promising results for future developments.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Training material models using gradient descent algorithms

High temperature design requires accurate constitutive models to describe material inelastic deformation and failure behavior. Oftentimes, calibrating accurate models devolves into the problem of fitting the model parameters against experimental test data. Here, we present the pyopmat package, an open source framework for calibrating constitutive models against experiment data subjected to various loading conditions using machine learning techniques. The package calculates the exact gradient of the model response with respect to the parameters using a combination of automatic differentiation and the adjoint method. Given this exact gradient, we compare the performance of several gradient-based optimization techniques in fitting realistic constitutive models against data. Here, we demonstrate the efficiency and accuracy of our package through example problems using both synthetic data, generated using known parameter sets, under monotonic and cyclic loading conditions and also with an example applying the techniques developed here to actual high temperature creep-fatigue test data.

36 MATERIALS SCIENCE↗

Dark Energy Survey Year 3 results: Cosmology from cosmic shear and robustness to modeling uncertainty

Here, this work and its companion paper, Amon et al. [Phys. Rev. D 105, 023514 (2022)], present cosmic shear measurements and cosmological constraints from over 100 million source galaxies in the Dark Energy Survey (DES) Year 3 data. We constrain the lensing amplitude parameter 𝑆 8 ≡𝜎 8 ⁢$\sqrt{Ω_{m}/0.3}$ at the 3% level in Λ⁢ CDM: 𝑆 8 =0.75⁢9$^{+0.025}_{−0.023}$ (68% CL). Our constraint is at the 2% level when using angular scale cuts that are optimized for the Λ⁢ CDM analysis: 𝑆 8 =0.77⁢2$^{+0.018}_{−0.017}$ (68% CL). With cosmic shear alone, we find no statistically significant constraint on the dark energy equation-of-state parameter at our present statistical power. We carry out our analysis blind, and compare our measurement with constraints from two other contemporary weak lensing experiments: the Kilo-Degree Survey (KiDS) and Hyper-Suprime Camera Subaru Strategic Program (HSC). We additionally quantify the agreement between our data and external constraints from the Cosmic Microwave Background (CMB). Our DES Y3 result under the assumption of Λ⁢ CDM is found to be in statistical agreement with Planck 2018, although favors a lower 𝑆 8 than the CMB-inferred value by 2.3⁢𝜎 (a 𝑝-value of 0.02). This paper explores the robustness of these cosmic shear results to modeling of intrinsic alignments, the matter power spectrum and baryonic physics. We additionally explore the statistical preference of our data for intrinsic alignment models of different complexity. The fiducial cosmic shear model is tested using synthetic data, and we report no biases greater than 0.3⁢𝜎 in the plane of 𝑆 8 ×Ω m caused by uncertainties in the theoretical models.

79 ASTRONOMY AND ASTROPHYSICS↗

Fast Gaussian Process Estimation for Large-Scale In Situ Inference using Convolutional Neural Networks

Exascale computing will bring with it significant I/O limitations. One foreseeable consequence of such restrictions is that the user can save only a small fraction of complex simulation data to disk for subsequent analysis. An alternative is to fit statistical models to data in situ, that is, inside the simulation as it runs. This option requires extremely fast statistical estimation to avoid slowing down the simulation. Gaussian processes (GPs) have state-of-the-art predictive performance for modeling spatial data. However, standard estimation methods for GPs scale quite poorly to large data sets as parameter estimation requires inverting a covariance matrix to the size of the data set. In the presented work, we use a convolutional neural network (CNN) to predict the GP parameters for a spatial data set, from a simulation or otherwise, rather than optimize the parameters directly. Here, our presented case study models spatial data from E3SM, the Department of Energy’s Exascale climate model. The CNN is trained on synthetic data simulated from GP models with known parameters and then applied to data from the climate simulation. In the presented examples, the neural network scheme produces parameter estimates that compare well with standard methods such as maximum likelihood estimation in predictive performance but is obtained four orders of magnitude faster.

big data↗

Synthetic spectra for Lyman- α forest analysis in the Dark Energy Spectroscopic Instrument

Synthetic data sets are used in cosmology to test analysis procedures, to verify that systematic errors are well understood and to demonstrate that measurements are unbiased. In this work we describe the methods used to generate synthetic datasets of Lyman-α quasar spectra aimed for studies with the Dark Energy Spectroscopic Instrument (DESI). In particular, we focus on demonstrating that our simulations reproduces important features of real samples, making them suitable to test the analysis methods to be used in DESI and to place limits on systematic effects on measurements of Baryon Acoustic Oscillations (BAO). We present a set of mocks that reproduce the statistical properties of the DESI early data set with good agreement. Additionally, we use a synthetic dataset to forecast the BAO scale constraining power of the completed DESI survey through the Lyman-α forest.

79 ASTRONOMY AND ASTROPHYSICS↗

Robust Infrasound Detection via Deep Learning (RIDDL)

RIDDL is a suite of software tools that enable AI/ML analysis of infrasound data. Methods and capabilities include construction, evaluation, and application of models for infrasound signal analysis, construction of synthetic data useful for construction and evaluation, as well as various other advanced data science tools enabling infrasound signal detection and categorization as well as downstream analysis methods such as localization and characterization of detected sources.

Blom, Philip↗

Identifying Heterogeneous Micromechanical Properties of Biological Tissues via Physics–Informed Neural Networks

The heterogeneous micromechanical properties of biological tissues have profound implications across diverse medical and engineering domains. However, identifying full-field heterogeneous elastic properties of soft materials using traditional engineering approaches is fundamentally challenging due to difficulties in estimating local stress fields. Recently, there has been a growing interest in data-driven models for learning full-field mechanical responses, such as displacement and strain, from experimental or synthetic data. However, research studies on inferring full-field elastic properties of materials, a more challenging problem, are scarce, particularly for large deformation, hyperelastic materials. Here, a physics-informed machine learning approach is proposed to identify the elasticity map in nonlinear, large deformation hyperelastic materials. This study reports the prediction accuracies and computational efficiency of physics-informed neural networks (PINNs) in inferring the heterogeneous elasticity maps across materials with structural complexity that closely resemble real tissue microstructure, such as brain, tricuspid valve, and breast cancer tissues. Further, the improved architecture is applied to three hyperelastic constitutive models: Neo-Hookean, Mooney Rivlin, and Gent. Furthermore, the improved network architecture consistently produces accurate estimations of heterogeneous elasticity maps, even when there is up to 10% noise present in the training data.

59 BASIC BIOLOGICAL SCIENCES↗

Source term estimation using noble gas and aerosol samples

Algorithms that estimate the location, time, and magnitude of a point-source atmospheric release using remotely sampled air concentrations typically use data for a single chemical or radioactive isotope. Here, a Bayesian algorithm is presented that uses data from multiple radioactive isotopes that are all released in the same short-duration event. Data from noble gas and aerosol samplers can be used simultaneously in the model. Application to a large synthetic data set using four isotopes shows the new algorithm generally gives more accurate location and time estimates than a comparable model using a single isotope.

54 ENVIRONMENTAL SCIENCES↗

Reconstruction of 2D line-integrated electron density using angular filter refractometry and a fast marching Eikonal solver

Refraction of an optical probe beam by a plasma can be measured with angular filter refractometry (AFR), which produces an image of the beam’s 2D spatial profile that contains intensity contours corresponding to curves of constant refraction angle. Further analysis is required to reconstruct the underlying line-integrated electron density. Most prior efforts to calculate density from AFR data have been limited to 1D analysis or forward-fitting techniques. Here, in this paper, we detail the use of a fast-marching Eikonal solver to directly invert AFR data and obtain the full 2D line-integrated electron density. The analysis method is first verified with synthetic data and then applied to experimental measurements of single and colliding plasma plumes collected at the OMEGA EP Laser Facility. The calculated densities agree with 1D results and are shown to be consistent with the original AFR measurements via forward modeling. We also discuss ways to improve the precision of this technique.

McCluskey, B. [Princeton Univ., NJ (United States)↗

First M87 Event Horizon Telescope Results. VII. Polarization of the Ring

In 2017 April, the Event Horizon Telescope (EHT) observed the near-horizon region around the supermassive black hole at the core of the M87 galaxy. These 1.3 mm wavelength observations revealed a compact asymmetric ring-like source morphology. This structure originates from synchrotron emission produced by relativistic plasma located in the immediate vicinity of the black hole. Here we present the corresponding linear-polarimetric EHT images of the center of M87. We find that only a part of the ring is significantly polarized. The resolved fractional linear polarization has a maximum located in the southwest part of the ring, where it rises to the level of ~15%. The polarization position angles are arranged in a nearly azimuthal pattern. We perform quantitative measurements of relevant polarimetric properties of the compact emission and find evidence for the temporal evolution of the polarized source structure over one week of EHT observations. The details of the polarimetric data reduction and calibration methodology are provided. We carry out the data analysis using multiple independent imaging and modeling techniques, each of which is validated against a suite of synthetic data sets. The gross polarimetric structure and its apparent evolution with time are insensitive to the method used to reconstruct the image. These polarimetric images carry information about the structure of the magnetic fields responsible for the synchrotron emission. Their physical interpretation is discussed in an accompanying publication.

79 ASTRONOMY AND ASTROPHYSICS↗

First Sagittarius A* Event Horizon Telescope Results. VII. Polarization of the Ring

Abstract The Event Horizon Telescope observed the horizon-scale synchrotron emission region around the Galactic center supermassive black hole, Sagittarius A* (Sgr A*), in 2017. These observations revealed a bright, thick ring morphology with a diameter of 51.8 ± 2.3 μas and modest azimuthal brightness asymmetry, consistent with the expected appearance of a black hole with mass M ≈ 4 × 106 M ⊙. From these observations, we present the first resolved linear and circular polarimetric images of Sgr A*. The linear polarization images demonstrate that the emission ring is highly polarized, exhibiting a prominent spiral electric vector polarization angle pattern with a peak fractional polarization of ∼40% in the western portion of the ring. The circular polarization images feature a modestly (∼5%–10%) polarized dipole structure along the emission ring, with negative circular polarization in the western region and positive circular polarization in the eastern region, although our methods exhibit stronger disagreement than for linear polarization. We analyze the data using multiple independent imaging and modeling methods, each of which is validated using a standardized suite of synthetic data sets. While the detailed spatial distribution of the linear polarization along the ring remains uncertain owing to the intrinsic variability of the source, the spiraling polarization structure is robust to methodological choices. The degree and orientation of the linear polarization provide stringent constraints for the black hole and its surrounding magnetic fields, which we discuss in an accompanying publication.

79 ASTRONOMY AND ASTROPHYSICS↗

Utah FORGE: Interferometric Synthetic Aperture Radar Data from 2023 and 2024

The dataset comprises Interferometric Synthetic Aperture Radar (InSAR) data from the TerraSAR-X and TanDEM-X satellite missions, covering the Utah FORGE site. This data includes interferometric pairs created using GMT-SAR processing software, chosen for their short orbital separations between May 1, 2023, and June 30, 2024. Included are various data and metadata, including Digital Elevation Models, unit vectors, and correlation coefficients. The dataset is packaged in several compressed tar files and formatted in NetCDF. To utilize this dataset, users will need software capable of handling NetCDF files and tools for decompressing tar files.

15 GEOTHERMAL ENERGY↗