Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

The Evaluation of Machine Learning Techniques for Isotope Identification Contextualized by Training and Testing Spectral Similarity

Precise gamma-ray spectral analysis is crucial in high-stakes applications, such as nuclear security. Research efforts toward implementing machine learning (ML) approaches for accurate analysis are limited by the resemblance of the training data to the testing scenarios. The underlying spectral shape of synthetic data may not perfectly reflect measured configurations, and measurement campaigns may be limited by resource constraints. Consequently, ML algorithms for isotope identification must maintain accurate classification performance under domain shifts between the training and testing data. To this end, four different classifiers (Ridge, Random Forest, Extreme Gradient Boosting, and Multilayer Perceptron) were trained on the same dataset and evaluated on twelve other datasets with varying standoff distances, shielding, and background configurations. A tailored statistical approach was introduced to quantify the similarity between the training and testing configurations, which was then related to the predictive performance. Wilcoxon signed-rank tests revealed that the OVR-wrapped XGB significantly outperformed the other algorithms, with confidence levels of 99.0% or above for the 133Ba, 60Co, 137Cs, and 152Eu sources. The findings from this work are significant as they outline techniques to promote the development of robust ML-based approaches for isotope identification.

domain adaptation↗

A semblance measure for model comparison

Algorithmic and computational advances have made it possible that geophysical survey and earth model design can be aided by many systematic trial inverse-modelling runs with synthetic data. Such may, for example, come up in machine-learning approaches. Automated image appraisal pertaining to such applications will involve common statistical tests for goodness-of-data fit as a primary evaluation method. However, solution non-uniqueness may render multiple images equivalent in terms of their data fit, requiring secondary categorizers. A logical choice for classifying synthetic-imaging results quantifies the goodness of model fit where a known reference model replaces the observational input. The task of model intercomparison in terms of measuring the resemblance to the reference model poses challenges to common distance-based metrics like root mean square error and mean absolute error. First, distance-based metrics can introduce spurious contributions when smooth models with fuzzy target contours are to be compared against a sharp reference. Second, large differences due to parameter-estimation overshoots can dominate distance metrics. Here, we propose a remedy that is referred to as semblance and is based on the idea of logistic functions, where a binary-dependent variable adds non-zero or zero accumulation terms for the, respectively, passing or failing of preset target thresholds. This classifying approach is amenable to an objective where model feature recognition is primary. Numerical comparisons to distance-based metrics provide evidence for the advantages of the semblance in view of this objective. Geophysical imaging in conjunction with machine-learning is seen as a benefitting upcoming application area.

58 GEOSCIENCES↗

DeFault: DEep‐Learning‐Based FAULT Delineation Using the IBDP Passive Seismic Data at the Decatur CO2 Storage Site

Abstract The carbon capture, utilization, and storage (CCUS) framework is an essential component in reducing greenhouse gas emissions, with its success hinging on the comprehensive knowledge of subsurface geology and geomechanics. Passive seismic event relocation and fault detection offer vital insights into subsurface structures and the ability to monitor fluid migration pathways. Accurate identification and localization of seismic events, however, face significant challenges, including the necessity for high‐quality seismic data and advanced computational methods. To address these challenges, we introduce a novel deep learning method, , specifically designed for passive seismic source relocation and fault delineating for passive seismic monitoring projects. By leveraging data domain‐adaptation, allows us to train a neural network with labeled synthetic data and apply it directly to field data. Using , the passive seismic sources are automatically clustered based on their recording time and spatial locations, and subsequently, faults and fractures are delineated accordingly. We demonstrate the efficacy of on a field case study involving injection related microseismic data from Decatur, Illinois area. Our approach accurately and efficiently relocated passive seismic events, identified faults and could aid in potential damage induced by seismicity. Our results highlight the potential of as a valuable tool for passive seismic monitoring, emphasizing its role in ensuring CCUS project safety. This research bolsters the understanding of subsurface characterization in CCUS, illustrating machine learning’s capacity to refine these methods. Ultimately, our work has significant implications for CCUS technology deployment, an essential strategy in combating climate change. Plain Language Summary In our quest to tackle climate change, we use a strategy known as carbon capture, utilization, and storage (CCUS) to keep greenhouse gases out of the atmosphere. This strategy relies heavily on our ability to understand what's happening deep under the earth's surface. To make sure we store super critical safely, we need to accurately map out the geological structure, especially faults, but this is tough without high‐quality data and complex computer programs. We've developed a new tool called “DeFault,” which uses advanced machine learning to improve how we find and map these underground features. “DeFault” is smart enough to learn from numerically simulated data and then apply what it’s learned to real‐world situations. It groups together seismic activity—tiny tremors and shifts in the earth—based on when and where they happen, which helps us spot where there might be cracks or faults. We tested “DeFault” in Illinois, where CO 2 is injected underground, and it successfully pinpointed where these tremors occurred and mapped out the faults, helping to prevent accidents accurately in the future. Our study shows that “DeFault” will be a powerful ally in making CCUS safer and more effective, especially for the Illinois Basin Decatur Project. Key Points Faults and fractures introduced by carbon storage can be monitored by passive seismicity DeFault algorithm enables an automatic process for accurate and efficient passive seismic event locating and clustering

58 GEOSCIENCES↗

MPACT Safeguards Modeling: FY25 Update

Sandia National Laboratories develops and maintains several open-source software packages to support material accountancy analyses. This includes the Material Accountancy Performance Indicator Toolkit (MAPIT), the Fissile Facility Flow Modeler (F3M) and the Separation and Safeguards Performance Model Library (SSPM-L). MAPIT is responsible for performing statistical safeguards analyses on bulk and itemized data from nuclear fuel cycle facilities and can operate on real or synthetic data. MAPIT is the only open-source software for such analyses. F3M is a library of modules, built in MATLAB Simulink, that contain pre made blocks to represent different generic fuel cycle processes. These blocks can be used together in a modular fashion to represent and simulate nuclear fuel cycle processes with the goal of improving facility-level accountancy during the design phase. F3M is also an open-source library. Finally, the SSPM-L library is a series of completed models built from F3M. The library includes facility models such as a generic PUREX facility and a fuel fabrication facility. The SSPM-L library is not open source, but is available to collaborators with a relevant use case. These tools include modeling and simulation pipelines to simulate nuclear fuel cycle facilities and the underlying software needed to simulate measurement uncertainty and perform statistical analyses. Together, these tools can perform end-to-end nuclear material accountancy analyses. This report documents the various improvements made to these tools in FY25. Specifically, we added new statistical test, new statistical modeling capabilities, new fuel cycle facility models, and launched a new open-source model component library.

97 MATHEMATICS AND COMPUTING↗

Explosion Discrimination Using Seismic Gradiometry and Spectral Filtering of Data

Here, we present a new method to discriminate between earthquakes and buried explosions using observed seismic data. The method is different from previous seismic discrimination algorithms in two main ways. First, we use seismic spatial gradients, as well as the wave attributes estimated from them (referred to as gradiometric attributes), rather than the conventional three-component seismograms recorded on a distributed array. The primary advantage of this is that a gradiometer is only a fraction of a wavelength in aperture compared with a conventional seismic array or network. Second, we use the gradiometric attributes as input data into a machine learning algorithm. The resulting discrimination algorithm uses the norms of truncated principal components obtained from the gradiometric data to distinguish the two classes of seismic events. Using high-fidelity synthetic data, we show that the data and gradiometric attributes recorded by a single seismic gradiometer performs as well as a conventional distributed array at the event type discrimination task.

58 GEOSCIENCES↗

Polyconvex neural network models of thermoelasticity

Machine-learning function representations such as neural networks have proven to be excellent constructs for constitutive modeling due to their flexibility to represent highly nonlinear data and their ability to incorporate constitutive constraints, which also allows them to generalize well to unseen data. Here, in this work, we extend a polyconvex hyperelastic neural network framework to (isotropic) thermo-hyperelasticity by specifying the thermodynamic and material theoretic requirements for an expansion of the Helmholtz free energy expressed in terms of deformation invariants and temperature. Different formulations which a priori ensure polyconvexity with respect to deformation and concavity with respect to temperature are proposed and discussed. The physics-augmented neural networks are furthermore calibrated with a recently proposed sparsification algorithm that not only aims to fit the training data but also penalizes the number of active parameters, which prevents overfitting in the low data regime and promotes generalization. The performance of the proposed framework is demonstrated on synthetic data, which illustrate the expected thermomechanical phenomena, and existing temperature-dependent uniaxial tension and tension-torsion experimental datasets.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Sparse-Data Deep Learning Strategies for Radiographic Non-Destructive Testing

Radiography is an imaging technique used in a variety of applications, such as medical diagnosis, airport security, and nondestructive testing. We present a deep learning system for extracting information from radiographic images. We perform various prediction tasks using our system, including material classification and regression on the dimensions of a given object that is being radiographed. Our system is designed to address the sparse-data issue for radiographic nondestructive testing applications. It uses a radiographic simulation tool for synthetic data augmentation, and it uses transfer learning with a pre-trained convolutional neural network model. Using this system, our preliminary results indicate that the object geometry regression task saw an improvement of 70% in the R-squared value when using a multi-regime model. In addition, we increase the performance of the object material classification tasks by utilizing data from different imaging systems. In particular, using neutron imaging improved the material classification accuracy by 20% when compared to x-ray imaging.

convolutional neural networks↗

Unified Wind-Wave Growth and Spectrum Functions for All Water Depths: Field Observations and Model Results

Abstract Wind-wave development is governed by the fetch- or duration-limited growth principle that is expressed as a pair of similarity functions relating the dimensionless elevation variance (wave energy) and spectral peak frequency to fetch or duration. Combining the pair of similarity functions, the fetch or duration variable can be removed to form a dimensionless function of elevation variance and spectral peak frequency, which is interpreted as the wave energy evolution with wave age. The relationship is initially developed for quasi-neural stability and quasi-steady wind forcing conditions. Further analyses show that the same fetch, duration, and wave-age similarity functions are applicable to unsteady wind forcing conditions, including rapidly accelerating and decelerating mountain gap wind episodes and tropical cyclone (TC) wind fields. Here it is shown that with the dimensionless frequency converted to dimensionless wavenumber using the surface wave dispersion relationship, the same similarity function is applicable in all water depths. Field data collected in shallow to deep waters and mild to TC wind conditions and synthetic data generated by spectrum model computations are assembled to illustrate the applicability. For the simulation work, the finite-depth wind-wave spectrum model and its shoaling function are formulated for variable spectral slopes. Given wind speed, wave age, and water depth, the measured and spectrum-computed significant wave heights and the associated growth parameters are in good agreement in forcing conditions from mild to TC winds and in all depths from deep ocean to shallow lake. Significance Statement This paper presents a growth function and spectrum model to describe wind-wave development in all water depths. Their applicability covers a wide range of wind forcing conditions including steady, accelerating, decelerating, and tropical cyclone events. Support for the unified spectrum model and growth function is presented with field observations and numerical computations.

Hwang, Paul A.↗

Reliable Measures of Spread in High Dimensional Latent Spaces

Understanding geometric properties of the latent spaces of natural language processing models allows the manipulation of these properties for improved performance on downstream tasks. One such property is the amount of data spread in a model’s latent space, or how fully the available latent space is being used. We demonstrate that the commonly used measures of data spread, average cosine similarity and a partition function min/max ratio I (V), do not provide reliable metrics to compare the use of latent space across data distributions. We propose and examine six alternative measures of data spread, all of which improve over these current metrics when applied to seven synthetic data distributions. Of our proposed measures, we recommend one principal component-based measure and one entropy-based measure that provide reliable, relative measures of spread and can be used to compare models of different sizes and dimensionalities.

97 MATHEMATICS AND COMPUTING↗

Airborne hyperspectral imaging of cover crops through radiative transfer process-guided machine learning

Cover cropping between cash crop growing seasons is a multifunctional conservation practice. Timely and accurate monitoring of cover crop traits, notably aboveground biomass and nutrient content, is beneficial to agricultural stakeholders to improve management and understand outcomes. Currently, there is a scarcity of spatially and temporally resolved information for assessing cover crop growth. Remote sensing has a high potential to fill this need, but conventional empirical regression operated with coarse-resolution multispectral data has large uncertainties. Therefore, this study utilized airborne hyperspectral imaging techniques and developed new process-guided machine learning approaches (PGML) for cover crop monitoring. Specifically, we deployed an airborne hyperspectral system covering visible to shortwave-infrared wavelengths (400–2400 nm) to acquire high spatial (0.5 m) and spectral (3–5 nm) resolution reflectance over 23 cover crop fields across Central Illinois in March and April of 2021. Airborne hyperspectral surface reflectance with high spectral and spatial resolution can be well matched with field data to quantify cover crop traits. Furthermore, the PGML models were pre-trained by synthetic data from soil-vegetation radiative transfer modeling (one million records), and then fine-tuned with field data of cover crop biomass and nutrient content. Results show that airborne hyperspectral data with PGML can achieve high accuracy to predict cover crop aboveground biomass (R 2 = 0.72, relative RMSE = 15.16%) and nitrogen content (R 2 = 0.69, relative RMSE = 16.59%) through leave-one-field-out cross-validation. Unlike the pure data-driven approach (e.g., partial least-squares regression), PGML incorporated radiative transfer knowledge and obtained higher predictive performance with fewer field data. Meanwhile, with field data for model fine-tuning, PGML predicted biomass more accurately than the inversion of radiative transfer models. Here we also found that the red edge has a high contribution in quantifying aboveground biomass and nitrogen content, followed by green and shortwave spectra. This study demonstrated the first attempt of utilizing hyperspectral remote sensing to accurately quantify cover crop traits. We highlight the strength of PGML in exploiting sensing data to quantify ecosystem variables to advance agroecosystem monitoring for sustainable agricultural management.

60 APPLIED LIFE SCIENCES↗

Subspace-Driven Learning for Anomaly Detection in Process Transients

Nuclear power plant (NPP) monitoring and diagnostic centers are actively investigating and implementing automated anomaly detection algorithms to help plants catch anomalies sooner, thereby preventing or reducing the duration of unexpected shutdowns. Current machine learning-based anomaly detection methods are expected to be highly effective during stable, full-power operations because NPPs typically operate as baseload power generators, meaning there are extensive operating data available from plant equipment. However, it is expected that anomaly detection methods will face significant challenges during transient conditions (i.e., when power output falls below full power) because plants only occasionally operate at these lower power levels, generating sparse transient operational data, and resulting in false alarms or missed detections. Here, to address this issue, transfer learning is used, which for this problem leverages knowledge (in the form of learned features) from stable, full-power operations to improve detection accuracy during transient conditions, even with limited data. In this effort, a novel subspace approach is developed to transfer a subset of the data features from full power operation to transients. This approach is validated through experiments using synthetic data and was found to outperform two baseline transfer learning approaches in anomaly detection performance across a range of amounts of transient data used in the training process.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Direct structural retrieval from gas-phase ultrafast diffraction data using a genetic algorithm

Ultrafast scattering techniques such as ultrafast electron diffraction and ultrafast x-ray diffraction have been utilized to elucidate the structural dynamics, reaction intermediates, and final products in molecular reactions following photoexcitation. The time-dependent structures are typically not directly retrieved from the experimental data, but they rely on comparison with calculations. The genetic algorithm (GA), a global optimization strategy, can be used to retrieve the molecular structures directly from diffraction patterns without any theoretical input. However, the robustness of the GA with respect to real experimental conditions such as a limited momentum transfer range, noise, and artifacts has not been studied in detail. In this work, we characterize the performance of the GA with simulated data that mimic realistic experimental conditions. We have developed and implemented a variant of the GA specific to diffraction measurements which performs better in the presence of imperfect data compared to the standard implementation of the GA. We demonstrate this method with both synthetic data and experimental ultrafast electron diffraction data on the UV-induced photodissociation of trifluoroiodomethane (C⁢F 3⁡ I) molecules.

74 ATOMIC AND MOLECULAR PHYSICS↗

Analytical Modeling of Exoplanet Transit Spectroscopy with Dimensional Analysis and Symbolic Regression

Abstract The physical characteristics and atmospheric chemical composition of newly discovered exoplanets are often inferred from their transit spectra, which are obtained from complex numerical models of radiative transfer. Alternatively, simple analytical expressions provide insightful physical intuition into the relevant atmospheric processes. The deep-learning revolution has opened the door for deriving such analytical results directly with a computer algorithm fitting to the data. As a proof of concept, we successfully demonstrate the use of symbolic regression on synthetic data for the transit radii of generic hot-Jupiter exoplanets to derive a corresponding analytical formula. As a preprocessing step, we use dimensional analysis to identify the relevant dimensionless combinations of variables and reduce the number of independent inputs, which improves the performance of the symbolic regression. The dimensional analysis also allowed us to mathematically derive and properly parameterize the most general family of degeneracies among the input atmospheric parameters that affect the characterization of an exoplanet atmosphere through transit spectroscopy.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine Learning-based Prediction of Departure from Nucleate Boiling Power for the PSBT Benchmark

Machine Learning (ML) has seen an exponential growth in its applications due to its advanced data driven prediction capabilities. The study presents a data-driven approach as a preliminary attempt to predict the power at which departure from nucleate boiling (DNB) occurs in pressurized water reactors (PWRs) by constructing an advanced ML algorithm that takes outlet pressure, inlet temperature and inlet mass flux as the input features. DNB is a critical heat flux (CHF) phenomenon seen in PWRs. The experimental data from the PWR subchannel and bundle tests (PSBT) benchmark is first used to train an artificial neural network (ANN) to predict the DNB power, which produces a root mean square error (RMSE) of 6.89 kW/m when tested on a blind subset of the PSBT data. Since the PSBT dataset is relatively small to train an accurate ANN, a data augmentation methodology based on generative adversarial networks (GANs) is used to expand the training dataset. By assuming that the real data follows a certain distribution, GANs try to learn that underlying distribution to generate similar synthetic data to augment the database and to improve the predictive capabilities of the ANN. The data generated from GANs are validated using 1-nearest neighbor and kernel maximum mean discrepancy. To further ensure data from GAN is similar to PSBT, the data is tested and filtered out using the sub-channel thermal-hydraulic code CTF. The results indicate that with the addition of 120 data points from GAN the RMSE reduces to 4.84 kW/m showing promising results for future developments.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Training material models using gradient descent algorithms

High temperature design requires accurate constitutive models to describe material inelastic deformation and failure behavior. Oftentimes, calibrating accurate models devolves into the problem of fitting the model parameters against experimental test data. Here, we present the pyopmat package, an open source framework for calibrating constitutive models against experiment data subjected to various loading conditions using machine learning techniques. The package calculates the exact gradient of the model response with respect to the parameters using a combination of automatic differentiation and the adjoint method. Given this exact gradient, we compare the performance of several gradient-based optimization techniques in fitting realistic constitutive models against data. Here, we demonstrate the efficiency and accuracy of our package through example problems using both synthetic data, generated using known parameter sets, under monotonic and cyclic loading conditions and also with an example applying the techniques developed here to actual high temperature creep-fatigue test data.

36 MATERIALS SCIENCE↗

Dark Energy Survey Year 3 results: Cosmology from cosmic shear and robustness to modeling uncertainty

Here, this work and its companion paper, Amon et al. [Phys. Rev. D 105, 023514 (2022)], present cosmic shear measurements and cosmological constraints from over 100 million source galaxies in the Dark Energy Survey (DES) Year 3 data. We constrain the lensing amplitude parameter 𝑆 8 ≡𝜎 8 ⁢$\sqrt{Ω_{m}/0.3}$ at the 3% level in Λ⁢ CDM: 𝑆 8 =0.75⁢9$^{+0.025}_{−0.023}$ (68% CL). Our constraint is at the 2% level when using angular scale cuts that are optimized for the Λ⁢ CDM analysis: 𝑆 8 =0.77⁢2$^{+0.018}_{−0.017}$ (68% CL). With cosmic shear alone, we find no statistically significant constraint on the dark energy equation-of-state parameter at our present statistical power. We carry out our analysis blind, and compare our measurement with constraints from two other contemporary weak lensing experiments: the Kilo-Degree Survey (KiDS) and Hyper-Suprime Camera Subaru Strategic Program (HSC). We additionally quantify the agreement between our data and external constraints from the Cosmic Microwave Background (CMB). Our DES Y3 result under the assumption of Λ⁢ CDM is found to be in statistical agreement with Planck 2018, although favors a lower 𝑆 8 than the CMB-inferred value by 2.3⁢𝜎 (a 𝑝-value of 0.02). This paper explores the robustness of these cosmic shear results to modeling of intrinsic alignments, the matter power spectrum and baryonic physics. We additionally explore the statistical preference of our data for intrinsic alignment models of different complexity. The fiducial cosmic shear model is tested using synthetic data, and we report no biases greater than 0.3⁢𝜎 in the plane of 𝑆 8 ×Ω m caused by uncertainties in the theoretical models.

79 ASTRONOMY AND ASTROPHYSICS↗

Fast Gaussian Process Estimation for Large-Scale In Situ Inference using Convolutional Neural Networks

Exascale computing will bring with it significant I/O limitations. One foreseeable consequence of such restrictions is that the user can save only a small fraction of complex simulation data to disk for subsequent analysis. An alternative is to fit statistical models to data in situ, that is, inside the simulation as it runs. This option requires extremely fast statistical estimation to avoid slowing down the simulation. Gaussian processes (GPs) have state-of-the-art predictive performance for modeling spatial data. However, standard estimation methods for GPs scale quite poorly to large data sets as parameter estimation requires inverting a covariance matrix to the size of the data set. In the presented work, we use a convolutional neural network (CNN) to predict the GP parameters for a spatial data set, from a simulation or otherwise, rather than optimize the parameters directly. Here, our presented case study models spatial data from E3SM, the Department of Energy’s Exascale climate model. The CNN is trained on synthetic data simulated from GP models with known parameters and then applied to data from the climate simulation. In the presented examples, the neural network scheme produces parameter estimates that compare well with standard methods such as maximum likelihood estimation in predictive performance but is obtained four orders of magnitude faster.

big data↗

Synthetic spectra for Lyman- α forest analysis in the Dark Energy Spectroscopic Instrument

Synthetic data sets are used in cosmology to test analysis procedures, to verify that systematic errors are well understood and to demonstrate that measurements are unbiased. In this work we describe the methods used to generate synthetic datasets of Lyman-α quasar spectra aimed for studies with the Dark Energy Spectroscopic Instrument (DESI). In particular, we focus on demonstrating that our simulations reproduces important features of real samples, making them suitable to test the analysis methods to be used in DESI and to place limits on systematic effects on measurements of Baryon Acoustic Oscillations (BAO). We present a set of mocks that reproduce the statistical properties of the DESI early data set with good agreement. Additionally, we use a synthetic dataset to forecast the BAO scale constraining power of the completed DESI survey through the Lyman-α forest.

79 ASTRONOMY AND ASTROPHYSICS↗