Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

DeFault: DEep‐Learning‐Based FAULT Delineation Using the IBDP Passive Seismic Data at the Decatur CO2 Storage Site

Abstract The carbon capture, utilization, and storage (CCUS) framework is an essential component in reducing greenhouse gas emissions, with its success hinging on the comprehensive knowledge of subsurface geology and geomechanics. Passive seismic event relocation and fault detection offer vital insights into subsurface structures and the ability to monitor fluid migration pathways. Accurate identification and localization of seismic events, however, face significant challenges, including the necessity for high‐quality seismic data and advanced computational methods. To address these challenges, we introduce a novel deep learning method, , specifically designed for passive seismic source relocation and fault delineating for passive seismic monitoring projects. By leveraging data domain‐adaptation, allows us to train a neural network with labeled synthetic data and apply it directly to field data. Using , the passive seismic sources are automatically clustered based on their recording time and spatial locations, and subsequently, faults and fractures are delineated accordingly. We demonstrate the efficacy of on a field case study involving injection related microseismic data from Decatur, Illinois area. Our approach accurately and efficiently relocated passive seismic events, identified faults and could aid in potential damage induced by seismicity. Our results highlight the potential of as a valuable tool for passive seismic monitoring, emphasizing its role in ensuring CCUS project safety. This research bolsters the understanding of subsurface characterization in CCUS, illustrating machine learning’s capacity to refine these methods. Ultimately, our work has significant implications for CCUS technology deployment, an essential strategy in combating climate change. Plain Language Summary In our quest to tackle climate change, we use a strategy known as carbon capture, utilization, and storage (CCUS) to keep greenhouse gases out of the atmosphere. This strategy relies heavily on our ability to understand what's happening deep under the earth's surface. To make sure we store super critical safely, we need to accurately map out the geological structure, especially faults, but this is tough without high‐quality data and complex computer programs. We've developed a new tool called “DeFault,” which uses advanced machine learning to improve how we find and map these underground features. “DeFault” is smart enough to learn from numerically simulated data and then apply what it’s learned to real‐world situations. It groups together seismic activity—tiny tremors and shifts in the earth—based on when and where they happen, which helps us spot where there might be cracks or faults. We tested “DeFault” in Illinois, where CO 2 is injected underground, and it successfully pinpointed where these tremors occurred and mapped out the faults, helping to prevent accidents accurately in the future. Our study shows that “DeFault” will be a powerful ally in making CCUS safer and more effective, especially for the Illinois Basin Decatur Project. Key Points Faults and fractures introduced by carbon storage can be monitored by passive seismicity DeFault algorithm enables an automatic process for accurate and efficient passive seismic event locating and clustering

58 GEOSCIENCES↗

MPACT Safeguards Modeling: FY25 Update

Sandia National Laboratories develops and maintains several open-source software packages to support material accountancy analyses. This includes the Material Accountancy Performance Indicator Toolkit (MAPIT), the Fissile Facility Flow Modeler (F3M) and the Separation and Safeguards Performance Model Library (SSPM-L). MAPIT is responsible for performing statistical safeguards analyses on bulk and itemized data from nuclear fuel cycle facilities and can operate on real or synthetic data. MAPIT is the only open-source software for such analyses. F3M is a library of modules, built in MATLAB Simulink, that contain pre made blocks to represent different generic fuel cycle processes. These blocks can be used together in a modular fashion to represent and simulate nuclear fuel cycle processes with the goal of improving facility-level accountancy during the design phase. F3M is also an open-source library. Finally, the SSPM-L library is a series of completed models built from F3M. The library includes facility models such as a generic PUREX facility and a fuel fabrication facility. The SSPM-L library is not open source, but is available to collaborators with a relevant use case. These tools include modeling and simulation pipelines to simulate nuclear fuel cycle facilities and the underlying software needed to simulate measurement uncertainty and perform statistical analyses. Together, these tools can perform end-to-end nuclear material accountancy analyses. This report documents the various improvements made to these tools in FY25. Specifically, we added new statistical test, new statistical modeling capabilities, new fuel cycle facility models, and launched a new open-source model component library.

97 MATHEMATICS AND COMPUTING↗

Analysis of data acquired by synthetic aperture radar over Dade County, Florida, and Acadia Parish, Louisiana

Results of digital processing of airborne X-band synthetic aperture radar (SAR) data acquired over Dade County, Florida, and Acadia Parish, Louisiana are presented. The goal was to investigate the utility of SAR data for land cover mapping and area estimation under the AgRISTARS Domestic Crops and Land Cover Project. In the case of the Acadia Paris study area, LANDSAT multispectral scanner (MSS) data were also used to form a combined SAR and MSS data set. The results of accuracy evaluation for the SAR, MSS, and SAR/MSS data using supervised classification show that the combined SAR/MSS data set results in an improved classification accuracy of the five land cover classes as compared with SAR-only and MSS-only data sets. In the case of the Dade County study area, the results indicate that both HH and VV polarization data are highly responsive to the row orientation of the row crop but not to the specific vegetation which forms the row structure. On the other hand, the HV polarization data are relatively insensitive to the orientation of row crop. Therefore, the HV polarization data may be used to discriminate the specific vegetation that forms the row structure.

Wu, S. T.↗

Generation and representation of synthetic smart meter data

Advanced energy algorithms running at big-data scale will be necessary to identify, realize, and verify energy savings to meet government and utility goals of building energy efficiency. Any algorithm must be well characterized and validated before it is trusted to run at these scales. Smart meter data from real buildings will ultimately be required for the development, testing, and validation of these energy algorithms and processes. However, for initial development and testing, smart meter data are difficult to work with due to privacy restrictions, noise from unknown sources, data accessibility, and other concerns which can complicate algorithm development and validation. This paper describes a new methodology to generate synthetic smart meter data of electricity use in buildings using detailed building energy modeling, which aims to capture the variability and stochastics of real energy use in buildings. The methodology can create datasets tailored to represent specific scenarios with known truth and controllable amounts of synthetic noise. Knowledge of ground truth also allows the development and validation of enhanced processes which leverage building metadata, such as building type or size (floor area), in addition to smart meter data. The methodology described in this paper includes the key influencing factors of real-world building energy use including weather data, occupant-driven loads, building operation and maintenance practices, and special events. Data formats to support workflows leveraging both synthetic meter data and associated metadata are proposed and discussed. Finally, example use cases of the synthetic meter data are described to illustrate potential applications.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

SEASAT synthetic-aperture radar data user's manual

The SEASAT Synthetic-Aperture Radar (SAR) system, the data processors, the extent of the image data set, and the means by which a user obtains this data are described and the data quality is evaluated. The user is alerted to some potential problems with the existing volume of SEASAT SAR image data, and allows him to modify his use of that data accordingly. Secondly, the manual focuses on the ultimate focuses on the ultimate capabilities of the raw data set and evaluates the potential of this data for processing into accurately located, amplitude-calibrated imagery of high resolution. This allows the user to decide whether his needs require special-purpose data processing of the SAR raw data.

Pravdo, S. H.↗

Discovery of Activities via Statistical Clustering of Fixation Patterns

Human behavior often consists of a series of distinct activities, each characterized by a unique pattern of interaction with the visual environment. This is true even in a restricted domain, such as a piloting an aircraft, where activities with distinct visual signatures might be things like communicating, navigating, and monitoring. We propose a novel analysis method for gaze-tracking data, to perform blind discovery of these hypothetical activities. The method is in some respects similar to recurrence analysis, but here we compare not individual fixations, but groups of fixations aggregated over a fixed time interval. The duration of this interval is a parameter that we will refer to as delta. We assume that the environment has been divided into a set of N different areas-of-interest (AOIs). For a given interval of time of duration delta, we compute the proportion of time spent fixating each AOI, resulting in an N-dimensional vector. These proportions can be converted to integer counts by multiplying by delta divided by the average fixation duration (another parameter that we fix at 280 milliseconds). We compare different intervals by computing the chi-square statistic. The p-value associated with the statistic is the likelihood of observing the data under the hypothesis that the data in the two intervals were generated by a single process with a single set of probabilities governing the fixation of each AOI. The method has been applied to approximately 100 hours of eye movement data collected from pilots in a high-fidelity B747 flight simulator, and the results have been compared to synthetic data in which the each activity is represented as first-order Markov process with random probabilities assigned to the AOIs. Randomly-generated synthetic activities can require thousands of fixations to be discriminated with statistical significance, while the human data can be clustered using averaging windows of some 10's of seconds, suggesting that the actual activities are much more narrowly focused than random Markov models.

activity analysis↗

Explosion Discrimination Using Seismic Gradiometry and Spectral Filtering of Data

Here, we present a new method to discriminate between earthquakes and buried explosions using observed seismic data. The method is different from previous seismic discrimination algorithms in two main ways. First, we use seismic spatial gradients, as well as the wave attributes estimated from them (referred to as gradiometric attributes), rather than the conventional three-component seismograms recorded on a distributed array. The primary advantage of this is that a gradiometer is only a fraction of a wavelength in aperture compared with a conventional seismic array or network. Second, we use the gradiometric attributes as input data into a machine learning algorithm. The resulting discrimination algorithm uses the norms of truncated principal components obtained from the gradiometric data to distinguish the two classes of seismic events. Using high-fidelity synthetic data, we show that the data and gradiometric attributes recorded by a single seismic gradiometer performs as well as a conventional distributed array at the event type discrimination task.

58 GEOSCIENCES↗

Polyconvex neural network models of thermoelasticity

Machine-learning function representations such as neural networks have proven to be excellent constructs for constitutive modeling due to their flexibility to represent highly nonlinear data and their ability to incorporate constitutive constraints, which also allows them to generalize well to unseen data. Here, in this work, we extend a polyconvex hyperelastic neural network framework to (isotropic) thermo-hyperelasticity by specifying the thermodynamic and material theoretic requirements for an expansion of the Helmholtz free energy expressed in terms of deformation invariants and temperature. Different formulations which a priori ensure polyconvexity with respect to deformation and concavity with respect to temperature are proposed and discussed. The physics-augmented neural networks are furthermore calibrated with a recently proposed sparsification algorithm that not only aims to fit the training data but also penalizes the number of active parameters, which prevents overfitting in the low data regime and promotes generalization. The performance of the proposed framework is demonstrated on synthetic data, which illustrate the expected thermomechanical phenomena, and existing temperature-dependent uniaxial tension and tension-torsion experimental datasets.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Sparse-Data Deep Learning Strategies for Radiographic Non-Destructive Testing

Radiography is an imaging technique used in a variety of applications, such as medical diagnosis, airport security, and nondestructive testing. We present a deep learning system for extracting information from radiographic images. We perform various prediction tasks using our system, including material classification and regression on the dimensions of a given object that is being radiographed. Our system is designed to address the sparse-data issue for radiographic nondestructive testing applications. It uses a radiographic simulation tool for synthetic data augmentation, and it uses transfer learning with a pre-trained convolutional neural network model. Using this system, our preliminary results indicate that the object geometry regression task saw an improvement of 70% in the R-squared value when using a multi-regime model. In addition, we increase the performance of the object material classification tasks by utilizing data from different imaging systems. In particular, using neutron imaging improved the material classification accuracy by 20% when compared to x-ray imaging.

convolutional neural networks↗

Unified Wind-Wave Growth and Spectrum Functions for All Water Depths: Field Observations and Model Results

Abstract Wind-wave development is governed by the fetch- or duration-limited growth principle that is expressed as a pair of similarity functions relating the dimensionless elevation variance (wave energy) and spectral peak frequency to fetch or duration. Combining the pair of similarity functions, the fetch or duration variable can be removed to form a dimensionless function of elevation variance and spectral peak frequency, which is interpreted as the wave energy evolution with wave age. The relationship is initially developed for quasi-neural stability and quasi-steady wind forcing conditions. Further analyses show that the same fetch, duration, and wave-age similarity functions are applicable to unsteady wind forcing conditions, including rapidly accelerating and decelerating mountain gap wind episodes and tropical cyclone (TC) wind fields. Here it is shown that with the dimensionless frequency converted to dimensionless wavenumber using the surface wave dispersion relationship, the same similarity function is applicable in all water depths. Field data collected in shallow to deep waters and mild to TC wind conditions and synthetic data generated by spectrum model computations are assembled to illustrate the applicability. For the simulation work, the finite-depth wind-wave spectrum model and its shoaling function are formulated for variable spectral slopes. Given wind speed, wave age, and water depth, the measured and spectrum-computed significant wave heights and the associated growth parameters are in good agreement in forcing conditions from mild to TC winds and in all depths from deep ocean to shallow lake. Significance Statement This paper presents a growth function and spectrum model to describe wind-wave development in all water depths. Their applicability covers a wide range of wind forcing conditions including steady, accelerating, decelerating, and tropical cyclone events. Support for the unified spectrum model and growth function is presented with field observations and numerical computations.

Hwang, Paul A.↗

A general rough-surface inversion algorithm: Theory and application to SAR data

Rough-surface inversion has significant applications in interpretation of SAR data obtained over bare soil surfaces and agricultural lands. Due to the sparsity of data and the large pixel size in SAR applications, it is not feasible to carry out inversions based on numerical scattering models. The alternative is to use parameter estimation techniques based on approximate analytical or empirical models. Hence, there are two issues to be addressed, namely, what model to choose and what estimation algorithm to apply. Here, a small perturbation model (SPM) is used to express the backscattering coefficients of the rough surface in terms of three surface parameters. The algorithm used to estimate these parameters is based on a nonlinear least-squares criterion. The least-squares optimization methods are widely used in estimation theory, but the distinguishing factor for SAR applications is incorporating the stochastic nature of both the unknown parameters and the data into formulation, which will be discussed in detail. The algorithm is tested with synthetic data, and several Newton-type least-squares minimization methods are discussed to compare their convergence characteristics. Finally, the algorithm is applied to multifrequency polarimetric SAR data obtained over some bare soil and agricultural fields. Results will be shown and compared to ground-truth measurements obtained from these areas. The strength of this general approach to inversion of SAR data is that it can be easily modified for use with any scattering model without changing any of the inversion steps. Note also that, for the same reason it is not limited to inversion of rough surfaces, and can be applied to any parameterized scattering process.

Moghaddam, M.↗

Reliable Measures of Spread in High Dimensional Latent Spaces

Understanding geometric properties of the latent spaces of natural language processing models allows the manipulation of these properties for improved performance on downstream tasks. One such property is the amount of data spread in a model’s latent space, or how fully the available latent space is being used. We demonstrate that the commonly used measures of data spread, average cosine similarity and a partition function min/max ratio I (V), do not provide reliable metrics to compare the use of latent space across data distributions. We propose and examine six alternative measures of data spread, all of which improve over these current metrics when applied to seven synthetic data distributions. Of our proposed measures, we recommend one principal component-based measure and one entropy-based measure that provide reliable, relative measures of spread and can be used to compare models of different sizes and dimensionalities.

97 MATHEMATICS AND COMPUTING↗

Airborne hyperspectral imaging of cover crops through radiative transfer process-guided machine learning

Cover cropping between cash crop growing seasons is a multifunctional conservation practice. Timely and accurate monitoring of cover crop traits, notably aboveground biomass and nutrient content, is beneficial to agricultural stakeholders to improve management and understand outcomes. Currently, there is a scarcity of spatially and temporally resolved information for assessing cover crop growth. Remote sensing has a high potential to fill this need, but conventional empirical regression operated with coarse-resolution multispectral data has large uncertainties. Therefore, this study utilized airborne hyperspectral imaging techniques and developed new process-guided machine learning approaches (PGML) for cover crop monitoring. Specifically, we deployed an airborne hyperspectral system covering visible to shortwave-infrared wavelengths (400–2400 nm) to acquire high spatial (0.5 m) and spectral (3–5 nm) resolution reflectance over 23 cover crop fields across Central Illinois in March and April of 2021. Airborne hyperspectral surface reflectance with high spectral and spatial resolution can be well matched with field data to quantify cover crop traits. Furthermore, the PGML models were pre-trained by synthetic data from soil-vegetation radiative transfer modeling (one million records), and then fine-tuned with field data of cover crop biomass and nutrient content. Results show that airborne hyperspectral data with PGML can achieve high accuracy to predict cover crop aboveground biomass (R 2 = 0.72, relative RMSE = 15.16%) and nitrogen content (R 2 = 0.69, relative RMSE = 16.59%) through leave-one-field-out cross-validation. Unlike the pure data-driven approach (e.g., partial least-squares regression), PGML incorporated radiative transfer knowledge and obtained higher predictive performance with fewer field data. Meanwhile, with field data for model fine-tuning, PGML predicted biomass more accurately than the inversion of radiative transfer models. Here we also found that the red edge has a high contribution in quantifying aboveground biomass and nitrogen content, followed by green and shortwave spectra. This study demonstrated the first attempt of utilizing hyperspectral remote sensing to accurately quantify cover crop traits. We highlight the strength of PGML in exploiting sensing data to quantify ecosystem variables to advance agroecosystem monitoring for sustainable agricultural management.

60 APPLIED LIFE SCIENCES↗

Subspace-Driven Learning for Anomaly Detection in Process Transients

Nuclear power plant (NPP) monitoring and diagnostic centers are actively investigating and implementing automated anomaly detection algorithms to help plants catch anomalies sooner, thereby preventing or reducing the duration of unexpected shutdowns. Current machine learning-based anomaly detection methods are expected to be highly effective during stable, full-power operations because NPPs typically operate as baseload power generators, meaning there are extensive operating data available from plant equipment. However, it is expected that anomaly detection methods will face significant challenges during transient conditions (i.e., when power output falls below full power) because plants only occasionally operate at these lower power levels, generating sparse transient operational data, and resulting in false alarms or missed detections. Here, to address this issue, transfer learning is used, which for this problem leverages knowledge (in the form of learned features) from stable, full-power operations to improve detection accuracy during transient conditions, even with limited data. In this effort, a novel subspace approach is developed to transfer a subset of the data features from full power operation to transients. This approach is validated through experiments using synthetic data and was found to outperform two baseline transfer learning approaches in anomaly detection performance across a range of amounts of transient data used in the training process.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Direct structural retrieval from gas-phase ultrafast diffraction data using a genetic algorithm

Ultrafast scattering techniques such as ultrafast electron diffraction and ultrafast x-ray diffraction have been utilized to elucidate the structural dynamics, reaction intermediates, and final products in molecular reactions following photoexcitation. The time-dependent structures are typically not directly retrieved from the experimental data, but they rely on comparison with calculations. The genetic algorithm (GA), a global optimization strategy, can be used to retrieve the molecular structures directly from diffraction patterns without any theoretical input. However, the robustness of the GA with respect to real experimental conditions such as a limited momentum transfer range, noise, and artifacts has not been studied in detail. In this work, we characterize the performance of the GA with simulated data that mimic realistic experimental conditions. We have developed and implemented a variant of the GA specific to diffraction measurements which performs better in the presence of imperfect data compared to the standard implementation of the GA. We demonstrate this method with both synthetic data and experimental ultrafast electron diffraction data on the UV-induced photodissociation of trifluoroiodomethane (C⁢F 3⁡ I) molecules.

74 ATOMIC AND MOLECULAR PHYSICS↗

Analytical Modeling of Exoplanet Transit Spectroscopy with Dimensional Analysis and Symbolic Regression

Abstract The physical characteristics and atmospheric chemical composition of newly discovered exoplanets are often inferred from their transit spectra, which are obtained from complex numerical models of radiative transfer. Alternatively, simple analytical expressions provide insightful physical intuition into the relevant atmospheric processes. The deep-learning revolution has opened the door for deriving such analytical results directly with a computer algorithm fitting to the data. As a proof of concept, we successfully demonstrate the use of symbolic regression on synthetic data for the transit radii of generic hot-Jupiter exoplanets to derive a corresponding analytical formula. As a preprocessing step, we use dimensional analysis to identify the relevant dimensionless combinations of variables and reduce the number of independent inputs, which improves the performance of the symbolic regression. The dimensional analysis also allowed us to mathematically derive and properly parameterize the most general family of degeneracies among the input atmospheric parameters that affect the characterization of an exoplanet atmosphere through transit spectroscopy.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine Learning-based Prediction of Departure from Nucleate Boiling Power for the PSBT Benchmark

Machine Learning (ML) has seen an exponential growth in its applications due to its advanced data driven prediction capabilities. The study presents a data-driven approach as a preliminary attempt to predict the power at which departure from nucleate boiling (DNB) occurs in pressurized water reactors (PWRs) by constructing an advanced ML algorithm that takes outlet pressure, inlet temperature and inlet mass flux as the input features. DNB is a critical heat flux (CHF) phenomenon seen in PWRs. The experimental data from the PWR subchannel and bundle tests (PSBT) benchmark is first used to train an artificial neural network (ANN) to predict the DNB power, which produces a root mean square error (RMSE) of 6.89 kW/m when tested on a blind subset of the PSBT data. Since the PSBT dataset is relatively small to train an accurate ANN, a data augmentation methodology based on generative adversarial networks (GANs) is used to expand the training dataset. By assuming that the real data follows a certain distribution, GANs try to learn that underlying distribution to generate similar synthetic data to augment the database and to improve the predictive capabilities of the ANN. The data generated from GANs are validated using 1-nearest neighbor and kernel maximum mean discrepancy. To further ensure data from GAN is similar to PSBT, the data is tested and filtered out using the sub-channel thermal-hydraulic code CTF. The results indicate that with the addition of 120 data points from GAN the RMSE reduces to 4.84 kW/m showing promising results for future developments.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Correction of instrumental distortion by analytical deconvolution of data

A general analytical theorem developed by van de Hulst (1946) for inverting the convolution integral is reviewed and illustrated both with synthetic data and with experimental data from time-of-flight measurements. If the undesired influence of an instrument used in an experimental measurement can be represented by the convolution integral, the original undistorted or true distribution may sometimes be recovered in postprocessing the data by means of deconvolution. Analytical deconvolution is achieved by using the coefficients from a power series representation of the distorted output distribution and a set of 'solving polynomials' which may be readily derived from the response function of the instrument.

Morton, D. C.↗