Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Expanded analysis of machine learning models for nuclear transient identification using TPOT

Industries around the world are becoming more and more data driven. The nuclear field is no exception with several different applications being proposed. One popular area of research is the use of machine learning in transient detection. This paper seeks to build upon a previous study which made use of the AutoML package TPOT to train traditional machine learning models to classify transient events occurring with a reactor. Synthetic data was once again collected using a GPWR reactor simulator. Data on 12 different events was collected using 15 different initial conditions. Here, a dataset consisting of over 100,000 data points was compiled and used to train 7 different machine learning models using a pre-defined TPOT dictionary with 12 different preprocessing techniques. Three of the trained models were able to produce validation results in the 90s with the expanded dataset. Once the models were trained, it was possible to look into where during the simulation, misclassifications occurred. Using these three models, analysis was done to determine if TPOT could be used to train models that were effective if important features were missing. The results from this were positive with the newly trained models scoring close to the original models. Finally, to conclude this study, the three high performing models were retrained using different random states to see if there was any major variation when different states were used.

42 ENGINEERING↗

Analysis of data acquired by synthetic aperture radar and LANDSAT Multispectral Scanner over Kershaw County, South Carolina, during the summer season

Data acquired by synthetic aperture radar (SAR) and LANDSAT multispectral scanner (MSS) were processed and analyzed to derive forest-related resources inventory information. The SAR data were acquired by using the NASA aircraft X-band SAR with linear (HH, VV) and cross (HV, VH) polarizations and the SEASAT L-band SAR. After data processing and data quality examination, the three polarization (HH, HV, and VV) data from the aircraft X-band SAR were used in conjunction with LANDSAT MSS for multisensor data classification. The results of accuracy evaluation for the SAR, MSS and SAR/MSS data using supervised classification show that the SAR-only data set contains low classification accuracy for several land cover classes. However, the SAR/MSS data show that significant improvement in classification accuracy is obtained for all eight land cover classes. These results suggest the usefulness of using combined SAR/MSS data for forest-related cover mapping. The SAR data also detect several small special surface features that are not detectable by MSS data.

Wu, S. T.↗

Inferring microbial co-occurrence networks from amplicon data: a systematic evaluation

Microbes commonly organize into communities consisting of hundreds of species involved in complex interactions with each other. 16S ribosomal RNA (16S rRNA) amplicon profiling provides snapshots that reveal the phylogenies and abundance profiles of these microbial communities. These snapshots, when collected from multiple samples, can reveal the co-occurrence of microbes, providing a glimpse into the network of associations in these communities. However, the inference of networks from 16S data involves numerous steps, each requiring specific tools and parameter choices. Moreover, the extent to which these steps affect the final network is still unclear. In this study, we perform a meticulous analysis of each step of a pipeline that can convert 16S sequencing data into a network of microbial associations. Through this process, we map how different choices of algorithms and parameters affect the co-occurrence network and identify the steps that contribute substantially to the variance. We further determine the tools and parameters that generate robust co-occurrence networks and develop consensus network algorithms based on benchmarks with mock and synthetic data sets. The Microbial Co-occurrence Network Explorer, or MiCoNE (available at https://github.com/segrelab/MiCoNE) follows these default tools and parameters and can help explore the outcome of these combinations of choices on the inferred networks. We envisage that this pipeline could be used for integrating multiple data sets and generating comparative analyses and consensus networks that can guide our understanding of microbial community assembly in different biomes.

16S rRNA↗

The Dark Energy Spectroscopic Instrument: one-dimensional power spectrum from first Ly α forest samples with Fast Fourier Transform

ABSTRACT We present the one-dimensional Ly α forest power spectrum measurement using the first data provided by the Dark Energy Spectroscopic Instrument (DESI). The data sample comprises 26 330 quasar spectra, at redshift z > 2.1, contained in the DESI Early Data Release and the first 2 months of the main survey. We employ a Fast Fourier Transform (FFT) estimator and compare the resulting power spectrum to an alternative likelihood-based method in a companion paper. We investigate methodological and instrumental contaminants associated with the new DESI instrument, applying techniques similar to previous Sloan Digital Sky Survey (SDSS) measurements. We use synthetic data based on lognormal approximation to validate and correct our measurement. We compare our resulting power spectrum with previous SDSS and high-resolution measurements. With relatively small number statistics, we successfully perform the FFT measurement, which is already competitive in terms of the scale range. At the end of the DESI survey, we expect a five times larger Ly α forest sample than SDSS, providing an unprecedented precise one-dimensional power spectrum measurement.

79 ASTRONOMY AND ASTROPHYSICS↗

A physical model for predicting bidirectional reflectances over bare soil

While most previous attempts to retrieve soil surface optical property characteristics have proceeded through a fitting of empirical functions to the data, an optimization technique is presently applied to a physically-based surface reflectance model developed for the study of planetary surfaces. This inversion procedure is shown to allow the direct estimation of the single-scattering coefficient, two parameters describing the 'hot spot' phenomenon, and two parameters describing the scattering phase function. A comparison of inversion technique results with both synthetic data and actual observations shows the model to be capable of predicting the observed bidirectional reflectances as well as directional-hemispherical reflectances; it can also build the complete radiance field over the upward hemisphere.

Pinty, Bernard↗

Bond Line Thickness Estimation in Composite Structures Using Multiple Inspection Techniques

Imaging and other nondestructive evaluation techniques are commonly used for material characterization and defect recognition in safety critical aerospace applications, with data fusion providing the framework for uncertainty quantification in these contexts. Most commonly, forward physics-based modeling predicts the response conditioned on material properties and defect assumptions, and probabilistic methods are used to infer the hidden state of the subject of the inspection from a combination of prior information, likelihoods, and inspection data. In this paper Bayesian methods are used to estimate bond thickness in lap joints comprised of aluminum adherends using a combination of infrared thermography and ultrasound. The concept of the conflation of probability distributions is applied to combine the posterior distributions derived from thermography and ultrasound and the quality of the fused estimates are compared against the individual estimates against synthetic data that was created to mimic the inspection of a lap joint comprised of aluminum adherends.

thermal nondestructive evaluation↗

Classification and Localization of Fracture-Hit Events in Low-Frequency Distributed Acoustic Sensing Strain Rate with Convolutional Neural Networks

Summary Distributed acoustic sensing (DAS) has been used in the oil and gas industry as an advanced technology for surveillance and diagnostics. Operators use DAS to monitor hydraulic fracturing activities, examine well stimulation efficacy, and estimate complex fracture system geometries. Particularly, low-frequency DAS can detect geomechanical events such as fracture hits because hydraulic fractures propagate and create strain rate variations in the rock. Analysis of DAS data today is mostly done post-job and subject to interpretation methods. However, the continuous and dense data stream generated live by DAS poses the opportunity for more efficient and accurate real-time data-driven analysis. The objective of this study is to develop a machine learning-based workflow that can identify and locate fracture-hit events in simulated strain rate responses correlated with low-frequency DAS data. In this paper, “fracture hit” refers to a hydraulic fracture originating from a stimulated well intersecting an offset well. We start with building a single fracture propagation model to produce strain rate patterns observed at a hypothetical monitoring well. This model is used to generate two sets of strain rate responses with one set containing fracture-hit events. The labeled synthetic data are then used to train a custom convolutional neural network (CNN) model for identifying the presence of fracture-hit events. The same model is trained again for locating the event with the output layer of the model replaced with linear units. We achieved near-perfect predictions for both event classification and localization. These promising results prove the feasibility of using CNN for real-time event detection from fiber-optic sensing data. Additionally, we use edge detection techniques to recognize fracture-hit event patterns in strain rate images. The fracture-hit location can be identified using recognized pixels in the image. The accuracy of edge detection-based location identification is also plausible, but edge detection is dependent on the assumption of pattern shape and image quality, hence it is less robust compared to CNN models. This comparison further supports the need for CNN applications in image-based real-time fiber-optic sensing event detection.

Engineering↗

Knowledge-guided graph machine learning for spatially distributed prediction of daily discharge and nitrogen export dynamics

Spatially distributed prediction of streamflow and nitrogen export dynamics is essential for precision management of agricultural watersheds. While temporal deep learning models such as Long Short-Term Memory (LSTM) have shown strong performance at basin scales, their ability to generalize spatially is limited by insufficient representation of spatial dependencies and flow paths, particularly under data-scarce conditions. To address this gap, we propose HydroGraphNet, a knowledge-guided graph machine learning framework that integrates process-based knowledge and explicit spatial learning into temporal modeling. This framework incorporates directed graph topology to encode watershed connectivity and upstream inflows, with mass balance constraints to improve physical consistency. To enhance generalization in sparsely monitored regions, HydroGraphNet is pretrained on synthetic data generated by the SWAT+ (Soil and Water Assessment Tool Plus) model. We evaluated HydroGraphNet in the Upper Sangamon River Basin (44 HUC-12 subwatersheds, 2001–2020) against two LSTM baselines: a lumped basin-level model and a distributed variant. When benchmarked on SWAT+ simulations in pretraining, HydroGraphNet improved test NSEs by 8.9% (discharge) and 13.7% (NO₃–N load) in temporal extrapolation, and by 27.1% and 34.7% in spatial extrapolation, relative to the Lumped LSTM baseline. After fine-tuning with USGS monitoring data, the model achieved mean test NSE (KGE) scores of 0.768 (0.861) for discharge and 0.626 (0.664) for NO₃–N load, substantially outperforming baselines. Attribution analysis further highlighted the importance of upstream inflow representation and graph-based spatial learning in capturing cross-subwatershed dependencies. The model also reproduced seasonal hydrological and biogeochemical patterns consistent with known processes, demonstrating its robustness and process fidelity for spatially distributed prediction. Altogether, HydroGraphNet advances the integration of physical knowledge and spatially explicit learning in hydrological modeling, offering a generalizable framework for distributed modeling to support spatially targeted water quality management in data-scarce watersheds.

54 ENVIRONMENTAL SCIENCES↗

Discovery of Activities via Statistical Clustering of Fixation Patterns

Human behavior often consists of a series of distinct activities, each characterized by a unique pattern of interaction with the visual environment. This is true even in a restricted domain, such as a pilot flying an airplane; in this case, activities with distinct visual signatures might be things like communicating, navigating, monitoring, etc. We propose a novel analysis method for gaze-tracking data, to perform blind discovery of these hypothetical activities. We compare, not individual fixations, but groups of fixations aggregated over a fixed time interval (Tau). We assume that the environment has been divided into a finite set of discrete areas-of-interest (AOIs). For a given time interval, we compute the proportion of time spent fixating each AOI, resulting in an N-dimensional vector, where N is the number of AOIs. These proportions can be converted to integer counts by multiplying by Tau divided by the average fixation duration, a parameter that we fix at 283 milliseconds. We compare different intervals by computing the chi-squared statistic. The p-value associated with the statistic is the likelihood of observing the data under the hypothesis that the data in the two intervals were generated by a single process with a single set of probabilities governing the fixation of each AOI. We cluster the intervals, first by merging adjacent intervals that are sufficiently similar, optionally shifting the boundary between non-merged intervals to maximize the difference. Then we compare and cluster non-adjacent intervals. The method is evaluated using synthetic data generated by a hand-crafted set of activities. While the method generally finds more activities than put into the simulation, we have obtained agreement as high as 80 percent between the inferred activity labels and ground truth.

Eye Movements↗

Regularizing INR with Diffusion Prior for Self-Supervised 3D Reconstruction OF Neutron Computed Tomography Data

Recently, generative diffusion priors have made huge strides as inverse problem solvers, including the ability to be adapted for inference on out-of-distribution data. Concurrently, implicit neural representations (INRs) have emerged as fast and lightweight inverse imaging solvers that are amenable to hybrid approaches that combine learned priors with traditional inverse problem formulations. In this paper, we present a diffusive computed tomography (CT) inversion framework for regularizing INRs called Diffusive INR (DINR), designed to enable high-quality reconstruction from sparse-view neutron CT. Pretrained purely on synthetic data, DINR is evaluated on simulated and experimentally obtained observations of concrete microstructures, where traditional reconstruction methods suffer substantial degradation when the number of views is reduced. Our approach delivers superior performance, reduces reconstruction artifacts, and achieves gains in PSNR and SSIM, enabling accurate micro-structural characterization even under extreme data limitations compared to state-of-the-art sparse-view reconstruction techniques.

Hossain, Maliha [ORNL]↗

Precise Dynamical Masses and Orbital Fits for β Pic b and β Pic c

We present a comprehensive orbital analysis to the exoplanets β Pictoris b and c that resolves previously reported tensions between the dynamical and evolutionary mass constraints on β Pic b. We use the Markov Chain Monte Carlo orbit code orvara to fit 15 years of radial velocities and relative astrometry (including recent GRAVITY measurements), absolute astrometry from Hipparcos and Gaia, and a single relative radial velocity measurement between β Pic A and b. We measure model-independent masses of 9.3{sub −2.5}{sup +2.6} M {sub Jup} for β Pic b and 8.3 ± 1.0 M {sub Jup} for β Pic c. These masses are robust to modest changes to the input data selection. We find a well-constrained eccentricity of 0.119 ± 0.008 for β Pic b, and an eccentricity of 0.21{sub −0.09}{sup +0.16} for β Pic c, with the two orbital planes aligned to within ∼05. Both planets’ masses are within ∼1σ of the predictions of hot-start evolutionary models and exclude cold starts. We validate our approach on N-body synthetic data integrated using REBOUND. We show that orvara can account for three-body effects in the β Pic system down to a level ∼5 times smaller than the GRAVITY uncertainties. Systematics in the masses and orbital parameters from orvara’s approximate treatment of multiplanet orbits are a factor of ∼5 smaller than the uncertainties we derive here. Future GRAVITY observations will improve the constraints on β Pic c’s mass and (especially) eccentricity, but improved constraints on the mass of β Pic b will likely require years of additional radial velocity monitoring and improved precision from future Gaia data releases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Commercial Synthetic Aperture Radar Data for Surface Deformation and Change

The commercial synthetic aperture radar (SAR) market is experiencing year-over-year growth and is currently capable of imaging anywhere in the world within an hour at X-band. Surface Deformation and Change (SDC) is a mission study at NASA to investigate innovative architectures beyond the NASA-ISRO SAR (NISAR) mission in the next decade. In this article we present preliminary findings on the technical capabilities of the currently available commercial SAR data, and investigate their applicability to SDC goals, such as imaging quality and retrieval of surface motion.

SAR↗

NaroNet: Discovery of tumor microenvironment elements from highly multiplexed images

Understanding the spatial interactions between the elements of the tumor microenvironment -i.e. tumor cells. fibroblasts, immune cells- and how these interactions relate to the diagnosis or prognosis of a tumor is one of the goals of computational pathology. We present NaroNet, a deep learning framework that models the multi-scale tumor microenvironment from multiplex-stained cancer tissue images and provides patient-level interpretable predictions using a seamless end-to-end learning pipeline. Trained only with multiplex-stained tissue images and their corresponding patient-level clinical labels, NaroNet unsupervisedly learns which cell phenotypes, cell neighborhoods, and neighborhood interactions have the highest influence to predict the correct label. To this end, NaroNet incorporates several novel and state-of-the-art deep learning techniques, such as patch-level contrastive learning, multi-level graph embeddings, a novel max-sum pooling operation, or a metric that quantifies the relevance that each microenvironment element has in the individual predictions. We validate NaroNet using synthetic data simulating multiplex-immunostained images where a patient label is artificially associated to the -adjustable- probabilistic incidence of different microenvironment elements. We then apply our model to two sets of images of human cancer tissues: 336 seven-color multiplex-immunostained images from 12 high-grade endometrial cancer patients; and 382 35-plex mass cytometry images from 215 breast cancer patients. In both synthetic and real datasets, NaroNet provides outstanding predictions of relevant clinical information while associating those predictions to the presence of specific microenvironment elements.

60 APPLIED LIFE SCIENCES↗

DESI DR1 Ly α 1D power spectrum: the Fast Fourier Transform estimator measurement

Here, we present the one-dimensional Lyman-α forest power spectrum measurement derived from the data release 1 (DR1) of the Dark Energy Spectroscopic Instrument (DESI). The measurement of the Lyman-α forest power spectrum along the line of sight from high-redshift quasar spectra provides information on the shape of the linear matter power spectrum, neutrino masses, and the properties of dark matter. In this work, we use a Fast Fourier Transform (FFT)-based estimator, which is validated on synthetic data in a companion paper. Compared to the FFT measurement performed on the DESI early data release, we improve the noise characterization with a cross-exposure estimator and test the robustness of our measurement using various data splits. We also refine the estimation of the uncertainties and now present an estimator for the covariance matrix of the measurement. Furthermore, we compare our results to previous high-resolution and eBOSS measurements. In another companion paper, we present the same DR1 measurement using the Quadratic Maximum Likelihood Estimator (QMLE). These two measurements are consistent with each other and constitute the most precise one-dimensional power spectrum measurement to date, while being in good agreement with results from the DESI early data release.

Lyman alpha forest↗

Modelling the Milky Way – I. Method and first results fitting the thick disc and halo with DES-Y3 data

ABSTRACT We present a technique to fit the stellar components of the Galaxy by comparing Hess Diagrams (HDs) generated from trilegal models to real data. We apply this technique, which we call mwfitting, to photometric data from the first 3 yr of the Dark Energy Survey (DES). After removing regions containing known resolved stellar systems such as globular clusters, dwarf galaxies, nearby galaxies, the Large Magellanic Cloud, and the Sagittarius Stream, our main sample spans a total area of ∼2300 deg2. We further explore a smaller subset (∼1300 deg2) that excludes all regions with known stellar streams and stellar overdensities. Validation tests on synthetic data possessing similar properties to the DES data show that the method is able to recover input parameters with a precision better than 3 per cent. We fit the DES data with an exponential thick disc model and an oblate double power-law halo model. We find that the best-fitting thick disc model has radial and vertical scale heights of 2.67 ± 0.09 kpc and 925 ± 40 pc, respectively. The stellar halo is fit with a broken power-law density profile with an oblateness of 0.75 ± 0.01, an inner index of 1.82 ± 0.08, an outer index of 4.14 ± 0.05, and a break at 18.52 ± 0.27 kpc from the Galactic centre. Several previously discovered stellar overdensities are recovered in the residual stellar density map, showing the reliability of mwfitting in determining the Galactic components. Simulations made with the best-fitting parameters are a promising way to predict Milky Way star counts for surveys such as the LSST and Euclid.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

First Sagittarius A* Event Horizon Telescope Results. IV. Variability, Morphology, and Black Hole Mass

In this paper we quantify the temporal variability and image morphology of the horizon-scale emission from Sgr A*, as observed by the EHT in 2017 April at a wavelength of 1.3 mm. We find that the Sgr A* data exhibit variability that exceeds what can be explained by the uncertainties in the data or by the effects of interstellar scattering. The magnitude of this variability can be a substantial fraction of the correlated flux density, reaching ~100% on some baselines. Through an exploration of simple geometric source models, we demonstrate that ring-like morphologies provide better fits to the Sgr A* data than do other morphologies with comparable complexity. We develop two strategies for fitting static geometric ring models to the time-variable Sgr A* data; one strategy fits models to short segments of data over which the source is static and averages these independent fits, while the other fits models to the full data set using a parametric model for the structural variability power spectrum around the average source structure. Both geometric modeling and image-domain feature extraction techniques determine the ring diameter to be 51.8 ± 2.3 μas (68% credible intervals), with the ring thickness constrained to have an FWHM between ~30% and 50% of the ring diameter. To bring the diameter measurements to a common physical scale, we calibrate them using synthetic data generated from GRMHD simulations. This calibration constrains the angular size of the gravitational radius to be ${4.8}_{-0.7}^{+1.4}$ μas, which we combine with an independent distance measurement from maser parallaxes to determine the mass of Sgr A* to be ${4.0}_{-0.6}^{+1.1}\times {10}^{6}$ M⊙.

79 ASTRONOMY AND ASTROPHYSICS↗

An Integrated Centroid Finding and Particle Overlap Decomposition Algorithm for Stereo Imaging Velocimetry

An integrated algorithm for decomposing overlapping particle images (multi-particle objects) along with determining each object s constituent particle centroid(s) has been developed using image analysis techniques. The centroid finding algorithm uses a modified eight-direction search method for finding the perimeter of any enclosed object. The centroid is calculated using the intensity-weighted center of mass of the object. The overlap decomposition algorithm further analyzes the object data and breaks it down into its constituent particle centroid(s). This is accomplished with an artificial neural network, feature based technique and provides an efficient way of decomposing overlapping particles. Combining the centroid finding and overlap decomposition routines into a single algorithm allows us to accurately predict the error associated with finding the centroid(s) of particles in our experiments. This algorithm has been tested using real, simulated, and synthetic data and the results are presented and discussed.

McDowell, Mark↗

FracML: A Machine Learning Based Tool to Quantify Reservoir Scale Fracture Network for CO2 Storage

Poster on “FRACML: A Machine Learning Based Tool to Quantify Reservoir Scale Fracture Network for CO2 Storage” for the CCUS 2025 conference held in Houston, Texas March 3-5, 2025. The accurate characterization of subsurface fracture networks is essential for the secure operation of carbon capture, utilization, and storage (CCUS) projects. A thorough understanding of the spatial distribution of subsurface faults and fractures is crucial for predicting CO2 plume evolution and minimizing risks such as potential leakage into overlying formations or induced seismicity. In this context, robust fracture network quantification plays a pivotal role in reservoir management, providing the data necessary to fine-tune operational parameters, and ensure the environmental and economic viability of CCUS projects. As part of the U.S. Department of Energy’s SMART (Science-informed Machine Learning for Accelerating Real-time Decisions in Subsurface Applications) initiative, we focused on the development and application of a machine learning-based tool (FRACML) designed to quantify and map fracture networks using real-world (non-synthetic) data from an active CO2 injection site. Our objective is to demonstrate the utility of this tool in improving operational efficiency and safety across CCUS sites.

artifical intelligence / machine learning (AI/ML)↗