Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

MOSAIC: a joint modeling methodology for combined circadian and non-circadian analysis of multi-omics data

Abstract Motivation Circadian rhythms are approximately 24-h endogenous cycles that control many biological functions. To identify these rhythms, biological samples are taken over circadian time and analyzed using a single omics type, such as transcriptomics or proteomics. By comparing data from these single omics approaches, it has been shown that transcriptional rhythms are not necessarily conserved at the protein level, implying extensive circadian post-transcriptional regulation. However, as proteomics methods are known to be noisier than transcriptomic methods, this suggests that previously identified arrhythmic proteins with rhythmic transcripts could have been missed due to noise and may not be due to post-transcriptional regulation. Results To determine if one can use information from less-noisy transcriptomic data to inform rhythms in more-noisy proteomic data, and thus more accurately identify rhythms in the proteome, we have created the Multi-Omics Selection with Amplitude Independent Criteria (MOSAIC) application. MOSAIC combines model selection and joint modeling of multiple omics types to recover significant circadian and non-circadian trends. Using both synthetic data and proteomic data from Neurospora crassa, we showed that MOSAIC accurately recovers circadian rhythms at higher rates in not only the proteome but the transcriptome as well, outperforming existing methods for rhythm identification. In addition, by quantifying non-circadian trends in addition to circadian trends in data, our methodology allowed for the recognition of the diversity of circadian regulation as compared to non-circadian regulation. Availability and implementation MOSAIC’s full interface is available at https://github.com/delosh653/MOSAIC. An R package for this functionality, mosaic.find, can be downloaded at https://CRAN.R-project.org/package=mosaic.find. Supplementary information Supplementary data are available at Bioinformatics online.

De los Santos, Hannah↗

Robust and optimal alignment of high-dimensional data using maximum likelihood estimation through a random sample consensus framework

Abstract Correcting spatial orientations of groups of high-dimensional data sets such that they are all in a consistent coordinate system is often a time-consuming and error-prone process. Automation of this process can be accomplished by using Generalized Procrustes Analysis to estimate the relative orientations among a population of high-dimensional data sets. A least squares Procrustes solution is applied through a maximum likelihood estimation and random sample consensus framework for robustness. The likelihood model is comprised of a mixture distribution where inliers are modeled using t -distribution and outliers from a uniform distribution. Applications will focus on a synthetic data set that emulates triaxial acceleration data and also real shock data from a population of triaxial accelerometers. Outliers represent either non-rigid body responses, environmental noise, and/or sensor and data acquisition issues. The intended application for the methodology is to robustly automate the rotation of populations of experimentally collected triaxial accelerometer data sets to a single global coordinate system.

LOSAC↗

Unlocking hidden information in sparse small-angle neutron scattering measurements

Hypothesis Small-Angle Neutron Scattering (SANS) is a powerful technique for studying soft matter systems such as colloids, polymers, and lyotropic phases, providing nanoscale structural insights. However, its effectiveness is limited by low neutron flux, leading to long acquisition times and noisy data. Here, we hypothesize that Bayesian statistical inference using Gaussian Process Regression (GPR) can reconstruct high-fidelity scattering data from sparse measurements by leveraging intensity smoothness and continuity. Experiments and Simulations The method was benchmarked computationally and validated through SANS experiments on various soft matter systems, including wormlike micelles, colloidal suspensions, polymeric structures, and lyotropic phases. GPR-based inference was applied to both experimental and synthetic data to evaluate its effectiveness in noise reduction and intensity reconstruction. Findings GPR significantly enhances SANS data quality and therefore reducing measurement times by up to two orders of magnitude. This cost-effective approach maximizes experimental efficiency, enabling high-throughput studies and real-time monitoring of dynamic systems. It is particularly beneficial for weakly scattering and time-sensitive studies. Beyond SANS, this framework applies to other low-SNR techniques, including laboratory-based small-angle X-ray scattering and various dynamical scattering methods. Furthermore, it offers transformative potential for compact neutron sources, enhancing their viability for structural analysis in resource-limited settings.

Small angle neutron scattering↗

A Benchmark to Test Generalization Capabilities of Deep Learning Methods to Classify Severe Convective Storms in a Changing Climate

Abstract This is a test case study assessing the ability of deep learning methods to generalize to a future climate (end of 21st century) when trained to classify thunderstorms in model output representative of the present‐day climate. A convolutional neural network (CNN) was trained to classify strongly rotating thunderstorms from a current climate created using the Weather Research and Forecasting model at high‐resolution, then evaluated against thunderstorms from a future climate and found to perform with skill and comparatively in both climates. Despite training with labels derived from a threshold value of a severe thunderstorm diagnostic (updraft helicity), which was not used as an input attribute, the CNN learned physical characteristics of organized convection and environments that are not captured by the diagnostic heuristic. Physical features were not prescribed but rather learned from the data, such as the importance of dry air at mid‐levels for intense thunderstorm development when low‐level moisture is present (i.e., convective available potential energy). Explanation techniques also revealed that thunderstorms classified as strongly rotating are associated with learned rotation signatures. Results show that the creation of synthetic data with ground truth is a viable alternative to human‐labeled data and that a CNN is able to generalize a target using learned features that would be difficult to encode due to spatial complexity. Most importantly, results from this study show that deep learning is capable of generalizing to future climate extremes and can exhibit out‐of‐sample robustness with hyperparameter tuning in certain applications.

54 ENVIRONMENTAL SCIENCES↗

DONKEY: A Flexible and Accurate Algorithm for Clustering

We propose an accurate clustering algorithm suitable for the varied and multidimensional data sets that correspond to temporal snapshots from on-the-fly nonadiabatic trajectory-based simulations of photoexcited dynamics. The algorithm approximates the underlying probability density function using variable kernel density estimation, with local maxima corresponding to cluster centers. Each data point is then assigned to one of the maxima by employing a maximization procedure. Finally, clusters artificially separated by minor fluctuations in the probability density are merged. The algorithm does not require parameter tuning, which ensures flexibility and reduces the risk of bias. It is tested on several synthetic data sets, where it consistently outperforms conventional clustering algorithms. As a final example, the algorithm is applied to the excited dynamics of the norbornadiene ⇌ quadricyclane (C 7 H 8 ) molecular photoswitch, demonstrating how distinct reaction pathways can be identified.

algorithms↗

Revisiting the Kinematics of the Cylinder Test

The cylinder expansion experiment is a well-established performance test for condensed explosives that is utilized routinely to determine pressure-energy-volume relationships for detonation products. The modern cylinder test employs optical interferometric techniques to measure velocity, as opposed to older realizations where streak cameras were commonplace. Despite their widespread use, questions sometimes remain as to what kind of data the velocity diagnostics in a cylinder test provide along with how to interpret such measurements and relate them to physical quantities of interest. Here in this study, equations are derived that fully describe the kinematics of the cylinder wall during expansion. These equations are then applied to experimental as well as synthetic data generated via a hydrodynamic simulation in order to verify the mathematical framework developed.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A generalized forward fit for neutron detectors with energy-dependent response functions

To date, most analysis of neutron time-of-flight data from inertial confinement fusion experiments has focused on the relatively small range of energies corresponding to the primary neutrons from DD and DT fusion, and have therefore employed instrument response functions (IRF’s) corresponding to monoenergetic 2.45-MeV or 14.03-MeV neutrons. For analysis of time-of-flight signals corresponding to broader ranges of neutron energies, accurate treatment of the data requires the use of an energy-dependent IRF. Here, this work describes interpolation of the IRF for neutrons of arbitrary energy, construction of an energy-dependent IRF, and application of this IRF in a forward fit via matrix multiplication. As an example of the application of this method, an analysis of synthetic data relevant to TT fusion experiments at the Omega Laser Facility is discussed. This example is used to illustrate the differences between a forward fit that uses an energy-dependent IRF and a forward fit that uses a monoenergetic IRF. Use of the energy-dependent IRF is shown to result in accurate inference of the fit parameters of interest.

47 OTHER INSTRUMENTATION↗

Application of an energy-dependent instrument response function to analysis of nTOF data from cryogenic DT experiments

Neutron time-of-flight (nTOF) detectors are used to diagnose the conditions present in inertial confinement fusion (ICF) experiments and basic laboratory physics experiments performed on an ICF platform. The instrument response function (IRF) of these detectors is constructed by convolution of two components: an x-ray IRF and a neutron interaction response. The shape of the neutron interaction response varies with incident neutron energy, changing the shape of the total IRF. Analyses of nTOF data that span a broad range of energies must account for this energy-dependence in order to accurately infer plasma parameters and nuclear properties in ICF experiments. This work briefly reviews a matrix multiplication approach to convolution which allows for an energy-dependent change in the shape of the IRF. This method is applied to synthetic data resembling symmetric cryogenic DT implosions to examine the effect of the energy-dependent IRF on the inferred areal density. Here, results of forward fits that infer ion temperatures and areal densities from nTOF data collected during cryogenic DT experiments on OMEGA are also discussed.

47 OTHER INSTRUMENTATION↗

RADEMACHER COMPLEXITY REGULARIZATION FOR CORRELATION-BASED MULTIVIEW REPRESENTATION LEARNING

Deep correlation-based multiview representation learning techniques have become increasingly popular methods for extracting highly correlated representations from multiview data. However, their ability to find highly complex mappings between the views can also lead to overfitting and overly correlated representations. In this work, we propose a regularizer for this specific problem, based on the Rademacher complexity of the DNNs, tailored for multiview correlation maximization. We demonstrate that the proposed regularization leads to less noisy representations in synthetic data and improved performance of downstream tasks in real-world multiview datasets.

Kuschel, Maurice↗

MONTE CARLO CROSS SECTION LOOKUP KERNEL FOR THE CEREBRAS WSE-2 IN CSL

This is a small kernel that was used to collect data for an upcoming paper. We would like to have the code be open source so that the reviewers (and then readers) of the paper can see the whole code, and can reproduce/verify our results. This is not a fully featured application, it cannot produce any useful simulation results, it just executes a small abstracted kernel using synthetic data. The purpose of the kernel is to understand the basic performance characteristics of an HPC kernel on novel AI accelerator architectures. The main kernel is written in the CSL coding language for use with the Cerebras WSE-2 AI accelerator. The kernel represented is the Monte Carlo cross section lookup kernel, which is a small kernel used by the Monte Carlo neutral particle transport algorithm. There is also a baseline kernel written in CUDA that we will include in the repository to form a basis for comparing the WSE-2 to GPU.

Tramm, John↗

Dispersive and nondispersive 𝐾-matrix formalisms

The modeling of coupled-channel effects has become increasingly important due to the availability of highly precise data for a large variety of hadronic (re)scattering processes. The 𝐾-matrix is a powerful, yet comparatively simple, method to describe scattering amplitudes, including coupled-channel effects, with the aim of interpreting experimental data. Throughout the literature, a range of dispersive and nondispersive 𝐾-matrix methods are employed. Here, we compare the dispersive and nondispersive formulations in the context of the N/D method. It is shown that the methods are equivalent in the physical region under 𝐾-matrix reparametrization. Differences away from the physical region are examined. Applications to synthetic data are used to illustrate the effects of model choices concerning form factors and the application of dispersion relations, with the goal of clarifying best practices. We find no clear preference with regard to dispersive modeling. In contrast, we find that interpretational ambiguity of the bare model parameters—and even of the form of the bare model—is endemic, and recommend a thorough sampling of data and model spaces to assess conclusion robustness.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Expanded analysis of machine learning models for nuclear transient identification using TPOT

Industries around the world are becoming more and more data driven. The nuclear field is no exception with several different applications being proposed. One popular area of research is the use of machine learning in transient detection. This paper seeks to build upon a previous study which made use of the AutoML package TPOT to train traditional machine learning models to classify transient events occurring with a reactor. Synthetic data was once again collected using a GPWR reactor simulator. Data on 12 different events was collected using 15 different initial conditions. Here, a dataset consisting of over 100,000 data points was compiled and used to train 7 different machine learning models using a pre-defined TPOT dictionary with 12 different preprocessing techniques. Three of the trained models were able to produce validation results in the 90s with the expanded dataset. Once the models were trained, it was possible to look into where during the simulation, misclassifications occurred. Using these three models, analysis was done to determine if TPOT could be used to train models that were effective if important features were missing. The results from this were positive with the newly trained models scoring close to the original models. Finally, to conclude this study, the three high performing models were retrained using different random states to see if there was any major variation when different states were used.

42 ENGINEERING↗

Inferring microbial co-occurrence networks from amplicon data: a systematic evaluation

Microbes commonly organize into communities consisting of hundreds of species involved in complex interactions with each other. 16S ribosomal RNA (16S rRNA) amplicon profiling provides snapshots that reveal the phylogenies and abundance profiles of these microbial communities. These snapshots, when collected from multiple samples, can reveal the co-occurrence of microbes, providing a glimpse into the network of associations in these communities. However, the inference of networks from 16S data involves numerous steps, each requiring specific tools and parameter choices. Moreover, the extent to which these steps affect the final network is still unclear. In this study, we perform a meticulous analysis of each step of a pipeline that can convert 16S sequencing data into a network of microbial associations. Through this process, we map how different choices of algorithms and parameters affect the co-occurrence network and identify the steps that contribute substantially to the variance. We further determine the tools and parameters that generate robust co-occurrence networks and develop consensus network algorithms based on benchmarks with mock and synthetic data sets. The Microbial Co-occurrence Network Explorer, or MiCoNE (available at https://github.com/segrelab/MiCoNE) follows these default tools and parameters and can help explore the outcome of these combinations of choices on the inferred networks. We envisage that this pipeline could be used for integrating multiple data sets and generating comparative analyses and consensus networks that can guide our understanding of microbial community assembly in different biomes.

16S rRNA↗

The Dark Energy Spectroscopic Instrument: one-dimensional power spectrum from first Ly α forest samples with Fast Fourier Transform

ABSTRACT We present the one-dimensional Ly α forest power spectrum measurement using the first data provided by the Dark Energy Spectroscopic Instrument (DESI). The data sample comprises 26 330 quasar spectra, at redshift z > 2.1, contained in the DESI Early Data Release and the first 2 months of the main survey. We employ a Fast Fourier Transform (FFT) estimator and compare the resulting power spectrum to an alternative likelihood-based method in a companion paper. We investigate methodological and instrumental contaminants associated with the new DESI instrument, applying techniques similar to previous Sloan Digital Sky Survey (SDSS) measurements. We use synthetic data based on lognormal approximation to validate and correct our measurement. We compare our resulting power spectrum with previous SDSS and high-resolution measurements. With relatively small number statistics, we successfully perform the FFT measurement, which is already competitive in terms of the scale range. At the end of the DESI survey, we expect a five times larger Ly α forest sample than SDSS, providing an unprecedented precise one-dimensional power spectrum measurement.

79 ASTRONOMY AND ASTROPHYSICS↗

Classification and Localization of Fracture-Hit Events in Low-Frequency Distributed Acoustic Sensing Strain Rate with Convolutional Neural Networks

Summary Distributed acoustic sensing (DAS) has been used in the oil and gas industry as an advanced technology for surveillance and diagnostics. Operators use DAS to monitor hydraulic fracturing activities, examine well stimulation efficacy, and estimate complex fracture system geometries. Particularly, low-frequency DAS can detect geomechanical events such as fracture hits because hydraulic fractures propagate and create strain rate variations in the rock. Analysis of DAS data today is mostly done post-job and subject to interpretation methods. However, the continuous and dense data stream generated live by DAS poses the opportunity for more efficient and accurate real-time data-driven analysis. The objective of this study is to develop a machine learning-based workflow that can identify and locate fracture-hit events in simulated strain rate responses correlated with low-frequency DAS data. In this paper, “fracture hit” refers to a hydraulic fracture originating from a stimulated well intersecting an offset well. We start with building a single fracture propagation model to produce strain rate patterns observed at a hypothetical monitoring well. This model is used to generate two sets of strain rate responses with one set containing fracture-hit events. The labeled synthetic data are then used to train a custom convolutional neural network (CNN) model for identifying the presence of fracture-hit events. The same model is trained again for locating the event with the output layer of the model replaced with linear units. We achieved near-perfect predictions for both event classification and localization. These promising results prove the feasibility of using CNN for real-time event detection from fiber-optic sensing data. Additionally, we use edge detection techniques to recognize fracture-hit event patterns in strain rate images. The fracture-hit location can be identified using recognized pixels in the image. The accuracy of edge detection-based location identification is also plausible, but edge detection is dependent on the assumption of pattern shape and image quality, hence it is less robust compared to CNN models. This comparison further supports the need for CNN applications in image-based real-time fiber-optic sensing event detection.

Engineering↗

Knowledge-guided graph machine learning for spatially distributed prediction of daily discharge and nitrogen export dynamics

Spatially distributed prediction of streamflow and nitrogen export dynamics is essential for precision management of agricultural watersheds. While temporal deep learning models such as Long Short-Term Memory (LSTM) have shown strong performance at basin scales, their ability to generalize spatially is limited by insufficient representation of spatial dependencies and flow paths, particularly under data-scarce conditions. To address this gap, we propose HydroGraphNet, a knowledge-guided graph machine learning framework that integrates process-based knowledge and explicit spatial learning into temporal modeling. This framework incorporates directed graph topology to encode watershed connectivity and upstream inflows, with mass balance constraints to improve physical consistency. To enhance generalization in sparsely monitored regions, HydroGraphNet is pretrained on synthetic data generated by the SWAT+ (Soil and Water Assessment Tool Plus) model. We evaluated HydroGraphNet in the Upper Sangamon River Basin (44 HUC-12 subwatersheds, 2001–2020) against two LSTM baselines: a lumped basin-level model and a distributed variant. When benchmarked on SWAT+ simulations in pretraining, HydroGraphNet improved test NSEs by 8.9% (discharge) and 13.7% (NO₃–N load) in temporal extrapolation, and by 27.1% and 34.7% in spatial extrapolation, relative to the Lumped LSTM baseline. After fine-tuning with USGS monitoring data, the model achieved mean test NSE (KGE) scores of 0.768 (0.861) for discharge and 0.626 (0.664) for NO₃–N load, substantially outperforming baselines. Attribution analysis further highlighted the importance of upstream inflow representation and graph-based spatial learning in capturing cross-subwatershed dependencies. The model also reproduced seasonal hydrological and biogeochemical patterns consistent with known processes, demonstrating its robustness and process fidelity for spatially distributed prediction. Altogether, HydroGraphNet advances the integration of physical knowledge and spatially explicit learning in hydrological modeling, offering a generalizable framework for distributed modeling to support spatially targeted water quality management in data-scarce watersheds.

54 ENVIRONMENTAL SCIENCES↗

Regularizing INR with Diffusion Prior for Self-Supervised 3D Reconstruction OF Neutron Computed Tomography Data

Recently, generative diffusion priors have made huge strides as inverse problem solvers, including the ability to be adapted for inference on out-of-distribution data. Concurrently, implicit neural representations (INRs) have emerged as fast and lightweight inverse imaging solvers that are amenable to hybrid approaches that combine learned priors with traditional inverse problem formulations. In this paper, we present a diffusive computed tomography (CT) inversion framework for regularizing INRs called Diffusive INR (DINR), designed to enable high-quality reconstruction from sparse-view neutron CT. Pretrained purely on synthetic data, DINR is evaluated on simulated and experimentally obtained observations of concrete microstructures, where traditional reconstruction methods suffer substantial degradation when the number of views is reduced. Our approach delivers superior performance, reduces reconstruction artifacts, and achieves gains in PSNR and SSIM, enabling accurate micro-structural characterization even under extreme data limitations compared to state-of-the-art sparse-view reconstruction techniques.

Hossain, Maliha [ORNL]↗

Precise Dynamical Masses and Orbital Fits for β Pic b and β Pic c

We present a comprehensive orbital analysis to the exoplanets β Pictoris b and c that resolves previously reported tensions between the dynamical and evolutionary mass constraints on β Pic b. We use the Markov Chain Monte Carlo orbit code orvara to fit 15 years of radial velocities and relative astrometry (including recent GRAVITY measurements), absolute astrometry from Hipparcos and Gaia, and a single relative radial velocity measurement between β Pic A and b. We measure model-independent masses of 9.3{sub −2.5}{sup +2.6} M {sub Jup} for β Pic b and 8.3 ± 1.0 M {sub Jup} for β Pic c. These masses are robust to modest changes to the input data selection. We find a well-constrained eccentricity of 0.119 ± 0.008 for β Pic b, and an eccentricity of 0.21{sub −0.09}{sup +0.16} for β Pic c, with the two orbital planes aligned to within ∼05. Both planets’ masses are within ∼1σ of the predictions of hot-start evolutionary models and exclude cold starts. We validate our approach on N-body synthetic data integrated using REBOUND. We show that orvara can account for three-body effects in the β Pic system down to a level ∼5 times smaller than the GRAVITY uncertainties. Systematics in the masses and orbital parameters from orvara’s approximate treatment of multiplanet orbits are a factor of ∼5 smaller than the uncertainties we derive here. Future GRAVITY observations will improve the constraints on β Pic c’s mass and (especially) eccentricity, but improved constraints on the mass of β Pic b will likely require years of additional radial velocity monitoring and improved precision from future Gaia data releases.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗