Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Low-pass spectral analysis of time-resolved serial femtosecond crystallography data

Low-pass spectral analysis (LPSA) is a recently developed dynamics retrieval algorithm showing excellent retrieval properties when applied to model data affected by extreme incompleteness and stochastic weighting. In this work, we apply LPSA to an experimental time-resolved serial femtosecond crystallography (TR-SFX) dataset from the membrane protein bacteriorhodopsin (bR) and analyze its parametric sensitivity. While most dynamical modes are contaminated by nonphysical high-frequency features, we identify two dominant modes, which are little affected by spurious frequencies. The dynamics retrieved using these modes shows an isomerization signal compatible with previous findings. We employ synthetic data with increasing timing uncertainty, increasing incompleteness level, pixel-dependent incompleteness, and photon counting errors to investigate the root cause of the high-frequency contamination of our TR-SFX modes. By testing a range of methods, we show that timing errors comparable to the dynamical periods to be retrieved produce a smearing of dynamical features, hampering dynamics retrieval, but with no introduction of spurious components in the solution, when convergence criteria are met. Using model data, we are able to attribute the high-frequency contamination of low-order dynamical modes to the high levels of noise present in the data. Finally, we propose a method to handle missing observations that produces a substantial dynamics retrieval improvement from synthetic data with a significant static component. Reprocessing of the bR TR-SFX data using the improved method yields dynamical movies with strong isomerization signals compatible with previous findings.

59 BASIC BIOLOGICAL SCIENCES↗

Maven: a multimodal foundation model for supernova science

Abstract A common setting in astronomy is the availability of a small number of high-quality observations, and larger amounts of either lower-quality observations or synthetic data from simplified models. Time-domain astrophysics is a canonical example of this imbalance, with the number of supernovae observed photometrically outpacing the number observed spectroscopically by multiple orders of magnitude. At the same time, no data-driven models exist to understand these photometric and spectroscopic observables in a common context. Contrastive learning objectives, which have grown in popularity for aligning distinct data modalities in a shared embedding space, provide a potential solution to extract information from these modalities. We present Maven, the first foundation model for supernova science. To construct Maven, we first pre-train our model to align photometry and spectroscopy from 0.5 M synthetic supernovae using a contrastive objective. We then fine-tune the model on 4702 observed supernovae from the Zwicky transient facility. Maven reaches state-of-the-art performance on both classification and redshift estimation, despite the embeddings not being explicitly optimized for these tasks. Through ablation studies, we show that pre-training with synthetic data improves overall performance. In the upcoming era of the Vera C. Rubin observatory, Maven will serve as a valuable tool for leveraging large, unlabeled and multimodal time-domain datasets.

Zhang, Gemma (ORCID:0000000280198082)↗

Terrestrial Water Mass Load Changes from Gravity Recovery and Climate Experiment (GRACE)

Recent studies show that data from the Gravity Recovery and Climate Experiment (GRACE) is promising for basin- to global-scale water cycle research. This study provides varied assessments of errors associated with GRACE water storage estimates. Thirteen monthly GRACE gravity solutions from August 2002 to December 2004 are examined, along with synthesized GRACE gravity fields for the same period that incorporate simulated errors. The synthetic GRACE fields are calculated using numerical climate models and GRACE internal error estimates. We consider the influence of measurement noise, spatial leakage error, and atmospheric and ocean dealiasing (AOD) model error as the major contributors to the error budget. Leakage error arises from the limited range of GRACE spherical harmonics not corrupted by noise. AOD model error is due to imperfect correction for atmosphere and ocean mass redistribution applied during GRACE processing. Four methods of forming water storage estimates from GRACE spherical harmonics (four different basin filters) are applied to both GRACE and synthetic data. Two basin filters use Gaussian smoothing, and the other two are dynamic basin filters which use knowledge of geographical locations where water storage variations are expected. Global maps of measurement noise, leakage error, and AOD model errors are estimated for each basin filter. Dynamic basin filters yield the smallest errors and highest signal-to-noise ratio. Within 12 selected basins, GRACE and synthetic data show similar amplitudes of water storage change. Using 53 river basins, covering most of Earth's land surface excluding Antarctica and Greenland, we document how error changes with basin size, latitude, and shape. Leakage error is most affected by basin size and latitude, and AOD model error is most dependent on basin latitude.

Seo, K.-W.↗

Deep Learning At Depth: Estimating subsurface parameters from geophysical monitoring data

Geophysical imaging techniques are a non-invasive way to image the subsurface and understand both subsurface solid (rock/soil) and fluid property distributions and their evolution in time. Inversions of the geophysical data, such as Electrical Resistance Tomography (ERT) data, are solved to estimate the subsurface property distributions, such as conductivity, and many inversion techniques smooth out sharp gradients in rock or fluid property distributions. Sharp gradients in subsurface properties tend to be present in situations with complex subsurface structures, which are common in many subsurface applications. We have successfully demonstrated that it is possible to inform, or constrain, inversions with neural networks trained on synthetic data with complex subsurface structures. Initial results suggest this process may be optimizable to yield property distributions that better represent the true property distributions than the same inversion process without the neural network constraint. Future work would optimize the neural network performance for this application and then apply the synthetic-data trained neural network to real data to understand the utility and performance of this technique for real data sets.

47 OTHER INSTRUMENTATION↗

Application of machine learning techniques for fast MeV x-ray spectra unfolding from filter stack spectrometer data

Recovery of MeV x-ray spectra from detector signals is difficult because the response matrix inversion is ill-conditioned and current methods are too slow for high-repetition-rate experiments. In this work, we make use of neural networks to unfold MeV x-ray spectra from measurements obtained with a filter stack spectrometer at rates of near 40 Hz. The neural network was trained on synthetic data and tested on both synthetic and experimental data, the latter obtained in two separate experiments performed at the Omega EP laser facility. We show here that this unfolding method has good performance on synthetic data and that it is a promising option for experimental data of up to 40 MeV. The accuracy on experimental data is verified by using a simple forward model to compare against measured values.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Bayesian Optimization of Non-Invariant Systems with Constraints Developed for Application to the ECR Ion Source VENUS

In this work, we consider the optimization of non-invariant systems with both safety and control constraints. We present a new approach based on Bayesian optimization for the dynamic, safe and controlled optimization of such systems. Although there are other possible use cases, we focus on the application to the electron cyclotron resonance ion source VENUS. From experimental data, we have observed that VENUS behaves to first order as a non-invariant dynamic system with moving areas of instability. Our novel approach aims at providing a tool that can maintain system optimization in a safe way. This is accomplished by making sure the objective function, the beam current in the case of VENUS, does not fall under an operational minimum, while simultaneously requiring the optimization to avoid areas where VENUS is unstable. We compare the result of our approach on synthetic data modeled to mimic the behavior of VENUS with two methods from the literature, a standard Bayesian optimizer and a safe Bayesian optimizer, both adapted to deal with dynamic systems. A cross Student T-test is conducted to show the significance of the improvement given by the new method we introduce here, regarding the two preexisting methods we compared to. The results of the tests conducted on synthetic data show that the proposed method succeeds at maintaining the system optimized and obeys the predefined constraints better than the literature methods explored.

Bayesian optimization↗

Rapid failure mode classification and quantification in batteries: A deep learning modeling framework

Unique, rapid identification and quantification of the dominant aging modes in lithium-ion batteries (LiBs) with early and non-specialized test data is a significant scientific challenge. Leveraging synthetic-data, deep-learning (DL) techniques have great potential to enable fast and robust classification and quantification of battery aging modes that produce different patterns of cell aging. This study, for the first time, presents a synthetic–data-based DL modeling framework for rapid and automatic classification and quantification of battery-aging modes and resultant aging with experimental validation. Availing synthetic dQ.dV -1 curves for ~26000 initial conditions and aging modes, the framework classified the dominant aging modes, for cells undergoing fast charge, in fewer than 100 cycles. Upon classification, the framework quantified the evolution of the aging modes, which were often nonuniform with cycling, for 22 gr/NMC532 pouch cells tested up to 600 cycles at different charging rates (1C–9C).

25 ENERGY STORAGE↗

AI-Driven Crack Detection for Remanufacturing Cylinder Heads Using Deep Learning and Engineering-Informed Data Augmentation

Detecting cracks in cylinder heads traditionally relies on manual inspection, which is time-consuming and susceptible to human error. As an alternative, automated object detection utilizing computer vision and machine learning models has been explored. However, these methods often face challenges due to a lack of sufficiently annotated training data, limited image diversity, and the inherently small size of cracks. Addressing these constraints, this paper introduces a novel automated crack-detection method that enhances data availability through a synthetic data generation technique. Unlike general data augmentation practices, our method involves copying cracks from one location to another, guided by both random and informed engineering decisions about likely crack formations due to cyclic thermomechanical loads. The innovative aspect of our approach lies in the integration of domain-specific engineering knowledge into the synthetic generation process, which substantially improves detection accuracy. We evaluate our method’s effectiveness using two metrics: the F2 score, which emphasizes recall to prioritize detecting all potential cracks, and mean average precision (MAP), a standard measure in object detection. Experimental results demonstrate that, without engineering insights, our method increases the F2 score from 0.40 to 0.65, while maintaining a stable MAP. Incorporating detailed engineering knowledge further enhances the F2 score to 0.70 and improves MAP to 0.57, representing increases of 63% and 43%, respectively. These results confirm that our approach not only mitigates the limitations of traditional data augmentation but also significantly advances the reliability and precision of crack detection in industrial settings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Recent advances and applications of deep learning methods in materials science

Deep learning (DL) is one of the fastest-growing topics in materials data science, with rapidly emerging applications spanning atomistic, image-based, spectral, and textual data modalities. DL allows analysis of unstructured data and automated identification of features. The recent development of large materials databases has fueled the application of DL methods in atomistic prediction in particular. In contrast, advances in image and spectral data have largely leveraged synthetic data enabled by high-quality forward models as well as by generative unsupervised DL methods. In this article, we present a high-level overview of deep learning methods followed by a detailed discussion of recent developments of deep learning in atomistic simulation, materials imaging, spectral analysis, and natural language processing. For each modality we discuss applications involving both theoretical and experimental data, typical modeling approaches with their strengths and limitations, and relevant publicly available software and datasets. We conclude the review with a discussion of recent cross-cutting work related to uncertainty quantification in this field and a brief perspective on limitations, challenges, and potential growth areas for DL methods in materials science.

36 MATERIALS SCIENCE↗

Tracking blobs in the turbulent edge plasma of a tokamak fusion device

Abstract The analysis of turbulence in plasmas is fundamental in fusion research. Despite extensive progress in theoretical modeling in the past 15 years, we still lack a complete and consistent understanding of turbulence in magnetic confinement devices, such as tokamaks. Experimental studies are challenging due to the diverse processes that drive the high-speed dynamics of turbulent phenomena. This work presents a novel application of motion tracking to identify and track turbulent filaments in fusion plasmas, called blobs, in a high-frequency video obtained from Gas Puff Imaging diagnostics. We compare four baseline methods (RAFT, Mask R-CNN, GMA, and Flow Walk) trained on synthetic data and then test on synthetic and real-world data obtained from plasmas in the Tokamak à Configuration Variable (TCV). The blob regime identified from an analysis of blob trajectories agrees with state-of-the-art conditional averaging methods for each of the baseline methods employed, giving confidence in the accuracy of these techniques. By making a dataset and benchmark publicly available, we aim to lower the entry barrier to tokamak plasma research, thereby greatly broadening the community of scientists and engineers who might apply their talents to this endeavor.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Machine learning analysis of high-repetition-rate two-dimensional Thomson scattering spectra from laser-produced plasmas

With the emergence of high-repetition-rate two-dimensional Thomson scattering (TS) measurements, improving spectral data analysis is a key area of interest. Here, we present a new way to derive the electron temperature and density of laser-driven blast waves in plasmas from their TS spectra with machine learning (ML). This analysis occurs in both the non-collective (α < 1) and collective (α > 1) scattering regimes with the goal of autonomously and more accurately determining T c and n e both where spectral data has been collected and to give the ability to predict these attributes in regions where data has not been collected. We introduce three ML models, one trained only on experimental data, one only on synthetic data, and one using transfer learning, and compare their speed and accuracy with the conventional TS inversion algorithms in the open source PlasmaPy python package.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

DRDMannTurb: A Python package for scalable, data-driven synthetic turbulence

Synthetic turbulence models (STMs) are used in wind engineering to generate realistic flow fields and are employed as inputs to industrial wind simulations. Examples include prescribing inlet conditions in large eddy simulations that model loads on wind turbines and tall buildings. We are interested in STMs capable of generating fluctuations based on prescribed second-moment statistics since such models can simulate environmental conditions that closely resemble on-site observations. To this end, the widely used Mann model (see Mann, 1994, 1998) is the inspiration for DRDMannTurb. The Mann model is described by three physical parameters: a magnitude parameter influencing the global variance of the wind field and corresponding to the Kolmogorov constant multiplied by the rate of viscous dissipation of the turbulent kinetic energy to the two-thirds, αϵ 2/3 , a turbulence length scale parameter L, and a nondimensional parameter Γ related to the lifetime of the eddies. A number of studies, as well as international standards (e.g., those by the International Electrotechnical Commission (IEC)), include recommended values for these three parameters with the goal of standardizing wind simulations according to observed energy spectra. Yet, having only three parameters, the Mann model faces limitations in accurately representing the diversity of observable spectra. This Python package enables users to extend the Mann model and more accurately fit field measurements through flexible neural network models of the eddy lifetime function. Following Keith et al. (2021), we refer to this class of models as Deep Rapid Distortion (DRD) models. DRDMannTurb also includes a general module implementing an efficient method for synthetic turbulence generation based on a domain decomposition technique. This technique is also described in Keith et al. (2021).

17 WIND ENERGY↗

Unraveling the Wrinkle in Time-Variable Sources with Lunes and Synthetic Seismic Data

In this report, we describe how to estimate the time-variable components of the seismic moment tensor and compare these estimates to the more conventional analysis that incorporates an assumption of the source time function (STF) across all components of the seismic moment tensor. The advantage of our method is that we are able to independently estimate the time-evolution of each component of the seismic moment tensor, which may help to resolve the complex source phenomena associated with buried explosions. By performing an eigen decomposition of the time-evolving seismic moment tensor components, we are able to plot the seismic mechanism as a trajectory on a lune diagram. This technique enables interpretation of the seismic mechanism as a function of time, as opposed to the more conventional analysis which assumes that the seismic mechanism is time invariant. Finally, we describe the differences between the seismic moment and the seismic moment rate STFs, how to implement each one in inversion schemes, and the relative strengths/weaknesses of each. Our key take-away is that we are able to distinguish nearly-overlapping sources with highly different mechanisms, such as an explosion immediately following an earthquake, by estimating moment rate from seismic data through a STF-invariant inversion for the full time-variable moment tensor.

58 GEOSCIENCES↗