Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Quantifying the effect of 3D models on moment tensor results using synthetic data

Moment tensors provide vital information on seismic source properties for seismic events. Moment tensors require seismic wavespeed models to compute Green’s functions, which measure the impulse response between a given source and receiver. Traditionally, researchers have used one-dimensional velocity models to calculate Green’s functions since 1D Green’s functions are computationally cheap to compute. Local 1D velocity models can also accurately model waveforms at short distances (< 500 km). However, 1D velocity models do not account for lateral heterogeneity, which can cause significant misfit in tectonically complex regions such as the Middle East (Covellone and Savage, 2012). Green’s functions calculated using 3D seismic wavespeed models have been shown to perform better in tectonically complex regions (Covellone and Savage, 2012; Kintner and Modrak, 2022), so we are interested in quantifying the effect of considering 3D structure on moment tensor inversion results. The Middle East is an ideal study area for a synthetic moment tensor test for two reasons. Firstly, the Middle East is a tectonically complex region that has been heavily studied. Secondly, the Middle East has significant seismic activity throughout the region, but imperfect station coverage due to limited open data through large swaths of the domain. The tectonic complexity and uneven station coverage will test real-world performance even in a synthetic experiment.

58 GEOSCIENCES↗

Synthetic data generation for machine learning model training for energy theft scenarios using cosimulation

Abstract Technical and non‐technical losses in distribution circuits result in significant economic costs to power utilities. One type of non‐technical loss is energy theft by various means including illegal tapping of feeders, bypassing the meter, and billing fraud. These losses are usually hard to detect, and can remain undetected for long periods of time. Machine learning models have been proven effective in detecting these conditions, but rely on the availability of large, good‐quality training data sets. The problem is exacerbated by the imbalanced nature of data related to these conditions—energy theft, though costly, is very rare. The available data sets generally have very few samples of theft with most of the data pertaining to normal operation. Such data sets are generally not suitable to train machine learning models. In this paper, an overview of energy theft detection techniques, the challenges with their data needs, and the limitations of current techniques to bridge such data limitations is presented. A co‐simulation framework is proposed to generate reliable training data for machine learning algorithms for theft detection. An example scenario is presented and a machine learning model is built to detect certain kinds of energy theft.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Deep residual networks for crystallography trained on synthetic data

The use of artificial intelligence to process diffraction images is challenged by the need to assemble large and precisely designed training data sets. To address this, a codebase called Resonet was developed for synthesizing diffraction data and training residual neural networks on these data. Here, two per-pattern capabilities of Resonet are demonstrated: (i) interpretation of crystal resolution and (ii) identification of overlapping lattices. Resonet was tested across a compilation of diffraction images from synchrotron experiments and X-ray free-electron laser experiments. Crucially, these models readily execute on graphics processing units and can thus significantly outperform conventional algorithms. While Resonet is currently utilized to provide real-time feedback for macromolecular crystallography users at the Stanford Synchrotron Radiation Lightsource, its simple Python-based interface makes it easy to embed in other processing frameworks. This work highlights the utility of physics-based simulation for training deep neural networks and lays the groundwork for the development of additional models to enhance diffraction collection and analysis.

36 MATERIALS SCIENCE↗

Synthetic Data and Graph Generation for Modeling Adversarial Activity (Final Project Report)

The Data and Graph Generation for Modeling Adversary Activity (MAA) project developed a methodology along with scalable graph modeling and generation tools to produce realistic large-scale background activity graphs with embedded adversarial activity pathways. The technical report presents PNNL methodology, released datasets, lessons learned, and recommendations to develop graph analytic algorithms for structure-only and attributed knowledge graphs.

97 MATHEMATICS AND COMPUTING↗

Adaptive Cybersecurity for Distributed Energy Resources (AdCyDER): Online Reinforcement Learning with Stackelberg-Optimized Defenses — Pipeline Architecture, Evaluation Methodology, and Findings from a Synthetic-Data Evaluation

This report documents the design and evaluation of an integrated online-learning pipeline developed within the AdCyDER project for Distributed Energy Resource (DER) cybersecurity. The pipeline couples a Reinforcement Learning (RL) attack classifier — which produces an attack-type probability distribution — with a Stackelberg game-theoretic (GT) defense selector that consumes those distributions alongside SME-encoded priors over (defense, attack) effectiveness pairings and perdefense costs to choose grid-health-preserving defenses. The objective is not attack classification per se but production of distributions that drive effective defense selection through the Stackelberg layer, learned from delayed grid-health feedback rather than labeled attack data. AdCyDER as a whole is broader than the work presented here; this report covers the specific RL/GT loop integration and its evaluation. We present the integrated pipeline (SCADA telemetry with Fronius inverter physics, Suricata IDS, time-windowed aggregation, per-facility LSTM classifier, Stackelberg optimizer, OpenC2 actuators), an experimental campaign of 28 eight-hour iterations across three baseline modes, and a pipeline-ordered diagnostic protocol. The protocol identifies two distinct failure modes within the loop: paired supervised ceilings on the same features establish that the deployed online RL classifier (macro F1 ≈ 0.07) sits at least 4.7× below a same-architecture supervised LSTM (≈ 0.34) and 10–11× below a linear feature-signal ceiling (≈ 0.70–0.79 depending on per-facility isolation), localizing the dominant failure to the training procedure; and the reward signal driving online updates carries weak directional coupling with classifier correctness in the methodology-expected direction (multi-lens convergent: top-decile P(true) records produce more frequent state changes and slightly larger improvements, top-vs-bot Cohen’s 𝑑 ≈ −0.19), but at effect magnitudes too small to drive gradient-based learning at the campaign sample size. The original learning hypothesis is not supported by the data. The primary contributions are the diagnostic methodology — proposed as a transferable falsification protocol for online RL/GT defense pipelines learning from delayed environmental reward — and the open, reproducible experimental infrastructure. We outline reward reformulation as the highest-priority aspirational next step given the underpowered-but-aligned Q6 reading, with hardware-in-the-loop evaluation as the broadest scope-expansion option.

Blakely, Benjamin [Argonne National Laboratory (AN↗

Intracardiac Electrical Imaging using the 12-lead ECG: A Machine Learning Approach using Synthetic Data

Current state-of-the-art techniques for non-invasive imaging of cardiac electrical phenomena require voltage recordings from dozens of different torso locations and anatomical models built from expensive medical diagnostic imaging procedures. Here this study aimed to assess if recent machine learning advances could alternatively reconstruct electroanatomical maps at clinically relevant resolutions using only the standard 12-lead electrocardiogram (ECG) as input. To that end, a computational study was conducted to generate a dataset of over 16000 detailed cardiac simulations, which was then used to train neural network (NN) architectures designed to exploit both spatial and temporal correlations in the ECG signal. Analysis over a validation set showed average errors in activation map reconstruction below 1.7 msec over 75 intracardiac locations. Furthermore, phenotypical patterns of activation and the morphology of the activation potential were correctly reconstructed. The approach offers opportunities to stratify patients non-invasively, both retrospectively and prospectively, using metrics otherwise only available through invasive clinical procedures.

59 BASIC BIOLOGICAL SCIENCES↗

Synthetic Infrasound Data for Machine Learning Detectors

Synthetic data is a powerful tool to generate large amounts of training data for machine learning models. The methods outlined in this report will be used to retrain the deep learning classifier for increased accuracy. Synthetic data will be useful to address the natural class imbalance between the different categories in the original ML work. Additionally, these tools will be applied for a variety of signal analysis methods that would use signals with a known signal-to-noise ratio for validation and testing.

58 GEOSCIENCES↗

A fast two-stage algorithm for non-negative matrix factorization in smoothly varying data

This article reports the study of algorithms for non-negative matrix factorization (NMF) in various applications involving smoothly varying data such as time or temperature series diffraction data on a dense grid of points. Utilizing the continual nature of the data, a fast two-stage algorithm is developed for highly efficient and accurate NMF. In the first stage, an alternating non-negative least-squares framework is used in combination with the active set method with a warm-start strategy for the solution of subproblems. In the second stage, an interior point method is adopted to accelerate the local convergence. The convergence of the proposed algorithm is proved. The new algorithm is compared with some existing algorithms in benchmark tests using both real-world data and synthetic data. Furthermore, the results demonstrate the advantage of the algorithm in finding high-precision solutions.

interior point method↗

Synthetic data-driven deep learning for label-free autonomous atomic force microscopy

Atomic force microscopy (AFM) is a widely used tool for nanoscale characterization across materials science, energy research, and biology. However, its adoption in high-throughput materials discovery and statistically driven studies remains limited by a strong dependence on expert operator input and by the scarcity of annotated experimental AFM datasets needed to enable data-driven automation. Here, we introduce SimuScan, a synthetic-data–driven framework that enables reliable AFM feature identification, segmentation, and targeted imaging without requiring large manually labeled experimental datasets. SimuScan generates tunable, high-fidelity synthetic AFM images of defined morphologies while incorporating realistic experimental artifacts, including tip–sample convolution, noise, flattening distortions, and surface debris. These datasets are shown to support scalable, label-free training of modern deep learning models for AFM analysis. When integrated into data-driven AFM workflows, SimuScan-trained models can locate and analyze nanoscale structures across large datasets and guide targeted follow-up imaging. We validate this approach on nanostructured surfaces, DNA assemblies, and bacterial cells, demonstrating robust generalization across diverse sample types with minimal operator intervention. More broadly, this work establishes a general strategy for generating explicitly conditioned, task-relevant synthetic data to improve the reliability of downstream models in autonomous microscopy.

Millan-Solsona, Ruben [Oak Ridge National Laborato↗