Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Updates in Developing a Prototype Science Pipeline and Full-Volume, Global Hyperspectral Synthetic Data Sets for NASA’s Earth System Observatory’s Upcoming Surface, Biology and Geology Mission

The Surface Biology and Geology (SBG) mission recently passed mission confirmation review and has entered phase A – design and development. SBG will acquire high resolution solar-reflected spectroscopy and thermal infrared observations at a data rate of ~2.5 TB/day and generate products at ~40 TB/day. Given that the per-day volume is greater than NASA’s total extant airborne hyperspectral data collection, collecting, processing, disseminating, and exploiting the SBG data present new challenges. To meet these challenges, we have developed a prototype science pipeline and a full-volume global hyperspectral synthetic data set to help prepare for SBG’s flight (see poster GC42D-0730). Our science pipeline is based on the science processing technology developed for NASA’s Kepler and TESS planet-hunting missions. The pipeline infrastructure, Ziggy, provides a scalable architecture for robust, repeatable, and replicable science and application products that can be run on a range of systems from a laptop to the cloud or a supercomputer. Ziggy is compliant with NASA Procedural Requirement (NPR) 7150.2C, is at a technical readiness level (TRL) of 7 and has been released to github.com/nasa/ziggy. We integrated Ziggy with EO-1/Hyperion workflows to build a prototype pipeline and ingested the 17-year mission archive that provides globally sampled visible through shortwave infrared spectra that are representative of SBG data types and volumes. We fully implemented the first stage and processed the entire 55 TB Hyperion data set from the raw data (Level 0) to top-of-the-atmosphere radiance (Level 1R). We are currently evaluating the ISOFIT atmospheric correction module to convert the L1R data to surface reflectance (Level 2) before reprocessing the full data set to L2. Crosschecks are being performed with RadCalNet as well as with coincident observations by AVIRIS. We are also investigating modern methods for georectifying the Hyperion scenes. Finally, we describe an analysis of the cost to conduct forward processing and reprocessing campaigns for SBG on HECC with dedicated compute and storage resources using the resurrected Hyperion pipeline as a proxy for full-volume SBG data. The analysis demonstrates that SBG L0 data can be processed to L2 on HECC with full reprocessing campaigns every two years for ~$2.6M over a 7-year lifespan. Moreover, 69% of the system capacity would be available for other activities, possibly enabling future open-source science activities, including algorithm development, L3+ processing, .etc.

ESD↗

Synthetic Data for Testing TRMM Radar Algorithms

Test data are required to test algorithms for the TRMM Precipitation Radar. These data are needed to test the design of the computer codes under development for the operational phase of the mission, and also to test and evaluate alternative or improved precipitation retrieval algorithms. Over a number of years we have developed and used a 3-dimensional radar model for simulating spaceborne precipitation radars. We have adapted this code to produce data files as close as possible to the TRMM file specifications. In this paper, we will describe the model as it is currently implemented, and show some samples of the synthetic data sets.

Jones, Jeffrey A.↗

Small target detection for search and rescue operations using distributed deep learning and synthetic data generation

It is important to find the target as soon as possible for search and rescue operations. Surveillance camera systems and unmanned aerial vehicles (UAVs) are used to support search and rescue. Automatic object detection is important because a person cannot monitor multiple surveillance screens simultaneously for 24 hours. Also, the object is often too small to be recognized by the human eye on the surveillance screen. This study used UAVs around the Port of Houston and fixed surveillance cameras to build an automatic target detection system that supports the US Coast Guard (USCG) to help find targets (e.g., person overboard). We combined image segmentation, enhancement, and convolution neural networks to reduce detection time to detect small targets. We compared the performance between the auto-detection system and the human eye. Our system detected the target within 8 seconds, but the human eye detected the target within 25 seconds. Our systems also used synthetic data generation and data augmentation techniques to improve target detection accuracy. This solution may help the search and rescue operations of the first responders in a timely manner.

Chow, Edward↗

An Investigation Into Possible Systematic Effects on Neutron Star Radius Estimates using NICER-like Synthetic Data

Neutron star cores contain the densest matter in the observable universe. The state of this matter is of interest in numerous fields, but laboratory experiments cannot explore this matter. Although the composition of the matter would be of great interest, macroscopic observables such as the neutron star mass-radius relation depend primarily on the equation of state (EOS). As a result, precise and reliable radius measurements would be valuable in constraining the EOS. However, most attempts at radius measurements are susceptible to systematic errors, meaning the inferred radius can be significantly biased even though the fit to the data appears to be statistically good. Previous studies suggested that radii inferred using X-ray data provided by NASA’s NICER mission may be more immune from such systematic errors. This is in part because, compared with previous measurements that obtained averaged spectra and fluxes, NICER observes millisecond pulsars by timing the photons so precisely, it is possible to obtain the spectrum as a function of rotational phase and see variations such as heated regions on the star rotate into and out of view hundreds of times per second. This extra information seems promising to break degeneracies and mitigate systematic errors, but a more in-depth study is necessary. We report the first steps of that study, in which we generate NICER-like synthetic data and determine the quality of fit and bias in radius obtained when we fit the data using a model different from the model used to generate the data.

Isiah Holt↗

A robust synthetic data generation framework for machine learning in high-resolution transmission electron microscopy (HRTEM)

Machine learning techniques are attractive options for developing highly-accurate analysis tools for nanomaterials characterization, including high-resolution transmission electron microscopy (HRTEM). However, successfully implementing such machine learning tools can be difficult due to the challenges in procuring sufficiently large, high-quality training datasets from experiments. In this work, we introduce Construction Zone, a Python package for rapid generation of complex nanoscale atomic structures which enables fast, systematic sampling of realistic nanomaterial structures and can be used as a random structure generator for large, diverse synthetic datasets. Using Construction Zone, we develop an end-to-end machine learning workflow for training neural network models to analyze experimental atomic resolution HRTEM images on the task of nanoparticle image segmentation purely with simulated databases. Further, we study the data curation process to understand how various aspects of the curated simulated data—including simulation fidelity, the distribution of atomic structures, and the distribution of imaging conditions—affect model performance across three benchmark experimental HRTEM image datasets. Using our workflow, we are able to achieve state-of-the-art segmentation performance on these experimental benchmarks and, further, we discuss robust strategies for consistently achieving high performance with machine learning in experimental settings using purely synthetic data. Construction Zone and its documentation are available at https://github.com/lerandc/construction_zone.

36 MATERIALS SCIENCE↗

Machine Learning Classification of Molten Salt Heat Exchanger Channel Plugging using Synthetic Data

This report addresses the requirements of Milestone M3.4 AI capability to identify and predict maintenance events. Development of digital twins (DT) for molten salt reactor (MSR) components is crucial for reducing operating and maintenance costs (O&M) and ensuring commercial viability of these reactors. Our focus is on development of DT for MSR primary system heat exchanger (HX), a critical component, the fault in which can reduce operating efficiency and force reactor shutdown. We are investigating the feasibility of a conceptual DT of HX consisting of internal distributed temperature sensing with fiber optics and machine learning (ML) algorithms to detect and localize faults. To determine the optimal approach to detection and localization of channel plugging, we benchmark seven different ML models: Logistic Regression, K-Nearest Neighbors (KNN), Gaussian Naïve Bayes, Support Vector Machines (SVM), Decision Tree Classifier, Random Forest Tree Classifier, and Feed-Forward Neural Network. ML algorithms are benchmarked using synthetic HX plugging data generated with computational fluid dynamics COMSOL software, with added brown noise to represent experimental noise. We show that the best performance is obtained with the Decision Tree classifier.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

A machine learning estimator trained on synthetic data for real-time earthquake ground-shaking predictions in Southern California

Abstract After large-magnitude earthquakes, a crucial task for impact assessment is to rapidly and accurately estimate the ground shaking in the affected region. To satisfy real-time constraints, intensity measures are traditionally evaluated with empirical Ground Motion Models that can drastically limit the accuracy of the estimated values. As an alternative, here we present Machine Learning strategies trained on physics-based simulations that require similar evaluation times. We trained and validated the proposed Machine Learning-based Estimator for ground shaking maps with one of the largest existing datasets (<100M simulated seismograms) from CyberShake developed by the Southern California Earthquake Center covering the Los Angeles basin. For a well-tailored synthetic database, our predictions outperform empirical Ground Motion Models provided that the events considered are compatible with the training data. Using the proposed strategy we show significant error reductions not only for synthetic, but also for five real historical earthquakes, relative to empirical Ground Motion Models.

Environmental Sciences & Ecology↗

Quantifying the effect of 3D models on moment tensor results using synthetic data

Moment tensors provide vital information on seismic source properties for seismic events. Moment tensors require seismic wavespeed models to compute Green’s functions, which measure the impulse response between a given source and receiver. Traditionally, researchers have used one-dimensional velocity models to calculate Green’s functions since 1D Green’s functions are computationally cheap to compute. Local 1D velocity models can also accurately model waveforms at short distances (< 500 km). However, 1D velocity models do not account for lateral heterogeneity, which can cause significant misfit in tectonically complex regions such as the Middle East (Covellone and Savage, 2012). Green’s functions calculated using 3D seismic wavespeed models have been shown to perform better in tectonically complex regions (Covellone and Savage, 2012; Kintner and Modrak, 2022), so we are interested in quantifying the effect of considering 3D structure on moment tensor inversion results. The Middle East is an ideal study area for a synthetic moment tensor test for two reasons. Firstly, the Middle East is a tectonically complex region that has been heavily studied. Secondly, the Middle East has significant seismic activity throughout the region, but imperfect station coverage due to limited open data through large swaths of the domain. The tectonic complexity and uneven station coverage will test real-world performance even in a synthetic experiment.

58 GEOSCIENCES↗

Synthetic data generation for machine learning model training for energy theft scenarios using cosimulation

Abstract Technical and non‐technical losses in distribution circuits result in significant economic costs to power utilities. One type of non‐technical loss is energy theft by various means including illegal tapping of feeders, bypassing the meter, and billing fraud. These losses are usually hard to detect, and can remain undetected for long periods of time. Machine learning models have been proven effective in detecting these conditions, but rely on the availability of large, good‐quality training data sets. The problem is exacerbated by the imbalanced nature of data related to these conditions—energy theft, though costly, is very rare. The available data sets generally have very few samples of theft with most of the data pertaining to normal operation. Such data sets are generally not suitable to train machine learning models. In this paper, an overview of energy theft detection techniques, the challenges with their data needs, and the limitations of current techniques to bridge such data limitations is presented. A co‐simulation framework is proposed to generate reliable training data for machine learning algorithms for theft detection. An example scenario is presented and a machine learning model is built to detect certain kinds of energy theft.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Deep residual networks for crystallography trained on synthetic data

The use of artificial intelligence to process diffraction images is challenged by the need to assemble large and precisely designed training data sets. To address this, a codebase called Resonet was developed for synthesizing diffraction data and training residual neural networks on these data. Here, two per-pattern capabilities of Resonet are demonstrated: (i) interpretation of crystal resolution and (ii) identification of overlapping lattices. Resonet was tested across a compilation of diffraction images from synchrotron experiments and X-ray free-electron laser experiments. Crucially, these models readily execute on graphics processing units and can thus significantly outperform conventional algorithms. While Resonet is currently utilized to provide real-time feedback for macromolecular crystallography users at the Stanford Synchrotron Radiation Lightsource, its simple Python-based interface makes it easy to embed in other processing frameworks. This work highlights the utility of physics-based simulation for training deep neural networks and lays the groundwork for the development of additional models to enhance diffraction collection and analysis.

36 MATERIALS SCIENCE↗

Adaptive Cybersecurity for Distributed Energy Resources (AdCyDER): Online Reinforcement Learning with Stackelberg-Optimized Defenses — Pipeline Architecture, Evaluation Methodology, and Findings from a Synthetic-Data Evaluation

This report documents the design and evaluation of an integrated online-learning pipeline developed within the AdCyDER project for Distributed Energy Resource (DER) cybersecurity. The pipeline couples a Reinforcement Learning (RL) attack classifier — which produces an attack-type probability distribution — with a Stackelberg game-theoretic (GT) defense selector that consumes those distributions alongside SME-encoded priors over (defense, attack) effectiveness pairings and perdefense costs to choose grid-health-preserving defenses. The objective is not attack classification per se but production of distributions that drive effective defense selection through the Stackelberg layer, learned from delayed grid-health feedback rather than labeled attack data. AdCyDER as a whole is broader than the work presented here; this report covers the specific RL/GT loop integration and its evaluation. We present the integrated pipeline (SCADA telemetry with Fronius inverter physics, Suricata IDS, time-windowed aggregation, per-facility LSTM classifier, Stackelberg optimizer, OpenC2 actuators), an experimental campaign of 28 eight-hour iterations across three baseline modes, and a pipeline-ordered diagnostic protocol. The protocol identifies two distinct failure modes within the loop: paired supervised ceilings on the same features establish that the deployed online RL classifier (macro F1 ≈ 0.07) sits at least 4.7× below a same-architecture supervised LSTM (≈ 0.34) and 10–11× below a linear feature-signal ceiling (≈ 0.70–0.79 depending on per-facility isolation), localizing the dominant failure to the training procedure; and the reward signal driving online updates carries weak directional coupling with classifier correctness in the methodology-expected direction (multi-lens convergent: top-decile P(true) records produce more frequent state changes and slightly larger improvements, top-vs-bot Cohen’s 𝑑 ≈ −0.19), but at effect magnitudes too small to drive gradient-based learning at the campaign sample size. The original learning hypothesis is not supported by the data. The primary contributions are the diagnostic methodology — proposed as a transferable falsification protocol for online RL/GT defense pipelines learning from delayed environmental reward — and the open, reproducible experimental infrastructure. We outline reward reformulation as the highest-priority aspirational next step given the underpowered-but-aligned Q6 reading, with hardware-in-the-loop evaluation as the broadest scope-expansion option.

Blakely, Benjamin [Argonne National Laboratory (AN↗