Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “synthetic data generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Electron-impact excitation data for W 2+ in support of tungsten spectroscopy and re-deposition measurements for magnetically-confined plasmas

Abstract To better understand plasma wall interactions involving tungsten, accurate atomic structure and electron-impact driven collisional processes for near-neutral ion stages of tungsten are required. Complementing existing work on neutral and singly ionised tungsten, atomic structure and collisional calculations for W 2+ electron-impact excitation have been completed. These excitation calculations are an important component of S/XB coefficients for near-neutral charge states, which may be used to spectroscopically infer re-deposition of tungsten at the plasma-solid boundary of fusion relevant devices. With W 2+ in particular having emission lines that can be observed at ultraviolet (UV) wavelengths, while higher charge states of tungsten are unlikely to have lines possible to observe outside of the vacuum UV range. The atomic structure was generated using the General-purpose Relativistic Atomic Structure Package (GRASP 0 ), implementing the Multi-configuration Dirac Fock approach. This structure was the basis for a subsequent Dirac R -matrix electron-impact excitation calculation to provide Maxwellian averaged rate coefficients. A synthetic spectrum was generated from this data using a collisional-radiative model to predict the strongest W III spectral lines and these lines were compared to emission from the Compact Toroidal Hybrid (CTH) plasma device. Several of the strongest W III lines are observed in CTH and agree well with the modelled line wavelengths and intensities, a table of these lines is provided that could be observed in other devices.

McCann, M. (ORCID:0000000215321240)↗

Synthetic Infrasound Data for Machine Learning Detectors

Synthetic data is a powerful tool to generate large amounts of training data for machine learning models. The methods outlined in this report will be used to retrain the deep learning classifier for increased accuracy. Synthetic data will be useful to address the natural class imbalance between the different categories in the original ML work. Additionally, these tools will be applied for a variety of signal analysis methods that would use signals with a known signal-to-noise ratio for validation and testing.

58 GEOSCIENCES↗

Estimation of Forest Aboveground Biomass from Derivatives of Vegetation-Structure Profiles

Several studies have found that the vertical Fourier transform of lidar, interferometric Synthetic Aperture Radar (SAR), and stereo photogrammetric profiles at empirically-determined spatial frequencies enables high-performance forest aboveground biomass (AGB) estimation. Linear combinations of real and imaginary parts of Fourier transforms of Tomographic (multi-baseline) SAR (TomoSAR) profiles, from Uninhabited Aerial Vehicle Synthetic Aperture Radar (UAVSAR) airborne data, generate ~20%-precision estimates of AGB in the Saskatchewan area of Canada. We found that this 20% precision can be improved to ~15%, a factor of 30% improvement in root mean square error (RMSE) if, in addition to using Fourier transforms of the profile itself, we use Fourier transforms of the spatial, vertical derivative of the profile. The formulation of this "derivative" algorithm is the subject of this paper.

Treuhaft, Robert↗

HITMAN

HITMAN (Hermite Interpolation of Trajectories and Measurement Synthesis for Analysis of Navigators) interpolates—or estimates the unknown values between known values—flight trajectories and generates synthetic inertial measurement unit (IMU) data using Hermite splines. This Python library provides modeling and simulation capabilities to synthesize inertial measurements from discrete trajectory points, enabling researchers to create exemplar datasets for evaluating navigation algorithms in various applications, including consumer devices like smartphones and vehicles. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Walker II, Michael [Sandia National Lab. (SNL-CA),↗

Frontal Slice Approaches for Tensor Linear Systems

Inspired by the row and column action methods for solving large-scale linear systems, in this work, we explore the use of frontal slices for solving tensor linear systems. In particular, this paper presents a novel approach for using frontal slices of a tensor $\mathcal{A}$ to solve tensor linear systems $\mathcal{A} ∗\mathcal{X} = \mathcal{B}$ where ∗ denotes the $t$-product. In addition, we consider variations of this method, including cyclic, block, and randomized approaches, each designed to optimize performance in different operational contexts. Our primary contribution lies in the development and convergence analysis of these methods. Experimental results on synthetically generated and real-world data, including applications such as image and video deblurring, demonstrate the efficacy of our proposed approaches and validate our theoretical findings.

Luo, Hengrui↗

Optimal Asteroid Mass Determination from Planetary Range Observations: A Study of a Simplified Test Model

Mars ranging observations are available over the past 10 years with an accuracy of a few meters. Such precise measurements of the Earth-Mars distance provide valuable constraints on the masses of the asteroids perturbing both planets. Today more than 30 asteroid masses have thus been estimated from planetary ranging data (see [1] and [2]). Obtaining unbiased mass estimations is nevertheless difficult. Various systematic errors can be introduced by imperfect reduction of spacecraft tracking observations to planetary ranging data. The large number of asteroids and the limited a priori knowledge of their masses is also an obstacle for parameter selection. Fitting in a model a mass of a negligible perturber, or on the contrary omitting a significant perturber, will induce important bias in determined asteroid masses. In this communication, we investigate a simplified version of the mass determination problem. Instead of planetary ranging observations from spacecraft or radar data, we consider synthetic ranging observations generated with the INPOP [2] ephemeris for a test model containing ~25000 asteroids. We then suggest a method for optimal parameter selection and estimation in this simplified framework.

asteroids↗

Contrasting Time-Frequency Representations for Unknown Waveform Detection

Identifying unseen electromagnetic waveforms is critical for many applications, like interference management, electronic warfare and spectrum management. Traditionally this is done using statistical methods for anomaly detection, which has evolved to deep learning models for identifying the unseen data, formally termed as open set recognition. Some prior methods use a generative model to emulate open set data, which face challenges in generating synthetic samples for open set while simultaneously selecting an optimal discriminator for accurate classification. To alleviate this issue, we propose a discriminative model that effectively combines time and frequency domain features of communication signals for accurate predictions. We further introduce a cosine similarity loss that makes the domain specific features unique to enhance the prediction rate. Additionally, our model avoids generic feature vectors by extracting class-specific features during training, resulting in improved class representation. The experiment results show that this combined feature approach with cosine loss outperforms single-domain models and improves accuracy by 10% over models without cosine loss.

99 - GENERAL AND MISCELLANEOUS↗

The meandering Gulf Stream as seen by the Geosat altimeter - Surface transport, position, and velocity variance from 73 deg to 46 deg W

Results are presented of an analysis of the surface geostrophic velocity field for the Gulf Stream region for the position, structure, and surface transport of the Gulf Stream for 2.5 yr of the Geosat altimeter Exact Repeat Mission. Synthetic data using a Gaussian velocity profile were generated and fit to the sea surface residual heights to create a synthetic mean sea surface height field and profiles of absolute geostrophic currents. An analysis of the model parameters and the actual geostrophic velocity profiles revealed two different flow regimes for the Gulf Stream connected by a narrow transition region coincident with the New England Seamount Chain. The upstream region was found to exhibit relatively straight Gulf Stream paths, long Eulerian time scales, and eastward propagating meanders. The downstream region had more large meanders, no consistent propagation direction, and shorter Eulerian time scales. A 25-percent reduction in surface transport occurred in the transition region, with a corresponding reduction in current speed and no change in Gulf Stream width.

Kelly, Kathryn A.↗

Augmented Reality Data Generation for Training Deep Learning Neural Network

One of the major challenges in deep learning is retrieving sufficiently large labeled training datasets, which can become expensive and time consuming to collect. A unique approach to training segmentation is to use Deep Neural Network (DNN) models with a minimal amount of initial labeled training samples. The procedure involves creating synthetic data and using image registration to calculate affine transformations to apply to the synthetic data. The method takes a small dataset and generates a highquality augmented reality synthetic dataset with strong variance while maintaining consistency with real cases. Results illustrate segmentation improvements in various target features and increased average target confidence.

Torres, Gil↗

Tree Canopy Characterization for EO-1 Reflective and Thermal Infrared Validation Studies: Rochester, New York

The tree canopy characterization presented herein provided ground and tree canopy data for different types of tree canopies in support of EO-1 reflective and thermal infrared validation studies. These characterization efforts during August and September of 2001 included stem and trunk location surveys, tree structure geometry measurements, meteorology, and leaf area index (LAI) measurements. Measurements were also collected on thermal and reflective spectral properties of leaves, tree bark, leaf litter, soil, and grass. The data presented in this report were used to generate synthetic reflective and thermal infrared scenes and images that were used for the EO-1 Validation Program. The data also were used to evaluate whether the EO-1 ALI reflective channels can be combined with the Landsat-7 ETM+ thermal infrared channel to estimate canopy temperature, and also test the effects of separating the thermal and reflective measurements in time resulting from satellite formation flying.

Ballard, Jerrell R., Jr.↗

MPEX AI Digital Twins

All magnetically confined plasma fusion power plant concepts (Tokamak, Spherical Tokamak, Stellarator, Mirror, ...) must exhaust the heat and plasma from the core confinement region to the material walls. The primary channel for this exhaust is through a plasma divertor which directs plasma along open magnetic field lines to a material target. The Material Plasma Exposure eXperiment (MPEX) illustrated in Figure 1, is a high-power, steady-state linear plasma device designed to produce the plasma material interaction (PMI) conditions of the divertor of future magnetic confinement fusion power plants: energy flux 20MW/m 2 , ion fluence 1031/m 2 , pulse duration 106 sec. These goals of plasma exposure in MPEX are well beyond those achieved in magnetic fusion experimental devices. Successfully achieving these high power steady state conditions for long pulses requires operational control of the heating and particle sources and the plasma flux to the walls and target. The MPEX AI Hot Spot Controller, proposed in this project, will help achieve the operational milestones of MPEX. The MPEX device will begin commissioning at the end of FY26. A smaller proto-MPEX was operated for 14,666 plasma discharges and will resume operation in September of 2025 as proto-MPEX-lite, with reduced capability, to test a new window for the Helicon plasma source. The proto-MPEX data has undergone surrogate modeling with machine learning methods (R. Archibald, 2022 IEEE International Conference on Big Data). This proto-MPEX data will be used to begin development of the AI digital twins described in this white paper. The scientific mission of MPEX is to qualify materials of different composition for use in the high energy and plasma flux conditions of a fusion power plant. The materials exposed in MPEX will in some cases be exposed to high neutron fluxes at other ORNL facilities to measure the changes to their PMI properties. The targets exposed in MPEX will be transported under vacuum to a Surface Analysis Station (SAS). The SAS will be equipped with the following diagnostics: Focused Ion Beam (FIB) for trench milling, 100-400 angstrom resolution scanning electron microscope (SEM), surface mapping x-ray spectrometer, high resolution camera, and a future upgrade to a laser induced breakdown spectroscopy quadruple mass spectrometer (LIBS-QMS). The MPEX experiments will generate diverse pre- and post-exposure measurement data of detailed material properties down to the crystal grain level in 3D for post-exposure assessment of PMI damage (e.g. cracking, melting, erosion and redeposition of the material). Physics models for the PMI, and how the material composition and manufacturing impact its performance under high energy plasma exposure, need to be validated with MPEX data to guide the selection of new candidate materials. Our vision for the MPEX AI Digital Twins project is to supply experimental and physics model simulation data to train Artificial Intelligence (AI) models for data processing, analysis, operational control, PMI and materials simulation to maximize the scientific output of the MPEX device. Ultimately, an AI digital twin of MPEX material assessment metrics for tested and synthetic material types with simulated PMI will be trained by the AI Modeling Teams on the experimental and physics simulation data submitted to the American Science Cloud by this project. A purely empirical search for the best material is inefficient given the finite number of samples that can be tested on MPEX. In order to expand the material properties database for training the MPEX Material Assessment AI Digital Twin, and to gain physics understanding of the PMI processes, physics models of the material properties and PMI processes are required. The physics simulations provide detailed simulation data, like impact angles for plasma ions, sputtering yields, transport of the ionized sputtered target material in the plasma, and redeposition locations. This simulation data expands the measurement data for deeper physics understanding. The experimental data is essential to validate the PMI and material structure simulation models. The validated models can then be used to generate new simulation data of MPEX material assessments for synthetic material compositions that have not been exposed in MPEX. These predictive simulations, plus the whole experimental dataset, will be used to train the MPEX Material Assessment AI Digital Twin allowing a rapid generative AI search for new materials with reduced PMI damage by interpolating the domain of the training set. These new optimum materials can be simulated with the physics codes and/or tested in MPEX. The ability of AI neural networks to interpolate multi-dimensional parameter spaces and generate virtual data is exploited for a more efficient search for optimum materials. The advent of the Transformational AI Models Consortium (TAIMC) is an opportunity to engage with state of the art private and public AI developers to achieve the goals of the AI digital twins and AI accelerated physics models proposed in this project. Our partners at ORNL from the Advance Scientific Computing Research (ASCR) organization will collaborate in accelerating the integrated plasma material interaction simulation framework. This simulation framework will provide a platform for generating simulation data across a range of physical fidelities, including hybrid methods that produce multi-fidelity results. This data will be leveraged for AI model development, both for generation of surrogates and the automation of simulation campaigns. A part of the research below will include collaborative efforts with the TAIMC to (i) adapt data storage approaches to ensure AI-readiness, (ii) provide a protypical exemplar to inform and exercise constructed workflows, and (iii) generate and share data, using the TAIMC unified AI data standard, for foundational models that will be trained from multiple sources across the DOE complex. We will also collaborate with the TAIMC, as well as the planned AI modeling teams, to develop approaches for reducing the cost of data generation. These include tailored multi-fidelity approaches as well as fine-tuning strategies to augment general, large-scale foundational models.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Online and Offline Data Quality Monitoring for the Mu2e Calorimeter

This thesis presents the design, implementation, and validation of a calorimeter Data Quality Monitoring (DQM) toolchain for the Mu2e experiment at Fermilab. Mu2e searches for charged lepton flavor violation via coherent muon-to-electron conversion in the field of an aluminum nucleus, $\mu^- Al \rightarrow e^-Al$, a process whose observation would constitute clear evidence of physics beyond the Standard Model. Achieving target sensitivity requires stringent control of detector performance and data integrity during acquisition, as subtle issues in readout configuration, data formatting, or electronics behavior can compromise reconstruction and bias downstream analyzes. To address these challenges, this work develops a multi-layer DQM approach spanning both raw data validation and reconstructed digi-level diagnostics. At the low level, a fragment analysis component performs word- and bit-field decoding of calorimeter readout blocks, enabling sanity checks of the expected structure and producing detailed error and integrity statistics useful for commissioning and troubleshooting. At the digi level, the CaloDigiDQM analyzer is implemented within the art framework and transforms each CaloDigiCollection into a structured hierarchy of ROOT histograms designed for fast drill-down diagnostics. The module generates coherent monitoring views at global, disk, board, and channel granularity, including occupancy, waveform-derived features (baseline, RMS, peak amplitude and position), and left-right sensor consistency metrics. Detector-aware channel-to-electronics mapping is performed through the conditions system (CaloDAQMap), ensuring that diagnostics remain aligned with hardware identifiers used in operations. For end-to-end testing without reliance on live DAQ data, a synthetic CaloDigi producer is developed to generate realistic waveforms with controlled noise and pulse shapes. The resulting system supports both offline ROOT-file production and online operation, including optional histogram streaming through otsdaq via ots::HistoSender. This toolchain provides a practical and scalable foundation for calorimeter commissioning and stable data collection, enabling early detection of anomalies and reducing operational risk for Mu2e.

Vakulenko, Mark [Drew U.] (ORCID:0009000276197818)↗

Discovery of Activities via Statistical Clustering of Fixation Patterns

Human behavior often consists of a series of distinct activities, each characterized by a unique pattern of interaction with the visual environment. This is true even in a restricted domain, such as a piloting an aircraft, where activities with distinct visual signatures might be things like communicating, navigating, and monitoring. We propose a novel analysis method for gaze-tracking data, to perform blind discovery of these hypothetical activities. The method is in some respects similar to recurrence analysis, but here we compare not individual fixations, but groups of fixations aggregated over a fixed time interval. The duration of this interval is a parameter that we will refer to as delta. We assume that the environment has been divided into a set of N different areas-of-interest (AOIs). For a given interval of time of duration delta, we compute the proportion of time spent fixating each AOI, resulting in an N-dimensional vector. These proportions can be converted to integer counts by multiplying by delta divided by the average fixation duration (another parameter that we fix at 280 milliseconds). We compare different intervals by computing the chi-square statistic. The p-value associated with the statistic is the likelihood of observing the data under the hypothesis that the data in the two intervals were generated by a single process with a single set of probabilities governing the fixation of each AOI. The method has been applied to approximately 100 hours of eye movement data collected from pilots in a high-fidelity B747 flight simulator, and the results have been compared to synthetic data in which the each activity is represented as first-order Markov process with random probabilities assigned to the AOIs. Randomly-generated synthetic activities can require thousands of fixations to be discriminated with statistical significance, while the human data can be clustered using averaging windows of some 10's of seconds, suggesting that the actual activities are much more narrowly focused than random Markov models.

activity analysis↗

Design and Characterization of a Transcriptional Repression Toolkit for Plants

Regulation of gene expression is essential for all life. Tools to manipulate the gene expression level have therefore proven to be very valuable in efforts to engineer biological systems. However, there are few well-characterized genetic parts that reduce gene expression in plants, commonly known as transcriptional repressors. We characterized the repression activity of a library consisting of repression motifs from approximately 25% of the members of the largest known family of repressors. Combining sequence information with our trans-regulatory function data, we next generated a library of synthetic transcriptional repression motifs with function predicted in advance. After characterizing our synthetic library, we demonstrated not only that many of our synthetic constructs were functional as repressors but also that our advance predictions of repression strength were better than random guesses. Finally, we assessed the functionality of known transcriptional repression motifs from a wide range of eukaryotes. Our study represents the largest plant repressor motif library experimentally characterized to date, providing unique opportunities for tuning transcription in plants.

59 BASIC BIOLOGICAL SCIENCES↗

SOC Microstructural Property Estimator

This pre-trained ML model is a tool that uses basic compositional parameters for porous solid oxide cell (SOC) electrodes - the phase fractions and mean particle/pore diameters – as inputs and uses them to estimate additional electrochemical performance parameters: active (i.e., connected) TPB density, all tortuosity factors, and phase pair specific interfacial areas. The electrode is assumed to be composed of two solid phases and a pore phase. The property calculations are performed using neural network regression models trained on a large bank of synthetic electrode microstructural data that NETL has generated using the program DREAM3D (that bank is also hosted on EDX: https://edx.netl.doe.gov/dataset/soc-synthetic-microstructure-bank). This means the generated parameters are based on training from actual measured properties from 3D microstructures, not estimated from geometric simplifications. This tool was developed and is intended to replace percolation theory calculations in models that use hypothetical electrode properties. An example use case would be running SOC performance simulations across a parametric sweep of electrode designs (e.g., varying phase fractions and particle sizes) and assessing how it impacts the electrochemical performance of the SOC. Within the parameter space of the training data (statistics of that parameter space is provided in the readme file), this model achieves sub-5% mean absolute percent errors, an order of magnitude less error than percolation theory across the same parameter space. However, be aware that this tool was developed with parametric simulations in mind, and users are encouraged to assess accuracy for their own specific use case rather than taking accuracy metrics at face value. More info, including a usage guide, is in the included readme file. This tool should be cited with the DOI number provided.

Electrode Microstructure↗

Adaptation of the ISCCP cloud detection algorithm to combined AVHRR and SMMR arctic data

The International Satellite Cloud Climatology Project (ISCCP) cloud detection algorithm is applied to artic data, and modifications are suggested. Both Advanced Very High Resolution Radiometer (AVHRR) and Scanning Multichannel Microwave Radiometer (SMMR) data are examined. Synthetic AVHRR and SMMR data are also generated. Modifications suggested include the use of snow and ice data sets for the estimation of surface parameters, additional AVHRR channels, and surface class characteristic values when clear sky values cannot be obtained. Greatest improvement in computed cloud fraction is realized over snow and ice surfaces; over other surfaces all versions perform similarly. Since the use of SMMR for surface analysis increases the computational burden, its use may be justified only over snow and ice-covered regions.

Key, J.↗

Target Optimal Aperture Selection

This document describes the computation of optimal pixels for planetary transit targets. The method we describe is based on that used in the End-to-End Model (ETEM), which simulates Kepler science output. An optimal aperture for a target is defined as the set of pixels which maximizes the signal-to-noise ratio (SNR) for that target. The optimal aperture for a target is determined by using catalog data to generate, for that target, a synthetic image with all background stars and signals and a second synthetic image with the target star's flux only. These images are incorporated into a noise model, from which the SNR of each pixel is computed. The pixels around a target are summed in an order that maximizes the SNR with each term. As dimmer and dimmer pixels contribute to this sum, the SNR reaches a maximum and further terms decrease the SNR. The set of pixels whose SNR sums to this maximum value defines the optimal aperture.

Bryson, Stephen↗

Aeroacoustic Study of a Subscale Large Civil Transport (STAR) Model – Part 2: Validation of Simulated Results

Aeroacoustic measurements of the 26%-scale, semispan Boeing 777-200 Subsonic Transport Aeroacoustic Research (STAR) model tested in the NASA Ames Research Center 40- by 80-foot wind tunnel were used to ascertain the efficacy of high-fidelity simulations to accurately predict noise from the landing gear of large commercial transports. The simulations, conducted with the lattice Boltzmann solver PowerFLOW®, used a digital replica of the STAR model with or without main landing gear deployed and slats and flaps set to their highest deflection angles to represent aircraft during landing. The computations were performed at a Mach number of 0.21, Reynolds number of 8.2 million based on the model mean aerodynamic chord, and other conditions prevalent during the STAR model test. Measured and computed surface pressures were in very good agreement at most port locations on the model, as were global force coefficients, indicating that the simulations captured the impact of main gear deployment on inboard flap loading. Noise sources produced by the main landing gear and high-lift devices were determined via source localization maps generated with CLEAN from synthetic and experimental data. In general, very good agreement between predicted and measured acoustic source location and relative strength was observed in the maps. Comparisons of far-field noise spectra obtained from the CLEAN deconvolution maps showed remarkable agreement between synthetic and experimental broadband noise at low and medium frequencies. Main landing gear sources for model-scale frequencies above 7,000 Hz could not be resolved with the spatial resolution used during the simulations.

airframe noise↗