Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Synthetic Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

An empirical approach to model selection: weak lensing and intrinsic alignments

ABSTRACT In cosmology, we routinely choose between models to describe our data, and can incur biases due to insufficient models or lose constraining power with overly complex models. In this paper, we propose an empirical approach to model selection that explicitly balances parameter bias against model complexity. Our method uses synthetic data to calibrate the relation between bias and the χ2 difference between models. This allows us to interpret χ2 values obtained from real data (even if catalogues are blinded) and choose a model accordingly. We apply our method to the problem of intrinsic alignments – one of the most significant weak lensing systematics, and a major contributor to the error budget in modern lensing surveys. Specifically, we consider the example of the Dark Energy Survey Year 3 (DES Y3), and compare the commonly used non-linear alignment (NLA) and tidal alignment and tidal torque (TATT) models. The models are calibrated against bias in the Ωm–S8 plane. Once noise is accounted for, we find that it is possible to set a threshold Δχ2 that guarantees an analysis using NLA is unbiased at some specified level Nσ and confidence level. By contrast, we find that theoretically defined thresholds (based on, e.g. p-values for χ2) tend to be overly optimistic, and do not reliably rule out cosmological biases up to ∼1–2σ. Considering the real DES Y3 cosmic shear results, based on the reported difference in χ2 from NLA and TATT analyses, we find a roughly $30{{\ \rm per\ cent}}$ chance that were NLA to be the fiducial model, the results would be biased (in the Ωm–S8 plane) by more than 0.3σ. More broadly, the method we propose here is simple and general, and requires a relatively low level of resources. We foresee applications to future analyses as a model selection tool in many contexts.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Forest biomass, canopy structure, and species composition relationships with multipolarization L-band synthetic aperture radar data

The effect of forest biomass, canopy structure, and species composition on L-band synthetic aperature radar data at 44 southern Mississippi bottomland hardwood and pine-hardwood forest sites was investigated. Cross-polarization mean digital values for pine forests were significantly correlated with green weight biomass and stand structure. Multiple linear regression with five forest structure variables provided a better integrated measure of canopy roughness and produced highly significant correlation coefficients for hardwood forests using HV/VV ratio only. Differences in biomass levels and canopy structure, including branching patterns and vertical canopy stratification, were important sources of volume scatter affecting multipolarization radar data. Standardized correction techniques and calibration of aircraft data, in addition to development of canopy models, are recommended for future investigations of forest biomass and structure using synthetic aperture radar.

Sader, Steven A.↗

Ice Phase Classification Made Easy with Score-Based Denoising

Accurate identification of ice phases is essential for understanding various physicochemical phenomena. However, such classification for structures simulated with molecular dynamics is complicated by the complex symmetries of ice polymorphs and thermal fluctuations. For this purpose, both traditional order parameters and data-driven machine learning approaches have been employed, but they often rely on expert intuition, specific geometric information, or large training data sets. In this work, we present an unsupervised phase classification framework that combines a score-based denoiser model with a subsequent model-free classification method to accurately identify ice phases. Further, the denoiser model is trained on perturbed synthetic data of ideal reference structures, eliminating the need for large data sets and labeling efforts. The classification step utilizes the smooth overlap of atomic position (SOAP) descriptors as the atomic fingerprint, ensuring Euclidean symmetries and transferability to various structural systems. Our approach achieves a remarkable 100% accuracy in distinguishing ice phases of test trajectories using only seven ideal reference structures of ice phases as model inputs. This demonstrates the generalizability of the score-based denoiser model in facilitating phase identification for complex molecular systems. The proposed classification strategy can be broadly applied to investigate structural evolution and phase identification for a wide range of materials, offering new insights into the fundamental understanding of water and other complex systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Constraining physical models at gigabar pressures

High-energy-density (HED) experiments in convergent geometry are able to test physical models at pressures beyond hundreds of millions of atmospheres. The measurements from these experiments are generally highly integrated and require unique analysis techniques to procure quantitative infor mation. This work describes a methodology to constrain the physics in convergent HED experiments by adapting the methods common to many other fields of physics. As an example, a mechanical model of an imploding shell is constrained by data from a thin-shelled direct-drive exploding-pusher experiment on the OMEGA Laser System using Bayesian inference, resulting in the reconstruction of the shell dynamics and energy transfer during the implosion. The model is tested by analyzing synthetic data from a 1-D hydrodynamics code and is sampled using a Markov chain Monte Carlo to generate the posterior distributions of the model parameters. As a result, the goal of this work is to demonstrate a general methodology that can be used to draw conclusions from a wide variety of HED experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Detection of Outliers in LiDAR Data Acquired by Multiple Platforms over Sorghum and Maize

High-resolution point cloud data acquired with a laser scanner from any platform contain random noise and outliers. Therefore, outlier detection in LiDAR data is often necessary prior to analysis. Applications in agriculture are particularly challenging, as there is typically no prior knowledge of the statistical distribution of points, plant complexity, and local point densities, which are crop-dependent. The goals of this study were first to investigate approaches to minimize the impact of outliers on LiDAR acquired over agricultural row crops, and specifically for sorghum and maize breeding experiments, by an unmanned aerial vehicle (UAV) and a wheel-based ground platform; second, to evaluate the impact of existing outliers in the datasets on leaf area index (LAI) prediction using LiDAR data. Two methods were investigated to detect and remove the outliers from the plant datasets. The first was based on surface fitting to noisy point cloud data via normal and curvature estimation in a local neighborhood. The second utilized the PointCleanNet deep learning framework. Both methods were applied to individual plants and field-based datasets. To evaluate the method, an F-score was calculated for synthetic data in the controlled conditions, and LAI, the variable being predicted, was computed both before and after outlier removal for both scenarios. Results indicate that the deep learning method for outlier detection is more robust than the geometric approach to changes in point densities, level of noise, and shapes. The prediction of LAI was also improved for the wheel-based vehicle data based on the coefficient of determination (R2) and the root mean squared error (RMSE) of the residuals before and after the removal of outliers.

36 MATERIALS SCIENCE↗

Constraining Physical Models at Gigabar Pressures

High-energy-density (HED) experiments in convergent geometry are able to test physical models at pressures beyond hundreds of millions of atmospheres. The measurements from these experiments are generally highly integrated and require unique analysis techniques to procure quantitative information. This work describes a methodology to constrain the physics in convergent HED experiments by adapting the methods common to many other fields of physics. As an example, a mechanical model of an imploding shell is constrained by data from a thin-shelled direct-drive exploding-pusher experiment on the OMEGA Laser System using Bayesian inference, resulting in the reconstruction of the shell dynamics and energy transfer during the implosion. The model is tested by analyzing synthetic data from a 1-D hydrodynamics code and is sampled using a Markov chain Monte Carlo to generate the posterior distributions of the model parameters. The goal of this work is to demonstrate a general methodology that can be used to draw conclusions from a wide variety of HED experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Basic Diagnosis and Prediction of Persistent Contrail Occurrence using High-resolution Numerical Weather Analyses/Forecasts and Logistic Regression. Part I: Effects of Random Error

Straightforward application of the Schmidt-Appleman contrail formation criteria to diagnose persistent contrail occurrence from numerical weather prediction data is hindered by significant bias errors in the upper tropospheric humidity. Logistic models of contrail occurrence have been proposed to overcome this problem, but basic questions remain about how random measurement error may affect their accuracy. A set of 5000 synthetic contrail observations is created to study the effects of random error in these probabilistic models. The simulated observations are based on distributions of temperature, humidity, and vertical velocity derived from Advanced Regional Prediction System (ARPS) weather analyses. The logistic models created from the simulated observations were evaluated using two common statistical measures of model accuracy, the percent correct (PC) and the Hanssen-Kuipers discriminant (HKD). To convert the probabilistic results of the logistic models into a dichotomous yes/no choice suitable for the statistical measures, two critical probability thresholds are considered. The HKD scores are higher when the climatological frequency of contrail occurrence is used as the critical threshold, while the PC scores are higher when the critical probability threshold is 0.5. For both thresholds, typical random errors in temperature, relative humidity, and vertical velocity are found to be small enough to allow for accurate logistic models of contrail occurrence. The accuracy of the models developed from synthetic data is over 85 percent for both the prediction of contrail occurrence and non-occurrence, although in practice, larger errors would be anticipated.

Duda, David P.↗

Data Science and Urban Air Mobility: Challenges and Opportunities

Aviation is broadly a combination of aircraft, airspace and airports. The data science life cycle comprises of five steps - capture, maintain, process, analyze and communicate. The presentation introduces the legacy of conventional aviation research in the context of the data science life cycle to motivate the challenges with Urban Air Mobility, a field that is quite nascent. A summary of recent research will be presented to highlight the innovative ways to address the challenges. Examples provided will include the generation of synthetic data, encounter models from simulations, and leveraging novel and diverse data sets from traditional transportation and non-aviation sources, to analyze problems of operation in urban airspace. Finally, opportunities will be identified for further exploration, niche development and filling the gaps in the field of data science for UAM.

Urban Air Mobility↗

Data Science and Urban Air Mobility: Challenges and Opportunities

Aviation is broadly a combination of aircraft, airspace and airports. The data science life cycle comprises of five steps - capture, maintain, process, analyze and communicate. The presentation introduces the legacy of conventional aviation research in the context of the data science life cycle to motivate the challenges with Urban Air Mobility, a field that is quite nascent. A summary of recent research will be presented to highlight the innovative ways to address the challenges. Examples provided will include the generation of synthetic data, encounter models from simulations, and leveraging novel and diverse data sets from traditional transportation and non-aviation sources, to analyze problems of operation in urban airspace. Finally, opportunities will be identified for further exploration, niche development and filling the gaps in the field of data science for UAM.

Urban Air Mobility↗

Data Science Challenges for Urban Air Mobility

Aviation is a combination of aircraft, airspace and airports. The data science life cycle comprises of five steps - capture, maintain, process, analyze and communicate. The presentation introduces the legacy of conventional aviation research in the context of the data science life cycle to motivate the challenges with Urban Air Mobility, a field that is quite nascent. A summary of recent research will be presented to highlight the innovative ways to address the challenges. Examples provided will include the generation of synthetic data, encounter models from simulations, and leveraging novel and diverse data sets from traditional transportation and non-aviation sources, to analyze problems of operation in urban airspace. Finally, opportunities will be identified for further exploration, niche development and filling the gaps in the field of data science for UAM.

Data Science↗

Data Science Challenges for Urban Air Mobility

Aviation is a combination of aircraft, airspace and airports. The data science life cycle comprises of five steps - capture, maintain, process, analyze and communicate. The presentation introduces the legacy of conventional aviation research in the context of the data science life cycle to motivate the challenges with Urban Air Mobility, a field that is quite nascent. A summary of recent research will be presented to highlight the innovative ways to address the challenges. Examples provided will include the generation of synthetic data, encounter models from simulations, and leveraging novel and diverse data sets from traditional transportation and non-aviation sources, to analyze problems of operation in urban airspace. Finally, opportunities will be identified for further exploration, niche development and filling the gaps in the field of data science for UAM.

Data Science↗

Machine learning based unfolding of x-ray spectra from filter stack spectrometer data

We demonstrate the application of neural networks to perform x-ray spectra unfolding from data collected by filter stack spectrometers. A filter stack spectrometer consists of a series of filter-detector pairs, where the detectors behind each filter measure the energy deposition through each layer as photo-stimulated luminescence (PSL). The network is trained on synthetic data, assuming x-rays of energies < 1 MeV and of two different distribution functions (Maxwellian and Gaussian) and the corresponding measured PSL values obtained from five different filter stack spectrometer designs. Predicted unfolds of single distributions are near identical reproductions of the ground truth spectra, with differences in the values lower than 20% at the higher energy end in some cases. The neural network has also demonstrated robustness to experimental measurement errors of < 5% and some capability of performing unfolds for linear combinations of the two distributions without previous training. The network can perform unfolds at rates > 1 Hz, ideal for application to some high-repetition-rate systems.

47 OTHER INSTRUMENTATION↗

Interpolation Models and Error Bounds for Verifiable Scientific Machine Learning

This repository contains python scripts and numerical data accompanying the paper: "Leveraging Interpolation Models and Error Bounds for Verifiable Scientific Machine Learning," Tyler Chang, Andrew Gillette, Romit Maulik, 2024. The following subdirectories are included: - "interpolants" contains our interpolation scripts used for all studies - "experiments" contains scripts demonstrating our experiments with synthetic data - "airfoil" contains scripts demonstrating our experiments with the publicly available UIUC airfoil dataset. Further instructions are provided in READMEs within the sub-directories.

Gillette, Andrew↗

Federated Machine Learning-Based Anomaly Detection System for Synchrophasor Network Using Heterogeneous Data Sets: Preprint

Synchrophasor technology is widely deployed in the energy management system to monitor the grid health at micro level and perform necessary corrective actions in real time; however, integrated phasor devices and data aggregators are exposed to several cybersecurity threats. This paper proposes a federated ML(FML)-based ADS to detect several data integrity attacks in the synchrophasor network. The proposed approach integrates the horizontal FML technique and consists of substation-based local models and a control center-based global model. The proposed methodology includes training local models using heterogeneous data sets that include network and grid information and updating the global model through multiple iterations by sharing model gradients. Finally, the trained global model is applied to identify cyberattacks, normal operation, and physical events. To validate the proof of concept, we used synthetic data sets generated by Mississippi State University and Oak Ridge National Laboratory for training and testing the classification models using the National Renewable Energy Laboratory's high performance computing resources. Our experimental results, computed through several performance measures, reveal that the proposed approach shows consistent performance during the binary, three-class, and multiclass classifications while ensuring privacy of synchrophasor data.

anomaly detection system↗

Evaluation of Methods for Causal Discovery in Hydrometeorological Systems

Understanding causal relations is of utmost importance in hydrology and climate research for systems identification, prediction, and understanding systems behavior in a changing climate. Traditionally, researchers in hydrometeorology attempted to study causal questions by conducting controlled experiments using numerical models. This approach, however, in most cases of interest provides uncertain results because the models are approximate representation of the natural system. An alternative approach that has recently drawn significant attention in several fields is to infer causal relations from purely observational data. It possesses several traits to its utility particularly in hydrometeorology due to the rapid accumulation of in situ and remotely sensed data records. The first objective of this study is to present a brief description of four causal discovery methods (Granger causality, Transfer Entropy, graph-based algorithms, and Convergent Cross Mapping) with special emphasis on the assumptions on which they are built. Second, using synthetic data generated from a hydrological model, we assess their performance in retrieving causal information taking into account sensitivity to sample size and presence of noise. Last, we use causal analysis to examine and formulate hypotheses on causal drivers of evapotranspiration in a shrubland region during summer and winter seasons. An interpretation of the hypotheses based on canopy seasonal dynamics and evapotranspiration processes is presented. It is hoped that the results presented here can be useful in guiding researchers studying hydrometeorological systems as to which causal method is most appropriate to the characteristics of the system under study.

54 ENVIRONMENTAL SCIENCES↗

A Markov chain Monte Carlo (MCMC) Bayesian inference approach to analyze apparent activation barriers and reaction orders from microreactor data

Statistical analysis of steady-state catalytic kinetic data is often limited by data sparsity due to the slow pace at which the data is collected. Data sparsity and limitations in statistical analysis make it difficult to differentiate between mechanistic models and catalytic sites. A Bayesian inference tool is reported for catalysis researchers to estimate error in the determination of reaction orders from steady state microreactor data. The benefits of a Bayesian inference approach are discussed, as an alternative to the more common frequentist approach. The approach incorporates prior knowledge of the system and the data collected to form an error estimate on reaction orders. We investigated the effects of three distinct data treatments—individual fitting of trials, pooled analysis, and constrained regression methods—on the precision and uncertainty of reaction order determinations. To assess the robustness of our findings, we conducted sensitivity analyses to evaluate the influence of Bayesian parameters on uncertainty estimation. Additionally, we utilized synthetic data to illustrate how data quality impacts the precision of uncertainty assessments. We show Bayesian analysis can obtain a more precise estimation of error with a sparse data set than a frequentist analysis. Finally, this work provides strong evidence that the adoption of Bayesian analysis of kinetic data may help researchers make more precise arguments as to the strength of their evidence for a particular mechanistic hypothesis, or in comparing across different catalysts.

42 ENGINEERING↗

A comprehensive guide to CAN IDS data and introduction of the ROAD dataset

Although ubiquitous in modern vehicles, Controller Area Networks (CANs) lack basic security properties and are easily exploitable. A rapidly growing field of CAN security research has emerged that seeks to detect intrusions or anomalies on CANs. Producing vehicular CAN data with a variety of intrusions is a difficult task for most researchers as it requires expensive assets and deep expertise. To illuminate this task, we introduce the first comprehensive guide to the existing open CAN intrusion detection system (IDS) datasets. We categorize attacks on CANs including fabrication (adding frames, e.g., flooding or targeting and ID), suspension (removing an ID’s frames), and masquerade attacks (spoofed frames sent in lieu of suspended ones). We provide a quality analysis of each dataset; an enumeration of each datasets’ attacks, benefits, and drawbacks; categorization as real vs. simulated CAN data and real vs. simulated attacks; whether the data is raw CAN data or signal-translated; number of vehicles/CANs; quantity in terms of time; and finally a suggested use case of each dataset. State-of-the-art public CAN IDS datasets are limited to real fabrication (simple message injection) attacks and simulated attacks often in synthetic data, lacking fidelity. In general, the physical effects of attacks on the vehicle are not verified in the available datasets. Only one dataset provides signal-translated data but is missing a corresponding “raw” binary version. This issue pigeon-holes CAN IDS research into testing on limited and often inappropriate data (usually with attacks that are too easily detectable to truly test the method). The scarcity of appropriate data has stymied comparability and reproducibility of results for researchers. As our primary contribution, we present the Real ORNL Automotive Dynamometer (ROAD) CAN IDS dataset, consisting of over 3.5 hours of one vehicle’s CAN data. ROAD contains ambient data recorded during a diverse set of activities, and attacks of increasing stealth with multiple variants and instances of real (i.e. non-simulated) fuzzing, fabrication, unique advanced attacks, and simulated masquerade attacks. To facilitate a benchmark for CAN IDS methods that require signal-translated inputs, we also provide the signal time series format for many of the CAN captures. Our contributions aim to facilitate appropriate benchmarking and needed comparability in the CAN IDS research field.

97 MATHEMATICS AND COMPUTING↗

Development of a deep learning based automated data analysis for step-filter x-ray spectrometers in support of high-repetition rate short-pulse laser-driven acceleration experiments

We present a deep learning based framework for real-time analysis of a differential filter based x-ray spectrometer that is common on short-pulse laser experiments. The analysis framework was trained with a large repository of synthetic data to retrieve key experimental metrics, such as slope temperature. With traditional analysis methods, these quantities would have to be extracted from data using a time-intensive and manual analysis. Furthermore, this framework was developed for a specific diagnostic, but may be applicable to a wide variety of diagnostics common to laser experiments and thus will be especially crucial to the development of high-repetition rate (HRR) diagnostics for HRR laser systems that are coming online.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗