Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “evaluation datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

A Comparison of Passive Microwave Emission Models for Estimating Brightness Temperature at L- and P-band Under Bare and Vegetated Soil Conditions

P-band radiometry has been demonstrated to have a deeper sensing depth than at L-band, making the consideration of multi-layer microwave interactions necessary. Additionally, the scattering and phase interference effects are different at P-band, requiring a re-consideration of the need for coherent models. However, the impact remains to be clarified, and understanding the validity and limitations of these models at both L-band and P-band is crucial for their refinement and application. Therefore, two general categories of microwave emission models, including two stratified coherent models (Njoku and Wilhite) and four incoherent models (conventional tau-omega model and three multi-layer models being zero-order, first-order, and incoherent solution), were intercompared for the first time on the same dataset. This evaluation utilized observations of L-band and P-band radiometry under different land cover conditions from a tower-based experiment in Victoria, Australia. Model estimations of brightness temperature (TB) were consistent with measurements, with the lowest root mean square error (RMSE) at P-band V-polarization under corn (2 K) and the highest RMSE at L-band H-polarization under bare soil (13 K). Coherent models performed slightly better than incoherent models under bare soil (3 K less RMSE), while the opposite was true under vegetated soil conditions (1 K less RMSE). Coherent and incoherent models showed maximum differences (3 K at P-band, 2 K at L-band), correlating strongly with soil moisture variations at 0-10 cm. Findings suggest that coherent and incoherent models perform similarly; thus, incoherent models may be preferable for estimating TB at L- and P-band due to reduced computational complexity.

Soil moisture profile↗

Real-Time Anomaly Detection for Searches Beyond the Standard Model in the ProtoDUNE Horizontal Drift Detector

This paper summarizes work conducted throughout a SULI internship at Fermi National Accelerator Laboratory focused on building an unsupervised machine learning model for real-time anomaly detection in ProtoDUNE Horizontal Drift. Using simulated data, we trained an autoencoder model on a pure cosmic dataset, and evaluated it on both cosmic and neutrino events—making the model an anomaly detector. The goal was to make a model which matches or exceeds the current ADC Simple Window trigger algorithm so that our model can perform at the same rate but provide sensitivity to potential beyond-the-Standard-Model (BSM) signatures. In the end, we were able to construct a model which slightly exceeds the capabilities of the ADC Simple Window while remaining completely unsupervised, achieving 31.9 ± 0.2% (26.6 ± 0.2%) ν efficiency at 5 Hz (2 Hz), a 3.6 (3.2) percentage point increase. Additionally, 17.5 ± 0.3% (18.3 ± 0.3%) of the events that passed the autoencoder at 5 Hz (2 Hz) were missed by the current trigger algorithm. Future work will investigate alternative normalization methods, including quantile transformation, and evaluate the model on ProtoDUNE-HD detector-glitch data if that data becomes available.

Wilson, C. [Cincinnati U., RWC]↗

VIC-Global Parameter Dataset Sensitivity with the Variable Infiltration Capacity Model: Evaluating the importance of dynamic land surface parameters when using the VIC-Global parameter dataset

Accurate prediction of runoff is essential to water resources management, flood risk assessment, and ecosystem protection. However, many hydrological models still have relatively substantial limitations when representing the influence of land use and land cover (LULC) on runoff generation and routing. Changes in LULC, such as deforestation, urban expansion, agricultural intensification, and wetland loss, have been shown to alter the water balance at the land surface through fundamental hydrologic processes (e.g., interception, infiltration, evapotranspiration, and soil storage). However, it remains an open question what the exact magnitude and timing of these impacts are for the spatial and temporal scales commonly used in engineering applications. In this analysis we focus on one aspect of recent LULC change for assessing human impacts, which is urbanization. Specifically we seek to determine the impacts of urbanization on the magnitude and timing of surface runoff and baseflow in HUC-12 basins in Clark County, Nevada which has experienced rapid urbanization. We use the Variable Infiltration Capacity (VIC) hydrology model with a widely used off-the-shelf dataset of land surface parameters, VIC-Global, both of which have been commonly used in the past for water and energy balance modeling for large scale hydrologic studies. We examine two scenarios where the first scenario removes all urbanized land cover and parameterizes those areas of the basins as barren or open shrubland. The second scenario tests the opposite case where all areas of the basins are classified as urban regardless of their present classification. The results from the VIC model show there is a low sensitivity for daily surface runoff between scenarios. The daily baseflow values indicate similar low sensitivity to the classification change during specific periods, but then have substantial differences during other period when large precipitation events are occurring. This is likely due to the assumed parameter values for the urban land cover classification made by the VIC-Global dataset. Using a static land cover parameterization is reasonable for large domain hydrology models that are being used for near-term planning horizons (<30 years). However, longer planning horizons where feedbacks between the atmosphere and land surface are important, especially in transient climate situations, considerations for how to update land surface parameters should be incorporated.

42 ENGINEERING↗

Model Assumptions and Data Characteristics: Impacts on Domain Adaptation in Building Segmentation

Studies on domain adaptation (DA) for remote sensing (RS) imagery analysis lack consistency in selection and description of evaluation scenarios. Without properly characterizing datasets, model assumptions, and evaluation scenarios, it is difficult to objectively compare DA methods and reach conclusions about their suitability across different applications. With this motivation, this work seeks to empirically assess to which extent the interaction between data characteristics and model assumptions influences the effectiveness of DA methods. Using the widely explored task of building footprint segmentation as a case study, we perform a large-scale study across over 200 DA scenarios that include variations across view angles, areas observed, and sensors used for data acquisition. Rather than adopting different model architectures or optimization criteria, we contrast the performances of two DA methods based on adversarial learning that differ only in their assumptions about source and target domains. Informed by metadata and data characteristics unveiled using traditional computer vision (CV) techniques as well as pretrained deep models, we provide a detailed meta-analysis of experiments highlighting the importance of accurately considering data assumptions for DA in RS segmentation tasks. As demonstrated by a “cherry-picking” exercise, different claims regarding which model is best could be made by selecting different subsets of evaluation scenarios. While well-calibrated assumptions can be beneficial, mismatching assumptions can lead to negative biases in DA applications. Furthermore, this study intends to motivate the community toward more consistent evaluation protocols while providing recommendations and insights toward creating novel benchmark datasets, documenting data characteristics, application-specific knowledge, and model assumptions.

42 ENGINEERING↗

Improving Earth Science Dataset Search with Publication

The NASA Goddard Earth Sciences Data and Information Services Center (GESDISC) archives a large number of Earth observational datasets. Thousands of the publications are created each year based on these datasets. The content of these publications can be used for discovery of the datasets based on the characteristics of applicational research. We leverage the content of these publications to retrieve the information about phenomena and domains where measurements from the datasets were utilized through linking these publications and dataset in Knowledge Graph. We retrieve phenomena and domain information using SWEET ontology and produce the set of keywords that are linked to the datasets. Further, we evaluate this link strength according to the frequency of dataset usage in the papers mentioning these keywords. We demonstrate how this linkage can improve dataset search by comparing the search results obtained from Common Metadata Repository (CMR) search and the publications based data.

Kristina Stoyanova↗

The HydroBio Dataset: a new data resource for evaluating existing and potential hydropower capacity and freshwater biodiversity in the conterminous United States

Hydropower is a critical source of affordable and reliable electricity and energy system stability services in the United States. Opportunities to expand US hydropower production include retrofitting existing non-powered dams to produce power, retrofitting existing hydropower dams to improve efficiency or increase capacity, or constructing new hydropower infrastructure on currently unregulated river reaches. We created the HydroBio Dataset, which summarizes existing and potential hydropower capacity and freshwater biodiversity at the sub-basin scale in the conterminous US to contextualize existing and potential grid contributions with the freshwater ecosystems in which dams are situated. We demonstrate a use-case of this dataset by rescaling and comparing potential non-powered dam nominal capacity to rarity-threat-weighted freshwater species richness for sub-basins where both types of data exist. On average, normalized freshwater biodiversity exceeded normalized potential non-powered dam nominal capacity in these sub-basins. Potential non-powered dam nominal capacity was concentrated in sub-basins in the Upper Mississippi and Ohio hydrologic regions while freshwater biodiversity was concentrated in the South Atlantic-Gulf, Ohio, and Tennessee hydrologic regions. Additionally, non-powered dams and existing hydropower dams are located in sub-basins with similar indices of freshwater biodiversity. The HydroBio Dataset adds an additional ecological dimension of context to our understanding of current and potential future US hydropower capabilities and is a valuable decision support tool for stakeholders tasked with balancing gains in services to the US power grid with the public and environmental benefits of freshwater ecosystems.

Biodiversity↗

CyberGAN: Generating High-fidelity Cybersecurity Data With Generative Adversarial Networks

Machine learning for cyber defense offers the promise of detecting adversarial activity against the ground data systems managing critical space assets. A fundamental challenge facing machine learning research in cybersecurity is the lack of high-fidelity, shareable datasets for robust evaluation and testing of machine learning-based solutions. High-fidelity, real-world datasets are necessary for reliable benchmarking of nominal system behavior and malicious activity. Unfortunately, such realistic datasets of both nominal and adversarial activity are rarely shared publicly by data owners due to security and privacy concerns. Besides, the available adversarial data is sparse, which makes training models on malicious activity much harder. This situation has impeded and continues to impede the research and successful adoption of machine learning methods for cyber defense. Researchers have dealt with this problem by generating data within a low-fidelity lab environment, using classified and thus unshareable datasets, or downloading low-fidelity public datasets made available by others. We propose an innovative solution to the problem by employing machine learning methods to generate high-fidelity data. Specifically, we propose the use of Generative Adversarial Networks (GANs) to generate high-fidelity data for cybersecurity purposes. GANs have found successful image processing and natural language applications, but have not yet been investigated for cyber data generation. Our proposed approach first involves training the `discriminator' network of the GAN with a sample of real-world data consisting of malicious and nominal samples. We then use the `generator' network to generate new high-fidelity data samples consisting of an appropriate mix of malicious and nominal activity. We demonstrate applications of our architecture by generating high-fidelity cybersecurity data containing both malicious and nominal samples. We thoroughly evaluate the fidelity of our generated data using heuristics and evaluate its usefulness for machine learning applications using three different datasets. Overall, our approach results in high-fidelity, shareable datasets.

Zhang, Yuening↗

PDF Entity Annotation Tool (PEAT)

While different text mining approaches – including the use of Artificial Intelligence (AI) and other machine based methods - continue to expand at a rapid pace, the tools used by researchers to create the labeled datasets required for training, modeling, and evaluation remain rudimentary. Labeled datasets contain the target attributes the machine is going to learn; for example, training an algorithm to delineate between images of a car or truck would generally require a set of images with a quantitative description of the underlying features of each vehicle type. Development of labeled textual data that can be used to build natural language machine learning models for scientific literature is not currently integrated into existing manual workflows used by domain experts. Published literature is rich with important information, such as different types of embedded text, plots, and tables that can all be used as inputs to train ML/natural language processing (NLP) models, when extracted and prepared in machine readable formats. Currently, both normalized data extraction of use to domain experts and extraction to support development of ML/NLP models are labor intensive and cumbersome manual processes. Automatic extraction of data and information from formats such as PDFs that are optimized for layout and human readability, not machine readability. The PDF (Portable Document Format) Entity Annotation Tool (PEAT) was developed with the goal of allowing users to annotate publications within their current print format, while also allowing those annotations to be captured in a machine-readable format. One of the main issues with traditional annotation tools is that they require transforming the PDF into plain text to facilitate the annotation process. While doing so lessens the technical challenges of annotating data, the user loses all structure and provenance that was inherent in the underlying PDF. Also, textual data extraction from PDFs can be an error prone process. Challenges include identifying sequential blocks of text and a multitude of document formats (multiple columns, font encodings, etc.). As a result of these challenges, using existing tools for development of NLP/ML models directly from PDFs is difficult because the generated outputs are not interoperable. We created a system that allows annotations to be completed on the original PDF document structure, with no plain text extraction. The result is an application that allows for easier and more accurate annotations. In addition, by including a feature that grants the user the ability to easily create a schema, we have developed a system that can be used to annotate text for different domain-centric schemas of relevance to subject matter experts. Different knowledge domains require distinct schemas and annotation tags to support machine learning.

97 MATHEMATICS AND COMPUTING↗

Assessment of Current Jet Noise Prediction Capabilities

An assessment was made of the capability of jet noise prediction codes over a broad range of jet flows, with the objective of quantifying current capabilities and identifying areas requiring future research investment. Three separate codes in NASA s possession, representative of two classes of jet noise prediction codes, were evaluated, one empirical and two statistical. The empirical code is the Stone Jet Noise Module (ST2JET) contained within the ANOPP aircraft noise prediction code. It is well documented, and represents the state of the art in semi-empirical acoustic prediction codes where virtual sources are attributed to various aspects of noise generation in each jet. These sources, in combination, predict the spectral directivity of a jet plume. A total of 258 jet noise cases were examined on the ST2JET code, each run requiring only fractions of a second to complete. Two statistical jet noise prediction codes were also evaluated, JeNo v1, and Jet3D. Fewer cases were run for the statistical prediction methods because they require substantially more resources, typically a Reynolds-Averaged Navier-Stokes solution of the jet, volume integration of the source statistical models over the entire plume, and a numerical solution of the governing propagation equation within the jet. In the evaluation process, substantial justification of experimental datasets used in the evaluations was made. In the end, none of the current codes can predict jet noise within experimental uncertainty. The empirical code came within 2dB on a 1/3 octave spectral basis for a wide range of flows. The statistical code Jet3D was within experimental uncertainty at broadside angles for hot supersonic jets, but errors in peak frequency and amplitude put it out of experimental uncertainty at cooler, lower speed conditions. Jet3D did not predict changes in directivity in the downstream angles. The statistical code JeNo,v1 was within experimental uncertainty predicting noise from cold subsonic jets at all angles, but did not predict changes with heating of the jet and did not account for directivity changes at supersonic conditions. Shortcomings addressed here give direction for future work relevant to the statistical-based prediction methods. A full report will be released as a chapter in a NASA publication assessing the state of the art in aircraft noise prediction.

Hunter, Craid A.↗

Report on the AAPM grand challenge on deep generative modeling for learning medical image statistics

Abstract Background The findings of the 2023 AAPM Grand Challenge on Deep Generative Modeling for Learning Medical Image Statistics are reported in this Special Report. Purpose The goal of this challenge was to promote the development of deep generative models for medical imaging and to emphasize the need for their domain‐relevant assessments via the analysis of relevant image statistics. Methods As part of this Grand Challenge, a common training dataset and an evaluation procedure was developed for benchmarking deep generative models for medical image synthesis. To create the training dataset, an established 3D virtual breast phantom was adapted. The resulting dataset comprised about 108 000 images of size 512 512. For the evaluation of submissions to the Challenge, an ensemble of 10 000 DGM‐generated images from each submission was employed. The evaluation procedure consisted of two stages. In the first stage, a preliminary check for memorization and image quality (via the Fréchet Inception Distance [FID]) was performed. Submissions that passed the first stage were then evaluated for the reproducibility of image statistics corresponding to several feature families including texture, morphology, image moments, fractal statistics, and skeleton statistics. A summary measure in this feature space was employed to rank the submissions. Additional analyses of submissions was performed to assess DGM performance specific to individual feature families, the four classes in the training data, and also to identify various artifacts. Results Fifty‐eight submissions from 12 unique users were received for this Challenge. Out of these 12 submissions, 9 submissions passed the first stage of evaluation and were eligible for ranking. The top‐ranked submission employed a conditional latent diffusion model, whereas the joint runners‐up employed a generative adversarial network, followed by another network for image superresolution. In general, we observed that the overall ranking of the top 9 submissions according to our evaluation method (i) did not match the FID‐based ranking, and (ii) differed with respect to individual feature families. Another important finding from our additional analyses was that different DGMs demonstrated similar kinds of artifacts. Conclusions This Grand Challenge highlighted the need for domain‐specific evaluation to further DGM design as well as deployment. It also demonstrated that the specification of a DGM may differ depending on its intended use.

Radiology, Nuclear Medicine & Medical Imaging↗

MAC Europe 1991 campaign: AIRSAR/AVIRIS data integration for agricultural test site classification

During summer 1991, multi-sensor data were acquired over the Italian test site 'Otrepo Pavese', an agricultural flat area in Northern Italy. This area has been the Telespazio pilot test site for experimental activities related to agriculture applications. The aim of the investigation described in the following paper is to assess the amount of information contained in the AIRSAR (Airborne Synthetic Aperture Radar) and AVIRIS (Airborne Visible/Infrared Imaging Spectrometer) data, and to evaluate classification results obtained from each sensor data separately and from the combined dataset. All classifications are examined by means of the resulting confusion matrices and Khat coefficients. Improvements of the classification results obtained by using the integrated dataset are finally evaluated.

Sangiovanni, S.↗

Creating a Repository of Publication Citations for a Data Center

Tracking dataset citations in scientific publications provide multiple benefits: obtaining citation indices for quantitative evaluation of the dataset scientific impact, learning about dataset usage in applied sciences, credits to dataset creators, datasets co-citation relationships and many more.

Infometrics↗

Improved Estimates of Pentad Precipitation through the Merging of Independent Precipitation Datasets

Three independent, quasi-global, gridded datasets of precipitation (a rain gauge-based dataset, the satellite-only component of the NASA Integrated Multi-satellitE Retrievals for Global Precipitation Measurement mission [IMERG] Final Run precipitation product, and precipitation estimates derived from NASA Soil Moisture Active Passive [SMAP] soil moisture retrievals), are objectively combined into a single pentad precipitation dataset at 36-km resolution using a unique approach based on extended triple collocation. The quality of each of the four datasets is then evaluated against independent observations. When a global land surface model at 36-km resolution is integrated four times, once utilizing the merged precipitation forcing and once with each of the three contributing datasets, the near-surface soil moisture variations produced with the merged forcing validate best against independent satellite-based soil moisture fields. In addition, the merged dataset is found to be more consistent, relative to each contributor, with estimates of air temperature variations across the globe. The merged dataset thus appears to draw successfully on the complementary strengths of each contributor: the particularly high quality of the rain gauge-based dataset in areas of high gauge density, the more uniform accuracy across the globe of the IMERG data, and the moderate accuracy, particularly in semi-arid regions, of the soil moisture retrieval-based data. Plain Language Summary Obtaining measurements of precipitation across the globe can be challenging. Rain gauges in some ways provide the most accurate measurements, but gauges are absent in many parts of the world, and even where they exist, they only measure precipitation at the gauge itself and therefore may not provide an accurate large-scale average. Satellite-based estimates of precipitation largely overcome these problems, but such data have their own issues, notably a “snapshot” (rather than a time-average) character of the measurements and difficulty associated with interpreting the measured radiances in the presence of complex land surfaces. In the present paper, we use a novel approach to generate a “merged” dataset, one that optimally combines the gauge precipitation information and the satellite-based precipitation information with a third set of estimates derived from soil moisture retrievals. The merged precipitation dataset and each of the three contributors (aggregated here to 5-day averages at a spatial resolution of about 36-km) are then evaluated for consistency with independent geophysical fields. The merged dataset is found to perform best, a clear indication that it takes proper advantage of the complementary strengths of each contributor and, accordingly, that the presented approach for merging the different contributors is indeed viable.

Precipitation↗

PV Module Operating Temperature - Data and Resources

The Photovoltaic Systems Evaluation Laboratory (PSEL) at Sandia National Laboratories (SNL) in Albuquerque, NM has an extensive test site where PV modules and other system components are deployed and monitored for testing and evaluation. For this dataset PV Performance Labs has assembled one year of measurements from the Systems Long-Term Evaluation (SLTE) project (formerly known as PV Lifetime) providing the main variables needed to investigate and validate PV module operating temperature models: irradiance, ambient temperature, wind speed and back-of-module temperature. For use with more advanced thermal modeling, an estimate of down-welling long-wave radiation is also included.

14 SOLAR ENERGY↗

AI-Ready Data Pilot Project Report

The proliferation of artificial intelligence in scientific research has created an urgent need to define "AI-ready data" for researchers and, more importantly, provide resources to help them produce AI-ready data. At Pacific Northwest National Laboratory, we conducted a pilot study with three data scientists evaluating three CSV datasets from different scientific domains, followed by semi-structured interviews capturing assessment practices. Our findings reveal that AI-readiness evaluation is intuition-based, with practitioners asking "How fast can I go from raw data to my machine learning pipeline?" Data scientists consistently prioritized workflow efficiency, human interpretability, and quality stewardship signals. From these insights, we developed a practical evaluation framework comprising data requirements, metadata standards, and validation tests that provides actionable criteria for producing and curating AI-ready datasets, addressing the gap between theoretical understanding and practical implementation.

97 MATHEMATICS AND COMPUTING↗

Using pile-up collisions as an abundant source of low-energy hadronic physics processes in ATLAS and an extraction of the jet energy resolution

During the 2015–2018 data-taking period, the Large Hadron Collider delivered proton-proton bunch crossings at a centre-of-mass energy of 13 TeV to the ATLAS experiment at a rate of roughly 30 MHz, where each bunch crossing contained an average of 34 independent inelastic proton-proton collisions. The ATLAS trigger system selected roughly 1 kHz of these bunch crossings to be recorded to disk. Offline algorithms then identify one of the recorded collisions as the collision of interest for subsequent data analysis, and the remaining collisions are referred to as pile-up. Pile-up collisions represent a trigger-unbiased dataset, which is evaluated to have an integrated luminosity of 1.33 pb -1 in 2015–2018. This is small compared with the normal trigger-based ATLAS dataset, but when combined with vertex-by-vertex jet reconstruction it provides up to 50 times more dijet events than the conventional single-jet-trigger-based approach, and does so without adding any additional cost or requirements on the trigger system, readout, or storage. The pile-up dataset is validated through comparisons with a special trigger-unbiased dataset recorded by ATLAS, and its utility is demonstrated by means of a measurement of the jet energy resolution in dijet events, where the statistical uncertainty is significantly reduced for jet transverse momenta below 65 GeV.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Variation in forest root image annotation by experts, novices, and AI

Abstract Background The manual study of root dynamics using images requires huge investments of time and resources and is prone to previously poorly quantified annotator bias. Artificial intelligence (AI) image-processing tools have been successful in overcoming limitations of manual annotation in homogeneous soils, but their efficiency and accuracy is yet to be widely tested on less homogenous, non-agricultural soil profiles, e.g., that of forests, from which data on root dynamics are key to understanding the carbon cycle. Here, we quantify variance in root length measured by human annotators with varying experience levels. We evaluate the application of a convolutional neural network (CNN) model, trained on a software accessible to researchers without a machine learning background, on a heterogeneous minirhizotron image dataset taken in a multispecies, mature, deciduous temperate forest. Results Less experienced annotators consistently identified more root length than experienced annotators. Root length annotation also varied between experienced annotators. The CNN root length results were neither precise nor accurate, taking ~ 10% of the time but significantly overestimating root length compared to expert manual annotation ( p = 0.01). The CNN net root length change results were closer to manual ( p = 0.08) but there remained substantial variation. Conclusions Manual root length annotation is contingent on the individual annotator. The only accessible CNN model cannot yet produce root data of sufficient accuracy and precision for ecological applications when applied to a complex, heterogeneous forest image dataset. A continuing evaluation and development of accessible CNNs for natural ecosystems is required.

Handy, Grace↗

Investigating Scientific Data Change with User Research Methods

Scientific datasets are continually expanding and changing due to fluctuations with instruments, quality assessment and quality control processes, and modifications to software pipelines. Datasets include minimal information about these changes or their effects requiring scientists manually assess modifications through a number of labor intensive and ad-hoc steps. The Deduce project is investigating data change to develop metrics, methods, and tools that will help scientists systematically identify and make decisions around data changes. Currently, there is a lack of understanding, and common practices, for identifying and evaluating changes in datasets since systematically measuring and managing data change is under explored in scientific work. We are conducting user research to address this need by exploring scientist's conceptualizations, behaviors, needs, and motivations when dealing with changing datasets. Our user research utilizes multiple methods to produce foundational, generative insights and evaluate research products produced by our team. In this paper, we detail our user research process and outline our findings about data change that emerge from our studies. Our work illustrates how scientific software teams can push beyond just usability testing user interfaces or tools to better probe the underlying ideas they are developing solutions to address.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗