Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Application of a Dataset-Publication Knowledge Graph for Improving Earth Science Data Search

Finding a dataset at a NASA data center that is the best fit for the researcher’s application presents a challenge, not only for a novice user but for an experienced one, due to the data complexity and a multitude of choices of the existing data. Users often search for the data based on the application they are interested in, their research domain, phenomena, research topic, etc. As existing dataset metadata may not cover these search terms, the user may not obtain the most relevant results for their purpose. This problem was addressed by leveraging the content of the titles and abstracts of the research papers that utilize NASA datasets. For this, features from the paper titles and abstracts were extracted, and then a knowledge graph (KG) was used to link these features to the datasets used in that paper. The search for the datasets was tested by querying this knowledge graph through various terms extracted from Earth Science ontologies such as Semantic Web for Earth and Environment Technology (SWEET), and it was shown that this KG search outperforms the existing search that exclusively queries the dataset metadata.

Kristina Stoyanova↗

Global Total Ozone Recovery Trends Attributed to Ozone-Depleting Substance (ODS) Changes Derived From Five Merged Ozone Datasets

We report on updated trends using different merged zonal mean total ozone datasets from satellite and ground-based observations for the period from 1979 to 2020. This work is an update of the trends reported in Weber et al. (2018) using the same datasets up to 2016. Merged datasets used in this study include NASA MOD v8.7 and NOAA Cohesive Data (COH) v8.6, both based on data from the series of Solar Backscatter Ultraviolet (SBUV), SBUV-2, and Ozone Mapping and Profiler Suite (OMPS) satellite instruments (1978–present), as well as the Global Ozone Monitoring Experiment (GOME)-type Total Ozone – Essential Climate Variable (GTO-ECV) and GOME-SCIAMACHY-GOME-2 (GSG) merged datasets (both 1995–present), mainly comprising satellite data from GOME, SCIAMACHY, OMI, GOME-2A, GOME-2B, and TROPOMI. The fifth dataset consists of the annual mean zonal mean data from ground-based measurements collected at the World Ozone and Ultraviolet Radiation Data Centre (WOUDC). Trends were determined by applying a multiple linear regression (MLR) to annual mean zonal mean data. The addition of 4 more years consolidated the fact that total ozone is indeed slowly recovering in both hemispheres as a result of phasing out ozone-depleting substances (ODSs) as mandated by the Montreal Protocol. The near-global (60° S–60° N) ODS-related ozone trend of the median of all datasets after 1995 was 0.4 ± 0.2 (2σ) %/decade, which is roughly a third of the decreasing rate of 1.5 ± 0.6 %/decade from 1978 until 1995. The ratio of decline and increase is nearly identical to that of the EESC (equivalent effective stratospheric chlorine or stratospheric halogen) change rates before and after 1995, confirming the success of the Montreal Protocol. The observed total ozone time series are also in very good agreement with the median of 17 chemistry climate models from CCMI-1 (Chemistry-Climate Model Initiative Phase 1) with current ODS and GHG (greenhouse gas) scenarios (REF-C2 scenario). The positive ODS-related trends in the Northern Hemisphere (NH) after 1995 are only obtained with a sufficient number of terms in the MLR accounting properly for dynamical ozone changes (Brewer–Dobson circulation, Arctic Oscillation (AO), and Antarctic Oscillation (AAO)). A standard MLR (limited to solar, Quasi-Biennial Oscillation (QBO), volcanic, and El Niño–Southern Oscillation (ENSO)) leads to zero trends, showing that the small positive ODS-related trends have been balanced by negative trend contributions from atmospheric dynamics, resulting in nearly constant total ozone levels since 2000.

Total Column Ozone Trends↗

Multi-Decadal Nitrogen Dioxide and Derived Products from Satellites (MINDS) Datasets Released by NASA GES DISC and Their Applications for Air Quality

Nitrogen dioxide (NO2), a pervasive air pollutant, comes from vehicles, power plants, industrial emissions, and off-road sources such as construction or lawn and gardening equipment. The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) curates many remote sensing datasets with NO2 retrievals, which have been utilized for air quality research and applications. The remotely-sensed datasets include those generated by the Ozone Monitoring Instrument (OMI) on the Aura satellite, the TROPOspheric Monitoring Instrument (TROPOMI) onboard the Copernicus Sentinel-5 Precursor (S5P), and the Ozone Mapping and Profiling Suite (OMPS) Nadir-Mapper (NM) instrument on the Suomi National Polar-orbiting Partnership (S- NPP). In collaboration with the NASA Making Earth System Data Records for Use in Research Environments (MEaSUREs) Multi-Decadal Nitrogen Dioxide and Derived Products from Satellites (MINDS) project, the GES DISC recently released MINDS datasets. The NASA MEaSUREs MINDS project aims to develop long-term NO2 global data records by adapting a consistent retrieval algorithm to multiple instrument measurements. Long-term data records will be achieved by applying consistent retrieval approaches to multiple satellite instruments, including OMI (2004 - ); the Global Ozone Monitoring Experiment (GOME, 1995-2011) onboard the second European Remote Sensing satellite (ERS-2); the Scanning Imaging Spectrometer for Atmospheric Cartography (SCIAMACHY, 2002-2012) onboard the ENVIronmental SATellite (ENVISAT); GOME-2 on the Meteorological Operational satellites (MetOp-A and MetOp-B, 2006 - ); and TROPOMI onboard the Copernicus S5P (2017 - ). The long-term record (1995 to present) of MINDS datasets makes them very useful for air quality trend studies. Some MINDS datasets with high spatial resolution of only a few kilometers can be used for air quality research and applications at regional scales. In this presentation, we will introduce all of the MINDS products and services, and demonstrate use cases of MINDS data for studying air quality. We will also present a few other NO2 datasets acquired from NASA’s Health and Air Quality Applied Sciences Team (HAQAST), to be archived and distributed by the GES DISC, and highlight some of their applications for air quality and health.

Feng Ding↗

Enhancing Dataset Discovery With Knowledge Graph Link Prediction Techniques

● In the evolving landscape of open science, the ability to navigate and discover pertinent datasets is increasingly significant. This primarily hinges on the presence of detailed metadata, delineating the dataset’s content, and potential spheres of application. ● The GES DISC datasets are characterized by science keywords to enable dataset discovery in web search interfaces. ● A problem may arise where a dataset lacks a science keyword that it otherwise should have. ● Machine learning techniques such as link prediction can be used to detect these missing science keywords by estimating the probability of new links forming between dataset and keyword nodes.

machine learning↗

Expanding Accurate Person Recognition to New Altitudes and Ranges: The BRIAR Dataset

Face recognition technology has advanced significantly in recent years due largely to the availability of large and increasingly complex training datasets for use in deep learning models. These datasets, however, typically comprise images scraped from news sites or social media plat-forms and, therefore, have limited utility in more advanced security, forensics, and military applications. These applications require lower resolution, longer ranges, and ele-vated viewpoints. To meet these critical needs, we collected and curated the first and second subsets of a large multi-modal biometric dataset designed for use in the research and development (R&D) of biometric recognition technolo-gies under extremely challenging conditions. Thus far, the dataset includes more than 350,000 still images and over 1,300 hours of video footage of approximately 1,000 sub-jects. To collect this data, we used Nikon DSLR cameras, a variety of commercial surveillance cameras, specialized long-rage R&D cameras, and Group 1 and Group 2 UAV platforms. The goal is to support the development of algorithms capable of accurately recognizing people at ranges up to 1,000 m and from high angles of elevation. These ad-vances will include improvements to the state of the art in face recognition and will support new research in the area of whole-body recognition using methods based on gait and anthropometry. This paper describes methods used to col-lect and curate the dataset, and the dataset's characteristics at the current stage.

Brogan, Joel↗

ProvSec: Open Cybersecurity System Provenance Analysis Benchmark Dataset with Labels

System provenance forensic analysis has been studied by a large body of research work. This area needs fine granularity data such as system calls along with event fields to track the dependencies of events. While prior work on security datasets has been proposed, we found a useful dataset of realistic attacks and details that are needed for high-quality provenance tracking is lacking. We created a new dataset of eleven vulnerable cases for system forensic analysis. It includes the full details of system calls including syscall parameters. Realistic attack scenarios with real software vulnerabilities and exploits are used. For each case, we created two sets of benign and adversary scenarios which are manually labeled for supervised machine-learning analysis. In addition, we present an algorithm to improve the data quality in the system provenance forensic analysis. We demonstrate the details of the dataset events and dependency analysis of our dataset cases.

97 MATHEMATICS AND COMPUTING↗

Resampling and data augmentation for short-term PV output prediction based on an imbalanced sky images dataset using convolutional neural networks

Integrating photovoltaics (PV) into electricity grids is challenged by potentially large fluctuations in power generation. In recent years, sky image-based PV output prediction using convolutional neural networks (CNNs) has emerged as a promising approach to forecasting fluctuations. A key challenge is imbalanced sky image datasets: because of the geography of solar PV system installations, sky image datasets are often rich in sunny condition data but deficient in cloudy condition data. This imbalance contrasts with the fact that model errors are dominated by cloudy condition performance. In this study, we attempt to remedy this by exploring the enrichment and augmentation of an imbalanced sky images dataset for two PV output prediction tasks: nowcasting (predicting concurrent PV output) and forecasting (predicting 15-minute-ahead future PV output). We empirically examine the efficacy of using different resampling and data augmentation approaches to create a rebalanced dataset for model development. A three-stage greedy search is used to determine the optimal resampling approach, data augmentation techniques and over-sampling rate. The results show that for the nowcast problem, resampling and data augmentation can effectively enhance the model performance, reducing overall root mean squared error (RMSE) by an average of 4%, or a 15 std. (standard deviation) of improvement compared to the variability of the baseline model. In contrast, the treatment RMSE for the forecast problem nearly always overlaps the baseline performance at the ± 2 std. level. The optimal resampling approach expands on the original dataset by over-sampling the minority cloudy data, with the best results from large over-sampling rate (e.g., 4 ~ 6 times over-sampling of cloudy images).

14 SOLAR ENERGY↗

Uncertainty Assessment of the NASA Earth Exchange Global Daily Downscaled Climate Projections (NEX-GDDP) Dataset

The NASA Earth Exchange Global Daily Downscaled Projections (NEX-GDDP) dataset is comprised of downscaled climate projections that are derived from 21 General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 5 (CMIP5) and across two of the four greenhouse gas emissions scenarios (RCP4.5 and RCP8.5). Each of the climate projections includes daily maximum temperature, minimum temperature, and precipitation for the periods from 1950 through 2100 and the spatial resolution is 0.25 degrees (approximately 25 km x 25 km). The GDDP dataset has received warm welcome from the science community in conducting studies of climate change impacts at local to regional scales, but a comprehensive evaluation of its uncertainties is still missing. In this study, we apply the Perfect Model Experiment framework (Dixon et al. 2016) to quantify the key sources of uncertainties from the observational baseline dataset, the downscaling algorithm, and some intrinsic assumptions (e.g., the stationary assumption) inherent to the statistical downscaling techniques. We developed a set of metrics to evaluate downscaling errors resulted from bias-correction ("quantile-mapping"), spatial disaggregation, as well as the temporal-spatial non-stationarity of climate variability. Our results highlight the spatial disaggregation (or interpolation) errors, which dominate the overall uncertainties of the GDDP dataset, especially over heterogeneous and complex terrains (e.g., mountains and coastal area). In comparison, the temporal errors in the GDDP dataset tend to be more constrained. Our results also indicate that the downscaled daily precipitation also has relatively larger uncertainties than the temperature fields, reflecting the rather stochastic nature of precipitation in space. Therefore, our results provide insights in improving statistical downscaling algorithms and products in the future.

climate projection↗

Uncertainty Assessment of the NASA Earth Exchange Global Daily Downscaled Climate Projections (NEX-GDDP) Dataset

The NASA Earth Exchange Global Daily Downscaled Projections (NEX-GDDP) dataset is comprised of downscaled climate projections that are derived from 21 General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 5 (CMIP5) and across two of the four greenhouse gas emissions scenarios (RCP4.5 and RCP8.5). Each of the climate projections includes daily maximum temperature, minimum temperature, and precipitation for the periods from 1950 through 2100 and the spatial resolution is 0.25 degrees (approximately 25 km by 25 km). The GDDP dataset has received warm welcome from the science community in conducting studies of climate change impacts at local to regional scales, but a comprehensive evaluation of its uncertainties is still missing. In this study, we apply the Perfect Model Experiment framework (Dixon et al. 2016) to quantify the key sources of uncertainties from the observational baseline dataset, the downscaling algorithm, and some intrinsic assumptions (e.g., the stationary assumption) inherent to the statistical downscaling techniques. We developed a set of metrics to evaluate downscaling errors resulted from bias-correction ("quantile-mapping"), spatial disaggregation, as well as the temporal-spatial non-stationarity of climate variability. Our results highlight the spatial disaggregation (or interpolation) errors, which dominate the overall uncertainties of the GDDP dataset, especially over heterogeneous and complex terrains (e.g., mountains and coastal area). In comparison, the temporal errors in the GDDP dataset tend to be more constrained. Our results also indicate that the downscaled daily precipitation also has relatively larger uncertainties than the temperature fields, reflecting the rather stochastic nature of precipitation in space. Therefore, our results provide insights in improving statistical downscaling algorithms and products in the future.

general circulation model (GCM)↗

PNNL DataHub Project: Omics Lethal Human Viruses Project Profiling of the Host Response to Ebola Virus Infection, Processed Experimental Dataset Catalog

Ebola virus (EBOV) is high risk biological agent, classified as a Category A priority pathogen (Flaviviridae) by the National Institute of Allergy and Infectious Diseases (NIAID), known to cause hemorrhagic fever with high mortality rates in humans. Lethal host-pathogen invasion mechanisms and the cellular intricacies behind these fatal infections still remain unclear. The NIAID Modeling Host Responses to Understand Severe Human Virus Infections Research Program project (2013-2018) aimed to develop an improved comprehensive understanding of the host response to a suite of viruses causing lethal infections leveraging a systems biology approach. Herein, PNNL sub-projects provide a never before released comprehensive infectious disease collection of primary and secondary transformation multi-Omics data profiling a series of priority pathogen primary experimental studies for enhanced open-access to viral Omics datasets and project lifecycle metadata. Secondary host-pathogen viral dataset downloads contain one or more statistically processed (normalization data transformation) quantitative dataset collections resulting in qualitative expression analyses of primary host-pathogen experimental study designs. Leveraging unique high-resolution Omics capabilities for proteomics (P), metabolomics (M), lipidomics (L), and transcriptomics (T) dataset downloads each have a direct relationship to a primary sample submission corresponding to a specific Ebola virus [NCBITAXON:186536] (Zaire/Makona or Zaire/Mayinga) experimental infection study. Human host samples types include peripheral blood mononuclear cells isolated from blood plasma ["PBMC", BTO:0001025], human hepatoma carcinoma cells ["HUH", BTO:0001950], human umbilical vein endothelial cells ["HUVEC", BTO:0001949], immortalized human hepatocyte cells ["IHH", BTO:0006147], and human histiocytic lymphoma cells ["U937", BTO:0001412].

59 BASIC BIOLOGICAL SCIENCES↗

PNNL DataHub Project: Omics Lethal Human Viruses Project Profiling of the Host Response to Influenza A Virus Infection, Processed Experimental Dataset Catalog

Influenza A virus (IAV) is a high risk biological agent, classified as a Category C priority pathogen (Orthomyxoviridae) by the National Institute of Allergy and Infectious Diseases (NIAID), and is known to cause severe respiratory disease with high mortality rates in humans. Lethal host-pathogen invasion mechanisms and the cellular intricacies behind these fatal infections still remain unclear. The NIAID Modeling Host Responses to Understand Severe Human Virus Infections Research Program project (2013-2018) aimed to develop an improved comprehensive understanding of the host response to a suite of viruses causing lethal infections leveraging a systems biology approach. Herein, PNNL sub-projects provide a never before released comprehensive infectious disease collection of primary and secondary transformation multi-Omics data profiling a series of priority pathogen primary experimental studies for enhanced open-access to viral Omics datasets and project lifecycle metadata. Secondary host-pathogen viral dataset downloads contain one or more statistically processed (normalization data transformation) quantitative dataset collections resulting in qualitative expression analyses of primary host-pathogen experimental study designs. Leveraging unique high-resolution Omics capabilities for proteomics (P), metabolomics (M), lipidomics (L), and transcriptomics (T) dataset downloads each have a direct relationship to a primary sample submission corresponding to a specific Influenza A virus [NCBITAXON:11320] experimental infection study. Host sample types include human lung adenocarcinoma cells ["Calu-3", BTO:0002750] and whole mouse lung [BTO:0000763] tissue collections.

59 BASIC BIOLOGICAL SCIENCES↗

PNNL DataHub Project: Omics Lethal Human Viruses Project Profiling of the Host Interferon-Stimulated Response to Virus Infection, Processed Experimental Dataset Catalog

Human Interferon (IFN) alpha, beta, and gamma participate in the body's natural immune response to lethal virus infection and disease.The NIAID Modeling Host Responses to Understand Severe Human Virus Infections Research Program project (2013 - 2018) aimed to develop an improved comprehensive understanding of the host response to a suite of viruses causing lethal infections leveraging a systems biology approach. The NIAID Modeling Host Responses to Understand Severe Human Virus Infections Research Program project (2013-2018) aimed to develop an improved comprehensive understanding of the host response to a suite of viruses causing lethal infections leveraging a systems biology approach. Herein, PNNL sub-projects provide a never before released comprehensive infectious disease collection of primary and secondary transformation multi-Omics data profiling a series of priority pathogen primary experimental studies for enhanced open-access to viral Omics datasets and project lifecycle metadata. Secondary host-associated viral dataset downloads contain one or more statistically processed (normalization data transformation) quantitative dataset collections resulting in qualitative expression analyses of primary host-pathogen experimental study designs. Transcriptomics (T) dataset downloads each have a direct relationship to a primary sample submission corresponding to a specific Human Interferon (IFN), interferon alpha (IFNα), interferon beta (IFNβ), and/or interferon gamma (IFNγ) stimulated response to an experimental virus infection treatment study. Host sample types include cerebellum ["CB", BTO:0000232], cortical neurons ["CN", BTO:0004102], cortex ["CT", BTO:0000233], dendritic cells ["DC", BTO:0002042], granule cell neurons ["GCN", BTO:0003393], lymph node ["LN", BTO:0000784], and serum ["SE", BTO:0001239] from mouse (Mus musculus) tissue collections.

59 BASIC BIOLOGICAL SCIENCES↗

Omics Lethal Human Viruses Project Profiling of the Host Response to MERS-CoV Infection, Processed Experimental Dataset Catalog

Middle East Respiratory Syndrome coronavirus (MERS-CoV) is classified as a Category C priority pathogen (Coronaviridae) by the National Institute of Allergy and Infectious Diseases (NIAID), and is known to cause severe respiratory disease with high mortality rates in humans. Lethal host-pathogen invasion mechanisms and the cellular intricacies behind these fatal infections still remain unclear. The NIAID Modeling Host Responses to Understand Severe Human Virus Infections Research Program project (2013-2018) aimed to develop an improved comprehensive understanding of the host response to a suite of viruses causing lethal infections leveraging a systems biology approach. Herein, PNNL sub-projects provide a never before released comprehensive infectious disease collection of primary and secondary transformation multi-Omics data profiling a series of priority pathogen primary experimental studies for enhanced open-access to viral Omics datasets and project lifecycle metadata. Secondary host-pathogen viral dataset downloads contain one or more statistically processed (normalization data transformation) quantitative dataset collections resulting in qualitative expression analyses of primary host-pathogen experimental study designs. Leveraging unique high-resolution Omics capabilities for proteomics (P), metabolomics (M), lipidomics (L), and transcriptomics (T) dataset downloads each have a direct relationship to a primary sample submission corresponding to a specific MERS-CoV [NCBITAXON:1335626] experimental infection study. Host sample types include human lung adenocarcinoma cells ["Calu-3", BTO:0002750], human bronchial epithelial cells ["Calu-3 clone 2B4"; BTO:0002022], primary human fibroblasts ["FB"; BTO:0000452], primary human airway epithelial cells ["HAE"; BTO:0005571], human microvascular endothelial cells ["HMVE"; BTO:0003123], and whole mouse lung [BTO:0000763] tissue collections.

59 BASIC BIOLOGICAL SCIENCES↗

PNNL DataHub Project: Omics Lethal Human Viruses Project Profiling of the Host Response to West Nile Virus Infection, Processed Experimental Dataset Catalog

West Nile virus (WNV) is classified as a Category B priority pathogen (mosquito-borne Flavivirus) by the National Institute of Allergy and Infectious Diseases (NIAID), and are known to cause severe infections in humans where lethal host-associated mechanisms are not clearly defined. The NIAID Modeling Host Responses to Understand Severe Human Virus Infections Research Program project (2013 - 2018) aimed to develop an improved comprehensive understanding of the host response to a suite of viruses causing lethal infections leveraging a systems biology approach. The NIAID Modeling Host Responses to Understand Severe Human Virus Infections Research Program project (2013-2018) aimed to develop an improved comprehensive understanding of the host response to a suite of viruses causing lethal infections leveraging a systems biology approach. Herein, PNNL sub-projects provide a never before released comprehensive infectious disease collection of primary and secondary transformation multi-Omics data profiling a series of priority pathogen primary experimental studies for enhanced open-access to viral Omics datasets and project lifecycle metadata. Secondary host-pathogen viral dataset downloads contain one or more statistically processed (normalization data transformation) quantitative dataset collections resulting in qualitative expression analyses of primary host-pathogen experimental study designs. Leveraging unique high-resolution Omics capabilities for proteomics (P), metabolomics (M), lipidomics (L), and transcriptomics (T) dataset downloads each have a direct relationship to a primary sample submission corresponding a specific West Nile virus [NCBITAXON:11082] (WNV-NY99 382) experimental infection study. Host sample types include cerebellum ["CB", BTO:0000232], cortical neurons ["CN", BTO:0004102], cortex ["CT", BTO:0000233], dendritic cells ["DC", BTO:0002042], granule cell neurons ["GCN", BTO:0003393], lymph node ["LN", BTO:0000784], and serum ["SE", BTO:0001239] from mouse (Mus musculus) tissue collections.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluation of normalization strategies for mass spectrometry-based multi-omics datasets

Introduction Data normalization is crucial for multi-omics integration, reducing systematic errors and maximizing the likelihood of discovering true biological variation. Most studies assess normalization for a single omics type or use datasets from separate experiments. Few address time-course data, where normalization might bias temporal differentiation. In this study, we compared common normalization methods and a machine learning approach, Systematical Error Removal using Random Forest (SERRF), using multi-omics datasets generated from the same experiment—even from the same cell lysate. Objectives To develop a straightforward process to assess normalization effects and identify the most robust methods across multi-omics datasets. Methods We analyzed metabolomics, lipidomics, and proteomics datasets from primary human cardiomyocytes and motor neurons exposed to acetylcholine-active compounds over time. Normalization effectiveness was evaluated based on improvement in QC features consistency and observing the change in treatment and time-related variance. Results Probabilistic Quotient Normalization (PQN) and Locally Estimated Scatterplot Smoothing (LOESS) QC were identified as optimal for metabolomics and lipidomics, while PQN, Median, and LOESS normalization excelled for proteomics. These methods consistently enhanced QC feature consistency in metabolomics and lipidomics, and preserved time-related variance or treatment-related variance in proteomics, demonstrating their effectiveness and robustness. SERRF normalization, applied only to metabolomics in this study, outperformed other methods in some datasets but inadvertently masked treatment-related variance in others. Conclusion Our evaluation identified PQN and LoessQC as the top methods for metabolomics and lipidomics, and PQN, Median, and Loess normalization for proteomics, in multi-omics integration in a temporal study.

60 APPLIED LIFE SCIENCES↗

Assessment of Outliers in Alloy Datasets Using Unsupervised Techniques

We report advancements in data analytics techniques have enabled complex, disparate datasets to be leveraged for alloy design. Identifying outliers in a dataset can reduce noise, identify erroneous and/or anomalous records, prevent overfitting, and improve model assessment and optimization. In this work, two alloy datasets (9-12% Cr ferritic martensitic steels, and austenitic stainless steels) have been assessed for outliers using unsupervised techniques and supplemented with domain knowledge. Principal component analysis and k-means clustering were applied to the data, and points were assessed as outliers based on their distance away from other points in the cluster and from other points in the dataset. The outlier characteristics were investigated to determine both cluster-specific and overall trends in the properties of the outlier points. The approach demonstrated here is extensible to other alloy datasets for outlier identification and evaluation to improve the reliability of machine learning and modeling predictions for advanced alloy design.

36 MATERIALS SCIENCE↗

Development of a Benchmark Eddy Flux Evapotranspiration Dataset for Evaluation of Satellite-Driven Evapotranspiration Models Over the CONUS

A large sample of ground-based evapotranspiration (ET) measurements made in the United States, primarily from eddy covariance systems, were post-processed to produce a benchmark ET dataset. The dataset was produced primarily to support the intercomparison and evaluation of the OpenET satellite-based remote sensing ET (RSET) models and could also be used to evaluate ET data from other models and approaches. OpenET is a web-based service that makes field-delineated and pixel-level ET estimates from well-established RSET models readily available to water managers, agricultural producers, and the public. The benchmark dataset is composed of flux and meteorological data from a variety of providers covering native vegetation and agricultural settings. Flux footprint predictions were developed for each station and included static flux footprints developed based on average wind direction and speed, as well as dynamic hourly footprints that were generated with a physically based model of upwind source area. The two footprint prediction methods were rigorously compared to evaluate their relative spatial coverage. Data from all sources were post-processed in a consistent and reproducible manner including data handling, gap-filling, temporal aggregation, and energy balance closure correction. The resulting dataset included 243,048 daily and 5,284 monthly ET values from 194 stations, with all data falling between 1995 and 2021. We assessed average daily energy imbalance using 172 flux sites with a total of 193,021 days of data, finding that overall turbulent fluxes were understated by about 12% on average relative to available energy. Multiple linear regression analyses indicated that daily average latent energy flux may be typically understated slightly more than sensible heat flux. This dataset was developed to provide a consistent reference to support evaluation of RSET data being developed for a wide range of applications related to water accounting and water resources management at field to watershed scales.

54 ENVIRONMENTAL SCIENCES↗

A dataset of eco-evidence tools to inform early-stage environmental impact assessments of hydropower development

The datasets described herein provide the foundation for a decision support prototype (DSP) toolkit aimed at assisting stakeholders in determining evidence of which aspects of river ecosystems have been impacted by hydropower. The DSP toolkit and its application are presented and described in the article “Evidence-based indicator approach to guide preliminary environmental impact assessments of hydropower development” [1]. Development of the DSP and the output for decision support centralize around 42 river function indicators describing the dimensionality of river ecosystems through six main categories: biota and biodiversity, water quality, hydrology, geomorphology, land cover, and river connectivity. Three main tools are represented in the DSP: A science-based questionnaire (SBQ), an environmental envelope model (EEM), and a river function linkage assessment tool (RFLAT). The SBQ is a structured survey-style questionnaire whose objective is to provide evidence of which indicators have been impacted by hydropower. Based on a global literature review, 140 questions were developed from general hypotheses regarding the impacts of dams on rivers. The EEM is a model to predict the likelihood of hydropower impacting indicators based on a several variables. The intended use of the EEM is for situations of new hydropower development where results of the SBQ are incomplete or highly uncertain. The EEM was developed through the compilation of a dataset containing attributes of dams, reservoirs, and geospatial information on environmental concerns, which was combined with data on ecological indicators documented at those sites through literature review. The model operates through 247 “envelopes” and weighting factors, representing the individual effect of each variable on each indicator, all available through spreadsheets. Finally, the RFLAT is a tool to examine causal relationships amongst indicators. Inter-indicator relationships were hypothesized based on literature review and summarized into node and edge datasets to represent the structure of a graphical network. Bayes theorem was used estimate conditional probabilities of inter-indicator relationships based on the output of the SBQ. Nodes and edges were imported into R programming environment to visualize ecological indicator networks. The datasets can be expanded upon and enriched with more detailed questions for the SBQ, building upon the EEM with to develop more sophisticated models, and identifying new relationships for the RFALT. Additionally, once the tools are applied to numerous hydropower developments, the output of the tools (e.g. evidence of impacted indicators) becomes a very useful dataset for meta-analyses of hydropower impacts.

13 HYDRO ENERGY↗