Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “datasets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Development of a 95-Year Solar Dataset for Resource Adequacy Studies

Long-term high-resolution solar data provides enhanced understanding of variability of solar generation and enhances our ability to develop strategies for a resilient and reliable electric grid under high deployment of solar energy. Therefore, it is important to develop long-term synthetic datasets that can provide multiple occurrences of various severe weather scenarios that are expected to test the limits of resource adequacy under scenarios contain various energy generation sources. Examples of such scenarios could be long periods of high temperatures when demand for electricity is high or periods where high winds could lead to a shut-down of transmission lines for long periods of time to ensure fire safety. NREL has developed the first version of such a dataset covering a 95-year period covering 2006-2100 at a 4km hourly resolution. This dataset contains all variables necessary to calculate solar generation. During development of this dataset, we focused on creating unbiased, high-resolution solar irradiance through statistical downscaling methods, using Regional Climate Model (RCM) simulations from the North American Coordinated Regional Climate Downscaling Experiment (NA-CORDEX) as input. The National Solar Radiation Database (NSRDB) containing over 25 years of observations was used to calibrate the statistical downscaling models. This presentation will outline the primary steps in developing this dataset, including (1) regridding RCM data to a common grid at 20-km resolution, (2) correcting RCM biases with NSRDB, (3) applying temporal and spatial downscaling methods to generate high-resolution (4-km, hourly) solar and ancillary data. Additionally, we will present an evaluation of the downscaled data against the NSRDB across various zones in the CONUS. Lastly, we will present a user guide for accessing the datasets.

14 SOLAR ENERGY↗

Ontology-Enriched Specifications Enabling Findable, Accessible, Interoperable, and Reusable Marine Metagenomic Datasets in Cyberinfrastructure Systems

Marine microbial ecology requires the systematic comparison of biogeochemical and sequence data to analyze environmental influences on the distribution and variability of microbial communities. With ever-increasing quantities of metagenomic data, there is a growing need to make datasets Findable, Accessible, Interoperable, and Reusable (FAIR) across diverse ecosystems. FAIR data is essential to developing analytical frameworks that integrate microbiological, genomic, ecological, oceanographic, and computational methods. Although community standards defining the minimal metadata required to accompany sequence data exist, they haven’t been consistently used across projects, precluding interoperability. Moreover, these data are not machine-actionable or discoverable by cyberinfrastructure systems. By making ‘omic and physicochemical datasets FAIR to machine systems, we can enable sequence data discovery and reuse based on machine-readable descriptions of environments or physicochemical gradients. In this work, we developed a novel technical specification for dataset encapsulation for the FAIR reuse of marine metagenomic and physicochemical datasets within cyberinfrastructure systems. This includes using Frictionless Data Packages enriched with terminology from environmental and life-science ontologies to annotate measured variables, their units, and the measurement devices used. This approach was implemented in Planet Microbe, a cyberinfrastructure platform and marine metagenomic web-portal. Here, we discuss the data properties built into the specification to make global ocean datasets FAIR within the Planet Microbe portal. We additionally discuss the selection of, and contributions to marine-science ontologies used within the specification. Finally, we use the system to discover data by which to answer various biological questions about environments, physicochemical gradients, and microbial communities in meta-analyses. This work represents a future direction in marine metagenomic research by proposing a specification for FAIR dataset encapsulation that, if adopted within cyberinfrastructure systems, would automate the discovery, exchange, and re-use of data needed to answer broader reaching questions than originally intended.

59 BASIC BIOLOGICAL SCIENCES↗

The executive disruption model of tinnitus distress: Model validation in two independent datasets using factor score regression

This study presents the executive disruption model (EDM) of tinnitus distress and subsequently validates it statistically using two independent datasets (the Construction Dataset: n = 96 and the Validation Dataset: n = 200). The conceptual EDM was first operationalised as a structural causal model (construction phase). Then multiple regression was used to examine the effect of executive functioning on tinnitus-related distress (validation phase), adjusting for the additional contributions of hearing threshold and psychological distress. For both datasets, executive functioning negatively predicted tinnitus distress score by a similar amount (the Construction Dataset: β = −3.50, p = 0.13 and the Validation Dataset: β = −3.71, p = 0.02). Theoretical implications and applications of the EDM are subsequently discussed; these include the predictive nature of executive functioning in the development of distressing tinnitus, and the clinical utility of the EDM.

Clarke, Nathan A.↗

Panta Rhei benchmark dataset: socio-hydrological data of paired events of floods and droughts

As the adverse impacts of hydrological extremes increase in many regions of the world, a better understanding of the drivers of changes in risk and impacts is essential for effective flood and drought risk management and climate adaptation. However, there is currently a lack of comprehensive, empirical data about the processes, interactions, and feedbacks in complex human–water systems leading to flood and drought impacts. Here we present a benchmark dataset containing socio-hydrological data of paired events, i.e. two floods or two droughts that occurred in the same area. The 45 paired events occurred in 42 different study areas and cover a wide range of socio-economic and hydro-climatic conditions. The dataset is unique in covering both floods and droughts, in the number of cases assessed and in the quantity of socio-hydrological data. The benchmark dataset comprises (1) detailed review-style reports about the events and key processes between the two events of a pair; (2) the key data table containing variables that assess the indicators which characterize management shortcomings, hazard, exposure, vulnerability, and impacts of all events; and (3) a table of the indicators of change that indicate the differences between the first and second event of a pair. The advantages of the dataset are that it enables comparative analyses across all the paired events based on the indicators of change and allows for detailed context- and location-specific assessments based on the extensive data and reports of the individual study areas. The dataset can be used by the scientific community for exploratory data analyses, e.g. focused on causal links between risk management; changes in hazard, exposure and vulnerability; and flood or drought impacts. The data can also be used for the development, calibration, and validation of socio-hydrological models. The dataset is available to the public through the GFZ Data Services (Kreibich et al., 2023, https://doi.org/10.5880/GFZ.4.4.2023.001).

54 ENVIRONMENTAL SCIENCES↗

ClimateNet: an expert-labeled open dataset and deep learning architecture for enabling high-precision analyses of extreme weather

Abstract. Identifying, detecting, and localizing extreme weather events is a crucial first step in understanding how they may vary under different climate change scenarios. Pattern recognition tasks such as classification, object detection, and segmentation (i.e., pixel-level classification) have remained challenging problems in the weather and climate sciences. While there exist many empirical heuristics for detecting extreme events, the disparities between the output of these different methods even for a single event are large and often difficult to reconcile. Given the success of deep learning (DL) in tackling similar problems in computer vision, we advocate a DL-based approach. DL, however, works best in the context of supervised learning – when labeled datasets are readily available. Reliable labeled training data for extreme weather and climate events is scarce. We create “ClimateNet” – an open, community-sourced human-expert-labeled curated dataset that captures tropical cyclones (TCs) and atmospheric rivers (ARs) in high-resolution climate model output from a simulation of a recent historical period. We use the curated ClimateNet dataset to train a state-of-the-art DL model for pixel-level identification – i.e., segmentation – of TCs and ARs. We then apply the trained DL model to historical and climate change scenarios simulated by the Community Atmospheric Model (CAM5.1) and show that the DL model accurately segments the data into TCs, ARs, or “the background” at a pixel level. Further, we show how the segmentation results can be used to conduct spatially and temporally precise analytics by quantifying distributions of extreme precipitation conditioned on event types (TC or AR) at regional scales. The key contribution of this work is that it paves the way for DL-based automated, high-fidelity, and highly precise analytics of climate data using a curated expert-labeled dataset – ClimateNet. ClimateNet and the DL-based segmentation method provide several unique capabilities: (i) they can be used to calculate a variety of TC and AR statistics at a fine-grained level; (ii) they can be applied to different climate scenarios and different datasets without tuning as they do not rely on threshold conditions; and (iii) the proposed DL method is suitable for rapidly analyzing large amounts of climate model output. While our study has been conducted for two important extreme weather patterns (TCs and ARs) in simulation datasets, we believe that this methodology can be applied to a much broader class of patterns and applied to observational and reanalysis data products via transfer learning.

54 ENVIRONMENTAL SCIENCES↗

Benchmark datasets for SARS-CoV-2 surveillance bioinformatics

Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the cause of coronavirus disease 2019 (COVID-19), has spread globally and is being surveilled with an international genome sequencing effort. Surveillance consists of sample acquisition, library preparation, and whole genome sequencing. This has necessitated a classification scheme detailing Variants of Concern (VOC) and Variants of Interest (VOI), and the rapid expansion of bioinformatics tools for sequence analysis. These bioinformatic tools are means for major actionable results: maintaining quality assurance and checks, defining population structure, performing genomic epidemiology, and inferring lineage to allow reliable and actionable identification and classification. Additionally, the pandemic has required public health laboratories to reach high throughput proficiency in sequencing library preparation and downstream data analysis rapidly. However, both processes can be limited by a lack of a standardized sequence dataset. We identified six SARS-CoV-2 sequence datasets from recent publications, public databases and internal resources. In addition, we created a method to mine public databases to identify representative genomes for these datasets. Using this novel method, we identified several genomes as either VOI/VOC representatives or non-VOI/VOC representatives. To describe each dataset, we utilized a previously published datasets format, which describes accession information and whole dataset information. Additionally, a script from the same publication has been enhanced to download and verify all data from this study.

60 APPLIED LIFE SCIENCES↗

Expanding Accurate Person Recognition to New Altitudes and Ranges: The BRIAR Dataset

Face recognition technology has advanced significantly in recent years due largely to the availability of large and increasingly complex training datasets for use in deep learning models. These datasets, however, typically comprise images scraped from news sites or social media plat-forms and, therefore, have limited utility in more advanced security, forensics, and military applications. These applications require lower resolution, longer ranges, and ele-vated viewpoints. To meet these critical needs, we collected and curated the first and second subsets of a large multi-modal biometric dataset designed for use in the research and development (R&D) of biometric recognition technolo-gies under extremely challenging conditions. Thus far, the dataset includes more than 350,000 still images and over 1,300 hours of video footage of approximately 1,000 sub-jects. To collect this data, we used Nikon DSLR cameras, a variety of commercial surveillance cameras, specialized long-rage R&D cameras, and Group 1 and Group 2 UAV platforms. The goal is to support the development of algorithms capable of accurately recognizing people at ranges up to 1,000 m and from high angles of elevation. These ad-vances will include improvements to the state of the art in face recognition and will support new research in the area of whole-body recognition using methods based on gait and anthropometry. This paper describes methods used to col-lect and curate the dataset, and the dataset's characteristics at the current stage.

Brogan, Joel↗

ProvSec: Open Cybersecurity System Provenance Analysis Benchmark Dataset with Labels

System provenance forensic analysis has been studied by a large body of research work. This area needs fine granularity data such as system calls along with event fields to track the dependencies of events. While prior work on security datasets has been proposed, we found a useful dataset of realistic attacks and details that are needed for high-quality provenance tracking is lacking. We created a new dataset of eleven vulnerable cases for system forensic analysis. It includes the full details of system calls including syscall parameters. Realistic attack scenarios with real software vulnerabilities and exploits are used. For each case, we created two sets of benign and adversary scenarios which are manually labeled for supervised machine-learning analysis. In addition, we present an algorithm to improve the data quality in the system provenance forensic analysis. We demonstrate the details of the dataset events and dependency analysis of our dataset cases.

97 MATHEMATICS AND COMPUTING↗

Resampling and data augmentation for short-term PV output prediction based on an imbalanced sky images dataset using convolutional neural networks

Integrating photovoltaics (PV) into electricity grids is challenged by potentially large fluctuations in power generation. In recent years, sky image-based PV output prediction using convolutional neural networks (CNNs) has emerged as a promising approach to forecasting fluctuations. A key challenge is imbalanced sky image datasets: because of the geography of solar PV system installations, sky image datasets are often rich in sunny condition data but deficient in cloudy condition data. This imbalance contrasts with the fact that model errors are dominated by cloudy condition performance. In this study, we attempt to remedy this by exploring the enrichment and augmentation of an imbalanced sky images dataset for two PV output prediction tasks: nowcasting (predicting concurrent PV output) and forecasting (predicting 15-minute-ahead future PV output). We empirically examine the efficacy of using different resampling and data augmentation approaches to create a rebalanced dataset for model development. A three-stage greedy search is used to determine the optimal resampling approach, data augmentation techniques and over-sampling rate. The results show that for the nowcast problem, resampling and data augmentation can effectively enhance the model performance, reducing overall root mean squared error (RMSE) by an average of 4%, or a 15 std. (standard deviation) of improvement compared to the variability of the baseline model. In contrast, the treatment RMSE for the forecast problem nearly always overlaps the baseline performance at the ± 2 std. level. The optimal resampling approach expands on the original dataset by over-sampling the minority cloudy data, with the best results from large over-sampling rate (e.g., 4 ~ 6 times over-sampling of cloudy images).

14 SOLAR ENERGY↗

PNNL DataHub Project: Omics Lethal Human Viruses Project Profiling of the Host Response to Ebola Virus Infection, Processed Experimental Dataset Catalog

Ebola virus (EBOV) is high risk biological agent, classified as a Category A priority pathogen (Flaviviridae) by the National Institute of Allergy and Infectious Diseases (NIAID), known to cause hemorrhagic fever with high mortality rates in humans. Lethal host-pathogen invasion mechanisms and the cellular intricacies behind these fatal infections still remain unclear. The NIAID Modeling Host Responses to Understand Severe Human Virus Infections Research Program project (2013-2018) aimed to develop an improved comprehensive understanding of the host response to a suite of viruses causing lethal infections leveraging a systems biology approach. Herein, PNNL sub-projects provide a never before released comprehensive infectious disease collection of primary and secondary transformation multi-Omics data profiling a series of priority pathogen primary experimental studies for enhanced open-access to viral Omics datasets and project lifecycle metadata. Secondary host-pathogen viral dataset downloads contain one or more statistically processed (normalization data transformation) quantitative dataset collections resulting in qualitative expression analyses of primary host-pathogen experimental study designs. Leveraging unique high-resolution Omics capabilities for proteomics (P), metabolomics (M), lipidomics (L), and transcriptomics (T) dataset downloads each have a direct relationship to a primary sample submission corresponding to a specific Ebola virus [NCBITAXON:186536] (Zaire/Makona or Zaire/Mayinga) experimental infection study. Human host samples types include peripheral blood mononuclear cells isolated from blood plasma ["PBMC", BTO:0001025], human hepatoma carcinoma cells ["HUH", BTO:0001950], human umbilical vein endothelial cells ["HUVEC", BTO:0001949], immortalized human hepatocyte cells ["IHH", BTO:0006147], and human histiocytic lymphoma cells ["U937", BTO:0001412].

59 BASIC BIOLOGICAL SCIENCES↗

PNNL DataHub Project: Omics Lethal Human Viruses Project Profiling of the Host Response to Influenza A Virus Infection, Processed Experimental Dataset Catalog

Influenza A virus (IAV) is a high risk biological agent, classified as a Category C priority pathogen (Orthomyxoviridae) by the National Institute of Allergy and Infectious Diseases (NIAID), and is known to cause severe respiratory disease with high mortality rates in humans. Lethal host-pathogen invasion mechanisms and the cellular intricacies behind these fatal infections still remain unclear. The NIAID Modeling Host Responses to Understand Severe Human Virus Infections Research Program project (2013-2018) aimed to develop an improved comprehensive understanding of the host response to a suite of viruses causing lethal infections leveraging a systems biology approach. Herein, PNNL sub-projects provide a never before released comprehensive infectious disease collection of primary and secondary transformation multi-Omics data profiling a series of priority pathogen primary experimental studies for enhanced open-access to viral Omics datasets and project lifecycle metadata. Secondary host-pathogen viral dataset downloads contain one or more statistically processed (normalization data transformation) quantitative dataset collections resulting in qualitative expression analyses of primary host-pathogen experimental study designs. Leveraging unique high-resolution Omics capabilities for proteomics (P), metabolomics (M), lipidomics (L), and transcriptomics (T) dataset downloads each have a direct relationship to a primary sample submission corresponding to a specific Influenza A virus [NCBITAXON:11320] experimental infection study. Host sample types include human lung adenocarcinoma cells ["Calu-3", BTO:0002750] and whole mouse lung [BTO:0000763] tissue collections.

59 BASIC BIOLOGICAL SCIENCES↗

PNNL DataHub Project: Omics Lethal Human Viruses Project Profiling of the Host Interferon-Stimulated Response to Virus Infection, Processed Experimental Dataset Catalog

Human Interferon (IFN) alpha, beta, and gamma participate in the body's natural immune response to lethal virus infection and disease.The NIAID Modeling Host Responses to Understand Severe Human Virus Infections Research Program project (2013 - 2018) aimed to develop an improved comprehensive understanding of the host response to a suite of viruses causing lethal infections leveraging a systems biology approach. The NIAID Modeling Host Responses to Understand Severe Human Virus Infections Research Program project (2013-2018) aimed to develop an improved comprehensive understanding of the host response to a suite of viruses causing lethal infections leveraging a systems biology approach. Herein, PNNL sub-projects provide a never before released comprehensive infectious disease collection of primary and secondary transformation multi-Omics data profiling a series of priority pathogen primary experimental studies for enhanced open-access to viral Omics datasets and project lifecycle metadata. Secondary host-associated viral dataset downloads contain one or more statistically processed (normalization data transformation) quantitative dataset collections resulting in qualitative expression analyses of primary host-pathogen experimental study designs. Transcriptomics (T) dataset downloads each have a direct relationship to a primary sample submission corresponding to a specific Human Interferon (IFN), interferon alpha (IFNα), interferon beta (IFNβ), and/or interferon gamma (IFNγ) stimulated response to an experimental virus infection treatment study. Host sample types include cerebellum ["CB", BTO:0000232], cortical neurons ["CN", BTO:0004102], cortex ["CT", BTO:0000233], dendritic cells ["DC", BTO:0002042], granule cell neurons ["GCN", BTO:0003393], lymph node ["LN", BTO:0000784], and serum ["SE", BTO:0001239] from mouse (Mus musculus) tissue collections.

59 BASIC BIOLOGICAL SCIENCES↗

Omics Lethal Human Viruses Project Profiling of the Host Response to MERS-CoV Infection, Processed Experimental Dataset Catalog

Middle East Respiratory Syndrome coronavirus (MERS-CoV) is classified as a Category C priority pathogen (Coronaviridae) by the National Institute of Allergy and Infectious Diseases (NIAID), and is known to cause severe respiratory disease with high mortality rates in humans. Lethal host-pathogen invasion mechanisms and the cellular intricacies behind these fatal infections still remain unclear. The NIAID Modeling Host Responses to Understand Severe Human Virus Infections Research Program project (2013-2018) aimed to develop an improved comprehensive understanding of the host response to a suite of viruses causing lethal infections leveraging a systems biology approach. Herein, PNNL sub-projects provide a never before released comprehensive infectious disease collection of primary and secondary transformation multi-Omics data profiling a series of priority pathogen primary experimental studies for enhanced open-access to viral Omics datasets and project lifecycle metadata. Secondary host-pathogen viral dataset downloads contain one or more statistically processed (normalization data transformation) quantitative dataset collections resulting in qualitative expression analyses of primary host-pathogen experimental study designs. Leveraging unique high-resolution Omics capabilities for proteomics (P), metabolomics (M), lipidomics (L), and transcriptomics (T) dataset downloads each have a direct relationship to a primary sample submission corresponding to a specific MERS-CoV [NCBITAXON:1335626] experimental infection study. Host sample types include human lung adenocarcinoma cells ["Calu-3", BTO:0002750], human bronchial epithelial cells ["Calu-3 clone 2B4"; BTO:0002022], primary human fibroblasts ["FB"; BTO:0000452], primary human airway epithelial cells ["HAE"; BTO:0005571], human microvascular endothelial cells ["HMVE"; BTO:0003123], and whole mouse lung [BTO:0000763] tissue collections.

59 BASIC BIOLOGICAL SCIENCES↗

PNNL DataHub Project: Omics Lethal Human Viruses Project Profiling of the Host Response to West Nile Virus Infection, Processed Experimental Dataset Catalog

West Nile virus (WNV) is classified as a Category B priority pathogen (mosquito-borne Flavivirus) by the National Institute of Allergy and Infectious Diseases (NIAID), and are known to cause severe infections in humans where lethal host-associated mechanisms are not clearly defined. The NIAID Modeling Host Responses to Understand Severe Human Virus Infections Research Program project (2013 - 2018) aimed to develop an improved comprehensive understanding of the host response to a suite of viruses causing lethal infections leveraging a systems biology approach. The NIAID Modeling Host Responses to Understand Severe Human Virus Infections Research Program project (2013-2018) aimed to develop an improved comprehensive understanding of the host response to a suite of viruses causing lethal infections leveraging a systems biology approach. Herein, PNNL sub-projects provide a never before released comprehensive infectious disease collection of primary and secondary transformation multi-Omics data profiling a series of priority pathogen primary experimental studies for enhanced open-access to viral Omics datasets and project lifecycle metadata. Secondary host-pathogen viral dataset downloads contain one or more statistically processed (normalization data transformation) quantitative dataset collections resulting in qualitative expression analyses of primary host-pathogen experimental study designs. Leveraging unique high-resolution Omics capabilities for proteomics (P), metabolomics (M), lipidomics (L), and transcriptomics (T) dataset downloads each have a direct relationship to a primary sample submission corresponding a specific West Nile virus [NCBITAXON:11082] (WNV-NY99 382) experimental infection study. Host sample types include cerebellum ["CB", BTO:0000232], cortical neurons ["CN", BTO:0004102], cortex ["CT", BTO:0000233], dendritic cells ["DC", BTO:0002042], granule cell neurons ["GCN", BTO:0003393], lymph node ["LN", BTO:0000784], and serum ["SE", BTO:0001239] from mouse (Mus musculus) tissue collections.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluation of normalization strategies for mass spectrometry-based multi-omics datasets

Introduction Data normalization is crucial for multi-omics integration, reducing systematic errors and maximizing the likelihood of discovering true biological variation. Most studies assess normalization for a single omics type or use datasets from separate experiments. Few address time-course data, where normalization might bias temporal differentiation. In this study, we compared common normalization methods and a machine learning approach, Systematical Error Removal using Random Forest (SERRF), using multi-omics datasets generated from the same experiment—even from the same cell lysate. Objectives To develop a straightforward process to assess normalization effects and identify the most robust methods across multi-omics datasets. Methods We analyzed metabolomics, lipidomics, and proteomics datasets from primary human cardiomyocytes and motor neurons exposed to acetylcholine-active compounds over time. Normalization effectiveness was evaluated based on improvement in QC features consistency and observing the change in treatment and time-related variance. Results Probabilistic Quotient Normalization (PQN) and Locally Estimated Scatterplot Smoothing (LOESS) QC were identified as optimal for metabolomics and lipidomics, while PQN, Median, and LOESS normalization excelled for proteomics. These methods consistently enhanced QC feature consistency in metabolomics and lipidomics, and preserved time-related variance or treatment-related variance in proteomics, demonstrating their effectiveness and robustness. SERRF normalization, applied only to metabolomics in this study, outperformed other methods in some datasets but inadvertently masked treatment-related variance in others. Conclusion Our evaluation identified PQN and LoessQC as the top methods for metabolomics and lipidomics, and PQN, Median, and Loess normalization for proteomics, in multi-omics integration in a temporal study.

60 APPLIED LIFE SCIENCES↗

Assessment of Outliers in Alloy Datasets Using Unsupervised Techniques

We report advancements in data analytics techniques have enabled complex, disparate datasets to be leveraged for alloy design. Identifying outliers in a dataset can reduce noise, identify erroneous and/or anomalous records, prevent overfitting, and improve model assessment and optimization. In this work, two alloy datasets (9-12% Cr ferritic martensitic steels, and austenitic stainless steels) have been assessed for outliers using unsupervised techniques and supplemented with domain knowledge. Principal component analysis and k-means clustering were applied to the data, and points were assessed as outliers based on their distance away from other points in the cluster and from other points in the dataset. The outlier characteristics were investigated to determine both cluster-specific and overall trends in the properties of the outlier points. The approach demonstrated here is extensible to other alloy datasets for outlier identification and evaluation to improve the reliability of machine learning and modeling predictions for advanced alloy design.

36 MATERIALS SCIENCE↗

Development of a Benchmark Eddy Flux Evapotranspiration Dataset for Evaluation of Satellite-Driven Evapotranspiration Models Over the CONUS

A large sample of ground-based evapotranspiration (ET) measurements made in the United States, primarily from eddy covariance systems, were post-processed to produce a benchmark ET dataset. The dataset was produced primarily to support the intercomparison and evaluation of the OpenET satellite-based remote sensing ET (RSET) models and could also be used to evaluate ET data from other models and approaches. OpenET is a web-based service that makes field-delineated and pixel-level ET estimates from well-established RSET models readily available to water managers, agricultural producers, and the public. The benchmark dataset is composed of flux and meteorological data from a variety of providers covering native vegetation and agricultural settings. Flux footprint predictions were developed for each station and included static flux footprints developed based on average wind direction and speed, as well as dynamic hourly footprints that were generated with a physically based model of upwind source area. The two footprint prediction methods were rigorously compared to evaluate their relative spatial coverage. Data from all sources were post-processed in a consistent and reproducible manner including data handling, gap-filling, temporal aggregation, and energy balance closure correction. The resulting dataset included 243,048 daily and 5,284 monthly ET values from 194 stations, with all data falling between 1995 and 2021. We assessed average daily energy imbalance using 172 flux sites with a total of 193,021 days of data, finding that overall turbulent fluxes were understated by about 12% on average relative to available energy. Multiple linear regression analyses indicated that daily average latent energy flux may be typically understated slightly more than sensible heat flux. This dataset was developed to provide a consistent reference to support evaluation of RSET data being developed for a wide range of applications related to water accounting and water resources management at field to watershed scales.

54 ENVIRONMENTAL SCIENCES↗

A dataset of eco-evidence tools to inform early-stage environmental impact assessments of hydropower development

The datasets described herein provide the foundation for a decision support prototype (DSP) toolkit aimed at assisting stakeholders in determining evidence of which aspects of river ecosystems have been impacted by hydropower. The DSP toolkit and its application are presented and described in the article “Evidence-based indicator approach to guide preliminary environmental impact assessments of hydropower development” [1]. Development of the DSP and the output for decision support centralize around 42 river function indicators describing the dimensionality of river ecosystems through six main categories: biota and biodiversity, water quality, hydrology, geomorphology, land cover, and river connectivity. Three main tools are represented in the DSP: A science-based questionnaire (SBQ), an environmental envelope model (EEM), and a river function linkage assessment tool (RFLAT). The SBQ is a structured survey-style questionnaire whose objective is to provide evidence of which indicators have been impacted by hydropower. Based on a global literature review, 140 questions were developed from general hypotheses regarding the impacts of dams on rivers. The EEM is a model to predict the likelihood of hydropower impacting indicators based on a several variables. The intended use of the EEM is for situations of new hydropower development where results of the SBQ are incomplete or highly uncertain. The EEM was developed through the compilation of a dataset containing attributes of dams, reservoirs, and geospatial information on environmental concerns, which was combined with data on ecological indicators documented at those sites through literature review. The model operates through 247 “envelopes” and weighting factors, representing the individual effect of each variable on each indicator, all available through spreadsheets. Finally, the RFLAT is a tool to examine causal relationships amongst indicators. Inter-indicator relationships were hypothesized based on literature review and summarized into node and edge datasets to represent the structure of a graphical network. Bayes theorem was used estimate conditional probabilities of inter-indicator relationships based on the output of the SBQ. Nodes and edges were imported into R programming environment to visualize ecological indicator networks. The datasets can be expanded upon and enriched with more detailed questions for the SBQ, building upon the EEM with to develop more sophisticated models, and identifying new relationships for the RFALT. Additionally, once the tools are applied to numerous hydropower developments, the output of the tools (e.g. evidence of impacted indicators) becomes a very useful dataset for meta-analyses of hydropower impacts.

13 HYDRO ENERGY↗