Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data repository”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

University Data Management Pilot Utilizing the Nuclear Research Data System

Background In 2022, the Office of Science and Technology Policy (OSTP) issued a memo that significantly reshaped the landscape of access to federally funded research. The memo mandated that all taxpayer-funded research be made available to the public without delay upon publication, without an embargo period, superseding the 2013 OSTP public access policy. This public access policy promotes transparency and the democratization of knowledge, ensuring that the fruits of scientific endeavors funded by federal agencies could be immediately accessed and built upon by scientists, educators, students, and the public at large. To implement the requirements of the OSTP guidance and DOE Public Access Plan, the Office of Nuclear Energy (NE) has implemented public access plan guidance and has identified several areas where better data management practices would further expand public access to important nuclear energy related scientific data, reports, and other technical products. Significant NE supported efforts are already underway for data management and public access to important nuclear energy related data.1 2 To address gaps in data management practices, and improve retention and accessibility of data, NE is actively exploring enhanced data management options utilizing its high-performance computing resources administered by its Nuclear Scientific User Facility Program. A newly piloted system, the Nuclear Research Data System (NRDS) acts as a portal for data collection and dissemination. Nuclear Energy University Program Research and Development Portfolio According to Web of Science, NEUP has produced 2,345 journal publication that have been cited more than 61,000 times3 and countless conference proceedings. These publications are publicly available through OSTI.gov and in the open literature. Additional scientific and technical products including project milestones that are not publications and NEUP project final reports are vetted through OSTI.gov and released once reviewed and approved by DOE. Since 2009, NEUP has awarded close to 1,000 different R&D projects in technical areas across the NE research programs. As of June 2023, 512 NEUP reports are publicly available on OSTI. The underlying data for projects is still held at universities, and data transfer, co-location, and dissemination has not occurred in a systematic way. NEUP data is currently accessible through myriad university-based data repositories, or through direct requests to PIs. The program identified this patchwork of repositories, or often lack of publicly available data, as a significant barrier to an organized, accessible, and comprehensive solution to sharing data with the larger nuclear energy community. Approach The goal of this pilot project is to establish a pathway to a consolidated long-term repository for NEUP project data. To accomplish this goal, the pilot strives to accomplish the following objectives: Establish data collection standards, including a standard set of required supplementary information to contextualize and support raw data files. Work with the HPC group collect and upload information and to modify the NRDS system, as needed, to support a standardized approach. Resolve potential barriers to successful roll out of an expanded data collection strategy, including modifying data management plan guidelines and establishing a document and data release process that accounts for potential intellectual property and/or export control concerns. Results Overall, the pilot was successful in collecting 8,982 raw and processes data files, 220 reports, 56 calibration files, and 5,931 other supplementary documents. Supplementary documents included experimental plans, methods, journal publications and conference proceedings, milestone reports, and final reports. Figure 2 shows the number of data sets and supplementary project information provided by each project. Projects has significantly different input, depending on experimental data produced and completeness of the datasets provided.

Data collection↗

Hydropower Capacity Factor Trends & Analytics for the United States

This data repository contains all code, input data, and data generated for Turner et al. (2024)—“Hydropower capacity factors trending down in the United States”. File descriptions: – hydro-cf-trends-inputs.zip: Full set of input data used in this study, organized for direct entry into “/data” directory of hydro-cf-trends data processing pipeline. – hydro-cf-trends.zip: Full data processing pipeline, coded using the R {targets} framework. This is a snapshot release (v1.0) of the code repository stored at https://code.ornl.gov/turnersw/hydro-cf-trends/. – hydro-cf-trends-results.zip: Provides all dam level results required to reproduce results and graphics in Turner et al. (2024). Dams are identified by the “complxID” (root of the hydropower plant ID in the Existing Hydropower Assets Database, inherited from HILARRI). Results include: • dam_CF_trends.csv: Table of long-term trends in annualized capacity factors for 610 dams and modeled annualized capacity factors for 362 modeled dams (naturalized and assimilated flows). • dam_annualized_CF_gen.csv: Annualized time series of the following variables for each of 610 hydropower dams with nameplate > 5MW – Reported nameplate capacity (MW) – Implied maximum annual generation (MWh) – Reported net generation (MWh) – Computed annual capacity factor – Modeled annual capacity factor (362 modeled plants only)

13 HYDRO ENERGY↗

Pathways to a Sustainable Aviation Ecosystem: Flight DNA: An Anonymized Aviation Data Tool and Repository

The National Renewable Energy Laboratory (NREL) has deep experience developing secure, national data repositories, which it augments with analysis, technology and market expertise, high-performance computing, and innovative data visualization. By adapting the architecture of existing mobility databases (i.e., Fleet DNA, the Transportation Secure Data Center, the National Fuel Cell Technology Evaluation Center), NREL can build a powerful aviation data clearinghouse - Flight DNA - that helps stakeholders navigate the web of pitfalls and possibilities generated by new aviation technologies.

aviation↗

AmeriFlux BASE data pipeline to support network growth and data sharing

Abstract AmeriFlux is a network of research sites that measure carbon, water, and energy fluxes between ecosystems and the atmosphere using the eddy covariance technique to study a variety of Earth science questions. AmeriFlux’s diversity of ecosystems, instruments, and data-processing routines create challenges for data standardization, quality assurance, and sharing across the network. To address these challenges, the AmeriFlux Management Project (AMP) designed and implemented the BASE data-processing pipeline. The pipeline begins with data uploaded by the site teams, followed by the AMP team’s quality assurance and quality control (QA/QC), ingestion of site metadata, and publication of the BASE data product. The semi-automated pipeline enables us to keep pace with the rapid growth of the network. As of 2022, the AmeriFlux BASE data product contains 3,130 site years of data from 444 sites, with standardized units and variable names of more than 60 common variables, representing the largest long-term data repository for flux-met data in the world. The standardized, quality-ensured data product facilitates multisite comparisons, model evaluations, and data syntheses.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

ESS-DIVE guidelines for archiving terrestrial model data

This dataset contains supporting documents and images for ESS-DIVE terrestrial model data archiving guidelines.Terrestrial models are broadly defined as numerical models that couple both land dynamics and energy, water, carbon, or nutrient fluxes. We created these guidelines based on input from the U.S. Department of Energy’s Biological and Environmental Research land modeling community. The guidelines are intended to help modelers determine which components of their terrestrial model data associated with publication should be archived. Based on input from the land modeling community, the guidelines recommend archiving both model input and testing data, as well as code, script, and metadata. The guidelines also recommend archiving model data output, depending on the limitations set by data repositories. Lastly, we provide recommendations for bundling data files for publication as well as a discussion about tools that can facilitate model data archiving and reuse.This dataset is an archive of the associated GitHub repository for our model archiving guidelines (https://github.com/ess-dive-community/essdive-model-data-archiving-guidelines). The ‘README.pdf’ file gives a general introduction to the guidelines, and the ‘instructions.pdf’ file provides more detailed steps for following the guidelines. We also provide 2 figures in this data package: 1) a decision tree (model_data_guidelines_decision_tree.png) that can help users determine which components of their model data to archive. and 2) the ‘model_data_guidelines_flmd.png’ file depicts the different files that can be archived in addition to the model data itself. Lastly, we include 3 digitized tables from our associated manuscript and 3 CSV files with anonymized input from DOE scientists about the importance of different aspects of model data archiving from which we developed the guidelines.Dataset updates for v1.1.0: We updated this data package on 2021-11-22 in response to review comments on our related manuscript. In this update we removed one figure so that the model archiving guidelines are conveyed in text rather than an image. We updated the file-level metadata (FLMD) figure to be in accord with the most recent FLMD recommendations. We made minor edits to the README file to update the recommended citation and added two co-authors. We also added 6 new data files (3 are anonymized input from DOE scientists that helped to inform guidelines, and 3 are digitized tables from our manuscript.

54 ENVIRONMENTAL SCIENCES↗

The Colorado East River Community Observatory Data Collection

Abstract The U.S. Department of Energy's (DOE) Colorado East River Community Observatory (ER) in the Upper Colorado River Basin was established in 2015 as a representative mountainous, snow‐dominated watershed to study hydrobiogeochemical responses to hydrological perturbations in headwater systems. The ER is characterized by steep elevation, geologic, hydrologic and vegetation gradients along floodplain, montane, subalpine, and alpine life zones, which makes it an ideal location for researchers to understand how different mountain subsystems contribute to overall watershed behaviour. The ER has both long‐term and spatially‐extensive observations and experimental campaigns carried out by the Watershed Function Scientific Focus Area (SFA), led by Lawrence Berkeley National Laboratory, and researchers from over 30 organizations who conduct cross‐disciplinary process‐based investigations and modelling of watershed behaviour. The heterogeneous data generated at the ER include hydrological, genomic, biogeochemical, climate, vegetation, geological, and remote sensing data, which combined with model inputs and outputs comprise a collection of datasets and value‐added products within a mountainous watershed that span multiple spatiotemporal scales, compartments, and life zones. Within 5 years of collection, these datasets have revealed insights into numerous aspects of watershed function such as factors influencing snow accumulation and melt timing, water balance partitioning, and impacts of floodplain biogeochemistry and hillslope ecohydrology on riverine geochemical exports. Data generated by the SFA are managed and curated through its Data Management Framework. The SFA has an open data policy, and over 70 ER datasets are publicly available through relevant data repositories. A public interactive map of data collection sites run by the SFA is available to inform the broader community about SFA field activities. Here, we describe the ER and the SFA measurement network, present the public data collection generated by the SFA and partner institutions, and highlight the value of collecting multidisciplinary multiscale measurements in representative catchment observatories.

54 ENVIRONMENTAL SCIENCES↗

The disCO2ver Platform: Curating Data and Tools for Geologic Carbon Sequestration and Deep Subsurface Research Systems

The U.S. DOE National Energy Technology Laboratory has invested 12+ years of development into the data repository and digital laboratory, the Energy Data eXchange (EDX, edx.netl.doe.gov). Supporting a variety of research areas across the DOE Office of Fossil Energy and Carbon Management, the platform has successfully curated and preserved thousands of data products from DOE research. The Carbon Storage Program has successfully supported data curation, upload, and publishing of data products on EDX for many years, demonstrating a success story of how resources like EDX can effectively help with long term preservation and publishing of DOE data products. EDX continues to shift towards cloud-supported infrastructure, taking a hybrid approach combining on-premises compute and storage integrated with cloud-hosted services. The integration of cloud compute and hybrid architecture enables the development of EDX-hosted platforms that tailor the data and tools hosted on them to a specific community, enables implementation of machine learning tools for data discovery and filtering, and enables the hosting of virtual (online user interface) tools. Geologic carbon sequestration (GCS) research continues to scale up in response to the current administration goals to reduce greenhouse gas emissions and transition the energy economy. Over the last year, EDX’s disCO2ver platform has been developed in response to the need for access to data products and tools to support the scaling up of GCS research. disCO2ver provides access to data resources and tools, produced by DOE and outside authoritative external resources. The platform also provides a user-access control component for the virtualization and cloud hosting of tools. Tools that need to be virtualized, to eliminate the need for users to download the tool and use local compute resources, is essential to supporting big-data analysis and machine learning that is becoming common place in carbon storage modeling, risk analysis, and data publishing practices. This talk will review the EDX’s disCO2ver platform and the current work ongoing to curate data and tools to support GCS and deep subsurface systems research.

Morkner, Paige↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

PDBx/mmCIF Ecosystem: Foundational Semantic Tools for Structural Biology

PDBx/mmCIF, Protein Data Bank Exchange (PDBx) macromolecular Crystallographic Information Framework (mmCIF), has become the data standard for structural biology. With its early roots in the domain of small-molecule crystallography, PDBx/mmCIF provides an extensible data representation that is used for deposition, archiving, remediation, and public dissemination of experimentally determined three-dimensional (3D) structures of biological macromolecules by the Worldwide Protein Data Bank (wwPDB, wwpdb.org). Extensions of PDBx/mmCIF are similarly used for computed structure models by ModelArchive (modelarchive.org), integrative/hybrid structures by PDB-Dev (pdb-dev.wwpdb.org), small angle scattering data by Small Angle Scattering Biological Data Bank SASBDB (sasbdb.org), and for models computed generated with the AlphaFold 2.0 deep learning software suite (alphafold.ebi.ac.uk). Community-driven development of PDBx/mmCIF spans three decades, involving contributions from researchers, software and methods developers in structural sciences, data repository providers, scientific publishers, and professional societies. Having a semantically rich and extensible data framework for representing a wide range of structural biology experimental and computational results, combined with expertly curated 3D biostructure data sets in public repositories, accelerates the pace of scientific discovery. Herein, we describe the architecture of the PDBx/mmCIF data standard, tools used to maintain representations of the data standard, governance, and processes by which data content standards are extended, plus community tools/software libraries available for processing and checking the integrity of PDBx/mmCIF data. Use cases exemplify how the members of the Worldwide Protein Data Bank have used PDBx/mmCIF as the foundation for its pipeline for delivering Findable, Accessible, Interoperable, and Reusable (FAIR) data to many millions of users worldwide.

59 BASIC BIOLOGICAL SCIENCES↗

WELLBASE - An Interactive Platform for Wellbore Material Assessment

This project seeks to build an open-source wellbore material data repository with adequate material performance and contextual data to support Geological Carbon Storage (GCS). By appropriately evaluating the data types as mentioned earlier made available by the WELLBASE tool, stakeholders can make more informed decisions regarding well selections, risk assessment, and economic analysis for geologic carbon storage projects. Advanced Natural Language Processing models and other custom python scripts will be deployed in an automated process to extract unstructured data from documents, reports, and web applications and subsequently parse to more usable formats. The processed data will then be integrated into a robust and comprehensive database architecture, optimizing data accessibility, and usability for analytical purposes. The final data products will be accessible through a user-friendly visualization platform that will allow users to query and visualize the data, as well as download data in usable formats.

Tetteh, Daniel A.↗

Performance and wake flow characterization of a 1:8.7-scale reference USDOE MHKF1 hydrokinetic turbine to establish a verification and validation test database

As hydrokinetic turbine technologies continue to advance towards commercialization, public datasets on the performance characteristics for these devices and their flow field effects are invaluable to advance our understanding of these technologies and to validate analytical and numerical models. The Applied Research Laboratory at The Pennsylvania State University (ARL Penn State) collaborated with Sandia National Laboratories and the University of California at Davis to design, fabricate (at a 1:8.7 scale), and experimentally test a novel hydrokinetic turbine rotor design to provide an open platform and dataset for further study and development. The water tunnel test of this three-bladed, horizontal-axis rotor recorded power production, blade loading, the near-wake flow, cavitation effects, and noise generation. These state-of-the-art measurements demonstrate much of the complex physics associated with the flow through an unducted, horizontal-axis turbine, and they elucidate the performance characteristics and flow field effects at an unprecedented fidelity, accuracy and resolution. Measurements of powering coefficients (power, torque and thrust) as a function of tip-speed-ratio were performed. The dataset also includes unsteady measurements of driveshaft loading, blade strain, tower pressures, and radiated noise. Detailed flow mapping using laser Doppler velocimetry, and planar and stereo particle image velocimetry includes measurements of mean velocity and Reynolds stresses. Although the wake measurements are limited to less than half a diameter, they reveal the complex flow patterns in the near-wake structure of the rotor. The full database, available at the United States Department of Energy’s marine and hydrokinetic data repository, includes tunnel and model Computer Aided Design geometry files and inflow data sufficient for a “Model-the-Test” computational Verification and Validation study.

13 HYDRO ENERGY↗

A Comprehensive Northern Hemisphere Particle Microphysics Data Set From the Precipitation Imaging Package

Microphysical observations of precipitating particles are critical data sources for numerical weather prediction models and remote sensing retrieval algorithms. However, obtaining coherent data sets of particle microphysics is challenging as they are often unindexed, distributed across disparate institutions, and have not undergone a uniform quality control process. This work introduces a unified, comprehensive Northern Hemisphere particle microphysical data set from the National Aeronautics and Space Administration precipitation imaging package (PIP), accessible in a standardized data format and stored in a centralized, public repository. Data is collected from 10 measurement sites spanning 34° latitude (37°N–71°N) over 10 years (2014–2023), which comprise a set of 1,070,000 precipitating minutes. The provided data set includes measurements of a suite of microphysical attributes for both rain and snow, including distributions of particle size, vertical velocity, and effective density, along with higher-order products including an approximation of volume-weighted equivalent particle densities, liquid equivalent snowfall, and rainfall rate estimates. The data underwent a rigorous standardization and quality assurance process to filter out erroneous observations to produce a self-describing, scalable, and achievable data set. Case study analyses demonstrate the capabilities of the data set in identifying physical processes like precipitation phase-changes at high temporal resolution. Bulk precipitation characteristics from a multi-site intercomparison also highlight distinct microphysical properties unique to each location. This curated PIP data set is a robust database of high-quality particle microphysical observations for constraining future precipitation retrieval algorithms, and offers new insights toward better understanding regional and seasonal differences in bulk precipitation characteristics.

54 ENVIRONMENTAL SCIENCES↗

EDGE COVID-19: a web platform to generate submission-ready genomes from SARS-CoV-2 sequencing efforts

Abstract Summary Genomics has become an essential technology for surveilling emerging infectious disease outbreaks. A range of technologies and strategies for pathogen genome enrichment and sequencing are being used by laboratories worldwide, together with different and sometimes ad hoc, analytical procedures for generating genome sequences. A fully integrated analytical process for raw sequence to consensus genome determination, suited to outbreaks such as the ongoing COVID-19 pandemic, is critical to provide a solid genomic basis for epidemiological analyses and well-informed decision making. We have developed a web-based platform and integrated bioinformatic workflows that help to provide consistent high-quality analysis of SARS-CoV-2 sequencing data generated with either the Illumina or Oxford Nanopore Technologies (ONT). Using an intuitive web-based interface, this workflow automates data quality control, SARS-CoV-2 reference-based genome variant and consensus calling, lineage determination and provides the ability to submit the consensus sequence and necessary metadata to GenBank, GISAID and INSDC raw data repositories. We tested workflow usability using real world data and validated the accuracy of variant and lineage analysis using several test datasets, and further performed detailed comparisons with results from the COVID-19 Galaxy Project workflow. Our analyses indicate that EC-19 workflows generate high-quality SARS-CoV-2 genomes. Finally, we share a perspective on patterns and impact observed with Illumina versus ONT technologies on workflow congruence and differences. Availability and implementation https://edge-covid19.edgebioinformatics.org, and https://github.com/LANL-Bioinformatics/EDGE/tree/SARS-CoV2. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

LinkML: an open data modeling framework

Background Scientific research relies on well-structured, standardized data; however, much of it is stored in formats such as free-text lab notebooks, nonstandardized spreadsheets, or data repositories. This lack of structure challenges interoperability, making data integration, validation, and reuse difficult. Findings LinkML (Linked Data Modeling Language) is an open framework that simplifies the process of authoring, validating, and sharing data. LinkML can describe a range of data structures, from flat, list-based models to complex, interrelated, and normalized models that utilize polymorphism and compound inheritance. It offers an approachable syntax that is not tied to any one technical architecture and can be integrated seamlessly with many existing frameworks. The LinkML syntax provides a standard way to describe schemas, classes, and relationships, allowing modelers to build well-defined, stable, and optionally ontology-aligned data structures. Once defined, LinkML schemas may be imported into other LinkML schemas. These key features make LinkML an accessible platform for interdisciplinary collaboration and a reliable way to define and share data semantics. Conclusions LinkML helps reduce heterogeneity, complexity, and the proliferation of single-use data models while simultaneously enabling compliance with FAIR (Findable, Accessible, Interoperable, and Reusable) data standards. LinkML has seen increasing adoption in various fields, including biology, chemistry, biomedicine, microbiome research, finance, electrical engineering, transportation, and commercial software development. In short, LinkML makes implicit models explicitly computable and allows data to be standardized at their origin. LinkML documentation and code are available at https://linkml.io/.

AI-ready data↗

A database of refractive indices and dielectric constants auto-generated using ChemDataExtractor

The ability to auto-generate databases of optical properties holds great potential for advancing optical research, especially with regards to the data-driven discovery of optical materials. An optical property database of refractive indices and dielectric constants is presented, which comprises a total of 49,076 refractive index and 60,804 dielectric constant data records on 11,054 unique chemicals. The database was auto-generated using the state-of-the-art natural language processing software, ChemDataExtractor, using a corpus of 388,461 scientific papers. The data repository offers a representative overview of the information on linear optical properties that resides in scientific papers from the past 30 years. Public availability of these data will enable a quick search for the optical property of certain materials. The large size of this repository will accelerate data-driven research on the design and prediction of optical materials and their properties. To the best of our knowledge, this is the first auto-generated database of optical properties from a large number of scientific papers. We provide a web interface to aid the use of our database.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Functional characterization of prokaryotic dark matter: the road so far and what lies ahead

Eight-hundred thousand to one trillion prokaryotic species may inhabit our planet. Yet, fewer than two-hundred thousand prokaryotic species have been described. This uncharted fraction of microbial diversity, and its undisclosed coding potential, is known as the “microbial dark matter” (MDM). Next-generation sequencing has allowed to collect a massive amount of genome sequence data, leading to unprecedented advances in the field of genomics. Still, harnessing new functional information from the genomes of uncultured prokaryotes is often limited by standard classification methods. These methods often rely on sequence similarity searches against reference genomes from cultured species. This hinders the discovery of unique genetic elements that are missing from the cultivated realm. It also contributes to the accumulation of prokaryotic gene products of unknown function among public sequence data repositories, highlighting the need for new approaches for sequencing data analysis and classification. Increasing evidence indicates that these proteins of unknown function might be a treasure trove of biotechnological potential. Here, we outline the challenges, opportunities, and the potential hidden within the functional dark matter (FDM) of prokaryotes. We also discuss the pitfalls surrounding molecular and computational approaches currently used to probe these uncharted waters, and discuss future opportunities for research and applications.

59 BASIC BIOLOGICAL SCIENCES↗

Sapflow and xylem water isotopes from Snodgrass Mountain, East River Watershed, Colorado USA

This dataset includes sapflux and stable water isotopes of soil water and xylem water for aspen, fir and spruce trees along the Snodgrass Mountain transect in the East River Watershed, Colorado USA. The purpose of generating this dataset was to understand: (1) the total flux of water being used by trees in the watershed, (2) separate the component of transpiration that was derived from recent precipitation vs. older water such as winter precipitation of groundwater and (3) understand how total water use and water sources for the trees varies between species and position on a hillslope. The data were collected from May 2019 until October 2022. The sap flux data were collected using ICT SFM1 sensors and are presented in both units of cm h^-1 and as kg h^-1 by multiplying the sap flux by the sapwood area of the tree. All sap flow data has been been corrected using estimates of wounding diameter, water content of wood and sap wood depth. The xylem water isotope data were collected approximately weekly from each of the trees instrumented with sap flux. The stems were collected and the water extracted using classic cryogenic methods. Isotope measurements were done using a Picarro 2140i analyzer. We provide an estimate of the Seasonal Origin Index for each measurements following Allen et al., 2019 (10.5194/hess-23-1199-2019) where values of -1 equate to trees relying on winter precipitation and +1 tree relying on summer precipitation. We also provide the simultaneous sap flux for each isotope measurement when this data was available. This is an update to an earlier data repository with the same name that only included 2019 data. The new dataset was posted in January 2023 and is inclusive of the original 2019 data but now includes 2021 and 2022 data. Please note there was an error in the units of transpiration in the original dataset. It was incorrectly listed as mm h-1 when it should have been kg h-1. This change was made in May 2024.

54 ENVIRONMENTAL SCIENCES↗

Projections of Hourly Meteorology by Balancing Authority Based on the IM3/HyperFACETS Thermodynamic Global Warming (TGW) Simulations

This dataset contains 40 years (1980-2019) of historical hourly meteorology and 80 years (2020-2099) of projected hourly meteorology for 54 Balancing Authorities (BAs) in the conterminous United States. Details about the scenarios and variables included in this dataset are in the readme.pdf file. This dataset is derived from the IM3/HyperFACETS Thermodynamic Global Warming (TGW) simulations (https://doi.org/10.57931/1885756). More details on the TGW approach can be found at: https://tgw-data.msdlive.org/. If you use this dataset please also cite the raw TGW dataset (Jones, A. D., Rastogi, D., Vahmani, P., Stansfield, A., Reed, K., Thurber, T., Ullrich, P., & Rice, J. S. (2022). IM3/HyperFACETS Thermodynamic Global Warming (TGW) Simulation Datasets (v1.0.0) [Data set]. MSD-LIVE Data Repository. https://doi.org/10.57931/1885756). To go from the TGW data to these BA-level aggregated data we first averaged the raw gridded data by county in the United States. That intermediate data step is also stored in MSD-LIVE (https://doi.org/10.57931/1960548). We then population-weight the county-level hourly data in order to create population-weighted meteorology time series for each BA. Historical populations are from the United States Census Bureau and the evolving future populations are based on Shared Socioeconomic Pathways (SSPs) 3 and 5. The four climate scenarios crossed with the two SSPs yield eight different future projections for each BA: rcp45cooler_ssp3, rcp45cooler_ssp5, rcp45hotter_ssp3, rcp45hotter_ssp5, rcp85cooler_ssp3, rcp85cooler_ssp5, rcp85hotterssp3, rcp85hotterssp5. For the historical period and each of the eight future scenarios the dataset has hourly estimates of the population-weighted average of five meteorological variables: Temperature, specific humidity, shortwave radiation, longwave radiation, and wind speed. All times are in Coordinated Universal Time (UTC). The code to go from the raw TGW data to county-level and then BA-level projections is available at: https://github.com/IMMM-SFA/im3components/tree/main/im3components/wrf_to_tell.

Balancing Authority↗