Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Metadata Normalization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Xanthos-Lake Dataset

The Xanthos-Lake v1.0 dataset provides the input data, trained machine-learning models, and simulation outputs needed to characterize lake water balance, snow and ice conditions, and mixing-layer temperature within the Xanthos global hydrological modeling framework. The dataset supports lake representation across a wide range of lake sizes and hydroclimatic conditions by combining xLSIM, a basin-specific machine-learning emulator of lake snow, ice, ice-cover fraction, and mixing-layer temperature, with the Xanthos-Lake water-balance model. The archive contains NetCDF datasets used to train and evaluate xLSIM, trained model weights, processed meteorological and lake-property inputs, and basin- and lake-category-specific simulation outputs. These materials are organized into four primary data groups, described below. Snowice_model_inputs: Contains the NetCDF input data used to train xLSIM. The xLSIM machine-learning framework uses three lake-based datasets. The meteorological forcing dataset provides monthly relative humidity, specific humidity, surface wind speed, maximum and minimum air temperature, downward longwave and shortwave radiation, snowfall, surface air pressure, and total precipitation. Lake surface area is included as an additional static predictor. The target-state dataset provides lake ice thickness, snow depth, snow cover, and lake mixing-layer temperature, while a companion lake-surface dataset provides the lake ice-cover fraction. Before training, ice thickness and snow depth are converted from meters to centimeters, mixing-layer temperature is converted from kelvin to degrees Celsius and constrained to nonnegative values, and ice-cover fraction is converted from a fraction to a percentage. The predictor variables are normalized using statistics calculated across the selected lakes and time steps. Snowice_model_outputs: Contains the NetCDF outputs generated by xLSIM. For each basin, xLSIM produces a file containing observed and predicted lake-state variables for the training, validation, and testing periods. The modeled variables include lake ice thickness, snow depth, snow cover, mixing-layer temperature, and lake ice-cover fraction. For basins without a sufficiently persistent snow-and-ice signal, the emulator predicts only mixing-layer temperature. The outputs also include training and validation loss histories, the selected model configuration, identifiers of the lakes used in training, and SHAP-based feature-importance information at the global, lake, and seasonal-regime levels. The trained machine-learning model weights are provided separately within the dataset archive. Together, these files support model evaluation and subsequent coupling with the Xanthos-Lake water-balance framework. XanthosLAKES: Contains the NetCDF input data used by the Xanthos-Lake framework. Monthly meteorological inputs include relative and specific humidity, downward shortwave and longwave radiation, mean, maximum, and minimum air temperature, wind speed, precipitation, snowfall, and surface air pressure. Static lake-property datasets provide lake identifiers, geographic locations, surface area, volume, mean depth, elevation, drainage area, fetch, outlet-routing information, and associated Xanthos grid-cell attributes. Separate bathymetric datasets provide the coefficients of the area–depth and volume–depth relationships for each aggregated lake unit. GLEV-based records provide observed lake surface area and evaporation data used to initialize lake states, define reference conditions, and calibrate and evaluate the model. Xanthos-Lake Outputs: Contains the basin- and lake-category-specific NetCDF outputs generated by Xanthos-Lake. Monthly variables include lake surface area, storage volume, outlet discharge, evaporation rate, evaporation volume, lake–groundwater exchange, lake inflow, ice thickness, snow depth, snow-cover fraction, ice-cover fraction, and mixing-layer temperature. The files also contain lake-specific calibration and validation statistics, including normalized root-mean-square error, mean absolute error, Nash–Sutcliffe efficiency, Kling–Gupta efficiency, and percent bias. Stored calibrated and derived parameters include the weir discharge coefficient, fractional freeboard, groundwater exchange coefficient, reference water level, corresponding reference surface area and storage volume, weir-width adjustment factor, and the fraction of routed inflow entering the lake. Basin identifiers, lake category, simulation period, calibration and validation periods, and parameter-schema information are retained as NetCDF metadata.

Abeshu, Guta [Pacific Northwest National Laborator↗

COMPASS-FME Synoptic Sites Level 1 Sensor Data v1-2

This is the version 1-2 Level 1 (L1) data release for COMPASS-FME environmental sensors located at our synoptic field sites. COMPASS-FME is studying sites in two distinct regions, the Chesapeake Bay and the Western Lake Erie Basin. We established the network at seven "synoptic" (observational) sites along the Chesapeake Bay and Lake Erie coastlines, collectively generating over three million observations per month, to track and comprehend environmental changes where land and water intersect. Additionally, the two regions provide an interesting contrast of saltwater and freshwater coasts that allow us to differentiate the impacts of inundation and coastal water chemistries in two nationally important coastal systems.L1 data are close to raw, but are units-transformed and have out-of-instrument-bounds and out-of-service flags added. Duplicates and missing data are removed but otherwise these data are not filtered, and have not been subject to any additional algorithmic or human QA/QC. Any scientific analyses of L1 data should be performed with care. **This dataset will be updated quarterly with new data for the duration of the project**This dataset includes:- An overall dataset README file that describes the current version, gives citation and contact information, etc.- Site- and year-specific folders, each holding up to 12 CSV (comma separated value) data files for each site and plot in that year.- Metadata files within each site-year folder provide full information on data units, expected ranges, contact information, as well as a general description of the site.- Environmental sensor types that appear in the data files include weather (ClimaVUE50, CS, RM Young, and LI instruments in the graphs below); soil conditions (TEROS12); soil redox state (Redox); groundwater variables (AquaTROLL200 and AquaTROLL600); open water sondes (Exo); tree sap velocity (Sapflow); and system voltage and state (Datalogger). Data are normally logged every 15 minutes. Please see v1-2 Synoptic L1 Sensor Package Quick Start.pdf for detailed information on data package structure, temporal coverage, and versioning.

54 ENVIRONMENTAL SCIENCES↗

PPI DataHub Project Data Package: S. elongatus PCC 7942 Limited Proteolysis and Thermal Proteome Profiling Structural Proteomics (JM-PB-DP3)

The purpose of this experiment was to investigate structural alterations in proteins involved in central carbon metabolism and photosynthetic electron transfer pathways in Synechococcus elongatus PCC 7942. Sample data was obtained from S. elongatus cell lysates using three complementary mass spectrometry (MS) techniques using limited proteolysis (LiP-MS), thermal proteome profiling (TPP-MS), and redox enrichment (Redox-MS) in evaluating alterations solvent accessibility and structural stability caused by light perturbation at the molecular level. Experimentally processed sample data for LiP and TPP proteomic datasets were derived from the same cell culture stock, prepared simultaneously in parallel, and acquired by mass spectrometry. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files, computed outputs, and supporting metadata materials. Experimental samples processed for LiP-MS label-free quantification (LFQ) or TPP-MS tandem mass tag (TMT) 10-plex were acquired using a Q-Exactive HF-X mass spectrometer and processed/compiled using either MSGF+ (v2024.03.26) or ​​​​PlexedPiper for proteome evaluation. Additional software supporting downstream proteomic analysis include FragPipe (v.4.0), MSFragger (v.22.1), and an adapted Microbial Isolate LiP Analysis Workflow (located at Zenodo). Processed proteomic data downloads include a sample naming key, normalized quantification results files, and processed protein annotated abundance files.

59 BASIC BIOLOGICAL SCIENCES↗

Vegetation classification map and covariates associated with NEON AOP survey, East River, CO 2018

This package includes geospatial data layers developed to investigate how environmental gradients—specifically topography and near-surface soil properties—drive the spatial arrangement of dominant plant communities in mountainous watersheds. The geospatial products, which support the analysis of these ecological relationships, are derived from airborne hyperspectral and LiDAR datasets acquired by the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP), in conjunction with an extensive ground field campaign conducted in summer 2018. This work is part of the DOE Watershed Function Science Focus Area (SFA) and features geospatial datasets developed based on observations and ground data collected at East River, Colorado, in collaboration with the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP) survey in June 2018. Classification Map: - Classification Map (PNG, GeoTIFF): Derived from hyperspectral and LiDAR airborne data using a machine learning approach. - Class Code Mapper (CSV): Associates pixel values with corresponding vegetation/non-vegetation classes. - Classification Reference Data (CSV): Reference data used in the machine learning procedure. LiDAR-Derived Products: - Topographical Metrics (GeoTIFFs): Elevation, slope, curvature, TWI, TPI, solar insolation, and canopy height model (CHM), smoothed with a 5x5 pixel window. Vegetation Indices: - GeoTIFFs of NDVI, NDNI, NDWI: Vegetation indices derived from hyperspectral data. Urban Masks: - Urban Mask (GeoTIFF): Applied to the mapping to convert bare soil classes to urban classes. Software Compatibility: GeoTIFFs: Can be visualized with GIS software or libraries that support GeoTIFF images. CSV Files: Can be opened with any software that handles comma-separated values. The FLMD file provides details and links to the source datasets used to derive the products. The manuscript (in the Method session) provides details on how each product was derived. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. Update on 2026-03-25: Since the original dataset publication date of 02/28/2020, this package has a new classification map derived by an improved methodology. This update also includes additional ground data that improved the representation of some of the communities. See the methods for further details on what has changed between versions.

2018 NEON and 2025 CHESS Campaigns↗

WHONDRS River Corridor Sediment and Water Geochemistry and In Situ Sensor Data from 7 Perennial and 7 Intermittent Streams across San Antonio, Texas (v3)

This dataset supports a broader study examining the effects of intermittency on sediment respiration. The dataset provides sediment and surface water geochemistry and in situ sensor data from 7 perennial and 7 intermittent streams in San Antonio, Texas. Each stream/site was visited both in summer during base flow (July-September 2023) and winter during peak flow (January-February 2024). Related data were collected and will be published separately in collaboration with A. Veach. The data package was originally published in April 2025. It was updated in June 2025 (v2; modified and new files) and September 2025 (v3; modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of two folders of field photos and videos, one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) international generic sample number (IGSN) mapping file; (6) field protocol; (7) a subfolder with sample data; and (8) a subfolder with sensor data. The sample data subfolder contains (1) surface water and sediment dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) surface water and sediment total nitrogen data and averages; (3) sediment grain size data; (4) sediment iron (II) data and averages; (5) wet sediment mass, dry sediment mass, water mass, and wet sediment volume in incubation and sediment ICR vials; (7) sediment incubation respiration rate data and averages; (8) normalized respiration rate data and averages; (9) methods codes; (10) sediment percent carbon and nitrogen; (11) sediment X-ray diffraction (XRD) data; (12) gravimetric moisture and averages; (13) a subfolder with sediment incubation respiration data, scripts, and plots; (14) surface water and sediment FTICR methods; and (15) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains five subfolders, one containing the sediment .xml data files, one containing the water .xml files, one containing the sediment CoreMS output files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). The sensor data subfolder contains (1) a subfolder with miniDOT dissolved oxygen and temperature data and plots; (2) miniDOT dissolved oxygen and temperature summary data; and (3) miniDOT installation methods. All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4. CORRECTION: The data processing methods for FTICR described in “v3_WHONDRS_AV1_Methods_Codes.csv” mistakenly indicate that users should process the data in Formultitude. The corrected description should read: “Both unprocessed and processed data are provided to allow users flexibility in data processing. Instructions and scripts for processing the data using CoreMS are included.” CORRECTION: Carbon and nitrogen content are reported as percentages. The current column headers "01395_C_percent_per_mg" and "01397_N_percent_per_mg" are incorrect. These should read "01395_C_percent" and "01397_N_percent" and will be corrected in the next version of this data package.

54 ENVIRONMENTAL SCIENCES↗

Dissolved Oxygen and Temperature Data from the Hyporheic Zone of the East River Watershed July 2017 to October 2018

Dissolved oxygen (DO) is critical for aquatic ecosystems. Our focus is on the long-term DO dynamics in hyporheic zone of rivers, which are a function of both transport (hydrologic exchange between river and hyporheic zone) and uptake by biogeochemical reactions or respiration. The study site is the alpine East River watershed in Colorado, USA, meander A downstream from the pump house. Opti O2 probes were deployed in the water column and directly within the river-bed at 10, 20, and 35 cm depth (38°55'23.38"N, 106°57'4.21"W, 9046.88m) to monitor DO and temperature. A continuous data stream from July 24, 2017 to Oct 24, 2018 was autonomously telemetered to the cloud. This 14-month DO and temperature time series were obtained without any servicing for maintenance or data downloads; additionally the ability to remotely verify probe performance during field deployment was essential to confirm data validity during winter freeze-in and hydrological/weather events, such as spring melt and summer monsoons. We investigate the variations in dissolved oxygen dynamics of this snow-pack dominated watershed during a comparatively low flow water year (2018) and a relatively normal water year (2017), enabled by distinctive, in-situ, high frequency (∆t = 5min) sensors that provided a continuous time-series from the undisturbed study site over multiple seasons.

54 ENVIRONMENTAL SCIENCES↗

Maps of land surface phenology derived from PlanetScope data, 2018-2022, Teller, Kougarok, and Council, Seward Peninsula

Remote sensing maps of land surface phenology derived from PlanetScope (Planet Team, 2017) normalized differential greenness index (NDGI; Yang et al., 2019) time series data. These maps include four phenological timing metric - start of spring (SOS), end of spring (EOS), start of fall (SOF), and end of fall (EOF), corresponding NDGI value at the four phenological timings, and annual maximum and minimum NDGI. This package includes maps for Next-Generation Ecosystem Experiment Arctic (NGEE Arctic)’s Teller Mile Marker (MM) 27, Kougarok MM64, and Council MM 71 watersheds. Maps of 5 years from 2018 to 2022 are included in this dataset. The map data and metadata are provided as image (ENVI) and text (*.txt, *hdr) formats. Additional supporting map quicklooks are provided as GIS *.kml files. These datasets are provided in support of Yang et al., (In revision) “Fine-scale Landscape Characteristics, Vegetation Composition, and Snowmelt Timing Control Phenological Heterogeneity across Arctic Tundra Landscapes”The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Untargeted metabolite data from a root surface in a rhizobox

Raw data is provided from samples analyzed using separate reverse phase chromatographic methods on a high performance liquid chromatograph with mass spectrometry. These porewater samples were collected from a microdialysis which generated samples along the surface of a growing A. sative root (all_hc). This data was used to answer questions connecting rhizosphere metabolite (putatively identified metabolites, hc_putative_norm) changes over time (root growth) with changes in the surrounding rhizosphere biogeochemistry (DOC, redox, pH). Rhizosphere biogeochemistry values are provided in the hc_putative_norm file as averages over their respective range of time that they were collected at. The hc_putative_norm file also contains all normalized values over only the intensity values collected for putatively identified metabolites.

54 ENVIRONMENTAL SCIENCES↗

Normalizing Resource Identifiers using Lexicons in the Global Change Information System: Linking Earth Science Identifiers, Concepts, and Communities

Earth Science informatics involves collaboration between multiple groups of people with diverse specializations and goals,often using variations in terminology to refer to common resources. The uniformity of the resource identifiers often does not cross organizational boundaries. Because of this, permanent, widely used, unambiguous identifiers for resources are elusive. We examine real world cases of changing and inconsistent identifiers which inherently work against persistence and uniformity. We also present a solution which mediates factors in these situations; namely the creation of lexicons:mappings of sets of terms to URIs which are curated within the Global Change Information System (GCIS). We discuss aspects of the GCIS which facilitate the use of lexicons: an information model which disambiguates resources, a RESTful API which provides metadata through content-negotiation, and a strategy for long term curation of URIs, including mechanisms for handling changes to URIs and variations in terms used by different communities while providing persistent URIs and preserving relationships between resources We provide working definitions of terms,contexts, and lexicons, and relate them to the practical challenges of disambiguation and curation. We also discuss the mechanisms employed and architecture of the GCIS, and how these choices facilitate representation of persistent identifiers and mappings of them to identifiers used colloquially within various earth science communities of practice.

Linkded Data↗

WHONDRS River Corridor Sediment and Water Geochemistry and In Situ Sensor Data from Machine-Learning-Informed Sites across the Contiguous United States (v6)

This dataset supports a broader study examining hyporheic zone respiration rates to improve predictive models at a contiguous United States (CONUS) scale. The CONUS-Scale Model-Sample Study (CM) was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Sampling began in April 2022 and ended in October 2023. In addition to the widely distributed CONUS sites, a more spatially focused sampling occurred in the Yakima River Basin, WA in summer 2022. Data from this more spatially intensive sampling occurred under the label “Second Spatial Study (SSS)” and were also included in the machine learning models. Other data types collected from SSS that were not part of CM were published in a separate data package (https://data.ess-dive.lbl.gov/view/doi:10.15485/1969566). This data package was originally published in February 2023. It was updated in June 2023 (v2; new and modified files); December 2023 (v3; new and modified files); June 2024 (v4; new and modified files); April 2024 (v5; new and modified files); and September 2025 (v6; modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of two folders of field photos and videos, one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) international generic sample number (IGSN) mapping file; (6) field protocols; (7) a subfolder with sample data; and (8) a subfolder with sensor data. The sample data subfolder contains (1) surface water and sediment dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) surface water and sediment total nitrogen data and averages; (3) surface water major cations and anions and averages; (4) sediment grain size data; (5) sediment iron (II) data and averages; (6) wet sediment mass, dry sediment mass, water mass, and wet sediment volume in incubation and sediment ICR vials; (7) sediment incubation respiration rate data and averages; (8) normalized respiration rate data and averages; (9) methods codes; (10) sediment specific surface area; (11) sediment percent carbon and nitrogen; (12) sediment gravimetric moisture and averages; (15) sediment X-ray diffraction (XRD) data; (16) sediment adenosine triphosphate (ATP) and averages; (17) a subfolder with sediment incubation respiration data, scripts, and plots; (18) surface water and sediment FTICR methods; and (19) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains five subfolders, one containing the sediment .xml data files, one containing the water .xml files, one containing the sediment CoreMS output files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS).The sensor data subfolder contains (1) a subfolder with miniDOT dissolved oxygen and temperature data and plots; (2) miniDOT dissolved oxygen and temperature summary data; and (3) miniDOT installation methods. All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4. CORRECTION: Carbon and nitrogen content are reported as percentages. The current column headers "01395_C_percent_per_mg" and "01397_N_percent_per_mg" are incorrect. These should read "01395_C_percent" and "01397_N_percent" and will be corrected in the next version of this data package. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES↗

Fine-Root Ecology Database (FRED): A Global Collection of Root Trait Data with Coincident Site, Vegetation, Edaphic, and Climatic Data, Version 4.

To address the need for a centralized root trait database, we compiled the Fine-Root Ecology Database (FRED) from published and unpublished data sources. We have continued to add to the FRED database since the release of FRED 1.0 in 2017, followed by 2.0 in 2018, and 3.0 in 2021. This new release of FRED 4.0 now has 213,941 observations of 238 root traits, for a combined total of roughly 3.4 million data fields for root traits and ancillary data together. FRED 4.0 has 39.8% more root trait observations than FRED 3.0 and a 34.4% increase in unique data sources. This release of FRED 4.0 also includes significant increases in geographic regions that have long been underrepresented in global datasets, notably in the tropical low latitudes. Ancillary data on associated site, vegetation, edaphic, and climatic conditions from across the globe have also increased concurrently with root trait observations. FRED is focused on fine roots (traditionally defined as roots less than 2 mm in diameter), as coarse roots are studied using different methodology, often at very different scales, and have different traits and trait interpretations. Despite this fine-root focus, FRED accepts data collected from roots of all sizes and contains observations of many root classes including coarse roots. Data collection will continue for the foreseeable future. The FRED4_Entire_Database_2026.csv file is the flat csv data file for FRED 4.0, and the FRED4_dd.csv file is the data dictionary of all columns available in FRED, including column IDs, column names, definitions, and unit (where applicable).

54 ENVIRONMENTAL SCIENCES↗

SPRUCE Whole Ecosystem Warming (WEW) Environmental Data and Water Table Summaries, Marcell Experimental Forest, Minnesota, 2015-2024

This data set contains observations of photosynthetically active radiation (PAR), precipitation, soil temperature, soil volumetric water content, air temperature, relative humidity, and normalized water table depth that are summarized on a daily, weekly, monthly, and annual basis for each of the SPRUCE plots. Observations span 2015-2024. This dataset draws on several datasets (Hanson et al. 2016; Hanson et al. 2020; and Warren, unpublished data) and compiles these environmental observations into useful formats for data analysis. These environmental metrics can be used to understand the environmental conditions inside SPRUCE environmental chambers throughout the durations of the experiment and can be paired with other data for modeling and analysis. R code used to generate these files is provided as part of the data package. This dataset contains four data files in comma separate (.csv) format and a compressed folder (*.zip) containing three R (*.r) scripts. Additional metadata are provided: one data dictionary and a file-level metadata file in comma separate (.csv) format and a user guide in PDF (*.pdf) format. User note: Users must cite the original dataset/s along with this dataset when publishing any analyses using this dataset. Details on the dataset used to compile each variable are available in the header row of the files and in the user guide.

air temperature↗

LAI, EVI, NDVI, and kNDVI in 23 pantropical forests affected by 21 cyclones

Statement of purpose: Cyclones alter the function and composition of tropical forests, making effects of intensifying cyclones on carbon-rich forests a critical topic of study. Here, we quantified cyclone-induced damage and recovery of 21 cyclone disturbances affecting 23 pantropical forest sites between 1988-2017 utilizing leaf area index (LAI), enhanced vegetation index (EVI), normalized difference vegetation index (NDVI), and transformed NDVI (kNDVI) values from Google Earth Engine. Field observations collected in a meta-analysis (Bomfim et al., 2022, in review) were used to ground-truth and test effects of soil resource availability and disturbance factors on damage and recovery. This meta-analysis also served as the basis to begin vegetation index extraction, utilizing unique site and date combinations, from tropical forests effect by cyclone disturbances. We began collecting NDVI (5km resolution) from the NOAA Climate Data Record (CDR) of AVHRR Normalized Difference Vegetation Index (NDVI), Version 5 data product (Vermote, 2019) for all case studies included, 42. Next, we began extracting Landsat data from Landsat 4, 5, and 8, courtesy of the U.S. Geological Survey, in search of higher resolution data. We selected a 3 by 3 Landsat pixel area, leading to a 90m resolution data extraction. The specific imagery used includes Landsat 4 USGS Landsat 4 TM Collection 1 Tier 1 TOA (top of atmosphere) Reflectance, Landsat 5 USGS Landsat 5 TM (thematic mapper) Collection 1 Tier 1 TOA Reflectance, and Landsat 8 USGS Landsat 8 Collection 1 Tier 1 TOA Reflectance. Within Google Earth Engine, we selected the date and location (latitude and longitude), calculated NDVI, kNDVI, and EVI utilizing Landsat bands (see metadata_NGEE-tropics_cyclones), and extracted post- and pre-cyclone values for each case study to calculate cyclone-induced change in the vegetative index. Due to limited spatial resolution of Landsat remote sensing data, MODIS products were investigated next. First, the MOD13Q1.006 Terra Vegetation Indices 16-Day Global 250m product was used to extract 250m EVI and NDVI (Didan, 2015) and then the MCD15A3H.006 MODIS Leaf Area Index/FPAR 4-Day Global 500m product product was used to extract LAI 500m (Myneni et al., 2015). Pre- and post-cyclone values, change in the vegetative index, and standard deviation for all values are included in the main csv (see case_study_data.csv) for all vegetative indices collected, including LAI 500m, EVI 250m, NDVI 250m, NDVI 90m, kNDVI 90m, EVI 90m, and NDVI 5km. Lastly, recovery values were calculated utilizing a standardization method (see metadata_NGEE-tropics_cyclones) and recovery values for MODIS (see MODIS_recovery.csv) and Landsat (Landsat_recovery.csv) data are included.

54 ENVIRONMENTAL SCIENCES↗

Estimating snow cover from high-resolution satellite imagery by thresholding blue wavelengths: Supporting Data

The extent and duration of snow cover is predicted to be altered as the climate changes. Developing high-resolution estimates of snow cover change is crucial for estimating changes in snow cover and the effects of these changes on watershed and ecosystems processes. Remote sensing tools have been a common method for rapidly mapping snow covered area (SCA) across a landscape. The most common remote sensing method for estimating SCA uses satellite-based calculations of the normalized difference snow index (NDSI), which relies on spectral measurements in the shortwave-infrared wavelengths (SWIR). NDSI is effective at catchment- to regional-scale estimates of SCA, but due to spatial resolution limitations of SWIR measurements, NDSI cannot be used to assess fine-scale SCA. In this work, we develop a new algorithm, called the Blue Snow Threshold (BST) algorithm, that maps high-resolution SCA by calculating a threshold on the blue wavelengths from high-resolution satellite imagery. This data package includes Orthorectified IKONOS-2 imagery (IkonosTestImage.tif) from August 14, 2004 at 1.00 meters Ground Sample Distance for Cook Inlet, Alaska (59.966414 , -152.982975). The Blue Snow Threshold algorithm (BST.py) was then used to produce a snow cover estimate (IkonosTestImage_BST.tif) for this study area. Additional imagery metadata is included in the ImageInfo.txt file. See Thaler et al., 2023 (https://doi.org/10.1016/j.rse.2022.113403) for more information about the BST algorithm.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a 15-year research effort (2012-2027) to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Electrical Resistivity Tomography data from 2016 to 2018 at the Lower Montane site in the East River Watershed, Colorado

This dataset contains time-lapse Electrical Resistivity Tomography (ERT) data along a transect located on the northeast-facing hillslope at the lower montane site (Pumphouse site) in the upper East River Watershed. The monitoring dataset covers the period from November 2, 2016, to August 6, 2018. In addition, the archive also contains a baseline dataset from October 9, 2016. The ERT transect consisted of 128 electrodes with an electrode spacing of 1.25 m. The acquisition system was located in the middle of the transect, about 50 m on one side, and included an MPT (Multi-Phase Technologies) ERT system, a mini computer, and batteries with solar panels. Acquisition occurred daily under normal circumstances. The first 16 electrodes (from the upper end of the transect) could not be used after the cable was damaged during the 2017–2018 winter. Also, due to multiple failures in the power system, the temporal resolution of the data is much lower in 2018 compared to 2016 and 2017. The data have been processed and used in Dafflon et al., 2023, and the baseline dataset was used in Falco et al., 2019 (see reference list). This archive contains the measurements (ER.zip containing csv files) for each of the 326 acquisition times and a filtered version where only electrodes 17 to 128 are included (ERT_sm.zip containing csv files). The archive also contains the baseline dataset and two acquisitions with full reciprocals (ERT_RB.zip containing csv files), as well as all the raw MPT files (ERT_raw_MTP.zip). The geometry (electrode position and elevation) is provided in Universal Transverse Mercator (UTM) 13N Geoid2012AB in the file named ERT_Location.csv. The archive contains 1 *.csv data files, four *.zip files, and three metadata *.csv files. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Remote Sensing Time Series Product Tool

The TSPT (Time Series Product Tool) software was custom-designed for NASA to rapidly create and display single-band and band-combination time series, such as NDVI (Normalized Difference Vegetation Index) images, for wide-area crop surveillance and for other time-critical applications. The TSPT, developed in MATLAB, allows users to create and display various MODIS (Moderate Resolution Imaging Spectroradiometer) or simulated VIIRS (Visible/Infrared Imager Radiometer Suite) products as single images, as time series plots at a selected location, or as temporally processed image videos. Manually creating these types of products is extremely labor intensive; however, the TSPT development tool makes the process simplified and efficient. MODIS is ideal for monitoring large crop areas because of its wide swath (2330 km), its relatively small ground sample distance (250 m), and its high temporal revisit time (twice daily). Furthermore, because MODIS imagery is acquired daily, rapid changes in vegetative health can potentially be detected. The new TSPT technology provides users with the ability to temporally process high-revisit-rate satellite imagery, such as that acquired from MODIS and from its successor, the VIIRS. The TSPT features the important capability of fusing data from both MODIS instruments onboard the Terra and Aqua satellites, which drastically improves cloud statistics. With the TSPT, MODIS metadata is used to find and optionally remove bad and suspect data. Noise removal and temporal processing techniques allow users to create low-noise time series plots and image videos and to select settings and thresholds that tailor particular output products. The TSPT GUI (graphical user interface) provides an interactive environment for crafting what-if scenarios by enabling a user to repeat product generation using different settings and thresholds. The TSPT Application Programming Interface provides more fine-tuned control of product generation, allowing experienced programmers to bypass the GUI and to create more user-specific output products, such as comparison time plots or images. This type of time series analysis tool for remotely sensed imagery could be the basis of a large-area vegetation surveillance system. The TSPT has been used to generate NDVI time series over growing seasons in California and Argentina and for hurricane events, such as Hurricane Katrina.

Predos, Don↗

Blodgett 13C–labeled litter incubation 2016-2019

The dataset is from 13C-labelled (stable isotope of carbon) root-litter in-situ field incubation experiment based on the whole-soil warming experiment at the Blodgett Forest Research Station, CA, USA. The files are in both ".csv" and ".xlsx" versions, and can be opened in "maCOS numbers", and "Microsoft Excel". The files includes several sheets with all the data published in the paper: Sun, B., Zosso, C., Wiesenberg, G. L. B., Pegoraro, E., Torn, M. S., and Schmidt, M. W. I.: Warming accelerates the decomposition of root-derived hydrolysable lipids in a temperate forest and is depth- and compound class-dependent, SOIL, 11, 1077–1093, https://doi.org/10.5194/soil-11-1077-2025, 2025. This dataset includes bulk soil carbon, nitrogen, delta 13C values, the normalized concentration (to organic carbon) of hydrolysable lipids identified, the absolute concentration (normalized to bulk soil) of hydrolysable lipids, hydrolysable lipids recovery, and weighted 13C-excess of bulk soil carbon, and weighted 13C-excess of each compound class in hydrolysable lipids. These data aim to answer two research questions: 1) How will warming affect the decomposition of 13C-labelled root-litter at different depth? 2) Will the decomposition of root-derived hydrolysable lipids under warming differ among different compound classes? The experiment sites located on the foothills of the Sierra Nevada near Georgetown, CA (120°3904000W; 38°5404300 N) at 1370m above see level. The Blodgett Forest is a mixed-coniferous forest. The site has a Mediterranean climate with a mean annual air temperature of 12.5 °C and a mean annual precipitation of 1774mm.

54 ENVIRONMENTAL SCIENCES↗

NLR HPC Kestrel Jobs Data

Overview: Anonymized job-level records from the Kestrel HPC system at the National Laboratory of the Rockies (NLR). Each record represents a Slurm batch job with scheduling metadata, resource requests, utilization, energy estimates, and efficiency metrics. Sensitive fields (user, account, job name, submit line, working directory, submit script, and job type) are replaced with 7-character cryptographic hashes. System & Timeframe: Kestrel is located at the NLR campus. Standard compute nodes have 104 cores and 256 GB RAM; bigmem nodes have 2,000 GB. GPU nodes (gpu-h100 partition) use NVIDIA H100 GPUs. Data covers jobs submitted August 2023 through December 2025. Funding provided by the U.S. Department of Energy, EERE. Files: esif.hpc.kestrel.job-anon.zip — Anonymized job records (Hive-partitioned Parquet) datacard.md — Full dataset documentation ~11 million rows, 50 variables. Readable with PyArrow, pandas, DuckDB, Apache Spark, or any Parquet-compatible tool. Data Collection: Jobs collected via sacct with timezone-aware export (SLURM_TIME_FORMAT="%Y-%m-%dT%H:%M:%S%z"), loaded into PostgreSQL. Calculated columns updated via database triggers and batch functions. All timestamps use timestamptz and correctly handle DST transitions. Preprocessing: Anonymization of name, user, account, submit_line, work_dir, submit_script, and job_type via 7-char hex hashes Derived columns: queue_wait, cpu_eff, max/min/avg_mem_eff, energy estimates Simplified job state mapping (e.g., "CANCELLED by 132357" → "CANCELLED") Boolean flags: python_job, reframe_job Temporal decomposition: year, month, day, day_of_week, hour, minute from submit_time Shared node tracking: shared_job_count, nodes_shared, jobs_shared Key Variables: Scheduling: job_id, partition, state_simple, submit_time, start_time, end_time, queue_wait Resources: nodes_req/used, processors_req/used, memory_req, wallclock_req/used, gpus_requested Efficiency: cpu_eff, max/min/avg_mem_eff Energy: cpu_energy_tdp_estimated_max/used_watt_hours, consumed_energy_raw_joules, consumed_energy_raw_watt_hours Sharing: shared_job_count, nodes_shared, jobs_shared Partitions: short, standard, debug, gpu-h100 Job States: CANCELLED, COMPLETED, FAILED, PENDING, RUNNING QoS Levels: normal, high Important Notes: Timestamps include timezone offsets; DST transitions are handled correctly, though adding intervals across DST boundaries requires offset adjustment shared_job_count reflects physical node co-residency, not use of the shared partition Job step records and raw Slurm JSONB fields are excluded Do not attempt to re-identify individuals from hashed fields

97 MATHEMATICS AND COMPUTING↗