Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Packages”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Data and scripts associated with “When do Riverine Systems 'Feel the Burn'? Simulating How Burn Extent and Severity Modulate Hydrologic Controls on Biogeochemical Export” (v2)

This data package is associated with the publication “When do Riverine Systems 'Feel the Burn'? Simulating How Burn Extent and Severity Modulate Hydrologic Controls on Biogeochemical Export” published in Water Resources Research (Wampler et al. 2025; preprint: https://doi.org/10.22541/essoar.174438106.63564767/v1). This study used the Soil and Water Assessment Tool (SWAT), a processed based model to explore the impacts of area burned and burn severity on streamflow, nitrate, and dissolved organic carbon (DOC) in two test basins: a semi-arid, mixed land use basin and a humid, primarily forested basin. We developed 1800 wildfire scenarios that we ran in each basin: 20 different burn extents (5 to 100% by 5%), 3 different burn severities (low, moderate, and high), and 30 different post-fire precipitation scenarios. We also ran an additional 30 scenarios associated with no wildfire for the 30 post-fire precipitation scenarios. For each scenario we were interested in the change in runoff ratio (streamflow) and average concentration and annual loads (nitrate and DOC) across the wildfire scenarios. This data package contains the data and scripts required to build SWAT models for the two test basins, create and run the wildfire scenarios, and generate the data summaries and figures used in the associated manuscript. This data package was originally published in March 2025. It was updated in January 2026 (v2; new and modified files) to include the final files after the manuscript went through reviews. See the change history section below for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.

54 ENVIRONMENTAL SCIENCES↗

Model associated with: "Thermodynamic control on the decomposition of organic matter across different electron acceptors"

This model data package is associated with the publication “Thermodynamic control on the decomposition of organic matter across different electron acceptors” submitted to Soil Biology and Biochemistry (Zheng et al., 2023; https://doi.org/10.1016/j.soilbio.2024.109364).In this research, a thermodynamic modeling framework is built to flexibly incorporate both organic matter (OM) molecules and electron acceptors for estimating potential free energy release from various redox reactions and to further predict reaction rates based on Microbial Transition State Theory. The model package includes scripts for thermodynamic modeling and postprocessing. Input Fourier-transform ion cyclotron resonance (FTICR) data are from a previous experimental study (Boye et al., 2018), and model outputs are free energy predictions and stoichiometric coefficients associated with all possible redox reactions.This data package is associated with the project GitHub repository found at MM_bioenergetic_modeling.This data package contains four folders (Input_FTICR, Model, Output, and Output_processing), a file-level metadata (FLMD) csv, and a data dictionary (dd) csv. Please see Zheng_bioenergetic_modeling_flmd.csv for a list of all files contained in this data package and descriptions for each. The Zheng_bioenergetic_modeling_dd.csv file describes the csv column headers. The “Model” folder contains scripts to run energy balance calculations for each electron acceptor. The “Output” folder contains csv files with stoichiometric information from model simulations. And the "Output_processing" folder contains scripts for reaction rate calculations and to generate plots.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts from: “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”

This data package includes data and scripts from the manuscript “Denoising autoencoder for reconstructing sensor observation data and predicting evapotranspiration: noisy and missing values repair and uncertainty quantification”.The study addressed common challenges faced in environmental sensing and modeling, including uncertain input data, missing sensor observations, and high-dimensional datasets with interrelated but redundant variables. Point-scaled meteorological and soil sensor observations were perturbed with noises and missing values, and denoising autoencoder (DAE) neural networks were developed to reconstruct the perturbed data and further predict evapotranspiration. This study concluded that (1) the reconstruction quality of each variable depends on its cross-correlation and alignment to the underlying data structure, (2) uncertainties from the models were overall stronger than those from the data corruption, and (3) there was a tradeoff between reducing bias and reducing variance when evaluating the uncertainty of the machine learning models.This package includes:(1) Four ipython scripts (.ipynb): “DAE_train.ipynb” trains and evaluates DAE neural networks, “DAE_predict.ipynb” makes predictions from the trained DAE models, “ET_train.ipynb” trains and evaluates ET prediction neural networks, and “ET_predict.ipynb” makes predictions from trained ET models.(2) One python file (.py): “methods.py” includes all user-defined functions and python codes used in the ipython scripts.(3) A “sub_models” folder that includes five trained DAE neural networks (in pytorch format, .pt), which could be used to ingest input data before being fed to the downstream ET models in ‘ET_train.ipynb” or ‘ET_predict.ipynb’.(4) Two data files (.csv). Daily meteorological, vegetation, and soil data is in “df_data.csv”, where “df_meta.csv” contains the location and time information of “df_data.csv”. Each row (index) in “df_meta.csv” corresponds to each row in “df_data.csv”. These data files are formatted to follow the data structure requirements and be directly used in the ipython scripts, and they have been shuffled chronologically to train machine learning models. The meteorological and soil data was collected using point sensors between 2019-2023 at(4.a) Three shrub-dominated field sites in East River, Colorado (named “ph1”, “ph2” and “sg5” in “df_meta.csv”, where “ph1” and “ph2” were located at PumpHouse Hillslopes, and “sg5” was at Snodgrass Mountain meadow) and(4.b) One outdoor, mesoscale, and herbaceous-dominated experiment in Berkeley, California (named “tb” in “df_meta.csv”, short for Smartsoils Testbed at Lawrence Berkeley National Lab).- See "df_data_dd.csv" and "df_meta_dd.csv" for variable descriptions and the Methods section for additional data processing steps. See "flmd.csv" and "README.txt" for brief file descriptions.- All ipython scripts and python files are written in and require PYTHON language software.

54 ENVIRONMENTAL SCIENCES↗

Organic Matter Composition in June 2023 and September 2023 Across the McKenzie Sub-Basin Impacted by the 2020 Holiday Farm Fire

This dataset represents results from a field study aiming to understand the variability in post-fire responses of dissolved organic matter and determine drivers of post-fire responses. Samples were collected at 58 sites within the McKenzie River Watershed (Oregon, USA) that were upstream, within, and downstream of the Holiday Farm Fire burn perimeter. The samples were collected in June 2023 and September 2023 during storm events, approximately 3 years post-fire. Samples were characterized for benezenepolycarboxylic acids (BPCA) and ultra-high resolution mass spectrometry. Dissolved organic carbon and optics (absorbance and fluorescence) data can be found in a separate data packages (https://ir.library.oregonstate.edu/concern/datasets/zc77sz60m, https://ir.library.oregonstate.edu/concern/datasets/mc87q034m). Related data from a subset of sites from 2020-2022 can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1869708 and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2478546. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset contains (1) file-level metadata; (2) data dictionary; (3) data package readme; (4) metadata; (5) methods information; (6) benzene polycarboxylic acid (BPCA) concentration data; (7) Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) methods; (8) folder of high resolution characterization of organic matter via 12 Tesla FTICR-MS data generated through the Environmental Molecular Sciences Laboratory (EMSL; https://www.pnnl.gov/environmental-molecular-sciences-laboratory). This package contains the following file types: csv, xml, pdf.

54 ENVIRONMENTAL SCIENCES↗

Data and Scripts associated with a manuscript on ecosystem responses to wildfires in the Columbia River Basin

This data package is associated with the publication “Ecosystem leaf area, gross primary production, and evapotranspiration responses to wildfire in the Columbia River Basin” submitted to Biogeosciences (Shi et al., 2024; doi: 10.22541/au.171053013.30286044/v1). In this research, data products, leaf area index (LAI), gross primary production (GPP), and evapotranspiration (ET), from the Moderate Resolution Imaging Spectroradiometer (MODIS) are used to quantify the resistance and resilience of different ecosystem types in the Columbia River Basin (CRB). A machine learning algorithm, random forest (RF), was used to examine the impacts of precipitation, vapor pressure deficit (VPD), and burn severity from Monitoring Trends in Burn Severity (MTBS) on ecosystem resilience. The data package includes the processed MODIS data products, precipitation, VPD, and burn severity in 138 fire regions in CRB and the input files for RF model training. This data package includes six folders. The MODIS products are included in three MODIS_* folders with shell scripts for data clipping and *ncl files for data processing: (1) “/MODIS_LAI_CRB”; (2) “/MODIS_GPP_CRB”; and (3) “/MODIS_ET_CRB”. All the processed data for each fire event are NetCDF formatted. The MTBS burn severity data and the shell and *ncl scripts used for data processing are in the folder named (4) “MTBS_fire”. The ERA meteorological fields and the data processing scritps are in (5) “ERA_Var_CR”. All the scripts for figure development are in the format of *ncl and in the folder (6) “paper_scripts”. See the file ending in “flmd.csv” for a list of all files contained in this data package and descriptions for each. Tabular column headers and units are described in the data dictionary file ending in “dd.csv”.

54 ENVIRONMENTAL SCIENCES↗

Pre-dawn leaf water potential, San Lorenzo, Panama, 2020

This data package contains predawn leaf water potential (LWP) data for leaves sampled in the San Lorenzo forest canopy crane site in Panama (PA-SLZ) from January to March 2020. Data were collected from 31 species, from top of canopy and vertical profiles within the canopy. All data and metadata are presented in .csv files. The protocol details are provided as *.pdf. See related data packages from the BNL 2020 field campaign for leaf optical properties and gas exchange measurements. Sample information including canopy elevation and leaf area index (LAI) can be found in the related “Leaf and canopy traits” data package.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with “Riverine dissolved organic matter transformations increase with watershed area, water residence time, and Damköhler numbers in nested watersheds” (v2)

This data package is associated with the publication “Riverine dissolved organic matter transformations increase with watershed area, water residence time, and Damköhler numbers in nested watersheds” submitted to Biogeochemistry by Ryan et al., 2024 (DOI: https://doi.org/10.1007/s10533-024-01169-5). This study aims to investigate fundamental and transferable drivers of dissolved organic matter (DOM) diversity across five nested watersheds within the contiguous United States. DOM diversity was explored using ultrahigh-resolution Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS). The samples and the unprocessed FTICR-MS data used in this study are publicly available on the Environmental System Science Data Infrastructure for a Virtual Ecosystem (ESS-DIVE) data repository (see DOIs below). The data for the Willamette, Gunnison, Connecticut, and Deschutes basins were collected as part of a collaboration between the Watershed Rules of Life (WROL) project and Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS). The data for the Yakima River basin (YRB) was collected by the PNNL River Corridor SFA. The raw, unprocessed FTICR-MS data with additional (meta)data can be found at doi:10.15485/1895159 for WROL samples and doi:10.15485/1898912 for YRB samples. This data package contains the processed data used in the associated manuscript. This package also contains ancillary geospatial, hydrological, and geochemical information that supports the interpretation of the FTICR-MS data within Ryan et al., 2024. This data package is associated with the GitHub repository found at https://github.com/WHONDRS-Hub/rcsfa-RC4-WROL-YRB_DOM_Diversity. This data package was originally published August 2024. It was updated January 2025 (modified files). See the change history in the readme more details. At the directory level, the data package is comprised of three folders: (1) data, (2) output, and (3) src; and five additional files including the data dictionary (file ending in "_dd.csv”) and file-level metadata (file ending in “_flmd.csv”). The “src” folder contains the scripts used to process the FTICR data, conduct the analyses, and produce the manuscript figures. The inputs for these scripts are in the “data” folder and the returned outputs in the “output” folder. Inputs include temporal and spatial metadata associated with the sampling efforts, processed FTICR data, and total and normalized putative biochemical transformations per sample. Outputs include cleaned and combined data presented as tables, descriptive statistics, and plots. The file-level metadata file lists all files contained in this data package and descriptions for each. The data dictionary describes the units and definitions for each tabular data column or row header.

54 ENVIRONMENTAL SCIENCES↗

Steptoe Valley NV Data Compilation: Understanding a Stratigraphic Hydrothermal Resource through Geophysical Imaging

Sandia National Laboratories partnered with a multi-disciplinary group of subject matter experts to evaluate a stratigraphic geothermal resource in Steptoe Valley, Nevada using both established and novel geophysical imaging techniques. Provided here are a compilation of newly acquired data over the area and select modeling efforts. This encompasses a 3D geological model (inclusive of full Leapfrog files, Leapfrog viewer files, and XYZ data for faults and stratigraphy) with embedded geophysical modeling, controlled-source electromagnetic (CSEM) and magnetotelluric (MT) data packages, aqueous spring geochemistry data, seismic reflection interpretations, and a gravity data package. The stratigraphic reservoir in Steptoe Valley was previously discovered during oil and gas exploration. Subsequent studies, such as the Nevada Play Fairway Analysis, added data which further highlighted potential resource targets in the basin. Geophysical surveys, complimented with refined geologic mapping and geochemical sampling, were deployed to further characterize the resource. The resulting 3D geologic interpretation, conceptual model refinements, and reservoir simulations suggest that a power-capable reservoir is economically accessible in the Paleozoic carbonates of the deep/central basin. Additional geophysical characterization and exploration drilling efforts are recommended to calibrate interpretation and determine where/how to potentially develop the Steptoe resource. The geophysical tools, interpretations, lessons learned, and publicly available data generated by this study establish an exploration methodology to inform decisions for successful development of stratigraphic reservoirs.

15 GEOTHERMAL ENERGY↗

Models, data, and scripts associated with “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning”

This data package is associated with the publication “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning’’ submitted to the Journal of Geophysical Research: Machine Learning and Computation (Scheibe et al. 2024). River sediment respiration observations are expensive and labor intensive to obtain and there is no physical model for predicting this quantity. The Worldwide Hydrobiogeochemisty Observation Network for Dynamic River Systems (WHONDRS) observational data set (Goldman et al.; 2020) is used to train machine learning (ML) models to predict respiration rates at unsampled sites. This repository archives training data, ML models, predictions, and model evaluation results for the purposes of reproducibility of the results in the associated manuscript and community reuse of the ML models trained in this project. One of the key challenges in this work was to find an optimum configuration for machine learning models to work with this feature-rich (i.e. 100+ possible input variables) data set. Here, we used a two-tiered approach to managing the analysis of this complex data set: 1) a stacked ensemble of ML models that can automatically optimize hyperparameters to accelerate the process of model selection and tuning and 2) feature permutation importance to iteratively select the most important features (i.e. inputs) to the ML models. The major elements of this ML workflow are modular, portable, open, and cloud-based, thus making this implementation a potential template for other applications. This data package is associated with the GitHub repository found at Please see the file level metadata (flmd; “sl-archive-whondrs_flmd.csv”) for a list of all files contained in this data package and descriptions for each. Please see the data dictionary (dd; “sl-archive-whondrs_dd.csv”) for a list of all column headers contained within comma separated value (csv) files in this data package and descriptions for each. The GitHub repository is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning models trained on the data in “input_data”; (3) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; (4) “examples” contains the visualization of the results in this repository including plotting scripts for the manuscript (e.g., model evaluation, FPI results) and scripts for running predictions with the ML models (i.e., reusing the trained ML models); (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. Furthermore, depending on the number of features used to train the ML models, the preprocessing and postprocessing scripts, and their intermediate results, can also be different branch-to-branch. The “main-*” branches are meant to be starting points (i.e. trunks) for each model branch (i.e. sprouts). Please see the Branch Navigation section in the top-level README.md in the GitHub repository for more details. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please the top-level README.md in the GitHub repository for more details on the automation.

13C↗

Machining of Alloy 709 Creep-fatigue Specimens from G. O. Carlson Heat

Alloy 709 has been selected as the next candidate material for Section III, Division 5 qualification in the American Society of Mechanical Engineers (ASME) Boiler and Pressure Vessel Code (BPVC) for elevated-temperature nuclear construction. The qualification data package requires an assortment of information and material data including tensile, creep, and creep-fatigue performance at elevated temperatures from different material heats. The goal of the project is to generate a dataset to support qualification of Alloy 709 material for ASME BPVC design code. Working towards the project goal, Idaho National Laboratory needs to conduct a series of tests on three different heats to support the A709 code case development. At present, there is a gap in the data package for one of the heats: Heat number 58776 manufactured by G.O. Carlson. To address this data gap, Argonne National Laboratory transmitted five plates of A709 heat 58776 manufactured by G.O. Carlson to Idaho National Laboratory. These plates were solution annealed at 1150°C and heat treated at 775°C for 10 hours. The objective of this specification is to procure a series of creep-fatigue specimens to support the qualification data package. The creep-fatigue specimen design captures the cyclic material performance at elevated temperature. The material performance data generated from specimens machined herein will support the ASME BPVC code case development and establish design limits, and design life curves for Alloy 709 material.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Surface water and groundwater FTICR-MS, NPOC, and TN from nine wetlands and three upland wells at the Tanglewood Biological Station, Alabama

This dataset supports a broader study examining wetland hydrobiogeochemical responses to flood disturbance and the subsequent impacts on watershed nutrient export. The study was designed following ICON (integrated, coordinated, open, and networked) principles. Samples were collected from nine wetlands and three upland wells at the Tanglewood Biological Station, Alabama in August 2024 and February 2025, during the dry and wet season, respectively. The contents include geochemistry (dissolved organic carbon measured as non-purgeable organic carbon; total dissolved nitrogen) and organic matter characterization (FTICR-MS). Related water level data from the same locations can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/2530253. Additional geochemistry will be published in a separate data package. For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions.This dataset is comprised of (1) a folder containing environmental context photos; (2) file-level metadata; (3) data dictionary; (4) field metadata; (5) readme; (6) international generic sample number (IGSN) mapping file; (7) the field protocol; and (8) a subfolder with sample data. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) total nitrogen data and averages; (3) methods codes; and (4) a subfolder of 12 Tesla (12T) Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) data. All files are .csv, .pdf, .jpeg, or .jpg.

54 ENVIRONMENTAL SCIENCES↗

Model scripts associated with “Revisiting controls on hyporheic respiration with knowledge-guided machine learning at continental scale”

NOTE: The manuscript associated with this data package is currently in review. The data/scripts may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final scripts and additional metadata. This data package is associated with the publication “Revisiting controls on hyporheic respiration with knowledge-guided machine learning at continental scale” submitted to Environmental Science & Technology (Zheng et al. 2026). The project combines mechanistic process modeling with knowledge-guided machine learning (KGML) to evaluate how organic matter chemistry, microbial biomass, and physical substrate accessibility regulate realized respiration rates across river corridors. All data used in this paper have been previously published and can be accessed at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719 (Goldman et al., 2020). This data package contains 3 R-markdown (Rmd) preprocessing scripts for the previously published data and subsequent modelling workflows. The full workflow with input and output data can be found in the associated GitHub repository at https://github.com/jianqiuz/KGML-WHONDRS.

Biogeochemistry↗

Geospatial Information, Metadata, and Maps for Global River Corridor Science Focus Area Sites (v5)

This dataset provides geospatial information, metadata, and maps for the Pacific Northwest National Laboratory (PNNL) River Corridor Science Focus Area (RC-SFA; https://www.pnnl.gov/projects/river-corridor) sites. The RC-SFA works to transform understanding of spatial and temporal dynamics in river corridor hydrobiogeochemical functions from molecular reaction to watershed and basin scales. The knowledge we gain is used to formulate and test hypotheses and to improve mechanistic representation of river corridor processes and their response to disturbances in multiscale models of integrated hydrobiogeochemical function. The data provided includes Site ID, latitude, longitude, stream name, and common ID (COMID) for sites used across the RC-SFA. The COMID can be used to find and download data from NHDPlus (https://www.epa.gov/waterdata/nhdplus-national-hydrography-dataset-plus) and other platforms. The sites included are non-exhaustive. Sites (including past sites) will be added to this data package in the future. Data generated from the RC SFA can be accessed at https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA. This data package was originally published in April 2023. It was updated in June 2023 (v2; modified files), December 2023 (v3; modified files), January 2025 (v4; modified files), and December 2025 (v5; modified files). See the change history section in the readme for more details. This dataset is comprised of one main data folder. The data folder consists of (1) file-level metadata; (2) data dictionary; (3) readme; (4) methods codes; (5) geospatial information for all RC SFA sites including International Generic Sample Number (IGSN); (6) maps of all sites and sites in Washington State, USA; and (7) a subfolder with the shapefile of all sites. All files are .csv, .pdf, .shp, .cpg, .dbf, .prj, .qmd, or .shx. We thank the Confederated Tribes and Bands of the Yakama Nation for access to field locations where some data were collected in Washington state. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Data for “Fast-decaying plant litter enhances soil carbon in temperate forests, but not through microbial physiological traits”

This data package contains data and code used in the paper “Fast-decaying plant litter enhances soil carbon in temperate forests, but not through microbial physiological traits”. This paper details results from two studies: 1) a laboratory leaf litter incubation experiment (lab experiment) and 2) a multi-site observational field study (field study). Both studies were designed to test the relationships among litter quality, microbial physiological traits, and mineral-associated soil carbon (C). In the lab experiment, we incubated 16 temperate tree litters of differing chemical quality with isotopically distinct soil, measuring microbial physiological traits and the flow of litter-derived C into the mineral-associated soil C pool. In the field study, we sampled soils (0-5 cm) across six eastern US temperate forests, measuring microbial physiological traits, soil abiotic properties, and leaf litter chemistry. The analysis data are provided in two .csv files corresponding to either the lab experiment or field study. Both files contain data on litter chemistry, microbial physiological traits (growth and turnover rates, and carbon use efficiency [CUE]), and the mineral-associated soil C pool. The lab experiment dataset additionally contains litter-derived versus soil-derived soil C, respiration, and litter decomposition parameters. The field study dataset additionally contains site-level climatic information, ectomycorrhizal dominance of plots, and plot-level soil properties. Also provided is the R code and output reproducing the results in the related publication. Analyses were originally performed using R version 3.6.1 using the package "lavaan" and the packages listed on lines 5-7 in the file "data.R".

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with a manuscript on a meta-analysis synthesizing stream biogeochemical response to wildfires across space and time (v2)

This data package is associated with the publication “Catchment characteristics modulate the influence of wildfires on nitrate and dissolved organic carbon in lotic systems across space and time: A meta-analysis” submitted to Global Biogeochemical Cycles (Cavaiani et al. 2025). This study uses meta-analytical techniques to evaluate the effect of wildfire on in-stream responses in burned and unburned watersheds. The study aims to provide additional insight into the range of responses and net influences that wildfires have on hydro-biogeochemistry across broad spatial scales, burn extents, and the persistence of water-quality change. This study compiles data and metadata from 18 total publications that includes 1) surface water geochemistry data (dissolved organic carbon; nitrate), 2) climate classifications, 3) year of the wildfire, 4) the time lag between when the fire occurred and when the sampling occurred, and 5) study design of the publication. In total, this meta-analysis draws data that spans 8 climate guilds, 3 biomes, 62 watersheds, and 20 unique wildfires. See Sites_meta_data.csv for citations of the papers used in this meta-analysis. All R scripts and the associated data can also be found on GitHub at This data package was originally published in March 2024. It was updated in April 2025 (v2; new and modified files). See the change history section in the readme for more details. This data package contains five primary folders that include the following: (1) inputs; (2) output for analysis; (3) initial plots; (4) R scripts; and (5) GIS data. The data package also contains a data dictionary (dd) that provides column header definitions and a file-level metadata (flmd) file that describes every file. The “inputs” folder contains a list of all publications identified during the formal web search and an indication of whether each publication was included in the final analysis. Additionally, it includes site-level metadata, catchment characteristics, and GIS data for all publications included in the final analysis. The “Output_for_analysis” folder contains all data frames and figures generated from each R script used for additional data analysis. The “initial_plots” folder includes all exploratory figures that will be included in a supplemental and figures that will be submitted with the manuscript for publication. The “R_scripts” folder contains the scripts that perform all the data manipulations, statistical analyses, and plots. The “gis_data” folder includes shape files for each fire included in this meta-analysis. This data package contains the following file types: csv, pdf, jpeg, cpg, dbf, prj, shp, shp.ea.iso.xml, shp.iso.xml, shx.

54 ENVIRONMENTAL SCIENCES↗

Arctic shrub root traits, northern Alaska, summer 2017 (Version 2.0)

This data package contains root trait data collected from 170 plots of rapidly expanding shrub genera (Alnus, Betula, and Salix) and a widespread sedge (Eriophorum vaginatum) along a latitudinal and temperature gradient in northern Alaska. The trait data were collected in July 2017 and include root architecture (root diameter and branching patterns), mycorrhizal colonization (%), nitrogen concentration (%), delta 15N (per mil), and vertical root biomass. These raw data support a submitted manuscript that examines the distribution and interspecific variations of absorptive root traits of shrubs and graminoids across the graminoid-dominated nutrient-poor arctic tundra and reveals how deciduous shrub expansion affects plant nutrient acquisition strategies in tundra ecosystems. Data are presented by site (n=5) and patch (shrub or sedge plot) in separate csv files. The location data are provided in the “plot coordinates” file; all other files contain the data in the file title. Detailed methods are in Chen et al. (2020).2023/07/20 Update: The latest version of this dataset publication is version 2.0. The latest version of the data package was updated to include the alder nodule biomass dataset (alder nodule biomass.csv) and associated metadata (metadata_alder nodule biomass.csv). The name of the previous metadata file was updated (metadata_ root traits and biomass.csv ) to distinguish it from the new metadata file.

54 ENVIRONMENTAL SCIENCES↗

Data for "Relative Reactivity and Bioavailability of Mercury Sorbed to or Coprecipitated with Aged Iron Sulfides"

The potential for inorganic mercury (Hg) to be converted to methylmercury depends, in part, on the chemical form of Hg and its bioavailability to anaerobic microorganisms that can methylate Hg. In anaerobic settings, Hg can be associated with sulfide phases, including ferrous iron sulfide (FeS), which can sorb or coprecipitated with Hg. The objective of this study was to determine if the aging state of FeS alters the Hg coordination environment as well as the reactivity and bioavailability of sorbed and coprecipitated Hg species. FeS particles were synthesized with and without Hg2+ and aged in anaerobic conditions for multiple time frames spanning from 1 hour to 1 month. For FeS particles synthesized without Hg, Hg2+ was subsequently sorbed to the FeS for 1 day. Analysis of Hg speciation of these materials by X-ray absorption near edge spectroscopy revealed a predominance of 4-coodinate Hg-S species in the sorbed Hg-FeS solids and a mixture of 2- and 4-coordinate Hg-S in the coprecipitated Hg-FeS. The leaching potential of the Hg was assessed by exposing the particles to a solution of dissolved glutathione (a thiolate-based Hg chelator). As expected, the sorbed Hg-FeS released more soluble Hg compared to the co-precipitated Hg-FeS. However, when these particles were exposed to Desulfovibrio desulfuricans ND132 (a known Hg methylator), more Hg was methylated from the co-precipitated Hg-FeS than the sorbed Hg-FeS, consistent with expectations from the Hg-S coordination state and inconsistent with the selective leaching results. Overall, these results suggest that the bioavailability of particulate Hg cannot be easily discerned by leaching potential into bulk solution. Rather, bioavailability entails more subtle interactions at particle-cell interfaces and perhaps correlates with the local Hg-S coordination state in the particles. This data package contains data shown in the referenced publication, including all figures. Each figure is available in the respective *.csv file and a summary of all files are also included in the the xlsx file. Additional formatted data include X-ray diffraction data of the FeS solids are also available in a HPF format (Panalytical software), X-ray absorption spectroscopy data as *.prj files (Athena), and X-ray fluorescence data in its raw data form.

54 ENVIRONMENTAL SCIENCES↗

Stochastic simulations of temporally and spatially variable Fe cycling within a floodplain aquifer

This data package contains model input and output files for local and global sensitivity analysis (SA) of a reactive transport model simulating redox cycling within a floodplain aquifer. The reaction network is implemented in CrunchFlow and focuses on spatially and temporally heterogeneous Fe cycling, though it contains several other species and redox pathways (37 aqueous species and 7 minerals in total). The package contains data from 600 Monte Carlo simulations (for global SA) and 104 one-at-a-time simulations (for local SA). It includes all input files required to run the CrunchFlow model, as well as limited model output – .csv files containing time series of primary species concentrations at 21 grid cells within the domain. These time series files contain the following species: H+, CO32-, SO42-, Cl-, Ca2+, Mg2+, Na+, K+, Fe(II), Fe(III), HS-, CH3COO-, O2(aq), S(aq), and Br-.These data were used to demonstrate the application of a novel form of global SA, distance-based generalized sensitivity analysis (DGSA) to a reactive transport model. The 2D model simulates the export of reduced species from pockets of fine-grained, organic-rich sediments embedded in a coarse sand aquifer. Stochastic simulations jointly varied 17 key model input parameters across several orders of magnitude, including the spatial variability of each parameter. DGSA on the model results reveal that the amplitude and variability of exported Fe(II) are most sensitive to the interaction between sand permeability and individual reaction rates. By contrast, the propagation of reducing conditions downgradient from the fine-grained lenses depends most heavily on the rate of dissolved organic carbon and sulfate release from the fines.

54 ENVIRONMENTAL SCIENCES↗