Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Reporting Format”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Model Inputs, Outputs, and Scripts associated with: “Spatial microbial respiration variations in the hyporheic zones within the Columbia River Basin”

This data package is associated with the publication “Spatial microbial respiration variations in the hyporheic zones within the Columbia River Basin” published in the Journal of Geophysical Research: Biogeosciences (Son et al. 2022) available at doi: 10.1029/2021JG006654. This data package includes the key model inputs/outputs of the river corridor model for the Columbia River Basin (CRB) and the model source codes, which were used in the manuscript. The model is a carbon-nitrogen-coupled river corridor model (RCM), and the model is used to quantify hyporheic zone (HZ) aerobic and anaerobic respiration at the NHDPLUS stream reach scales. The RCM used in this study combines empirical substrate models derived from observations and three microbially driven reactions to compute respiration of the HZ for each National Hydrography Dataset (NHD) reach within the CRB. The reactions in HZs of each NHD reach include anaerobic respiration and two-step anaerobic respiration via denitrification. Our HZ respiration estimates are limited to the lotic (or flowing) stream/river systems, and do not account for the respiration process in water column. Note that the RCM only simulates the HZ’s contribution to the dissolved carbon dioxide (CO2) concentrations in the streams, and the CO2 emissions to the atmosphere are not modelled. The model computes at hourly timesteps because of the fast reaction rates. The key input data of the model are exchange flux, residence time, and stream solute (dissolved organic carbon (DOC), dissolved oxygen (DO), and nitrate concentrations). These inputs are constant over time and represent long-term averaged values.This modeling framework successfully quantified HZ respiration components over multiple scales. It revealed key mechanisms driving the spatial variation of HZ aerobic and anaerobic respiration in reaches with varying hydrologic and substrate conditions. Thus, this modeling study offers a testing hypothesis in different river system (e.g., climate and biomes) for the HZ respiration processes, and can be used as a sampling design tool for large-scale HZ experimental studies.This dataset contains five folders: (1) model_inputs, (2) model_outputs, (3) Rscripts, (4) figures, and (5) model_codes. It also contains a readme, file level metadata (FLMD), and data dictionary (dd). Please see the FLMD for a list of all the files contained in this data package and descriptions for each. The model_inputs folder contains the model inputs used to drive the model simulations. The model_outputs folder contains key model output files from the river corridor model. The Rscripts folder contains the Rscripts for pre- and post- processing model results. The figures folder contains the raw figures associated with the manuscript. The model_codes folder includes key model source codes/input files. All files are .jpg, .jpeg, .out, .e, .od, .dat, .sub, .F90, .0, .R, .sbx, .cpg, .sbn, .shx, .shp, .dbf, .prj, .tfw, .tif, .xml, .pdf, or .csv.

54 ENVIRONMENTAL SCIENCES↗

Electrical Resistivity Tomography data from 2016 to 2018 at the Lower Montane site in the East River Watershed, Colorado

This dataset contains time-lapse Electrical Resistivity Tomography (ERT) data along a transect located on the northeast-facing hillslope at the lower montane site (Pumphouse site) in the upper East River Watershed. The monitoring dataset covers the period from November 2, 2016, to August 6, 2018. In addition, the archive also contains a baseline dataset from October 9, 2016. The ERT transect consisted of 128 electrodes with an electrode spacing of 1.25 m. The acquisition system was located in the middle of the transect, about 50 m on one side, and included an MPT (Multi-Phase Technologies) ERT system, a mini computer, and batteries with solar panels. Acquisition occurred daily under normal circumstances. The first 16 electrodes (from the upper end of the transect) could not be used after the cable was damaged during the 2017–2018 winter. Also, due to multiple failures in the power system, the temporal resolution of the data is much lower in 2018 compared to 2016 and 2017. The data have been processed and used in Dafflon et al., 2023, and the baseline dataset was used in Falco et al., 2019 (see reference list). This archive contains the measurements (ER.zip containing csv files) for each of the 326 acquisition times and a filtered version where only electrodes 17 to 128 are included (ERT_sm.zip containing csv files). The archive also contains the baseline dataset and two acquisitions with full reciprocals (ERT_RB.zip containing csv files), as well as all the raw MPT files (ERT_raw_MTP.zip). The geometry (electrode position and elevation) is provided in Universal Transverse Mercator (UTM) 13N Geoid2012AB in the file named ERT_Location.csv. The archive contains 1 *.csv data files, four *.zip files, and three metadata *.csv files. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Model Inputs, Outputs, and Scripts associated with: “Combined effects of stream hydrology and land use on basin-scale hyporheic zone denitrification in the Columbia River Basin”

This data package is associated with the publication “Combined effects of stream hydrology and land use on basin‐scale hyporheic zone denitrification in the Columbia River Basin”, published in Water Resource Research (Son et al.2022) available at https://doi.org/10.1029/2021WR031131. This data package includes the key model inputs/outputs of the river corridor model for the Columbia River Basin (CRB) and the model source codes used in the manuscript. The model is a carbon-nitrogen-coupled river corridor model (RCM), and the model is used to quantify hyporheic zone (HZ) denitrification at the NHDPLUS stream reach scales. The RCM used in this study combines empirical substrate models derived from observations and three microbially driven reactions, including two-step denitrification and aerobic respiration, are considered within the HZ. The key input data of the model are exchange flux, residence time, and stream solute (dissolved organic carbon (DOC), dissolved oxygen (DO), and nitrate concentrations). These inputs are constant over time and represent long-term averaged values. This study uses the RCM to explore the spatial patterns of HZ denitrification across reaches with different sizes and land use in the CRB. Our main objective is to use the RCM as a virtual reality model, and the machine-learning models as surrogates that encapsulate the complexities of the physics-based model while identifying the importance of different variables that are not evident in the model conceptualization. We do not include a direct comparison of the modeled HZ denitrification and measurements; however, the RCM can capture the overall spatial patterns of the HZ denitrification because the model inputs and its reaction networks are based on well-established theory and a physical-based model. The combination of the model-based predictions and a machine-learning approach (e.g., random forest) is used to improve our understanding of what variables of the model are associated with spatial patterns of the modeled denitrification across reaches with different sizes and land uses, and to develop a proxy model using measurable variables to reproduce the simulated patterns.This dataset contains five folders: (1) model_inputs, (2) model_outputs, (3) Rscripts, (4) figures, and (5) model_codes. It also contains a readme, file level metadata (FLMD), and data dictionary (dd). Please see the FLMD for a list of all the files contained in this data package and descriptions for each. The model_inputs folder contains the model inputs used to drive the model simulations. The model_outputs folder contains key model output files from the river corridor model. The Rscripts folder contains the Rscripts for pre- and post- processing model results. The figures folder contains the raw figures associated with the manuscript. The model_codes folder includes key model source codes/input files. All files are .jpg, .jpeg, .out, .e, .od, .dat, .sub, .F90, .0, .R, .sbx, .cpg, .sbn, .shx, .shp, .dbf, .prj, .tfw, .tif, .xml, .pdf, or .csv.

54 ENVIRONMENTAL SCIENCES↗

Data associated with “Different methods of estimating riverbed sediment grain size diverge at the basin scale ” (v2)

This data package is associated with the publication “Different methods of estimating riverbed sediment grain size diverge at the basin scale” published in Frontiers in Earth Science (Regier et al., 2025). The distribution of sediment grain size in streams and rivers is often quantified by the median grain size (d50), a key metric for understanding and predicting hydrologic and biogeochemical function of streams and rivers. Manual methods to measure d50 are time-consuming and ignore larger grains, while model-based methods to estimate d50 often over-generalize basin characteristics, and therefore cannot accurately represent site-scale heterogeneity. Here, we apply a machine learning-enabled photogrammetry methodology (You Only Look Once, or YOLO) for estimating d50 for grains > 2 mm based on images collected from streams and rivers throughout the Yakima River Basin (YRB). To understand how such methods may help bridge the gaps in resolution and accuracy between manual and catchment characteristics model-based d50 estimates, we compared YOLO d50 values to manual and model-based estimates across the YRB. We found distinct differences among methods for d50 averages and variability, and relationships between d50 estimates and basin characteristics. Source images can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/1892052. This data package was originally published in May 2023. It was updated August 2025 (v2; new and modified files). File and folder names were not revised to indicate changes. See the change history section in the readme for more details. In addition to the readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of one main data folder containing (1) file-level metadata; (2) data dictionary; (3) readme; (4) and subfolders containing data, figures, and scripts. The data folder contains datasets used for the analyses in the manuscript in image, text-delimited or geospatially-referenced formats. The figures folder contains the figures from the manuscript in different formats. The scripts folder contains all of the scripts used to complete the analyses in the manuscript. All files are .csv, .rds, .dbf, .prj, .shp, .shx, .jpg, .png, .R, .Rproj, or .pdf. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with "Coupled primary production and respiration in a large river contrasts with smaller rivers and streams."

This data package is associated with the publication "Coupled primary production and respiration in a large river contrasts with smaller rivers and streams." in review at Limnology and Oceanography (Roley et al. 2023). This study focuses on understanding ecosystem metabolism for the Hanford Reach of the Columbia River in Washington state, a free-flowing stretch with a substantial discharge. Large rivers have been overlooked compared to small and medium rivers due to the challenges associated with measurements. Our study presents novel ways to address these challenges and highlights that metabolism patterns in large rivers differ from those observed in small-medium rivers and requires the application of knowledge and tools beyond those implemented for smaller rivers.This data package includes the data and R scripts for the analyses described in Roley et al. 2023. It includes dissolved oxygen and temperature data from a dissolved oxygen HOBO sensor, light data collected from the National Solar Radiation Database (https://nsrdb.nrel.gov/) and hydrologic variables estimated from the MASS-1 model (Niehus et al.; 2014). It also includes metabolism estimates (gross primary production, ecosystem respiration, and net ecosystem production) estimated via streamMetabolizer (Appling et al.; 2018). All analyses in the paper can be replicated with these data and scripts.The data package is comprised of one main data folder. The folder includes (1) file-level metadata (flmd); (2) a data dictionary (dd) for each data file; (3) data files; and (4) R scripts for metabolism estimates and data analysis. All files are .R, .csv, or .pdf.

54 ENVIRONMENTAL SCIENCES↗

Artificial intelligence models, photos, and data associated with the manuscript “Quantifying Streambed Grain Size, Uncertainty, and Hydrobiogeochemical Parameters Using Machine Learning Model YOLO” (v2)

This data package is associated with the manuscript “Quantifying Streambed Grain Size, Uncertainty, and Hydrobiogeochemical Parameters Using Machine Learning Model YOLO” published in Water Resources Research (Chen et al., 2024). This data package includes the training, validation, testing, and prediction data used by the artificial intelligence (AI) model for automated grain size and hydro-biogeochemistry quantification using streambed photos. The grain size data are extracted for each photo using You Look Only Once (YOLO), a pre-trained object detection model. This data package was originally published in October 2023. It was updated August 2025 (v2; new and modified files). File and folder names were not revised to indicate changes. See the change history section in the readme for more details. Please see flmd.csv for a list of all files contained in this data package and descriptions for each. Please see dd.csv for a data dictionary that defines the column headers of .csv files in the data package. This dataset is comprised of one data folder containing (1) file-level metadata; (2) data dictionary; (3) readme; and (4) six subfolders. Subfolders 1 to 4 include the training, validation, testing, and prediction data. Subfolder 5_Summary includes the summary results of different combinations of training, validation, testing, and prediction data. Subfolder 6_SupplementalData includes additional data downloaded from public sources (Kaufman et al., 2023a; Kaufman et al., 2023b; Garefalakis et al., 2023; Mair et al., 2024; https://github.com/river-corridors-sfa/Geospatial_variables). In total, the data package includes 110 folders and 44,283 files. These files include 9,047 .jpg photos, 1 .png photo, 3 .tif photos; 26,639 photo labels and individual grain sizes and probability from AI (.txt); 8,447 grain size distribution data (.dat); and 126 CSV files for results summary, and 14 required metadata files (.xlsx). The summary CSV files contain 68 columns and approximately 2,200 rows that represent photo names, site locations, recording time, GPS coordinates, grains sizes (D10, D50, D60, and D84), number of grains, and additional hydro-biogeochemical data such as water depth, flow velocity, Manning’s coefficient, friction factor, hydraulic conductivity, permeability, streambed interstitial velocity magnitude, mass transfer rate, and nitrate uptake velocity. The photos were obtained from 75 sites in the Yakima River Basin and the Columbia River shorelines, and other associated data from samples and sensors obtained when the photos were taken are publicly available (Fulton et al. 2022; Grieger et al. 2023). All files are .csv, .txt, .dat, .jpg, or .pdf. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with the manuscript evaluating the hydrologic responses of the Pacific Northwest watersheds to wildfires (v2)

This data package is associated with the publication “Evaluating Post-fire Watershed Response to Varying Burn Severity and Precipitation Regimes Using Fully-distributed and Integrated Hydrologic Models” submitted to Journal of Hydrology (Li et al. 2025). In this study, we employed the Advanced Terrestrial Simulator (ATS), an integrated watershed model that couples surface flow, subsurface flow, and canopy biophysical processes, to investigate post-fire hydrologic responses in a few selected watersheds with varying burn severity.The data package contains the required input data (meteorological forcing, Leaf Area Index, wildfire burn severities, etc.) to run the model, configuration files, the Jupyter notebooks in Python to pre-process and post-process data, the figures in the manuscript, and the modeling output files. The variables include watershed-averaged evapotranspiration, watershed-averaged surface/subsurface/canopy water content, and river discharge at watershed outlet.The data package contains a file-level metadata that lists and describes all the files contained in the data package (ATS_flmd.csv), a data dictionary file that defines columns headers across all csv files contained in the data package (ATS_dd.csv), a data package level readme file (the current file), and four zipped folders.The ‘data’ folder provides data needed to run the model in .h5, .i2s, .xyz, .shp, and .exo formats. The sub-folders are for each data types. The ‘model’ folder provides input files (.xml format) and essential model outputs. Each sub-folder provides the files from each simulated watershed. The ‘notebooks’ folder provides the Jupyter notebooks (.ipynb format) for pre- and post- processing model files, and for producing the figures in the manuscript. The ‘figures’ folder provides the figures associated with manuscript in .pdf and .png formats.The ‘model’ folder and the ‘data’ folder have been split into 5GB-large pieces using the Linux command ‘split -b 5120m model.zip model.zip.’ and ‘split -b 5120m data.zip data.zip.’, respectively. They can be merged back using the Linux command ‘cat model.zip.* > model.zip’ and ‘cat data.zip.* > data.zip’, respectively.

54 ENVIRONMENTAL SCIENCES↗

Data, scripts, and figures associated with a manuscript studying impact of climate and topography on post-fire vegetation recovery.

This data package is associated with the publication “Impact of Topography and Climate on Post-fire Vegetation Recovery Across Different Burn Severity and Land Cover Types through Machine Learning” submitted to Remote Sensing of Environment (Zahura et al. 2023). In this research, a machine learning algorithm, random forest (RF), was utilized to examine the impact of climate and topography on post-fire vegetation recovery. We used enhanced vegetation index (EVI) to examine varying burn severity and land cover types. The data package includes the input files for RF model training, outputs from model predictions and analysis, and python scripts to run the model, analyze the results to understand model performance and interpretability, and plot manuscript figures. This data package contains three folders (Data, Scripts, and Figures), a file-level metadata (FLMD) csv, and a data dictionary (dd) csv. Please see Postfire_recovery_flmd.csv for a list of all files contained in this data package and descriptions for each. The data dictionary (Postfire_recovery_dd.csv) describes the csv column headers. The “Data” folder provides all the inputs and outputs to train the RF model, evaluate performance, and interpret predictions. The “Scripts” folder contains python scripts and jupyter notebooks for model training and result analysis. The “Figures” folder includes the figures used in the manuscript in “.png” and “.jpg” format.

54 ENVIRONMENTAL SCIENCES↗

Scripts and data associated with a manuscript linking soil and sediment elemental composition with dissolved organic matter chemistry across CONUS

This data package provides scripts and geochemical data for a manuscript titled “Linkages between mineral element composition of soils and sediments with hyporheic zone dissolved organic matter chemistry across the contiguous United States” (preprint: doi: 10.22541/essoar.169447343.31694990/v1). This data is associated with the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS, https://whondrs.pnnl.gov) and is an extension of the Summer 2019 Sampling campaign which crowdsourced samples from rivers and sediment across the continental United States. Data from this study can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1603775 and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719. The main objective of this manuscript was to couple sediment water extractable dissolved organic matter chemistry, defined by ultra-high resolution mass spectrometry, with localized sediment elemental composition and watershed scale soil elemental characteristics. This data package contains one main folder with four subfolders. The main data folder contains (1) readme; (2) data dictionary (dd); (3) file-level metadata (flmd); (4) an R markdown to reproduce manuscript figures and analyses; (5) a pdf of instructions to reproduce NGS interpolations with ArcGIS software; and (6) a python script to reproduce NGS extrapolations with python. The four subfolders contain files required to reproduce NGS extrapolations include (1) ‘CONUS_boundaries’ containing boundary layers (.shp) for the Continental United States; (2) ‘ngs_project’ containing files (.shp) with point level NGS soil elemental data (Grossman et al., 2004); (3) ‘raster_outputs’ containing the interpolated raster output files for various soil elements; and (4) ‘NGS_Chemistry_Final’ contain final extracted soil elemental data.

54 ENVIRONMENTAL SCIENCES↗

Data and Scripts associated with “Lambda-PFLOTRAN: Workflow for Incorporating Organic Matter Chemistry Informed by Ultra High Resolution Mass Spectrometry into Biogeochemical Modeling.”

This data package is associated with the publication “Lambda-PFLOTRAN: Workflow for Incorporating Organic Matter Chemistry Informed by Ultra High Resolution Mass Spectrometry into Biogeochemical Modeling” submitted to Geoscientific Model Development (Muller et al., 2024). In this manuscript, organic matter chemistry and thermodynamics are directly connected to reactive transport simulators through the newly developed Lambda-PFLOTRAN (Parallel Reactive Flow and Transport model) workflow tool that succinctly incorporates organic matter chemistry data generated from Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) into reaction networks to simulate aerobic respiration of the organic matter and the resulting biogeochemistry. Lambda-PFLOTRAN is a python-based workflow, executed through a Jupyter Notebook interface, that digests raw FTICR-MS data, develops a representative reaction network based on substrate-explicit thermodynamic modeling (also termed lambda modeling due to its key thermodynamic parameter λ used therein), and completes a biogeochemical simulation with the open source, reactive flow, and transport code PFLOTRAN. This data package contains Jupyter Notebook based workflows for two test cases for running biogeochemical simulations of organic matter oxidation identified by FTICR-MS. It contains four primary folders (workflow, data, src, and analysis), a file-level metadata file (Muller_2024_Lambda_PFLOTRAN_Manuscript_Data_Package_flmd.csv) that lists all the files contained in this data package with a short description of each, and a data dictionary (Muller_2024_Lambda_PFLOTRAN_Manuscript_Data_Package_dd.csv) file that describes the tabular column headers. The ‘workflow’ folder contains the Jupyter Notebook based workflows for running the lambda analysis, PFLOTRAN simulation, sensitivity analysis and parameter estimation. The ‘data’ folder contains the FTICR-MS data, initial conditions, and incubation data for test cases 1 and 2 in folders titled ‘WHONDRS’ and ‘Colloids’, respectively. The data folder also has a ‘Database’ folder containing a reaction network for bulk organic matter (assumed to be CH2O) and a general database for PFLOTRAN (hanford_rxn_network). The CH2O reaction network defines bulk organic matter oxidation. Biogeochemical simulations are completed for both the lambda binned organic matter and bulk organic matter reaction networks. The ‘hanford_rxn_network’ database includes information required for PFLTORAN simulations including ion size, molar mass, and charge of the aqueous species, gases, and minerals phases. The ‘src’ folder contains python source codes for performing lambda analysis, PFLOTRAN simulation, sensitivity analysis and parameter estimation. The ‘analysis’ folder contains outputs from the test cases 1 and 2 including lambda analysis, PFLOTRAN runs and the calibration results.

54 ENVIRONMENTAL SCIENCES↗

Data and Scripts Associated with the Manuscript “Water Column Respiration in the Yakima River Basin is Explained by Temperature, Nutrients and Suspended Solids”

This data package is associated with the publication “Water Column Respiration in the Yakima River Basin is Explained by Temperature, Nutrients and Suspended Solids” published in EGU Biogeochemistry (Laan et al. 2025). In this research, water column respiration (ERwc) data, surface water chemistry data, organic matter (OM) chemistry data, and publicly available geospatial data were used in analysis to evaluate the variability in ERwc at 47 sites across the Yakima River basin in Washington, USA. In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. The data package includes the data inputs, and outputs, and R scripts to reproduce all the analyses performed in the manuscript and create manuscript figures. The data package is comprised of three main folders (Code, Data, and Figures). The Code folder is comprised of four scripts and three analysis-specific subfolders that contain the R scripts to perform the analyses described in the publication and create publication figures. The Data folder is comprised of two “.csv” files and four subfolders that contain data input and output files. The Published_Data folder contains a readme that directs the user to download the appropriate files and add to this folder when using scripts. The Figures folder includes figures from the manuscript in “.pdf” and “.png” formats and a folder with intermediate figure files. This data package is associated with a GitHub repository which can be found at https://github.com/river-corridors-sfa/rcsfa-RC2-SPS-ERwc. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Geophysical and Environmental Monitoring Data, and Subsurface Flow Modelling Results for Chicken Bone Meadow, Mt. Snodgrass, Crested Butte, CO

This dataset includes geoelectrical monitoring data acquired between October 2021 and November 2022, soil moisture and temperature data, groundwater data obtained from borehole SNIB covering the period from June 2021 to September 2022, and hydrological modelling results. The data were acquired to investigate how variations in bedrock type and topography, and vegetation cover control subsurface flow dynamics. To provide insights into the subsurface flow dynamics and their controls, a monitoring transect was installed at the Chicken Bone Meadow, Mt. Snodgrass, Crested Butte, CO, measuring the spatio-temporal variations of soil moisture, soil and snow temperature, subsurface electrical resistivity variations, and groundwater dynamics. Field data are organized in a folder structure, with Electrical Resistivity Tomography (ERT) data being provided as one file per measurement, and data of the soil moisture and temperature sensors being provided as text files covering the entire monitoring period. The ‘Locations.csv’ file contains the location of all sensors, given in NAD83 – UTM Zone 13N. ERT monitoring data has been processed to filter data based on reciprocal errors (data with errors > 30% were removed), a linear error model was fitted to each survey, and to ensure a constant set of measurements for time-lapse inversion, filtered data were interpolated and assigned a 100% measurement error. Soil moisture and temperature data were acquired at 15 min intervals, and averaged to provide 1h data. Weather data and borehole data (groundwater depth, conductivity and temperature) were acquired at 30 min intervals, and are provided as daily measurements; all measurements are averaged, except of precipitation values, which are given as daily accumulation. The hydrological model was set up along the ERT monitoring transect, and net infiltration was used as surface boundary condition and derived from the weather data. Four different results are provided, (1) results for a parameterization using hydraulic permeability and porosity as derived from the ERT data through petrophysical relationships, and (2) three simplified model results, using 1 to 3 geological layers above the bedrock. Modelling was performed using PFLOTRAN, and for each model the PFLOTRAN input files are provided. The result files include weekly hydrological modelling results (e.g., saturation, velocities, pressures), as well as the model parameterization. The dataset additionally includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.

54 ENVIRONMENTAL SCIENCES↗

Radon Isotopes and Stable Water Isotopes from Coal Creek Watershed, Colorado (2021)

The radon isotope and stable water isotope data for Coal Creek Watershed, Colorado, consists of d2H, d18O, and 222Rn values from samples collected at 8 stream location along Coal Creek, samples from 7 groundwater springs within the watershed, and precipitation isotope samples collected by Next Generation Water Observing System (NGWOS) from a collector within the watershed. All stream and spring samples were collected between June and October, 2021, and precipitation isotope samples were collected between November 2020 and September 2021. These data were collected to evaluate how groundwater contributions to Coal Creek originating from a fractured hillslope and alluvial fan respond to summer monsoon rains and seasonal drying. Understanding of groundwater-surface water interactions in montane systems in critical for the future of water availability in the Western US as groundwater contributions are expected to become more important for sustaining summer stream flows. This data package contains: (1) a csv of all radon samples; (2) a csv of all stream and spring isotope samples; (3) a csv of precipitation isotope samples; and (4) a csv of locations for each sampling site. The dataset additionally includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.

54 ENVIRONMENTAL SCIENCES↗

Continuous soil temperature measurements from 2019-10-4 to 2020-10-4, Teller road Mile 27, Seward Peninsula, Alaska

The dataset contains depth-resolved soil temperature measured at 45 discrete locations in a watershed located along the Nome-Teller road at Mile 27 in Seward Peninsula, Alaska. The dataset was generated to understand the local heterogeneity of soil thermal dynamics and their controls in a discontinuous permafrost region. At each location, temperatures were measured by a distributed temperature profiling probe designed based on Dafflon et al (2022). The dataset includes a description of the probe locations in the "Probe_locations.csv" file and the 45 data files (Soil_temperatures_*.csv). Metadata files include data descriptions (_dd.csv) for tabular data. All included files are listed and described in NGA513_flmd.csv.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

Models, data, and scripts associated with “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning”

This data package is associated with the publication “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning’’ submitted to the Journal of Geophysical Research: Machine Learning and Computation (Scheibe et al. 2024). River sediment respiration observations are expensive and labor intensive to obtain and there is no physical model for predicting this quantity. The Worldwide Hydrobiogeochemisty Observation Network for Dynamic River Systems (WHONDRS) observational data set (Goldman et al.; 2020) is used to train machine learning (ML) models to predict respiration rates at unsampled sites. This repository archives training data, ML models, predictions, and model evaluation results for the purposes of reproducibility of the results in the associated manuscript and community reuse of the ML models trained in this project. One of the key challenges in this work was to find an optimum configuration for machine learning models to work with this feature-rich (i.e. 100+ possible input variables) data set. Here, we used a two-tiered approach to managing the analysis of this complex data set: 1) a stacked ensemble of ML models that can automatically optimize hyperparameters to accelerate the process of model selection and tuning and 2) feature permutation importance to iteratively select the most important features (i.e. inputs) to the ML models. The major elements of this ML workflow are modular, portable, open, and cloud-based, thus making this implementation a potential template for other applications. This data package is associated with the GitHub repository found at Please see the file level metadata (flmd; “sl-archive-whondrs_flmd.csv”) for a list of all files contained in this data package and descriptions for each. Please see the data dictionary (dd; “sl-archive-whondrs_dd.csv”) for a list of all column headers contained within comma separated value (csv) files in this data package and descriptions for each. The GitHub repository is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning models trained on the data in “input_data”; (3) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; (4) “examples” contains the visualization of the results in this repository including plotting scripts for the manuscript (e.g., model evaluation, FPI results) and scripts for running predictions with the ML models (i.e., reusing the trained ML models); (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. Furthermore, depending on the number of features used to train the ML models, the preprocessing and postprocessing scripts, and their intermediate results, can also be different branch-to-branch. The “main-*” branches are meant to be starting points (i.e. trunks) for each model branch (i.e. sprouts). Please see the Branch Navigation section in the top-level README.md in the GitHub repository for more details. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please the top-level README.md in the GitHub repository for more details on the automation.

13C↗

Data and scripts associated with a manuscript investigating dissolved organic matter and microbial community linkages across seven globally distributed rivers

This data package is associated with the publication “Meta-metabolome ecology reveals that geochemistry and microbial functional potential are linked to organic matter development across seven rivers” submitted to Science of the Total Environment. This data package includes the data necessary to replicate the analyses presented within the manuscript to investigate dissolved organic matter (DOM) development across broad spatial distances and within divergent biomes. Specifically, we included the Fourier transform ion cyclotron mass spectrometry (FTICR-MS) data, geochemistry data, annotated metagenomic data, and results from ecological null modeling analyses in this data package. Additionally, we included the scripts necessary to generate the figures from the manuscript. Complete metagenomic data associated with this data package can be found at the National Center for Biotechnology (NCBI) under Bioproject PRJNA946291. This dataset consists of (1) four folders; (2) a file-level metadata (flmd) file; (3) a data dictionary (dd) file; (4) a factor sheet describing samples; and (5) a readme. The FTICR Data folder contains (1) the processed Fourier transform ion cyclotron mass spectrometry (FTICR-MS) data; (2) a transformation-weighted characteristics dendrogram generated from the FTICR-MS data; and (3) the script used to generate all FTICR-MS related figures. The Geochemical Data folder contains (1) the single geochemistry data file and (2) the R script responsible for generating associated figures. The Metagenomic Data folder contains (1) annotation information across different levels; (2) carbohydrate active enzyme (CAZyme) information from the dbCAN database (Yin et al., 2012); (3) phylogenetic tree data (FASTAs, alignments, and tree file); and (4) the scripts necessary to analyze all of these data and generate figures. The Null Modeling Data folder contains (1) data generated during null modeling for each river and all rivers combined and (2) the R scripts necessary to process the data. All files are .csv, .pdf, .tsv, .tre, .faa, .afa, .tree, or .R.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with a manuscript on a meta-analysis synthesizing stream biogeochemical response to wildfires across space and time (v2)

This data package is associated with the publication “Catchment characteristics modulate the influence of wildfires on nitrate and dissolved organic carbon in lotic systems across space and time: A meta-analysis” submitted to Global Biogeochemical Cycles (Cavaiani et al. 2025). This study uses meta-analytical techniques to evaluate the effect of wildfire on in-stream responses in burned and unburned watersheds. The study aims to provide additional insight into the range of responses and net influences that wildfires have on hydro-biogeochemistry across broad spatial scales, burn extents, and the persistence of water-quality change. This study compiles data and metadata from 18 total publications that includes 1) surface water geochemistry data (dissolved organic carbon; nitrate), 2) climate classifications, 3) year of the wildfire, 4) the time lag between when the fire occurred and when the sampling occurred, and 5) study design of the publication. In total, this meta-analysis draws data that spans 8 climate guilds, 3 biomes, 62 watersheds, and 20 unique wildfires. See Sites_meta_data.csv for citations of the papers used in this meta-analysis. All R scripts and the associated data can also be found on GitHub at This data package was originally published in March 2024. It was updated in April 2025 (v2; new and modified files). See the change history section in the readme for more details. This data package contains five primary folders that include the following: (1) inputs; (2) output for analysis; (3) initial plots; (4) R scripts; and (5) GIS data. The data package also contains a data dictionary (dd) that provides column header definitions and a file-level metadata (flmd) file that describes every file. The “inputs” folder contains a list of all publications identified during the formal web search and an indication of whether each publication was included in the final analysis. Additionally, it includes site-level metadata, catchment characteristics, and GIS data for all publications included in the final analysis. The “Output_for_analysis” folder contains all data frames and figures generated from each R script used for additional data analysis. The “initial_plots” folder includes all exploratory figures that will be included in a supplemental and figures that will be submitted with the manuscript for publication. The “R_scripts” folder contains the scripts that perform all the data manipulations, statistical analyses, and plots. The “gis_data” folder includes shape files for each fire included in this meta-analysis. This data package contains the following file types: csv, pdf, jpeg, cpg, dbf, prj, shp, shp.ea.iso.xml, shp.iso.xml, shx.

54 ENVIRONMENTAL SCIENCES↗

iButton and Tinytag snow/ground interface temperature measurements at Teller 27 and Kougarok 64 from 2022-2023, Seward Peninsula, Alaska

Snow/ground interface temperature measurements were collected at the NGEE Arctic Teller Road Site at mile marker 27 (TL_MM27) and at the Kougarok Road Site at mile marker 64 (KG_MM64) on the Seward Peninsula, Alaska. Data were collected between October 1, 2022 to September 18, 2023 using iButton Link DS1921G-F5# Thermochron miniature temperature sensors (https://www.ibuttonlink.com/products/ds1921g) and Tinytag TGP-4017 internal sensors (https://www.micronmeters.com/product/tgp-4017-internal-sensor-40-to-85-c-40-f-to-185-f) deployed across the Kougarok and Teller sites. These sensors are a cost-efficient way to collect snowpack temperatures at a higher spatial resolution than what is normally achieved. iButton data were collected every 4 hours, while Tinytag data were collected every 30 minutes. In total, data were collected from 196 iButtons and 26 Tinytags. This dataset contains four *.csv files of near-ground surface temperatures at various locations throughout each study site and two *.kml files of sensor locations. Data were collected throughout the snow cover season so that snowpack characteristics could be derived using the temperature data. Sensors were placed both inside and outside of vegetation to better capture the spatial variability of snow properties across each domain.The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic) was a 15-year research effort (2012-2027) to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy’s Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy’s Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗