Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “dictionaries”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

2013-2014 Greater Fairbanks, Alaska, Transportation Survey

# 2013–2014 Greater Fairbanks, Alaska, Transportation Survey The 2013–2014 Greater Fairbanks Transportation Survey obtained behavior data for regional travel demand modeling. The planning region in Alaska comprised the North Star Borough—known as the PM2.5 nonattainment region—and included the cities of Fairbanks and North Pole. ## Data Collection Agency The Alaska Department of Transportation and Public Facilities conducted the survey. ## Methodology Data collection occurred in two phases: the first in fall 2013 and the second in winter 2014. The first phase employed address-based sampling to recruit more than 1,700 households for a one-day personal travel survey, and a sub-sample participated with global position system (GPS) and on-board diagnostic (OBD) loggers installed in their vehicles (282 vehicles) for one week. The purpose of phase one was to better understand the impact of vehicle emissions on air quality in the PM2.5 nonattainment region. Many of the households participating in the vehicle GPS/OBD portion of phase one were asked to participate in phase two. ## Drive Cycle Processing and Filtering NREL has developed a GPS data filtration routine to filter erroneous data points in individual drive cycles sourced from GPS devices mounted in vehicles. Second-by-second drive cycle data collected from GPS-instrumented vehicles during this survey have passed through NREL's drive cycle processing and filtering routines. ## Survey Records Survey records include 135 households. ## More Information For more information about the survey, see the [Greater Fairbanks Transportation Survey Final Report](https://www.nrel.gov/media/docs/libraries/tsdc/greater-fairbanks-transportation-survey-final-report.pdf?sfvrsn=16ecd45b_1). ## Transportation Data For details on available travel survey data and variable definitions, see the [data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/akdot_data_dictionary.pdf?sfvrsn=16778e11_1). NREL-generated drive cycle data are also available for this survey. For details on available data and variable definitions, see the [drive cycle data dictionary](https://www.nrel.gov/media/docs/libraries/tsdc/drive_cycles_data_dictionary.pdf?sfvrsn=7de7e888_1). Transportation data are available as zipped files. [Download Winzip](http://www.winzip.com/downwz.htm).

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Fungal and bacterial growth variation due to drought and nitrogen addition experimental treatments. Loma Ridge Experimental Project. 2010-2012

Terrestrial ecosystem models assume that microbial communities respond instantaneously, or are immediately resilient, to environmental change. Here we tested this assumption by quantifying the resilience of a leaf litter community to changes in precipitation or nitrogen availability. By manipulating composition within a global change experiment, we decoupled the legacies of abiotic parameters versus that of the microbial community itself. After one rainy season, more variation in fungal composition could be explained by the original microbial inoculum than the litterbag environment (18% versus 5.5% of total variation). This compositional legacy persisted for 3 years, when 6% of the variability in fungal composition was still explained by the microbial origin. In contrast, bacterial composition was generally more resilient than fungal composition. Microbial functioning (measured as decomposition rate) was not immediately resilient to the global change manipulations; decomposition depended on both the contemporary environment and rainfall the year prior. Finally, using metagenomic sequencing, we showed that changes in precipitation, but not nitrogen availability, altered the potential for bacterial carbohydrate degradation, suggesting why the functional consequences of the two experiments may have differed. Predictions of how terrestrial ecosystem processes respond to environmental change may thus be improved by considering the legacies of microbial communities. This data package includes ten csv files (five data files and their corresponding data dictionaries) and one file-level metadata excel file. Data files contains information about which plots were exposed to treatments related to drought and nitrogen, information about litter bags reciprocal transplants manipulation for water input and nitrogen, detail information about water addition, precipitation records, and litter variables collected. Data dictionary files include detail explanation for each column in the data files. The file-level metadata file describes each file mentioned above. All the analyses were done using the R software.

54 ENVIRONMENTAL SCIENCES↗

QA/QC-ed Groundwater Level Time Series in PLM-1 and PLM-6 Monitoring Wells, East River, Colorado (2016-2022)

This data set contains QA/QC-ed (Quality Assurance and Quality Control) water level data for the PLM1 and PLM6 wells. PLM1 and PLM6 are location identifiers used by the Watershed Function SFA project for two groundwater monitoring wells along an elevation gradient located along the lower montane life zone of a hillslope near the Pumphouse location at the East River Watershed, Colorado, USA. These wells are used to monitor subsurface water and carbon inventories and fluxes, and to determine the seasonally dependent flow of groundwater under the PLM hillslope. The downslope flow of groundwater in combination with data on groundwater chemistry (see related references) can be used to estimate rates of solute export from the hillslope to the floodplain and river. QA/QC analysis of measured groundwater levels in monitoring wells PLM-1 and PLM-6 included identification and flagging of duplicated values of timestamps, gap filling of missing timestamps and water levels, removal of abnormal/bad and outliers of measured water levels. The QA/QC analysis also tested the application of different QA/QC methods and the development of regular (5-minute, 1-hour, and 1-day) time series datasets, which can serve as a benchmark for testing other QA/QC techniques, and will be applicable for ecohydrological modeling. The package includes a Readme file, one R code file used to perform QA/QC, a series of 8 data csv files (six QA/QC-ed regular time series datasets of varying intervals (5-min, 1-hr, 1-day) and two files with QA/QC flagging of original data), and three files for the reporting format adoption of this dataset (InstallationMethods, file level metadata (flmd), and data dictionary (dd) files).QA/QC-ed data herein were derived from the original/raw data publication available at Williams et al., 2020 (DOI: 10.15485/1818367). For more information about running R code file (10.15485_1866836_QAQC_PLM1_PLM6.R) to reproduce QA/QC output files, see README (QAQC_PLM_readme.docx). This dataset replaces the previously published raw data time series, and is the final groundwater data product for the PLM wells in the East River. Complete metadata information on the PLM1 and PLM6 wells are available in a related dataset on ESS-DIVE: Varadharajan C, et al (2022). https://doi.org/10.15485/1660962. These data products are part of the Watershed Function Scientific Focus Area collection effort to further scientific understanding of biogeochemical dynamics from genome to watershed scales. 2022/09/09 Update: Converted data files using ESS-DIVE’s Hydrological Monitoring Reporting Format. With the adoption of this reporting format, the addition of three new files (v1_20220909_flmd.csv, V1_20220909_dd.csv, and InstallationMethods.csv) were added. The file-level metadata file (v1_20220909_flmd.csv) contains information specific to the files contained within the dataset. The data dictionary file (v1_20220909_dd.csv) contains definitions of column headers and other terms across the dataset. The installation methods file (InstallationMethods.csv) contains a description of methods associated with installation and deployment at PLM1 and PLM6 wells. Additionally, eight data files were re-formatted to follow the reporting format guidance (er_plm1_waterlevel_2016-2020.csv, er_plm1_waterlevel_1-hour_2016-2020.csv, er_plm1_waterlevel_daily_2016-2020.csv, QA_PLM1_Flagging.csv, er_plm6_waterlevel_2016-2020.csv, er_plm6_waterlevel_1-hour_2016-2020.csv, er_plm6_waterlevel_daily_2016-2020.csv, QA_PLM6_Flagging.csv). The major changes to the data files include the addition of header_rows above the data containing metadata about the particular well, units, and sensor description. 2023/01/18 Update: Dataset updated to include additional QA/QC-ed water level data up until 2022-10-12 for ER-PLM1 and 2022-10-13 for ER-PLM6. Reporting format specific files (v2_20230118_flmd.csv, v2_20230118_dd.csv, v2_20230118_InstallationMethods.csv) were updated to reflect the additional data. R code file (QAQC_PLM1_PLM6.R) was added to replace the previously uploaded HTML files to enable execution of the associated code. R code file (QAQC_PLM1_PLM6.R) and ReadMe file (QAQC_PLM_readme.docx) were revised to clarify where original data was retrieved from and to remove local file paths.

54 ENVIRONMENTAL SCIENCES↗

Transport and Retention of Particulate Organic Matter in Sand: Lab Experiments and Modelling

Flow-through reactor (FTR) experiments were conducted at three different downward vertical flow rates to study the transport and retention of particulate organic matter (Chlorella powder) in riverbed sediments. This data package includes data files (.csv) regarding effluent collected from the FTRs during the experiments (time, duration, volume of each sample collection; concentration of nonreactive bromide tracer; concentration of suspended Chlorella determined by absorbance measurements) and sediment (sand) slices collected from the FTRs after the experiments (mass of Chlorella retained in 1 cm depth intervals determined by loss-on-ignition method). More detailed descriptions of data are provided in the data dictionary, and detailed methods are available in the associated MSc thesis listed in ‘Related References.’ The data dictionary and file level meta data (FLMD) are included in the data package as both .csv and .xlsx files. This data package also includes timelapse videos (.mp4) of the experiments, and modelling scripts associated with the experiments. The modelling scripts are included as MATLAB ‘live scripts’ (.mlx, which can be run using MATLAB), and are also included as .pdf and .html files with output plots included (which can be viewed without a MATLAB license and installation). The data files and modelling scripts are organized into a .zip folder for each of the three FTR experiments (three flow rates).

54 ENVIRONMENTAL SCIENCES↗

Data associated with “Different methods of estimating riverbed sediment grain size diverge at the basin scale ” (v2)

This data package is associated with the publication “Different methods of estimating riverbed sediment grain size diverge at the basin scale” published in Frontiers in Earth Science (Regier et al., 2025). The distribution of sediment grain size in streams and rivers is often quantified by the median grain size (d50), a key metric for understanding and predicting hydrologic and biogeochemical function of streams and rivers. Manual methods to measure d50 are time-consuming and ignore larger grains, while model-based methods to estimate d50 often over-generalize basin characteristics, and therefore cannot accurately represent site-scale heterogeneity. Here, we apply a machine learning-enabled photogrammetry methodology (You Only Look Once, or YOLO) for estimating d50 for grains > 2 mm based on images collected from streams and rivers throughout the Yakima River Basin (YRB). To understand how such methods may help bridge the gaps in resolution and accuracy between manual and catchment characteristics model-based d50 estimates, we compared YOLO d50 values to manual and model-based estimates across the YRB. We found distinct differences among methods for d50 averages and variability, and relationships between d50 estimates and basin characteristics. Source images can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/1892052. This data package was originally published in May 2023. It was updated August 2025 (v2; new and modified files). File and folder names were not revised to indicate changes. See the change history section in the readme for more details. In addition to the readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of one main data folder containing (1) file-level metadata; (2) data dictionary; (3) readme; (4) and subfolders containing data, figures, and scripts. The data folder contains datasets used for the analyses in the manuscript in image, text-delimited or geospatially-referenced formats. The figures folder contains the figures from the manuscript in different formats. The scripts folder contains all of the scripts used to complete the analyses in the manuscript. All files are .csv, .rds, .dbf, .prj, .shp, .shx, .jpg, .png, .R, .Rproj, or .pdf. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Aerobic respiration controls on shale weathering, Geochimica et Cosmochimica Acta, 2023: Dataset

This data package was generated in order to support the development of a deep-time weathering model and to assess the coupling between shale weathering and aerobic respiration in the paper “Aerobic respiration controls on shale weathering” by Stolze et al., Geochimica et Cosmochimica Acta (2023). The package contains two csv files providing the average CO2(g) concentration profiles [ppm] and mineral concentration profiles [wt%], respectively. The CO2(g) concentration profiles were measured in the vicinity of the monitoring well PLM2 between January 2018 and April 2019. The gas samples were collected in the unsaturated zone to a depth of 1.52 m. The mineral concentration profiles were determined by X-Ray diffraction (XRD). The XRD measurements were performed on sub-core samples collected in the monitoring well PLM3 down to a depth of 7.01 m. The dataset additionally includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.Update on 2024-05-28: Revised versions of the CSV data files (CO2_data_GCA_Stolze_et_al_2023.csv and XRD_data_GCA_Stolze_et_al_2023.csv) were made to apply ESS-DIVE's CSV reporting format guidelines. Updated versions of the File Level Metadata (v2_20240528_flmd.csv) and Data Dictionary (v2_20240528_dd.csv) files were updated to reflect the changes made to the CSV files.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with a manuscript investigating impacts of solid phase extraction on freshwater organic matter optical signatures and mass spectrometry pairing

This data package is associated with the publication “Investigating the impacts of solid phase extraction on dissolved organic matter optical signatures and the pairing with high-resolution mass spectrometry data in a freshwater system” submitted to “Limnology and Oceanography: Methods.” This data is an extension of the River Corridor and Watershed Biogeochemistry SFA’s Spatial Study 2021 (https://doi.org/10.15485/1898914). Other associated data and field metadata can be found at the link provided. The goal of this manuscript is to assess the impact of solid phase extraction (SPE) on the ability to pair ultra-high resolution mass spectrometry data collected from SPE extracts with optical properties collected on ambient stream samples. Forty-seven samples collected from within the Yakima River Basin, Washington were analyzed dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC), absorbance, and fluorescence. Samples were subsequently concentrated with SPE and reanalyzed for each measurement. The extraction efficiency for the DOC and common optical indices were calculated. In addition, SPE samples were subject to ultra-high resolution mass spectrometry and compared with the ambient and SPE generated optical data. Finally, in addition to this cross-platform inter-comparison, we further performed and intra-comparison among the high-resolution mass spectrometry data to determine the impact of sample preparation on the interpretability of results. Here, the SPE samples were prepared at 40 milligrams per liter (mg/L) based on the known DOC extraction efficiency of the samples (ranging from ~30 to ~75%) compared to the common practice of assuming the DOC extraction efficiency of freshwater samples at 60%. This data package folder consists of one main data folder with one subfolder (Data_Input). The main data folder contains (1) readme; (2) data dictionary (dd); (3) file-level metadata (flmd); (4) final data summary output from processing script; and (5) the processing script. The R-markdown processing script (SPE_Manuscript_Rmarkdown_Data_Package.rmd) contains all code needed to reproduce manuscript statistics and figures (with the exception of that stated below). The Data_Input folder has two subfolders: (1) FTICR and (2) Optics. Additionally, the Data_Input folder contains dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data (SPS_NPOC_Summary.csv) and relevant supporting Solid Phase Extraction Volume information (SPS_SPE_Volumes.csv). Methods information for the optical and FTICR data is embedded in the header rows of SPS_EEMs_Methods.csv and SPS_FTICR_Methods.csv, respectively. In addition, the data dictionary (SPS_SPE_dd.csv), file level metadata (SPS_SPE_flmd.csv), and methods codes (SPS_SPE_Methods_codes.csv) are provided. The FTICR subfolder contains all raw FTICR data as well as instructions for processing. In addition, post processed FTICR molecular information (Processed_FTICRMS_Mol.csv) and sample data (Processed_FTICRMS_Data.csv) is provided that can be directly read into R with the associated R-markdown file. The Optics subfolder contains all Absorbance and Fluorescence Spectra. Fluorescence spectra have been blank corrected, inner filter corrected, and undergone scatter removal. In addition, this folder contains Matlab code used to make a portion of Figure 1 within the manuscript, derive various spectral parameters used within the manuscript, and used for parallel factor analysis (PARAFAC) modeling. Spectral indices (SPS_SpectralIndices.csv) and PARAFAC outputs (SPS_PARAFAC_Model_Loadings.csv and SPS_PARAFAC_Sample_Scores.csv) are directly read into the associated R-markdown file. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Artificial intelligence models, photos, and data associated with the manuscript “Quantifying Streambed Grain Size, Uncertainty, and Hydrobiogeochemical Parameters Using Machine Learning Model YOLO” (v2)

This data package is associated with the manuscript “Quantifying Streambed Grain Size, Uncertainty, and Hydrobiogeochemical Parameters Using Machine Learning Model YOLO” published in Water Resources Research (Chen et al., 2024). This data package includes the training, validation, testing, and prediction data used by the artificial intelligence (AI) model for automated grain size and hydro-biogeochemistry quantification using streambed photos. The grain size data are extracted for each photo using You Look Only Once (YOLO), a pre-trained object detection model. This data package was originally published in October 2023. It was updated August 2025 (v2; new and modified files). File and folder names were not revised to indicate changes. See the change history section in the readme for more details. Please see flmd.csv for a list of all files contained in this data package and descriptions for each. Please see dd.csv for a data dictionary that defines the column headers of .csv files in the data package. This dataset is comprised of one data folder containing (1) file-level metadata; (2) data dictionary; (3) readme; and (4) six subfolders. Subfolders 1 to 4 include the training, validation, testing, and prediction data. Subfolder 5_Summary includes the summary results of different combinations of training, validation, testing, and prediction data. Subfolder 6_SupplementalData includes additional data downloaded from public sources (Kaufman et al., 2023a; Kaufman et al., 2023b; Garefalakis et al., 2023; Mair et al., 2024; https://github.com/river-corridors-sfa/Geospatial_variables). In total, the data package includes 110 folders and 44,283 files. These files include 9,047 .jpg photos, 1 .png photo, 3 .tif photos; 26,639 photo labels and individual grain sizes and probability from AI (.txt); 8,447 grain size distribution data (.dat); and 126 CSV files for results summary, and 14 required metadata files (.xlsx). The summary CSV files contain 68 columns and approximately 2,200 rows that represent photo names, site locations, recording time, GPS coordinates, grains sizes (D10, D50, D60, and D84), number of grains, and additional hydro-biogeochemical data such as water depth, flow velocity, Manning’s coefficient, friction factor, hydraulic conductivity, permeability, streambed interstitial velocity magnitude, mass transfer rate, and nitrate uptake velocity. The photos were obtained from 75 sites in the Yakima River Basin and the Columbia River shorelines, and other associated data from samples and sensors obtained when the photos were taken are publicly available (Fulton et al. 2022; Grieger et al. 2023). All files are .csv, .txt, .dat, .jpg, or .pdf. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Data, scripts, and figures associated with a manuscript studying impact of climate and topography on post-fire vegetation recovery.

This data package is associated with the publication “Impact of Topography and Climate on Post-fire Vegetation Recovery Across Different Burn Severity and Land Cover Types through Machine Learning” submitted to Remote Sensing of Environment (Zahura et al. 2023). In this research, a machine learning algorithm, random forest (RF), was utilized to examine the impact of climate and topography on post-fire vegetation recovery. We used enhanced vegetation index (EVI) to examine varying burn severity and land cover types. The data package includes the input files for RF model training, outputs from model predictions and analysis, and python scripts to run the model, analyze the results to understand model performance and interpretability, and plot manuscript figures. This data package contains three folders (Data, Scripts, and Figures), a file-level metadata (FLMD) csv, and a data dictionary (dd) csv. Please see Postfire_recovery_flmd.csv for a list of all files contained in this data package and descriptions for each. The data dictionary (Postfire_recovery_dd.csv) describes the csv column headers. The “Data” folder provides all the inputs and outputs to train the RF model, evaluate performance, and interpret predictions. The “Scripts” folder contains python scripts and jupyter notebooks for model training and result analysis. The “Figures” folder includes the figures used in the manuscript in “.png” and “.jpg” format.

54 ENVIRONMENTAL SCIENCES↗

Timeseries Unlabeled and Labeled Photos, Modeled Stream Elevation, and (Meta)Data of Variably Inundated Streams Across The Yakima River Basin, Washington, United States (v2)

This dataset is associated with the “River Monitoring Photos” (RMP) study and subsequent manuscript (Bao et al. 2025. Monitoring river flow status using low-cost wildlife camera and image segmentation artificial intelligence doi: 10.1016/j.envsoft.2025.106715). Game camera timeseries photos were collected to evaluate stream variable inundation via changes in width. A subset of photos was labeled for training the YOLOv8 and Mask2Former models and used to segment water surface fractions from all the game camera photos.This data package was originally published in March 2024. It was updated in October 2025 (v2) to add additional photos and files associated with the manuscript (i.e., processed data, labeled photos, and trained models). For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.In addition to a readme, this data package also includes two file-level metadata (FLMD) files that describes each file and two data dictionaries (DD) that describe all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) readme; (4) field metadata; and (5) folders containing game camera photos and manuscript-associated files. Each Yakima River Basin site has a folder that contains subfolders for each month photos were collected. There is also a folder for files associated with the manuscript which has subfolders for labeled data, trained models, Yakima River Basin site water surface fractions, and USGS site water surface fractions. All files are .csv, .json, .txt, .yaml, .pth, .pt, or .pdf. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Data associated with the manuscript “Radiative impact of record-breaking wildfires from integrated ground-based data” collected in Richland, Washington in September 2020

This data package is associated with the publication “Radiative impact of record-breaking wildfires from integrated ground-based data” submitted to Nature Scientific Reports (Kassianov et al., 2024). Data from ground-based measurements of shortwave and spectrally resolved irradiance and aerosol optical depth (AOD) in the visible and near-infrared spectral ranges were assessed to quantify the radiative impact of the September 2020 wildfires that occurred in the Western United States. Data were collected in September 2020 by several ground-based instruments at the Atmospheric Measurements Laboratory (AML) located in Richland, Washington (46.3451, -119.2792). These data include (1) Aerosol Optical Depth (AOD); (2) spectrally resolved and shortwave (SW) irradiances; (3) backscatter profiles; (4) total sky images; and (5) near-surface ambient air temperatures.The data package consists of five sub-directories: (1) “AML_Ceilometer_”; (2)” AML_CSPHOT_”; (3) “AML_MFRSR_irradiances_”; (4) “AML_SW_irradiances_and_Temp_”; (5) “AML_TSI_images_”; and 6 files stored at the directory level, including the readme, file-level metadata file, and data dictionary. The file-level metadata file (the file ending in “_flmd.csv”) lists all files contained in this data package and descriptions for each. The data dictionary (the file ending in “_dd.csv”) describes each tabular column header’s unit, definition, and structure. Below are descriptions of each sub-directory:“AML_Ceilometer_” includes ceilometer data collected at the AML. These files contain the corresponding narratives of data. Details related to the ceilometer data can be found in Morris (2016). “AML_CSPHOT_” includes ascii files with high-temporal resolution (about 10-15 min) AML CSPHOT data and their daily-averaged counterparts. These two files contain the corresponding narratives of data. Details related to the CSPHOT data can be found in Gregory (2011). “AML_MFRSR_irradiances_” includes ascii files with the AML MFRSR-measured diffuse, normal, and total spectrally resolved irradiance. Details related to the MFRSR data can be found in Hodges and Michalsky (2016) and Koontz et al. (2013). “AML_SW_irradiances_+_Temp_” includes near-surface ambient air temperature and SW irradiances, namely direct normal, diffuse hemispherical, and total hemispheric (global), measured at the AML. These files also incorporate the corresponding narratives of data. Details related to the SW irradiances can be found in Andreas et al. (2018). “AML_TSI_images_” includes Total Sky Images (TSIs) collected at the AML. Details related to the TSI data can be found in Morris (2005).

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with “Riverine dissolved organic matter transformations increase with watershed area, water residence time, and Damköhler numbers in nested watersheds” (v2)

This data package is associated with the publication “Riverine dissolved organic matter transformations increase with watershed area, water residence time, and Damköhler numbers in nested watersheds” submitted to Biogeochemistry by Ryan et al., 2024 (DOI: https://doi.org/10.1007/s10533-024-01169-5). This study aims to investigate fundamental and transferable drivers of dissolved organic matter (DOM) diversity across five nested watersheds within the contiguous United States. DOM diversity was explored using ultrahigh-resolution Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS). The samples and the unprocessed FTICR-MS data used in this study are publicly available on the Environmental System Science Data Infrastructure for a Virtual Ecosystem (ESS-DIVE) data repository (see DOIs below). The data for the Willamette, Gunnison, Connecticut, and Deschutes basins were collected as part of a collaboration between the Watershed Rules of Life (WROL) project and Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS). The data for the Yakima River basin (YRB) was collected by the PNNL River Corridor SFA. The raw, unprocessed FTICR-MS data with additional (meta)data can be found at doi:10.15485/1895159 for WROL samples and doi:10.15485/1898912 for YRB samples. This data package contains the processed data used in the associated manuscript. This package also contains ancillary geospatial, hydrological, and geochemical information that supports the interpretation of the FTICR-MS data within Ryan et al., 2024. This data package is associated with the GitHub repository found at https://github.com/WHONDRS-Hub/rcsfa-RC4-WROL-YRB_DOM_Diversity. This data package was originally published August 2024. It was updated January 2025 (modified files). See the change history in the readme more details. At the directory level, the data package is comprised of three folders: (1) data, (2) output, and (3) src; and five additional files including the data dictionary (file ending in "_dd.csv”) and file-level metadata (file ending in “_flmd.csv”). The “src” folder contains the scripts used to process the FTICR data, conduct the analyses, and produce the manuscript figures. The inputs for these scripts are in the “data” folder and the returned outputs in the “output” folder. Inputs include temporal and spatial metadata associated with the sampling efforts, processed FTICR data, and total and normalized putative biochemical transformations per sample. Outputs include cleaned and combined data presented as tables, descriptive statistics, and plots. The file-level metadata file lists all files contained in this data package and descriptions for each. The data dictionary describes the units and definitions for each tabular data column or row header.

54 ENVIRONMENTAL SCIENCES↗

Temporal Study 2022-2024: Sensor-Based Time Series of Surface Water Temperature, Specific Conductance, Total Dissolved Solids, Turbidity, Chlorophyll A, and Dissolved Oxygen from across Multiple Watersheds in the Yakima River Basin in Washington, USA

This dataset supports a broader study examining the drivers of temporal variability in sediment respiration rates in the Yakima River Basin. The dataset provides periodic (bi-weekly or monthly) in situ hydrological and water chemistry sensor data, handheld sensor water chemistry data, general environmental context photos, and field metadata collected at six sites across the Yakima River Basin in Washington, USA. Sample and sensor data from previous years (2021-2022) can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1898912 and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1892054, respectively. Related sample data from 2022-2024 are available at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2562910. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions This dataset contains a folder of environmental context photographs and videos and (1) file-level metadata; (2) data dictionary; (3) readme; (4) field metadata; (5) field protocols; (6) international generic sample number (IGSN) mapping file; (7) handheld sensor data; and (8) two sensor subfolders. Each sensor subfolder (BarotrollAtm and MantaRiverData) contains a subfolder containing sensor time series data and plots. The BarotrollAtm Data subfolder contains In Situ Rugged BaroTROLL sensor pressure and air temperature data. The MantaRiverData subfolder contains Eureka Manta+ 35B multisonde temperature, specific conductance, and chlorophyll A. All files are .csv, .pdf, .jpg, .jpeg, .mp4, .png, or .mov.

54 ENVIRONMENTAL SCIENCES↗

WHONDRS laboratory time series moisture manipulative experiment from soil core layers across eastern contiguous US: time series aerobic respiration, geochemistry, and aggregates

This dataset supports a broader study examining the effects of wetting and drying on soil layers across the eastern contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata. Samples were collected as part of a collaboration between WHONDRS (Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems; https://whondrs.pnnl.gov) and MONet (Molecular Observation Network; https://www.emsl.pnnl.gov/monet). The field samples (soil cores) were labeled as MEL_##_COR and subsequent subsamples begin with MEL_##. Additional subsamples were taken for the laboratory experiment and were labeled as EL_##. The labels from the MEL field samples and the EL subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EL_01 is a subsample from MEL_01). See the critical details section below for more details on sample naming and experimental design.For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions.This dataset is comprised of (1) a folder containing environmental context photos; (2) file-level metadata; (3) data dictionary; (4) field metadata; (5) readme; (6) international generic sample number (IGSN) mapping file; and (7) a subfolder with soil sample data from field samples and the incubation experiment. The sample data subfolder contains (1) effect size; (2) gravimetric moisture from field samples and incubation experiment; (3) respiration rates, raw dissolved oxygen values, and plots; (4) specific conductance, pH, and temperature from the incubation; (5) soil aggregates; (6) a summary containing median values of each data type for each treatment (wet and dry) in the incubation; (7) a summary containing averages for each data type of each soil layer; and (8) methods codes. All files are .csv, .pdf, .jpeg, or .jpg.

54 ENVIRONMENTAL SCIENCES↗

Terrestrial laser scanning data (Levels 0 and 1) for Pasoh, Malaysia, Sep 2024

This data package contains data from terrestrial laser scanning (TLS) at the Pasoh Forest Reserve, Malaysia. The Pasoh Forest Reserve is a facility of the Forest Research Institute Malaysia, and contains evergreen lowland dipterocarp forest. The Next-Generation Ecosystem Experiments Tropics (NGEE-Tropics) study areas at Pasoh were established to study how different species respond to climatic variation and soil water availability. Two study areas were chosen representing different topography and species. The TLS data archived here were collected to provide detailed, three-dimensional information about forest structure. Specifically, data were collected to allow tree-level characterization of woody structure and leaf area for 12 focal trees with FloraPulse and sap flux sensors, facilitating estimation of woody biomass and leaf area to allow upscaling of water content and transpiration data to the tree-level. Scan positions were not selected to provide consistent data for non-focal trees with the study areas. This data package contains the following data: - High-level files document further details of the campaign and data package: 1_CampaignSummary.csv provides details about the campaign and study site, 2_ScanAreasDetail.csv provides details about each separate scan area (groups of scans post-processed into a single point cloud), 3_TerrestrialLidarSensor.csv provides further technical details about the Riegl VZ-400i TLS sensor, TLS_CSV_dd.csv is a CSV Data Dictionary providing information about the fields in CSV files following the ESS-DIVE CSV File Formatting Guidelines Reporting Format, TLS_flmd.csv is a File Level Metadata file providing information about each file in the data package following the ESS-DIVE File Level Metadata Reporting Format, and README.txt is a text file describing the overall project and file structure. - Level 0 data are the raw data (.PROJ folders) as recorded by the Riegl VZ-400i TLS instrument before scan co-registration and post-processing with the Riegl's proprietary RiSCAN PRO software, which requires a license. - Level 1 data contain post-processed, co-registered data from each scan area. The "PointClouds" folder for each scan area contains a .las file with 1 cm resolution point cloud data exported from RiSCAN PRO. These are the main files likely to be of interest to most users and can be further processed with any software capable of manipulating .las files (e.g. Python, R CloudCompare). The "Project Information" folder contains log files from post-processing in RiSCAN PRO that may be of interest to users who want to see detailed records of post-processing, including all PDF reports generated by RiSCAN PRO. The "ScanPositions" folder contains information about the final position of all TLS scans, after post-processing, in multiple formats. The file ScanPositions_*.csv provides final geo-referenced scan positions, and the file SOP_backup_*.csv can be used in RiSCAN PRO to restore the co-registered scan positions if users wish to re-process raw data (Level 0 .PROJ folders) with RiSCAN PRO software (e.g., subsample to a different resolution, exclude a certain scan position, or apply different filters on reflectance or deviation values) without redoing time-consuming co-registration steps.

54 ENVIRONMENTAL SCIENCES↗

SPRUCE: Peat Core Sample Collection Metadata, Marcell Experimental Forest, Minnesota, August 2024

This data set contains metadata associated with peat core samples collected from the Spruce and Peatland Responses Under Changing Environments (SPRUCE) experiment in August 2024. This sample metadata contains no analytical results and is a reference for analytical datasets. To ensure accessibility and discoverability, each sample was assigned an International Generic Sample Number (IGSN), a persistent identifier, using System for Earth and Extraterrestrial Sample Registration (SESAR). These samples were used for downstream analysis by multiple teams of researchers the results of which will be reported separately. This dataset contains one data file in comma separate (.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma separate (.csv) format and a user guide in PDF (*.pdf) format. An aliquot of most samples is stored at Oak Ridge National Laboratory and may be available for further analysis. Access this collection event on SESAR https://doi.org/10.58052/IEJ9B00VQ. To inquire about obtaining archived samples for analysis, reach out using the Contact Sample Owner form located on the bottom of the landing page in SESAR. Note: Only dried and ground material from C Cores are available for new analysis.

Birkebak, Joshua [ORNL] (ORCID:0009000955611494)↗

Surface water and groundwater FTICR-MS, NPOC, and TN from nine wetlands and three upland wells at the Tanglewood Biological Station, Alabama

This dataset supports a broader study examining wetland hydrobiogeochemical responses to flood disturbance and the subsequent impacts on watershed nutrient export. The study was designed following ICON (integrated, coordinated, open, and networked) principles. Samples were collected from nine wetlands and three upland wells at the Tanglewood Biological Station, Alabama in August 2024 and February 2025, during the dry and wet season, respectively. The contents include geochemistry (dissolved organic carbon measured as non-purgeable organic carbon; total dissolved nitrogen) and organic matter characterization (FTICR-MS). Related water level data from the same locations can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/2530253. Additional geochemistry will be published in a separate data package. For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions.This dataset is comprised of (1) a folder containing environmental context photos; (2) file-level metadata; (3) data dictionary; (4) field metadata; (5) readme; (6) international generic sample number (IGSN) mapping file; (7) the field protocol; and (8) a subfolder with sample data. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) total nitrogen data and averages; (3) methods codes; and (4) a subfolder of 12 Tesla (12T) Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) data. All files are .csv, .pdf, .jpeg, or .jpg.

54 ENVIRONMENTAL SCIENCES↗

Pyrogenic Organic Matter Laboratory Experiment: Aerobic Respiration and Geochemistry from Variably Inundated Stream Sediments (v3)

This dataset supports a broader study examining the effects of variable inundation and pyrogenic organic matter on ecosystem respiration. The dataset provides data generated from a laboratory batch experiment investigating the interaction between variable inundation conditions (wet and dry sediment) and pyrogenic organic matter (burned and unburned treatments). The contents include time series dissolved oxygen, sediment geochemistry data, and field metadata (including qualitative information on instream and river corridor characteristics). This data package was originally published in November 2025. It was updated in April 2026 (v2; new and modified files) and May 2026 (v3; modified files). See the change history section in the readme for more details For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) international generic sample number (IGSN) mapping file; (5) readme; (6) field protocol; (7) sample name metadata; (8) an environmental context picture for the dry and inundated sampling locations; and (9) a subfolder with sample data from the sediment incubation experiment. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC); (2) total nitrogen (TN); (3) gravimetric moisture; (4) partial pressure and production rates of carbon dioxide, methane, and nitrous oxide; (5) field wet sediment mass, dry sediment mass, water mass, and field wet sediment volume in incubation and sediment NPOC/TN vials; (6) methods codes; (7) respiration rates, pH, and temperature from after the incubation, raw time series dissolved oxygen and temperature, and a subfolder containing associated plots and scripts; (8) ions; (9) FTICR-MS methods; and (10) a subfolder of 12 Tesla (12T) FTICR-MS data. This folder contains the CoreMS processed data and three subfolders, one containing the .xml files, one containing the CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .html, .Rmd, .py, .cal, .json, or .jpg.

54 ENVIRONMENTAL SCIENCES↗