Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Level 2 data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Pilot Heavy-Duty Electric Vehicle Deployment for Anchorage, Alaska, Municipal Solid Waste Collection

Through a grant awarded by the U.S. Department of Energy, the Municipality of Anchorage initiated a pilot program in their Solid Waste Services (SWS) department to add heavy-duty electric trucks to its vehicle fleet. The project involves the purchase and deployment of a Peterbilt 220EV electric box truck and two Peterbilt 520EV heavy-duty electric refuse trucks. The Alaska Center for Energy and Power at the University of Alaska Fairbanks performed data analysis. Data collected include telemetry data from both types of electric trucks, charging data from the 520EV telemetry data and a Level 2 charger, and facility-level electric use data.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Pilot Heavy-Duty Electric Vehicle Deployment for Anchorage, Alaska, Municipal Solid Waste Collection

Through a grant awarded by the U.S. Department of Energy, the Municipality of Anchorage initiated a pilot program in their Solid Waste Services (SWS) department to add heavy-duty electric trucks to its vehicle fleet. The project involves the purchase and deployment of a Peterbilt 220EV electric box truck and two Peterbilt 520EV heavy-duty electric refuse trucks. The Alaska Center for Energy and Power at the University of Alaska Fairbanks performed data analysis. Data collected include telemetry data from both types of electric trucks, charging data from the 520EV telemetry data and a Level 2 charger, and facility-level electric use data.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Data and scripts associated with “Point-scale organic-matter decomposition in streambeds is weakly associated with reach-scale respiration”

This data package is associated with “Point-scale organic-matter decomposition in streambeds is weakly associated with reach-scale respiration” published in EGU Biogeosciences (Stegen et al., 2026; https://doi.org/10.5194/bg-23-3981-2026). It contains cotton strip decomposition rates (Kcd and Kdd) collected across the Yakima River Basin (YRB), Washington, USA. These data were collected to support a broader study examining the drivers of spatial variability in sediment respiration rates in the Yakima River Basin. Associated data used in analysis, metadata, and field protocols can be accessed at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689, https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1969566, and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1987520. This data package is associated with the repository found at https://github.com/river-corridors-sfa/rcsfa-ST-2B-SSS-cotton-strip. A preliminary version of this data package was published in December 2025 at the time of manuscript submission. It was updated in June 2026, at the time of manuscript acceptance, to include additional metadata (this readme, data dictionary, and file level metadata). The data did not change. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This data package consists of (1) readme; (2) data dictionary (dd); (3) file level metadata (flmd); and (4) four folders: (1) R-scripts; (2) figures; (3) outputs from the scripts; and (4) published data. The published data folder contains a readme directing the user to download data in order to run the R-scripts. All files are .csv, .pdf, .R, .Rmd, and .txt. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES

Levoglucosan data from five coastal streams impacted by the 2020 CZU Lightning Complex Fires, California, United States

This dataset includes levoglucosan data for five coastal California (United States) streams impacted by the 2020 CZU Lightning Complex Fires which burned from August 16th through September 22nd. Levoglucosan is a highly soluble and biolabile fraction of pyrogenic carbon. The five watersheds (San Lorenzo River, Pescadero Creek, Majors Creek, Laguna Creek, and Scott Creek) were impacted by the fires with watersheds experiencing a range of burn severity and extents. Grab samples were collected from each stream between October 2020 and May 2021, targeting both baseflow and event flow hydrologic conditions. Additional biogeochemistry data (i.e., organic and black carbon concentrations) can be found in a separate data package (https://doi.org/10.4211/hs.421c0226bb38460c8393d67fe0c4f802). This data package consists of one main data folder that contains (1) readme; (2) file-level metadata; (3) data dictionary; (4) field metadata with international generic sample numbers (IGSN); (5) methods codes; and (6) levoglucosan data. All files are .csv or .pdf.

2020 CZU Lightning Complex Fires

Fitness For Service Assessment of a Corroded Heat Exchanger

Within the Fermi National Accelerator complex, there exist various water systems that support accelerator operations. One of these systems is extremely vital to the operation of the machine; that is the cooling system. The cooling system consists of nine relatively large heat exchangers that take untreated pond water and use it to cool the process fluid that further cools machine components. Over the 30 years these heat exchangers have been in operation, they have undergone significant material loss on the channels. This material loss, due to various forms of corrosion such as galvanic and microbiologically influenced corrosion (MIC) and possibly others, has deteriorated more than 80% of the nominal wall thickness of some of the exchangers and placed them in a questionable state. ASME FFS-1 (API 579) has been applied to address the condition of the heat exchangers due to their noncompliance with the governing code, BPVC Sec. VIII Div. 1. The assessments encompassed ASME FFS-1 parts 4: General Metal Loss and 9: Crack Like Flaw using level 1, 2, and 3 analysis techniques based on inspection data obtained by API 510 inspections. Level 1 and 2 assessments were deemed unfit for the corroded regions due to their location relative to a major structural discontinuity (channel to tube-sheet joint), so a level 3 analysis was conducted according to ASME Sec. VIII Div. 2 (design by analysis) rules for pressure vessels. Supplemental information included pond water tests to determine an accurate future corrosion allowance due to lacking inspection history. A leak before break (LBB) route was chosen to evaluate the possibility of leaking prior to the onset of failure. The analysis of one heat exchanger shows that the possibility the channel will develop a pinhole leak over 2.5 more years of operation should not be overlooked, but burst was unlikely from operation. The use of fracture mechanics show, that if a through-wall crack were to develop, it would not propagate further than the channel geometry and cause a leak not greater than 35 GPM. Using ASME Section XI Code Case N-705-1, allowing us to operate with a leak until the next outage given certain operating conditions and developing a leak mitigation procedure, this heat exchanger is deemed fit-for-service.

Humenik, Alex [Fermilab]

WHONDRS laboratory time series moisture manipulative experiment from soil core layers across eastern contiguous US: time series aerobic respiration, geochemistry, and aggregates

This dataset supports a broader study examining the effects of wetting and drying on soil layers across the eastern contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata. Samples were collected as part of a collaboration between WHONDRS (Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems; https://whondrs.pnnl.gov) and MONet (Molecular Observation Network; https://www.emsl.pnnl.gov/monet). The field samples (soil cores) were labeled as MEL_##_COR and subsequent subsamples begin with MEL_##. Additional subsamples were taken for the laboratory experiment and were labeled as EL_##. The labels from the MEL field samples and the EL subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EL_01 is a subsample from MEL_01). See the critical details section below for more details on sample naming and experimental design.For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions.This dataset is comprised of (1) a folder containing environmental context photos; (2) file-level metadata; (3) data dictionary; (4) field metadata; (5) readme; (6) international generic sample number (IGSN) mapping file; and (7) a subfolder with soil sample data from field samples and the incubation experiment. The sample data subfolder contains (1) effect size; (2) gravimetric moisture from field samples and incubation experiment; (3) respiration rates, raw dissolved oxygen values, and plots; (4) specific conductance, pH, and temperature from the incubation; (5) soil aggregates; (6) a summary containing median values of each data type for each treatment (wet and dry) in the incubation; (7) a summary containing averages for each data type of each soil layer; and (8) methods codes. All files are .csv, .pdf, .jpeg, or .jpg.

54 ENVIRONMENTAL SCIENCES

Surface water and groundwater FTICR-MS, NPOC, and TN from nine wetlands and three upland wells at the Tanglewood Biological Station, Alabama

This dataset supports a broader study examining wetland hydrobiogeochemical responses to flood disturbance and the subsequent impacts on watershed nutrient export. The study was designed following ICON (integrated, coordinated, open, and networked) principles. Samples were collected from nine wetlands and three upland wells at the Tanglewood Biological Station, Alabama in August 2024 and February 2025, during the dry and wet season, respectively. The contents include geochemistry (dissolved organic carbon measured as non-purgeable organic carbon; total dissolved nitrogen) and organic matter characterization (FTICR-MS). Related water level data from the same locations can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/2530253. Additional geochemistry will be published in a separate data package. For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions.This dataset is comprised of (1) a folder containing environmental context photos; (2) file-level metadata; (3) data dictionary; (4) field metadata; (5) readme; (6) international generic sample number (IGSN) mapping file; (7) the field protocol; and (8) a subfolder with sample data. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) total nitrogen data and averages; (3) methods codes; and (4) a subfolder of 12 Tesla (12T) Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) data. All files are .csv, .pdf, .jpeg, or .jpg.

54 ENVIRONMENTAL SCIENCES

FATHOMS-RAG: A Framework for the Assessment of Thinking and Observation in Multimodal Systems that use Retrieval Augmented Generation

Retrieval-augmented generation (RAG) has emerged as a promising paradigm for improving factual accuracy in large language models (LLMs). We introduce a benchmark designed to evaluate RAG pipelines as a whole, evaluating a pipelines ability to ingest several modalities of information. We present (1) a curated dataset of 93 questions designed to evaluate a pipeline's ability to ingest textual data, tables, images, multimodal data, and cross-document multimodal data; (2) a phrase-level recall metric for correctness; (3) a nearest-neighbor embedding classifier in an attempt to classify pipeline hallucinations; (4) a comparative evaluation of 2 pipelines built with open-source retrieval mechanisms and 4 closed-source foundational models; and (5) a third-party human evaluation of the alignment of our correctness and hallucination metrics. We find that closed-source pipelines significantly outperform open-source pipelines in both the correctness and halucination metrics, with a wider performance gap in questions relying on multimodal and cross-document information. We also find after a human evaluation of our correctness and hallucination metric compared with our questions and pipeline responses, average agreement was 4.62 for correctness 4.53 for hallucination detection on a 1-5 Likert scale with 5 being strongly agree with our determination.

Hildebrand, Samuel [ORNL] (ORCID:0009000465963104)

Investigating the opioid epidemic across the United States: Associations between county-level characteristics and overdose mortality

The opioid crisis remains a critical public health challenge in the United States. Despite national efforts that reduced opioid prescribing by nearly 44% between 2011 and 2021, opioid overdose deaths more than tripled during the same period. This alarming trend reflects a major shift in the crisis, with illegal opioids now driving the majority of overdose deaths instead of prescription opioids. Although supply-side factors fueling this transition have been widely studied, the structural and community-level conditions that shape overdose mortality are less well understood. To help address this gap, this study has three primary objectives: (1) overcome structural gaps in national data to construct a complete nationwide county-level dataset from 2010 to 2022; (2) using data analysis, identify and investigate spatiotemporal anomalies in overdose mortality; and (3) using two machine-learning models, quantify the importance of thirteen social vulnerability variables in predicting overdose mortality. Our results identify unemployment and limited vehicle access as key county-level predictors of overdose mortality. Higher levels of these vulnerabilities are associated with elevated mortality, whereas lower levels are associated with reduced mortality. These findings highlight factors that may be relevant for public health planning and policy prioritization within the context of the opioid crisis.

Anomaly analysis

WHONDRS 2016 Sediment Organic Matter Characterization Data from Streams across HJ Andrews Experimental Forest, Oregon

This dataset supports a broader synoptic effort to map morphological, hydrological, chemical, and biological conditions across a fifth-order mountain stream network. Samples were generated through a collaborative synoptic sampling effort in 2016. The dataset provides sediment Fourier Transform Ion Cyclotron Resonance Mass Spectrometry (FTICR-MS) from 60 sites across the HJ Andrews Experimental Forest, Oregon (https://andrewsforest.oregonstate.edu). Related data were collected as part of the event and were published separately in collaboration with other team members. The data are available at http://www.hydroshare.org/resource/ea6c0832885a46c3939e7bb22e48e754 and are described within https://doi.org/10.5194/essd-11-1567-2019 (Ward et al., 2019). The hydroshare data package contains processed FTICR-MS data from the samples included in this data package. The data were processed via Formultitude (previously called Formularity; https://github.com/PNNL-Comp-Mass-Spec/Formultitude). However, we have re-processed the data using Core-MS and included it in this data package. Additional related data collected in 2025 from a similar effort can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3023310 and http://www.hydroshare.org/resource/b274c4a234bf4b12b7cb8a54a696c629. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) a folder of sample data; (2) data dictionary; (3) file-level metadata; (4); (5) coordinates; and (6) readme. The sample data subfolder contains 12 Tesla (12T) FTICR-MS data. This folder contains the processed data and three subfolders, one containing the .xml files, one containing the CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .Rmd, .py, .cal, or .json.

Biogeochemistry

Direct Feed High-Level Waste APPS Model Glass Testing (DFHLW APPS) Matrix, Phase 2

This report summarizes the data collected during the batching and melting of a second matrix of Direct Feed High-Level Waste (DFHLW) glasses generated using the preliminary enhanced waste glass models (EWG2.5) and the Britton and Anderson (2024) preliminary DFHLW feed vector. The purpose of these glasses is two-fold: 1. Validate EWG2.5 glass calculations being used in the Aspen Process Performance Simulation (APPS) model. 2. Evaluate and ultimately improve the glass property models and formulation methods used for design of DFHLW glasses as part of an iterative process of data collection and model refinement. Some of the 16 APPS2 glasses tested did not satisfy all target property constraints due to the limited data on DFHLW glass supporting the EWG2.5 models. • One glass, APPS2-10, formed nepheline on canister centerline cooling (CCC) heat-treatment and failed the product consistency test (PCT) response limits. This glass also had high B and Cr release rates for the toxicity characteristic leaching procedure (TCLP). All other glasses were found to satisfy the PCT and TCLP constraints for both quenched and CCC samples. • One glass, APPS2-08, had higher than acceptable viscosity due to magnetite crystallization. • One glass, APPS2-09, formed greater than 2 vol% crystals at 950 °C. As the glass design criterion was that the temperature at 2 vol% crystal (T 2% ) be less than 950 °C, only one glass failed the criteria. However, this criterion is being reevaluated. Four additional glasses formed crystal fractions between 1 and 2 vol% at 950 °C (APPS2-03, -08, -12, and -14). • Four glasses – APPS2-01, -02, -04, and -16 – failed the Monofrax K-3 refractory neck corrosion (k neck ) design limit of 0.04 in. at 1208 °C for 6 d. This is another criterion being reevaluated. Four additional glasses (APPS2-05, -06, -11, and -13) exhibited 0.025 = k neck = 0.04 in. • All 16 glasses passed the sulfur solubility and TCLP constraints. The measured property values were compared to predicted values using EWG2.5 and a selection of other existing models. A few models (e.g., electrical conductivity, TCLP) were found to be adequate for designing DFHLW glasses in the near future, while others require refits or offsets. It is recommended that new property models be developed for EWG3.0, as a large amount of DFHLW glass property data (> 14 × existing data) is expected to be collected in the compositional spaces where no data was previously available. To enable near-term calculations and formulations for designing DFHLW glasses and processing rate estimations, a formulation algorithm with minor modifications will be developed, EWG2.6.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

Recent Experience with the CMS Data Management System

The CMS[1] experiment manages a large-scale data infrastructure, currently handling over 200 PB of disk and 500 PB of tape storage and transferring more than 1 PB of data per day on average between various WLCG[2] sites. Utilizing Rucio[3] for high-level data management, FTS[4] for data transfers, and a variety of storage and network technologies at the sites, CMS confronts inevitable challenges due to the system’s growing scale and evolving nature. Key challenges include managing transfer and storage failures, optimizing data distribution across different storages based on production and analysis needs, implementing necessary technology upgrades and migrations, and efficiently handling user requests. The data management team has established comprehensive monitoring to supervise this system and has successfully addressed many of these challenges. The team’s efforts aim to ensure data availability and protection, minimize failures and manual interventions, maximize transfer throughput and resource utilization, and provide reliable user support. This paper details the operational experience of CMS with its data management system in recent years, focusing on the encountered challenges, the effective strategies employed to overcome them and the ongoing challenges as we prepare for future demands.

Öztürk, Hasan [CERN]

CMB low multipole alignments across WMAP and Planck data releases

ABSTRACT The first observations of the cosmic microwave background (CMB) from NASA's Wilkinson Microwave Anisotropy Probe (WMAP) led to finding ‘alignment’ anomalies not expected from fluctuations in the isotropic cosmological model. We study the data of all 8 full-sky public releases since then to test for anomalous alignments and shapes of the first 60 multipoles, i.e. over the range $2\le l \le 61$. We use rotationally invariant and covariant statistics to test isotropy of all subsequent WMAP data releases, along with those from the ESA’s Planck mission. Anomalous alignments among the multipoles $l=1, 2, 3$ are very consistent and robust. More alignments are detected, some of them new, while significance is diluted by the large range of the search. Power entropy, a measure of the randomness of the multipoles, is consistently anomalous at about $2\sigma$ level or better across all data releases. It appears that the CMB is not as random as the cosmological principle predicts on large angular scales.

Patel, Sanjeet Kumar

Data-driven Community-centered Resilient Assessment and Planning Toolkit for Nexus of Energy and Water (DCRAPT-NEW)

Urban areas, including Detroit and Pittsburgh, have suffered significant dual outages of the electrical and water infrastructure in the past decade due, in part, to the increasing number of extreme weather events. With increasing temperatures and rainfall intensity, these regions need to prepare for increasing extreme events through community-based energy and water resilience analysis, planning, and enhancement. This project developed a suite of open-source, open-access, community-centered, data-driven assessment and distributed energy resource (DER) and planning tools for energy and water resilience enhancement in urban areas. Through establishing a multi-level community awareness and engagement mechanism and a comprehensive collection of power outage and flooding data, an innovative group of community energy and water resilience assessment and planning tools have been developed for a wide range of users with differing and variable sets of data available to them. The developed tools include (1) DOE EAGLE-I data-driven, deep-learning assisted resilience assessment and DER planning tools at the county level with socioeconomic factors incorporated; (2) Utility annual power outage data-driven tools for long term resilience assessment and DER planning and 15-min power outage data-driven tools for short term resilience assessment and planning; (3) Detailed engineering tools for energy and water systems resilience assessment and planning when the system topology and component fragility curves are available; (4) Alternative Resiliency Metric Calculation that extracts and separates outage and restoration processes; and (5) Co-optimization tools that evaluate the resilience of the power and sewage system and allow users to conduct joint planning with energy and wastewater systems. The developed tools provide planners, decision-makers, and stakeholders with powerful capabilities to systematically evaluate system/community resilience and optimal and actionable guidance for enhancing resilience while prioritizing DER investments. The tools have been used and validated in Detroit and Pittsburgh and can be used in other areas of the nation. In addition, this project will (1) advance the knowledge and applications of machine-learning methods in analyzing and fusing different layers of information and generating meaningful data points such as generating rare weather events; (2) significantly improve the energy and water resilience of the identified communities in Detroit and Pittsburgh and prepare for more frequent and severe weather conditions; (3) help communities assess extreme weather event impacts and address short-term and long-term resilience-related issues The developed tools have been made public via GitHub and demonstrated to community stakeholders and utility companies via the two annual workshops and numerous community engagement meetings. The project outcomes are also disseminated through publications in various journals and conference proceedings, and presentations at top conferences.

13 HYDRO ENERGY

Data and scripts associated with a manuscript modeling microbial regulation of priming effects

This data package is associated with the publication “Modeling Microbial Regulatory Feedback in Organic Matter Decomposition Identifies Copiotrophic Traits as Key Drivers of Positive Priming” published as a preprint on BioRXiv by Ahamed et al. (2026); https://doi.org/10.1101/2024.08.11.607483. The package contains MATLAB scripts and saved simulation outputs used to implement a cybernetic model of microbial regulation during complex organic matter (OM) decomposition governing priming effects. It includes models of (i) single microbial functional groups (copiotrophic or oligotrophic degraders) and (ii) binary consortia composed of degraders and non-degraders with contrasting or common growth traits. Simulation results were generated using Monte Carlo analyses, with randomized key model parameters across a range of environmental mixing fractions of complex and labile OM. The dataset was created to provide a transparent and reusable computational framework for systematically exploring how microbial growth traits, metabolic regulation, and community composition influence OM decomposition dynamics and priming effects. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes the variable definitions. This package includes: (1) annotated MATLAB code implementing the system of ordinary differential equations and cybernetic control laws; (2) saved output files containing data (e.g., biomass, substrates, enzyme levels, priming metrics); and (3) scripts for processing saved outputs and regenerating figures. Specifically, the data package contains three main MATLAB scripts: runPrimingModel.m, runPlotData.m, and runPlotSuppFigS1.m, along with this readme and supporting documentation. Users should begin with runPrimingModel.m, which contains the annotated code implementing the system of ordinary differential equations and cybernetic control laws. This script runs the Monte Carlo simulations of microbial OM decomposition and allows users to modify microbial trait definitions, adjust parameter distributions, or define new community configurations. Simulation outputs are automatically saved as .mat files in the folder named SavedData, which stores all pre-generated results included in this package. The second script, runPlotData.m, reads files from the SavedData folder and processes them to regenerate the figures presented in the manuscript. The third script, runPlotSuppFigS1.m, specifically generates Figure S1 in the Supplementary Material of the manuscript. The package also includes the aforementioned files in non-proprietary .txt format. If users intend to use them, they should first save the files in their respective .m or .mat formats prior to execution in MATLAB.

Biomass concentration

Rural EVSE Planning and Analysis

The dataset includes detailed anonymized public charging station usage from several rural stations on the ChargePoint and Shell Recharge Solutions (formerly Greenlots) networks situated in and around Athens, Ohio, a rural Appalachian community. Both Level 2 and DC fast charging stations are represented. Historical data in the set date back to 2019; additional data will be uploaded semiannually until the project's completion in 2023. Each charging session recorded includes information on date and time, location, charging station level, session duration, energy delivered, and fuel savings.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Data and scripts associated with “Moisture content modulates DOM thermodynamic regulation of oxygen consumption in drying streambed sediments”

This data package is associated with the publication “Moisture content modulates DOM thermodynamic regulation of oxygen consumption in drying streambed sediments” published in Scientific Reports (Garayburu-Caruso et al., 2026). The package contains processed data products and scripts used to quantify how drying and re-inundation of riverbed sediments influence dissolved organic matter (DOM) thermodynamic properties and their relationship with sediment oxygen (O₂) consumption across 33 stream sites in the contiguous United States. The data package contains DOM thermodynamic metrics (e.g., Gibbs free energy of carbon oxidation and thermodynamic efficiency), and O₂ consumption along with watershed-scale climate and land-cover metrics used as explanatory variables in the analyses. Underlying unprocessed and processed ultrahigh-resolution mass spectrometry data, oxygen consumption rates from laboratory moisture-manipulation experiments, within-sample environmental properties, sediment moisture content and contextual field measurements are archived separately at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2428003 (Laan et al., 2024) and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689 (Forbes et al.,2023). A preliminary version of this data package was published in February 2026 at the time of manuscript submission. It was updated in June 2026, at the time of manuscript acceptance, to include the finalized data and additional metadata (readme, data dictionary, and file level metadata). For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. At the top level, the data package is organized into five main folders: (1) Data, (2)Figures, (3) Map, (4) GAM_Reulsts, and (5) src. The Data folder contains analysis-ready tabular files with oxygen consumption rates, DOM thermodynamic properties by site and treatment, site-level environmental variables, watershed-scale metrics, and other derived variables referenced in the manuscript. The Figures folder contains static image files associated with the main text and supplemental figures, while the Map folder includes spatial data and map-layer files used to create the sampling-location map. The GAM results folder contains the results for each of the general additive model (GAM).The src folder contains R scripts used to perform data processing, statistical analyses (including clustering, generalized additive models, and threshold analysis), and figure generation. This data package is associated with a GitHub repository found at https://github.com/WHONDRS-Hub/ECA_DOM_Thermodynamics.

Dissolved organic matter