Engineering Papers⌕ Search

Engineering topics

Scheibe, Timothy D.

Publications and source records attributed to Scheibe, Timothy D..

Laboratory time series moisture manipulative experiment from sediment across the contiguous US: time series aerobic respiration and geochemistry (v2)

This dataset supports a broader study examining the effects of wetting and drying on hyporheic zone respiration across the contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata (including qualitative information on instream and river corridor characteristics). Samples were collected as part of the WHONDRS CONUS-Scale Model-Sample Study (CM). This study was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. The data package associated with the CM study is available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689. CM sampling began in April 2022 and ended in October 2023. This study uses subsamples from a subset of CM samples collected between June 2022 and June 2023. The original field samples were labeled as CM_###. Subsequent subsamples for this study were labeled as EC_###. The labels from the field samples and the EC subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EC_001 is a subsample from CM_001). See the critical details section below for more details on sample naming. This data package was originally published in August 2024. It was updated in February 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) field protocol; and a (6) a subfolder with sediment sample data from the incubation experiment. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC); (2) total nitrogen (TN); (3) adenosine triphosphate (ATP); (4) percent carbon and nitrogen; (5) effect size; (6) iron (II); (7) gravimetric moisture; (8) respiration rates and raw dissolved oxygen values; (9) specific conductance; (10) pH; (11) temperature; (12) a summary containing median values of each data type for each treatment (wet and dry); (13) methods codes; (14) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla FTICR-MS data. This folder contains three subfolders, one containing the sediment .xml data files, one containing the sediment CoreMS output files, the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .ref, or .xml.

54 ENVIRONMENTAL SCIENCES↗

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES↗

Sinuosity‐Driven Hyporheic Exchange: Hydrodynamics and Biogeochemical Potentials

Abstract Hydrologic exchange processes are critical for ecosystem services along river corridors. Meandering contributes to this exchange by driving channel water, solutes, and energy through the surrounding alluvium, a process called sinuosity‐driven hyporheic exchange. This exchange is embedded within and modulated by the regional groundwater flow (RGF), which compresses the hyporheic zone and potentially diminishes its overall impact. Quantifying the role of sinuosity‐driven hyporheic exchange at the reach‐to‐watershed scale requires a mechanistic understanding of the interplay between drivers (meander planform) and modulators (RGF) and its implications for biogeochemical transformations. Here, we use a 2D, vertically integrated numerical model for flow, transport, and reaction to analyze sinuosity‐driven hyporheic exchange systematically. Using this model, we propose a dimensionless framework to explore the role of meander planform and RGF in hydrodynamics and how they constrain nitrogen cycling. Our results highlight the importance of meander topology for water flow and age. We demonstrate how the meander neck induces a shielding effect that protects the hyporheic zone against RGF, imposing a physical constraint on biogeochemical transformations. Furthermore, we explore the conditions when a meander acts as a net nitrogen source or sink. This transition in the net biogeochemical potential is described by a handful of dimensionless physical and biogeochemical parameters that can be measured or constrained from literature and remote sensing. This work provides a new physically based model that quantifies sinuosity‐driven hyporheic exchange and biogeochemical reactions, a critical step toward their representation in water quality models and the design and assessment of river restoration strategies.

54 ENVIRONMENTAL SCIENCES↗

Thermodynamic control on the decomposition of organic matter across different electron acceptors

The increasing availability of high-resolution characterization of natural organic matter (OM) data has shifted the paradigm of lumped descriptions of OM components and potential microbial activities. Our recent development of a substrate-explicit thermodynamic model uniquely enables incorporating complex OM pools to formulate biogeochemical reaction models based on their elemental compositions. While this previous work facilitates prediction of aerobic respiration of complex OM, it is equally imperative to consider the role of non-oxygenic electron acceptors in regulating OM turnover and the fate of carbon. In this study, we significantly expand our previous model by flexibly incorporating both detailed OM chemistry and electron acceptors other than oxygen. Here, our modeling analysis has revealed substantial variations in the energy status of OM molecules across different soils, which drive the co-occurrence of different electron-accepting processes. We demonstrated the effectiveness of the proposed model using a consistency check with experimental data. Through systematic evaluation of the impact of diverse chemical inputs (both electron donors and acceptors) on OM decomposition, the new model also revealed how key microbial growth parameters such as carbon use efficiency (CUE) and reaction rates vary across different electron-accepting processes. Our model provides a unified framework integrating thermodynamic and kinetic constraints on microbial metabolic activities. It complements traditional kinetic models, which are often designed solely to capture mass fluxes. We conclude that thermodynamic modeling emerges as a powerful tool for describing the mechanisms underlying the interplay between microbial growth and OM chemistry and cycling across different electron acceptors, enhancing our ability to project complex ecosystem behaviors in dynamic environments.

59 BASIC BIOLOGICAL SCIENCES↗

Model associated with: "Thermodynamic control on the decomposition of organic matter across different electron acceptors"

This model data package is associated with the publication “Thermodynamic control on the decomposition of organic matter across different electron acceptors” submitted to Soil Biology and Biochemistry (Zheng et al., 2023; https://doi.org/10.1016/j.soilbio.2024.109364).In this research, a thermodynamic modeling framework is built to flexibly incorporate both organic matter (OM) molecules and electron acceptors for estimating potential free energy release from various redox reactions and to further predict reaction rates based on Microbial Transition State Theory. The model package includes scripts for thermodynamic modeling and postprocessing. Input Fourier-transform ion cyclotron resonance (FTICR) data are from a previous experimental study (Boye et al., 2018), and model outputs are free energy predictions and stoichiometric coefficients associated with all possible redox reactions.This data package is associated with the project GitHub repository found at MM_bioenergetic_modeling.This data package contains four folders (Input_FTICR, Model, Output, and Output_processing), a file-level metadata (FLMD) csv, and a data dictionary (dd) csv. Please see Zheng_bioenergetic_modeling_flmd.csv for a list of all files contained in this data package and descriptions for each. The Zheng_bioenergetic_modeling_dd.csv file describes the csv column headers. The “Model” folder contains scripts to run energy balance calculations for each electron acceptor. The “Output” folder contains csv files with stoichiometric information from model simulations. And the "Output_processing" folder contains scripts for reaction rate calculations and to generate plots.

54 ENVIRONMENTAL SCIENCES↗

Models, data, and scripts associated with “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning”

This data package is associated with the publication “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning’’ submitted to the Journal of Geophysical Research: Machine Learning and Computation (Scheibe et al. 2024). River sediment respiration observations are expensive and labor intensive to obtain and there is no physical model for predicting this quantity. The Worldwide Hydrobiogeochemisty Observation Network for Dynamic River Systems (WHONDRS) observational data set (Goldman et al.; 2020) is used to train machine learning (ML) models to predict respiration rates at unsampled sites. This repository archives training data, ML models, predictions, and model evaluation results for the purposes of reproducibility of the results in the associated manuscript and community reuse of the ML models trained in this project. One of the key challenges in this work was to find an optimum configuration for machine learning models to work with this feature-rich (i.e. 100+ possible input variables) data set. Here, we used a two-tiered approach to managing the analysis of this complex data set: 1) a stacked ensemble of ML models that can automatically optimize hyperparameters to accelerate the process of model selection and tuning and 2) feature permutation importance to iteratively select the most important features (i.e. inputs) to the ML models. The major elements of this ML workflow are modular, portable, open, and cloud-based, thus making this implementation a potential template for other applications. This data package is associated with the GitHub repository found at Please see the file level metadata (flmd; “sl-archive-whondrs_flmd.csv”) for a list of all files contained in this data package and descriptions for each. Please see the data dictionary (dd; “sl-archive-whondrs_dd.csv”) for a list of all column headers contained within comma separated value (csv) files in this data package and descriptions for each. The GitHub repository is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning models trained on the data in “input_data”; (3) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; (4) “examples” contains the visualization of the results in this repository including plotting scripts for the manuscript (e.g., model evaluation, FPI results) and scripts for running predictions with the ML models (i.e., reusing the trained ML models); (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. Furthermore, depending on the number of features used to train the ML models, the preprocessing and postprocessing scripts, and their intermediate results, can also be different branch-to-branch. The “main-*” branches are meant to be starting points (i.e. trunks) for each model branch (i.e. sprouts). Please see the Branch Navigation section in the top-level README.md in the GitHub repository for more details. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please the top-level README.md in the GitHub repository for more details on the automation.

13C↗

Data and Scripts Associated with the Manuscript “Water Column Respiration in the Yakima River Basin is Explained by Temperature, Nutrients and Suspended Solids”

This data package is associated with the publication “Water Column Respiration in the Yakima River Basin is Explained by Temperature, Nutrients and Suspended Solids” published in EGU Biogeochemistry (Laan et al. 2025). In this research, water column respiration (ERwc) data, surface water chemistry data, organic matter (OM) chemistry data, and publicly available geospatial data were used in analysis to evaluate the variability in ERwc at 47 sites across the Yakima River basin in Washington, USA. In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. The data package includes the data inputs, and outputs, and R scripts to reproduce all the analyses performed in the manuscript and create manuscript figures. The data package is comprised of three main folders (Code, Data, and Figures). The Code folder is comprised of four scripts and three analysis-specific subfolders that contain the R scripts to perform the analyses described in the publication and create publication figures. The Data folder is comprised of two “.csv” files and four subfolders that contain data input and output files. The Published_Data folder contains a readme that directs the user to download the appropriate files and add to this folder when using scripts. The Figures folder includes figures from the manuscript in “.pdf” and “.png” formats and a folder with intermediate figure files. This data package is associated with a GitHub repository which can be found at https://github.com/river-corridors-sfa/rcsfa-RC2-SPS-ERwc. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Data and scripts associated with: “Burn severity and vegetation type control phosphorus concentration, molecular composition, and mobilization”

This data package is associated with the publication “Burn severity and vegetation type control phosphorus concentration, molecular composition, and mobilization” published in European Geophysical Union - Biogeosciences (Barnes et al. 2025). This study investigates how phosphorus (P) biogeochemistry is altered by burn severity in contrasting types of vegetation chars. This data package documents the workflow used to process and generate the main figures and statistics in the manuscript. The R scripts reference minimally processed P nuclear magnetic resonance (P-NMR) and X-ray absorption near edge structure (P-XANES) data, as well as fully processed data including total elemental composition of the solid chars, total elemental composition of the char leachates (particulate and aqueous phases), and leachate aqueous phase molybdate reactive P concentration. These source data and associated metadata can be found on ESS-DIVE at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1894135 (Grieger et al. 2022; v3). Files and scripts included in this data package finish the processing workflow for P-NMR and P-XANES data. These data can be used to gain a better understanding of bulk chemical changes in chars and their leachates, as well as detailed molecular changes to P. This data package is associated with the GitHub repository found at https://github.com/river-corridors-sfa/rcsfa-RC3-BSLE_P. This data package is comprised of a “data” folder and a series of data processing and analysis scripts. Details on how to recreate the workflow can be found in the Critical Details section of the readme and the “workflow_readme.md” file. The file-level metadata file (file ending in “flmd.csv”) lists all files contained in this data package and descriptions for each. The data dictionary (file ending in “dd.csv”) describes all tabular data columns and their respective definitions and units.

54 ENVIRONMENTAL SCIENCES↗

Schneider Springs Fire Study 2023 for Ecosystem Respiration Rates: Surface Water Chemistry and Hydrologic Sensor Data across the Yakima River Basin, Washington, USA (v2)

This dataset supports a broader study examining the drivers of spatial variability in wildfire impacts across the Yakima River Basin. Data provided within this dataset were generated from sample collection across 17 total sites (8 sites affected by a recent wildfire, 9 sites unaffected by a recent wildfire) within multiple rivers throughout the Yakima River Basin in Washington, USA from May-July 2023. Fire affected sites are defined as those affected by the 2021 Schneider Springs Fire, based on the drainage area of the streams being within the 2021 Schneider Springs Fire burn perimeter or not (Figure 1, below). The contents include surface water geochemistry data (dissolved organic carbon; total dissolved nitrogen; total suspended solids); short-term sonde data (specific conductivity; turbidity; pH; chlorophyll A; temperature); stream depth data; stream velocity; manual chamber open channel respiration data; sensor time-series data (oxygen; water pressure; barometric pressure); field metadata (including qualitative information on in stream and river corridor characteristics); and environmental context photos taken in the field. The dataset also includes a summary file of the sensor data and plots of the sensor data. Sensors were only recovered at 15 out of the 17 sites, and not all sensors were recovered at all 15 sites (see Methods section for more details), therefore all data does not exist at all sites. Data from a 2022 study at the same sites, as well as additional sites, can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/1969566. The data package was originally published in November 2023. It was updated in June 2025 (v2; modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one folder with field photos and one main data folder with two subfolders. The main data folder consists of (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) field protocol; (5) readme; (6) international generic sample number (IGSN) mapping file; and (7) stream depth and averages. The sensor data subfolder consists of (1) sensor installation methods summary; (2) stream velocity; and (3) six subfolders. The BarotrollAtm (barometric pressure; temperature), DepthHOBO (water pressure; temperature), MantaRiver (specific conductivity; turbidity; pH; chlorophyll A; temperature), EXO (specific conductivity; pH; temperature), miniDOT (dissolved oxygen; temperature), and miniDOTManualChamber (dissolved oxygen; temperature) contain time-series data, plots, and summary files. The sample data subfolder consists of (1) total suspended solids (TSS) data; (2) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (3) total dissolved nitrogen (TN) data and averages; and (4) methods codes. All files are .csv, .pdf, .jpg, .jpeg, or .mov.

54 ENVIRONMENTAL SCIENCES↗

Pore-resolved investigation of turbulent open channel flow over a randomly packed permeable sediment bed

Pore-resolved direct numerical simulations are performed to investigate the interactions between streamflow turbulence and groundwater flow through a randomly packed porous sediment bed for three permeability Reynolds numbers, Re K = 2.56 , 5.17 and 8.94, representative of natural stream or river systems. Time–space averaging is used to quantify the Reynolds stress, form-induced stress, mean flow and shear penetration depths, and mixing length at the sediment–water interface (SWI). Here, the mean flow and shear penetration depths increase with Re K and are found to be nonlinear functions of non-dimensional permeability. The peaks and significant values of the Reynolds stresses, form-induced stresses, and pressure variations are shown to occur in the top layer of the bed, which is also confirmed by conducting simulations of just the top layer as roughness elements over an impermeable wall. The probability distribution functions (p.d.f.s) of normalized local bed stress are found to collapse for all Reynolds numbers, and their root-mean-square fluctuations are assumed to follow logarithmic correlations. The fluctuations in local bed stress and resultant drag and lift forces on sediment grains are mainly a result of the top layer; their p.d.f.s are symmetric with heavy tails, and can be well represented by a non-Gaussian model fit. The bed stress statistics and the pressure data at the SWI potentially can be used in providing better boundary conditions in modelling of incipient motion and reach-scale transport in the hyporheic zone.

turbulence simulation↗