Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Machine learning-enabled model-data integration for predicting subsurface water storage

Subsurface water storage (SWS) is a key variable of the climate system and a storage component for precipitation and radiation anomalies, inducing persistence in the climate system. It plays a critical role in climate-change projections and can mitigate the impacts of climate change on ecosystems. However, because of the difficult accessibility of the underground, hydrologic properties and dynamics of SWS are poorly known. Direct observations of SWS are limited, and accurate incorporation of SWS dynamics into Earth system land models remains challenging. We propose a machine learning-enabled model-data integration framework to improve the SWS prediction at local to conus scales in a changing climate by leveraging all the available observation and simulation resources, as well as to inform the model development and guide the observation collection. The accurate prediction will enable an optimal decision of water management and land use and improve the ecosystem's resilience to the climate change.

Lu, Dan↗

Model Data Archive for Manuscript Titled "Evaluation of a Coupled Surface–Subsurface Hydrologic Model Using Dense Water‑Level Sensors in a Mixed Urban–Rural Watershed"

This archive provides scripts, input files, and datasets used for the implementation and evaluation of a fully coupled surface–subsurface hydrologic model in the Neches River Basin, southeast Texas. The study uses the Advanced Terrestrial Simulator (ATS) to simulate coupled surface–subsurface hydrologic processes over a mixed urban–rural watershed and evaluates model performance using a dense network of 136 in situ water-level sensors, nine U.S. Geological Survey (USGS) stream gauges, and SSEBop-derived evapotranspiration estimates during the period October 2014–June 2024. The workflow is implemented primarily in Python 3 using the Watershed Workflow package. The Jupyter notebooks can be executed using open-source software such as Anaconda JupyterLab or Visual Studio Code. Other data files include TXT, CSV, XML, SHP, TIF, NetCDF, HDF5, and ExodusII files, which can be processed using the provided Python scripts. ATS input files are provided in XML format and can be edited using any commonly used text editor. This archive contains: *Scripts and input files used to generate the ATS model setup, including watershed discretization, mesh generation, parameter mapping, and model configuration. *Jupyter notebooks used for preprocessing observational data, evaluating streamflow, water levels, and evapotranspiration, computing performance metrics, and generating the figures presented in the manuscript. *ATS simulation outputs and processed observational datasets, including OneRain and DD6 water-level sensors, USGS streamflow observations, GIS data, and supporting spatial datasets used throughout the study.

Dense water-level sensor network↗

Bayesian model-data comparison incorporating theoretical uncertainties

Accurate comparisons between theoretical models and experimental data are critical for scientific progress. However, inferred physical model parameters can vary significantly with the chosen physics model, highlighting the importance of properly accounting for theoretical uncertainties. In this Letter, we present a Bayesian framework that explicitly quantifies these uncertainties by statistically modeling theory errors, guided by qualitative knowledge of a theory’s varying reliability across the input domain. We demonstrate the effectiveness of this approach using two systems: a simple ball drop experiment and multi-stage heavy-ion simulations. In both cases incorporating model discrepancy leads to improved parameter estimates, with systematic improvements observed as additional experimental observables are integrated.

Bayesian methods↗

Utah FORGE: 2024 Discrete Fracture Network Model Data

The Utah FORGE 2024 Discrete Fracture Network (DFN) Model dataset provides a set of files representing discrete fracture network modeling for the FORGE site near Milford, Utah. The dataset includes four distinct DFN model file sets, each corresponding to different time frames and modeling approaches in 2024. These models characterize both natural and induced fractures in the geothermal reservoir, which consists of crystalline granitic and metamorphic rock approximately 8,000 feet below the ground surface. The dataset includes a reference DFN model from February 2024 that incorporates planar fractures and well trajectories, as well as upscaled permeability, porosity, compressibility, and storage values on specified grids. Additionally, there are models based on new microseismic (MEQ) data from May and July 2024, including fracture planes fitted to the latest MEQ catalog datasets, tensile fractures from hydraulic stimulation, and an alternative connected DFN for modeling purposes. Coordinate data is provided in both global and local frames, with detailed instructions on the transformations used to align with principal stress orientations. The dataset also includes notes and calculation files for estimating fracture sizes and differences between various fracture sets. There are subfolders for Global Coordinates and Local Coordinates. To move from the global to the local coordinate frame, fractures and wells were a) rotated 20 degrees counterclockwise looking down about the global point (335376.400482041, 4263189.99998761, 250.093546450195) to better align with the principal stresses; and b) translated by (-335408.68, -4263010.9, 1150). Upscaled permeability values using the _XYZ suffix show directions with respect to the global XYZ coordinate frame, while those using the _IJK suffix are aligned with local coordinate frame.

15 GEOTHERMAL ENERGY↗

Block Island Acoustic Propagation Modeling Data

This dataset contains acoustic propagation model outputs, computational subroutines, and analysis tools for underwater sound propagation in the Block Island region, including parabolic equation (PE) model results, visualization products, and comprehensive modeling software tools.

17 WIND ENERGY↗

Hydrologic Model Data for the East Fork Poplar Creek Watershed Simulated with the Advanced Terrestrial Simulator (ATS): Streamflow and Network Expansion–Contraction Dynamics

This dataset supports hydrologic modeling and stream network expansion–contraction analysis for the East Fork Poplar Creek (EFPC) Watershed in Tennessee. It includes a Jupyter notebook for model setup, model configuration files, simulation outputs, and derived products used to evaluate model performance and investigate stream dynamics under varying hydrologic conditions. The dataset was generated using the Watershed Workflow Python package and the Advanced Terrestrial Simulator (ATS), enabling integrated surface–subsurface hydrologic simulations using a stream-aligned mesh. Outputs include high-resolution time series of streamflow, active network length, water table depth, and related hydrologic variables. Also included are spatially explicit stream persistency indices and classifications of reaches as perennial or non-perennial. These data facilitate reproducibility and support further research on stream intermittency and variability in network extent.The model data archive is organized in following directories:1) model_setup_inputsContains the Watershed Workflow Jupyter notebooks (accessed through any open source code editor), selected input datasets, and resulting ATS input files, including XML files (access through any open source code editor), computational mesh (.exo files can be viewed using Paraview), and meteorological forcing files (.h5 files can be accessed through h5py python package and HDFView open source software). 2) model_outputsIncludes ATS simulation outputs relevant to this study. Time series of spatially integrated or averaged variables (e.g., streamflow, water table depth) are provided as CSV files. Select spatial fields (e.g., ponded depth and water table depth) are saved as pickled Python objects to reduce file size, and can be accessed through pickle package in Python. Key geometry objects from Watershed Workflow—such as the surface mesh and river tree—are also included to support analysis of streamflow persistency and expansion–contraction dynamics. These files can also be accessed through Watershed Workflow Python package.3) model_evaluationProvides observed streamflow time series and field survey-based flow regime classifications used to evaluate model performance. Jupyter notebooks for processing ATS outputs and comparing model predictions with observations to build confidence in the model prior to scientific analysis are also included.4) Q_L_relationshipsContains workflows for generating time series of discharge, active network length, and related hydrologic variables used in the stream network expansion–contraction analysis. Includes routines for delineating baseflow-dominated periods. For each catchment, notebooks and processed data (as pickled DataFrames accessed through Pandas Python package) are provided. 5) figure_scriptsProvides the Jupyter notebooks used to generate the figures presented in the paper.

54 ENVIRONMENTAL SCIENCES↗

Northern Pacific Turbulence Intensity Model Data in Observational Space

The dataset archives model-simulated turbulence intensity and meteorological profiles and timeseries at the lidar buoys sites off the coast of California (Humboldt and Morro Bay). The simulated data are interpolated in time and/or space according to observed quantities. The simulations were carried out for the north Pacific region using the revised Weather Research and Forecasting (WRF) model version 4.2 that incorporates the implementation of online turbulence intensity (TI) calculations (Tai et al. 2023). The simulated atmospheric profiles near the Shell Exploration and Production Corporation's Tension Leg Platforms Ursa and Mars are archived. Physics parameterizations chosen for the simulations include the Thompson microphysics parameterization, Mellor-Yamada-Nakanishi Niino (MYNN) boundary layer parameterization, Mellor-Yamada-Janjic surface layer parameterization, Unified Noah land-surface parameterization, and the RRTMG longwave and shortwave radiation parameterization. Initial and boundary conditions are taken from NOAA’s High-Resolution Rapid Refresh (HRRR) product. The JPL 0.01-degree Level 4 Multiscale Ultrahigh Resolution (MUR) Global Foundation Sea Surface Temperature (SST) Analysis (V4.1) data are used as the model’s SST forcing.

17 WIND ENERGY↗

Gulf of Mexico Turbulence Intensity Model Data in Observational Space

The dataset archives model-simulated turbulence intensity and meteorological profiles and timeseries at the Shell Exploration and Production Corporation's Tension Leg Platforms Ursa and Mars. The simulated data are interpolated in time and/or space according to observed quantities. The simulations were carried out for the Gulf of Mexico region using the revised Weather Research and Forecasting (WRF) model version 4.2 that incorporates the implementation of online turbulence intensity (TI) calculations (Tai et al. 2023). The simulated atmospheric profiles near the Shell Exploration and Production Corporation's Tension Leg Platforms Ursa and Mars are archived. Physics parameterizations chosen for the simulations include the Thompson microphysics parameterization, Mellor-Yamada-Nakanishi Niino (MYNN) boundary layer parameterization, Mellor-Yamada-Janjic surface layer parameterization, Unified Noah land-surface parameterization, and the RRTMG longwave and shortwave radiation parameterization. Initial and boundary conditions are taken from NOAA’s High-Resolution Rapid Refresh (HRRR) product. The JPL 0.01-degree Level 4 Multiscale Ultrahigh Resolution (MUR) Global Foundation Sea Surface Temperature (SST) Analysis (V4.1) data are used as the model’s SST forcing.

17 WIND ENERGY↗

Data-model files associated with the manuscript "Modeling the Effects of Wetland Restoration on Coastal Hydrology: A Case Study of Elkhorn Slough Watershed, California"

This package contains the data, simulation setups, notebooks and figures used in “Modeling the Effects of Wetland Restoration on Coastal Hydrology: A Case Study of Elkhorn Slough Watershed, California” (Xu et al., 2025). In this study, we selected Elkhorn Slough, a tidal estuary, in California, to investigate the impact of wetland restoration and sea level rise on coastal hydrology using the process-based coastal hydrologic model, Advanced Terrestrial Simulator (ATS), informed by site-specific data. We designed a novel modeling workflow for incorporating wetland restoration features into land cover and soil properties for the model parameterization. The validation results demonstrate a strong agreement between modeled and observed data. We studied the characteristics of coastal watershed hydrology, then focused on the surface water dynamics at two wetland sites within Elkhorn Slough, a reference site and a restored site. Our simulation results indicate that the restored site successfully maintains surface elevation, resulting in reduced surface inundation. We also examined the impact of wetland restoration under expected sea level rise over the next few decades. The low-lying Yampah Marsh, the reference site, is likely to be inundated due to future sea level rise when highest tides arrive; while a higher percentage of Hester Marsh, the restored site, would retain marsh vegetation in coming decades, regardless of tidal conditions. Our study provides important information for examining the outcome of restoration practices that include surface elevation in tidal wetlands under climate changes.Several files can be found from this data package.1. README.md: This file describes the title, journal, co-authors, abstract, repository structure and model version.2. Simulation_Setups.zip: The file contains the model configuration files (XML format) for ATS. 3. Notebooks.zip: The file contains the Jupyter notebooks for generating the pre- and post-restoration meshes and the meshes of future scenarios. 4. Figures.zip: The file contains the figures used in the manuscript.5. Data.zip: The file contains the data used to drive the model simulations, including watershed and wetlands boundaries, mesh files and references to additional datasets (e.g., meteorological forcing, tidal dataset, DEMs, land cover, soil properties). Also, it contains water level observations at the restored wetland.

54 ENVIRONMENTAL SCIENCES↗

The future of Earth system prediction: Advances in model-data fusion

Predictions of the Earth system, such as weather forecasts and climate projections, require models informed by observations at many levels. Some methods for integrating models and observations are very systematic and comprehensive (e.g., data assimilation), and some are single purpose and customized (e.g., for model validation). We review current methods and best practices for integrating models and observations. We highlight how future developments can enable advanced heterogeneous observation networks and models to improve predictions of the Earth system (including atmosphere, land surface, oceans, cryosphere, and chemistry) across scales from weather to climate. As the community pushes to develop the next generation of models and data systems, there is a need to take a more holistic, integrated, and coordinated approach to models, observations, and their uncertainties to maximize the benefit for Earth system prediction and impacts on society.

54 ENVIRONMENTAL SCIENCES↗

Aerial and Processed Model Data Representing As-built Conditions in Coastal Port Arthur, Texas in May 2025

This dataset was collected by the Co-Design Team of the Southeast Texas Urban Integrated Field Lab, a research initiative led by the University of Texas at Austin and funded by the U.S. Department of Energy. The broader project focuses on developing climate-resilient design solutions for the Beaumont–Port Arthur region, with more information available at www.setx-uifl.org. Our team conducted aerial surveys of the Port Arthur coastal neighborhood in May 2025, before the start of construction scheduled for Summer 2026. These pre-construction datasets are designed to facilitate comparative analyses, including pre- and post-construction assessments and simulated inundation scenario evaluations. Aerial images were captured using DroneDeploy autonomous flight systems, with imagery processed through the DroneDeploy engine. All original aerial photographs are provided in JPG format and organized in zipped folders by area. The processed data package includes: 3D surface models Orthomosaics Geospatial and topographic mappings Point clouds For guidance on file contents, structure, and recommended usage, please refer to the included README file.

2D mapping↗

Model data for Flood Frequency Analysis using Stochastic Storm Transposition and an Integrated Surface-Subsurface Hydrological Model

This archived provides scripts and input files used for the implementation of a novel approach to conduct process-based Flood Frequency Analysis using a Stochastic Storm Transposition (SST) and an Integrated Surface-Subsurface Hydrological Model (ISSHM). As a proof-of-concept, this study uses the ISSHM, Advanced Terrestrial Simulator (Amanzi-ATS) model, and the SST model, RainyDay, to conduct flood frequency analysis by simulating the flood response to 5,000 annual synthetic storm events in a ~2000 km2 Southeast Texas watershed.The Watershed Workflow package is implemented in Python3. The Jupyter notebooks can be executed through multiple open-source tools, for example, Anaconda Jupyter Lab, VS Studio Code, etc. Other data files include TXT, CSV, DAT, SBATCH, SHP, TIF, NetCDF, and HDF5 files, which can be read through Python scripts. The input files for the ATS model and RainyDay model have .XML and .SST extensions, respectively, and can be edited in any commonly used text editors.This archive contains:* Scripts and data files essential for generating the ATS model input. It uses the Watershed Workflow package to produce both mesh and ATS input files. * Jupyter notebooks designated for the ATS model evaluation, covering both long-term simulations and 40 rainfall-runoff events.* Input files required to simulate SST storm events using RainyDay.

54 ENVIRONMENTAL SCIENCES↗

Downscaled Earth System Model Data for Resilient Energy System Planning

The second-generation Sup3rCC dataset provides high-resolution meteorological data generated through the downscaling of multiple earth system models (ESMs) from the Coupled Model Intercomparison Project Phase 6 (CMIP6). This downscaling is performed through application of a generative machine learning approach called Super-Resolution for Renewable Resource Data (sup3r). This dataset builds on the first-generation Sup3rCC data by applying improved bias correction methods and adding downscaled precipitation to the output variables. In this presentation, we explore the output characteristics of the dataset and various validation analyses. We also present and discuss plans for the integration of this data into power system planning models using a decision-making under deep uncertainty (DMDU) methodology.

97 MATHEMATICS AND COMPUTING↗

A clustering-based approach to ocean model–data comparison around Antarctica

The Antarctic Continental Shelf seas (ACSS) are a critical, rapidly changing element of the Earth system. Analyses of global-scale general circulation model (GCM) simulations, including those available through the Coupled Model Intercomparison Project, Phase 6 (CMIP6), can help reveal the origins of observed changes and predict the future evolution of the ACSS. However, an evaluation of ACSS hydrography in GCMs is vital: previous CMIP ensembles exhibit substantial mean-state biases (reflecting, for example, misplaced water masses) with a wide inter-model spread. Because the ACSS are also a sparely sampled region, grid-point-based model assessments are of limited value. Our goal is to demonstrate the utility of clustering tools for identifying hydrographic regimes that are common to different source fields (model or data), while allowing for biases in other metrics (e.g., water mass core properties) and shifts in region boundaries. We apply K-means clustering to hydrographic metrics based on the stratification from one GCM (Community Earth System Model version 2; CESM2) and one observation-based product (World Ocean Atlas 2018; WOA), focusing on the Amundsen, Bellingshausen and Ross seas. When applied to WOA temperature and salinity profiles, clustering identifies “primary” and “mixed” regimes that have physically interpretable bases. For example, meltwater-freshened coastal currents in the Amundsen Sea and a region of high-salinity shelf water formation in the southwestern Ross Sea emerge naturally from the algorithm. Both regions also exhibit clearly differentiated inner- and outer-shelf regimes. The same analysis applied to CESM2 demonstrates that, although mean-state model biases in water mass T–S characteristics can be substantial, using a clustering approach highlights that the relative differences between regimes and the locations where each regime dominates are well represented in the model. CESM2 is generally fresher and warmer than WOA and has a limited fresh-water-enriched coastal regimes. Given the sparsity of observations of the ACSS, this technique is a promising tool for the evaluation of a larger model ensemble (e.g., CMIP6) on a circum-Antarctic basis.

54 ENVIRONMENTAL SCIENCES↗

Model Data Archive Associated with Manuscript "Fire-altered Carbon Pools Create Disturbance Memory in Stream Dissolved Organic Carbon"

This data package supports the publication “Fire-altered Carbon Pools Create Disturbance Memory in Stream Dissolved Organic Carbon” by Li et al. (2026). The package contains processed model inputs, configuration files, restart files, simulation outputs, scripts, and visualization products used to evaluate post-fire dissolved organic carbon (DOC) dynamics in the Naches River Watershed, Washington, USA, following the 2021 Schneider Springs Fire. The modeling workflow couples ELM-BGC, the biogeochemistry-enabled Energy Exascale Earth System Model Land Model; ATS, the Advanced Terrestrial Simulator for integrated surface-subsurface hydrology; and PFLOTRAN, a reactive transport model for multicomponent aqueous geochemistry. Together, these models simulate how wildfire-induced changes in vegetation, litter, coarse woody debris, and soil organic matter influence DOC production, transport, and reaction from burned hillslopes to stream networks. The archive includes preprocessed meteorological, geospatial, hydrologic, and biogeochemical forcing data; ELM-BGC-derived DOC source terms; ATS mesh files; PFLOTRAN reactive-transport inputs; model configuration files; spin-up and transient restart files; watershed-scale diagnostic outputs; stream concentration time series; and figures or visualization files used to inspect and reproduce key results. File types include Hierarchical Data Format 5 (HDF5) files for gridded forcing and model-coupling data, model input and configuration files for ELM-BGC, ATS, and PFLOTRAN, restart and simulation-output files generated by the modeling workflow, tabular or time-series diagnostic outputs, scripts for post-processing and figure generation, and image or visualization products associated with the manuscript. Use of the package depends on the intended task. Re-running the simulations requires the relevant modeling software, including ELM-BGC, ATS, and PFLOTRAN as ATS's geochemical engine. Inspecting outputs and reproducing figures requires Python with scientific plotting libraries such as Matplotlib, and three-dimensional model outputs may be viewed with ParaView. Geographic information system files or maps may be inspected with ArcGIS Pro or comparable GIS software. The data package is intended to enable traceability, reuse, and partial reproduction of the coupled land-to-watershed hydro-biogeochemical modeling workflow used to test how wildfire disturbance affects terrestrial carbon pools and downstream DOC dynamics.

ATS↗

Data, model inputs, and analysis scripts associated with a manuscript on stream intermittency controls across spatial scales in Pacific Northwest watersheds

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript "Hydroclimatic Memory and Watershed Template Shape Stream Intermittency: Multi-scale Attribution Using Process-based Simulation and Explainable ML" by Niroula et al. (2026), submitted to Water Resources Research (WRR). The study investigates the dominant controls on stream intermittency across local, reach, and watershed scales using a coupled process-based simulation and explainable machine-learning framework. Long-term daily simulations from the Advanced Terrestrial Simulator (ATS) were used to generate wetness states and ponded-depth responses over river-corridor cells. These ATS outputs were then aggregated across scales and used to train XGBoost (eXtreme Gradient Boosting) models. SHAP (SHapley Additive exPlanations) was applied to quantify the relative importance of hydroclimatic forcings, watershed template attributes, and antecedent-memory effects in shaping intermittency behavior. The analysis is carried out for three contrasting Pacific Northwest watersheds: Oak Creek (OCW), American River Watershed (ARW), and H.J. Andrews (HJA). Across these testbeds, the package contains ATS-ready watershed inputs, ATS run configuration and selected output files, model-evaluation data products, intermittency-analysis datasets, machine-learning target-feature tables, SHAP outputs, and notebooks used to organize, analyze, and visualize results. At a high level, the package documents a workflow in which ATS provides the physically based simulation backbone and explainable machine learning is used as a post-processing attribution tool. The contents are intended to support interpretation of the manuscript figures and results, provide context for how intermittency metrics were generated at multiple scales, and preserve the key artifacts needed to understand and reuse the analysis workflow. The package contains a high-level directory summary file (`summary.txt`) and four main content folders (1) `evaluation_plots` contains evaluation figures and supporting evaluation datasets; (2) `intermittency_plots` contains intermittency-focused analysis notebook and prepared datasets; (3) `ml-training-and-shap_values_plots` contains ML training inputs, SHAP outputs, and figure-generation notebooks; and (4) `watershed_mesh_and_ats_input` contains ATS model setup materials, forcing inputs, geometry, and selected run files. More specifically, the `evaluation_plots` folder contains the notebook used for ATS evaluation plotting and site-specific evaluation datasets. These include evapotranspiration and water-balance products for three watersheds, as well as an Oak Creek field-measurement discharge file. The `intermittency_plots` folder contains the notebook used for intermittency analysis and the prepared datasets used to analyze intermittent and non-intermittent wetness behavior across the study watersheds. The `ml-training-and-shap_values_plots` folder contains notebooks and outputs for the machine-learning and explainability workflow. This includes the main XGBoost and SHAP notebook(s), a beeswarm plotting notebook, target-feature tables for machine-learning training, SHAP summary tables, and per-sample SHAP value archives. The `watershed_mesh_and_ats_input` folder contains ATS-related watershed inputs and supporting materials. This includes mesh and shape products, ATS-readable LAI and meteorological forcing inputs, selected ATS spinup and transient-run files, and a watershed workflow example notebook. Subdirectories are organized by watershed where applicable.All files are .cpg (codepage files), .csv (comma-separated values), .dbf (database files), .exo (Exodus mesh format), .h5 (HDF5 format), .ipynb (Jupyter notebooks), .pkl (Python pickle), .prj (projection files), .sh (shell scripts), .shp (shapefile geometry), .shx (shapefile index), .txt (text files), or .xml (markup data).

Advanced Terrestrial Simulator↗

Revolutionizing Materials Design: The Intersection of Quantum Mechanics and Data Modeling

The field of materials design is currently experiencing a notable evolution, driven by the convergence of sophisticated computational methodologies based on first principles and data-driven modeling approaches. I will review our recent endeavors employing AI/ML to expedite first-principles simulations and mitigate traditional methods' temporal and spatial limitations. Central to our efforts is developing and utilizing ML interatomic potentials (MLPs) across a diverse spectrum of materials. We show that MLPs serve as invaluable tools for navigating the complexities of the simulations, such as understanding the behavior of MgO at extreme environments of ~1 terapascal and temperatures >10,000 Kelvin. Moreover, we show that MLPs can provide precise details of the intricate dynamics governing the oxidation processes of binary alloy systems due to the competition between surface segregation and reconstruction tendencies. In summation, advancements in MLPs open the door to fresh possibilities in material modeling and, ultimately, discovery.

Saidi, Wissam↗

Models, data, and scripts associated with “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning”

This data package is associated with the publication “Prediction of Distributed River Sediment Respiration Rates using Community-Generated Data and Machine Learning’’ submitted to the Journal of Geophysical Research: Machine Learning and Computation (Scheibe et al. 2024). River sediment respiration observations are expensive and labor intensive to obtain and there is no physical model for predicting this quantity. The Worldwide Hydrobiogeochemisty Observation Network for Dynamic River Systems (WHONDRS) observational data set (Goldman et al.; 2020) is used to train machine learning (ML) models to predict respiration rates at unsampled sites. This repository archives training data, ML models, predictions, and model evaluation results for the purposes of reproducibility of the results in the associated manuscript and community reuse of the ML models trained in this project. One of the key challenges in this work was to find an optimum configuration for machine learning models to work with this feature-rich (i.e. 100+ possible input variables) data set. Here, we used a two-tiered approach to managing the analysis of this complex data set: 1) a stacked ensemble of ML models that can automatically optimize hyperparameters to accelerate the process of model selection and tuning and 2) feature permutation importance to iteratively select the most important features (i.e. inputs) to the ML models. The major elements of this ML workflow are modular, portable, open, and cloud-based, thus making this implementation a potential template for other applications. This data package is associated with the GitHub repository found at Please see the file level metadata (flmd; “sl-archive-whondrs_flmd.csv”) for a list of all files contained in this data package and descriptions for each. Please see the data dictionary (dd; “sl-archive-whondrs_dd.csv”) for a list of all column headers contained within comma separated value (csv) files in this data package and descriptions for each. The GitHub repository is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning models trained on the data in “input_data”; (3) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; (4) “examples” contains the visualization of the results in this repository including plotting scripts for the manuscript (e.g., model evaluation, FPI results) and scripts for running predictions with the ML models (i.e., reusing the trained ML models); (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. Furthermore, depending on the number of features used to train the ML models, the preprocessing and postprocessing scripts, and their intermediate results, can also be different branch-to-branch. The “main-*” branches are meant to be starting points (i.e. trunks) for each model branch (i.e. sprouts). Please see the Branch Navigation section in the top-level README.md in the GitHub repository for more details. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please the top-level README.md in the GitHub repository for more details on the automation.

13C↗