Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “github”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

HydraGNN_Predictive_GFM_2024 - Ensemble of predictive graph foundation models for ground state atomistic materials modeling

We provide the ensemble of fifteen pre-trained graph foundation models (GFMs) for atomistic materials modeling applications. Each one of the fifteen GFMs has been trained on five open-source datasets that (once aggregated) amount to over 154 million atomistic structures, which cover over two-thirds of the natural elements of the periodic table and that comprises a broad set of organic and inorganic compounds. This vast set of atomistic structures comprises ground state configurations that are dynamically stable (i.e., equilibrated structures with atomic forces approximately close to zero values) as well as dynamically unstable structures (i.e., non-equilibrium structures with non-negligible non-zero values of atomic forces). The ensemble of datasets aggregated does NOT include excited states. The datasets have been curated to remove atomistic structures with spectral norm of the force tensor above 100 eV/angstrom. Moreover, a linear term of the energy was computed for each dataset using a linear regression model that uses the chemical concentration of each natural element as regressor. The linear term predicted by the linear regression model has been subtracted from each original energy value to perform a re-alignment of the energy values across different electronic structures approximation theories performed to generate the diverse multi-source, multi-fidelity datasets. The folder "ADIOS_files" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "ADIOS_files" directory contains 6 sub-directories named as follows: - ANI1x-v3.bp - MPTrj-v3.bp - OC2020-20M-v3.bp - OC2020-v3.bp - OC2022-v3.bp - qm7x-v3.bp Each sub-directory contains the pre-processed datasets converted in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used to the development, training, and performance testing of the ensemble go predictive graph foundation models. Each GFM was developed using HydraGNN (https://github.com/ORNL/HydraGNN) as underlying graph neural network (GNN) architecture. The multi-task learning (MTL) capability of HydraGNN was used to simultaneously train the GFMs on labeled values for direct predictions of energy (a total system property of an atomistic structure that measures the chemical stability) and atomic forces (an atomic level property of an atomistic structure that measures the dynamical stability). The hyper parameters of the GFM have been tuned using scalable hyperparameter optimization (HPO) algorithms implemented in the software DeepHyper (https://github.com/deephyper/deephyper). The pre-training of each HPO trial was performed using distributed data parallelism (DDP) to scale the training across 128 compute nodes of the exascale OLCF supercomputer Frontier. Each HPO trial was trained only for 10 epochs and an early stopping was performed to avoid wasting significant computational resources on GNN architectures that were clearly underperforming. For each HPO trial, the 'omnistat' tool developed by (AMD Research - Advanced Micro Device) was used to measure the total energy consumption in kWh. The ensemble of GFMs was obtained by selecting the fifteen best performing HPO trials. Four models have been selected for their clear advantage in accuracy, and these are the GFMs with IDs 229, 156, 147, 260. Additional eleven models have been selected based on judicious balance between accuracy and energy consumption needed for training, and these are the GFMs with IDs 165, 78, 137, 1, 175, 171, 181, 67, 179, 167, 351. Each selected GFM of the ensemble was continued to cumulate a total of at most 30 epochs. In some cases, the total number of epochs actually performed was les than 30 due to two combined factors: (1) the size of the GFM (i.e., the number of model parameters to train) and (2) the total wall-clock time for which the computational resources could be allocated on OLCF-Frontier. The "Ensemble_of_models" directory contains 15 sub-directories named as follows: - gfm_0.229 - gfm_0.156 - gfm_0.147 - gfm_0.260 - gfm_0.165 - gfm_0.78 - gfm_0.137 - gfm_0.1 - gfm_0.175 - gfm_0.171 - gfm_0.181 - gfm_0.67 - gfm_0.179 - gfm_0.167 - gfm_0.351 Each one of these sub-directories refers to one of the fifteen HPO trials that have been selected to continue the pre-training with at most 30 epochs. With each sub-directory associated with a specific HPO trial, the following files can be found: - config.json: file for argument parsing to develop and train an HydraGNN architecture - gfm_0.ID_epoch_N.pk: file with model parameters for HPO ID trial after N epochs of training The ensemble of fifteen GFM architectures was used for (1) ensemble averaging to stabilize the predictions of energy and atomic forces after pre-training for post-processing analysis and (2) ensemble uncertainty quantification (UQ). The code used to develop, pre-train, and load the pre-trained models for post-processing analysis is available on the ORNL-GitHub at the following link: https://github.com/ORNL/HydraGNN/tree/Predictive_GFM_2024

36 MATERIALS SCIENCE↗

FeSi binary alloy electronic structure low-Si dataset (1024 atoms)

This dataset contaims the calculated atomic charge density, atomic magnetic moment, and total energy for 1600 configurations of iron-silicon (Fe-Si) binary alloys body-centered cubic (BCC) structures at 3, 6, and 9% Si. These large scale (1024 atom) ab initio calculations were produced with the LSMS code on the OLCF Summit supercomputer. LSMS GitHub repository: https://github.com/mstsuite/lsms

36 MATERIALS SCIENCE↗

Dataset for Leveraging CryoEM and AI-Driven Morphological Feature Analysis for Insights on Bacterial Structures

This repository hosts an AI-assisted image segmentation and analysis pipeline for Pantoea sp. YR343 cryo-electron microscopy (cryoEM) datasets. The workflow automates membrane thickness measurements, flagella detection, and field-of-view (FOV) screening from low-dose, high-resolution cryoEM micrographs eliminating the need for slow manual annotation. By integrating deep-learning based segmentation (YOLOv11) with quantitative post-processing, this toolkit provides a scalable and reproducible way to study bacterial morphology under hydrated, near-native conditions. The GitHub repository for AI-based tools for cryoEM bacteria ultrastructures can be found here: https://github.com/Sireesiru/Cryo-EM-Ultrastructures/tree/main

60 APPLIED LIFE SCIENCES↗

Dataset for Top Model Decision Tree: Selecting Segmentation Models for Reliable Quantitative Analysis in Low- and Ultralow-Dose CryoEM

Motivation Multiple deep learning model architectures can be used to segment bacterial membranes in cryoEM images. However, an AI-based tool advancement is often presented with only a single segmentation model for broad use, and this single model may show inconsistent results across datasets from different users. Here, we present the Top Model Decision Tree, a model screening framework to screen for the best model to generate bacterial inner and outer membrane masks based on user priorities. We use pre-trained segmentation models from YOLOv11, YOLO26, U-Net, Detectron2 and SAM3 fine-tuned on bacterial inner and outer membranes imaged with cryoEM. Run the Framework This notebook must be opened in Google Colab. Mount Google Drive and run with a GPU-based runtime. Open the notebook and follow steps to git clone in folders and files within this repository. There will be a repeating top_model_decision_tree.ipynb (notebook clone) that will not be used. Save your .png binary mask files and .csv table outputs within your Google Drive or download before closing the notebook. The models and all analysis/training scripts are available at [GitHub: https://github.com/Lynnicia/CryoEM_membranes_top_model_decision_tree and https://github.com/Sireesiru/Semantic-Segmentation-of-bacterial-cell-envelope-using-U-Nets.

59 BASIC BIOLOGICAL SCIENCES↗

Pangeo-Enabled ESM Pattern Scaling (PEEPS): A customizable dataset of emulated Earth System Model output

Emulation through pattern scaling is a well-established method of rapidly producing climate fields (like temperature or precipitation) from existing Earth System Model (ESM) output that, while inaccurate, is often useful for a variety of downstream purposes. Conducting pattern scaling has historically been a laborious process, in large part due to the increasing volume of ESM output data that has often required downloading and storing locally to train on. Here we describe the Pangeo-Enabled ESM Pattern Scaling (PEEPS) dataset, a repository of trained annual and monthly patterns from CMIP6 outputs. This manuscript describes and validates these updated patterns so that users can save effort calculating and reporting error statistics in manuscripts focused on the use of patterns. The trained patterns are available as NetCDF files on Zenodo for ease of use in the impact community, and are reproducible with the code provided via GitHub in both Jupyter notebook and Python script formats. Because all training data for the PEEPS data set is cloud-based, users do not need to download and house the ESM output data to reproduce the patterns in the zenodo archive, should that be more efficient. Validating the PEEPS data set on the CMIP6 archive for annual and monthly temperature, precipitation, and near-surface relative humidity, pattern scaling performs well over a variety of future scenarios except for regions in which there are strong, potentially nonlinear climate feedbacks. Although pattern scaling is normally conducted on annual mean ESM output data, it works equally well on monthly mean ESM output data. We identify several downstream applications of the PEEPS data set, including impacts assessment and evaluating certain types of Earth system uncertainties.

54 ENVIRONMENTAL SCIENCES↗

SEGUID v2: Extending SEGUID checksums for circular, linear, single- and double-stranded biological sequences

Background Synthetic biology involves combining different DNA fragments, each containing functional biological parts, to address specific problems. Fundamental gene-function research often requires cloning and propagating DNA fragments, such as those from the iGEM Parts Registry or Addgene, typically distributed as circular plasmids. Addgene’s repository alone offers around 150,000 plasmids. To ensure data integrity, cryptographic checksums can be calculated for the sequences. Each sequence has a unique checksum, making checksums useful for validation and quick lookups of associated annotations. For example, the SEGUID checksum uniquely identifies protein sequences with a 27-character string. Objectives The original SEGUID, while effective for protein sequences and single-stranded DNA (ssDNA), is not suitable for circular DNA since there is no natural starting position nor for double-stranded DNA (dsDNA) since two separate sequences are present. Challenges include how to uniquely represent linear dsDNA, circular ssDNA, and circular dsDNA. To meet these needs, we propose SEGUID v2, which extends the original SEGUID to handle additional types of sequences. Conclusions SEGUID v2 produces orientation and rotation invariant checksums for single-stranded, double-stranded, possibly staggered, linear, and circular DNA and RNA sequences. Customizable alphabets allow for other types of sequences. In contrast to the original SEGUID, which uses Base64, SEGUID v2 uses Base64url to encode the SHA-1 hash. This ensures SEGUID v2 checksums can be used as-is in filenames, regardless of platform, and in URLs, with minimal friction. Availability SEGUID v2 is readily available for major programming languages, distributed under the MIT license. JavaScript package seguid is available on npm, Python package seguid on PyPi, R package seguid on CRAN, and a Tcl script on GitHub. These tools, along with documentation, examples, and an online SEGUID Calculator , can be found at https://www.seguid.org .

Pereira, Humberto↗

GOOML Big Kahuna Forecast Modeling and Genetic Optimization Files

This submission includes example files associated with the Geothermal Operational Optimization using Machine Learning (GOOML) Big Kahuna fictional power plant, which uses synthetic data to model a fictional power plant. A forecast was produced using the GOOML data model framework and fictional input data, and a genetic optimization is included which determines optimal flash plant parameters. The inputs and outputs associated with the forecast and genetic optimization are included. The input and output files consist of data, configuration files, and plots. A link to the Physics-Guided Neural Networks (phygnn) GitHub repository is also included, which augments a traditional neural network loss function with a generic loss term that can be used to guide the neural network to learn physical or theoretical constraints. phygnn is used by the GOOML framework to help integrate its machine learning models into the relevant physics and engineering applications. Note that the data included in this submission are intended to provide a demonstration of GOOML's capabilities. Additional files that have not been released to the public are needed for users to run these models and reproduce these results. Units can be found in the readme data resource.

15 GEOTHERMAL ENERGY↗

Potential structures - Applications of Machine Learning Techniques to Geothermal Play Fairway Analysis in the Great Basin Region, Nevada

This submission contains shapefiles, geotiffs, and symbology for the revised-from-Play-Fairway potential structures/structural settings used in the Nevada Geothermal Machine Learning project. Layers include potential structural setting ellipses, centroids, and distance-to-centroid raster. A submission linking the full GitHub repository for our machine learning Jupyter Notebooks will appear in the related datasets section of this page once available.

15 GEOTHERMAL ENERGY↗

Subsurface Characterization and Machine Learning Predictions at Brady Hot Springs Results

Geothermal power plants typically show decreasing heat and power production rates over time. Mitigation strategies include optimizing the management of existing wells - increasing or decreasing the fluid flow rates across the wells - and drilling new wells at appropriate locations. The latter is expensive, time-consuming, and subject to many engineering constraints, but the former is a viable mechanism for periodic adjustment of the available fluid allocations. Data and supporting literature from a study describing a new approach combining reservoir modeling and machine learning to produce models that enable strategies for the mitigation of decreased heat and power production rates over time for geothermal power plants. The computational approach used enables translation of sets of potential flow rates for the active wells into reservoir-wide estimates of produced energy and discovery of optimal flow allocations among the studied sets. In our computational experiments, we utilize collections of simulations for a specific reservoir (which capture subsurface characterization and realize history matching) along with machine learning models that predict temperature and pressure timeseries for production wells. We evaluate this approach using an "open-source" reservoir we have constructed that captures many of the characteristics of Brady Hot Springs, a commercially operational geothermal field in Nevada, USA. Selected results from a reservoir model of Brady Hot Springs itself are presented to show successful application to an existing system. In both cases, energy predictions prove to be highly accurate: all observed prediction errors do not exceed 3.68% for temperatures and 4.75% for pressures. In a cumulative energy estimation, we observe prediction errors that are less than 4.04%. A typical reservoir simulation for Brady Hot Springs completes in approximately 4 hours, whereas our machine learning models yield accurate 20-year predictions for temperatures, pressures, and produced energy in 0.9 seconds. This paper aims to demonstrate how the models and techniques from our study can be applied to achieve rapid exploration of controlled parameters and optimization of other geothermal reservoirs. Includes a synthetic, yet realistic, model of a geothermal reservoir, referred to as open-source reservoir (OSR). OSR is a 10-well (4 injection wells and 6 production wells) system that resembles Brady Hot Springs (a commercially operational geothermal field in Nevada, USA) at a high level but has a number of sufficiently modified characteristics (which renders any possible similarity between specific characteristics like temperatures and pressures as purely random). We study OSR through CMG simulations with a wide range of flow allocation scenarios. Includes a dataset with 101 simulated scenarios that cover the period of time between 2020 and 2040 and a link to the published paper about this project, where we focus on the Machine Learning work for predicting OSR's energy production based on the simulation data, as well as a link to the GitHub repository where we have published the code we have developed (please refer to the repository's readme file to see instructions on how to run the code). Additional links are included to associated work led by the USGS to identify geologic factors associated with well productivity in geothermal fields. Below are the high-level steps for applying the same modeling + ML process to other geothermal reservoirs: 1. Develop a geologic model of the geothermal field. The location of faults, upflow zones, aquifers, etc. need to be accounted for as accurately as possible 2. The geologic model needs to be converted to a reservoir model that can be used in a reservoir simulator, such as, for instance, CMG STARS, TETRAD, or FALCON 3. Using native state modeling, the initial temperature and pressure distributions are evaluated, and they become the initial conditions for dynamic reservoir simulations 4....

15 GEOTHERMAL ENERGY↗

A tutorial review of machine learning-based model predictive control methods

Abstract This tutorial review provides a comprehensive overview of machine learning (ML)-based model predictive control (MPC) methods, covering both theoretical and practical aspects. It provides a theoretical analysis of closed-loop stability based on the generalization error of ML models and addresses practical challenges such as data scarcity, data quality, the curse of dimensionality, model uncertainty, computational efficiency, and safety from both modeling and control perspectives. The application of these methods is demonstrated using a nonlinear chemical process example, with open-source code available on GitHub. The paper concludes with a discussion on future research directions in ML-based MPC.

Wu, Zhe [Department of Chemical and Biomolecular E↗

Numerical Model of IProTech PIP WEC Device

iProTech PIP wave energy converter (WEC) is a slack moored, single hull device with no moving parts in the water, joints or bearings. This submission includes data of the simulation, reports, and code for the iProTech PIP (WEC) project. The organization of the data included in the provided archive is detailed below and in the data description of the archive. The data teamer-iprotech-nrel folder includes and explains matlab and python code developed to hydrodynamically model the PIP WEC device in WEC-Sim. The subfolders cover the following steps: 1) report: explanatory information on device geometry 2) pip_mesher: python code to generate mesh panels from device profile data 3) wec-sim_models: matlab code to run WEC-Sim The data uploaded is a snapshot as of 11/02/2121 of code residing in a Github repository administered by David Ogden of NREL.

16 TIDAL AND WAVE POWER↗

MASK4 Test Campaign for Sandia WaveBot Device

This data and report details the findings from a wave tank test focused on production of useful work of a wave energy converter (WEC) device. The experimental system and test were specifically designed to validate models for power transmission throughout the WEC system. Additionally, the validity of co-design informed changes to the power take-off (PTO) were assessed and shown to provide the expected improvements in system performance. These data describe the "MASK4" wave tank test of the Sandia WaveBot device. The WaveBot device has been tested a number of times in different permutations at the US Navy's Maneuvering and Sea Keeping (MASK) basin. Each test in this series is referred to as MASK1, MASK2, etc. The WaveBot device was first tested in one degree of freedom (heave) in 2016. This MASK1 test focused primarily on system identification and modeling. After MASK1, major modifications were performed to improve the overall real-time control and measurement system, improve the heave drive train, and add surge and pitch degrees of freedom. The second set of testing, which was broken up in to two stages: MASK2A and MASK2B, focused on bench testing and closed-loop control performance as well as nonlinear modeling. MASK3 then focused on multi-input, multi-output modeling and control for maximization of electrical power. The attached report presents the results from MASK4, which focuses on detailed modeling of the power conversion chain and validation co-design principles by way of the introduction of a magnetic spring. The test log, report, and data from the MASK4 test of the WaveBot augmented with a tunable magnetic spring. Processing codes can be found at the Github link below.

16 TIDAL AND WAVE POWER↗

MBARI-WEC September and October 2022 Field Data

This data is needed to simulate a model of the MBARI-WEC (Monterey Bay Aquarium Research Institute, Wave Energy Converter device) in a simulation environment (e.g. Gazebo) for 56 observation dates in the time between September and October 2022, and to compare the simulation outputs to the corresponding field data of the physical MBARI-WEC. There were 50 observations chosen in Sept and 6 observations in Oct. To help understand terms below, a summary of the system can be found at the github link in the downloads section below. The Gazebo MBARI-WEC model is also provided, should users wish to simulate using this platform. There are 4 *mat files included. ................................................................................................................................................................................................................................... Spectrum and Simulation Inputs: September2022_spectrum_siminputs.mat and October2022_spectrum_siminputs.mat has data needed for simulation inputs in table format. These include the ocean spectrum for an observation and operating parameters of the MBARI-WEC during that observation. They are organized as rows representing an observation and columns representing data. For example, for the September *mat there are 50 rows. The first 7 columns are Datetime, sig_waveheight, peak_period, mean_period, heaveconedoor_status, pistonpos_mean, and scale_factor: - Datetime is the date and time the observation occurred in PST - sig_waveheight is the significant wave height of the ocean spectrum during that observation in meters - peak_period is the peak period of the ocean spectrum during that observation in seconds - mean_period is the mean period of the ocean spectrum during that observation in seconds - heaveconedoor_status is the status of the heave cone doors where 0 represents the doors are open and 1 represents they are closed - pistonpos_mean is the mean position of the PTO ram (piston) in meters - scale_factor is an additional factor of 0.5 --1.4 applied to a default damping relationship The next columns are data needed to represent the ocean spectrum. First are the frequencies [Hz] labeled as "f0-f38", then the variance density [m2/Hz] labeled as "vardens0-vardens38". October2022_spectrum_siminputs.mat follows as a similar format as above, but includes a larger amount of ocean spectrum frequencies and variance density elements. ................................................................................................................................................................................................................................... Field data: The field data is found in MBARIWEC_septdata.mat and MBARIWEC_octdata.mat for the observations of September and October, respectively. These contain data in a struct format. The struct contains the following fields for each observation: PC_BattCurr, PC_LoadCurr, PC_RPM, PC_Voltage, SC_Range, SC_Velocity, DateTime, where: - PC_BattCurr is the current flowing to or from the onboard batteries in Amps - PC_LoadCurr is the current flowing to the load dump in Amps - PC_RPM is the electric/hydraulic motor shaft speed (directly coupled) in RPM - PC_Voltage is the bus voltage at the power converter in Volts - SC_Range is the PTO ram (piston) position in meters where 0 is fully retracted and 2.03 is fully extended - SC_Velocity is the PTO ram (piston) velocity in meters/sec - DateTime is the date and time of the sampled field data in each observation in PST - Electric Power is equal to: PC_Voltage*(PC_BattCurr + PC_LoadCurr) in Watts For example, upon loading MBARIWEC_octdata.mat, the aforementioned fields would be loaded, each with {6x1} cells for the 6 observations chosen in October. Within the first cell of e.g. SC_Range would be sampled data representing the field data of the MBARIWEC PTO piston position for say, one hour, of the first October observation. The corresponding field DateTime would...

16 TIDAL AND WAVE POWER↗

Data for "Soil carbon dynamics during drying vs. rewetting: Importance of antecedent moisture conditions"

This dataset contains data used for the paper "Soil carbon dynamics during drying vs. rewetting: importance of antecedent moisture conditions". The Related References field will be updated with a full citation when available.Soil moisture influences soil carbon dynamics, including microbial growth and respiration. The response of such ‘soil respiration’ to moisture changes is generally assumed to be linear and reversible, i.e. to depend only on the current moisture state. Current models thus do not account for antecedent soil moisture conditions when determining soil respiration or the available substrate pool. We conducted a laboratory incubation to determine how the antecedent conditions of drought and flood influenced soil organic matter (SOM) chemistry, bioavailability, and respiration. We sampled soils from an upland coastal forest, Beaver Creek, WA USA, and subjected them to drying and rewetting treatments. For the drying treatment, field moist soils were saturated and then dried to 75, 50, 35, and 5 % saturation. In the rewetting treatment, field moist soils were air-dried and then rewet to 35, 50, 75, and 100 % saturation. We measured respiration and water extractable organic carbon (WEOC) concentrations and used 1H-NMR and FT-ICR-MS to characterize the WEOC pool across the treatments. The drying vs. wetting treatment strongly influenced SOM bioavailability, as rewet soils (with antecedent drought) had greater WEOC concentrations and respiration fluxes compared to the drying soils (with antecedent flood). In addition, air-dry soils had the highest WEOC concentrations, and the NMR-resolved peaks showed a strong contribution of protein groups in these soils. Both NMR and FT-ICR-MS analyses indicated increased contribution of complex aromatic groups/molecules in the rewet soils, compared to the drying soils. We suggest that drying introduced organic matter into the WEOC pool via desorption of aromatic molecules and/or by microbial cell lysis, and this stimulated microbial mineralization rates. Our work indicates that even short-term shifts in antecedent moisture conditions can strongly influence soil C dynamics at the core scale. The predictive uncertainties in current soil models may be reduced by a more accurate representation of soil water and C persistence that includes a mechanistic and quantitative understanding of the impact of antecedent moisture conditions.This dataset contains a compressed (.zip) archive of the data and R scripts used for this manuscript. The dataset includes files in .csv and .txt format, which can be accessed and processed using MS Excel or R. NMR data are provided as raw output data (accessed in Bruker TopSpin or MestreNova) as well as the MestreNova-processed files. This archive can also be accessed on GitHub at https://github.com/kaizadp/hysteresis_and_soil_carbon (DOI: 10.5281/zenodo.4432885).

1H-NMR↗

Data for "Soil texture and environmental conditions influence the biogeochemical responses of soils to drought and flooding"

This dataset contains data used for the paper "Soil texture and environmental conditions influence the biogeochemical responses of soils to drought and flooding". The Related References field will be updated with a full citation when available.Climate change is intensifying the global water cycle, with increased frequency of drought and flood. Water is an important driver of soil carbon dynamics, and it is crucial to understand how moisture disturbances will affect carbon availability and fluxes in soils. Here we investigate the role of water in substrate-microbe connectivity and soil carbon cycling under extreme moisture conditions. We collected soils from Alaska, Florida, and Washington USA, and incubated them under Drought and Flood conditions. Drought had a stronger effect on soil respiration, pore-water carbon, and microbial community composition than flooding. Soil response was not consistent across sites, and was influenced by site-level pedological and environmental factors. Soil texture and porosity can influence microbial access to substrates through the pore network, driving the chemical response. Further, the microbial communities are adapted to the historic stress conditions at their sites and therefore show site-specific responses to drought and flood.This dataset contains a compressed (.zip) archive of the data and R scripts used for this manuscript. The dataset includes files in .csv and .txt format, which can be accessed and processed using MS Excel or R. This archive can also be accessed on GitHub at https://github.com/kaizadp/TES_3Soils_2021 (DOI: 10.5281/zenodo.4792655).

54 ENVIRONMENTAL SCIENCES↗

ESS-DIVE guidelines for archiving terrestrial model data

This dataset contains supporting documents and images for ESS-DIVE terrestrial model data archiving guidelines.Terrestrial models are broadly defined as numerical models that couple both land dynamics and energy, water, carbon, or nutrient fluxes. We created these guidelines based on input from the U.S. Department of Energy’s Biological and Environmental Research land modeling community. The guidelines are intended to help modelers determine which components of their terrestrial model data associated with publication should be archived. Based on input from the land modeling community, the guidelines recommend archiving both model input and testing data, as well as code, script, and metadata. The guidelines also recommend archiving model data output, depending on the limitations set by data repositories. Lastly, we provide recommendations for bundling data files for publication as well as a discussion about tools that can facilitate model data archiving and reuse.This dataset is an archive of the associated GitHub repository for our model archiving guidelines (https://github.com/ess-dive-community/essdive-model-data-archiving-guidelines). The ‘README.pdf’ file gives a general introduction to the guidelines, and the ‘instructions.pdf’ file provides more detailed steps for following the guidelines. We also provide 2 figures in this data package: 1) a decision tree (model_data_guidelines_decision_tree.png) that can help users determine which components of their model data to archive. and 2) the ‘model_data_guidelines_flmd.png’ file depicts the different files that can be archived in addition to the model data itself. Lastly, we include 3 digitized tables from our associated manuscript and 3 CSV files with anonymized input from DOE scientists about the importance of different aspects of model data archiving from which we developed the guidelines.Dataset updates for v1.1.0: We updated this data package on 2021-11-22 in response to review comments on our related manuscript. In this update we removed one figure so that the model archiving guidelines are conveyed in text rather than an image. We updated the file-level metadata (FLMD) figure to be in accord with the most recent FLMD recommendations. We made minor edits to the README file to update the recommended citation and added two co-authors. We also added 6 new data files (3 are anonymized input from DOE scientists that helped to inform guidelines, and 3 are digitized tables from our manuscript.

54 ENVIRONMENTAL SCIENCES↗

Data for: Spatial access and resource limitations control carbon mineralization in soils

This dataset contains data and code used for the paper "Spatial access and resource limitations control carbon mineralization in soils", https://doi.org/10.1016/j.soilbio.2021.108427. Core-scale soil carbon fluxes are ultimately regulated by pore-scale dynamics of substrate availability and microbial access. These are constrained by physicochemical and biochemical phenomena (e.g. spatial access and hydrologic connectivity, physical occlusion, adsorption-desorption with mineral surfaces, nutrient and resource limitations). We conducted an experiment to determine how spatial access and resource limitations influence core-scale water-soluble soil organic matter (SOM) mineralization, and how these are regulated by antecedent moisture conditions. Intact soil cores were incubated at field-moist vs. drought conditions, after which they were saturated from above (to simulate precipitation) or below (to simulate groundwater recharge). Soluble carbon (acetate) and nitrogen (nitrate) forms were added to some cores during the rewetting process to alleviate potential nutrient limitations. Soil respiration was measured during the incubation, after which pore water was extracted from the saturated soils and analyzed for water soluble organic carbon concentrations and characterization. Our results showed that carbon (C) amendments increased the cumulative carbon dioxide (CO2) evolved from the soil cores, suggesting that the soils were C-limited. Drought and rewetting increased soil respiration, and there was a greater abundance of complex aromatic molecules in pore waters sampled from these soils. This newly available substrate appeared to alleviate nutrient limitations on respiration, because there were no further respiration increases with subsequent C and N amendments. We had hypothesized that respiration would be influenced by wetting direction, as simulated precipitation would mobilize C from the surface. However, as a main effect, this response was seen only in the C-amended soils, indicating that surface-C may not have been bioavailable. At the pore scale (pore water samples), drought and the C, N amendments caused a net loss of identified molecules when the soils were rewet from below, whereas wetting from above caused a net increase in identified molecules, suggesting that fresh inputs stimulated the C-and N-limited microbial populations present deeper in the soil profile. Our experiment highlights the complex and interactive role of antecedent moisture conditions, wetting direction, and resource limitations in driving core-scale C fluxes.This dataset contains a compressed (.zip) archive of the data and R scripts used for this manuscript. The dataset includes files in .csv format, which can be accessed and processed using MS Excel or R. This archive can also be accessed on GitHub at https://github.com/kaizadp/TES_spatial_access_2021 (DOI: 10.5281/zenodo.5522938).

54 ENVIRONMENTAL SCIENCES↗

Soil pore network response to freeze-thaw cycles in permafrost aggregates

This dataset contains data used for the paper "Pore network response to freeze-thaw cycles in permafrost aggregates". The Related References field will be updated with a full citation when available.Climate change in Arctic landscapes may increase freeze-thaw frequency within the active layer as well as newly thawed permafrost. A highly disruptive process, freeze-thaw can deform soil pores and alter the architecture of the soil pore network with varied impacts to water transport and retention, redox conditions, and microbial activity. Our objective was to investigate how freeze-thaw cycles impacted the pore network of newly thawed permafrost aggregates to improve understanding of what type of transformations can be expected from warming Arctic landscapes. We measured the impact of freeze-thaw on pore morphology, pore throat diameter distribution, and pore connectivity with X-ray computed tomography (XCT) using six permafrost aggregates with sizes of 2.5 cm3 from a mineral soil horizon (Bw; 28-50 cm depths) in Toolik, Alaska. Freeze-thaw cycles were performed using a laboratory incubation consisting of five freeze-thaw cycles (-10˚C to 20˚C) over five weeks. Our findings indicated decreasing spatial connectivity of the pore network across all aggregates with higher frequencies of singly connected pores following freeze-thaw. Water-filled pores that were connected to the pore network decreased in volume while the overall connected pore volumetric fraction was not affected. Shifts in the pore throat diameter distribution were mostly observed in pore throats ranges of 100 microns or less with no corresponding changes to the pore shape factor of pore throats. Responses of the pore network to freeze-thaw varied with aggregate, suggesting that initial pore morphology may play a role in driving freeze-thaw response. Our research suggests that freeze-thaw alters the microenvironment of permafrost aggregates during the incipient stage of deformation following permafrost thaw, impacting soil properties and function in Arctic landscapes undergoing transition. This dataset contains a compressed (.zip) archive of the data and R scripts used for this manuscript. The dataset includes files in .csv format, which can be accessed and processed using MS Excel or R. This archive can also be accessed on GitHub at https://github.com/Erin-Rooney/XCT-freezethaw (DOI: 10.5281/zenodo.5816355).

54 ENVIRONMENTAL SCIENCES↗