Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “interpretable models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Role of Neutrals Versus Transport in Determining the Pedestal Density Structure: Final Technical Report

In fusion devices the plasma density plays a crucial role in determining the fusion reaction rate and has a direct impact on the fusion gain of a given device. This density is in general regulated by the particle sources and transport near the plasma edge, which give rise to an edge density pedestal. When predicting the performance of future devices, this density pedestal is often prescribed, rather than predicted, due to a lack of models which allow confident extrapolation. This project aims to advance these models through the focused validation of theoretical models related to the transport of fueling neutral particles, and through interpretive transport modeling in present day fusion plasmas, in which the penetration of neutrals is altered to better simulate future reactor-like conditions. Achievements in theory and model validation under this project have advanced our understanding of how much of the edge density profile is set by transport versus direct ionization, enabling interesting projections to future burning plasma devices.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Understanding the Drivers of Atlantic Multidecadal Variability using a Stochastic Model Hierarchy

The relative importance of ocean and atmospheric dynamics in generating Atlantic Multidecadal Variability (AMV) remains an open question. Comparisons between climate models with SLAB and fully-dynamic (FULL) ocean components are often used to explore this question, but cannot reveal how individual ocean processes generate these differences. We build a hierarchy of physically interpretable stochastic models to investigate the contribution of two upper-ocean processes to AMV: the role of seasonal variation and mixed-layer entrainment. This interpretability arises from the stochastic model’s simplified representation of sea surface temperature (SST), considering only the local upper ocean response to white-noise atmospheric forcing and its impact on surface heat exchange. We focus on understanding differences between SLAB and FULL non-eddy resolving pre-industrial control simulations of the Community Earth System Model 1 (CESM), and estimate the stochastic model parameters from each respective simulation. Despite its simplicity, the stochastic model reproduces temporal characteristics of SST variability in the SPG, including reemergence, seasonal-to-interannual persistence and power spectra. Furthermore, unrealistically persistent SST of the CESM-SLAB ocean simulation is reproduced in the equivalent stochastic model configuration where the mixed-layer depth (MLD) is constant. The stochastic model also reveals that vertical entrainment primarily damps SST variability, thus explaining why SLAB exhibits larger SST variance than FULL. Here, the stochastic model driven by temporally stochastic, spatially coherent forcing patterns reproduces the canonical AMV pattern. However, the amplitude of low-frequency variability remains underestimated, suggesting a role for ocean dynamics beyond entrainment.

54 ENVIRONMENTAL SCIENCES↗

Assessing Metal Ion Assignment Accuracy in Protein Data Bank Models via Elemental Spectroscopy

Accurate representation of metal ions in macromolecular structures is critical for chemical interpretation, computational modeling, and machine-learning methods that rely on Protein Data Bank (PDB) entries. However, the elemental identity of metals modeled in crystallographic structures is often inferred indirectly and rarely validated experimentally. Here, we combine Particle Induced X-ray Emission (PIXE) and X-ray Fluorescence Spectroscopy (XRFS) to determine the elemental composition of protein samples used to generate 70 deposited metalloprotein crystal structures. By analyzing the original protein material employed for crystallization, but before the addition of crystallization buffer solutions, we assess whether the modeled metal ions in deposited structures are consistent with experimentally detectable elemental content. We find that in a majority of cases, the metals modeled in the corresponding PDB entries are inconsistent with the metals present in the protein samples before crystallization, or that additional metals are present but not represented in the structural models. Spectroscopic results were integrated with automated crystallographic validation metrics, including real-space Z-difference (RSZD) analysis and systematic rerefinement, to evaluate atomic-number mismatch at metal sites. PIXE and XRFS show strong agreement for dominant elemental signals and provide complementary, scalable approaches for identifying suspect metal assignments. This work does not address physiological or functional metalation but instead highlights a widespread data integrity issue in deposited macromolecular structures, PDB-wide. These results establish an experimentally corroborated link between elemental identity and crystallographic validation metrics, enabling the large-scale detection of chemically inconsistent annotations in structural databases used for computational modeling and machine learning.

Crystallization↗

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis↗

Searches for new phenomena in events with two leptons, jets, and missing transverse momentum in 139 fb –1 of √s = 13 TeV $pp$ collisions with the ATLAS detector

Searches for new phenomena inspired by supersymmetry in final states containing an e + e – or μ + μ – pair, jets, and missing transverse momentum are presented. These searches make use of proton–proton collision data with an integrated luminosity of 139 fb –1 , collected during 2015–2018 at a centre-of-mass energy √s = 13 TeV by the ATLAS detector at the Large Hadron Collider. Two searches target the pair production of charginos and neutralinos. One uses the recursive-jigsaw reconstruction technique to follow up on excesses observed in 36.1 fb –1 of data, and the other uses conventional event variables. The third search targets pair production of coloured supersymmetric particles (squarks or gluinos) decaying through the next-to-lightest neutralino ($\tilde{χ}$$^{0}_{2}$) via a slepton ($\tilde{ℓ}$) or Z boson into ℓ + ℓ – $\tilde{χ}$$^{0}_{1}$ , resulting in a kinematic endpoint or peak in the dilepton invariant mass spectrum. The data are found to be consistent with the Standard Model expectations. Results are interpreted using simplified models and exclude masses up to 900 GeV for electroweakinos, 1550 GeV for squarks, and 2250 GeV for gluinos.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Programmatic Advantages of Linear Equivalent Seismic Models

Underground explosions nonlinearly deform the surrounding earth material and can interact with the free surface to produce spall. However, at typical seismological observation distances the seismic wavefield can be accurately modeled using linear approximations. Although nonlinear algorithms can accurately simulate very near field ground motions, they are computationally expensive and potentially unnecessary for far field wave simulations. Conversely, linearized seismic wave propagation codes are orders of magnitude faster computationally and can accurately simulate the wavefield out to typical observational distances. Thus, devising a means of approximating a nonlinear source in terms of a linear equivalent source would be advantageous both for scenario modeling and for interpretation of seismic source models that are based on linear, far-field approximations. This allows fast linear seismic modeling that still incorporates many features of the nonlinear source mechanics built into the simulation results so that one can have many of the advantages of both types of simulations without the computational cost of the nonlinear computation. In this report we first show the computational advantage of using linear equivalent models, and then discuss how the near-source (within the nonlinear wavefield regime) environment affects linear source equivalents and how well we can fit seismic wavefields derived from nonlinear sources.

58 GEOSCIENCES↗

Overview of IMPACT Data Acquisition System and Data Reduction Process

This report documents the development of the data acquisition system (DAS) and data reduction methodologies for the Irradiated Material Property Accelerated Characterization Test (IMPACT) experiment at the Advanced Test Reactor (ATR). The IMPACT experiment is designed to enable in-pile measurement of thermal conductivity in metallic nuclear fuels, specifically U-10Zr, using an instrumented thermal conductivity probe. The DAS supports both passive temperature monitoring and active thermal interrogation of the probe through controlled AC and DC excitation. Significant modifications to laboratory-scale systems were required to accommodate the higher resistance paths associated with the in-pile application. Custom electronics and relay-controlled measurement sequencing were developed to enable the measurement and sufficient power delivery to the sensing region. A reduced-order, axisymmetric thermal model based on the thermal quadrupoles method is presented to support data interpretation. This model enables efficient evaluation of transient heat transfer behavior and facilitates solution of the inverse problem required to extract thermal properties from measured signals. Multiple boundary condition formulations are discussed to address varying experimental time scales and geometries. Additionally, machine learning techniques are introduced to support data reduction and improve confidence in inverse solutions. Convolutional neural networks are applied to identify the presence of gas gaps and other evolving geometric features that significantly impact thermal response during irradiation. These efforts contribute to the broader integration of digital twin frameworks and real-time modeling capabilities within the Advanced Fuels Campaign.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Search for supersymmetry in final states with two or three soft leptons and missing transverse momentum in proton-proton collisions at $\sqrt{s}$ = 13 TeV

A search for supersymmetry in events with two or three low-momentum leptons and missing transverse momentum is performed. The search uses proton-proton collisions at $\sqrt{s}$ = 13 TeV collected in the three-year period 2016–2018 by the CMS experiment at the LHC and corresponding to an integrated luminosity of up to 137 fb -1 . The data are found to be in agreement with expectations from standard model processes. The results are interpreted in terms of electroweakino and top squark pair production with a small mass difference between the produced supersymmetric particles and the lightest neutralino. For the electroweakino interpretation, two simplified models are used, a wino-bino model and a higgsino model. Exclusion limits at 95% confidence level are set on $^{\sim0}_{χ2}/^{\sim±}_{χ1}$ masses up to 275 GeV for a mass difference of 10 GeV in the wino-bino case, and up to 205(150) GeV for a mass difference of 7.5 (3) GeV in the higgsino case. The results for the higgsino are further interpreted using a phenomenological minimal supersymmetric standard model, excluding the higgsino mass parameter μ up to 180 GeV with the bino mass parameter M 1 at 800 GeV. In the top squark interpretation, exclusion limits are set at top squark masses up to 540 GeV for four-body top squark decays and up to 480 GeV for chargino-mediated decays with a mass difference of 30 GeV.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

pyTCR: A tropical cyclone rainfall model for python

pyTCR is a climatology software package developed in the Python programming language. It integrates the capabilities of several legacy physical models and increases computational efficiency to allow rapid estimation of tropical cyclone (TC) rainfall consistent with the large-scale environment. Specifically, pyTCR implements a horizontally distributed and vertically integrated model [Zhu et al., 2013] for simulating rainfall driven by TCs. Along storm tracks, rainfall is estimated by computing the cross-boundary-layer, upward water vapor transport caused by different mechanisms including frictional convergence, vortex stretching, large-scale baroclinic effect (i.e., wind shear), topographic forcing, and radiative cooling [Lu et al., 2018]. The package provides essential functionalities for modeling and interpreting spatio-temporal TC rainfall data. pyTCR requires a limited number of model input parameters, making it a convenient and useful tool for analyzing rainfall mechanisms driven by TCs. To sample rare (most intense) rainfall events that are often of great societal interest, pyTCR adapts and leverages outputs from a statistical-dynamical TC downscaling model [Lin et al., 2023] capable of rapidly generating a large number of synthetic TCs given a certain climate. As a result, pyTCR significantly reduces computational effort and improves the efficiency in capturing extreme TC rainfall events at the tail of the distributions from limited datasets. Furthermore, the TC downscaling model is forced entirely by large-scale environmental conditions from reanalysis data or coupled General Circulation Models (GCMs), simplifying the projection of TC-induced rainfall and wind speed under future climate using pyTCR. Finally, pyTCR can be coupled with hydrological and wind models to assess risks associated with independent and compound events (e.g., storm surges and freshwater flooding).

54 ENVIRONMENTAL SCIENCES↗

Deciphering the Solvation Structure of Aqueous ZnCl 2 Solutions from X-ray Absorption Spectra Using the Interpretable Graph Neural Network

Machine learning (ML) provides powerful pathways for predicting spectroscopic observables from atomic structures, but its broader impact depends on making model predictions interpretable in terms of physical and chemical principles. Here, we introduce a physics-guided graph neural network (GNN) model that predicts Zn K-edge X-ray spectroscopy (XAS) spectra of aqueous ZnCl 2 solutions. Training data are generated from ab initio XAS calculations on molecular dynamics snapshots obtained using a machine learning interatomic potential. The GNN reproduces experimental spectra across concentrations from dilute (<0.1 m) to highly concentrated (30 m, “water-in-salt”) regimes and scales efficiently to large, disordered liquid systems beyond the reach of conventional ab initio approaches. Gradient-based attribution analysis reveals that the model learns physically meaningful structure-spectrum relationships. Ligand-specific attributions reflect orbital hybridization patterns and the origin of the excitations derived from the density functional theory. Bond-length attributions recover spectral shifts consistent with multiple-scattering theory. Finally, this work bridges data-driven prediction with electronic-structure theory, establishing a general paradigm for interpretable ML that links atomic structure, electronic structure, and spectroscopic observables.

25 ENERGY STORAGE↗

A Unified Workflow for Sensitivity-Based Kinetic Analysis in Microkinetic Models

Degrees of rate control (DRC), apparent activation energies, and apparent reaction orders are established local sensitivity diagnostics for interpreting microkinetic models, but applying them routinely to large mechanisms often requires substantial reaction-specific bookkeeping, perturbation design, and postprocessing. Here, in this study, we present a unified derivative-based workflow that evaluates these quantities from a single compiled reaction-network model and target-rate definition. For any user-provided microkinetic model, the workflow compiles the mechanism into stoichiometrically consistent mass-action rate equations, solves the surface dynamics, and uses automatic differentiation to compute sensitivities with respect to rate constants, temperature, and gas partial pressures. By combining their calculations in the same framework, the workflow clearly demonstrates the relationships between different DRCs and the apparent activation energy. Using existing examples of propylene partial oxidation and methane oxidation on Pd(100), we verify expected transient redistribution of rate control, distinguish net Campbell DRCs from one-sided directional sensitivities, and show how apparent activation energy can be reconstructed either from one-sided DRCs or from state-based DRCs while critical mechanistic insights are obtained consistently. In the methane oxidation case, a pathway-subset test further illustrates how a simplified mechanism preserves key kinetic signatures of a full model, showing the potential of our user-friendly tool for model construction beyond kinetic analysis.

36 MATERIALS SCIENCE↗

Deep learning of dynamically responsive chemical Hamiltonians with semiempirical quantum mechanics

Conventional machine-learning (ML) models in computational chemistry learn to directly predict molecular properties using quantum chemistry only for reference data. While these heuristic ML methods show quantum-level accuracy with speeds several orders of magnitude faster than traditional quantum chemistry methods, they suffer from poor extensibility and transferability; i.e., their accuracy degrades on large or new chemical systems. Incorporating quantum chemistry frameworks into the ML models directly solves this problem. Here we take the structure of semiempirical quantum mechanics (SEQM) methods to construct dynamically responsive Hamiltonians. SEQM methods use empirical parameters fitted to experimental properties to construct reduced-order Hamiltonians, facilitating much faster calculations than ab initio methods but with compromised accuracy. By replacing these static parameters with machine-learned dynamic values inferred from the local environment, we greatly improve the accuracy of the SEQM methods. Trained on molecular energies and atomic forces, these dynamically generated Hamiltonian parameters show a strong correlation with atomic hybridization and bonding. Trained with only about 60,000 small organic molecular conformers, the resulting model retains interpretability, extensibility, and transferability when testing on much larger chemical systems and predicting various molecular properties. Overall, this work demonstrates the virtues of incorporating physics-based descriptions with ML to develop models that are simultaneously accurate, transferable, and interpretable.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A procedure for rule extraction from a Self-Organising plasma disruption predictor for JET

In a previous paper, a Self-Organizing Map had proven to be able to identify the regions of the plasma operative space characterizing the pre-disruptive phase at JET without relying on any a priori information. One of the strengths of this disruption predictor lies in its inherent self-organization capability. The Self-Organizing Map discovers non-trivial relationships and captures the complicated interplay of device diagnostics on the internal plasma states directly from the experimental data. Moreover, the provided model allows the visualization of high-dimensional plasma parameters and facilitates easy interrogation of the model to understand the reasons behind its correlations. In this paper, an additional step is taken towards the interpretability of models for predicting disruptions by training a Decision Tree to classify the plasma states according to the interpretation provided by the Self-Organizing Map (stable or at high risk of disruptions). The Decision tree provides a set of rules which describe the transition of the plasma towards the pre-disruptive phase as visualized in the Self-Organizing Map. The obtained rules for the database explored in the study identify four regions in the map, two of which are at risk of disruption. These regions correspond to partitions of a 3D space based on the peaking factors of the core and divertor radiation, as well as the Locked Mode. The agreement between the Self-Organizing Map answers and the rules supplied by the Decision Tree is confirmed by the comparison of the performance exhibited by the two models in the prediction of disruptions.

Setzu, Samuele [Univ. of Cagliari, Monserrato, Cag↗

Steptoe Valley NV Data Compilation: Understanding a Stratigraphic Hydrothermal Resource through Geophysical Imaging

Sandia National Laboratories partnered with a multi-disciplinary group of subject matter experts to evaluate a stratigraphic geothermal resource in Steptoe Valley, Nevada using both established and novel geophysical imaging techniques. Provided here are a compilation of newly acquired data over the area and select modeling efforts. This encompasses a 3D geological model (inclusive of full Leapfrog files, Leapfrog viewer files, and XYZ data for faults and stratigraphy) with embedded geophysical modeling, controlled-source electromagnetic (CSEM) and magnetotelluric (MT) data packages, aqueous spring geochemistry data, seismic reflection interpretations, and a gravity data package. The stratigraphic reservoir in Steptoe Valley was previously discovered during oil and gas exploration. Subsequent studies, such as the Nevada Play Fairway Analysis, added data which further highlighted potential resource targets in the basin. Geophysical surveys, complimented with refined geologic mapping and geochemical sampling, were deployed to further characterize the resource. The resulting 3D geologic interpretation, conceptual model refinements, and reservoir simulations suggest that a power-capable reservoir is economically accessible in the Paleozoic carbonates of the deep/central basin. Additional geophysical characterization and exploration drilling efforts are recommended to calibrate interpretation and determine where/how to potentially develop the Steptoe resource. The geophysical tools, interpretations, lessons learned, and publicly available data generated by this study establish an exploration methodology to inform decisions for successful development of stratigraphic reservoirs.

15 GEOTHERMAL ENERGY↗

Integrating continuous atmospheric boundary layer and tower-based flux measurements to advance understanding of land-atmosphere interactions

The atmospheric boundary layer mediates the exchange of energy, matter, and momentum between the land surface and the free troposphere, integrating a range of physical, chemical, and biological processes and is defined as the lowest layer of the atmosphere (ranging from a few meters to 3 km). In this review, we investigate how continuous, automated observations of the atmospheric boundary layer can enhance the scientific value of co-located eddy covariance measurements of land-atmosphere fluxes of carbon, water, and energy, as are being made at FLUXNET sites worldwide. We highlight four key opportunities to integrate tower-based flux measurements with continuous, long-term atmospheric boundary layer measurements: (1) to interpret surface flux and atmospheric boundary layer exchange dynamics and feedbacks at flux tower sites, (2) to support flux footprint modelling, the interpretation of surface fluxes in heterogeneous terrain, and quality control of eddy covariance flux measurements, (3) to support regional-scale modeling and upscaling of surface fluxes to continental scales, and (4) to quantify land-atmosphere coupling and validate its representation in Earth system models. Adding a suite of atmospheric boundary layer measurements to eddy covariance flux tower sites, and supporting the sharing of these data to tower networks, would allow the Earth science community to address new emerging research questions, better interpret ongoing flux tower measurements, and would present novel opportunities for collaborations between FLUXNET scientists and atmospheric and remote sensing scientists.

54 ENVIRONMENTAL SCIENCES↗

Automated classification of big X-ray diffraction data using deep learning models

Abstract In current in situ X-ray diffraction (XRD) techniques, data generation surpasses human analytical capabilities, potentially leading to the loss of insights. Automated techniques require human intervention, and lack the performance and adaptability required for material exploration. Given the critical need for high-throughput automated XRD pattern analysis, we present a generalized deep learning model to classify a diverse set of materials’ crystal systems and space groups. In our approach, we generate training data with a holistic representation of patterns that emerge from varying experimental conditions and crystal properties. We also employ an expedited learning technique to refine our model’s expertise to experimental conditions. In addition, we optimize model architecture to elicit classification based on Bragg’s Law and use evaluation data to interpret our model’s decision-making. We evaluate our models using experimental data, materials unseen in training, and altered cubic crystals, where we observe state-of-the-art performance and even greater advances in space group classification.

Chemistry↗

Enhanced descriptor identification and mechanism understanding for catalytic activity using a data-driven framework: revealing the importance of interactions between elementary steps

We report accurate identification of descriptors for catalytic activities has long been essential to the in-depth understanding of catalysis and recently to set the basis for catalyst screening. However, commonly used methods suffer from low accuracy in predictability. This study reports an enhanced approach to accurately identify the descriptors from a kinetic dataset using a machine learning (ML) surrogate model. CO hydrogenation to methanol over Cu-based catalysts was taken as a case study. Our model captures not only the contribution from individual elementary steps but also the interaction between relevant steps within a reaction network, which was found to be essential for high accuracy. As a result, six effective descriptors are identified, which are accurate enough to ensure the trained gradient boosted regression (GBR) model for good prediction of the methanol turnover frequency (TOF) over metal (M)-doped Cu(111) model surfaces (M = Au, Cu, Pd, Pt, Ni). More importantly, going beyond the purely mathematical ML model, the catalytic role of each identified descriptor can be revealed by using model-agnostic interpretation tools, which enhances the insight into the promoting effect of alloying. The trained GBR model outperforms the conventional derivative-based methods in terms of both the predictability and the mechanism understanding. It opens alternative possibilities toward accurate descriptor-based rational catalyst optimization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

2D reactive transport model of shale chemical weathering and biogeochemical fluxes along a mountainous hillslope, East River Watershed, Colorado: Input files and simulation results

This data package contains input files and simulation results for a two-dimensional (2D) reactive transport model used to quantitatively analyze the coupled hydrological and biogeochemical processes governing shale weathering and associated biogeochemical fluxes under realistic environmental conditions in the high-elevation East River Watershed. These data support the conclusions presented in Stolze et al. (Water Resources Research, under review), "Model-based interpretation of solute exports and carbon partitioning during shale weathering in a mountainous hillslope". The model simulates atmospheric-subsurface gas exchange, subsurface water flow, and shale weathering processes under dynamic, year-scale conditions along a shale-underlain hillslope located in the East River watershed. The simulations were performed using the PFLOTRAN flow and reactive transport code and executed on the Perlmutter supercomputer to leverage its large-scale parallel computing capabilities. The data package contains two zipped folders, "model_input_files" and "simulation_results", and one readme.txt file. "model_input_files" contains the necessary input files to run the calibrated base-base model presented in Stolze et al. (Water Resources Research, under review). "simulation_results" contains a single hdf5 file ("Output_2D_hillslope_model.h5") which includes the results of simulation performed using the base-case model. This file can be opened with HDFView 3.1.4, Python, or MATLAB. "readme.txt" contains relevant information about the base-case model and provides guidelines on how to run the associated input files provided in the folder "model_input_files". Furthermore, readme.txt provides information regarding the model results provided in "Output_2D_hillslope_model.h5" such as matrix dimensionality and output units. Field datasets used to evaluate model performance were collected at three monitoring wells located along a hillslope transect (PLM1, PLM2, and PLM3). Dissolved ion concentration data were collected from November 2016 to October 2021 for Ca, Mg, DIC, Na, K, SO4 (Dong et al., 2025 - dic_npoc_data_2014_2024.zip - DOI:10.15485/1660459; Williams et al., 2025 - anion_data_2014_2024.zip - DOI:10.15485/1668054; Dong et al., 2025 - cation_data_2014_2024.zip - DOI:10.15485/1668055). Note that we used the files named er_PLM1_xx_yy, er_PLM2_xx_yy, and er_PLM3_xx_yy where xx stands for the name of the aqueous species and yy stands for the depth where the measurements were performed. Soil water content ([0 - 1] m) and water table depth were collected from November 2016 to October 2021 (Wan et al., 2024 - Dynamic_water_table__depthsFig2b.csv and Soil_water_content_Fig4e.csv - DOI:10.15485/2322567). Gaseous CO2 concentration were collected from October 2020 to December 2021(Wan et al., 2024 - Soil_CO2_concentrations_Fig4h.csv - DOI:10.15485/2322567) Gaseous CO2 flux from the subsurface to the atmosphere were collected in the vicinity of PLM2 from October 2019 to May 2022 (Wu et al., 2025). Soil microbial biomass concentration was measured from August 2016 to June 2017 (Sorensen et al., 2019 - 2017_East_River_Pumphouse_Microbial_Biomass__1_.csv - DOI:10.15485/1577267) All field data are published as CSV files compatible with Microsoft Excel, MATLAB, and Python, or as text files. The coordinates of the monitoring wells and the CO2(g) flux sensor in the coordinate system WGS84 are: -PLM1: [38.9197710 ; -106.9492750] -PLM2: [38.9201580 ; -106.9487170] -PLM3: [38.9207843 ; -106.9483668] -PLM4: 38.9210060 ; -106.9479528] -CO2(g) flux sensor: [38.9199180 ; -106.9489906] ------------------------------------------------------------------------------------------- This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231. This research used resources of the National Energy Research Scientific Computing Center (NERSC), a Department of Energy User Facility using NERSC award BER-ERCAP 23980, BER-ERCAP 28550, and BER-ERCAP 33789.

54 ENVIRONMENTAL SCIENCES↗