Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sampling methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

PySAGES: flexible, advanced sampling methods accelerated with GPUs

Abstract Molecular simulations are an important tool for research in physics, chemistry, and biology. The capabilities of simulations can be greatly expanded by providing access to advanced sampling methods and techniques that permit calculation of the relevant underlying free energy landscapes. In this sense, software that can be seamlessly adapted to a broad range of complex systems is essential. Building on past efforts to provide open-source community-supported software for advanced sampling, we introduce PySAGES, a Python implementation of the Software Suite for Advanced General Ensemble Simulations (SSAGES) that provides full GPU support for massively parallel applications of enhanced sampling methods such as adaptive biasing forces, harmonic bias, or forward flux sampling in the context of molecular dynamics simulations. By providing an intuitive interface that facilitates the management of a system’s configuration, the inclusion of new collective variables, and the implementation of sophisticated free energy-based sampling methods, the PySAGES library serves as a general platform for the development and implementation of emerging simulation techniques. The capabilities, core features, and computational performance of this tool are demonstrated with clear and concise examples pertaining to different classes of molecular systems. We anticipate that PySAGES will provide the scientific community with a robust and easily accessible platform to accelerate simulations, improve sampling, and enable facile estimation of free energies for a wide range of materials and processes.

Chemistry↗

Development of a phonon-based sampling method for thermal neutron scattering data

Simulations of reactor systems require access to accurate nuclear data. For many systems, thermal neutron scattering data can have large effects on the eigenvalue and neutron flux distributions. Inelastic thermal neutron scattering can excite or de-excite vibrational, rotational, and translational modes in a material, so thermal scattering evaluations are often obtained by summing over the number of phonons created/destroyed by a scattering event. In recent years, the thermal scattering cross sections and angular distributions have greatly improved in accuracy, but the format in which this data is delivered to simulation codes has remained virtually unchanged. Thermal scattering data is typically either compiled into large tables and sorted by incoming neutron energy, outgoing neutron energy, scattering angle, and material temperature, or represented as cumulative distribution functions of momentum exchange or energy exchange. Either method can be quite memory intensive when fine bins are used. In an effort to decrease the amount of space that processed thermal scattering data requires, an alternate format is proposed. The phonon-based sampling method introduced here can sample the number of phonons excited for each collision, the change in neutron energy, and the scattering angle while avoiding pre-computed angular bins and limiting the amount of data that is dependent on incoming energy. Through this method, the generation and storage of large interpolation tables is avoided, which could have benefits in both memory storage and accuracy. While the initial implementation of this method is slower than current alternatives, it is significantly more resistant to grid coarseness errors and has good potential for improvement. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Laser ultrasonic imaging of subsurface defects with the linear sampling method

Laser ultrasonics is a remote nondestructive evaluation technique suitable for real-time monitoring of fabrication processes in semiconductor metrology, advanced manufacturing, and other applications where non-contact, high fidelity measurements are required. Here we investigate laser ultrasonic data processing approaches to reconstruct images of subsurface side drilled holes in aluminum alloy specimens. We demonstrate through simulation that the model-based linear sampling method (LSM) can perform accurate shape reconstruction of single and multiple holes and produce images with well-defined boundaries. We experimentally confirm that LSM produces images that represent the internal geometric features of an object, some of which may be missed by conventional imaging.

42 ENGINEERING↗

An interoperable implementation of collective‐variable based enhanced sampling methods in extended phase space within the OpenMM package

Collective variable (CV)-based enhanced sampling techniques are widely used today for accelerating barrier-crossing events in molecular simulations. A class of these methods, which includes temperature accelerated molecular dynamics (TAMD)/driven-adiabatic free energy dynamics (d-AFED), unified free energy dynamics (UFED), and temperature accelerated sliced sampling (TASS), uses an extended variable formalism to achieve quick exploration of conformational space. These techniques are powerful, as they enhance the sampling of a large number of CVs simultaneously compared to other techniques. Extended variables are kept at a much higher temperature than the physical temperature by ensuring adiabatic separation between the extended and physical subsystems and employing rigorous thermostatting. Here, in this work, we present a computational platform to perform extended phase space enhanced sampling simulations using the open-source molecular dynamics engine OpenMM. The implementation allows users to have interoperability of sampling techniques, as well as employ state-of-the-art thermostats and multiple time-stepping. This work also presents protocols for determining the critical parameters and procedures for reconstructing high-dimensional free energy surfaces. As a demonstration, we present simulation results on the high dimensional conformational landscapes of the alanine tripeptide in vacuo, tetra-N-methylglycine (tetra-sarcosine) peptoid in implicit solvent, and the Trp-cage mini protein in explicit water.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Derivation and verification of the direct-sampling method for simulating Monte Carlo flight paths in tetrahedral meshes with linear finite-element cross sections

This paper provides a derivation of a direct-sampling approach for modeling continuously varying cross sections in tetrahedral-mesh-based Monte Carlo codes. Specifically, cross sections are spatially approximated using linear nodal finite elements. A linearization strategy is provided for non-linearly varying cross sections. The method is verified against seven analytical pure-absorber test problems. These test problems also highlight the benefit of using linear finite elements over element-wise-constant cross sections.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Radioisotope Analysis of Wastewater from Livermore Site Retention Tanks by Gel Laboratory Gross Alpha, Gross Beta and Tritium Sampling Method

Lawrence Livermore National Laboratory discharged approximately 4.4% of the City of Livermore’s total wastewater in 2022 (LLNL’s Annual Site Environmental Report, Chapter 5, 2022). This volume includes wastewater from Sandia National Laboratories (SNL) and some process wastewater from Site 300. Due to the high volume and constituents of the discharge, LLNL works alongside the City of Livermore under permit #1250, requiring wastewater generated to be monitored and sampled in accordance with permit limits. Process wastewater, from buildings with the highest risk to sewer, is collected by wastewater retention tanks throughout the Livermore Site and sampled prior to discharge. Domestic wastewater directly discharges to sanitary sewer. To maintain permit requirements and ensure proper wastewater discharge practices, an internal wastewater audit was conducted during the summer of 2023. Current wastewater practices, regulatory knowledge and risk management across various Livermore Site buildings were evaluated. Workspaces connected to a wastewater retention tank and sanitary sewer drains were major focus areas. Data collected from walk-throughs prompted further evaluation as many practices were reported to be done based on historical usage. A table of concerns was created to showcase reasons for auditing and proceeding action. An analysis of current and historical retention tank usage throughout LLNL Livermore Site buildings with radioisotope results over a 5-year period from 2019 to 2024, was done to assess building trends and any significant changes throughout the 5-year period. Analytes evaluated were Gross Alpha, Gross Beta and Tritium (GABT) of eleven buildings at the Livermore Site, posing the highest risk to sanitary sewer for radioisotopes.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

A comprehensive study of non-adaptive and residual-based adaptive sampling for physics-informed neural networks

Physics-informed neural networks (PINNs) have shown to be effective tools for solving both forward and inverse problems of partial differential equations (PDEs). PINNs embed the PDEs into the loss of the neural network using automatic differentiation, and this PDE loss is evaluated at a set of scattered spatio-temporal points (called residual points). The location and distribution of these residual points are highly important to the performance of PINNs. However, in the existing studies on PINNs, only a few simple residual point sampling methods have mainly been used. Here, we present a comprehensive study of two categories of sampling for PINNs: non-adaptive uniform sampling and adaptive nonuniform sampling. We consider six uniform sampling methods, including (1) equispaced uniform grid, (2) uniformly random sampling, (3) Latin hypercube sampling, (4) Halton sequence, (5) Hammersley sequence, and (6) Sobol sequence. We also consider a resampling strategy for uniform sampling. To improve the sampling efficiency and the accuracy of PINNs, we propose two new residual-based adaptive sampling methods: residual-based adaptive distribution (RAD) and residual-based adaptive refinement with distribution (RAR-D), which dynamically improve the distribution of residual points based on the PDE residuals during training. Hence, we have considered a total of 10 different sampling methods, including six non-adaptive uniform sampling, uniform sampling with resampling, two proposed adaptive sampling, and an existing adaptive sampling. We extensively tested the performance of these sampling methods for four forward problems and two inverse problems in many setups. Our numerical results presented in this study are summarized from more than 6000 simulations of PINNs. Here, we show that the proposed adaptive sampling methods of RAD and RAR-D significantly improve the accuracy of PINNs with fewer residual points for both forward and inverse problems. Furthermore, the results obtained in this study can also be used as a practical guideline in choosing sampling methods.

97 MATHEMATICS AND COMPUTING↗

Sample Preparation Methods for Targeted Single-Cell Proteomics

We compared three cell isolation and two proteomic sample preparation methods for single-cell and near-single-cell analysis. Whole blood was used to quantify hemoglobin (Hb) and glycated-Hb (gly-Hb) in erythrocytes using targeted mass spectrometry and stable isotope-labeled standard peptides. Each method differed in cell isolation and sample preparation as follows: 1) FACS and automated preparation in one-pot for trace samples (autoPOTS); 2) limited dilution via microscopy and a novel rapid one-pot sample preparation method that circumvented the need for the solid-phase extraction, low-volume liquid handling instrumentation and humidified incubation chamber; and 3) CellenONE-based cell isolation and the same one-pot sample preparation method used for limited dilution. Only the CellenONE device routinely isolated single-cells from which Hb was measured to be 540–660 amol per red blood cell (RBC), which was comparable to the calculated SI reference range for mean corpuscular hemoglobin (390–540 amol/RBC). FACSAria sorter and limited dilution could routinely isolate single-digit cell numbers, to reliably quantify CMV-Hb heterogeneity. Finally, we observed that repeated measures, using 5–25 RBCs obtained from N = 10 blood donors, could be used as an alternative and more efficient strategy than single RBC analysis to measure protein heterogeneity, which revealed multimodal distribution, unique for each individual.

59 BASIC BIOLOGICAL SCIENCES↗

Sampling Rare Events in Aqueous Systems Using Molecular Simulations

Birth of a new distinct phase is a phenomenon encountered in a myriad of processes, and has wide ranging consequences in material processing, biological self-assembly, separations and several other processes. Several phase transitions are nucleation driven. The nucleation events occur over nanosecond timescales and involve hundreds to thousands of molecules. These length and timescales are difficult to access in experiments, thereby making experimental studies of nucleation challenging. On the other hand, molecular simulations sample the nanosecond and nanometer scales making them ideal to study nucleation. However, nucleation is a rare event, meaning that the waiting time to observe one nucleation event is significant. This makes simulation studies of rare events challenging. The project focused on a multi-pronged approach to address such challenges to develop the next generation rare event sampling methods for molecular simulations. The key outcomes of our work include developing more effective methods for sampling rare events, utilizing machine learning to better elucidate nucleation mechanisms, development of software for easy implementation of the methodologies, and applications of the methods to realistic systems to push the method applicability beyond model systems. Overall, this work has enabled pushing the frontiers of molecular simulations to study rare events with a focus on nucleation in aqueous solutions.

36 MATERIALS SCIENCE↗

End-Use Savings Shapes Measure Documentation: Dispatch Schedule Generation for Demand Flexibility Measures

This supplemental document describes the methodology used for determining the dispatch timing of various EUSS demand flexibility measures. Demand flexibility measures are designed to reduce/dispatch electricity demand in buildings during especially beneficial/critical times. The method used in this work utilizes predictions of building loads to generate a schedule that reflects the periods when the building's daily peak load occurs to support decision making in demand flexibility measures. The dispatch schedule generation method described in this document creates an hourly schedule that includes a load dispatch (peak) window for each day for a whole year based on load prediction, with options using different prediction methods: perfect prediction, bin-sampling method, fixed schedule, and outdoor air temperature (OAT)-based prediction method. The perfect prediction method performs a simulation to obtain the annual load profile as predicted load, representing the scenario of perfect load prediction. The bin-sampling method (1) categorizes days into representative bins by temperature characteristics, (2) performs simulations on sample days from each of those bins to create representative (or predicted) load, and (3) assigns representative loads for all days in a year based on the bin categorization. The fixed schedule method defines uniform start and end time of peak window with assumed fixed daily peak time, for all days in a season or a year. The OAT-based prediction method uses the statistics of OAT (minimum and maximum) as the indicators of peak load, with specified delay response time from building loads to temperature. Given the load prediction, daily peak periods are determined as a time window with specified length in each day that include the predicted daily peak load and with a secondary rule such as maximizing energy saving potential. The dispatch schedule generation method is not a standalone measure and is intended to be combined with other demand flexibility measures that could leverage the peak schedule and apply demand controls on specific systems or devices for demand response, such as measures described in "Measure Documentation - Thermostat Control for Load Shedding" and "Measure Documentation - Thermostat Control for Load Shifting".

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Compressed sensing methods with applications to advanced air sampling

Environmental sampling methods developed by the Savannah River National Laboratory (SRNL) employ collectors with sorbent media tubes set at various locations to collect airborne emissions. Laboratory analyses of these tubes results in one-dimensional signals regarding what chemicals are being released and transported within the atmosphere. The analysis process is time consuming especially when analyzing a full year’s worth of tubes (hourly sample collection results in nearly 9,000 tubes per year). Using a signal processing method such as compressed sensing allows for recreation of the full signal while greatly reducing the number of analyzed samples required. Due to the sparsity of data retrieved from the air tubes, it is possible to use measurements a fraction of the size of the original data to gain much of the same information. This would improve the overall time and cost of analysis when modeling one-dimensional sampling signals.

54 ENVIRONMENTAL SCIENCES↗

Compressed Sensing Methods with Applications to Advanced Air Sampling [Poster]

Environmental sampling methods developed by the Savannah River National Laboratory (SRNL) employ collectors with sorbent media tubes set at various locations to collect airborne emissions. Laboratory analyses of these tubes results in one-dimensional signals regarding what chemicals are being released and transported within the atmosphere. The analysis process is time consuming especially when analyzing a full year’s worth of tubes (hourly sample collection results in nearly 9,000 tubes per year). Using a signal processing method such as compressed sensing allows for recreation of the full signal while greatly reducing the number of analyzed samples required. Due to the sparsity of data retrieved from the air tubes, it is possible to use measurements a fraction of the size of the original data to gain much of the same information. This would improve the overall time and cost of analysis when modeling one-dimensional sampling signals.

Campbell, Cassidy [Savannah River National Laborat↗

Using low volume eDNA methods to sample pelagic marine animal assemblages

Environmental DNA (eDNA) is an increasingly useful method for detecting pelagic animals in the ocean but typically requires large water volumes to sample diverse assemblages. Ship-based pelagic sampling programs that could implement eDNA methods generally have restrictive water budgets. Studies that quantify how eDNA methods perform on low water volumes in the ocean are limited, especially in deep-sea habitats with low animal biomass and poorly described species assemblages. Using 12S rRNA and COI gene primers, we quantified assemblages comprised of micronekton, coastal forage fishes, and zooplankton from low volume eDNA seawater samples (n = 436, 380–1800 mL) collected at depths of 0–2200 m in the southern California Current. We compared diversity in eDNA samples to concurrently collected pelagic trawl samples (n = 27), detecting a higher diversity of vertebrate and invertebrate groups in the eDNA samples. Differences in assemblage composition could be explained by variability in size-selectivity among methods and DNA primer suitability across taxonomic groups. The number of reads and amplicon sequences variants (ASVs) did not vary substantially among shallow (<200 m) and deep samples (>600 m), but the proportion of invertebrate ASVs that could be assigned a species-level identification decreased with sampling depth. Using hierarchical clustering, we resolved horizontal and vertical variability in marine animal assemblages from samples characterized by a relatively low diversity of ecologically important species. Low volume eDNA samples will quantify greater taxonomic diversity as reference libraries, especially for deep-dwelling invertebrate species, continue to expand.

59 BASIC BIOLOGICAL SCIENCES↗

Sample Preparation Method for Low-Level Total 129 I Measurements by ICP-MS

Trace-level measurements of iodine’s isotopic ( 129 I and 127 I) and chemical species distributions are needed for an accurate understanding of radioiodine migration in the Hanford subsurface. Pacific Northwest National Laboratory (PNNL) previously developed a novel analytical method for iodine characterization that uses ion chromatography (IC) coupled to inductively coupled mass spectrometry (ICP-MS). While the method can measure speciated forms of 129 I at levels below the drinking water standard, an interference from molybdenum (Mo) prevents the assay from quantifying the $\underline{total}$ 129 I concentration in many Hanford sample matrices. In this work, solvent extraction was evaluated as a sample preparation method for eliminating the Mo interference. A series of 10 experiments was conducted in which solutions containing known amounts of iodate or iodide were treated by solvent extraction, and the extracted solutions were analyzed for total iodine concentrations by ICP MS. Several extraction parameters such as reagent concentrations and chemical reaction times were systematically adjusted in attempts to optimize the extraction process. While solvent extraction was shown to be effective at removing Mo, there was a consistent inability to recover more than approximately 75% of the total iodine in most experiments. This would reduce the ability to detect 129 I at levels near the drinking water standard. Additionally, the extraction efficiencies in several experiments were highly variable, suggesting that solvent extraction could add significant uncertainty to radioiodine measurements. We recommend evaluating ion exchange as an alternative sample preparation approach in fiscal year (FY) 2024.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Intensity of sample processing methods impacts wastewater SARS-CoV-2 whole genome amplicon sequencing outcomes

Wastewater SARS-CoV-2 surveillance has been deployed since the beginning of the COVID-19 pandemic to monitor the dynamics in virus burden in local communities. Genomic surveillance of SARS-CoV-2 in wastewater, particularly efforts aimed at whole genome sequencing for variant tracking and identification, are still challenging due to low target concentration, complex microbial and chemical background, and lack of robust nucleic acid recovery experimental procedures. The intrinsic sample limitations are inherent to wastewater and are thus unavoidable. Here, we use a statistical approach that couples correlation analyses to a random forest-based machine learning algorithm to evaluate potentially important factors associated with wastewater SARS-CoV-2 whole genome amplicon sequencing outcomes, with a specific focus on the breadth of genome coverage. We collected 182 composite and grab wastewater samples from the Chicago area between November 2020 to October 2021. Samples were processed using a mixture of processing methods reflecting different homogenization intensities (HA + Zymo beads, HA + glass beads, and Nanotrap), and were sequenced using one of the two library preparation kits (the Illumina COVIDseq kit and the QIAseq DIRECT kit). Technical factors evaluated using statistical and machine learning approaches include sample types, certain sample intrinsic features, and processing and sequencing methods. The results suggested that sample processing methods could be a predominant factor affecting sequencing outcomes, and library preparation kits was considered a minor factor. Finally, a synthetic SARS-CoV-2 RNA spike-in experiment was performed to validate the impact from processing methods and suggested that the intensity of the processing methods could lead to different RNA fragmentation

60 APPLIED LIFE SCIENCES↗

Soil Water Retention and Hydraulic Conductivity Data and Model at Pump House in East River Watershed, Colorado 2019-2024

This data package includes soil water retention and hydraulic conductivity data and model fitting results from measurements of ex-situ soil samples and in-situ soil sensors near Pump House at Mount Crested Butte in the East River Watershed. Soil water retention curves (SWRC) characterize soil water content as a function of soil water potential. SWRC depends on soil texture and pore structure and can be used to describe the constraints on biogeochemical processes in terms of soil water availability. In this data package, the sample identification follows the format ER-X-Y, where ER refers to East River, X is the location identifier, and Y is the depth identifier at the same X (shallow Y=1). Specifically, ER-PHS, ER-LMC, ER-LMF, and ER-SMN are associated with ecohydrology sites under the East-Taylor Watershed Community Observatory Sites directory, and ER-RBTn (upslope n=1) are sampling transects during the 2019 Rootball Campaign. The sample and location information can be found in metadata.csv. Sampling and Measurements Each sample falls into one of the three sampling methods – (1) intact cores, (2) repacked samples, or (3) soil sensors – and one of the two measurement methods – (a) laboratory or (b) in-situ. Both intact cores and repacked samples were measured using the laboratory methods, which include measurements of soil water potential (HYPROP & WP4C, METER), saturated (KSAT, METER) and unsaturated hydraulic conductivity (HYPROP). The in-situ method uses a pair of co-located soil sensors to measure volumetric water content (TEROS12, METER) and soil water potential (TEROS21, METER), and the hydraulic conductivity was not measured. In comparison, the laboratory methods progress from full saturation to dry conditions, and the in-situ method includes both dry-to-wet and wet-to-dry cycles. The sampling and measurement methods for each sample can be found in metadata.csv, and more information about the measurements is detailed in the Methods section below. Models Retention and hydraulic conductivity data were fitted with four van-Genuchten-type models (specified by “model_name” column in the files): (1) traditional constrained van Genuchten model (“vG_constrained”), (2) traditional unconstrained van Genuchten model (“vG_unconstrained”), (3) PDI-variant of the constrained van Genuchten model (“vG_constrained_PDI”), and (4) PDI-variant of the unconstrained van Genuchten model (“vG_unconstrained_PDI”). The difference between the constrained (1: n) and the unconstrained (2: n, m) van Genuchten models is the number of pore-size distribution parameters in the model equations, giving the unconstrained model more degrees of freedom when fitting the data. Between the traditional and the PDI-variant models, model fitting differs the most at the dry end of the measurements. The traditional models allow infinite suction at the residual water content (water content does not drop below residual water content), and the PDI-variant models enforce a soil water potential value of pF=6.8 (~ -630 MPa) at oven-dryness (water content reaches 0). The inclusion of the van-Genuchten-type models is due to their common application. If other retention models are required, users can access the data in data.csv for further data fitting. More information about the models can be found in the Methods section below. Fitting Tasks The model fitting can be categorized into three levels of tasks (specified by “fitting_task” column in the files). Level 1 (“fit_retention”) only includes retention data fitting (the only level available for the in-situ method). Level 2 (“fit_retention_conductivity”) includes both retention and hydraulic conductivity data fitting, and the saturated hydraulic conductivity (Ks, a parameter of the hydraulic conductivity functions) is fixed by the measurements from KSAT. Level 3 (“fit_retention_conductivity_Ks”) also includes both retention and hydraulic conductivity data fitting, but Ks is a fitted parameter without the constraints from KSAT measurements. Among the same retention models (e.g. vG_constrained models of the same sample), level 1 should produce the best retention data fitting. Level 2 should have the highest misfit of the retention and hydraulic conductivity data, because the retention and hydraulic conductivity functions share common model parameters, and the unsaturated hydraulic conductivity (HYPROP) data fitting is subject to Ks measured independently by KSAT. Level 3 should have mid-level misfits of the retention and hydraulic conductivity data. While level 3 fits the hydraulic conductivity data better than level 2, the fitted Ks value might be unreasonable due to the lack of constraints at the wet end of the measurements. General recommendation when using this data package: (1) Choice of sampling methods: Intact cores and in-situ soil sensors could be prioritized because these sampling methods are less destructive. While the repacked samples were packed to the target bulk density (estimated post-sampling, when sample volume was known), these samples had altered pore structures. Nevertheless, intact cores might suffer from sample gaps that would lead to overestimation of Ks (sample gaps can be inferred from the “soil_sample_volume” column in metadata.csv when the value is < 249). In-situ method also has higher uncertainty in characterizing the wet end of the SWRC because of sensor limitations and the difficulty in reaching full saturation under natural conditions. (2) Choice of fitting tasks: When only retention data is needed, level 1 (“fit_retention”) should be prioritized. When both retention and hydraulic conductivity data are needed, level 2 (“fit_retention_conductivity”) could be prioritized. (3) Choice of models: This could depend on what the downstream models call for. If no specific model is required, model misfit could be used as a ranking criterion. Model misfit values in terms of RMSE can be found in model_parameters.csv. The following files are included in this data package: (1) metadata.csv – This file includes the general information of each sample, including location (description, geocoordinates, elevation), sampling and measurements details (method, depth, time or period, volume, instruments), and soil physical properties (bulk density, saturated hydraulic conductivity, only applicable to physical soil samples). (2) data.csv – This file includes soil water potential, volumetric water content, and unsaturated hydraulic conductivity data of each sample. Column “instrument” specifies the instrument (HYPROP, WP4C, or TEROS) used to perform the measurements. (3) model_fit.csv – This file includes soil water potential, volumetric water content, and unsaturated hydraulic conductivity fitted from the four models and three fitting tasks. Column “model_name” specifies the retention model used, and “fitting_task” specifies the level of data fitting. Missing values indicate that the variable does not apply to that fitting task. (4) model_parameters.csv – This file includes the fitted model parameters, model misfits, and conventional water content thresholds (field capacity and wilting point) from the four models and three fitting tasks. Column “model_name” specifies the retention model used, and “fitting_task” specifies the level of data fitting. Missing values indicate that the parameter does not apply to that model and/or that fitting task. (5) data_Ks.csv – This file includes the saturated hydraulic conductivity measurements from KSAT. (6) /figure/*.png – This folder includes three quick visualizations of the data, retention model fitting results and misfits, and hydraulic conductivity model fitting results, misfits, and parameters. The model fitting results are separated by samples and fitting tasks and colored by models. Zoom-in required. (7) /hyprop/*.bdhx – This folder includes proprietary hyprop files that require the free Labros SoilView-Analysis (METER) to open. Users can explore data fitting using other retention models (i.e. Brooks-Corey, Fredlund-Xing, Kosugi, bimodal models). Be aware that Ks value is pre-entered under “Fitting tab, Conductivity functions parameters” for level 2 fitting. If the value is lost, please refer to metadata.csv under “Ks” column. (8) Six file-level metadata that summarize file, header, column, and variable information of all files. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

EARTH SCIENCE > LAND SURFACE > SOILS↗

Soil Water Retention and Hydraulic Conductivity Data and Model at Snodgrass Mountain in East River Watershed, Colorado 2020-2025

This data package includes soil water retention and hydraulic conductivity data and model fitting results from measurements of ex-situ soil samples and in-situ soil sensors at Snodgrass Mountain. Soil water retention curves (SWRC) characterize soil water content as a function of soil water potential. SWRC depends on soil texture and pore structure and can be used to describe the constraints on biogeochemical processes in terms of soil water availability. In this data package, the sample identification follows the format SG-X-Y, where SG refers to Snodgrass Mountain, X is the location identifier, and Y is the depth identifier at the same X (shallow Y=1). Specifically, SG-EHS is associated with ecohydrology sites under the East-Taylor Watershed Community Observatory Sites directory, and SG-ERTn (upslope n=1) are points along the Snodgrass electrical resistivity tomography transect not associated with the existing site names in the directory. The sample and location information can be found in metadata.csv. Sampling and Measurements Each sample falls into one of the three sampling methods – (1) intact cores, (2) repacked samples, or (3) soil sensors – and one of the two measurement methods – (a) laboratory or (b) in-situ. Both intact cores and repacked samples were measured using the laboratory methods, which include measurements of soil water potential (HYPROP & WP4C, METER), saturated (KSAT, METER) and unsaturated hydraulic conductivity (HYPROP). The in-situ method uses a pair of co-located soil sensors to measure volumetric water content (TEROS12, METER) and soil water potential (TEROS21, METER), and the hydraulic conductivity was not measured. In comparison, the laboratory methods progress from full saturation to dry conditions, and the in-situ method includes both dry-to-wet and wet-to-dry cycles. The sampling and measurement methods for each sample can be found in metadata.csv, and more information about the measurements is detailed in the Methods section below. Models Retention and hydraulic conductivity data were fitted with four van-Genuchten-type models (specified by “model_name” column in the files): (1) traditional constrained van Genuchten model (“vG_constrained”), (2) traditional unconstrained van Genuchten model (“vG_unconstrained”), (3) PDI-variant of the constrained van Genuchten model (“vG_constrained_PDI”), and (4) PDI-variant of the unconstrained van Genuchten model (“vG_unconstrained_PDI”). The difference between the constrained (1: n) and the unconstrained (2: n, m) van Genuchten models is the number of pore-size distribution parameters in the model equations, giving the unconstrained model more degrees of freedom when fitting the data. Between the traditional and the PDI-variant models, model fitting differs the most at the dry end of the measurements. The traditional models allow infinite suction at the residual water content (water content does not drop below residual water content), and the PDI-variant models enforce a soil water potential value of pF=6.8 (~ -630 MPa) at oven-dryness (water content reaches 0). The inclusion of the van-Genuchten-type models is due to their common application. If other retention models are required, users can access the data in data.csv for further data fitting. More information about the models can be found in the Methods section below. Fitting Tasks The model fitting can be categorized into three levels of tasks (specified by “fitting_task” column in the files). Level 1 (“fit_retention”) only includes retention data fitting (the only level available for the in-situ method). Level 2 (“fit_retention_conductivity”) includes both retention and hydraulic conductivity data fitting, and the saturated hydraulic conductivity (Ks, a parameter of the hydraulic conductivity functions) is fixed by the measurements from KSAT. Level 3 (“fit_retention_conductivity_Ks”) also includes both retention and hydraulic conductivity data fitting, but Ks is a fitted parameter without the constraints from KSAT measurements. Among the same retention models (e.g. vG_constrained models of the same sample), level 1 should produce the best retention data fitting. Level 2 should have the highest misfit of the retention and hydraulic conductivity data, because the retention and hydraulic conductivity functions share common model parameters, and the unsaturated hydraulic conductivity (HYPROP) data fitting is subject to Ks measured independently by KSAT. Level 3 should have mid-level misfits of the retention and hydraulic conductivity data. While level 3 fits the hydraulic conductivity data better than level 2, the fitted Ks value might be unreasonable due to the lack of constraints at the wet end of the measurements. General recommendation when using this data package: (1) Choice of sampling methods: Intact cores and in-situ soil sensors could be prioritized because these sampling methods are less destructive. While the repacked samples were packed to the target bulk density (estimated post-sampling, when sample volume was known), these samples had altered pore structures. Nevertheless, intact cores might suffer from sample gaps that would lead to overestimation of Ks (sample gaps can be inferred from the “soil_sample_volume” column in metadata.csv when the value is < 249). In-situ method also has higher uncertainty in characterizing the wet end of the SWRC because of sensor limitations and the difficulty in reaching full saturation under natural conditions. (2) Choice of fitting tasks: When only retention data is needed, level 1 (“fit_retention”) should be prioritized. When both retention and hydraulic conductivity data are needed, level 2 (“fit_retention_conductivity”) could be prioritized. (3) Choice of models: This could depend on what the downstream models call for. If no specific model is required, model misfit could be used as a ranking criterion. Model misfit values in terms of RMSE can be found in model_parameters.csv. The following files are included in this data package: (1) metadata.csv – This file includes the general information of each sample, including location (description, geocoordinates, elevation), sampling and measurements details (method, depth, time or period, volume, instruments), and soil physical properties (bulk density, saturated hydraulic conductivity, only applicable to physical soil samples). (2) data.csv – This file includes soil water potential, volumetric water content, and unsaturated hydraulic conductivity data of each sample. Column “instrument” specifies the instrument (HYPROP, WP4C, or TEROS) used to perform the measurements. (3) model_fit.csv – This file includes soil water potential, volumetric water content, and unsaturated hydraulic conductivity fitted from the four models and three fitting tasks. Column “model_name” specifies the retention model used, and “fitting_task” specifies the level of data fitting. Missing values indicate that the variable does not apply to that fitting task. (4) model_parameters.csv – This file includes the fitted model parameters, model misfits, and conventional water content thresholds (field capacity and wilting point) from the four models and three fitting tasks. Column “model_name” specifies the retention model used, and “fitting_task” specifies the level of data fitting. Missing values indicate that the parameter does not apply to that model and/or that fitting task. (5) data_Ks.csv – This file includes the saturated hydraulic conductivity measurements from KSAT. (6) /figure/*.png – This folder includes three quick visualizations of the data, retention model fitting results and misfits, and hydraulic conductivity model fitting results, misfits, and parameters. The model fitting results are separated by samples and fitting tasks and colored by models. Zoom-in required. (7) /hyprop/*.bdhx – This folder includes proprietary hyprop files that require the free Labros SoilView-Analysis (METER) to open. Users can explore data fitting using other retention models (i.e. Brooks-Corey, Fredlund-Xing, Kosugi, bimodal models). Be aware that Ks value is pre-entered under “Fitting tab, Conductivity functions parameters” for level 2 fitting. If the value is lost, please refer to metadata.csv under “Ks” column. (8) Six file-level metadata that summarize file, header, column, and variable information of all files. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

EARTH SCIENCE > LAND SURFACE > SOILS↗