Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Science Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Heuristics for Relevancy Ranking of Earth Dataset Search Results

As the Variety of Earth science datasets increases, science researchers find it more challenging to discover and select the datasets that best fit their needs. The most common way of search providers to address this problem is to rank the datasets returned for a query by their likely relevance to the user. Large web page search engines typically use text matching supplemented with reverse link counts, semantic annotations and user intent modeling. However, this produces uneven results when applied to dataset metadata records simply externalized as a web page. Fortunately, data and search provides have decades of experience in serving data user communities, allowing them to form heuristics that leverage the structure in the metadata together with knowledge about the user community. Some of these heuristics include specific ways of matching the user input to the essential measurements in the dataset and determining overlaps of time range and spatial areas. Heuristics based on the novelty of the datasets can prioritize later, better versions of data over similar predecessors. And knowledge of how different user types and communities use data can be brought to bear in cases where characteristics of the user (discipline, expertise) or their intent (applications, research) can be divined. The Earth Observing System Data and Information System has begun implementing some of these heuristics in the relevancy algorithm of its Common Metadata Repository search engine.

science data management↗

Relevancy Ranking of Satellite Dataset Search Results

As the Variety of Earth science datasets increases, science researchers find it more challenging to discover and select the datasets that best fit their needs. The most common way of search providers to address this problem is to rank the datasets returned for a query by their likely relevance to the user. Large web page search engines typically use text matching supplemented with reverse link counts, semantic annotations and user intent modeling. However, this produces uneven results when applied to dataset metadata records simply externalized as a web page. Fortunately, data and search provides have decades of experience in serving data user communities, allowing them to form heuristics that leverage the structure in the metadata together with knowledge about the user community. Some of these heuristics include specific ways of matching the user input to the essential measurements in the dataset and determining overlaps of time range and spatial areas. Heuristics based on the novelty of the datasets can prioritize later, better versions of data over similar predecessors. And knowledge of how different user types and communities use data can be brought to bear in cases where characteristics of the user (discipline, expertise) or their intent (applications, research) can be divined. The Earth Observing System Data and Information System has begun implementing some of these heuristics in the relevancy algorithm of its Common Metadata Repository search engine.

science data management↗

Performance and Reliability Assessment of the U.S. Department of Energy Atmospheric Radiation Measurement (ARM) Data Advisor (ADA)

The Atmospheric Radiation Measurement (ARM) User Facility provides one of the world's largest openly accessible repositories of atmospheric observations through the ARM Data Discovery platform. Although the repository contains more than three decades of measurements collected from permanent observatories, mobile facilities, aircraft campaigns, and field experiments, identifying appropriate datasets can be challenging, particularly for new users unfamiliar with ARM instrumentation and datastream organization. To improve data accessibility, the ARM Data Center developed the ARM Data Advisor (ADA), an artificial intelligence-powered assistant designed to facilitate scientific data discovery, dataset interpretation, and user guidance. This report evaluates ADA's performance as a domain-specific scientific assistant using realistic atmospheric science workflows. The evaluation examines five key capabilities: data retrieval and curation efficiency, hallucination resistance, scientific reasoning, response to ambiguous queries, and content retention and session continuity. Representative prompts were developed to simulate typical interactions between researchers and the ARM Data Discovery platform, and ADA's responses were assessed for retrieval completeness, scientific accuracy, consistency, and practical usefulness. In these representative tests, ADA reduced the complexity of discovering and accessing ARM datasets by recommending appropriate datastreams, explaining instrumentation, interpreting metadata, and assisting with data processing workflows. ADA also exhibits strong domain knowledge of atmospheric science terminology and generally resists hallucination by acknowledging unavailable datasets and requesting clarification when appropriate. Overall, the results indicate that ADA represents a promising advancement in scientific data discovery within the ARM User Facility and has considerable potential to improve researcher productivity, particularly for new users and interdisciplinary scientists seeking efficient access to ARM observations.

Salvador, Christian [ORNL] (ORCID:0000000283287777↗

Data Cards for Standardized Metadata Across DOE-Aligned Data Initiatives: Toward Transparent, Interoperable, and Governed Dataset Documentation

As data-intensive research, advanced computing, and artificial intelligence become increasingly central to scientific and operational workflows, the need for consistent, transparent, and machine-actionable documentation has grown correspondingly. Multiple DOE-aligned communities—including Office of Science, Genesis Mission, American Science Cloud (AmSC), National Nuclear Security Administration (NNSA) stewardship and governance, and related cross-laboratory collaborations—have independently developed metadata practices to support discovery, access, reuse, repository deposit, and compliance.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Framework for Integrating Science Data Processing Algorithms Into Process Control Systems

A software framework called PCS Task Wrapper is responsible for standardizing the setup, process initiation, execution, and file management tasks surrounding the execution of science data algorithms, which are referred to by NASA as Product Generation Executives (PGEs). PGEs codify a scientific algorithm, some step in the overall scientific process involved in a mission science workflow. The PCS Task Wrapper provides a stable operating environment to the underlying PGE during its execution lifecycle. If the PGE requires a file, or metadata regarding the file, the PCS Task Wrapper is responsible for delivering that information to the PGE in a manner that meets its requirements. If the PGE requires knowledge of upstream or downstream PGEs in a sequence of executions, that information is also made available. Finally, if information regarding disk space, or node information such as CPU availability, etc., is required, the PCS Task Wrapper provides this information to the underlying PGE. After this information is collected, the PGE is executed, and its output Product file and Metadata generation is managed via the PCS Task Wrapper framework. The innovation is responsible for marshalling output Products and Metadata back to a PCS File Management component for use in downstream data processing and pedigree. In support of this, the PCS Task Wrapper leverages the PCS Crawler Framework to ingest (during pipeline processing) the output Product files and Metadata produced by the PGE. The architectural components of the PCS Task Wrapper framework include PGE Task Instance, PGE Config File Builder, Config File Property Adder, Science PGE Config File Writer, and PCS Met file Writer. This innovative framework is really the unifying bridge between the execution of a step in the overall processing pipeline, and the available PCS component services as well as the information that they collectively manage.

Mattmann, Chris A.↗

Infusion of AI/ML Technology into Operational NASA Data Systems

NASA has been developing a variety of Artificial Intelligence / Machine Learning technologies related to Earth Observations. In most cases, the full value of such a technology is realized when it is infused into an operational system. NASA’s Earth Science Data Systems program has been formulating repeatable methods to execute technology infusion. These efforts include the Advancing Collaborative Connections for Earth System Science (ACCESS) program, a Technology Infusion Playbook, and an assemblage of working groups investigating methods for infusion collaboration, community development, and capacity building. ESDS has also been executing a pathfinder activity to infuse a machine-learning-driven recommender of science keywords for Earth Observation datasets, which is intended to be used for metadata curation in the Earth Observation System Data and Information System.

C Lynnes↗

Soil physical and chemical measurements for topsoils collected during NEON campaign in East River, CO (06/14/2018-06/28/2018)

The package is part of the DOE Watershed Function Science Focus Area (SFA) project and includes soil physical and chemical measurements from topsoils collected at the East River, Colorado, in conjunction with the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP) survey conducted in June 2018. The soil measurements include soil bulk density, soil volumetric water content, soil microbial biomass C (Carbon), N (Nitrogen) and C:N (C to N ratio), soil DNA yield, soil total extractable organic C, soil total extractable N, soil extractable nitrate, soil extractable ammonium, soil dissolved inorganic N, soil dissolved organic N, soil pH, soil TOC400 (total organic carbon at 400°C), soil ROC (residual oxidizable carbon), soil TIC (total inorganic carbon), soil TOC (total organic carbon), soil TC (total carbon), soil N, soil OM (organic matter) loss on ignition. Additional associated site metadata can be found in the ESS-DIVE package 10.15485/1618130. The dataset includes (1) 2018_NEON_soil_physical_chemical_measurements.csv: soil physical and chemical measurements indexed by soil sample IGSNs; (2) samples.csv: sample metadata file used to register International Generic Sample Numbers (IGSNs); (3) flmd.csv: file level metadata file; and (4) dd.csv: data dictionary file. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS (Catchment Hydrology and ↗

Metagenome-assembled genomes from topsoils along a hillslope water gradient across early snowmelt to late summer in East River, CO

Drought is changing the American Mountain West at unprecedented rates with unknown consequences to soil microbiome composition and function. As a part of LBNL Watershed Science Focus Area (SFA), we investigated shifts in microbial community and transcriptional activity on a subalpine conifer-meadow transition zone throughout the summer of 2023 as soil dried down. This work took place in Crested Butte, CO on Snodgrass mountain, using a proxy for drought conditions.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal community at 0-10cm from three sites along a hillslope water gradient across five timepoints from early snowmelt to late summer. 42 metagenomes were sequenced at Joint Genome Institute (JGI) and can be found under the JGI GOLD (Genomes Online Database) sequencing project Gs0166660. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>70%) and contamination (<10%), and dereplicated at 95% ANI using drep. This dataset (1) a zip file of 157 MAGs (as fasta files, Gs0166660_bins_tar.gz), (2) sample metadata file with sample IGSNs (International Generic Sample Numbers) (samples.csv), (3) bounding box coordinates for the sampled locations (Gs0166660.kml), (4) metagenome assembly and coassembly metadata file listing IMG/M (Integrated Microbial Genomes/Metagenomes) metagenome accessions linking samples to metagenomes (EastRiver_Drought_ESSDive_Metadata.csv), (5) location metadata file (locations.csv), (6) file-level metadata file (flmd.csv) and (7) data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Prediction of Distributed River Sediment Respiration Rates Using Community-Generated Data and Machine Learning

River sediment microbial respiration is a key indicator of ecosystem functioning and the biogeochemical fluxes across this critical zone link surface and subsurface waters. As such, there is tremendous interest in measuring and mapping these respiration rates. Respiration observations are expensive and labor intensive; there is limited data available to the community. An open science, collaborative initiative is collecting samples for respiration rate analysis and multi-scale metadata; this evolving data set is being used for making machine learning (ML) predictions at unsampled sites to help inform continued community engagement. However, it is a challenge to find an optimum configuration for ML models to work with this feature-rich (i.e., 100+ possible input variables) data set. Here, we present results from a two-tiered approach to managing the analysis of this complex data set: (a) a stacked ensemble of models that automatically optimizes hyperparameters and manages the training of many models and (b) feature permutation importance to detect the most important features in the models. The major elements of this workflow are modular, portable, open, and cloud-based thus making this implementation a potential template for other applications. The models developed here predict that sediment organic matter chemistry is one of the most important features for predicting sediment respiration rate. Other larger-scale, important features fall into the categories of climatic, ecological, geological, and fluvial settings. Leveraging these larger-scale features to generate data-driven estimates of river sediment respiration rates reveals spatially consistent but heterogeneous patterns across the river network of the Columbia River Basin.

54 ENVIRONMENTAL SCIENCES↗

Montane Conifer, Aspen, Meadow, and Sagebrush Metagenome Resolved Genomes and Traits in East River Watershed, Colorado, USA

Climate change is driving vegetation shifts in mountain watersheds, with unknown impacts on biogeochemical cycles. We hypothesize that these shifts will reshape soil microbiomes and associated biogeochemical processes. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed microbiome and microbial functional trait differences between soils under conifer, aspen, forby meadows, and sagebrush across the East River Watershed, CO, controlling for elevation and aspect.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from soils 0-20cm in depth across three locations in the watershed—Headwaters, Upper Reaches, and Lower Reaches from August 3-11th 2016. Each location was further subdivided into two blocks, with one block on a west facing aspect, and two on the east aspect of the valley. Within blocks, two samples per vegetation type were taken (one at each depth). This resulted in 66 samples, which were sequenced at JGI and can be found under the Joint Genome Institute (JGI) Genomes Online Database (GOLD) sequencing project Gs0118068. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>75%) and contamination (<25%), and dereplicated at 95% ANI using drep. The dataset includes a zip file of 687 genomes (Vegtype_MAGS.zip), the accession numbers for the underlying metagenomes, a csv file with MAG quality metrics and taxonomy from Genome Taxonomy Database (GTDB) and National Center for Biotechnology Information (NCBI) taxonomic representative genome proteins (EastRiver_Vegtype_drep_genome_info.csv), and a file containing MAG quality metrics and taxonomy (gtdb_drep_bin_taxonomy.csv). The dataset additionally includes a sample metadata file (EastRiver_Vegtype_sample_metadata.csv), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a Google KML file for the sampled locations (sample_collection_sites.kml), a location metadata file (locations.csv), a file-level metadata file (flmd.csv), and a data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Soil Water Retention and Hydraulic Conductivity Data and Model at Pump House in East River Watershed, Colorado 2019-2024

This data package includes soil water retention and hydraulic conductivity data and model fitting results from measurements of ex-situ soil samples and in-situ soil sensors near Pump House at Mount Crested Butte in the East River Watershed. Soil water retention curves (SWRC) characterize soil water content as a function of soil water potential. SWRC depends on soil texture and pore structure and can be used to describe the constraints on biogeochemical processes in terms of soil water availability. In this data package, the sample identification follows the format ER-X-Y, where ER refers to East River, X is the location identifier, and Y is the depth identifier at the same X (shallow Y=1). Specifically, ER-PHS, ER-LMC, ER-LMF, and ER-SMN are associated with ecohydrology sites under the East-Taylor Watershed Community Observatory Sites directory, and ER-RBTn (upslope n=1) are sampling transects during the 2019 Rootball Campaign. The sample and location information can be found in metadata.csv. Sampling and Measurements Each sample falls into one of the three sampling methods – (1) intact cores, (2) repacked samples, or (3) soil sensors – and one of the two measurement methods – (a) laboratory or (b) in-situ. Both intact cores and repacked samples were measured using the laboratory methods, which include measurements of soil water potential (HYPROP & WP4C, METER), saturated (KSAT, METER) and unsaturated hydraulic conductivity (HYPROP). The in-situ method uses a pair of co-located soil sensors to measure volumetric water content (TEROS12, METER) and soil water potential (TEROS21, METER), and the hydraulic conductivity was not measured. In comparison, the laboratory methods progress from full saturation to dry conditions, and the in-situ method includes both dry-to-wet and wet-to-dry cycles. The sampling and measurement methods for each sample can be found in metadata.csv, and more information about the measurements is detailed in the Methods section below. Models Retention and hydraulic conductivity data were fitted with four van-Genuchten-type models (specified by “model_name” column in the files): (1) traditional constrained van Genuchten model (“vG_constrained”), (2) traditional unconstrained van Genuchten model (“vG_unconstrained”), (3) PDI-variant of the constrained van Genuchten model (“vG_constrained_PDI”), and (4) PDI-variant of the unconstrained van Genuchten model (“vG_unconstrained_PDI”). The difference between the constrained (1: n) and the unconstrained (2: n, m) van Genuchten models is the number of pore-size distribution parameters in the model equations, giving the unconstrained model more degrees of freedom when fitting the data. Between the traditional and the PDI-variant models, model fitting differs the most at the dry end of the measurements. The traditional models allow infinite suction at the residual water content (water content does not drop below residual water content), and the PDI-variant models enforce a soil water potential value of pF=6.8 (~ -630 MPa) at oven-dryness (water content reaches 0). The inclusion of the van-Genuchten-type models is due to their common application. If other retention models are required, users can access the data in data.csv for further data fitting. More information about the models can be found in the Methods section below. Fitting Tasks The model fitting can be categorized into three levels of tasks (specified by “fitting_task” column in the files). Level 1 (“fit_retention”) only includes retention data fitting (the only level available for the in-situ method). Level 2 (“fit_retention_conductivity”) includes both retention and hydraulic conductivity data fitting, and the saturated hydraulic conductivity (Ks, a parameter of the hydraulic conductivity functions) is fixed by the measurements from KSAT. Level 3 (“fit_retention_conductivity_Ks”) also includes both retention and hydraulic conductivity data fitting, but Ks is a fitted parameter without the constraints from KSAT measurements. Among the same retention models (e.g. vG_constrained models of the same sample), level 1 should produce the best retention data fitting. Level 2 should have the highest misfit of the retention and hydraulic conductivity data, because the retention and hydraulic conductivity functions share common model parameters, and the unsaturated hydraulic conductivity (HYPROP) data fitting is subject to Ks measured independently by KSAT. Level 3 should have mid-level misfits of the retention and hydraulic conductivity data. While level 3 fits the hydraulic conductivity data better than level 2, the fitted Ks value might be unreasonable due to the lack of constraints at the wet end of the measurements. General recommendation when using this data package: (1) Choice of sampling methods: Intact cores and in-situ soil sensors could be prioritized because these sampling methods are less destructive. While the repacked samples were packed to the target bulk density (estimated post-sampling, when sample volume was known), these samples had altered pore structures. Nevertheless, intact cores might suffer from sample gaps that would lead to overestimation of Ks (sample gaps can be inferred from the “soil_sample_volume” column in metadata.csv when the value is < 249). In-situ method also has higher uncertainty in characterizing the wet end of the SWRC because of sensor limitations and the difficulty in reaching full saturation under natural conditions. (2) Choice of fitting tasks: When only retention data is needed, level 1 (“fit_retention”) should be prioritized. When both retention and hydraulic conductivity data are needed, level 2 (“fit_retention_conductivity”) could be prioritized. (3) Choice of models: This could depend on what the downstream models call for. If no specific model is required, model misfit could be used as a ranking criterion. Model misfit values in terms of RMSE can be found in model_parameters.csv. The following files are included in this data package: (1) metadata.csv – This file includes the general information of each sample, including location (description, geocoordinates, elevation), sampling and measurements details (method, depth, time or period, volume, instruments), and soil physical properties (bulk density, saturated hydraulic conductivity, only applicable to physical soil samples). (2) data.csv – This file includes soil water potential, volumetric water content, and unsaturated hydraulic conductivity data of each sample. Column “instrument” specifies the instrument (HYPROP, WP4C, or TEROS) used to perform the measurements. (3) model_fit.csv – This file includes soil water potential, volumetric water content, and unsaturated hydraulic conductivity fitted from the four models and three fitting tasks. Column “model_name” specifies the retention model used, and “fitting_task” specifies the level of data fitting. Missing values indicate that the variable does not apply to that fitting task. (4) model_parameters.csv – This file includes the fitted model parameters, model misfits, and conventional water content thresholds (field capacity and wilting point) from the four models and three fitting tasks. Column “model_name” specifies the retention model used, and “fitting_task” specifies the level of data fitting. Missing values indicate that the parameter does not apply to that model and/or that fitting task. (5) data_Ks.csv – This file includes the saturated hydraulic conductivity measurements from KSAT. (6) /figure/*.png – This folder includes three quick visualizations of the data, retention model fitting results and misfits, and hydraulic conductivity model fitting results, misfits, and parameters. The model fitting results are separated by samples and fitting tasks and colored by models. Zoom-in required. (7) /hyprop/*.bdhx – This folder includes proprietary hyprop files that require the free Labros SoilView-Analysis (METER) to open. Users can explore data fitting using other retention models (i.e. Brooks-Corey, Fredlund-Xing, Kosugi, bimodal models). Be aware that Ks value is pre-entered under “Fitting tab, Conductivity functions parameters” for level 2 fitting. If the value is lost, please refer to metadata.csv under “Ks” column. (8) Six file-level metadata that summarize file, header, column, and variable information of all files. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

EARTH SCIENCE > LAND SURFACE > SOILS↗

Soil Water Retention and Hydraulic Conductivity Data and Model at Trail Creek in Taylor River Watershed, Colorado 2024-2025

This data package includes soil water retention and hydraulic conductivity data and model fitting results from measurements of ex-situ soil samples and in-situ soil sensors near Trail Creek. Soil water retention curves (SWRC) characterize soil water content as a function of soil water potential. SWRC depends on soil texture and pore structure and can be used to describe the constraints on biogeochemical processes in terms of soil water availability. In this data package, the sample identification follows the format TR-X-Y, where TR refers to Trail Creek, X is the treatment block identifier, and Y is the location identifier. Specifically, TR-ASCC1 is the control treatment block under the Adaptive Silviculture for Climate Change (ASCC) project, and TR-ASCC2 is the clear-cut treatment block. TR-ASCC-EHSn is associated with ecohydrology sites under the East-Taylor Watershed Community Observatory Sites directory, and TR-ASCC-ERTn (upslope n=1) are ecohydrology sites along the electrical resistivity tomography transects. The sample and location information can be found in metadata.csv, and the data from the soil sensors will be included in a future data version when the observation period becomes sufficiently long for data analysis. Sampling and Measurements Each sample falls into one of the two sampling methods – (1) intact cores or (2) soil sensors – and one of the two measurement methods – (a) laboratory or (b) in-situ. The intact cores were measured using the laboratory methods, which include measurements of soil water potential (HYPROP & WP4C, METER), saturated (KSAT, METER) and unsaturated hydraulic conductivity (HYPROP). The in-situ method uses a pair of co-located soil sensors to measure volumetric water content (TEROS12, METER) and soil water potential (TEROS21, METER), and the hydraulic conductivity was not measured. In comparison, the laboratory methods progress from full saturation to dry conditions, and the in-situ method includes both dry-to-wet and wet-to-dry cycles. The sampling and measurement methods for each sample can be found in metadata.csv, and more information about the measurements is detailed in the Methods section below. Models Retention and hydraulic conductivity data were fitted with four van-Genuchten-type models (specified by “model_name” column in the files): (1) traditional constrained van Genuchten model (“vG_constrained”), (2) traditional unconstrained van Genuchten model (“vG_unconstrained”), (3) PDI-variant of the constrained van Genuchten model (“vG_constrained_PDI”), and (4) PDI-variant of the unconstrained van Genuchten model (“vG_unconstrained_PDI”). The difference between the constrained (1: n) and the unconstrained (2: n, m) van Genuchten models is the number of pore-size distribution parameters in the model equations, giving the unconstrained model more degrees of freedom when fitting the data. Between the traditional and the PDI-variant models, model fitting differs the most at the dry end of the measurements. The traditional models allow infinite suction at the residual water content (water content does not drop below residual water content), and the PDI-variant models enforce a soil water potential value of pF=6.8 (~ -630 MPa) at oven-dryness (water content reaches 0). The inclusion of the van-Genuchten-type models is due to their common application. If other retention models are required, users can access the data in data.csv for further data fitting. More information about the models can be found in the Methods section below. Fitting Tasks The model fitting can be categorized into three levels of tasks (specified by “fitting_task” column in the files). Level 1 (“fit_retention”) only includes retention data fitting (the only level available for the in-situ method). Level 2 (“fit_retention_conductivity”) includes both retention and hydraulic conductivity data fitting, and the saturated hydraulic conductivity (Ks, a parameter of the hydraulic conductivity functions) is fixed by the measurements from KSAT. Level 3 (“fit_retention_conductivity_Ks”) also includes both retention and hydraulic conductivity data fitting, but Ks is a fitted parameter without the constraints from KSAT measurements. Among the same retention models (e.g. vG_constrained models of the same sample), level 1 should produce the best retention data fitting. Level 2 should have the highest misfit of the retention and hydraulic conductivity data, because the retention and hydraulic conductivity functions share common model parameters, and the unsaturated hydraulic conductivity (HYPROP) data fitting is subject to Ks measured independently by KSAT. Level 3 should have mid-level misfits of the retention and hydraulic conductivity data. While level 3 fits the hydraulic conductivity data better than level 2, the fitted Ks value might be unreasonable due to the lack of constraints at the wet end of the measurements. General recommendation when using this data package: (1) Choice of sampling methods: Intact cores might suffer from sample gaps that would lead to overestimation of Ks (sample gaps can be inferred from the “soil_sample_volume” column in metadata.csv when the value is < 249). In-situ method has higher uncertainty in characterizing the wet end of the SWRC because of sensor limitations and the difficulty in reaching full saturation under natural conditions. (2) Choice of fitting tasks: When only retention data is needed, level 1 (“fit_retention”) should be prioritized. When both retention and hydraulic conductivity data are needed, level 2 (“fit_retention_conductivity”) could be prioritized. (3) Choice of models: This could depend on what the downstream models call for. If no specific model is required, model misfit could be used as a ranking criterion. Model misfit values in terms of RMSE can be found in model_parameters.csv. The following files are included in this data package: (1) metadata.csv – This file includes the general information of each sample, including location (description, geocoordinates, elevation), sampling and measurements details (method, depth, time or period, volume, instruments), and soil physical properties (bulk density, saturated hydraulic conductivity, only applicable to physical soil samples). (2) data.csv – This file includes soil water potential, volumetric water content, and unsaturated hydraulic conductivity data of each sample. Column “instrument” specifies the instrument (HYPROP, WP4C, or TEROS) used to perform the measurements. (3) model_fit.csv – This file includes soil water potential, volumetric water content, and unsaturated hydraulic conductivity fitted from the four models and three fitting tasks. Column “model_name” specifies the retention model used, and “fitting_task” specifies the level of data fitting. Missing values indicate that the variable does not apply to that fitting task. (4) model_parameters.csv – This file includes the fitted model parameters, model misfits, and conventional water content thresholds (field capacity and wilting point) from the four models and three fitting tasks. Column “model_name” specifies the retention model used, and “fitting_task” specifies the level of data fitting. Missing values indicate that the parameter does not apply to that model and/or that fitting task. (5) data_Ks.csv – This file includes the saturated hydraulic conductivity measurements from KSAT. (6) /figure/*.png – This folder includes three quick visualizations of the data, retention model fitting results and misfits, and hydraulic conductivity model fitting results, misfits, and parameters. The model fitting results are separated by samples and fitting tasks and colored by models. Zoom-in required. (7) /hyprop/*.bdhx – This folder includes proprietary hyprop files that require the free Labros SoilView-Analysis (METER) to open. Users can explore data fitting using other retention models (i.e. Brooks-Corey, Fredlund-Xing, Kosugi, bimodal models). Be aware that Ks value is pre-entered under “Fitting tab, Conductivity functions parameters” for level 2 fitting. If the value is lost, please refer to metadata.csv under “Ks” column. (8) Six file-level metadata that summarize file, header, column, and variable information of all files. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

EARTH SCIENCE > LAND SURFACE > SOILS↗

Soil Water Retention and Hydraulic Conductivity Data and Model at Snodgrass Mountain in East River Watershed, Colorado 2020-2025

This data package includes soil water retention and hydraulic conductivity data and model fitting results from measurements of ex-situ soil samples and in-situ soil sensors at Snodgrass Mountain. Soil water retention curves (SWRC) characterize soil water content as a function of soil water potential. SWRC depends on soil texture and pore structure and can be used to describe the constraints on biogeochemical processes in terms of soil water availability. In this data package, the sample identification follows the format SG-X-Y, where SG refers to Snodgrass Mountain, X is the location identifier, and Y is the depth identifier at the same X (shallow Y=1). Specifically, SG-EHS is associated with ecohydrology sites under the East-Taylor Watershed Community Observatory Sites directory, and SG-ERTn (upslope n=1) are points along the Snodgrass electrical resistivity tomography transect not associated with the existing site names in the directory. The sample and location information can be found in metadata.csv. Sampling and Measurements Each sample falls into one of the three sampling methods – (1) intact cores, (2) repacked samples, or (3) soil sensors – and one of the two measurement methods – (a) laboratory or (b) in-situ. Both intact cores and repacked samples were measured using the laboratory methods, which include measurements of soil water potential (HYPROP & WP4C, METER), saturated (KSAT, METER) and unsaturated hydraulic conductivity (HYPROP). The in-situ method uses a pair of co-located soil sensors to measure volumetric water content (TEROS12, METER) and soil water potential (TEROS21, METER), and the hydraulic conductivity was not measured. In comparison, the laboratory methods progress from full saturation to dry conditions, and the in-situ method includes both dry-to-wet and wet-to-dry cycles. The sampling and measurement methods for each sample can be found in metadata.csv, and more information about the measurements is detailed in the Methods section below. Models Retention and hydraulic conductivity data were fitted with four van-Genuchten-type models (specified by “model_name” column in the files): (1) traditional constrained van Genuchten model (“vG_constrained”), (2) traditional unconstrained van Genuchten model (“vG_unconstrained”), (3) PDI-variant of the constrained van Genuchten model (“vG_constrained_PDI”), and (4) PDI-variant of the unconstrained van Genuchten model (“vG_unconstrained_PDI”). The difference between the constrained (1: n) and the unconstrained (2: n, m) van Genuchten models is the number of pore-size distribution parameters in the model equations, giving the unconstrained model more degrees of freedom when fitting the data. Between the traditional and the PDI-variant models, model fitting differs the most at the dry end of the measurements. The traditional models allow infinite suction at the residual water content (water content does not drop below residual water content), and the PDI-variant models enforce a soil water potential value of pF=6.8 (~ -630 MPa) at oven-dryness (water content reaches 0). The inclusion of the van-Genuchten-type models is due to their common application. If other retention models are required, users can access the data in data.csv for further data fitting. More information about the models can be found in the Methods section below. Fitting Tasks The model fitting can be categorized into three levels of tasks (specified by “fitting_task” column in the files). Level 1 (“fit_retention”) only includes retention data fitting (the only level available for the in-situ method). Level 2 (“fit_retention_conductivity”) includes both retention and hydraulic conductivity data fitting, and the saturated hydraulic conductivity (Ks, a parameter of the hydraulic conductivity functions) is fixed by the measurements from KSAT. Level 3 (“fit_retention_conductivity_Ks”) also includes both retention and hydraulic conductivity data fitting, but Ks is a fitted parameter without the constraints from KSAT measurements. Among the same retention models (e.g. vG_constrained models of the same sample), level 1 should produce the best retention data fitting. Level 2 should have the highest misfit of the retention and hydraulic conductivity data, because the retention and hydraulic conductivity functions share common model parameters, and the unsaturated hydraulic conductivity (HYPROP) data fitting is subject to Ks measured independently by KSAT. Level 3 should have mid-level misfits of the retention and hydraulic conductivity data. While level 3 fits the hydraulic conductivity data better than level 2, the fitted Ks value might be unreasonable due to the lack of constraints at the wet end of the measurements. General recommendation when using this data package: (1) Choice of sampling methods: Intact cores and in-situ soil sensors could be prioritized because these sampling methods are less destructive. While the repacked samples were packed to the target bulk density (estimated post-sampling, when sample volume was known), these samples had altered pore structures. Nevertheless, intact cores might suffer from sample gaps that would lead to overestimation of Ks (sample gaps can be inferred from the “soil_sample_volume” column in metadata.csv when the value is < 249). In-situ method also has higher uncertainty in characterizing the wet end of the SWRC because of sensor limitations and the difficulty in reaching full saturation under natural conditions. (2) Choice of fitting tasks: When only retention data is needed, level 1 (“fit_retention”) should be prioritized. When both retention and hydraulic conductivity data are needed, level 2 (“fit_retention_conductivity”) could be prioritized. (3) Choice of models: This could depend on what the downstream models call for. If no specific model is required, model misfit could be used as a ranking criterion. Model misfit values in terms of RMSE can be found in model_parameters.csv. The following files are included in this data package: (1) metadata.csv – This file includes the general information of each sample, including location (description, geocoordinates, elevation), sampling and measurements details (method, depth, time or period, volume, instruments), and soil physical properties (bulk density, saturated hydraulic conductivity, only applicable to physical soil samples). (2) data.csv – This file includes soil water potential, volumetric water content, and unsaturated hydraulic conductivity data of each sample. Column “instrument” specifies the instrument (HYPROP, WP4C, or TEROS) used to perform the measurements. (3) model_fit.csv – This file includes soil water potential, volumetric water content, and unsaturated hydraulic conductivity fitted from the four models and three fitting tasks. Column “model_name” specifies the retention model used, and “fitting_task” specifies the level of data fitting. Missing values indicate that the variable does not apply to that fitting task. (4) model_parameters.csv – This file includes the fitted model parameters, model misfits, and conventional water content thresholds (field capacity and wilting point) from the four models and three fitting tasks. Column “model_name” specifies the retention model used, and “fitting_task” specifies the level of data fitting. Missing values indicate that the parameter does not apply to that model and/or that fitting task. (5) data_Ks.csv – This file includes the saturated hydraulic conductivity measurements from KSAT. (6) /figure/*.png – This folder includes three quick visualizations of the data, retention model fitting results and misfits, and hydraulic conductivity model fitting results, misfits, and parameters. The model fitting results are separated by samples and fitting tasks and colored by models. Zoom-in required. (7) /hyprop/*.bdhx – This folder includes proprietary hyprop files that require the free Labros SoilView-Analysis (METER) to open. Users can explore data fitting using other retention models (i.e. Brooks-Corey, Fredlund-Xing, Kosugi, bimodal models). Be aware that Ks value is pre-entered under “Fitting tab, Conductivity functions parameters” for level 2 fitting. If the value is lost, please refer to metadata.csv under “Ks” column. (8) Six file-level metadata that summarize file, header, column, and variable information of all files. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

EARTH SCIENCE > LAND SURFACE > SOILS↗

Metagenome-assembled genomes from topsoils collected during NEON campaign in East River, CO (06/14/2018-06/28/2018)

The Watershed Function Science Focus Area (WF SFA) at Lawrence Berkeley National Lab is working to build a mechanistic understanding of the distribution and dynamics of biogeochemical processes in mountainous watersheds and their response to perturbation. In June 2018, the NEON (National Ecological Observatory Network) Airborne Observatory Platform (AOP) performed a taskable airborne imaging campaign to collect visible to shortwave infrared (VSWIR) imaging spectroscopy and LiDAR data across 330 km2 in the Upper East River at Crested Butte, CO. We conducted a parallel ground sampling campaign to sample vegetation traits, as well as soil physical, chemical, and microbiological characteristics. We collected these samples from 438 sites across 12 locations spanning much of the elevation, topographic, and geologic variability across the study area. A subset of 250 samples were used for soil metagenomics which is presented here. In addition, at each site, vegetation samples were collected to measure species-specific leaf water content and leaf mass area, foliar elemental composition and foliar CN stable isotope ratios. Soil samples were collected to measure soil physical properties which include bulk density and soil texture analysis. A suite of soil chemical properties was measured from the samples collected at each site, including pH, organic matter, concentrations exchangeable cations, total elemental composition, and the concentrations of extractable N pools (e.g. total free amino acids, ammonium, nitrate, dissolved organic N, and total dissolved N). Additionally, we have measured soil microbial biomass CN stoichiometry. Here, we present 1982 metagenome-assembled genomes (MAGs) for the bacterial and archaeal community from topsoil collected from during NEON 2018 campaign. All metagenomes were sequenced at JGI (Joint Genome Institute) (GOLD Study ID: Gs0149986). Metagenomes were assembled using JGI Metagenome Workflow (10.1128/mSystems.00804-20). The dataset includes (1) zip files for 1982 MAG fasta files (neon_genomes1-5.tar.gz, split into 5 tarballs to keep tarballs under 0.5 GB), (2) neon_Gs0149986_samples_soilproperties_metagenomes.csv: the sample information together with the accession numbers for the underlying metagenomes and the associated soil physical and chemical measurements in NMDC (National Microbiome Data Collaborative) compliant format, (3) neon_Gs0149986.kml: location bounding box file for the sampled locations, (4) samples.csv: sample metadata file used to register Internationall Generic Sample Numbers (IGSNs), (5) flmd.csv: file level metadata file, and (6) dd.csv: data dictionary file. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

Mind the gap: Bridging the divide between AI aspirations and the reality of autonomous microscopy

What does materials science look like in the “Age of Artificial Intelligence?” Each material’s domain—synthesis, characterization, and modeling—has a different answer to this question, motivated by unique challenges and constraints. This work focuses on the tremendous potential of autonomous characterization within electron microscopy. We present our recent advancements in developing domain-aware, multimodal models for microscopy analysis capable of describing complex atomic systems. We then address the critical gap between the theoretical promise of autonomous microscopy and its current practical limitations, showcasing recent successes while highlighting the necessary developments to achieve robust, real-world autonomy.

2D materials↗

Data from TropiRoot 1.0 database: tropical root characteristics across environments

TropiRoot 1.0 is a new tropical root database with root characteristics across environment gradients. It has data extracted from 104 new sources, resulting in more than 8000 rows of data (either species or community data). Most of the data in TropiRoot 1.0 includes root characteristics such as root biomass, morphology, root dynamics, mass fraction, architecture, anatomy, physiology and root chemistry. This initiative represents an approximately 30% increase in the currently available data for tropical roots in the Fine Root Ecology Database (FRED). TropiRoot 1.0, contains root characteristics from 25 different countries where seven are located in Asia, six in South America, five in Central America and the Caribbean, four in Africa, two in North America, and 1 in Oceania. Due to the volume of data, when ancillary data was available, including soil data, these data was either extracted and included in the database or their availability was recorded in an additional column. Multiple contributors checked the entries for outliers during the collation process to ensure data quality. For text-based observations, we examined all cells to ensure that their content relates to their specific categories. For numerical observations, we ordered each numerical value from least to greatest and plotted the values, checking apparent outliers against the data in their respective sources and correcting or removing incorrect or impossible values. Some data (soil and aboveground) have different columns for the same variable presented in different units, including originally published units, but root characteristics data had units converted to match the ones reported in FRED. By filling a gap from global databases, TropiRoot 1.0 expands our knowledge of otherwise so far underrepresented regions, and our ability to assess global trends. This advancement can be used to improve tropical forest representation in vegetation models.

54 ENVIRONMENTAL SCIENCES↗

Hot Droughts and Forest Tree Dynamics in the Amazon - Statistical Models, Scripts, Data, and Outputs

This package contains data, outputs, equations, and R scripts for analyses for manuscript entitled "Hot droughts in the Amazon: A window to a future hypertropical climate" by J. Chambers et al., in particular it contains statistical models and analyses for the INPA BIONTE tree mortality study. The Models folder contains details for all statistical models in PDF files. The Scripts folder contains the R scripts for Bayesian Hierarchical Models (two text files) and SEMs (one text file) are separate and reasonably annotated. All data associated with these scripts are in the data folder. The Data folder contains two of the three CSV files used for the analyses and are called by the R scripts. Two of them are part of published datasets (`BIONTE_mortality-rates.csv` from Lima et al. 2024, DOI:10.15486/ngt/1898910 and `SPEI.csv` from Pastorello et al. 2023 DOI:10.15486/ngt/1958257) and also provided in this package for convenience (please see the corresponding datasets for usage and citation terms). The third dataset (`BIONTE_gapfilled_wd.csv`) contains sensitive information and can be obtained by contacting the manuscript lead author. The Outputs folder contains the two output files that provide extra information about the analyses. The file `figuresFeb2025d.pdf` contains all the figures from the manuscript - captions are in the manuscript. The file `ChambersMS.pdf` contains primary results from Bayesian statistical models, regression analyses, and validation steps applied to the tree mortality data from the INPA experiments. The document includes visual summaries, model diagnostics, and leave-one-out (LOO) validation results. A breakdown of file contents can be found in the README file that is part of this package.

54 ENVIRONMENTAL SCIENCES↗

Using UMM-Var and E2E to Improve the User Experience for Accessing NASA EOSDIS Data Sets

The UMM-Variables (Var) Metadata Model has been evolved to support an End-to-End Services (E2E) capability, which enables variable level subsetting, data transformation, and data reformatting. This talk will discuss what is new with the model, how users can get their metadata ready for the E2E capability, and include a demo of how the model is being used to drive and improve the user experience in Earthdata Search Client when accessing EOSDIS data sets.

Metadata Modeling↗