Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “metadata extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

bibcheck

SAND2026-16981O Bibcheck is designed to extract bibliographies from research papers and perform metadata searches to identify errors. It assists authors in checking their bibliographies for metadata errors during the writing process and helps reviewers identify errors in bibliographies of papers under review. The software uses large language models (LLMs) to extract bibliography entries from PDF documents, classifies the type of bibliography entry, and verifies referenced works. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Pearson, Carl [Sandia National Lab. (SNL-CA), Live↗

Responses of Drying-rewetting (Transient Soil Moisture) and Steady State Soil Moisture Incubation on Soil Organic Carbon Dynamics in Three US Soils, 2017

This data set contains measurements of soil characteristics (aggregate size distribution and mean size, total aggregate associated carbon, extractable organic C, and microbial biomass C), microbial respiration, and soil metabolite concentrations from a transient and steady soil moisture incubation experiment using soils of different textures (sandy, loamy, and clayey). The study investigated mechanisms driving the Birch effect (increased carbon mineralization pulses with wetting following a drying period) in differing soil textures. Three different soils of distinctly different textures were collected from 0-15cm depth in Georgia (sandy, 2017-05-01), Missouri (loamy, 2017-06-14), and Texas (clayey, December 2017). Soils were incubated for 140 days with destructive harvests done on days 1, 29, 33, 56, 112, 116, and 140 in transient soil moisture incubation and on days 1, 33, 116, and 140 in steady state soil moisture incubation. This dataset contains six data files in comma separate (.csv) format. Additional metadata are provided: six data dictionaries and a file-level metadata file in comma separate (*.csv) format and a user guide in PDF (*.pdf) format.

1-Methyladenosine concentration↗

Differential Organic Carbon Mineralization Responses to Soil Moisture in Three Different Soil Orders Under Mixed Forested System: Supporting Data

This data contains data from 90-day long incubation study which aimed to look at the soil moisture-texture relationship on soil organic carbon (SOC) cycling. Soils were collected from three distinct soil textures from mixed forests in 2017: sandy (Georgia, 2017-05-01), loamy (Missouri, 2017-06-14) and clayey (Texas, December 2017) were incubated at different soil moisture levels (air-dried, 25% water holding capacity (WHC), 50% WHC, 100% WHC and 175% WHC) at room temperature for a period of 90 days. Files contain microbial respiration, active and slow SOC pools, and their respective mineralization rates, extractable organic carbon (C), and C-acquiring extracellular enzymes. Findings from these data were used in Singh et al. (2021). This study aimed to examine the interactive effect of soil moisture and texture on SOC mineralization. Soil samples of three distinct textures (sandy, loamy, and clayey) were collected from mixed forests of Georgia, Missouri, and Texas, respectively. Soil cores of 5 cm diameter were collected from numerous random locations at each site from 0-15 cm depth after scraping the litter layer and mixed thoroughly to obtain a composite sample per site. Three additional soil cores were collected to determine the WHC using pressure plate extractors. Soil samples were composited, and triplicate soil samples were incubated in mason jars for a period of 90 days at room temperature under different moisture regimes: air dried, 25% WHC, 50% WHC, at WHC and 100% saturation. Soil respiration was measured weekly, and destructive sampling was conducted at 1, 15, 60, and 90 days to determine extractable organic C, C acquiring enzyme activity, and active and slow SOC pools with their respective mineralization rates. The C acquiring enzyme activity was the total activity of α-glucosidase, β-glucosidase, cellobiohydrolase, and β-xylosidase enzymes. Gas samples for microbial respiration measurements were collected from headspace of incubation jars through the sampling ports on the lids and then analyzed using a Shimadzu Gas Chromatograph (GC-2014). Prior to sampling, the vials were evacuated. Blank correction was also done by collecting gas samples from empty incubation jars. Double pool exponential decay model was used in SigmaPlot to determine the active and slow SOC pools and their mineralization rates (Farrar et al., 2012; Jagadamma et al., 2014). The C-acquiring extracellular enzymes were measured using the microplate method by German et al., (2011). Microbial community structure was determined using the phospholipid fatty acid (PLFA) and neutral lipid fatty acid (NLFA) analyses (Buyer and Sasser, 2012). This dataset has seven data files provided in comma-separate (*.csv) format. Additional metadata are provided: seven data dictionaries and a file-level metadata file in comma separate (*.csv) format and a user guide in PDF (*.pdf) format.

Carbon acquiring enzyme activity↗

Reproductive and leaf litterfall fluxes in forest ecosystem sites globally (1950-2022)

Forest allocation of net primary productivity (NPP) to reproduction is poorly quantified globally, despite its critical role in forest regeneration and a well-supported trade-off with allocation to growth. Although field measurements of total NPP are rare, our work finds that a proxy for reproductive carbon allocation constructed from leaf (L) and reproductive (R) litterfall fluxes, R/(R+L), is strongly correlated with R/NPP, facilitating analysis across a wide range of sites where biometric estimates of NPP are not available (R² = 0.85; Hanbury-Brown et al., 2022, Ward et al., in prep). To investigate relationships between ecosystem-scale reproductive allocation (RA) and climate, soil fertility, and stand age gradients, we conducted a literature search and synthesized 824 observations of annual average leaf and reproductive litterfall fluxes across forest sites globally. The zip file includes 1) a folder Data/ containing the litterfall data ("GlobalForestRA_data.csv") and metadata ("GlobalForestRA_metadata.doc") files. The data file includes geographic coordinates, long-term mean annual temperature and precipitation (1970-2000, extracted from WorldClim2.1), leaf and reproductive litterfall fluxes, sampling interval and protocols, forest characteristics (dominant leaf morphology, information pertaining to forest age and successional stage, and disturbance history) and soil properties (% sand, %silt, %clay, total phosphorus (P), nitrogen (N), cation exchange capacity (CEC) and pH) extracted from SoilGrids250 and from on-site measurements, where available. The metadata file contains information about each variable reported in the data file, including data sources, processing methods, and all references. The Data folder contains two additional files used to create Figure 1; these are described in greater detail in the README.2) R scripts GloalForestRA_analysis.r and GlobalForestRA_SI.r and a folder /Functions used to produce results, figures, and tables in the manuscript Ward et al. (in press)3) a README file describing how the data and R scripts can be used to reproduce statistical results, figures, and tables found in the manuscript. Ward et al. (in press)This repository can also be found at: https://github.com/r-ward/Global_Analysis_ForestRA.Ward, R.E., Zhang-Zheng, H. Aernethy, K., Adu-Bredu, S., Arroyo, L., Bailey, A. et al. (in press). Forest age rivals climate to explain reproductive allocation patterns in forest ecosystems globally. Ecology Letters. Hanbury-Brown, A.R., Ward, R.E. & Kueppers, L.M. (2022). Forest regeneration within Earth system models: current process representations and ways forward. New Phytol., 235, 20–40.Ward et al. (2025), Forest age rivals climate to explain reproductive allocation patterns in forest ecosystems globally, in prep.

54 ENVIRONMENTAL SCIENCES↗

Whole metagenome sequencing and 16S rRNA gene amplicon analyses reveal the complex microbiome responsible for the success of enhanced in-situ reductive dechlorination (ERD) of a tetrachloroethene-contaminated Superfund site

The North Railroad Avenue Plume (NRAP) Superfund site in New Mexico, USA exemplifies successful chlorinated solvent bioremediation. NRAP was the result of leakage from a dry-cleaning that operated for 37 years. The presence of tetrachloroethene biodegradation byproducts, organohalide respiring genera (OHRG), and reductive dehalogenase (rdh) genes detected in groundwater samples indicated that enhanced reductive dechlorination (ERD) was the remedy of choice. This was achieved through biostimulation by mixing emulsified vegetable oil into the contaminated aquifer. This report combines metagenomic techniques with site monitoring metadata to reveal new details of ERD. DNA extracts from groundwater samples collected prior to and at four, 23 and 39 months after remedy implementation were subjected to whole metagenome sequencing (WMS) and 16S rRNA gene amplicon (16S) analyses. The response of the indigenous NRAP microbiome to ERD protocols is consistent with results obtained from microcosms, dechlorinating consortia, and observations at other contaminated sites. WMS detects three times as many phyla and six times as many genera as 16S. Both techniques reveal abundance changes in Dehalococcoides and Dehalobacter that reflect organohalide form and availability. Methane was not detected before biostimulation but appeared afterwards, corresponding to an increase in methanogenic Archaea. Assembly of WMS reads produced scaffolds containing rdh genes from Dehalococcoides, Dehalobacter, Dehalogenimonas, Desulfocarbo, and Desulfobacula. Anaerobic and aerobic cometabolic organohalide degrading microbes that increase in abundance include methanogenic Archaea, methanotrophs, Dechloromonas, and Xanthobacter, some of which contain hydrolytic dehalogenase genes. Aerobic cometabolism may be supported by oxygen gradients existing in aquifer microenvironments or by microbes that produce O 2 via microbial dismutation. The NRAP model for successful ERD is consistent with the established pathway and identifies new taxa and processes that support this syntrophic process. This project explores the potential of metagenomic tools (MGT) as the next advancement in bioremediation.

59 BASIC BIOLOGICAL SCIENCES↗

Leaf phenology data at The Morton Arboretum Forestry Plots 2019-2023

We are collecting long-term leaf phenology data at The Morton Arboretum to determine seasonal patterns of leaf production in trees. This data on leaf phenology will be integrated with other ongoing data streams to create a connection between above- and below-ground tree processes. This data package contains raw and smooth outputs from phenology data, as well as extracted phenophase dates (i.e., start, peak, and end of season): the raw and smooth outputs from the PhenoCam GUI can be found in the "leafRaw.csv" and "leafSmooth.csv" files, respectively, and the extracted phenophase dates can be found in the "leafPhenophaseDates2019-2023.csv" file. Extracted phenophase dates for evergreen species in 2023 are currently unavailable, and the files will be updated once they are extracted. Additional information on units and other file-level metadata can be found within each data file's respective data dictionary, and metadata for each of the 23 surveyed plots can be found within the "Location_metadata.csv" file. While the "leafRaw.csv" and the "leafSmooth.csv" files contain all data for all species, the "leafPhenophaseDates2019-2023.csv" file currently excludes the dates for evergreen species in 2023. Another version of the file will be added as those dates are extracted.

54 ENVIRONMENTAL SCIENCES↗

Multispectral and thermal surface imagery and surface elevation mosaics (camspec-air)

This dataset contains high resolution image products (orthomosaics) acquired from midsized uncrewed aerial systems, which have been processed for value added quality. The instrument itself, a multispectral imager, the Altum by Micasense, captures 6 spectral bands (red, green blue, NIR, red edge, and LWIR/thermal1) as radiance, which is converted to reflectance. The code used to develop these images first uses tools from the Micasense python library2 to apply dark level corrections, row gradient corrections, and radiometric corrections. Next it uses the processing API from Agisoft Metashape software to align and mosaic the processed imagery, following the processes developed by the USGS' structure from motion workflow documentation3. Captures at different altitudes (recorded in MSL) produce an orthomosaic, a tif image containing information related to the 6 spectral bands, and a digital elevation model (DEM), a tif image containing information related to the elevation of the surveyed terraine. Metadata included in every image can be used to extract lat, lon, and reflectance values. 1https://www.arm.gov/publications/tech_reports/handbooks/doe-sc-arm-tr-281.pdf 2https://micasense.github.io/imageprocessing/MicaSense%20Image%20Processing%20Setup.html 3https://pubs.usgs.gov/of/2021/1039/ofr20211039.pdf

54 ENVIRONMENTAL SCIENCES↗

Multispectral and thermal surface imagery and surface elevation mosaics - SGP July 2022

This data set contains high-resolution image products (orthomosaics) acquired from midsized uncrewed aerial systems that have been processed for value-added quality. The instrument itself, a multispectral imager, the Altum by Micasense, captures six spectral bands (red, green blue, NIR, red edge, and LWIR/thermal1) as radiance, which is converted to reflectance via custom code. The code used to develop these images first uses tools from the Micasense Python library2 to apply dark level corrections, row gradient corrections, and radiometric corrections. Next, it uses the processing API from Agisoft Metashape software to align and mosaic the processed imagery, following the processes developed by the USGS' structure from motion workflow documentation.3 Captures at different altitudes (recorded in MSL) produce an orthomosaic, a tif image containing information related to the six spectral bands, and a digital elevation model (DEM), a tif image containing information related to the elevation of the surveyed terraine. Metadata included in every image can be used to extract lat, lon, and reflectance values. 1 https://www.arm.gov/publications/tech_reports/handbooks/doe-sc-arm-tr-281.pdf 2 https://micasense.github.io/imageprocessing/MicaSense%20Image%20Processing%20Setup.html 3 https://pubs.usgs.gov/of/2021/1039/ofr20211039.pdf

54 ENVIRONMENTAL SCIENCES↗

A Data Processing Pipeline To Extract A Knowledge Graph From Heterogeneous Data For Socio-technical Analysis Of Critical Infrastructure Influence

The code is written in Python and consists of the following pipeline that is implemented in Apache Airflow. This pipeline intends to understand the companies that are directly or indirectly involved with a type of critical infrastructure system at some point in that system's lifecycle. The pipeline takes a configuration file that specifies a list of initial companies to consider, a geographic region of interest, and a set of SEC form types as well as other data sources (e.g. CrunchBase) from which to extract entities and relations. There are four main components to this pipeline as currently implemented: Entity Extraction, Network Construction, Analysis, and Visualization. First, Entity Extraction, is implemented as the `topear-extract_organizations` Apache Airflow workflow. Given an initial query that specifies a geographic region of interest and a time interval, the software will extract CI facilities of interest and organizations that have a direct influence relationship to those facilities (e.g. ownership). During the course of the LDRD, we focused on Electric Vehicle charging stations and this information is available via the Department of Energy (DOE) database on fueling stations maintained by NREL. Within the context of the DOE CESER project, we have focused on Battery Energy Storage Systems (BESS). Second, the Network Extraction component will iteratively construct a social network graph given the set of organizations and people extracted in the previous step. Organizations (and eventually People if desired) are then fed as a query to the `topgear-construct_social_network` Apache Airflow workflow which given a set of initial companies and data sets (e.g. SEC EDGAR form types, OpenCorporates, Crunchbase). This Airflow workflow will iteratively query such data sources to discover relationships with new organizations and people. For example, this module can iteratively query SEC EDGAR for metadata that documents the number of each type of form for the given set of companies and their location. This forms metadata represents a catalog of data sources from SEC EDGAR for the extracted social network knowledge graph. The pipeline then downloads these forms from the website and saves them in a build directory for further processing. These documents are then parsed for entities and relations. Again, we note that in additional to SEC data sources, this step can also pull in information on organizations via API services such as CrunchBase and OpenCorporates or bulk data sources. At the end of this step, the resultant social network, the Critical Infrastructure network, and the edges that encode relationships between organizations and CI facilities, form the Adversarial Socio-Technical Network (ASTN) that informs the analysis. Third, the Analysis component processes these generated ASTN. Previously, that has included the ability to compare prevalence of different vendors for a given infrastructure component type across different regions as well as identify common public and private investors across those vendors. This was demonstrated for EV Charging Stations across several different metropolitan areas within an IEEE PES GridEdge publication. More recently, we have looked at ways to identify infrastructure owners and operators of BESS with the most nameplate capacity across different states as well as other indictors of risk resulting from changes in ownership over time. Finally, the Visualization component consists of an HTML/CSS/JS framework by which users can interact geospatial, operational, and organizational relationships across a given portfolio of Critical Infrastructure facilities. The objective is to provide a library of UI/UX modules that can be repurposed for stakeholder-specific dashboards. All of the modules are related via a common event model that enables UI actions in one view to percolate across the other views.

Weaver, Gabriel [Idaho National Laboratory (INL), ↗

Optical emissivity dataset of multi-material heterogeneous designs generated with automated figure extraction

Optical device design is typically an iterative optimization process based on a good initial guess from prior reports. Optical properties databases are useful in this process but difficult to compile because their parsing requires finding relevant papers and manually converting graphical emissivity curves to data tables. Here, we present two contributions: one is a dataset of thermal emissivity records with design-related parameters, and the other is a software tool for automated colored curve data extraction from scientific plots. We manually collected 64 papers with 176 figures reporting thermal emissivity and automatically retrieved 153 colored curve data records. The automated figure analysis software pipeline uses Faster R-CNN for axes and legend object detection, EasyOCR for axes numbering recognition, and k-means clustering for colored curve retrieval. Additionally, we manually extracted geometry, materials, and method information from the text to add necessary metadata to each emissivity curve. Finally, we analyzed the dataset to determine the dominant classes of emissivity curves and determine the underlying design parameters leading to a type of emissivity profile.

47 OTHER INSTRUMENTATION↗

Data and scripts associated with a manuscript investigating impacts of solid phase extraction on freshwater organic matter optical signatures and mass spectrometry pairing

This data package is associated with the publication “Investigating the impacts of solid phase extraction on dissolved organic matter optical signatures and the pairing with high-resolution mass spectrometry data in a freshwater system” submitted to “Limnology and Oceanography: Methods.” This data is an extension of the River Corridor and Watershed Biogeochemistry SFA’s Spatial Study 2021 (https://doi.org/10.15485/1898914). Other associated data and field metadata can be found at the link provided. The goal of this manuscript is to assess the impact of solid phase extraction (SPE) on the ability to pair ultra-high resolution mass spectrometry data collected from SPE extracts with optical properties collected on ambient stream samples. Forty-seven samples collected from within the Yakima River Basin, Washington were analyzed dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC), absorbance, and fluorescence. Samples were subsequently concentrated with SPE and reanalyzed for each measurement. The extraction efficiency for the DOC and common optical indices were calculated. In addition, SPE samples were subject to ultra-high resolution mass spectrometry and compared with the ambient and SPE generated optical data. Finally, in addition to this cross-platform inter-comparison, we further performed and intra-comparison among the high-resolution mass spectrometry data to determine the impact of sample preparation on the interpretability of results. Here, the SPE samples were prepared at 40 milligrams per liter (mg/L) based on the known DOC extraction efficiency of the samples (ranging from ~30 to ~75%) compared to the common practice of assuming the DOC extraction efficiency of freshwater samples at 60%. This data package folder consists of one main data folder with one subfolder (Data_Input). The main data folder contains (1) readme; (2) data dictionary (dd); (3) file-level metadata (flmd); (4) final data summary output from processing script; and (5) the processing script. The R-markdown processing script (SPE_Manuscript_Rmarkdown_Data_Package.rmd) contains all code needed to reproduce manuscript statistics and figures (with the exception of that stated below). The Data_Input folder has two subfolders: (1) FTICR and (2) Optics. Additionally, the Data_Input folder contains dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data (SPS_NPOC_Summary.csv) and relevant supporting Solid Phase Extraction Volume information (SPS_SPE_Volumes.csv). Methods information for the optical and FTICR data is embedded in the header rows of SPS_EEMs_Methods.csv and SPS_FTICR_Methods.csv, respectively. In addition, the data dictionary (SPS_SPE_dd.csv), file level metadata (SPS_SPE_flmd.csv), and methods codes (SPS_SPE_Methods_codes.csv) are provided. The FTICR subfolder contains all raw FTICR data as well as instructions for processing. In addition, post processed FTICR molecular information (Processed_FTICRMS_Mol.csv) and sample data (Processed_FTICRMS_Data.csv) is provided that can be directly read into R with the associated R-markdown file. The Optics subfolder contains all Absorbance and Fluorescence Spectra. Fluorescence spectra have been blank corrected, inner filter corrected, and undergone scatter removal. In addition, this folder contains Matlab code used to make a portion of Figure 1 within the manuscript, derive various spectral parameters used within the manuscript, and used for parallel factor analysis (PARAFAC) modeling. Spectral indices (SPS_SpectralIndices.csv) and PARAFAC outputs (SPS_PARAFAC_Model_Loadings.csv and SPS_PARAFAC_Sample_Scores.csv) are directly read into the associated R-markdown file. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

ESS-DIVE Unoccupied Aerial Systems (UAS) Reporting Format v1

Here we present documentation of the ESS-DIVE reporting format for Unoccupied Aerial System (UAS) data and metadata. This reporting format provides guidance to data contributors on how to store data to maximize their discoverability, facilitate their efficient reuse, and add value to individual datasets. For data users, the reporting format will better allow data repositories to optimize data search and extraction, and more readily integrate similar data into harmonized synthesis products. The reporting format provides templates and guidance for the reporting of metadata for UAS experimental campaigns, individual flights, platform and sensor description. To improve data access and discoverability, the reporting format proposes a data description scheme of Levels based on the degree of processing, where Level 0 includes raw data, through to Level 3 being derived data end products. A range of examples of data types for each Level are given, with suggested file naming schemes. The reporting format presented here is intended to form a foundation for future development that will accommodate new UAS technologies and approaches to data access and use in the future. The reporting format documentation is maintained and updated on the ESS-DIVE Community Space GitHub at https://github.com/ess-dive-community/essdive-uas. This data package is the first published version of this reporting format, and comprises a zip file of the complete content of https://github.com/ess-dive-community/essdive-uas v1.0. The zip contains the reporting format description, instructions and variable definitions in GitHub markdown language (*.md) and metadata templates in csv format. The reporting format is designed to be compatible with other ESS-DIVE formats, and it is specifically recommended that this reporting format be used in conjunction with the File-level metadata (FLMD) and comma separated values (csv) reporting formats for submission to the ESS-DIVE repository.

54 ENVIRONMENTAL SCIENCES↗

Biogeochemistry of Pond B (Savannah River Site, South Carolina, USA): Sediment Core, Total extraction data, Pond B Savannah River Site July 2019. Subsurface Biogeochemistry of Actinides SFA

Pond B at Savannah River Site (SRS, South Carolina) is a monomictic reservoir that received SRS R reactor cooling water from 1961–1964. Previous studies conducted between the 1980s–1990s on the water column and sediments of Pond B measured trace amounts of Pu (33 MBq 238Pu and 430 MBq 239,240Pu), 241Am, and 137Cs. Since then, the pond has been relatively isolated and the radionuclide concentrations have not been monitored over time. Herein, about 30 years after the last publication on Pond B, we are re-evaluating the geochemistry and radionuclide distribution within Pond B at four locations along a horizontal transect from the inlet to outlet.This study investigated the distribution of anthropogenic radionuclides Pu-239 and Cs-137 along with total organic carbon, iron, and trace element in contaminated sediments of Pond B at the Savannah River Site (SRS). Pond B received reactor cooling water from 1961 to 1964, and trace amounts of Pu-239 and Cs-137 during operations. Our study collected sediment cores to determine concentrations of Pu-239, Cs-137, and major and minor elements in solid phase, pore water and an electrochemical method was used on wet cores to determine dissolved elemental concentrations.

54 ENVIRONMENTAL SCIENCES↗

Data from a four-day long microcosm experiment addressing the destabilization of artificial mineral-associated organic matter by model root exudates embedded in a soil matrix from the Rocky Mountain Biological Laboratory (Gothic, CO, USA), 2019

This dataset provides data collected during a four-day long laboratory soil microcosm experiment testing the efficacy of root exudate-driven mineral-associated organic matter destabilization. This dataset contains four data files in comma-separate values (*.csv). The files provide the metadata and the experimental results on microbial respiration, MAOM-derived respiration, and sequential mineral-extractions. This data was used to produce the figures in Bölscher et al., 2026. The results of the experiment can be found in the open access article Bölscher et al., 2026 (https://doi.org/10.1016/j.soilbio.2026.110276). Abstract: Mineral-associated organic matter (MAOM) is often considered stable, but root exudates can destabilize MAOM via various pathways. Theory and model system studies suggest that direct MAOM destabilization by strong ligands, like oxalic acid, or reducing agents, like catechol, is more effective than indirect, microbial-mediated MAOM destabilization, stimulated by less reactive compounds like glucose. Here, we demonstrate that the presence of a soil matrix alters the efficacy of exudate-driven MAOM destabilization pathways. Glucose and catechol destabilized significantly greater amounts of MAOM from ferrihydrite and aluminum hydroxide (Al (OH)3) embedded in a soil matrix than oxalic acid. Our findings indicate that indirect, microbial-mediated MAOM destabilization may play a larger role than direct MAOM destabilization in soil environments.

Destabilization↗

Quarterly Soil Core and Root Analyses from the Missouri Ozarks AmeriFlux (MOFLUX) Site, Ashland, Missouri, 2017-2023

This dataset contains quarterly soil core measurements from the Missouri Ozarks AmeriFlux (MOFLUX) site located at the University of Missouri’s Thomas H. Baskett Wildlife Research and Education Area near Ashland, Missouri. These data will be used to parameterize an ensemble of MOFLUX-optimized soil carbon-nitrogen models, used to simulate carbon (C) and nitrogen (N) cycling responses to future hydroclimatic scenarios and the trajectory of soil C stocks with concomitant forest decline. Beginning in 2017, eight soil cores were collected approximately quarterly near plot 1 of the southeast transect, near the automated soil respiration flux chambers, from 0–15 cm depth. Data are currently available through 2023 (2017-06-14 to 2023-11-13); additional observations will be appended to this dataset as they become available. Cores were analyzed for gravimetric moisture content, pH, total carbon and nitrogen, texture, microbial biomass carbon and nitrogen, and extractable dissolved organic carbon and nitrogen. This dataset contains one data file in comma separate (*.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma separate (*.csv) format and a user guide in PDF (*.pdf) format.

54 ENVIRONMENTAL SCIENCES↗

APPL Hyperspectral_Imaging_Dataset_for_Heritability_Analysis_in_Populus_trichocarpa

This dataset contains hyperspectral imaging data collected at the Advanced Plant Phenotyping Laboratory (APPL) at Oak Ridge National Laboratory. Natural variants of Populus trichocarpa were imaged using a high-throughput hyperspectral phenotyping pipeline to quantify spectral reflectance traits for downstream quantitative genetics analyses. The dataset includes hyperspectral image files and derived reflectance data products suitable for extracting spectral features across the measured wavelength range (e.g., VNIR and/or SWIR, depending on instrument configuration), along with associated sample metadata (e.g., genotype identifiers, experimental design factors, and imaging run identifiers). These data were generated to support analyses of broad-sense heritability of hyperspectral traits and their relationships with biochemical phenotypes (including lignin traits from Py-MBMS).

APPL↗

CHESS 2025: Crown polygons and extracted reflectance for field sampling sites

This dataset contains (1) crown polygons for each tree, meadow, and shrub site sampled in the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS) campaign (in geojson format, .geojson) and (2) extracted reflectance, uncertainty, and shade estimates for each crown polygon from the 2018 National Ecological Observatory Network (NEON) and 2025 CHESS campaigns. (in CSV format, .csv). Additional metadata are provided in a data dictionary describing column names and definitions (dd.csv), and in a file-level metadata file (flmd.csv). Crown polygons were manually delineated for each site in the 2025 campaign using a combination of field-collected GPS data (doi:10.15485/3022418), RGB (red, green, blue) and false color reflectance mosaics (doi:10.15485/3013535), and LiDAR-derived (Light Detection and Ranging) canopy height (CHM) and digital surface (DSM) models (DOI and citation to be added upon publication). Where there was misalignment between the spectrometer- and LiDAR-derived data products, polygons prioritized alignment with the spectrometer-derived data products. Polygons were delineated conservatively to only select pixels representative of vegetation samples collected in the field. Crown polygons for 2018 are published at (doi:10.15485/1618130) and were developed using the same protocol. For each polygon, all pixels from all flightlines were extracted where the pixel centroid was contained within the polygon. For each pixel, we extracted the surface reflectance, uncertainty, and shade estimates. Details on the extracted datasets are available at (doi:10.15485/3013527, doi:10.15485/3013535). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: This research was carried out at the Jet Propulsion Laboratory, California Institute of Technology, under a contract with the National Aeronautics and Space Administration (80NM0018D0004) and was funded by EMIT Extended Mission Phase E Science.

2018 NEON and 2025 CHESS Campaigns↗

HydroDCM: Hydrological Domain-Conditioned Modulation for Cross-Reservoir Inflow Prediction

Deep learning models have shown promise in reservoir inflow prediction, yet their performance often deteriorates when applied to different reservoirs due to distributional differences, referred to as the domain shift problem. Domain generalization (DG) solutions aim to address this issue by extracting domain-invariant representations that mitigate errors in unseen domains. However, in hydrological settings, each reservoir exhibits unique inflow patterns, while some metadata beyond observations like spatial information exerts indirect but significant influence. This mismatch limits the applicability of conventional DG techniques to many-domain hydrological systems. To overcome these challenges, we propose HydroDCM, a scalable DG framework for cross-reservoir inflow forecasting. Spatial metadata of reservoirs is used to construct pseudo-domain labels that guide adversarial learning of invariant temporal features. During inference, HydroDCM adapts these features through light-weight conditioning layers informed by the target reservoir’s metadata, reconciling DG’s invariance with location-specific adaptation. Experiment results on 30 real-world reservoirs in the Upper Colorado River Basin demonstrate that our method substantially outperforms state-of-the-art DG baselines under many-domain conditions and remains computationally efficient.

Hu, Pengfei [ORNL] (ORCID:0009000367130950)↗