Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “metadata extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Changuinola peat soil characteristics and gas emission raw data October 2019

This dataset comprises radiocarbon and geochemical measurements from peat and porewater samples collected across various depths at a site in Bocas del Toro, Panama. The study focuses on carbon cycling dynamics in tropical peatlands by examining carbon isotopic signatures (¹⁴C and ¹³C) and elemental compositions of bulk peat, dissolved organic carbon (DOC), carbon dioxide (CO₂), and methane (CH₄). Key parameters include radiocarbon ages and isotopic ratios (δ¹³C) of bulk peat, concentrations of carbon (%C) and nitrogen (%N), and radiocarbon content of porewater gases and dissolved organic carbon (DOC). The data provide insights into the vertical and spatial distribution of carbon sources and possible preservation and decomposition processes within tropical peat profiles, offering critical information for understanding carbon storage and greenhouse gas emissions in these ecosystems.This dataset is comprised of one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) carbon isotopic signatures (¹⁴C and ¹³C); (5) concentrations of carbon (%C) and nitrogen (%N); (6) radiocarbon content of porewater carbon dioxide (CO₂), and methane (CH₄) ; (7) porewater DOC; (8) bulk peat sampling protocol; (9) porewater sampling protocol; (10) porewater gas collection methods; and (11) gas extraction methods. All files are in .csv format and can be opened with any software that supports this file types.

54 ENVIRONMENTAL SCIENCES↗

1H-NMR characterization of soil dissolved organic matter from soil samples in control and warming plots in Blodgett Forest, CA (2014 and 2018)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of Lawrence Berkeley National Laboratory Terrestrial Ecosystem Science Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM (soil organic matter) decomposition and stabilization. This package contains metabolite data obtained through 1H nuclear magnetic resonance (NMR) spectroscopy on water-extracted soils. Soil samples were collected in 2014/06/03 and 2018/06/04 from 3 replicated paired plots that had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. The following files are included: (1) nmr_h2o_data_raw.csv: raw data, (2) nmr_h2o_data_processed.csv: computed compound concentrations and metadata, (3) nmr_h2o_compound_metadata.csv: compound metadata, (4) nmr_h2o_sample_metadata.csv: sample metadata

1H-NMR (nucleic magnetic resonance) spectroscopy↗

FTICR-MS Data from Multi-continent River Water and Sediment and from Coastal River Fresh and Saline Sediment Associated with: “Dissolved Organic Matter Functional Trait Relationships are Conserved Across Rivers”

This data package is associated with the publication “Dissolved Organic Matter Functional Trait Relationships are Conserved Across Rivers” submitted to PNAS (Stegen et al., 2023). The study aims to understand large-scale spatial structure of the dissolved organic matter (DOM) thermodynamic traits and inter-trait relationships by investigating (1) river water and sediments collected along 97 rivers spanning 3 continents and (2) coastal sediment collected from fresh and saline locations in Pacific and Gulf/Atlantic rivers. Sediment extracts and water samples were analyzed using ultrahigh resolution Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS). This dataset is comprised of three folders (1) Coastal, (2) WHONDR_S19S, and (3) Data_Dictionaries. Coastal contains (1) a subfolder with processed FTICR-MS data as csv files and sample collection metadata, (2) a subfolder with R scripts used to process the data and create associated figures, (3) a subfolder with the raw, unprocessed FTICR-MS data as .xml files, and (4) a readme file with more information about the dataset and instructions for using Formularity (https://omics.pnl.gov/software/formularity). WHONDRS_S19S contains (1) a csv file with processed FTICR data, (2) a csv with sample collection metadata, (3) a csv with sample geospatial data, (4) a csv with simulated lambda model outputs, (5) a subfolder with R scripts used to process the data and create associated figures, and (6) a readme file with more information regarding WHONDRS raw FTICR data and processing scripts. Data_Dictionaries contains data dictionaries for each csv file in the data package. The 97 global river corridors were part of a WHONDRS (https://whondrs.pnnl.gov) study. The raw, unprocessed FTICR-MS data with additional data can be found at doi:10.15485/1729719 for sediments and doi:10.15485/1603775 for water. This data package contains the processed data used in the associated manuscript. The coastal data has not been previously published, and this data package contains both the raw and processed data. Version 3 of this data package published February 2023 includes updates to the title of the manuscript, additional data and data dictionary and updated scripts linked to new analysis.

54 ENVIRONMENTAL SCIENCES↗

Carbon Storage Technical Viability Approach (CS TVA) Database

The Carbon Storage Technical Viability Approach (CS TVA) database was developed to support the implementation of the CS TVA Matrix to a national data availability assessment for technically viable carbon storage. This database leverages the efforts of multiple adjacent and overlapping databases by non-redundantly combining the databases into a single database along with additionally providing tags facilitating the CS TVA. The non-redundant aspect of the database permits an accurate assessment of the concentration of available data, aiding in spatial and categorical data gaps analysis relative to the individual CS TVA Matrix Components. Version 2.0 of the database is an expansion of Version 1.0. Version 2.0 was created to include additional data gathered to fill gaps in the existing data set. Downloading the CS TVA v2.0 database will result in two separate databases, the version 1.0 original .gdb, and a second addendum .gdb with the new data gathered, together these two databases make up v2.0. Please see the ReadMe file below for full details, metadata information, use disclaimer, and attributions.

Coal↗

SPRUCE Quantitative PCR (qPCR) of Microbial Gene Copy Numbers, 2021-2022

This dataset provides the results for quantitative polymerase chain reaction (qPCR) of peat samples collected from ambient and experimental plots in the Spruce and Peatland Responses Under Climatic and Environmental Change (SPRUCE) experiment site in June and August of 2021, and June of 2022. SPRUCE is located within the Marcell Experimental Forest in northern Minnesota, USA. The dataset includes bacterial, archaeal, fungal gene copy numbers, along with corresponding logarithmic values, at 11 depth increments of two-meter deep peat cores taken from 12 sampling sites locations inside SPRUCE plots (10 chambered and 2 ambient plots). The sampling, sample prep and analysis followed standard methods outlined in prior publications (Wilson et al. 2016; Kluber et al. 2020) except that a higher yielding Omega Bio-Tek Mag-Bind Environmental DNA 96 Kit was used for extractions and DNA was quantified using Qubit dsDNA High Sensitivity Assay Kit. qPCR subsamples of peat cores from the SPRUCE plots characterize changes in the abundance and composition of microbial communities of peat seasonally showing how composition varies under multiple levels of experimental peat warming and atmospheric CO2 concentrations. This dataset contains one data file in comma-separated values (.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma-separated values (.csv) format and a user guide in PDF (*.pdf) format. On 2026-07-07 this dataset was updated to add three columns to the data file: ‘Fungal_copy_dry’, ‘Log_fungal_copy_dry’, ‘Fungal_copy_wet’. No previously released data values were altered. Additionally, the abstract, data dictionary, and user guide were updated, and a file-level metadata file was added.

Archaea↗

NIST: Soil Respiration, Moisture, Temperature, Chemistry; and Fine Root Measurements from a Transect Through a Forest Edge, Gaithersburg, Maryland, 2017-2021

This dataset contains soil respiration, moisture, temperature, and chemistry, as well as fine root measurements from the National Institute of Standards and Technology (NIST) Forested Optical Reference for Evaluating Sensor Technology (FOREST) research facility at Gaithersburg, Maryland. Measurements were taken at an existing transect array that begins in a grassy meadow, crosses a sharp forest edge, then a small stream, and finally extends upwards in the interior of the forest at the top of a ridge. There are 6 different landscape positions replicated across three transects in the array. Soil respiration was measured during growing seasons in 2017-2019 (2017-06-02 to 2020-02-27). Pedons (1 m3) were isolated from surrounding tree roots using trenching and a fabric to inhibit root ingrowth. Flux measurements inside the pedons were thus assumed to represent heterotrophic only respiration in 2019, and these fluxes were paired with nearby fluxes assumed to represent total respiration. Deep vertical probes measured volumetric moisture content and temperature at the same points in the array every 10 cm in depth to either 90 cm or 120 cm total depth, at 15 minute intervals, from 2019-2021 (2019-07-09 to 2021-09-10). Soil core samples were collected from each of the array points for three different months in early- to mid-2019 (2019-03-19 to 2019-07-10), at three depths each. Soils were analyzed for gravimetric moisture content; pH; total carbon, nitrogen, and phosphorus; texture; microbial biomass carbon, nitrogen, and phosphorus; extractable dissolved organic carbon, nitrogen, and phosphorus; extractable nitrate and ammonia; and extracellular hydrolytic enzyme activities. The fine roots were separated from the cores and segregated by plant functional type (grass or tree species) and if they were dead or alive. Fine roots were then measured for length, surface area, diameter, and dry mass. This dataset contains four data files in comma separated (*.csv) format. These data serve to deepen our understanding of root and soil processes at forest edges and in transitional zones.

54 ENVIRONMENTAL SCIENCES↗

Organic matter concentration and composition of experimentally burned open air and muffle furnace vegetation chars across differing burn severity and feedstock types from Pacific Northwest, USA (v4).

This dataset represents results from an experimental study designed to compare how the chemical composition of organic matter changes across different burn conditions and feedstock materials. The dataset provides both solid and dissolved phase bulk concentration and organic matter characterization data from experimentally generated chars. Chars were created in a closed muffle furnace or on an open burn table from four different feedstock species representing vegetation commonly impacted by fire regimes across the Pacific Northwest, USA. This data can be used to compare how different burn conditions may influence resultant organic matter chemistry and help further our understanding of potential biogeochemical impacts on river corridors post-fire. This dataset is comprised of one data package readme, one data dictionary (dd), one file level metadata (flmd), fourteen burn table videos, burn table video metadata and three folders containing (A) data; (B) metadata and protocols; and (C) photos. The folder names and the file name of the data package readme include a version number which will be updated with future iterations of this data package. The data folder includes (1) solid carbon and solid nitrogen; (2) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) and total dissolved nitrogen (TN); (3) pH; (4) thermocouple time series temperature; (5) methods codes; (6) installation methods; (7) excitation emissions matrix (EEM) methods information; (8) a folder of excitation emissions matrix (EEM) fluorescence and absorbance spectra in dissolved organic matter and EEMs processing instructions; (9) solid state carbon-13 and solution state phosphorus nuclear magnetic resonance (13-C NMR and 31-P NMR) data and methods; (10) benzene polycarboxylic acid (BPCA) concentration and stable isotope data; (11) FTICR-MS methods; (12) Inductively coupled plasma (ICP) data for total calcium, magnesium, iron, aluminum, potassium, phosphorus, sodium, and sulfur along with sodium hydroxide-ethylenediaminetetraacetic acid (sodium hydroxide-EDTA) extractable calcium, magnesium, iron, aluminum, potassium, phosphorus, and sulfur; (13) a folder of phosphorus, carbon, and nitrogen X-ray absorption near edge structure (P-XANES, N-XANES, C-XANES) data for samples and standards; (14) P-XANES, N-XANES, C-XANES methods; (15) molybdate reactive phosphorus; and (16) folder of high resolution characterization of organic matter via 21 Tesla Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) generated through the Environmental Molecular Sciences Laboratory (EMSL; https://www.pnnl.gov/environmental-molecular-sciences-laboratory). The FTICR folder contains .txt data files and a subfolder containing instruments for using Formularity (https://omics.pnl.gov/software/formularity) and an R script to process the data based on the user's specific needs. The metadata and protocols folder includes (1) international geo-sample number (IGSN) mapping file (2) burn and laboratory metadata; (3) burn protocol; (4) laboratory protocol; (5) vegetation collection metadata; and (6) vegetation collection protocol. The folder contains photos of the solid chars. All files are .csv, .txt, .pdf, .jpg, .jpeg, .R, .ref, or .mp4. The data package was originally published October 2022 (v1). It was updated April 2023 (v2; new data files), September 2023 (v3; new and corrected data files), and September 2024 (v4; new and added/updated files). Metadata files were also updated to reflect these changes. See the change history section in the readme for more details.

54 ENVIRONMENTAL SCIENCES↗

Compilation of Experimental Yield Data for Spontaneous Fission of 252 Cf

We present a comprehensive compilation and curation of experimental fission yield (FY) data for the spontaneous fission of 252 Cf, extracted from the EXFOR database. The compilation follows a structured methodology developed for prior compilations of neutron-induced fission yields, and incorporates both independent (IFY) and cumulative (CFY) yields. A total of 62 datasets were reviewed, with entries spanning from 1955 to 2021. A significant portion of the literature reports pre-neutron emission yields, which were excluded from the present compilation due to limitations in format compatibility. Each accepted dataset was processed into a standardized JSON format, including metadata, uncertainties, and bibliographic references. Where available, decay radiation information was used to update the FY data using the latest ENSDF evaluations; 237 data points were corrected accordingly. These corrections are fully traceable and preserve original values. The result is a curated dataset suitable for use in nuclear data evaluations. This work is part of an ongoing effort to modernize the handling of FY data and provide evaluators with high-quality, machine-readable experimental inputs

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Mapping use cases and dataset needs for benchmarking buildings data

A perennial challenge in buildings research is the lack of high-quality datasets that can be relied upon for a wide array of tasks, including model calibration and improving energy efficiency and load flexibility. Instrumenting a building for data collection is resource intensive, so it is important to be methodical in the approach and ensure that resulting data are flexible and useful for a broad range of analyses. This study aims to fill the gaps in characterizing potential use cases for buildings datasets and mapping them to dataset needs using a well-defined data infrastructure. Here, we have developed a systematic mapping strategy between buildings dataset needs and use cases to help streamline the processes of efficiently targeting datasets, designing building sensing systems, and determining buildings research use cases. We selected 14 prospective use cases and 11 refined buildings data categories for developing the preliminary dataset-needs-to-use-cases mapping matrix (‘DN-UC mapping matrix’) with generic ‘Tags’—a detailed sub-level of data categories extracted by justifying the needs of an aspect of the datasets to use cases. We present two example applications of the developed mapping matrix to demonstrate use of the mapping matrix and its effectiveness.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Artificial intelligence models, photos, and data associated with the manuscript “Quantifying Streambed Grain Size, Uncertainty, and Hydrobiogeochemical Parameters Using Machine Learning Model YOLO” (v2)

This data package is associated with the manuscript “Quantifying Streambed Grain Size, Uncertainty, and Hydrobiogeochemical Parameters Using Machine Learning Model YOLO” published in Water Resources Research (Chen et al., 2024). This data package includes the training, validation, testing, and prediction data used by the artificial intelligence (AI) model for automated grain size and hydro-biogeochemistry quantification using streambed photos. The grain size data are extracted for each photo using You Look Only Once (YOLO), a pre-trained object detection model. This data package was originally published in October 2023. It was updated August 2025 (v2; new and modified files). File and folder names were not revised to indicate changes. See the change history section in the readme for more details. Please see flmd.csv for a list of all files contained in this data package and descriptions for each. Please see dd.csv for a data dictionary that defines the column headers of .csv files in the data package. This dataset is comprised of one data folder containing (1) file-level metadata; (2) data dictionary; (3) readme; and (4) six subfolders. Subfolders 1 to 4 include the training, validation, testing, and prediction data. Subfolder 5_Summary includes the summary results of different combinations of training, validation, testing, and prediction data. Subfolder 6_SupplementalData includes additional data downloaded from public sources (Kaufman et al., 2023a; Kaufman et al., 2023b; Garefalakis et al., 2023; Mair et al., 2024; https://github.com/river-corridors-sfa/Geospatial_variables). In total, the data package includes 110 folders and 44,283 files. These files include 9,047 .jpg photos, 1 .png photo, 3 .tif photos; 26,639 photo labels and individual grain sizes and probability from AI (.txt); 8,447 grain size distribution data (.dat); and 126 CSV files for results summary, and 14 required metadata files (.xlsx). The summary CSV files contain 68 columns and approximately 2,200 rows that represent photo names, site locations, recording time, GPS coordinates, grains sizes (D10, D50, D60, and D84), number of grains, and additional hydro-biogeochemical data such as water depth, flow velocity, Manning’s coefficient, friction factor, hydraulic conductivity, permeability, streambed interstitial velocity magnitude, mass transfer rate, and nitrate uptake velocity. The photos were obtained from 75 sites in the Yakima River Basin and the Columbia River shorelines, and other associated data from samples and sensors obtained when the photos were taken are publicly available (Fulton et al. 2022; Grieger et al. 2023). All files are .csv, .txt, .dat, .jpg, or .pdf. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Comprehensive Database of Environmental Mitigations Extracted from FERC-Licensed Hydropower Projects Using Artificial Intelligence Techniques, 1998-2023

This dataset provides a comprehensive inventory of environmental mitigation measures required by Federal Energy Regulatory Commission (FERC) licensed hydropower facilities from 461 licenses that were issued from 1998 to 2023. These licenses constitute 446 of the 1015 FERC projects that were active at the end of 2023. 17,612 mentions of environmental mitigations were identified and categorized in 128 unique categories. Mitigations were identified using a Natural Language Processing (NLP) approach, specifically with a Bidirectional Encoder Representations from Transformer (BERT) model. Model-derived results were then reviewed and updated by a subject matter expert as needed. This dataset introduces important enhancements to previous efforts to inventory environmental mitigations, such as including associated license text for each mitigation, tracking the number of instances a mitigation was identified within a license, and providing improved location information. These enhancements significantly expand the dataset's utility, offering greater analytical capabilities and ensuring reproducibility. The dataset is downloadable as a zip file containing the metadata and dataset files.

Ruggles, Thomas [Oak Ridge National Laboratory (OR↗

Fe(III) reducing bacterial activities in Old Woman Creek wetland sediments, June 2023

To evaluate the Fe(III) reducing microbiological activities in Old Woman Creek Nature Preserve (OWC) wetland sediments, we incubated OWC sediments under anoxic and oxic conditions and with or without Fe(III) amendment [as hydrous ferric oxide (HFO)]. No Fe(III) reduction was observed in heat-deactivated incubations. In non-sterile anoxic incubations, measurement of 0.5 M HCl-extractable Fe(II) indicated that Fe(III) reduction occurred in both Fe(III)-amended and -unamended incubations, indicating that abundant Fe(III) is associated with the OWC sediments. Little Fe(II) accumulated in solution, indicating that the most biogenic Fe(II) adsorbs to the sediments. When air was added to the headspace of non-sterile incubations, Fe(III) reduction was halted and any biogenic Fe(II) that accumulated was oxidized. These experiments were used to guide preparation and analyses of incubations to determine if electrochemical measuements can be used to detect microbiological activities in contrasting terminal electron accepting regimes (i.e., aerobic and Fe(III) reducing conditions). Data package includes methods and data from experiments, including dissolved anion concentrations, dissolved Fe(II) concentrations, and 0.5 M HCl-extractable Fe(II) concentrations. All files are either .txt or .csv and can be opened by any plain text editor application.

EARTH SCIENCE↗

The catalog-to-cosmology framework for weak lensing and galaxy clustering for LSST

We present TXPipe, a modular, automated and reproducible pipeline for ingesting catalog data and performing all the calculations required to obtain quality-assured two-point measurements of lensing and clustering, and their covariances, with the metadata necessary for parameter estimation. The pipeline is developed within the Rubin Observatory Legacy Survey of Space and Time (LSST) Dark Energy Science Collaboration (DESC), and designed for cosmology analyses using LSST data. In this paper, we present the pipeline for the so-called 3x2pt analysis -- a combination of three two-point functions that measure the auto- and cross-correlation between galaxy density and shapes. We perform the analysis both in real and harmonic space using TXPipe and other LSST-DESC tools. We validate the pipeline using Gaussian simulations and show that it accurately measures data vectors and recovers the input cosmology to the accuracy level required for the first year of LSST data under this simplified scenario. We also apply the pipeline to a realistic mock galaxy sample extracted from the CosmoDC2 simulation suite (Korytov et al. 2019). TXPipe establishes a baseline framework that can be built upon as the LSST survey proceeds. Furthermore, the pipeline is designed to be easily extended to science probes beyond the 3x2pt analysis.

79 ASTRONOMY AND ASTROPHYSICS↗

Metagenome-assembled genomes from topsoils collected during NEON campaign in East River, CO (06/14/2018-06/28/2018)

The Watershed Function Science Focus Area (WF SFA) at Lawrence Berkeley National Lab is working to build a mechanistic understanding of the distribution and dynamics of biogeochemical processes in mountainous watersheds and their response to perturbation. In June 2018, the NEON (National Ecological Observatory Network) Airborne Observatory Platform (AOP) performed a taskable airborne imaging campaign to collect visible to shortwave infrared (VSWIR) imaging spectroscopy and LiDAR data across 330 km2 in the Upper East River at Crested Butte, CO. We conducted a parallel ground sampling campaign to sample vegetation traits, as well as soil physical, chemical, and microbiological characteristics. We collected these samples from 438 sites across 12 locations spanning much of the elevation, topographic, and geologic variability across the study area. A subset of 250 samples were used for soil metagenomics which is presented here. In addition, at each site, vegetation samples were collected to measure species-specific leaf water content and leaf mass area, foliar elemental composition and foliar CN stable isotope ratios. Soil samples were collected to measure soil physical properties which include bulk density and soil texture analysis. A suite of soil chemical properties was measured from the samples collected at each site, including pH, organic matter, concentrations exchangeable cations, total elemental composition, and the concentrations of extractable N pools (e.g. total free amino acids, ammonium, nitrate, dissolved organic N, and total dissolved N). Additionally, we have measured soil microbial biomass CN stoichiometry. Here, we present 1982 metagenome-assembled genomes (MAGs) for the bacterial and archaeal community from topsoil collected from during NEON 2018 campaign. All metagenomes were sequenced at JGI (Joint Genome Institute) (GOLD Study ID: Gs0149986). Metagenomes were assembled using JGI Metagenome Workflow (10.1128/mSystems.00804-20). The dataset includes (1) zip files for 1982 MAG fasta files (neon_genomes1-5.tar.gz, split into 5 tarballs to keep tarballs under 0.5 GB), (2) neon_Gs0149986_samples_soilproperties_metagenomes.csv: the sample information together with the accession numbers for the underlying metagenomes and the associated soil physical and chemical measurements in NMDC (National Microbiome Data Collaborative) compliant format, (3) neon_Gs0149986.kml: location bounding box file for the sampled locations, (4) samples.csv: sample metadata file used to register Internationall Generic Sample Numbers (IGSNs), (5) flmd.csv: file level metadata file, and (6) dd.csv: data dictionary file. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

SPRUCE FT-ICR MS, Bulk Chemistry, and Mass Loss from Litter Decomposition Study in Experimental Plots, Marcell Experimental Forest, Minnesota, 2015-2017

This dataset contains molecular, bulk chemical, and mass loss measurements from a litter decomposition study at the Spruce and Peatland Responses Under Changing Environments (SPRUCE) experimental site within the Marcell Experimental Forest in northern Minnesota, USA. This site is in a Sphagnum spp. ombrotrophic bog forest. Litterbags were deployed into the peat in September 2015 across three warming levels (+0, +4.5, and +9°C) under ambient and elevated carbon dioxide (CO₂ - +500 ppm) and retrieved after roughly 0.5, 1, and 2 years of field incubation (2015-09-23 to 2017-08-02). Litterbags containing six peatland litter types: black spruce needles (Picea mariana - SPL), spruce fine roots (SPR), Sphagnum angustifolium (ANG), Sphagnum magellanicum (MAG), Labrador tea leaves (Rhododendron groenlandicum - LTL), and Labrador tea roots (LTR). Molecular composition of water-soluble organic matter extracts was characterized using Fourier Transform Ion Cyclotron Resonance Mass Spectrometry (FT-ICR MS) at 9.4 Tesla, operated in negative ion mode with electrospray ionization, providing molecular formula assignments and compound-class distributions across the decomposition time series. Bulk chemical characterization included elemental analysis (percent carbon, nitrogen, and phosphorus) and Fourier Transform Infrared Spectroscopy (FTIR) to quantify functional group composition. Litter mass loss was tracked gravimetrically at each retrieval interval, expressed as percent mass remaining relative to initial dry mass for each litter type and treatment combination. These data are valuable for understanding how vegetation shifts driven by increased atmospheric CO2 and temperature in peatlands alter litter inputs and organic matter stabilization trajectories, with implications for projecting and modeling peatland carbon cycling. This dataset contains two data files in comma-separated value (.csv) format. Additional metadata are provided: two data dictionaries and a file-level metadata file in comma separate (.csv) format and a user guide in PDF (*.pdf) format.

decomposition↗

CHESS 2025: Waveform LiDAR data from NEON AOP surveys

This dataset provides Level 1 (L1) full-waveform light detection and ranging (LiDAR) data collected for the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS). These data were acquired to enable characterization of vegetation structure and other three-dimensional features of the land surface, and to evaluate structural changes that may have occurred between a prior LiDAR acquisition in 2018 and the 2025 overflight. Waveform LiDAR data can provide more detailed information about objects on the ground than discrete point clouds typically do, and they are often used for granular target segmentation and characterization of subcanopy vegetation. The data were acquired over three study domains in the Upper Gunnison river basin: the upper East River watershed (CRBU); Almont Triangle and Taylor Canyon (ALMO); and Upper Taylor River watershed (UPTA) between 2025-06-13 and 2025-07-15. LiDAR data were acquired using the Optech Galaxy Prime Airborne LiDAR Terrain Mapper onboard the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP). These are the primary waveform LiDAR data delivered by NEON and are provided per flightline in compressed Pulsewaves format, an open-source binary file standard. A Pulsewaves object comprises a two files: a pulse (.pls) file, which stores the geographic origin, outgoing vector, and metadata for every laser pulse emitted by the scanner, and a wave file (.wvs), which stores the sequential amplitude samples of the outgoing pulse and the returning signals. The files are published here in their compressed forms (.plz, .wvz). All waveform data were processed following the theoretical workflow described in the NEON L0-to-L1 Waveform LiDAR Algorithm Theoretical Basis Document (Krause and Goulden 2022a); however, the Pulsewaves output format differs from a legacy format described in that document. Waveform amplitude samples are recorded at 1 nanosecond intervals. All coordinates are provided in meters. Horizontal coordinates are referenced in Universal Transverse Mercator (UTM) zone 13N and the World Geodetic System (WGS) 1984 ensemble datum. Elevations are referenced to Geoid12A. Waveform data for the UPTA survey area were collected without incident and the published records are complete. However, both the ALMO and CRBU collections experienced issues that resulted in incomplete data for those areas. On collection day 2018-06-16 a hardware failure caused the waveform digitizer to lose data from the eastern edge of the ALMO site (Figure 22). The waveform data for flightlines 2–20 could not be extracted from the digitizer, and the data proved unrecoverable. As a result, a portion of the site does not have coverage with waveform data. Although no hardware failure was observed during collection over the CRBU area, final waveform files generated by vendor software contained only ~25% of the expected number of return pulses. After discovery, NEON initiated troubleshooting with the vendor. The root cause of the data ablation had not been identified at the time of publication. Additional data will be published in an update to this package if further recovery proves successful. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

A High-Throughput Computing Infrastructure to Generate Custom, Open Community Geothermal Datasets

The most significant challenge facing geothermal research, development, and deployment is a lack of comprehensive datasets describing the geological and economical properties of North America. Automated knowledge base construction, the process of designing algorithms to analyze text and images to programmatically build new datasets, is one possible solution to this problem. The xDD library of full-text scientific articles (https://xdd.wisc.edu) is one of the largest collections of open and controlled-access scientific documents available for knowledge base construction in the world, but it has been underutilized by experts in geothermal research. The xDD development team attributed the lack of engagement by software developers and geothermal researchers to two perceived shortcomings of the system. First, the workflow for obtaining data from xDD for local development and testing of data mining applications was unnecessarily abstruse and required significant manual intervention by xDD systems administrators. Second, although xDD already held articles from a broad cross-section of scientific literature with an emphasis on the geosciences, it did not have an explicit set of geothermal research documents that could serve as the nucleus of a geothermal data mining application. To address these issues, the Automated Data Extraction PlaTform (ADEPT) was proposed to extend the data distribution capabilities of the xDD system. The ADEPT extension added the following four key features to xDD: 1) integration of National Geothermal Data System (NGDS) documents into the xDD library to provide an explicitly geothermally-themed collection; 2) improved RESTful (i.e., https-protocol driven) web services for external partners to access xDD data for machine learning application development; 3) a web platform for end-users and xDD administrators to coordinate the development of data mining applications from the initial step of browsing available documents to the final stage of deploying a production-quality machine learning application on high-throughput computing infrastructure; and 4) the development of demonstration data mining applications to illustrate the new workflow to potential collaborators. A total of 21,674 geothermal documents from NGDS were fully ingested into the xDD library and the associated metadata is publicly available through the xDD web services; furthermore, the ADEPT web platform is now publicly accessible and fully live at https://xdd.wisc.edu/adept/.

15 GEOTHERMAL ENERGY↗

Technical Report on Subsurface Monitoring of the Brady Hot Spring Geothermal Site, Nevada, based upon Full Waveform Inversion

Abilities to accurately characterize the subsurface in a geothermal setting is key to assess and support production. An important element of geothermal reservoir monitoring is also the ability to investigate fluid transport within fracture network. This report focuses on improving subsurface imaging and monitoring in geothermal settings using full waveform inversion based on the adjoint method and time-lapse imaging. To assess our method, we rely on a dense seismic dataset collected in 2016 at the Brady Hot Springs geothermal site in Nevada for the DOE-funded project Poroelastic Tomography by Adjoint Inverse Modeling of Data from Seismology, Geodesy, and Hydrology. This dataset captures subsurface changes across four stages of geothermal power plant operations, which involve varying rates of fluid injection and extraction. Two velocity models were previously derived from this dataset using different methods: one based on travel times and another on sweep interferometry. Our first step is to refine these models using adjoint tomography, which has been applied successfully at global and regional-scales but is less common at the reservoir-scale. Two approaches are then explored for time-lapse analysis: directly comparing refined tomographic models from different stages or backpropagating waveform differences relative to a baseline tomographic model. The main take away is that both approaches highlight similar reservoir behaviors, but the latter approach is more computationally effective in capturing small-scale changes in subsurface properties. For this work, we leverage the use of Salvus (www.mondaic.com), an end-to-end seismic imaging solution, relying on the spectral element method to compute forward and adjoint simulations, and developed by Mondaic Ltd. It includes integrated workflow management that handles waveform and metadata, launches simulations, computes waveform misfits and adjoint sources, and iterates for model updates by nonlinear optimization.

15 GEOTHERMAL ENERGY↗