Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “metadata quality”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Data for Kim et al., "Variations in the optical and molecular composition of dissolved organic matter exported from coastal wetlands"

Knowledge about sources and composition of marsh-derived dissolved organic matter (DOM) is critical for understanding the role of marshes in coastal biogeochemical cycling and the fate of marsh-derived DOM in the ocean. To investigate tidal variability in composition of marsh-derived DOM, Kim et al. examined the optical and molecular characteristics of hourly surface water samples at three tidal creeks in the Chesapeake Bay. Groundwater samples along the terrestrial landscape gradient as well as estuarine water from the adjacent estuary at each site were also collected to help resolve sources of surface water DOM. Samples were collected in summer 2024 at three sites – SWH: Sweet Hall Marsh, GCW: Kirkpatrick Marsh, and GWI: Goodwin Islands – which are part of synoptic sites in the Chesapeake Bay region of the COMPASS-FME (Coastal Observations, Mechanisms, and Predictions Across Systems and Scales - Field, Measurements, and Experiments) project. Surface water samples were collected hourly over a 48-hour period at each site. Groundwater and estuarine water samples were collected once. This dataset includes- Surface water depth and salinity- Dissolved organic carbon (DOC) and total dissolved nitrogen (TDN) concentrations- Optical indices and relative composition of parallel factor analysis (PARAFAC) components- High resolution mass spectrometry data.

54 ENVIRONMENTAL SCIENCES↗

Introduction

This report provides a comprehensive overview of metadata to describe sensor signals in wastewater treatment plants and methods to obtain such metadata. In this introduction, we explain the original motivation behind the MetaCO task group. This includes a description of historical challenges (data volume, data velocity) for which mature technology is now available, and newer challenges, which relate to data structure (data variety) and data quality (veracity). We conclude the chapter with an expression of gratitude to all involved.

Aguado, Daniel↗

DAISY: A Rapid Approach to Evaluating Marine Energy Converter Sound (Final Technical Report)

This project’s objective was to improve the quality of acoustic information about marine energy converters that could be collected from groups of drifting hydrophones, while reducing the costs of deployment and data analysis. This was achieved through technology development addressing four focus areas: (1) minimizing flow-noise and self-noise, (2) integrating metadata streams into a single data acquisition system, (3) developing post-processing routines to facilitate rapid data review, and (4) enabling objective identification of marine energy converter sound against a backdrop of ambient noise using time-delay-of-arrival localization.

16 TIDAL AND WAVE POWER↗

FAIR Data Meets FAIR Software

Modern scientific research is increasingly defined by the interplay between data, software, and the workflows that connect them. Yet while the FAIR (Findable, Accessible, Interoperable, Reusable) principles have become foundational for scientific data stewardship, the same level of structure and expectation has only recently begun to extend to research software. This talk covers why and how FAIR principles are being applied to data and software to support data reuse. It outlines the gaps in current sharing norms, the growing federal emphasis on persistent identifiers and public access, and the opportunities created when datasets, computational workflows, code, and models are linked through rich, standardized metadata. Practical implementation pathways for the EIC and JLab communities are described, including datacards for structured dataset documentation and provenance-aware workflows. By aligning data lifecycle management with FAIR-aligned software practices, the scientific community can advance toward autonomous knowledge graphs, generative workflows, and high-quality, AI-ready scientific datasets.

McSpadden, Diana [Thomas Jefferson National Accele↗

Groundwater table elevation and temperature from 2015 to 2024 at the Lower Montane site in the East River Watershed, Colorado.

This groundwater level elevation and temperature data package is aimed at improving the predictive understanding of hydro-biogeochemical processes at the lower montane site in the East River Watershed, Colorado. The dataset is obtained using pressure transducers placed in shallow wells in the floodplain. This dataset contains data from wells with Location ID's ER-DOW (alias DO1West), ER-DOE (alias DO2East), ER-MBA1 (alias M1Bend1), ER-MBA2 (alias M1Bend2), ER-UPW (alias UP1West), ER-UPM (alias UP2), ER-UPE (alias UP3East). Another dataset contains the data from wells with Location ID's ER-CPA1 to ER-CPA6. Each file contains the water level elevation and the water temperature. Water level elevation has been obtained using the barometric pressure from the pressure transducer (Hobos sensor) in the well, barometric pressure from a sensor in air located at the same site (lower montane), depth from top-of-casing (TOC) to sensor measurement point, and TOC elevation. Data have been checked with a few measurements of water table depths. A real-time kinematic (RTK) global positioning system (GPS) has been used to survey the TOC (data in file Well_Location.csv). The water level elevation is given in UTM13N Geoid2012AB. While depth to water level is not present in the data files, it can be easily calculated with the TOC and distance to ground provided in the GPS coordinate file. The dataset quality is discussed in Collection/Analysis section of the methods. Time-series of measurements were initially added to the archive for the period 2015 to 2019, and later updated with time-series until 2024 (end of data collection). The dataset contains 8 *.csv data files, and 3 *.csv metadata files. Feel free to contact the author with any questions or collaboration interests. The publication year was updated from "2020" to "2025" to reflect the revised version of this dataset.

54 ENVIRONMENTAL SCIENCES↗

Groundwater table elevation and temperature from 2015 to 2024 across Meander C at the Lower Montane site in the East River Watershed, Colorado.

This groundwater level elevation and temperature data package is aimed at improving the predictive understanding of hydro-biogeochemical processes at the lower montane site in the East River Watershed, Colorado. The dataset is obtained using pressure transducers placed in shallow wells in the floodplain. This dataset contains data from wells ER-CPA1 to ER-CPA6. Another dataset contains the data from wells at nearby Locations. Each file contains the water level elevation and the water temperature. Water level elevation has been obtained using the barometric pressure from the pressure transducer (Hobos sensor) in the well, barometric pressure from a sensor in air located at the same site (lower montane), depth from top-of-casing (TOC) to sensor measurement point, and TOC elevation. Data have been checked with a few measurements of water table depths. A real-time kinematic (RTK) global positioning system (GPS) has been used to survey the TOC (data in file Well_Location.csv). The water level elevation is given in UTM13N Geoid2012AB. While depth to water level is not present in the data files, it can be easily calculated with the TOC and distance to ground provided in the GPS coordinate file. The dataset quality is discussed in Collection/Analysis section of the methods. Time-series of measurements were initially added to the archive for the period 2015 to 2019, and later updated with time-series until 2024 (end of data collection). The dataset contains 7 *.csv data files, and 3 *.csv metadata files. Feel free to contact the author with any questions or collaboration interests. The publication year was updated from "2020" to "2025" to reflect the revised version of this dataset.

54 ENVIRONMENTAL SCIENCES↗

Automating the Analysis of Large Language Models Responses through Zero-Shot Question Answering

Recent advancements in Large Language Models (LLMs) have shown significant potential in various applications, yet their evaluation, particularly in zero-shot question answering scenarios, remains a challenging task. In this study, our objective was to explore precision metrics for Large Language Models (LLM) and design and implement a software pipeline to automatically evaluate LLMs' outputs under zero-shot question answering. Zero-shot question answering involves a model providing answers to questions about topics it hasn't seen during training. It leverages the principles of zero-shot learning by relying on semantic understanding and generalization from related knowledge. The data used was metadata from medical databases on congenital heart disease. We explored eleven LLM metrics and selected three for our evaluation: BLEU, BERTScore, and MoverScore. BLEU calculates a score based on the overlap of n-grams (contiguous sequences of n items, typically words) between the machine-generated translation and the reference translations. Higher BLEU scores indicate better correspondence between the machine-generated and human-generated translations. BERTScore is a metric used to evaluate the quality of machine-generated text by measuring the similarity of token embeddings produced by BERT (Bidirectional Encoder Representations from Transformers) between the generated text and reference text. MoverScore is a metric that quantifies the dissimilarity between the distributions of word embeddings from machine-generated text and reference text, emphasizing semantic similarity over exact token overlap. We also introduced HBKI, a composite metric summarizing these approaches. We tested five models —GPT-3, Llama-2, Gemini 1.5 Pro, Solar 10.7B, and Mixtral-8x7b. Our software pipeline, designed and implemented using Object-Oriented Programming principles, allows users to customize the selection and extraction of features for topics of interest in their own research. Our results show that MoverScore delivered the most precise evaluation of the LLM's outputs, while Mixtral-8x7b achieved the best overall performance in extracting metadata from the databases.

97 MATHEMATICS AND COMPUTING↗

Lessons Learned from AskGDR: Usage and Impact Analysis of the Geothermal Data Repository's AI Research Assistant: Preprint

In October of 2024, the Department of Energy's (DOE) Geothermal Data Repository (GDR) team officially launched AskGDR, an AI research assistant resulting from the integration of a Large Language Model (LLM) with the metadata and supporting documents associated with GDR datasets. AskGDR allows GDR users to ask deeper questions about the origin of datasets, the methods used to collect them, and the findings they help support. Using Retrieval Augmented Generation (RAG), AskGDR can be used to summarize findings spread across dozens of papers and technical reports or to extract relevant information describing a single data field. However, generative AI is experimental. The National Renewable Energy Laboratory (NREL) has been collecting metrics on AskGDR and documenting lessons learned during its deployment. This paper will outline the efficacy and impact of AskGDR through analysis of its use, operating costs, number and types of questions asked, and the quality of answers provided.

15 GEOTHERMAL ENERGY↗

VirJenDB: a FAIR (meta)data and bioinformatics platform for all viruses

High-throughput sequencing has generated an unprecedented volume of data. However, researcher-submitted data in repositories requires extensive curation and quality control for reuse. These tasks are hindered by the multiplicity of repositories, the sheer volume of the data, and the complexity of virus (meta)data curation. To address these challenges, VirJenDB offers a user-friendly platform to facilitate versioned, community-driven curation, and ontology development. Virus sequences were ingested from 16 sources, including ~200 fields of metadata or standards, covering taxonomy, sample, and host information. Up to 85 metadata fields have undergone at least one round of curation, and are linked to 15.4 million virus sequences, with 88 % from those infecting eukaryotes and the remaining infecting prokaryotes. Subsets were created, including a novel collection of 0.91 million viral operational taxonomic unit (vOTU) sequences across all viruses, while keeping the original sequences from each vOTU to facilitate downstream analyses, e.g. sequence variation. The VirJenDB web portal (https://www.virjendb.org) provides HTTPS and Application Programming Interface (API) access to the sequence datasets and metadata, offering a search engine, filtering, download, visualizations, and documentation. VirJenDB aims to connect the phage and eukaryotic virus research communities by supporting webtool integration, meta-analyses, and metadata schema extensions.

Saghaei, Shahram↗

Meteorological Variables and Energy Fluxes at the Pumphouse Site, Crested Butte, CO 2017-2019

This data contains output from the pumphouse eddy covariance tower that includes shortwave radiation, longwave radiation, net radiation, air temperature, relative humidity, as well as sensible, latent, and ground heat fluxes. Also included is calculated evapotranspiration from the latent heat flux and the latent heat of vaporization. All data are on a daily timestep and displayed in Mountain Time. The data has been processed, and Quality Assurance / Quality Control (QA/QC) was done, but any daily gaps in the data have not been filled in. This research was funded by the Department of Energy and performed as part of the Watershed Function Scientific Focus Area. This research aimed to constrain evapotranspiration in a high-elevation catchment.The dataset includes one comma-separated values (CSV) data file (EddyCovariance_MeteorlogicalVariables_CrestedButtePumphouse.csv). Additionally, three metadata CSV files are included: (1) location metadata file (locations.csv), which contains location metadata and coordinates; (2) a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata; and (3) a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Dated soil C–N–P profiles, water quality, and chamber fluxes across Ohio and Michigan wetlands (2024–2025)

This dataset includes dated soil core chemistry (bulk density, phosphorus, nitrogen and carbon concentrations), water quality, and chamber flux measurements collected from wetlands in the Midwest United States—12 sites in Ohio, one site in Indiana, one site in Michigan—collected in the spring or summer of 2024 or 2025, all in (.csv) format. These data were generated to examine how wetland restoration, management activities, and time since restoration affect biogeochemical processes, carbon sequestration, nutrient accumulation, water quality, and greenhouse gas emissions. Specifically, these data aim to investigate how restored wetlands differ from natural wetlands in terms of carbon, nitrogen, phosphorus dynamics, as well as carbon dioxide and methane fluxes. Also included are surface and porewater quality parameters and chamber flux measurements across these different wetlands. Sampling was conducted at various sites representing a range of restoration stages, from about 4 years post-restoration up to 105 years post-restoration, and also includes a natural wetland used as a reference in Michigan. These data can be used to determine carbon sequestration rates, nutrient cycling, and to enhance our understanding of biogeochemical responses to wetland restoration in temperate ecosystems. This data package contains (1) a csv file (Water_Quality.csv) containing water quality data (dissolved organic carbon, total dissolved nitrogen, and temperature) organized by location; (2) a csv file (Soil_C_N_P_Seq.csv) containing carbon, nitrogen, and phosphorus concentrations at each soil level and time of each soil level, as well as their sequestration rates; (3) a csv file (CH4_CO2_Flux.csv) including methane and carbon dioxide fluxes that were measured with a chamber; (4) a file-level metadata (FLMD.csv) file that lists each file contained in the dataset with associated metadata; (5) a data dictionary (DD.csv) file that contains terms/column headers used throughout the files along with a definition, units, and data type; and (6) a locations metadata file (Location_metadata.csv).

Earth Science > Atmosphere > Atmospheric Chemistry↗

Cation Data for the East River Watershed, Colorado (2014-2025)

This data package contains mean values for cation concentration for water samples taken from the East River Watershed in Colorado. Inductively coupled plasma mass spectrometry (ICP-MS) has been used to measure the concentrations of elements of interest simultaneously for the East River Watershed, Colorado groundwater and surface water samples to inform insights on the biogeochemistry processes within the watershed. The East River is part of the Watershed Function Scientific Focus Area (WFSFA) located in the Upper Colorado River Basin, United States. For samples collected prior to 06-16-2021, the instrumentation, Elan DRC II, PerkinElmer SCIEX, automatically switches among the three models necessary to analyze all 37 elements. These 37 elements include: (1) Lithium (Li), Beryllium (Be), Boron (B), Sodium (Na), Magnesium (Mg), Aluminium (Al), Silicon (Si), Phosphorus (P), Titanium (Ti), Cobalt (Co), Nickel (Ni), Copper (Cu), Zinc (Zn), Germanium (Ge), Arsenic (As), Rubidium (Rb), Strontium (Sr), Zirconium (Zr), Molybdenum (Mo), Silver (Ag), Cadmium (Cd), Tin (Sn), Antimony (Sb), Caesium (Cs), Barium (Ba), Europium (Eu), Lead (Pb), Thorium (Th), Uranium (U) using standard model, argon Ar as reaction gas, (2) Potassium (K), Calcium (Ca), Vanadium (V), Chromium (Cr), Manganese (Mn), Iron (Fe) using dynamic reaction cell (DRC) model, ammonia NH3 as reaction gas, and (3) Phosphorus (P) and Selenium (Se) using DRC model, oxygen O2 as reaction gas. Note for the samples with higher concentrations of chloride (Cl-), asenic (As) concentrations were analysed with DRC model (oxygen O2 as reaction gas) to avoid the interference of chloride. For samples collected on and after 06-16-2021, an advanced Agilent 8900 triple quadrupole inductively coupled plasma mass spectrometry system (Agilent 8900 QQQ ICP-MS, Agilent Technologies) has been used to measure the concentrations of interested 36 elements simultaneously for environmental samples, including (1) Lithium (Li), Beryllium (Be) and Boron (B) using standard no gas mode, (2) Sodium (Na), Magnesium (Mg), Aluminium (Al) Phosphorus (P), Potassium (K), Chromium (Cr), Manganese (Mn), Iron (Fe), Cobalt (Co), Nickel (Ni), Copper (Cu), Zinc (Zn), Germanium (Ge), Arsenic (As), Rubidium (Rb), Strontium (Sr), Zirconium (Zr), Molybdenum (Mo), Silver (Ag), Cadmium (Cd), Tin (Sn), Antimony (Sb), Cesium (Cs), Barium (Ba), Europium (Eu), Lead (Pb), Thorium (Th) and Uranium (U) using standard helium (He) collision mode, (3) Titanium (Ti) and Vanadium (V) using high Energy (HEHe) helium (He) collision mode, and (4) Silicon (Si), Calcium (Ca) and Selenium (Se) using standard H2 reaction mode. All samples were prepared/diluted with 2% (v/v) ultrapure nitric acid in Milli-Q water (18.2 mega ohm-cm), and analyzed under a rigorous quality assurance and quality control (QA/QC) process. This data package contains (1) a zip file (cation_data_2014_2025.zip) containing a total of 5,849 files: 5.848 data files of cation data from across the Lawrence Berkeley National Laboratory (LBNL) Watershed Function Scientific Focus Area (SFA) which is reported in .csv files per location and a locations.csv (1 file) with latitude and longitude for each location; (2) a file-level metadata (v6_20260901_flmd.csv) file that lists each file contained in the dataset with associated metadata; (3) a data dictionary (v6_20260901_dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; (4) PDF and docx files for the detemination of Method Detection Limits (MDLs) for ICP-MS PerkinElmer DRC II instrumentation (Detemination_of_Method_Detection_Limits__MDLs__for_ICP_MS__PerkinElmer_Elan_DRC_II__LBL_Bldg74_Lab214D) for samples before November 2021; (5) PDF and docx files for the determination of MDLs for ICP-MS Agilent 8900 QQQ instrumentation (ICP_MS_Analysis_detection_limits_and_QA_QC_WenmingDong_updated_2026-08-06) for samples November 2021 and onward. Missing values within the anion data files are noted as either "-9999" or "0.0" for not detectable (N.D.) data. There are a total of 113 locations containing cation data. Update on 2021-04-11: Added Detemination of Method Detection Limits (MDLs) for ICP-MS document, which can be accessed as a PDF or with Microsoft Word. Update on 2022-06-10: versioned updates to this dataset was made along with these changes: (1) updated cation data for all locations up to 2021-12-31, (2) removal of units from column headers in datafiles, (3) added row underneath headers to contain units of variables, (4) removed suffix and prefix on two variables (“aqberylliumion_asberyllium” and “aqlithiumion_aslithium”), (5) added -9999 for empty numerical cells, and (6) the addition of the file-level metadata (flmd.csv) and data dictionary (dd.csv) were added to comply with the File-Level Metadata Reporting Format. Update on 2022-09-09: Updates were made to reporting format specific files (file-level metadata and data dictionary) to correct swapped file names, add additional details on metadata descriptions on both files, add a header_row column to enable parsing, and add version number and date to file names (v2_20220909_flmd.csv and v2_20220909_dd.csv). Update on 2023-08-08: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-01-05. The file level metadata and data dictionary files were updated to reflect the additional data added. Update on 2024-03-11: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-10-16. Further, revisions to the data files were made to remove incorrect data points (from 1970 and 2001). The reporting format specific files were updated to reflect the additional data added. Updated versions of the PDF and docx files for determination of MDLs for ICP-MS data were added to this dataset for samples starting in November 2021. Update on 2025-05-15: Updates were made to both the data files and reporting format specific files. New available cation data was added, up until the end of WY2024 (September 30, 2024). International Generic Sample Numbers (IGSNs), when registered, were added to the data files. The reporting format specific files were updated to reflect the additional data added. Update on 2026-09-01: Updates were made to both the data files and reporting format specific files. New available cation data was added, up until the end of WY2025 (September 30, 2025). Updated versions, as of 2026-08-06, of the PDF and docx files for determination of MDLs for ICP-MS data were added to this dataset for samples starting in November 2021.

54 ENVIRONMENTAL SCIENCES↗

Dissolved Inorganic Carbon and Dissolved Organic Carbon Data for the East River Watershed, Colorado (2015-2025)

This data package contains mean values for dissolved organic carbon (DOC) and dissolved inorganic carbon (DIC) for water samples taken from the East River Watershed in Colorado. The East River is part of the Watershed Function Scientific Focus Area (WFSFA) located in the Upper Colorado River Basin, United States. DOC and DIC concentrations in water samples were determined using a TOC-VCPH analyzer (Shimadzu Corporation, Japan). DOC was analyzed as non-purgeable organic carbon (NPOC) by purging HCl-acidified samples with carbon-free air to remove DIC prior to measurement. After the acidified sample has been sparged, it is injected into a combustion tube filled with oxidation catalyst heated to 680 oC. The DOC in samples is combusted to CO2 and measured by a non-dispersive infrared (NDIR) detector. The peak area of the analog signal produced by the NDIR detector is proportional to the DOC concentration of the sample. DIC was determined by acidifying the samples with HCl first, and then purging with carbon-free air to release CO2 for analysis by NDIR detector. Total dissolved nitrogen (TDN) was analyzed using a Shimadzu Total Nitrogen Module (TNM-L) combined with the TOC-L analyzer (Shimadzu Corporation, Japan). TNM-L is a non-specific measurement of total nitrogen (TN). All nitrogen species in samples are combusted to nitrogen monoxide and nitrogen dioxide, then reacted with ozone to form an excited state of nitrogen dioxide. Upon returning to ground state, light energy is emitted. Then, TDN is measured using a chemiluminescence detector. All data reported are the mean values upon minimum of three replicate measurements, with a relative standard deviation < 3%. All samples were analyzed under a rigorous quality assurance and quality control (QA/QC) process. This data package contains (1) a zip file (dic_npoc_data_2014-2025.zip) containing a total of 337 files: 336 data files of DIC and NPOC data from across the Lawrence Berkeley National Laboratory (LBNL) Watershed Function Scientific Focus Area (SFA) which is reported in .csv files per location and a locations.csv (1 file) with latitude and longitude for each location; (2) a file-level metadata (v6_20250901_flmd.csv) file that lists each file contained in the dataset with associated metadata; (3) a data dictionary (v6_20250901_dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; and (4) PDF and docx files for the determiniation of Method Detection Limits (MDLs) for DIC and NPOC data, which has been updated in 2026-08. Missing values within the anion data files are noted as either "-9999" or "0.0" for not detectable (N.D.) data. There are a total of 113 locations containing DIC/NPOC data. Update on 2020-10-07: Updated the data files to remove times from the timestamps, so that only dates remain. The data values have not changed. Update on 2021-04-11: Added Determination of Method Detection Limits (MDLs) for DIC, NPOC and TDN Analyses document, which can be accessed as a PDF or with Microsoft Word.Update on 6/10/2022: versioned updates to this dataset was made along with these changes: (1) updated dissolved inorganic carbon and dissolved organic carbon data for all locations up to 2021-12-31, (2) removal of units from column headers in datafiles, (3) added row underneath headers to contain units of variables, (4) restructure of units to comply with CSV reporting format requirements, (5) added -9999 for empty numerical cells, and (6) the addition of the file-level metadata (flmd.csv) and data dictionary (dd.csv) were added to comply with the File-Level Metadata Reporting Format. Update on 2022-09-09: Updates were made to reporting format specific files (file-level metadata and data dictionary) to correct swapped file names, add additional details on metadata descriptions on both files, add a header_row column to enable parsing, and add version number and date to file names (v2_20220909_flmd.csv and v2_20220909_dd.csv). Update on 2023-08-08: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-01-05. The file level metadata and data dictionary files were updated to reflect the additional data added. Update on 2024-03-11: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-11-21. Further, revisions to the data files were made to remove incorrect data points (from 1970 and 2001). The reporting format specific files were updated to reflect the additional data added. Revised versions of the PDF and docx files for determination of MDLs for DIC and NPOC were added to replace previous versions. Update on 2025-05-15: Updates were made to both the data files and reporting format specific files. New available DIC and NPOC data was added, up until the end of WY2024 (September 30, 2024). International Generic Sample Numbers (IGSNs), when registered, were added to the data files. The reporting format specific files were updated to reflect the additional data added. Update on 2026-09-01: Updates were made to both the data files and reporting format specific files. New available DIC and NPOC data was added, up until the end of WY2025 (September 30, 2025). Updated versions, as of 2026-08-10, of the PDF and docx files for determination of MDLs for DIC and NPOC data were added to this dataset.

54 ENVIRONMENTAL SCIENCES↗

Total Dissolved Nitrogen and Ammonia Data for the East River Watershed, Colorado (2015-2025)

This data package contains mean values for total dissolved nitrogen (TDN) and ammonia concentrations for water samples taken from the East River Watershed in Colorado. The East River is part of the Watershed Function Scientific Focus Area (WFSFA) located in the Upper Colorado River Basin, United States. TDN was analyzed using a Shimadzu Total Nitrogen Module (TNM-1) combined with the TOC-VCSH analyzer (Shimadzu Corporation, Japan). TNM-1 is a non-specific measurement of total nitrogen (TN). All nitrogen species in samples are combusted to nitrogen monoxide and nitrogen dioxide, then reacted with ozone to form an excited state of nitrogen dioxide. Upon returning to ground state, light energy is emitted. Then, TDN is measured using a chemiluminescence detector. Ammonia was determined using a Lachat's QuikChem 8500 Series 2 Flow Injection Analysis System (LACHAT Instruments, QuckChem 8500 series 2, Automated Ion Analyzer, Loveland, Colorado). When ammonia in water samples is heated (60 degrees C) with salicylate and hypochlorite in an alkaline phosphate buffer, an emerald green color is produced which is proportional to the ammonia concentration. The color is intensified by the addition of nitroprusside. Ethylenediaminetetraacetic acid (EDTA) is added to the buffer to prevent the interference of metal ions (Ca, Mg, and Fe etc.). Ammonia-N is then determined by LACHAT flow injection and a colorimetric assay at an absorbance wavelength 660 nm. (Reference: LACHAT Instruments: QuickChem Method 90-107-06-3-A, Determination of Ammonia by Flow Injection Analysis (High Throughput, Salicylate Method/DCIC) (Multi Matrix method). Written by Lynn Egan (Application group), February 08, 2011.) All files are labeled by location and variable, and data reported are the mean values upon replicate measurements. All samples were analyzed under a rigorous quality assurance and quality control (QA/QC) process as detailed in the methods. This data package contains (1) a zip file (tdn_ammonia_data_2015-2025.zip) containing a total of 299 files: 298 data files of ammonia and TDN data from across the Lawrence Berkeley National Laboratory (LBNL) Watershed Function Scientific Focus Area (SFA) which is reported in .csv files per location and a locations.csv (1 file) with latitude and longitude for each location; (2) a file-level metadata (v7_20260901_flmd.csv) file that lists each file contained in the dataset with associated metadata; (3) a data dictionary (v7_20260901_dd.csv) file that contains terms/column_headers used throughout the files along with a definition, units, and data type; (4) PDF and docx files for the determination of Method Detection Limits (MDLs) for TDN data, which has been updated in 2026-08; and (5) PDF and docx files for the detemination of Method Detection Limits (MDLs) for Ammonia and the Interferences by LACHAT Flow Injection Analysis. Missing values within the anion data files are noted as either "-9999" or "0.0" for not detectable (N.D.) data. There are a total of 105 locations containing TDN and Ammonia-N data. Update 2020-10-07: Updated the data files to remove times from the timestamps, so that only dates remain. The data values have not changed. Update 2021-04-11: Added Determination of Method Detection Limits (MDLs) for DIC, NPOC and TDN Analyses and Determination of Method Detection Limit for Ammonia and the Interferences by LACHAT Flow Injection Analysis documents, which can be accessed as PDFs or with Microsoft Word.Update on 6/10/2022: versioned updates to this dataset was made along with these changes: (1) updated total dissolved nitrogen and ammonia data for all locations up to 2021-12-31, (2) removal of units from column headers in datafiles, (3) added row underneath headers to contain units of variables, (4) restructure of units to comply with CSV reporting format requirements, (5) added -9999 for empty numerical cells, and (6) the addition of the file-level metadata (flmd.csv) and data dictionary (dd.csv) were added to comply with the File-Level Metadata Reporting Format. Update on 2022-09-09: Updates were made to reporting format specific files (file-level metadata and data dictionary) to correct swapped file names, add additional details on metadata descriptions on both files, add a header_row column to enable parsing, and add version number and date to file names (v2_20220909_flmd.csv and v2_20220909_dd.csv). Update on 2022-12-20: Updates were made to both the data files and reporting format specific files. Units were listed incorrectly, but have been fixed to reflect correct units (ug/L). File level metadata (flmd) and data dictionary (dd) files were updated to reflect the updated versions of these files. Available data was added up until 2022-06-01. Update on 2023-08-08: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-01-05. The file level metadata and data dictionary files were updated to reflect the additional data added. Update on 2024-03-11: Updates were made to both the data files and reporting format specific files. New available anion data was added, up until 2023-10-27. Further, revisions to the data files were made to remove incorrect data points (from 1970 and 2001). The reporting format specific files were updated to reflect the additional data added. Revised versions of the PDF and docx files for determination of MDLs for TDN were added to replace previous versions. Update on 2025-05-15: Updates were made to both the data files and reporting format specific files. New available TDN and Ammonia-N data was added, up until the end of WY2024 (September 30, 2024). International Generic Sample Numbers (IGSNs), when registered, were added to the data files. The reporting format specific files were updated to reflect the additional data added. Update on 2026-09-01: Updates were made to both the data files and reporting format specific files. New available TDN and Ammonia-N data was added, up until the end of WY2025 (September 30, 2025). Updated versions, as of 2026-08-10, of the PDF and docx files for determination of MDLs for TDN data were added to this dataset.

54 ENVIRONMENTAL SCIENCES↗

Compilation of Experimental Yield Data for Spontaneous Fission of 252 Cf

We present a comprehensive compilation and curation of experimental fission yield (FY) data for the spontaneous fission of 252 Cf, extracted from the EXFOR database. The compilation follows a structured methodology developed for prior compilations of neutron-induced fission yields, and incorporates both independent (IFY) and cumulative (CFY) yields. A total of 62 datasets were reviewed, with entries spanning from 1955 to 2021. A significant portion of the literature reports pre-neutron emission yields, which were excluded from the present compilation due to limitations in format compatibility. Each accepted dataset was processed into a standardized JSON format, including metadata, uncertainties, and bibliographic references. Where available, decay radiation information was used to update the FY data using the latest ENSDF evaluations; 237 data points were corrected accordingly. These corrections are fully traceable and preserve original values. The result is a curated dataset suitable for use in nuclear data evaluations. This work is part of an ongoing effort to modernize the handling of FY data and provide evaluators with high-quality, machine-readable experimental inputs

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A Prototype Software to Demonstrate a Data Catalog for Hanford Environmental Datasets

Ensuring that data on long-term environmental remediation at the Hanford Site is high-quality, traceable, and easily accessible is an ongoing challenge, complicated by decades of data collection, multiple contractors maintaining data sources, and the wide range of data types. A centralized data catalog, known as the Hanford Environmental Information and Data Index (HEIDI), has been under development as part of the Hanford Environmental Data Management (HEDM) program to address these challenges. HEIDI fulfills a critical need to bring together a wide range of data types and sizes from multiple authoritative data sources, while documenting the data pedigree and quality information (i.e., traceable to the data source/originator). This document describes additional development and maturation of the HEIDI prototype. Key accomplishments included deploying the catalog software, Esri Geoportal Server, on a server accessible to Hanford Local Area Network users, conducting cybersecurity evaluations, investigating integrated authentication solutions, and conducting functional testing of the catalog prototype. The server-based deployment enabled targeted feedback, leading to enhancements including improved accessibility features and an expanded metadata schema. Specifications for the server-based deployment of the prototype catalog and the HEIDI metadata schema are provided in this document to support subsequent HEIDI deployment by the U.S. Department of Energy Richland Operations Office.

54 ENVIRONMENTAL SCIENCES↗

BSEC ecohydrological and water quality fluxes from RHESSys Simulations in USGS gauged watersheds

Baltimore Environmental Social Collaborative (BSEC) Water and Water Quality Simulations from RHESSys Model The repository contains RHESSys (Tague & Band, 2004; source code) simulated ecohydrological and nutrient (nitrogen only) fluxes at daily, basin-average (RHESSys_basin_output) and monthly, grid (RHESSys_patch_output) levels. We currently simulated the following 8 watersheds in Baltimore: Dead Run Baisman Run Scotts Level Branch Moores Run Powder Mill Run Maidens Choice Run Stony Run The watershed boundaries of all studied watersheds are stored in Watershed_Boundary folder. Variables and their units are listed in the metadata. Spatial projection, NAD83 / UTM zone 18N (EPSG:26918) is used for patch-level, netCDF-format files. For more information, please contact Ruoyu Zhang (rz3jr@virginia.edu).

Baltimore MD↗

CROCUS Dual Polarization Ceilometer Data at Northeastern Illinois University Rooftop

This dataset is from the Department of Energy Office of Science funded project, CROCUS Urban Integrated Field Laboratory (https://crocus-urban.org/). The dual-polarization ceilometer (Vaisala CL61) is an autonomous lidar system operating at 910 nm wavelength, providing valuable measurements for understanding atmospheric boundary layer evolution, air quality, and cloud-aerosol interactions. The CL61 measures the backscattered signal intensity alternating between parallel- and cross-polarization signals. The unique depolarization measurement capability improves discrimination between different particle types, such as liquid droplets, ice crystals, and aerosols. The depolarization is highly dependent on the scatterer shape and orientation (spherical vs non-spherical particles), with the linear depolarization ratio providing a measure of dominant backscatter signal component from atmospheric particles at various heights, essentially allowing discrimination between liquid and solid particles. With its efficient optical system, CL61’s improved signal-to-noise ratio compared to traditional ceilometers allows studying detailed vertical profiles of aerosols and clouds up to 15 km height.Datasets are stored in a netCDF data format, and we we encourage users to make use of the associated toolkits available from Unidata (https://www.unidata.ucar.edu/software/netcdf/), Project Pythia (https://foundations.projectpythia.org/core/data-formats/netcdf-cf.html), and our “Instrument Cookbooks” (https://crocus-urban.github.io/instrument-cookbooks) for more information on how to process the metadata-rich datasets.

54 ENVIRONMENTAL SCIENCES↗