Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Soil nitrogen mineralization rates, nutrient stocks, stable isotopes, and water volumetric measurements across terrestrial-aquatic interfaces from three wetlands at the Tanglewood Biological Station, Alabama

This dataset supports a broader study investigating wetland hydrologic and biogeochemical responses to inundation events. Soil samples were collected across four sampling events along terrestrial-aquatic gradients at three wetland sites located within the Tanglewood Biological Station in Alabama from April 2024 to June 2025. The contents in this data package include soil in-situ nitrogen mineralization rates (measurements of net nitrification, net ammonification, and net mineralization), nutrient stocks (total carbon, total nitrogen, and organic matter), stable isotopes (carbon and nitrogen), and water volumetric measurements (water-filled pore space). Water level data related to each wetland location can be found at https://data.ess-dive.lbl.gov/view/doi:10.15485/2530253 (Kirker et al., 2024), related water geochemistry data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3001967 (Forbes et al., 2025), and related surface water sediment chemistry data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/3377325 (Molina Serpas et al., 2026). In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) field metadata and international generic sample numbers (IGSNs); (4) readme; (5) the field protocol; and (6) a subfolder with sample data. The sample data subfolder contains (1) net nitrification rate, (2) net ammonification rate, (3) areal net mineralization rate, (4) percent organic matter, (5) water-filled pore space, (6) total carbon content, (7) total nitrogen content, (8) stable carbon isotope (delta carbon-13), and (9) stable nitrogen isotope (delta nitrogen-15), and (10) methods codes. All files are .csv or .pdf.

13-C

Integrated Hourly Meteorological Database of 20 Meteorological Stations (1981-2022) for Watershed Function SFA Hydrological Modeling

This dataset contains (a) a script “R_met_integrated_for_modeling.R”, and (b) associated input CSV files: 3 CSV files per location to create a 5-variable integrated meteorological dataset file (air temperature, precipitation, wind speed, relative humidity, and solar radiation) for 19 meteorological stations and 1 location within Trail Creek from the modeling team within the East River Community Observatory as part of the Watershed Function Scientific Focus Area (SFA). As meteorological forcings varied across the watershed, a high-frequency database is needed to ensure consistency in the data analysis and modeling. We evaluated several data sources, including gridded meteorological products and field data from meteorological stations. We determined that our modeling efforts required multiple data sources to meet all their needs. As output, this dataset contains (c) a single CSV data file (*_1981-2022.csv) for each location (20 CSV output files total) containing hourly time series data for 1981 to 2022 and (d) five PNG files of time series and density plots for each variable per location (100 PNG files). Detailed location metadata is contained within the Integrated_Met_Database_Locations.csv file for each point location included within this dataset, obtained from Varadharajan et al., 2023 doi:10.15485/1660962. This dataset also includes (e) a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and (f) a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. Review the (g) ReadMe_Integrated_Met_Database.pdf file for additional details on the script, methods, and structure of the dataset.The script integrates Northwest Alliance for Computational Science and Engineering’s PRISM gridded data product, National Oceanic and Atmospheric Administration’s NCEP-NCAR Reanalysis 1 gridded data product (through the `RCNEP` R package, Kemp et al., doi:10.32614/CRAN.package.RNCEP), and analytical-based calculations. Further, this script downscales the input data into hourly frequency, which is necessary for the modeling efforts.

54 ENVIRONMENTAL SCIENCES

Mineralogy of floodplain sediments from Meanders C, O, and Z in the East River Watershed, CO, USA

This dataset includes bulk X-ray diffraction data from floodplain sediments collected as a part of the Watershed Function Scientific Focus Area (SFA) located in the Upper Colorado River Basin. The data were collected in order to investigate the role of biogeochemical cycling and other river corridor processes on riverine export of solutes. Sediment cores were collected from Meander C, Meander O, and Meander Z in July 2016 to September 2017 to depths of approximately 40-95 cm. Sample metadata including locations, depths, and sample dates are included in a csv file ("sample_list_and_locations.csv"). The file "diffraction_data.csv" contains raw diffraction data, and mineral quantification is in the file "mineral_abundance.csv". This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type.

54 ENVIRONMENTAL SCIENCES

Metagenome-assembled genomes from soil samples in control and warming plots in Blodgett Forest, CA (2014-2021)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of LBNL (Lawrence Berkeley National Laboratory) TES (Terrestrial Ecosystem Science) Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM decomposition and stabilization.Here, we present metagenome-assembled genomes (MAGs) for the bacterial and archaeal community from soil depth profiles collected from 2014 to 2021 from three paired control and warming plots. We collected soil samples across a range of depth profiles (spanning surface to 90 cm deep) from three paired control and warming plots from a temperate mixed forest in Northern California. Each paired plot had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. 101 soil metagenomes were sequenced at JGI (Joint Genome Institute) and UCSF (University of California San Francisco) Center for Advanced Technology and can be found under the JGI (Joint Genome Institute) GOLD (Genomes Online Database) Sequencing project Gs0151586 and NCBI (National Center for Biotechnology Information) Projects PRJNA1225762 and PRJEB39497. Metagenomes were assembled using JGI (Joint Genome Institute) Metagenome Workflow (10.1128/mSystems.00804-20). For each metagenome, the assembled contigs were binned into genomes using 3 binning algorithms (cocacola, metabat, and maxbin) and the resulting bins were consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>50%) and contamination (<25%), and dereplicated at 99% ANI (average nucleotide identity) using dRep (https://github.com/MrOlm/drep).The dataset includes a zip file of 2321 MAG (Metagenome Assembled Genome) fasta files, the accession numbers for the underlying metagenomes, and a csv file with MAG (Metagenome Assembled Genome) quality metrics and taxonomic classification (GTDB -Genome Taxonomy Database-RS220). This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. A sample metadata file (samples.csv) that contains site information has also been included.

54 ENVIRONMENTAL SCIENCES

Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces

These data are from Bandopadhyay et al., "Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces". This study aims to understand the soil microbial ecology along terrestrial-aquatic interfaces of a freshwater and estuarine region and how it relates to organic matter. We analyzed soil microbial (16S rRNA gene) and organic matter (Fourier-transform ion cyclotron resonance mass spectrometry, FTICR-MS) composition from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie (freshwater) and Chesapeake Bay (estuarine) regions. This dataset includes 16S rRNA gene amplicon data (only processed file types included here) and organic matter composition from FTICR-MS data (raw and processed files included here) from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie and Chesapeake Bay regions. These sites are part of the COMPASS-FME project (https://compass.pnnl.gov/FME/COMPASSFME). File formats and software needed to access files: 16S rRNA gene amplicon data: These files follow the format reported here https://ess-dive.gitbook.io/amplicon-sequencing-reporting-format#updates-in-v1.0.1. As per this format, there are four file types reported: 1. Taxon tables (also called sequence-by-sample or OTU (operational taxonomic unit)/ESV (exact sequence variant) tables) : available in a .txt file format and accessible using TextEdit or MS Excel. 2. Representative sequences (also called consensus sequences) : available in a .fasta format and accessible using TextEdit. 3. Sequencing metadata : available in a MS Excel workbook file format and CSV file format 4. Bioinformatic metadata : available in a MS Excel workbook file format and CSV file format FTICR-MS data: 1. Raw data converted to a processed file with intensities of the peaks in the given samples : available in a MS Excel CSV file format 2. Processed file used in analyses and visualizations (appended as icr_long_) : available in a MS Excel CSV file format 3. Metadata file for ICR features (appended as icr_meta) : available in a MS Excel CSV file format

54 ENVIRONMENTAL SCIENCES

Site and endmember spectra of terrestrial vegetation and soils for the Colorado Headwaters Ecological Spectroscopy Study, June-July 2025

This dataset provides site and endmember spectra collected during the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS) campaign. The site spectra were collected to help validate airborne hyperspectral data acquired by the National Ecological Observatory Network's aerial observation platform (NEON AOP). Endmember spectra were collected to augment existing spectral libraries with additional samples of bare surfaces and non-photosynthetic vegetation. All measurements were acquired with an Analytical Spectral Devices (ASD) FieldSpec4 Hi-Res NG (Next Generation) spectroradiometer, which records radiance at 1nm (nanometer) intervals from the ultraviolet to the short-wave infrared (350-2500 nm). The dataset includes spectra measured at meadow sites where the CHESS team also collected vegetation samples for trait analyses. The site spectra were collected with the ASD FieldSpec4 palm grip attachment using an 8° field-of-view foreoptic. Site spectra are integrated measurements of the entire surface within the foreoptic’s field of view. For site-level spectra, the sun is the illumination source. A Spectralon panel mounted on a tripod was used for instrument optimization and white reference measurements for all site spectra. Site spectra were acquired within two hours of solar noon and within 48 hours of a NEON AOP overflight. Site spectra are labeled by date, sampling area, and site number according to the naming conventions of the CHESS campaign’s data management plan. The dataset also contains endmember spectra in the following categories: photosynthetic vegetation (PV), non-photosynthetic vegetation (NPV), bare (soil/rock), and flowers. Endmember measurements were acquired using either the contact probe or the leaf clip attachments of the ASD FieldSpec4. In these configurations, the bulb inside the spectrometer provides the light source for the measurements. The spectrometer was optimized and white reference measurements were recorded using the circular white pucks attached to the contact probe and leaf clip. Because they do not rely on solar illumination, contact probe and leaf clip measurements were collected during a broader time frame than the palm grip site spectra. Some endmembers were measured at CHESS meadow sites, while others were collected within the larger sampling area or in nearby locations (e.g. Gothic Townsite) with similar characteristics. Radiance, reflectance, and metadata files are split into three subfolders according to measurement type: proximal/palm grip (prx), contact probe (cp), and leaf clip (lc). Radiance spectra are provided in ASD file format (.asd file extension). All ASD files can be opened using the provided scripts. Metadata is provided in two formats: CSV file format (no geolocation) and GEOJSON file format (includes geolocation for each spectra). The dataset includes a set of pre-processed reflectance spectra as CSV files (yyyymmdd_rfl.csv). The python scripts and jupyter notebook used to calculate reflectance spectra from the ASD radiance data is included here and was previously published at: https://doi.org/10.3334/ORNLDAAC/2446. There is also a folder of JPEG photographs corresponding to selected spectra. We include a protocol document with detailed steps for ASD FieldSpec4 assembly and operations. This data additionally contains a file level metadata (flmd.csv) and data dictionary (dd.csv) file. Geospatial information: Geospatial data for mapping measurement site locations are in the files CHESS_polygons_lai_UTM.geojson, CHESS_polygons_shrub_UTM.geojson, and CHESS_polygons_meadow_UTM.geojson in the companion geospatial package for the 2025 CHESS campaign, ‘CHESS 2025: Location data for field observations and sampling’ (Henderson et al., 2026). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: This research was carried out at the Jet Propulsion Laboratory, California Institute of Technology, under a contract with the National Aeronautics and Space Administration (80NM0018D0004) and was funded by EMIT Extended Mission Phase E Science.

2018 NEON and 2025 CHESS Campaigns

SPRUCE Aboveground Vegetation Coverage in Root Ingrowth Core Plots, Marcell Experimental Forest, Minnesota, August 2022

This dataset contains vegetation survey measurements from root ingrowth core plots (Määttä et al. 2025) inside SPRUCE Experiment plots at the Marcell Experimental Forest in northern Minnesota. Vegetation surveys were conducted in 0.25 meter2 plots containing root ingrowth cores on August 8th and 9th, 2022 (2022-08-08 to 2022-08-09). The warming and elevated carbon dioxide (CO2) treatments in the dataset include the full treatment gradient: +0 degrees Celsius (C) (+0 and +500 parts per million (ppm) elevated CO2), +2.25 degrees C (+0 and +500 ppm), +4.5 degrees C (+0 and +500 ppm elevated CO2), +6.75 degrees C (+0 and +500 ppm) and +9 degrees C (+0 and +500 ppm elevated CO2) for both hummocks and hollows. This dataset includes measurements of the height and absolute coverage (%) for each vascular plant and moss species, as well as organic litter and dead overstory vascular plants, and the distance from the grid center to the nearest tree and the species of the nearest tree. These data were used as species-specific aboveground plant metadata for assessing the warming and elevated CO2 response of fine roots across different peatland microtopographical features (hummocks and hollows) and plant functional types (shrub, spruce and larch). This dataset contains one data file in comma separate (.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma separate (.csv) format and a user guide in PDF (*.pdf) format.

ESS-DIVE CSV File Formatting Guidelines Reporting

Data and scripts associated with “Moisture content modulates DOM thermodynamic regulation of oxygen consumption in drying streambed sediments”

This data package is associated with the publication “Moisture content modulates DOM thermodynamic regulation of oxygen consumption in drying streambed sediments” published in Scientific Reports (Garayburu-Caruso et al., 2026). The package contains processed data products and scripts used to quantify how drying and re-inundation of riverbed sediments influence dissolved organic matter (DOM) thermodynamic properties and their relationship with sediment oxygen (O₂) consumption across 33 stream sites in the contiguous United States. The data package contains DOM thermodynamic metrics (e.g., Gibbs free energy of carbon oxidation and thermodynamic efficiency), and O₂ consumption along with watershed-scale climate and land-cover metrics used as explanatory variables in the analyses. Underlying unprocessed and processed ultrahigh-resolution mass spectrometry data, oxygen consumption rates from laboratory moisture-manipulation experiments, within-sample environmental properties, sediment moisture content and contextual field measurements are archived separately at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/2428003 (Laan et al., 2024) and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689 (Forbes et al.,2023). A preliminary version of this data package was published in February 2026 at the time of manuscript submission. It was updated in June 2026, at the time of manuscript acceptance, to include the finalized data and additional metadata (readme, data dictionary, and file level metadata). For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. At the top level, the data package is organized into five main folders: (1) Data, (2)Figures, (3) Map, (4) GAM_Reulsts, and (5) src. The Data folder contains analysis-ready tabular files with oxygen consumption rates, DOM thermodynamic properties by site and treatment, site-level environmental variables, watershed-scale metrics, and other derived variables referenced in the manuscript. The Figures folder contains static image files associated with the main text and supplemental figures, while the Map folder includes spatial data and map-layer files used to create the sampling-location map. The GAM results folder contains the results for each of the general additive model (GAM).The src folder contains R scripts used to perform data processing, statistical analyses (including clustering, generalized additive models, and threshold analysis), and figure generation. This data package is associated with a GitHub repository found at https://github.com/WHONDRS-Hub/ECA_DOM_Thermodynamics.

Dissolved organic matter

Timeseries Photos of a Variably Inundated Stream: Umtanum Creek, Washington, United States

This dataset is associated with a broader study using game camera timeseries photos collected to evaluate stream variable inundation via changes in width (i.e. wet fraction). Four game cameras were deployed along Umtanum Creek (Washington, United States) to track changes in stream inundation over time. Drone imagery was collected at the same location on October 18, 2024 which was used to construct a digital elevation model (DEM) of the streambed topography. The associated paper and data can be found at https://doi.org/10.1016/j.envsoft.2025.106715 (Bao et al., 2025a)) and https://doi.org/10.15485/2589885 (Bao et al., 2025b), respectively. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to this readme, this data package also includes a file-level metadata (FLMD) files that describes each file and a data dictionaries (DD) that describe all column/row headers and variable definitions. This dataset is comprised of (1) file-level metadata; (2) data dictionary; (3) readme; (4) field metadata; (5) field protocol; and (5) folders containing game camera photos. Game camera photos are organized into folders for each camera (CDL, CUL, CDR, CUR; see readme for information on camera naming) by the month photos were collected. All files are .csv, .jpg, or .pdf.

AI image segmentation

Dated soil C–N–P profiles, water quality, and chamber fluxes across Ohio and Michigan wetlands (2024–2025)

This dataset includes dated soil core chemistry (bulk density, phosphorus, nitrogen and carbon concentrations), water quality, and chamber flux measurements collected from wetlands in the Midwest United States—12 sites in Ohio, one site in Indiana, one site in Michigan—collected in the spring or summer of 2024 or 2025, all in (.csv) format. These data were generated to examine how wetland restoration, management activities, and time since restoration affect biogeochemical processes, carbon sequestration, nutrient accumulation, water quality, and greenhouse gas emissions. Specifically, these data aim to investigate how restored wetlands differ from natural wetlands in terms of carbon, nitrogen, phosphorus dynamics, as well as carbon dioxide and methane fluxes. Also included are surface and porewater quality parameters and chamber flux measurements across these different wetlands. Sampling was conducted at various sites representing a range of restoration stages, from about 4 years post-restoration up to 105 years post-restoration, and also includes a natural wetland used as a reference in Michigan. These data can be used to determine carbon sequestration rates, nutrient cycling, and to enhance our understanding of biogeochemical responses to wetland restoration in temperate ecosystems. This data package contains (1) a csv file (Water_Quality.csv) containing water quality data (dissolved organic carbon, total dissolved nitrogen, and temperature) organized by location; (2) a csv file (Soil_C_N_P_Seq.csv) containing carbon, nitrogen, and phosphorus concentrations at each soil level and time of each soil level, as well as their sequestration rates; (3) a csv file (CH4_CO2_Flux.csv) including methane and carbon dioxide fluxes that were measured with a chamber; (4) a file-level metadata (FLMD.csv) file that lists each file contained in the dataset with associated metadata; (5) a data dictionary (DD.csv) file that contains terms/column headers used throughout the files along with a definition, units, and data type; and (6) a locations metadata file (Location_metadata.csv).

Earth Science > Atmosphere > Atmospheric Chemistry

A machine learning method of modern urban building energy modeling: A case study of Chicago

Urban-scale building energy modeling is vital for urban planning. However, it can be challenging to assimilate reliable non-geometry building data for urban-scale modeling without extensive investment. Here, this study introduces a novel approach to developing modern urban-scale building energy stock data using geographic information systems and machine learning algorithms without necessarily requiring pre-supplied non-geometric metadata. The proposed framework integrates building footprint and height data to estimate gross floor areas, and matches each building to a pool of candidate records from ComStock or ResStock—filtered to the same county and ranked by geometric similarity—demonstrate a proof-of-concept case study in Chicago for predicting energy use intensity (EUI) using scalable datasets. The model achieved a mean bias error (MBE) of 0.08 kWh/m² and root mean square error (RMSE) of 14.84 kWh/m² under full metadata input for EUI prediction. With only location inputs, the model captured 69.2 % of EUI within predicted ranges. These results demonstrate the model’s potential to support early-stage urban planning, identify candidates for energy-efficient retrofits. By removing the dependency on detailed pre-surveys or extensive building metadata, the approach overcomes a key barrier in traditional urban-scale building energy modeling, illustrating a pathway toward broader and more cost-effective application, though further multi-city validation and improved treatment of pre-1925 buildings are needed.

Energy Use Intensity

mzPeak: Designing a Scalable, Interoperable, and Future-Ready Mass Spectrometry Data Format

Advances in mass spectrometry (MS) instrumentation, such as higher resolution, faster scan speeds, and improved sensitivity, have significantly increased the volume and complexity of data. The growing adoption of imaging and ion mobility further amplifies these challenges across MS-based omics fields, including proteomics, metabolomics, and lipidomics. While these technologies unlock new possibilities, they also present significant challenges in data management, storage, and accessibility. Existing open formats, such as the XML-based community standards mzML and imzML, struggle to meet the demands of modern MS workflows due to their large file sizes, slow data access, and limited metadata support. Vendor-specific formats, while optimized for proprietary instruments, lack interoperability, comprehensive metadata support and long-term archival reliability. This white paper lays the groundwork for mzPeak, a next-generation community data format designed to address these challenges and support high-throughput, multi-dimensional MS workflows. By adopting a hybrid model that combines efficient binary storage for numerical data and both human and machine-readable metadata storage, mzPeak will reduce file sizes, accelerate data access, and offer a scalable, adaptable solution for evolving MS technologies. For researchers, mzPeak will enable enhanced interoperability across platforms, seamless support for complex workflows including ion mobility and MS imaging, and faster data access compared to existing community formats such as mzML. Its design will ensure data is managed in compliance with regulatory standards, essential for applications such as precision medicine and chemical safety, where long-term data integrity and accessibility are critical. For vendors, mzPeak provides a streamlined, open alternative to proprietary formats, reducing the burden of regulatory compliance while aligning with the industry's push for transparency and standardization. By offering a high-performance, interoperable solution, mzPeak positions vendors to meet customer demands for sustainable data management tools which will be able to handle emerging and future data types and workflows. mzPeak aspires to become the cornerstone of MS data management, empowering researchers, vendors, and developers to innovate and collaborate more effectively.

data formats

Identifying genomic data use with the Data Citation Explorer

Increases in sequencing capacity, combined with rapid accumulation of publications and associated data resources, have increased the complexity of maintaining associations between literature and genomic data. As the volume of literature and data have exceeded the capacity of manual curation, automated approaches to maintaining and confirming associations among these resources have become necessary. Here we present the Data Citation Explorer (DCE), which discovers literature incorporating genomic data that was not formally cited. This service provides advantages over manual curation methods including consistent resource coverage, metadata enrichment, documentation of new use cases, and identification of conflicting metadata. The service reduces labor costs associated with manual review, improves the quality of genome metadata maintained by the U.S. Department of Energy Joint Genome Institute (JGI), and increases the number of known publications that incorporate its data products. The DCE facilitates an understanding of JGI impact, improves credit attribution for data generators, and can encourage data sharing by allowing scientists to see how reuse amplifies the impact of their original studies.

59 BASIC BIOLOGICAL SCIENCES

ATAT: Astronomical Transformer for time series and Tabular data

Context. The advent of next-generation survey instruments, such as theVera C. RubinObservatory and its Legacy Survey of Space and Time (LSST), is opening a window for new research in time-domain astronomy. The Extended LSST Astronomical Time-Series Classification Challenge (ELAsTiCC) was created to test the capacity of brokers to deal with a simulated LSST stream. Aims. Our aim is to develop a next-generation model for the classification of variable astronomical objects. We describe ATAT, the Astronomical Transformer for time series And Tabular data, a classification model conceived by the ALeRCE alert broker to classify light curves from next-generation alert streams. ATAT was tested in production during the first round of the ELAsTiCC campaigns. Methods. ATAT consists of two transformer models that encode light curves and features using novel time modulation and quantile feature tokenizer mechanisms, respectively. ATAT was trained on different combinations of light curves, metadata, and features calculated over the light curves. We compare ATAT against the current ALeRCE classifier, a balanced hierarchical random forest (BHRF) trained on human-engineered features derived from light curves and metadata. Results. When trained on light curves and metadata, ATAT achieves a macro F1 score of 82.9 ± 0.4 in 20 classes, outperforming the BHRF model trained on 429 features, which achieves a macro F1 score of 79.4 ± 0.1. Conclusions. The use of transformer multimodal architectures, combining light curves and tabular data, opens new possibilities for classifying alerts from a new generation of large etendue telescopes, such as theVera C. RubinObservatory, in real-world brokering scenarios.

Astronomy & Astrophysics

NEPATEC2.0: NEPA Text Corpus v2.0

The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for organizing and sharing NEPA-related information across agencies. Much of the information gathered during NEPA reviews is written into documents such as categorical exclusions, environmental assessments, and environmental impact statements, then filed in predominately independent agency file stores that may or may not be publicly accessible. The application of metadata and data standards, such as those recommended by the Council on Environmental Quality (CEQ), to NEPA documents offers a shared vocabulary and structure for key entities like projects, processes, and documents that can streamline information exchange and enhance collaboration across systems. In this work, we publicly release NEPATEC2.0, an expanded corpus of NEPA documents with associated metadata. NEPATEC2.0 encompasses approximately 120,000 documents from 60,000 projects prepared by more than 60 different agencies. Modeled to align with CEQ metadata standards, NEPATEC2.0 promotes consistency in environmental reviews and supports the ongoing effort to modernize permitting technologies by facilitating more transparent, efficient, and data-driven decision-making. Importantly, NEPATEC2.0 demonstrates the possibilities and limitations of large language model-based prompting to extract information from NEPA documents at scale.

environmental review

Drifting Acoustic Measurements around C-Power's SeaRay WEC

The repository contains underwater noise measurements and associated metadata collected around C-Power's SeaRay wave energy converter on July 15, 2024 and July 16, 2024 while it was deployed at the U.S. Navy's Wave Energy Test Site (WETS) in Kaneohe, HI. Measurements were obtained using Drifting Acoustic Instrumentation SYstems (DAISYs). DAISYs consist of a surface expression connected to a hydrophone recording package by a tether. Both elements are instrumented to provide metadata (e.g., position, orientation, and depth). Information about how to build DAISYs is available at https://www.pmec.us/research-projects/daisy. The repository's primary content is a compressed archive (.zip format), containing multiple MATLAB binary data files (.mat format). The structure of each file is included in the repository as a Word document (Data Description MHK-DR.docx). Each file contains time series information for a single DAISY deployment (file naming convention: WETS_DAISY_[Drift #].mat) consisting of processed hydrophone data and associated metadata. During these measurements, C-Power's SeaRay was located at approximately 21.48112 N, 157.74451 W.

16 TIDAL AND WAVE POWER

1H-NMR characterization of soil dissolved organic matter from soil samples in control and warming plots in Blodgett Forest, CA (2014 and 2018)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of Lawrence Berkeley National Laboratory Terrestrial Ecosystem Science Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM (soil organic matter) decomposition and stabilization. This package contains metabolite data obtained through 1H nuclear magnetic resonance (NMR) spectroscopy on water-extracted soils. Soil samples were collected in 2014/06/03 and 2018/06/04 from 3 replicated paired plots that had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. The following files are included: (1) nmr_h2o_data_raw.csv: raw data, (2) nmr_h2o_data_processed.csv: computed compound concentrations and metadata, (3) nmr_h2o_compound_metadata.csv: compound metadata, (4) nmr_h2o_sample_metadata.csv: sample metadata

1H-NMR (nucleic magnetic resonance) spectroscopy

AIACHNE's contribution for Nuclear Energy Agency Working Party on International Nuclear Data Evaluation Co-operation Subgroup 50

The AIACHNE (AI/ML Informed cAlifornium CHi Nuclear data Experiment) project aims at designing an experiment for the 252 Cf Prompt Fission Neutron Spectrum (PFNS) that explores systematic biases in an experimental database retrieved from the EXFOR databases. To that end, machine learning (ML) methods were applied to pint-point measurement features likely related to bias. From that information, we selected a feature that should be explored by the AIACHNE experiment. Measurement features are metadata encapsulating all pertinent information about the physical measurement and analysis techniques. Examples are, for instance, what neutron and fission detectors were used for the physical metadata, and what background reduction techniques were employed for analysis techniques. Such metadata were retrieved both from EXFOR entries as well as the literature of data sets described in detail in Reference 2 (at the end of the article).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS