Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “dataset citation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Citation network datasets for benchmarking spiking graph neural networks on experimental neuromorphic hardware

Spiking neural networks (SNNs) running on neuromorphic computers offer an energy-efficient alternative for AI tasks. Recently, spiking graph neural networks (S-GNNs) have been shown to produce encouraging results on benchmark citation network datasets such as Cora, CiteSeer, and PubMed for node classification tasks. These S-GNNs were run on SNN simulators only because they contain up to tens of thousands of neurons and up to millions of synapses, translating poorly to neuromorphic hardware. Therefore, in this paper, we create a suite of benchmark datasets from the CiteSeer dataset that can be accommodated on current neuromorphic hardware platforms. Our contribution consists of a collection of three datasets. First, we have an induced subgraph of CiteSeer, which we call MiniSeer, containing 2110 papers, 3604 binary features, and 6 topics. Second, MicroSeer is a very small dataset consisting of 84 papers, 1227 features, and 6 topics. Lastly, BiteSeer is a collection of 15 binary classification datasets. We present creation of these datasets along with accuracies, running times, and spike counts when simulated. We believe that our results in this paper will be used by the neuromorphic community to benchmark, test, and develop neuromorphic hardware and simulators.

Zhu, Kevin [George Mason University, Virginia]↗

PNNL INFRARED REFRACTIVE INDEX (n/k) DATASET FOR SEVEN PAH SOLIDS AT ROOM TEMPERATURE

This dataset is an open-source repository of spectral data measured at Pacific Northwest National Laboratory (PNNL). This database provides quantitative values for the complex index of refraction for seven polycyclic aromatic hydrocarbon (PAH) solids. A list of the chemicals is available in the readme file. These spectra consist of the optical constants, i.e., the real, n(ν), and imaginary, k(ν), refractive indices, over the spectral range from 7,800 to 400 cm-1 (1.28 – 25 μm). The conditions under which the individual data were acquired are described in the associated metadata files, and the user is strongly encouraged to read and understand this information to ensure the data are used appropriately for your application. Recommended Citation for Dataset Jessica M Salcido, Jeremy D. Erickson, Ashley M. Bradley, Russell G. Tonkyn, Timothy J. Johnson and Tanya L. Myers. 2026. PNNL INFRARED REFRACTIVE INDEX (n/k) DATASET FOR SEVEN PAH SOLIDS AT ROOM TEMPERATURE. [Data Set] PNNL DataHub. INSERT DOI License Information This work is marked with CC0 1.0: https://creativecommons.org/publicdomain/zero/1.0/. The authors do request that you appropriately cite the dataset when referencing or using the dataset.

Salcido, Jessica Marie Ortola↗

Hot Droughts and Forest Tree Dynamics in the Amazon - Statistical Models, Scripts, Data, and Outputs

This package contains data, outputs, equations, and R scripts for analyses for manuscript entitled "Hot droughts in the Amazon: A window to a future hypertropical climate" by J. Chambers et al., in particular it contains statistical models and analyses for the INPA BIONTE tree mortality study. The Models folder contains details for all statistical models in PDF files. The Scripts folder contains the R scripts for Bayesian Hierarchical Models (two text files) and SEMs (one text file) are separate and reasonably annotated. All data associated with these scripts are in the data folder. The Data folder contains two of the three CSV files used for the analyses and are called by the R scripts. Two of them are part of published datasets (`BIONTE_mortality-rates.csv` from Lima et al. 2024, DOI:10.15486/ngt/1898910 and `SPEI.csv` from Pastorello et al. 2023 DOI:10.15486/ngt/1958257) and also provided in this package for convenience (please see the corresponding datasets for usage and citation terms). The third dataset (`BIONTE_gapfilled_wd.csv`) contains sensitive information and can be obtained by contacting the manuscript lead author. The Outputs folder contains the two output files that provide extra information about the analyses. The file `figuresFeb2025d.pdf` contains all the figures from the manuscript - captions are in the manuscript. The file `ChambersMS.pdf` contains primary results from Bayesian statistical models, regression analyses, and validation steps applied to the tree mortality data from the INPA experiments. The document includes visual summaries, model diagnostics, and leave-one-out (LOO) validation results. A breakdown of file contents can be found in the README file that is part of this package.

54 ENVIRONMENTAL SCIENCES↗

NEWTS State-Level Database Dashboard

The National Energy Water Treatment and Speciation (NEWTS) State-Level Database Dashboard enables quick visualization and exploration of energy-related wastewater data collected from state environmental agencies and research projects across the United States. The dashboard features a subset of the NEWTS Integrated Dataset (version 1.0), which was built to provide stakeholders with a unified and standardized database to support wastewater treatment and critical minerals exploration. The full NEWTS Integrated Dataset and citations for the original data resources can be found here: https://edx.netl.doe.gov/dataset/newts-integrated-dataset-version-1-0

abandoned mine drainage↗

Unlocking Scholarly Insights: Leveraging Machine Learning Approaches for Citation Analysis and Intent Classification

Publicly funded organizations, notably institutions like the Los Alamos National Laboratory (LANL), are deeply vested in acquiring robust productivity metrics to gauge the entirety of their research output. Motivated by the imperative to enhance institutional productivity assessment, this study investigates the utilization of Large Language Models (LLM) such as BERT-based models, as well as local LLaMa-30b-instruct and Mixtral-8x7b-instruct architectures for classifying type of URL referenced resources in academic papers such as software, dataset, as well as authorship intent. Challenges in discerning resource types from context are highlighted, along with the potential of BERT and LLMs to address these challenges. Through comprehensive analysis, this research unveils a notable surge in documents featuring URL citations, indicative of the escalating importance of digital resources in scholarly publications. Moreover, citations to datasets and software demonstrate consistent growth over time, underscoring their increasing significance. Our findings also reveal that LANL authors contribute substantially to accessible science, comprising about 10% of dataset and software mentions in LANL

Large Language Models, BERT, citation classificati↗

COMPASS-FME Synoptic Sites Level 1 Sensor Data v2-1

This is the version 2-1 Level 1 (L1) data release for COMPASS-FME environmental sensors located at our synoptic field sites. COMPASS-FME is studying sites in two distinct regions, the Chesapeake Bay and the Western Lake Erie Basin. We established the network at seven "synoptic" (observational) sites along the Chesapeake Bay and Lake Erie coastlines, collectively generating over three million observations per month, to track and comprehend environmental changes where land and water intersect. Additionally, the two regions provide an interesting contrast of saltwater and freshwater coasts that allow us to differentiate the impacts of inundation and coastal water chemistries in two nationally important coastal systems. L1 data are close to raw, but are units-transformed and have out-of-instrument-bounds, out-of-service, and outlier flags added. Duplicates and missing data are removed but otherwise these data are not filtered, and have not been subject to any additional algorithmic or human QA/QC. Any scientific analyses of L1 data should be performed with care. **This dataset will be updated quarterly with new data for the duration of the project** This dataset includes: - An overall dataset README file that describes the current version, gives citation and contact information, etc. - Site- and year-specific folders, each holding up to 12 CSV (comma separated value) data files for each site and plot in that year. - Metadata files within each site-year folder provide full information on data units, expected ranges, contact information, as well as a general description of the site. - Environmental sensor types that appear in the data files include weather (ClimaVUE50, CS, RM Young, and LI instruments in the graphs below); soil conditions (TEROS12); soil redox state (Redox); groundwater variables (AquaTROLL200 and AquaTROLL600); open water sondes (Exo); tree sap velocity (Sapflow); and system voltage and state (Datalogger). Data are normally logged every 15 minutes. Please see v2-0 Synoptic L1 Sensor Package Quick Start.pdf for detailed information on data package structure, temporal coverage, and versioning. This dataset was updated 2026-03-12: (i) data now go through 2025-12-31 (previous end was 2025-06-30) and (ii) dataset and file names updated to “…v2-1” (previously was “v2-0”).

54 ENVIRONMENTAL SCIENCES↗

COMPASS-FME Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) Experiment Level 1 Sensor Data v2-1

This is the version 2-1 Level 1 (L1) data release for COMPASS-FME environmental sensors located at our Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) experimental site. This manipulative, ecosystem-scale TEMPEST experiment addresses the potential for freshwater and estuarine-water disturbance events to alter tree function, species composition, and ecosystem processes in a deciduous coastal forest in MD, USA. The experiment uses a large-unit (2000 m2), un-replicated experimental design, with three 50 m × 40 m plots serving as control, freshwater, and estuarine-water treatments. L1 data are close to raw, but are units-transformed and have out-of-instrument-bounds, out-of-service, and outlier flags added. Duplicates and missing data are removed but otherwise these data are not filtered, and have not been subject to any additional algorithmic or human QA/QC. Any scientific analyses of L1 data should be performed with care. **This dataset will be updated quarterly with new data for the duration of the project** This dataset includes: - An overall dataset README file that describes the current version, gives citation and contact information, etc. - Site- and year-specific folders, each holding variable-specific CSV (comma separated value) data files for each site and plot in that year. - Metadata files within each site-year folder provide full information on data units, expected ranges, contact information, detailed flood times, as well as a general description of the site. - Environmental sensor types that appear in the data files include weather (ClimaVUE50, CS, RM Young, and LI instruments in the graphs below); soil conditions (TEROS12); soil redox state (Redox); groundwater variables (AquaTROLL200 and AquaTROLL600); open water sondes (Exo); tree sap velocity (Sapflow); and system voltage and state (Datalogger). Data are normally logged every 15 minutes. Please see v2-1 TEMPEST L1 Sensor Package Quick Start.pdf for detailed information on data package structure, temporal coverage, and versioning. The TEMPEST flood events occurred on the following dates. They lasted for ~10 hours each day and delivered ~80,000 gallons to each plot; many data streams are available at 1 or 5 minute frequency during these periods. * Tests: Aug 25 (fresh plot) and Sep 9 (salt plot), 2021 * TEMPEST 1: June 22, 2022 * TEMPEST 2: June 6-7, 2023 * TEMPEST 3: June 11-13, 2024 This dataset was updated 2026-03-12: (i) data now go through 2025-12-31 (previous end was 2025-06-30) and (ii) dataset and file names updated to “…v2-1” (previously was “v2-0”).

54 ENVIRONMENTAL SCIENCES↗

COMPASS-FME Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) Experiment Level 2 Sensor Data v2-1

This is the version v2-1 Level 2 (L2) data release for COMPASS-FME environmental sensors located at our Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) experimental site. This manipulative, ecosystem-scale TEMPEST experiment addresses the potential for freshwater and estuarine-water disturbance events to alter tree function, species composition, and ecosystem processes in a deciduous coastal forest in MD, USA. The experiment uses a large-unit (2000 m2), un-replicated experimental design, with three 50 m × 40 m plots serving as control, freshwater, and estuarine-water treatments. Level 2 (L2) data consist of sensor observations from the COMPASS-FME synoptic sites, TEMPEST, and DELUGE. Compared to the L1 data, these are more consistent (always 15-minute timestamps for the entire year); better QA/QC’d (out of bounds, out of service, and extreme outlier values are removed); and more complete, with a gap-filled time series available alongside the main observations, and additional derived (calculated) variables. L2 data are intended to be rapidly and easily usable in analyses and simulations. However, algorithmic outlier identification always carries the risk of removing valid data, and Level 1 data may be more suitable for analyses that focus on variability or extreme events. This dataset includes: - An overall dataset README file that describes the current version, gives citation and contact information, etc. - Site- and year-specific folders, each holding variable-specific Parquet (a high performance, space efficient format; see https://parquet.apache.org) data files for each site and plot in that year. - Metadata files within each site-year folder provide full information on data units, expected ranges, contact information, detailed flood times, as well as a general description of the site. - Environmental sensor types that appear in the data files include weather (ClimaVUE50, CS, RM Young, and LI instruments in the graphs below); soil conditions (TEROS12); soil redox state (Redox); groundwater variables (AquaTROLL200 and AquaTROLL600); open water sondes (Exo); tree sap velocity (Sapflow); and system voltage and state (Datalogger). Data are reported every 15 minutes. Please see v2-1 TEMPEST L2 Sensor Package Quick Start.pdf for detailed information on data package structure, temporal coverage, and versioning. Data files are in Apache Parquet, a high performance, space efficient format for tabular data. These files can be read using R's `arrow` package (https://arrow.apache.org/docs/r/), with similar tools available in other languages. The TEMPEST flood events occurred on the following dates. They lasted for ~10 hours each day and delivered ~80,000 gallons to each plot; many data streams are available at 1 or 5 minute frequency during these periods. * Tests: Aug 25 (fresh plot) and Sep 9 (salt plot), 2021 * TEMPEST 1: June 22, 2022 * TEMPEST 2: June 6-7, 2023 * TEMPEST 3: June 11-13, 2024

EARTH SCIENCE > ATMOSPHERE > ATMOSPHERIC TEMPERATU↗

COMPASS-FME Synoptic Sites Level 2 Sensor Data v2-1

This is the version 2-1 Level 2 (L2) data release for COMPASS-FME environmental sensors located at our synoptic field sites. COMPASS-FME is studying sites in two distinct regions, the Chesapeake Bay and the Western Lake Erie Basin. We established the network at seven "synoptic" (observational) sites along the Chesapeake Bay and Lake Erie coastlines, collectively generating over three million observations per month, to track and comprehend environmental changes where land and water intersect. Additionally, the two regions provide an interesting contrast of saltwater and freshwater coasts that allow us to differentiate the impacts of inundation and coastal water chemistries in two nationally important coastal systems. Level 2 (L2) data consist of sensor observations from the COMPASS-FME synoptic sites, TEMPEST, and DELUGE. Compared to the L1 data, these are more consistent (always 15-minute timestamps for the entire year); better QA/QC’d (out of bounds, out of service, and extreme outlier values are removed); and more complete, with a gap-filled time series available alongside the main observations, and additional derived (calculated) variables. L2 data are intended to be rapidly and easily usable in analyses and simulations. However, algorithmic outlier identification always carries the risk of removing valid data, and Level 1 data may be more suitable for analyses that focus on variability or extreme events. This dataset includes: - An overall dataset README file that describes the current version, gives citation and contact information, etc. - Site- and year-specific folders, each holding variable-specific Parquet (a high performance, space efficient format; see https://parquet.apache.org) data files for each site and plot in that year. - Metadata files within each site-year folder provide full information on data units, expected ranges, contact information, detailed flood times, as well as a general description of the site. - Environmental sensor types that appear in the data files include weather (ClimaVUE50, CS, RM Young, and LI instruments in the graphs below); soil conditions (TEROS12); soil redox state (Redox); groundwater variables (AquaTROLL200 and AquaTROLL600); open water sondes (Exo); tree sap velocity (Sapflow); and system voltage and state (Datalogger). Data are reported every 15 minutes. Data files are in Apache Parquet, a high performance, space efficient format for tabular data. These files can be read using R's `arrow` package (https://arrow.apache.org/docs/r/), with similar tools available in other languages. Please see v2-1 L2 Sensor Package QStart.pdf for detailed information on data package structure, temporal coverage, and versioning.

EARTH SCIENCE > ATMOSPHERE > ATMOSPHERIC TEMPERATU↗

Soil microbiome resilience to short-term (30 days, 90 days) and long-term (1000 days) drought

This dataset contains data used for the paper "Drought duration does not impact soil microbiome resilience". The Related References will be updated with a full citation when available. Increasing global droughts exert large but poorly understood effects on the microbial communities and ecology of soil. Microbial communities generally show resilience and return to pre-drought conditions when short-term droughted soils are rewet; soils exposed to long-term drought, however, often show a lag upon rewetting, after which microbial communities may or may not return to their pre-stressed conditions. Though short-term droughts have been widely studied, long-term drought manipulation experiments remain rare, especially those that compare microbial response to short-term and long-term drought in tandem. We conducted a 1000-day drought simulation in controlled laboratory conditions with soil cores collected from a tidal freshwater ecosystem in Washington state, USA, and subsequently exposed them to rewetting for two weeks. We also included short-term (30-day and 90-day) drought and rewet treatments to directly compare microbial community and organic matter responses across drought durations. We found distinct microbial taxa belonging to Firmicutes and Actinobacteria enriched after the 1000-day drought, but not after the short-term droughts. While we hypothesized that the microbial community would recover from a short-term drought after rewetting to resemble pre-drought conditions, our results revealed community dissimilarities between rewet and pre-drought conditions across all drought durations. These findings suggest unique microbial life history strategies within certain microbial phyla that make them successful colonizers during an extended drought period, and the influence of environmental and physiological context on microbial responses to rewetting. The 16SrRNA gene amplicon dataset contains processed DNA sequences in the form of an ASV table with raw unrarefied read counts and representative sequences in .fasta format as described in the ESS-DIVE amplicon sequence reporting format (https://ess-dive.gitbook.io/amplicon-sequencing-reporting-format/instructions). The Fourier Transform Ion Cyclotron Resonance Mass Spectrometry (FTICR-MS) dataset consists of processed files containing presence absence data of molecular formulae and molecular characterization of FTICR resolved peaks. The Nuclear Magnetic Resonance (NMR) dataset contains files relevant to NMR spectra and peaks. A sample key file and a sample metadata file is included for the FTICR/NMR and 16S dataset respectively.

1000-day drought↗

CHESS 2025: Crown polygons and extracted reflectance for field sampling sites

This dataset contains (1) crown polygons for each tree, meadow, and shrub site sampled in the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS) campaign (in geojson format, .geojson) and (2) extracted reflectance, uncertainty, and shade estimates for each crown polygon from the 2018 National Ecological Observatory Network (NEON) and 2025 CHESS campaigns. (in CSV format, .csv). Additional metadata are provided in a data dictionary describing column names and definitions (dd.csv), and in a file-level metadata file (flmd.csv). Crown polygons were manually delineated for each site in the 2025 campaign using a combination of field-collected GPS data (doi:10.15485/3022418), RGB (red, green, blue) and false color reflectance mosaics (doi:10.15485/3013535), and LiDAR-derived (Light Detection and Ranging) canopy height (CHM) and digital surface (DSM) models (DOI and citation to be added upon publication). Where there was misalignment between the spectrometer- and LiDAR-derived data products, polygons prioritized alignment with the spectrometer-derived data products. Polygons were delineated conservatively to only select pixels representative of vegetation samples collected in the field. Crown polygons for 2018 are published at (doi:10.15485/1618130) and were developed using the same protocol. For each polygon, all pixels from all flightlines were extracted where the pixel centroid was contained within the polygon. For each pixel, we extracted the surface reflectance, uncertainty, and shade estimates. Details on the extracted datasets are available at (doi:10.15485/3013527, doi:10.15485/3013535). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: This research was carried out at the Jet Propulsion Laboratory, California Institute of Technology, under a contract with the National Aeronautics and Space Administration (80NM0018D0004) and was funded by EMIT Extended Mission Phase E Science.

2018 NEON and 2025 CHESS Campaigns↗

CHESS 2025: Field-collected vegetation attributes and site photos

This dataset represents field observations of vegetation samples collected as part of the Colorado Headwaters Ecological Spectroscopy Study (CHESS) during June and July of 2025. Samples were collected in the field using tablet computers and digital forms, with target data differing by sample type (individual trees, individual shrubs, or 1-meter square plots of meadow and subshrub vegetation). Field samples were collected within 72 hours of airborne data collection using the National Ecological Observatory Network’s Aerial Observation Platform (NEON AOP). The NEON AOP collected waveform LiDAR (Light Detection and Ranging) and imaging spectrometer data in 426 spectral bands from the visible to shortwave infrared. Remote sensing data for the project is available on ESS-DIVE (DOI and citation to be added upon publication). Field data collected included canopy height and per-species horizontal proportional cover for meadow plots, species identity and height information for shrubs, as well as species identity, height, diameter at breast height, and health assessment information for trees. Photos of the focal site and surrounding landscape were taken for all sampling sites and are included in this archive. Green leaves or needles were collected for plant trait and foliar chemistry analysis. This data is archived separately (DOI and citation to be added upon publication). High-precision geospatial data for each sample (crown perimeter polygons for trees and shrubs, plot boundaries for meadow plots) is available here (Henderson et al., 2026). Field and remote sensing protocols largely followed those of a previous field and airborne imaging campaign performed in 2018 (described in Chadwick et al. 2020). Field data from the 2018 campaign can be found here (Chadwick et al., 2020 doi:10.15485/1618130). Because different field measurements were taken for meadow, shrub, and tree sites, data from these three sample types are archived as separate tables (chess_meadow_site_cleaned.csv, chess_shrub_site_cleaned.csv, chess_tree_site_cleaned.csv). Meadow proportional cover data is stored in a separate table (chess_meadow_cover_cleaned.csv). Taxonomy was treated identically between sample types, and the dataset shares a common set of voucher specimens (chess_voucher_IDs_cleaned.csv), as well as a single species list (chess_species_list_cleaned.csv). All taxonomic determinations were performed to the species level, and adhere to the Global Biodiversity Information Facility (GBIF) backbone taxonomy as of January 10th, 2026 (GBIF Secretariat 2023). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

Carbon Transport and Storage Planning and Viability Support Tools

The EDX disCO2ver Carbon Transport and Storage Planning and Viability Support Tools are made up of the Carbon Storage Planning Inquiry Tool (CS PlanIT, Justman et al. 2024) and the Carbon Storage Technical Viability Approach Support Tool (CS TVA). Together, these tools support data access to support understanding data availability to support planning efforts for carbon transport and storage. The Carbon Storage Planning Inquiry Tool (CS PlanIT) is an online web mapping application designed to help users explore, query, and evaluate multiple data layers to support and accelerate carbon storage resource and feasibility assessments and planning efforts. CS PlanIT currently contains a range of datasets associated with geologic, technical, and infrastructure factors. The data sets can be filtered geographically for an area of interest to update statistics and charts within the dashboard. The dashboard is divided into different sections called widgets, relating to different steps in the carbon storage planning process. The resources in this submission include a link to PlanIT, as well as a data catalog and link to user documentation. The original citation for the CS PlanIT tool, which has now been integrated into the toolset here, was: - Devin Justman, Scott Pantaleone, Maneesh Sharma, Lucy Romeo, Paige Morkner, CS PlanIT (Carbon Storage Planning Inquiry Tool) , 6/28/2024, https://edx.netl.doe.gov/dataset/cs-planit-carbon-storage-planning-inquiry-tool, DOI: 10.18141/2377953 The Carbon Storage Technical Viability Approach Support (CS TVA) Tool displays spatial data availability for the many components of Geologic Carbon Storage (GCS) technical viability assessment (Creason al 2025). Identifying sites suitable for GCS requires evaluating the intersection of myriad factors, including reservoir conditions, subsurface and surface hazards, infrastructure requirements, and energy community metrics. The technical viability of a site can only be confirmed for instances where all these factors have data available, and where those data support viability. Additional Resources related to the Technical Viability Assessment Tool: - Julia Mulhern, Casey White, Araceli Lara, Neyda Cordero Rodriguez, Zachary Jackson, Jacob Shay, Gabriel Creason, MacKenzie Mark-Moser, Paige Morkner, Kelly Rose, Carbon Storage Technical Viability Approach (CS TVA) Database, 3/26/2025, https://edx.netl.doe.gov/dataset/edx4ccs-carbon-storage-technical-viability-approach-database , DOI:10.18141/1984655 - Julia Mulhern, MacKenzie Mark-Moser, Gabriel Creason, Casey White, Araceli Lara, Neyda Cordero Rodriguez, Zach Jackson, Paige Morkner, Kelly Rose, Carbon Storage Technical Viability Approach (CS TVA) Matrix, 3/27/2025, https://edx.netl.doe.gov/dataset/carbon-storage-technical-viability-approach-cs-tva-matrix , DOI: 10.18141/2539979 - Gabriel Creason, Zach Jackson, Neyda Cordero Rodriguez, Julia Mulhern, Casey White, Araceli Lara, MacKenzie Mark-Moser, Paige Morkner, Kelly Rose, Carbon Storage Technical Viability Approach (CS TVA) Data Availability Results Database, 3/27/2025, https://edx.netl.doe.gov/dataset/carbon-storage-technical-viability-approach-cs-tva-data-availability-results-database, DOI:10.18141/2538557

Carbon storage↗

Recommendations for developing, documenting, and distributing data products derived from NEON data

The National Ecological Observatory Network (NEON) provides over 180 distinct data products from 81 sites (47 terrestrial and 34 freshwater aquatic sites) within the United States and Puerto Rico. These data products include both field and remote sensing data collected using standardized protocols and sampling schema, with centralized quality assurance and quality control (QA/QC) provided by NEON staff. Such breadth of data creates opportunities for the research community to extend basic and applied research while also extending the impact and reach of NEON data through the creation of derived data products—higher level data products derived by the user community from NEON data. Derived data products are curated, documented, reproducibly-generated datasets created by applying various processing steps to one or more lower level data products—including interpolation, extrapolation, integration, statistical analysis, modeling, or transformations. Derived data products directly benefit the research community and increase the impact of NEON data by broadening the size and diversity of the user base, decreasing the time and effort needed for working with NEON data, providing primary research foci through the development via the derivation process, and helping users address multidisciplinary questions. Creating derived data products also promotes personal career advancement to those involved through publications, citations, and future grant proposals. However, the creation of derived data products is a nontrivial task. Here we provide an overview of the process of creating derived data products while outlining the advantages, challenges, and major considerations.

54 ENVIRONMENTAL SCIENCES↗

Towards a RAG-based summarization for the Electron Ion Collider

Abstract The complexity and sheer volume of information — encompassing documents, papers, data, and other resources — from large-scale experiments demand significant time and effort to navigate, making the task of accessing and utilizing these varied forms of information daunting, particularly for new collaborators and early-career scientists.To tackle this issue, a Retrieval Augmented Generation (RAG)-based Summarization AI for EIC (RAGS4EIC) is under development. This AI-Agent not only condenses information but also effectively references relevant responses, offering substantial advantages for collaborators. Our project involves a two-step approach: first, querying a comprehensive vector database containing all pertinent experiment information; second, utilizing a Large Language Model (LLM) to generate concise summaries enriched with citations based on user queries and retrieved data. We describe the evaluation methods that use RAG assessments (RAGAs) scoring mechanisms to assess the effectiveness of responses. Furthermore, we describe the concept of prompt template based instruction-tuning which provides flexibility and accuracy in summarization. Importantly, the implementation relies on LangChain [1], which serves as the foundation of our entire workflow. This integration ensures efficiency and scalability, facilitating smooth deployment and accessibility for various user groups within the Electron Ion Collider (EIC) community. This innovative AI-driven framework not only simplifies the understanding of vast datasets but also encourages collaborative participation, thereby empowering researchers. As a demonstration, a web application has been developed to explain each stage of the RAG Agent development in detail. The application can be accessed athttps://rags4eic-ai4eic.streamlit.app.[A tagged version of the source code can be found inhttps://github.com/ai4eic/EIC-RAG-Project/releases/tag/AI4EIC2023_PROCEEDING.]

Instruments & Instrumentation↗

Deep-learning methods for contrast enhancement and artifact reduction in cryo-electron tomography: a systematic analysis of the state of the art and proposed improvements

Cryo-electron tomography (cryo-ET) has emerged as the preferred technique for visualizing the organization of macromolecular complexes in situ and resolving their structures at subnanometre resolution [Tegunov et al. (2021)View full citation, Nat. Methods, 18, 186–193]. Despite improvements in data quality as a result of advances in detector technology, microscope stability and stage precision, the analysis and interpretation of tomograms remains challenging due to a low signal-to-noise ratio and reconstruction artifacts stemming from experimental constraints in specimen tilt during data collection resulting in a missing wedge in the Fourier space. Recently, self-supervised deep-learning methods have been proposed for contrast enhancement and reduction of resolution anisotropy in reconstructed tomograms. Here, we evaluate several state-of-the-art deep-learning methods which aim to improve the interpretability of cryo-ET reconstructions, with a focus on their performance on downstream tasks of template matching, sub­tomogram averaging and segmentation. We propose new training architectures and a loss function based on Fourier shell correlation that show improved performance over the standard U-Net with L1/L2 losses. We demonstrate our analysis on four diverse experimental datasets: purified 80S ribosomes, in situ Chlamydomonas reinhardtii, immature HIV-1 virus-like particles and INS-1E cells.

contrast enhancement↗