Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “metadata extraction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

NEPATEC v2.0: Standardized Metadata and Text Corpus of National Environmental Policy Act Documents

The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for organizing and sharing NEPA-related information across agencies. Much of the information gathered during NEPA reviews is written into documents such as categorical exclusions, environmental assessments, and environmental impact statements, then filed in predominately independent agency file stores that may or may not be publicly accessible. The application of metadata and data standards, such as those recommended by the Council on Environmental Quality (CEQ), to NEPA documents offers a shared vocabulary and structure for key entities like projects, processes, and documents that can streamline information exchange and enhance collaboration across systems. In this work, we publicly release NEPATEC2.0, an expanded corpus of NEPA documents with associated metadata. NEPATEC2.0 encompasses approximately 120,000 documents from 60,000 projects prepared by more than 60 different agencies. Modeled to align with CEQ metadata standards, NEPATEC2.0 promotes consistency in environmental reviews and supports the ongoing effort to modernize permitting technologies by facilitating more transparent, efficient, and data-driven decision-making. Importantly, NEPATEC2.0 demonstrates the possibilities and limitations of large language model-based prompting to extract information from NEPA documents at scale.

54 ENVIRONMENTAL SCIENCES↗

Interrelationships among methods of estimating microbial biomass across multiple soil orders and biomes: Supporting data

This dataset contains environmental and soil measurements from 18 different locations across the globe including the SPRUCE experiment site and multiple sampling depths, with 17 of these locations having samples processed between 2012-2013 and one location (SPRUCE) collected in 2021 and processed in 2022. Environmental measurements include: mean annual temperature, mean annual precipitation, and 30-day presampling temperature. Soil physicochemical measurements include: particle size analysis (PSA), pH, gravimetric moisture content (GMC), bulk soil carbon (C) and nitrogen (N), total organic C and N, C:N ratio, and dissolved organic carbon (DOC). Soil biological measurements include: microbial biomass carbon (MBC) measured through chloroform fumigation extraction (CFE), gene copy numbers (GCN) of bacteria, fungi, and archaea measured through quantitative polymerase chain reaction (qPCR), DNA yield measured through Nanodrop spectrophotometry, and phospholipid fatty acids (PLFA) of bacteria and fungi measured through PLFA analysis. This data set contains one file in comma separate (*.csv) format.

archaea gene copy number↗

An Automated Approach to Labelling Datasets in Earth Science Publications

NASA Data Active Archive Centers, orDAACs, ingest, store, and distribute dataacquired from satellites, ground systems as well asreanalysis models. Many authors use this datain their research. However, most of the datasets usedin Earth Science Publications are not citedcorrectly or not cited at all. Thus, there is no directlink between the datasets used and thescientific publications which reference them. Thisleads to issues with reproducibility of theresults, attribution of the research results, anddiscovery of new datasets. This project began byexploring various methods of automatically labellingGoddard Earth Sciences Data andInformation Services Center (GES DISC) datasets usingSupervised Machine Learning and EarthData Search Common Metadata Repository (CMR) queries.The ultimate goal was to create alibrary of citations that utilized automated citationlabeling to directly link the researchpublications to the data they use. Supervised MachineLearning approaches struggled due to thelimited amount of labelled training data to learnfrom. Increasing the volume of training data isdifficult as it requires subject matter experts todevote time to manually reviewing journalarticles and determining the datasets used. The CMRqueries were inconsistent because theunderlying metadata is continuously being updated.Thus, it is hard to generalize theeffectiveness of the CMR results as they are dependenton the internal state of CMR. Theseapproaches helped inform the decision to transitionthe project into using a Knowledge Graph.Another key aspect of this project focused on theautomated extraction of features (platform,instrument, variables, etc) and explicit citationsfrom within Earth Science Publications. Theseautomated extractions were used to classify researchpapers based on their platform/instrumentcouples. This information was input into the CitationManagement System for GES DISC. Theseplatform/instrument couples also provide an additionalfacet that can be searched on the GESDISC website.

Edward Jahoda↗

Changuinola peat soil characteristics and gas emission raw data October 2019

This dataset comprises radiocarbon and geochemical measurements from peat and porewater samples collected across various depths at a site in Bocas del Toro, Panama. The study focuses on carbon cycling dynamics in tropical peatlands by examining carbon isotopic signatures (¹⁴C and ¹³C) and elemental compositions of bulk peat, dissolved organic carbon (DOC), carbon dioxide (CO₂), and methane (CH₄). Key parameters include radiocarbon ages and isotopic ratios (δ¹³C) of bulk peat, concentrations of carbon (%C) and nitrogen (%N), and radiocarbon content of porewater gases and dissolved organic carbon (DOC). The data provide insights into the vertical and spatial distribution of carbon sources and possible preservation and decomposition processes within tropical peat profiles, offering critical information for understanding carbon storage and greenhouse gas emissions in these ecosystems.This dataset is comprised of one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) carbon isotopic signatures (¹⁴C and ¹³C); (5) concentrations of carbon (%C) and nitrogen (%N); (6) radiocarbon content of porewater carbon dioxide (CO₂), and methane (CH₄) ; (7) porewater DOC; (8) bulk peat sampling protocol; (9) porewater sampling protocol; (10) porewater gas collection methods; and (11) gas extraction methods. All files are in .csv format and can be opened with any software that supports this file types.

54 ENVIRONMENTAL SCIENCES↗

1H-NMR characterization of soil dissolved organic matter from soil samples in control and warming plots in Blodgett Forest, CA (2014 and 2018)

The pathways of carbon transport and loss through and from soils—soil organic matter (SOM) depolymerization to dissolved organic carbon and mineralization to carbon dioxide (CO2)—are fundamentally driven by microbial activity, which is strongly regulated by environmental conditions. As part of Lawrence Berkeley National Laboratory Terrestrial Ecosystem Science Belowground Biogeochemistry Science Focus Area (SFA), we have established a novel whole-soil long-term warming experiment at the University of California (UC) Blodgett Forest Research Station (Sierra Nevada) in 2014, where we study the role of biogeochemical, microbial and geochemical process interactions in SOM (soil organic matter) decomposition and stabilization. This package contains metabolite data obtained through 1H nuclear magnetic resonance (NMR) spectroscopy on water-extracted soils. Soil samples were collected in 2014/06/03 and 2018/06/04 from 3 replicated paired plots that had been subjected to experimental warming since June 2014 to simulate a predicted climate change scenario for northern California. The following files are included: (1) nmr_h2o_data_raw.csv: raw data, (2) nmr_h2o_data_processed.csv: computed compound concentrations and metadata, (3) nmr_h2o_compound_metadata.csv: compound metadata, (4) nmr_h2o_sample_metadata.csv: sample metadata

1H-NMR (nucleic magnetic resonance) spectroscopy↗

FTICR-MS Data from Multi-continent River Water and Sediment and from Coastal River Fresh and Saline Sediment Associated with: “Dissolved Organic Matter Functional Trait Relationships are Conserved Across Rivers”

This data package is associated with the publication “Dissolved Organic Matter Functional Trait Relationships are Conserved Across Rivers” submitted to PNAS (Stegen et al., 2023). The study aims to understand large-scale spatial structure of the dissolved organic matter (DOM) thermodynamic traits and inter-trait relationships by investigating (1) river water and sediments collected along 97 rivers spanning 3 continents and (2) coastal sediment collected from fresh and saline locations in Pacific and Gulf/Atlantic rivers. Sediment extracts and water samples were analyzed using ultrahigh resolution Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS). This dataset is comprised of three folders (1) Coastal, (2) WHONDR_S19S, and (3) Data_Dictionaries. Coastal contains (1) a subfolder with processed FTICR-MS data as csv files and sample collection metadata, (2) a subfolder with R scripts used to process the data and create associated figures, (3) a subfolder with the raw, unprocessed FTICR-MS data as .xml files, and (4) a readme file with more information about the dataset and instructions for using Formularity (https://omics.pnl.gov/software/formularity). WHONDRS_S19S contains (1) a csv file with processed FTICR data, (2) a csv with sample collection metadata, (3) a csv with sample geospatial data, (4) a csv with simulated lambda model outputs, (5) a subfolder with R scripts used to process the data and create associated figures, and (6) a readme file with more information regarding WHONDRS raw FTICR data and processing scripts. Data_Dictionaries contains data dictionaries for each csv file in the data package. The 97 global river corridors were part of a WHONDRS (https://whondrs.pnnl.gov) study. The raw, unprocessed FTICR-MS data with additional data can be found at doi:10.15485/1729719 for sediments and doi:10.15485/1603775 for water. This data package contains the processed data used in the associated manuscript. The coastal data has not been previously published, and this data package contains both the raw and processed data. Version 3 of this data package published February 2023 includes updates to the title of the manuscript, additional data and data dictionary and updated scripts linked to new analysis.

54 ENVIRONMENTAL SCIENCES↗

Carbon Storage Technical Viability Approach (CS TVA) Database

The Carbon Storage Technical Viability Approach (CS TVA) database was developed to support the implementation of the CS TVA Matrix to a national data availability assessment for technically viable carbon storage. This database leverages the efforts of multiple adjacent and overlapping databases by non-redundantly combining the databases into a single database along with additionally providing tags facilitating the CS TVA. The non-redundant aspect of the database permits an accurate assessment of the concentration of available data, aiding in spatial and categorical data gaps analysis relative to the individual CS TVA Matrix Components. Version 2.0 of the database is an expansion of Version 1.0. Version 2.0 was created to include additional data gathered to fill gaps in the existing data set. Downloading the CS TVA v2.0 database will result in two separate databases, the version 1.0 original .gdb, and a second addendum .gdb with the new data gathered, together these two databases make up v2.0. Please see the ReadMe file below for full details, metadata information, use disclaimer, and attributions.

Coal↗

SPRUCE Quantitative PCR (qPCR) of Microbial Gene Copy Numbers, 2021-2022

This dataset provides the results for quantitative polymerase chain reaction (qPCR) of peat samples collected from ambient and experimental plots in the Spruce and Peatland Responses Under Climatic and Environmental Change (SPRUCE) experiment site in June and August of 2021, and June of 2022. SPRUCE is located within the Marcell Experimental Forest in northern Minnesota, USA. The dataset includes bacterial, archaeal, fungal gene copy numbers, along with corresponding logarithmic values, at 11 depth increments of two-meter deep peat cores taken from 12 sampling sites locations inside SPRUCE plots (10 chambered and 2 ambient plots). The sampling, sample prep and analysis followed standard methods outlined in prior publications (Wilson et al. 2016; Kluber et al. 2020) except that a higher yielding Omega Bio-Tek Mag-Bind Environmental DNA 96 Kit was used for extractions and DNA was quantified using Qubit dsDNA High Sensitivity Assay Kit. qPCR subsamples of peat cores from the SPRUCE plots characterize changes in the abundance and composition of microbial communities of peat seasonally showing how composition varies under multiple levels of experimental peat warming and atmospheric CO2 concentrations. This dataset contains one data file in comma-separated values (.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma-separated values (.csv) format and a user guide in PDF (*.pdf) format. On 2026-07-07 this dataset was updated to add three columns to the data file: ‘Fungal_copy_dry’, ‘Log_fungal_copy_dry’, ‘Fungal_copy_wet’. No previously released data values were altered. Additionally, the abstract, data dictionary, and user guide were updated, and a file-level metadata file was added.

Archaea↗

NASA Taxonomies for Searching Problem Reports and FMEAs

Many types of hazard and risk analyses are used during the life cycle of complex systems, including Failure Modes and Effects Analysis (FMEA), Hazard Analysis, Fault Tree and Event Tree Analysis, Probabilistic Risk Assessment, Reliability Analysis and analysis of Problem Reporting and Corrective Action (PRACA) databases. The success of these methods depends on the availability of input data and the analysts knowledge. Standard nomenclature can increase the reusability of hazard, risk and problem data. When nomenclature in the source texts is not standard, taxonomies with mapping words (sets of rough synonyms) can be combined with semantic search to identify items and tag them with metadata based on a rich standard nomenclature. Semantic search uses word meanings in the context of parsed phrases to find matches. The NASA taxonomies provide the word meanings. Spacecraft taxonomies and ontologies (generalization hierarchies with attributes and relationships, based on terms meanings) are being developed for types of subsystems, functions, entities, hazards and failures. The ontologies are broad and general, covering hardware, software and human systems. Semantic search of Space Station texts was used to validate and extend the taxonomies. The taxonomies have also been used to extract system connectivity (interaction) models and functions from requirements text. Now the Reconciler semantic search tool and the taxonomies are being applied to improve search in the Space Shuttle PRACA database, to discover recurring patterns of failure. Usual methods of string search and keyword search fall short because the entries are terse and have numerous shortcuts (irregular abbreviations, nonstandard acronyms, cryptic codes) and modifier words cannot be used in sentence context to refine the search. The limited and fixed FMEA categories associated with the entries do not make the fine distinctions needed in the search. The approach assigns PRACA report titles to problem classes in the taxonomy. Each ontology class includes mapping words - near-synonyms naming different manifestations of that problem class. The mapping words for Problems, Entities and Functions are converted to a canonical form plus any of a small set of modifier words (e.g. non-uniformity NOT + UNIFORM.) The report titles are parsed as sentences if possible, or treated as a flat sequence of word tokens if parsing fails. When canonical forms in the title match mapping words, the PRACA entry is associated with the corresponding Problem, Entity or Function in the ontology. The user can search for types of failures associated with types of equipment, clustering by type of problem (e.g., all bearings found with problems of being uneven: rough, irregular, gritty ). The results could also be used for tagging PRACA report entries with rich metadata. This approach could also be applied to searching and tagging failure modes, failure effects and mitigations in FMEAs. In the pilot work, parsing 52K+ truncated titles (the test cases that were available), has resulted in identification of both a type of equipment and type of problem in about 75% of the cases. The results are displayed in a manner analogous to Google search results. The effort has also led to the enrichment of the taxonomy, adding some new categories and many new mapping words. Further work would make enhancements that have been identified for improving the clustering and further reducing the false alarm rate. (In searching for recurring problems, good clustering is more important than reducing false alarms). Searching complete PRACA reports should lead to immediate improvement.

Malin, Jane T.↗

NIST: Soil Respiration, Moisture, Temperature, Chemistry; and Fine Root Measurements from a Transect Through a Forest Edge, Gaithersburg, Maryland, 2017-2021

This dataset contains soil respiration, moisture, temperature, and chemistry, as well as fine root measurements from the National Institute of Standards and Technology (NIST) Forested Optical Reference for Evaluating Sensor Technology (FOREST) research facility at Gaithersburg, Maryland. Measurements were taken at an existing transect array that begins in a grassy meadow, crosses a sharp forest edge, then a small stream, and finally extends upwards in the interior of the forest at the top of a ridge. There are 6 different landscape positions replicated across three transects in the array. Soil respiration was measured during growing seasons in 2017-2019 (2017-06-02 to 2020-02-27). Pedons (1 m3) were isolated from surrounding tree roots using trenching and a fabric to inhibit root ingrowth. Flux measurements inside the pedons were thus assumed to represent heterotrophic only respiration in 2019, and these fluxes were paired with nearby fluxes assumed to represent total respiration. Deep vertical probes measured volumetric moisture content and temperature at the same points in the array every 10 cm in depth to either 90 cm or 120 cm total depth, at 15 minute intervals, from 2019-2021 (2019-07-09 to 2021-09-10). Soil core samples were collected from each of the array points for three different months in early- to mid-2019 (2019-03-19 to 2019-07-10), at three depths each. Soils were analyzed for gravimetric moisture content; pH; total carbon, nitrogen, and phosphorus; texture; microbial biomass carbon, nitrogen, and phosphorus; extractable dissolved organic carbon, nitrogen, and phosphorus; extractable nitrate and ammonia; and extracellular hydrolytic enzyme activities. The fine roots were separated from the cores and segregated by plant functional type (grass or tree species) and if they were dead or alive. Fine roots were then measured for length, surface area, diameter, and dry mass. This dataset contains four data files in comma separated (*.csv) format. These data serve to deepen our understanding of root and soil processes at forest edges and in transitional zones.

54 ENVIRONMENTAL SCIENCES↗

Organic matter concentration and composition of experimentally burned open air and muffle furnace vegetation chars across differing burn severity and feedstock types from Pacific Northwest, USA (v4).

This dataset represents results from an experimental study designed to compare how the chemical composition of organic matter changes across different burn conditions and feedstock materials. The dataset provides both solid and dissolved phase bulk concentration and organic matter characterization data from experimentally generated chars. Chars were created in a closed muffle furnace or on an open burn table from four different feedstock species representing vegetation commonly impacted by fire regimes across the Pacific Northwest, USA. This data can be used to compare how different burn conditions may influence resultant organic matter chemistry and help further our understanding of potential biogeochemical impacts on river corridors post-fire. This dataset is comprised of one data package readme, one data dictionary (dd), one file level metadata (flmd), fourteen burn table videos, burn table video metadata and three folders containing (A) data; (B) metadata and protocols; and (C) photos. The folder names and the file name of the data package readme include a version number which will be updated with future iterations of this data package. The data folder includes (1) solid carbon and solid nitrogen; (2) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) and total dissolved nitrogen (TN); (3) pH; (4) thermocouple time series temperature; (5) methods codes; (6) installation methods; (7) excitation emissions matrix (EEM) methods information; (8) a folder of excitation emissions matrix (EEM) fluorescence and absorbance spectra in dissolved organic matter and EEMs processing instructions; (9) solid state carbon-13 and solution state phosphorus nuclear magnetic resonance (13-C NMR and 31-P NMR) data and methods; (10) benzene polycarboxylic acid (BPCA) concentration and stable isotope data; (11) FTICR-MS methods; (12) Inductively coupled plasma (ICP) data for total calcium, magnesium, iron, aluminum, potassium, phosphorus, sodium, and sulfur along with sodium hydroxide-ethylenediaminetetraacetic acid (sodium hydroxide-EDTA) extractable calcium, magnesium, iron, aluminum, potassium, phosphorus, and sulfur; (13) a folder of phosphorus, carbon, and nitrogen X-ray absorption near edge structure (P-XANES, N-XANES, C-XANES) data for samples and standards; (14) P-XANES, N-XANES, C-XANES methods; (15) molybdate reactive phosphorus; and (16) folder of high resolution characterization of organic matter via 21 Tesla Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) generated through the Environmental Molecular Sciences Laboratory (EMSL; https://www.pnnl.gov/environmental-molecular-sciences-laboratory). The FTICR folder contains .txt data files and a subfolder containing instruments for using Formularity (https://omics.pnl.gov/software/formularity) and an R script to process the data based on the user's specific needs. The metadata and protocols folder includes (1) international geo-sample number (IGSN) mapping file (2) burn and laboratory metadata; (3) burn protocol; (4) laboratory protocol; (5) vegetation collection metadata; and (6) vegetation collection protocol. The folder contains photos of the solid chars. All files are .csv, .txt, .pdf, .jpg, .jpeg, .R, .ref, or .mp4. The data package was originally published October 2022 (v1). It was updated April 2023 (v2; new data files), September 2023 (v3; new and corrected data files), and September 2024 (v4; new and added/updated files). Metadata files were also updated to reflect these changes. See the change history section in the readme for more details.

54 ENVIRONMENTAL SCIENCES↗

GMI-IPS: Python Processing Software for Aircraft Campaigns

NASA's Atmospheric Tomography Mission (ATom) seeks to understand the impact of anthropogenic air pollution on gases in the Earth's atmosphere. Four flight campaigns are being deployed on a seasonal basis to establish a continuous global-scale data set intended to improve the representation of chemically reactive gases in global atmospheric chemistry models. The Global Modeling Initiative (GMI), is creating chemical transport simulations on a global scale for each of the ATom flight campaigns. To meet the computational demands required to translate the GMI simulation data to grids associated with the flights from the ATom campaigns, the GMI ICARTT Processing Software (GMI-IPS) has been developed and is providing key functionality for data processing and analysis in this ongoing effort. The GMI-IPS is written in Python and provides computational kernels for data interpolation and visualization tasks on GMI simulation data. A key feature of the GMI-IPS, is its ability to read ICARTT files, a text-based file format for airborne instrument data, and extract the required flight information that defines regional and temporal grid parameters associated with an ATom flight. Perhaps most importantly, the GMI-IPS creates ICARTT files containing GMI simulated data, which are used in collaboration with ATom instrument teams and other modeling groups. The initial main task of the GMI-IPS is to interpolate GMI model data to the finer temporal resolution (1-10 seconds) of a given flight. The model data includes basic fields such as temperature and pressure, but the main focus of this effort is to provide species concentrations of chemical gases for ATom flights. The software, which uses parallel computation techniques for data intensive tasks, linearly interpolates each of the model fields to the time resolution of the flight. The temporally interpolated data is then saved to disk, and is used to create additional derived quantities. In order to translate the GMI model data to the spatial grid of the flight path as defined by the pressure, latitude, and longitude points at each flight time record, a weighted average is then calculated from the nearest neighbors in two dimensions (latitude, longitude). Using SciPya's Regular Grid Interpolator, interpolation functions are generated for the GMI model grid and the calculated weighted averages. The flight path points are then extracted from the ATom ICARTT instrument file, and are sent to the multi-dimensional interpolating functions to generate GMI field quantities along the spatial path of the flight. The interpolated field quantities are then written to a ICARTT data file, which is stored for further manipulation. The GMI-IPS is aware of a generic ATom ICARTT header format, containing basic information for all flight campaigns. The GMI-IPS includes logic to edit metadata for the derived field quantities, as well as modify the generic header data such as processing dates and associated instrument files. The ICARTT interpolated data is then appended to the modified header data, and the ICARTT processing is complete for the given flight and ready for collaboration. The output ICARTT data adheres to the ICARTT file format standards V1.1. The visualization component of the GMI-IPS uses Matplotlib extensively and has several functions ranging in complexity. First, it creates a model background curtain for the flight (time versus model eta levels) with the interpolated flight data superimposed on the curtain. Secondly, it creates a time-series plot of the interpolated flight data. Lastly, the visualization component creates averaged 2D model slices (longitude versus latitude) with overlaid flight track circles at key pressure levels. The GMI-IPS consists of a handful of classes and supporting functionality that have been generalized to be compatible with any ICARTT file that adheres to the base class definition. The base class represents a generic ICARTT entry, only defining a single time entry and 3D spatial positioning parameters. Other classes inherit from this base class; several classes for input ICARTT instrument files, which contain the necessary flight positioning information as a basis for data processing, as well as other classes for output ICARTT files, which contain the interpolated model data. Utility classes provide functionality for routine procedures such as: comparing field names among ICARTT files, reading ICARTT entries from a data file and storing them in data structures, and returning a reduced spatial grid based on a collection of ICARTT entries. Although the GMI-IPS is compatible with GMI model data, it can be adapted with reasonable effort for any simulation that creates Hierarchical Data Format (HDF) files. The same can be said of its adaptability to ICARTT files outside of the context of the ATom mission. The GMI-IPS contains just under 30,000 lines of code, eight classes, and a dozen drivers and utility programs. It is maintained with GIT source code management and has been used to deliver processed GMI model data for the ATom campaigns that have taken place to date.

Damon, M. R.↗

Restoration and Reexamination of Data from the Apollo 11, 12, 14, and 15 Dust, Thermal and Radiation Engineering Measurements Experiments

As part of an effort by the Lunar Data Node (LDN) we are restoring data returned by the Apollo Dust, Thermal, and Radiation Engineering Measurements (DTREM) packages emplaced on the lunar surface by the crews of Apollo 11, 12, 14, and 15. Also commonly known as the Dust Detector experiments, the DTREM packages measured the outputs of exposed solar cells and thermistors over time. They operated on the surface for up to nearly 8 years, returning data every 54 seconds. The Apollo 11 DTREM was part of the Early Apollo Surface Experiments Package (EASEP), and operated for a few months as planned following emplacement in July 1969. The Apollo 12, 14, and 15 DTREMs were mounted on the central station as part of the Apollo Lunar Surface Experiments Package (ALSEP) and operated from deployment until ALSEP shutdown in September 1977. The objective of the DTREM experiments was to determine the effects of lunar and meteoric dust, thermal stresses, and radiation exposure on solar cells. The LDN, part of the Geosciences Node of the Planetary Data System (PDS), operates out of the National Space Science Data Center (NSSDC) at Goddard Space Flight Center. The goal of the LDN is to extract lunar data stored on older media and/or in obsolete formats, restore the data into a usable digital format, and archive the data with PDS and NSSDC. For the DTREM data we plan to recover the raw telemetry, translate the raw counts into appropriate output units, and then apply calibrations. The final archived data will include the raw, translated, and calibrated data and the associated conversion tables produced from the microfilm, as well as ancillary supporting data (metadata) packaged in PDS format.

McBride, Marie J.↗

Compilation of Experimental Yield Data for Spontaneous Fission of 252 Cf

We present a comprehensive compilation and curation of experimental fission yield (FY) data for the spontaneous fission of 252 Cf, extracted from the EXFOR database. The compilation follows a structured methodology developed for prior compilations of neutron-induced fission yields, and incorporates both independent (IFY) and cumulative (CFY) yields. A total of 62 datasets were reviewed, with entries spanning from 1955 to 2021. A significant portion of the literature reports pre-neutron emission yields, which were excluded from the present compilation due to limitations in format compatibility. Each accepted dataset was processed into a standardized JSON format, including metadata, uncertainties, and bibliographic references. Where available, decay radiation information was used to update the FY data using the latest ENSDF evaluations; 237 data points were corrected accordingly. These corrections are fully traceable and preserve original values. The result is a curated dataset suitable for use in nuclear data evaluations. This work is part of an ongoing effort to modernize the handling of FY data and provide evaluators with high-quality, machine-readable experimental inputs

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Mapping use cases and dataset needs for benchmarking buildings data

A perennial challenge in buildings research is the lack of high-quality datasets that can be relied upon for a wide array of tasks, including model calibration and improving energy efficiency and load flexibility. Instrumenting a building for data collection is resource intensive, so it is important to be methodical in the approach and ensure that resulting data are flexible and useful for a broad range of analyses. This study aims to fill the gaps in characterizing potential use cases for buildings datasets and mapping them to dataset needs using a well-defined data infrastructure. Here, we have developed a systematic mapping strategy between buildings dataset needs and use cases to help streamline the processes of efficiently targeting datasets, designing building sensing systems, and determining buildings research use cases. We selected 14 prospective use cases and 11 refined buildings data categories for developing the preliminary dataset-needs-to-use-cases mapping matrix (‘DN-UC mapping matrix’) with generic ‘Tags’—a detailed sub-level of data categories extracted by justifying the needs of an aspect of the datasets to use cases. We present two example applications of the developed mapping matrix to demonstrate use of the mapping matrix and its effectiveness.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

SPASE: Current Uses, Tools, and Plans

The Space Physics Archive Search and Extract (SPASE) project is an international collaboration among Heliophysics (solar and space physics) groups concerned with data acquisition and archiving. Within this community there are a variety of old and new data centers, resident archives, "virtual observatories", etc. acquiring, holding, and distributing data. The main product of the SPASE group is an XML-based SPASE Data Model now in operational use to enable searches for and ultimate acquisition of data of interest to a researcher. The SPASE Data Model defines the content of resource descriptions (metadata). The intent is to describe all SCientifically usable Heliophysics data sets using the Data Model. Another product of the SPASE group, in collaboration with NASA's Virtual Observatories, is a set of tools and services which work with SPASE meta data. This includes Registry Services which can retrieve and render metadata using resource identifiers and facilitate the downloading of the data referenced by the meta data. The SPASE Data Model has also been used as a vocabulary in specialized data models. One example is the Heliophysics Event List Manager (HELM) model. The SPASE Data Model is also being expanded to provide the means for more detailed description of data sets with the aim of enabling more automated ingestion and use of the data through detailed format descriptions. The evolution is based on a number of lessons learned and feedback from our community. Some of the lessons learned are unique to Heliophysics, and some are common to the various data diSCiplines. We will discuss the present state of SPASE usage, the role the SPASE Data Model can play in speCialized data models and how we foresee the development direction in the future.

Thieman, J. R.↗

Artificial intelligence models, photos, and data associated with the manuscript “Quantifying Streambed Grain Size, Uncertainty, and Hydrobiogeochemical Parameters Using Machine Learning Model YOLO” (v2)

This data package is associated with the manuscript “Quantifying Streambed Grain Size, Uncertainty, and Hydrobiogeochemical Parameters Using Machine Learning Model YOLO” published in Water Resources Research (Chen et al., 2024). This data package includes the training, validation, testing, and prediction data used by the artificial intelligence (AI) model for automated grain size and hydro-biogeochemistry quantification using streambed photos. The grain size data are extracted for each photo using You Look Only Once (YOLO), a pre-trained object detection model. This data package was originally published in October 2023. It was updated August 2025 (v2; new and modified files). File and folder names were not revised to indicate changes. See the change history section in the readme for more details. Please see flmd.csv for a list of all files contained in this data package and descriptions for each. Please see dd.csv for a data dictionary that defines the column headers of .csv files in the data package. This dataset is comprised of one data folder containing (1) file-level metadata; (2) data dictionary; (3) readme; and (4) six subfolders. Subfolders 1 to 4 include the training, validation, testing, and prediction data. Subfolder 5_Summary includes the summary results of different combinations of training, validation, testing, and prediction data. Subfolder 6_SupplementalData includes additional data downloaded from public sources (Kaufman et al., 2023a; Kaufman et al., 2023b; Garefalakis et al., 2023; Mair et al., 2024; https://github.com/river-corridors-sfa/Geospatial_variables). In total, the data package includes 110 folders and 44,283 files. These files include 9,047 .jpg photos, 1 .png photo, 3 .tif photos; 26,639 photo labels and individual grain sizes and probability from AI (.txt); 8,447 grain size distribution data (.dat); and 126 CSV files for results summary, and 14 required metadata files (.xlsx). The summary CSV files contain 68 columns and approximately 2,200 rows that represent photo names, site locations, recording time, GPS coordinates, grains sizes (D10, D50, D60, and D84), number of grains, and additional hydro-biogeochemical data such as water depth, flow velocity, Manning’s coefficient, friction factor, hydraulic conductivity, permeability, streambed interstitial velocity magnitude, mass transfer rate, and nitrate uptake velocity. The photos were obtained from 75 sites in the Yakima River Basin and the Columbia River shorelines, and other associated data from samples and sensors obtained when the photos were taken are publicly available (Fulton et al. 2022; Grieger et al. 2023). All files are .csv, .txt, .dat, .jpg, or .pdf. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

The IPAC Image Subtraction and Discovery Pipeline for the Intermediate Palomar Transient Factory

We describe the near real-time transient-source discovery engine for the intermediate Palomar Transient Factory (iPTF), currently in operations at the Infrared Processing and Analysis Center (IPAC), Caltech. We coin this system the IPAC/iPTF Discovery Engine (or IDE). We review the algorithms used for PSF-matching, image subtraction, detection, photometry, and machine-learned (ML) vetting of extracted transient candidates. We also review the performance of our ML classifier. For a limiting signal-to-noise ratio of 4 in relatively unconfused regions, bogus candidates from processing artifacts and imperfect image subtractions outnumber real transients by approximately equal to 10:1. This can be considerably higher for image data with inaccurate astrometric and/or PSF-matching solutions. Despite this occasionally high contamination rate, the ML classifier is able to identify real transients with an efficiency (or completeness) of approximately equal to 97% for a maximum tolerable false-positive rate of 1% when classifying raw candidates. All subtraction-image metrics, source features, ML probability-based real-bogus scores, contextual metadata from other surveys, and possible associations with known Solar System objects are stored in a relational database for retrieval by the various science working groups. We review our efforts in mitigating false-positives and our experience in optimizing the overall system in response to the multitude of science projects underway with iPTF.

methods: analytical – methods: data analysis –↗