Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Catalog”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

CO2 pipeline data catalog

This "Data Catalog" is an excel spreadsheet of collected and catalogued spatial and non-spatial public datasets in relevance to and support of CO2 transport (pipelines) infrastructure development and risk modeling. Includes data associated with CO2 sources, sinks, and transportation.

CO2↗

DEEPEN Data Catalog for Magmatic Geothermal Systems in the United States

This data catalog contains information related to the Training Site Analysis for the Geothermica project "DE-risking Exploration of geothermal Plays in magmatic ENvironments (DEEPEN)." The DEEPEN project aims to reduce exploration risk for geothermal fluids in magmatic systems by developing improved an improved framework for interpretation of exploration data using the Play Fairway Analysis (PFA) methodology. The Training Site Analysis performed for DEEPEN leverages existing datasets to develop a customized PFA approach to exploration for multiple geothermal resource types in magmatic systems (conventional hydrothermal resources, supercritical fluid and superheated steam resources, and superhot EGS resources). This data catalog contains links to publicly available data files related to 8 training sites in the United States. US training sites are: the Cascades/Aleutians PFA project; the Hawaii PFA project, the Oregon Cascades PFA project, the Snake River Plain, Idaho PFA project, the Washington State PFA project, Newberry Volcano, Coso Geothermal Field, and the Geysers Geothermal field. This database contains an overview of these training sites, data sources, and links to publicly available exploration datasets. For the five PFA projects, details on exploration data related to PFA components (heat, fluid, permeability, sometimes seal) are provided, including a summary of data weighting methodologies.

15 GEOTHERMAL ENERGY↗

Evaluation of Data Catalog Software for Hanford Site Environmental Datasets

Environmental information and data underpin achievement of the U.S. Department of Energy (DOE) Office of Environmental Management (EM) mission at the Hanford Site. The Hanford Environmental Data Management (HEDM) Program is the DOE Richland Operations Office (RL) approach to develop and implement a formal program for managing environmental data and the associated records, materials, and systems at the Hanford Site. The current project, contract, organization, and contractor-specific efforts at managing environmental data sets are insufficient to provide orderly, long-term, site-wide access. A vital element to be created within the HEDM program plan is a catalog of data sources, called the Hanford Environmental Information and Data Index (HEIDI), that will enable long-term access and retrievability for the multiple independent sources of data that might otherwise be difficult to discover. This report compares leading open source and commercial data catalog platforms using criteria to assess the functionality needed to develop the HEIDI catalog of Hanford data sources that connects and exchanges data with established Hanford Local Area Network (HLAN) enterprise information technology systems. Proprietary platforms evaluated included ArcGIS Enterprise Sites, Junar, OpenDataSoft, and Socrata, and non-proprietary platforms included Energy Data eXchange (EDX), Comprehensive Knowledge Archive Network (CKAN), and DKAN (a Drupal-based open data portal based on CKAN). Capabilities supporting data discoverability, retrieval, and archival, as well as metadata standard requirements and integration into the HLAN were rated as either failing to meet requirements (F), meeting requirements (M), or exceeding requirements by delivering additional desired features (E). The lowest rating for any capability area was assigned as the overall rating for the platform. These findings enable DOE-RL and the contractors implementing the HEDM plan to focus on candidate tools likely to meet the requirements for implementing HEIDI. All of the platforms receiving an overall rating of ‘F’ were unable to be deployed on Hanford infrastructure or within dedicated cloud resources. A propriety software-as-a-service (SaaS) model of delivering a data catalog (e.g., found in software such as Junar and OpenDataSoft) favors consistency across customers at the expense of customization and configurable roles that are needed for Hanford work. Hosting data on a shared commercial platform places limits on dataset size (maximum of 240 Mb for OpenDataSoft), a significant limitation for HEIDI implementation. EDX, a government data catalog based on CKAN, received the ‘F’ rating due to an inability to incorporate authentication from HLAN into the system. Among platforms rated ‘M’ or ‘E’, only the Socrata platform had a SaaS delivery model. In contrast to other SaaS platforms, Socrata provided custom roles and gateways that allow local datasets to be incorporated into an online catalog. Socrata also complies with the Federal Risk and Authorization Management Program, a significant benefit for cloud-based management of Hanford data. The other platforms rated ‘M’ or ‘E’, ArcGIS Enterprise Sites, CKAN, and DKAN, provide fully self-hosted options, allowing for greater control and flexibility with the HEIDI catalog. These widely used tools have supportive communities of practice, extensive customization options, and demonstrated deployments that provide evidence that they can meet requirements, often deliver additional desired features, and work well with federal government systems. Completely customized alternatives built on a collection of applications were not evaluated because achieving similar performance to CKAN or DKAN requires substantial resources, especially in the absence of the active communities that have grown to support these tools. ArcGIS Enterprise Sites, Socrata, CKAN, and DKAN were evaluated as strong candidates for successful implementation with HEIDI.

54 ENVIRONMENTAL SCIENCES↗

A Prototype Software to Demonstrate a Data Catalog for Hanford Environmental Datasets

Ensuring that data on long-term environmental remediation at the Hanford Site is high-quality, traceable, and easily accessible is an ongoing challenge, complicated by decades of data collection, multiple contractors maintaining data sources, and the wide range of data types. A centralized data catalog, known as the Hanford Environmental Information and Data Index (HEIDI), has been under development as part of the Hanford Environmental Data Management (HEDM) program to address these challenges. HEIDI fulfills a critical need to bring together a wide range of data types and sizes from multiple authoritative data sources, while documenting the data pedigree and quality information (i.e., traceable to the data source/originator). This document describes additional development and maturation of the HEIDI prototype. Key accomplishments included deploying the catalog software, Esri Geoportal Server, on a server accessible to Hanford Local Area Network users, conducting cybersecurity evaluations, investigating integrated authentication solutions, and conducting functional testing of the catalog prototype. The server-based deployment enabled targeted feedback, leading to enhancements including improved accessibility features and an expanded metadata schema. Specifications for the server-based deployment of the prototype catalog and the HEIDI metadata schema are provided in this document to support subsequent HEIDI deployment by the U.S. Department of Energy Richland Operations Office.

54 ENVIRONMENTAL SCIENCES↗

Deliverable D12 – Distributed Wind Data Catalog Development Guide and Instruction Manual

Pacific Northwest National Laboratory and Technical University of Denmark completed this deliverable as part of Work Package 2: Data Information Catalog for Distributed Wind Research (WP2) for the International Energy Agency Wind Technology Collaboration Programme Task 41: Enabling Wind to Contribute to a Distributed Energy Future (IEA Wind Task 41). As the final deliverable for WP2, Deliverable D12 includes a data instruction guide for the IEA Wind Task 41 distributed wind data catalog. As such, this document includes: a step-by-step explanation of how the IEA Wind Task 41 data catalog was created, how it was populated, and how to use it; future options for the IEA Wind Task 41 data catalog, and a summary with recommendations for future work.

17 WIND ENERGY↗

HPDF Data Catalog and Lakehouse Demo Heuristic Evaluation and Roadshow Feedback Reports

The December 2025 roadshow gathered rapid feedback from community members on the proof of concept AmSC/HPDF data lakehouse and catalog solution 1. This provided the opportunity to demonstrate our current state of progress and gather input on workflows and AI agents. We recognize that the proofs of concept interfaces (OpenMetadata and Goose) are not intended for our end users so we have captured feedback to keep in mind as additional technical and conceptual design work is undertaken.

97 MATHEMATICS AND COMPUTING↗

Persistence Control of Engineered Functions in Complex Soil Microbiomes (PerCon SFA), Secure Biosystems Design Project Data Catalog at PNNL DataHub

The Persistence Control of Engineered Functions in Complex Soil Microbiomes Project (PerCon SFA) at Pacific Northwest National Laboratory (PNNL) is a Genomic Sciences Program Biosystems Design, Science Focus Area research project consortium. Collaborating across highly integrated institutions, PerCon SFA scientists are exploring how environmental niches can be sculpted using the mechanisms of genome reduction and metabolic addiction to drive secure rhizosphere community design for robust biomass cropping in challenging environments. The PerCon SFA DataHub project repository contains publication-relevant digital dataset and metadata DOI packages, enabling exploration and download of integrated experimental dataset catalogs publicly available to a global scientific community. and metadata repository allow for exploring and downloading integrated experimental biodesign omics dataset catalogs, including experimental protocols and/or workflows, raw and/or processed data, as required by the repository, and other relevant supporting materials and/or metadata required for research reproducibility and reporting.

59 BASIC BIOLOGICAL SCIENCES↗

Specifications and a Prototype Software to Demonstrate a Data Catalog for Hanford Datasets

Environmental management activities at the Hanford Site produce extensive data about site conditions, contaminants, and cleanup activities. Managing, archiving, and accessing that data requires a high degree of collaboration among site contractors and a high level of awareness by project managers and staff. The Hanford Site has a range of databases (e.g., Hanford Environmental Information System [HEIS]) and their associated user interfaces (e.g., Environmental Dashboard Application [EDA], Virtual Library [VL]), as well as other document management systems (e.g., Integrated Document Management System [IDMS]). However, Hanford lacks a single unified resource to find data (which itself comes in multiple formats) amongst the multiple disparate systems, not to mention ad hoc data not contained in an official repository/database.

54 ENVIRONMENTAL SCIENCES↗

Data Catalog Project - A Browsable, Searchable, Metadata System

Modern experiments are typically conducted by large, extended, where researchers rely on other team members to produce much of the data they use. The experiments record very large numbers of measurements which can be difficult for users to find, access and understand. We are developing a system for users to annotate their data products with structured metadata, providing data consumers with a discoverable, browsable data index. Machine understandable metadata captures the underlying semantics of the recorded data, which can then be consumed by both programs, and interactively by users. Collaborators can use these metadata to select and understand recorded measurements.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Datum: A Scientific Metadata Catalog

The data catalog market is currently flooded with a myriad of different products, but none serve the scientific community well. There are cloud-native tools like Databricks, Snowflake,to on-premise solutions like Collibra and Datahub. The common failing of all these tools however, is their inability to serve the scientific data community directly. Most catalogs are targeted towards financial, health, or user data - not sensor or scientific domain data. They also prioritize integrations that often don’t exist or are just starting to be used in the scientific realm - all while ignoring common scientific tools and file types. Datum is a catalog which targets the scientific data directly, including the tools and networks in which those tools are used. We work with the producers and consumers of the data where they are, targeting cloud and on-premise with a focus on classified networks. Datum is an Erlang/Elixir application. Technical Features Note: The features listed below are still under development and may change, slightly, upon final delivery of the product. File Formats - Datum has the ability to read additional metadata and provides processing pipelines for the following file formats: Plain Text, PDF, LaTeX, HTML, Open Document Format (.odt), XML, CSV/TSV (and other standard delimiters), OpenDocument Database and Spreadsheets, Geo-Referenced TIFF, Common Data Format, HDF/HDF5, LabView TDMS, Excel, DeltaTables, Parquet, Apache Iceberg, Apache Hudi and many others. Metadata Collection - Scanners for the local and networked file systems and cloud storage providers. Network integration with common databases such as MSSQL and MySQL. User Plugin System - Users are able to provide either file processing, metadata extraction, or sampling plugins in the programming language of their choice. Authentication/Authorization -: OIDC integration, SCIM provisioning and EntraID integration out of the box. Full user and group management system with a “least privilege” operating mode. Governance - Customizable data governance platform; dictate and enforce required metadata, enforce data embargos, and enforce user agreements and NDAs before data access. Ability to create health checks on data, rejecting abandoned or poorly curated data and automatically removing it from the search index. Ability for users to submit corrections. Search - Semantic search is a first class citizen. No licenses to expensive, external software required. Integrated use of vectors and vector-based search allows for AI agent integration at all levels of operation. Metadata Model - Display and control data’s lineage and connections to other data and data directories. Data is modeled after a filesystem - an organization instantly recognizable and navigable by most any user. CLI and SDK - Ships with a Command Line Interface (CLI) tool and with a fully-featured Python SDK. This allows for rapid and programmatic use of Datum by every level of user. Minimal Infrastructure - Datum ships as a single executable file and can be run on any operating system and most CPU architectures. Datum has no reliance on external databases, search indexing tools, or other outside services - and it runs equally well on edge computing devices, cloud services, or in a clustered HPC environment.

darrington, john↗

Mapping Inquiry Tool (MapIT) Database

The Mapping Inquiry Tool (MapIT) database consists of a geodatabase and data catalog of geologic, geophysical, structural, hydrologic, and contextual data, based on the data types to support geologic carbon storage activities and other subsurface energy systems resource assessments. The database was aggregated from publicly available data across the USA from state and federal entities. The database is structured by categories including rock unit geology, boundaries, national CS datasets, geophysical data, faults and structural data, infrastructure, surface hydrology, groundwater, and more. The data described in the data catalog is also available in the Mapping Inquiry Tool (https://edx.netl.doe.gov/dataset/mapping-inquiry-tool). Version 3 of the geodatabase and data catalog have been updated as of 5/17/2024. The database was published with a limited number of layers. The Catalog V3 contains many more resources than the geodatabase, documenting all layers that will be included in MapIT, and includes links to the original sources of the data. Within the catalog, in the final column, there is information about if the file is included in the geodatabase or not. Use the links provided in the catalog to download data directly from the original source if not included in the geodatabase. Four resources are included in this submission: 1. Geodatabase 2. ReadMe file 3. Catalog of data layers and additional data resources 4. Web link to a resource describing the motivation and reviewing the content of the geodatabase - DOE NETL Carbon Storage Site Mapping Inquiry Tool Database

carbon storage↗

Prospective Seal Unit Spatial Extent Database for U.S. Sedimentary Basins

The Prospective Seal Unit Spatial Extent Database for U.S. Sedimentary Basins contains a series of spatial datasets representing spatial extents of publicly available data for caprock and seal rock units within the Appalachian Basin, Denver-Julesburg Basin, Great Valley Basin (Sacramento and San Joaquin Basins), Illinois Basin, Michigan Basin, San Juan Basin, U.S. Gulf Coast Basin, and Williston Basin. The database is designed to support carbon storage feasibility and resources assessment for carbon transport and storage (CTS) projects while displaying the spatial extent of prospective seal units and provide a guide to the original data source. This database leverages publicly available data resources from authoritative sources (e.g. U.S. Geological Survey, State Geologic Surveys, and published reports), and aims to help guide users to understand the seal unit's spatial coverage and data gaps from the regional to sub-basin/field scale. The database is organized by seal unit/formation, including the spatial extent for data found to be available for the seal unit. The various datasets represented include spatial extents of the lithologic formation, depth to top structural contour maps, and thickness/isopach maps. Included in this submission are the following resources: 1. Geodatabase/Dataset: “prospective-seal-unit-extents-2025.gdb” 2. ReadMe: “readme-prospective-seal-unit-spatial-extent-dataset-2025.pdf” 3. Data Catalog: “prospective-seal-unit-spatial-extents-data-catalog-2025.xlsx” 4. Data Sources Key: “data-source.csv” Please see NETL disclaimers here: https://netl.doe.gov/home/disclaimer

Basin↗

Assessing Low-Temperature Geothermal Play Types: Relevant Data and Play Fairway Analysis Methods

This data catalog contains information on low temperature geothermal play types. The U.S. Department of Energy (DOE) Geothermal Technologies Office (GTO) supports the Geothermal Heating and Cooling Geospatial Datasets and Analysis project, conducted by the National Renewable Energy Laboratory (NREL). This project is part of a broader effort to demonstrate the multifaceted value of integrating geothermal power and geothermal heating and cooling technologies into national decarbonization strategies and community energy plans. There is a need to establish baseline low-temperature geothermal resource data sets and evaluate methods for deploying these technologies. This project aims to reduce exploration risk of low temperature geothermal systems by collecting baseline datasets that can be used for Play Fairway Analysis methodologies. This data catalog contains links to publicly available datasets from different sources that can be relevant for the low temperature geothermal systems. This submission contains data catalogs for Alaska, Hawaii, and the Conterminous United States, as well as a technical report on the methods used to classify and asses the geothermal play types.

15 GEOTHERMAL ENERGY↗

Cataloging Legacy Data from the Tritium Systems Test Assembly Program

The Tritium Systems Test Assembly (TSTA) at Los Alamos National Laboratory, operational from 1984 to 2001, was critical in advancing fusion fuel cycle technologies, including tritium storage, gas separation, and pumping. TSTA’s contributions, particularly in safe tritium operations, have influenced subsequent fusion projects. This paper discusses the ongoing effort to digitize and catalog TSTA’s historical data to create a searchable resource for the fusion research community. While the long-term objective is to develop a relational database for structured data management, the project remains in the early phase, with current efforts focused on scanning and indexing physical documents. Initial plans for database implementations are also presented, outlining key considerations for structure, query indexing, and standardization. As digitization progresses, future discussions will refine these implantation details to ensure an efficient and comprehensive system. This initiative aims to preserve critical legacy data, enhance the design of tritium system facilities, and support the next generation of fusion energy research.

42 ENGINEERING↗

Empowering Scientific Discovery Through Computing at the Advanced Photon Source

This paper explores the challenges and solutions for managing and processing the vast amount of data generated by the Advanced Photon Source (APS), a synchrotron light source facility producing ultra-bright x-rays for diverse scientific domains. With 68 experimental beamlines covering materials research, biology, and more, the APS serves a wide user base across academia, government, and industry. The ongoing upgrade of the APS storage ring and installation of new instruments will amplify data generation and processing demands. This paper discusses the approach to address these demands through automated data processing using standardized workflows that produce faster scientific insights. The APS Data Management System coordinates various data related tasks to manage storage, data transfer, metadata cataloging, data processing, and interfaces with tools provided by Globus. Through integration with the Argonne Leadership Computing Facility (ALCF), APS users can efficiently access high-performance computing resources. Standardized workflows have led to reduced computational burdens on scientists and greater accessibility of high performance computing resources. We demonstrate how standardization and collaboration enable scientists to rapidly convert raw data into meaningful scientific results, establishing a streamlined path from data collection to analysis and ultimately to publication.

Parraga, Hannah↗

Spatial Seal Database for Prospective Storage Resources in the USA

The goal of the Spatial Seal Database for Prospective Storage Resources in the USA is to provide relevant information and spatial extents of caprock and seal rocks. A lack of aggregated information is readily available that focuses on the caprock and seal units within sedimentary basins. The EPA class VI permit requires an assessment of the confining zone as part of submitting a permit. The data catalog of seal unit names and relevant properties with the seal spatial extent database aims to help provided important data for carbon storage based assessments. The data catalog and database are designed to show what seal data is available in a sedimentary basin and guide stakeholders to the original data source for those datasets.

Pantaleone, Scott↗

DESI 2024 III: baryon acoustic oscillations from galaxies and quasars

We present the DESI 2024 galaxy and quasar baryon acoustic oscillations (BAO) measurements using over 5.7 million unique galaxy and quasar redshifts in the range 0.1 < z < 2.1. Divided by tracer type, we utilize 300,017 galaxies from the magnitude-limited Bright Galaxy Survey with 0.1 < z < 0.4, 2,138,600 Luminous Red Galaxies with 0.4 < z < 1.1, 2,432,022 Emission Line Galaxies with 0.8 < z < 1.6, and 856,652 quasars with 0.8 < z < 2.1, over a ∼ 7,500 square degree footprint. The analysis was blinded at the catalog-level to avoid confirmation bias. All fiducial choices of the BAO fitting and reconstruction methodology, as well as the size of the systematic errors, were determined on the basis of the tests with mock catalogs and the blinded data catalogs. We present several improvements to the BAO analysis pipeline, including enhancing the BAO fitting and reconstruction methods in a more physically-motivated direction, and also present results using combinations of tracers. We employ a unified BAO analysis method across all tracers. We present a re-analysis of SDSS BOSS and eBOSS results applying the improved DESI methodology and find scatter consistent with the level of the quoted SDSS theoretical systematic uncertainties. With the total effective survey volume of ∼ 18 Gpc3, the combined precision of the BAO measurements across the six different redshift bins is ∼0.52%, marking a 1.2-fold improvement over the previous state-of-the-art results using only first-year data. We detect the BAO in all of these six redshift bins. The highest significance of BAO detection is 9.1σ at the effective redshift of 0.93, with a constraint of 0.86% placed on the BAO scale. We find that our observed BAO scales are systematically larger than the prediction of the Planck 2018-ΛCDM at z < 0.8. We translate the results into transverse comoving distance and radial Hubble distance measurements, which are used to constrain cosmological models in our companion paper.

79 ASTRONOMY AND ASTROPHYSICS↗

Community Requirements Meta-Analysis: Characterizing Needs and Opportunities for HPDF

This High Performance Data Facility (HPDF) Project is creating a new scientific user facility to provide advanced infrastructure for data-intensive science, supporting the DOE’s Office of Science (SC) community. HPDF’s mission is to enable and accelerate scientific discovery by delivering state-of-the-art data management infrastructure, capabilities, and tools. This meta-analysis examines the needs of the breadth of the SC community, captured in publicly available community reports or mission documents. The meta-analysis identifies and provides initial characterization of fifteen core requirements for the HPDF Project team to consider during the conceptual design phase. The fifteen requirements illustrate how scientific work among SC communities requires modern, seamless user experiences across the ASCR Ecosystem to advance the use of large volumes of heterogeneous data. The scientific community requires support for the missing middle of compute between local and HPC to interactively and collaboratively use growing datasets. Data producers and end users will benefit from enhanced data catalogs and portals that improve data access through advanced search of well curated data. The fifteen requirements are examined here organized across five themes for discussion. Examples in each theme illustrate the array of scientific needs that convey the important role that the fully realized and operational High Performance Data Facility will be able to play as an integral part of the evolving ASCR Ecosystem. Our amalgamated data tables from ESnet reports demonstrate ranges to the volumes of data HPDF must be concerned with, but limitations are inherent to this meta-analysis (see Key Challenges & Limitations). Feedback and validation of these requirements along with additional details and emergent community requirements will be gathered through user research and design activities.

97 MATHEMATICS AND COMPUTING↗