Cloud-Native Geospatial Data Formats
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.
The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.
The Department of Energy's (DOE) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and aiding users in accessing data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data.
Research on infrastructure resilience has produced promising methods to simulate and optimize complex networks to improve performance. However, restrictions on sharing infrastructure models and the steep cost of developing and maintaining infrastructure models presents a roadblock to adoption. To overcome this limitation, this research focuses on methods to create data-driven infrastructure models that will help improve infrastructure resilience and security. The analysis couples incomplete utility data, geospatial data, machine learning, and synthetic network generation methods to rapidly develop and update infrastructure models. The methods are validated using realistic utility models and site-specific data, with a focus on Puerto Rico due to its unique infrastructure challenges and available data. This research highlights promising opportunities for the use of synthetic network generation and machine learning to create infrastructure models when very little data is available. Results demonstrate that hybrid methods, which combine sparse utility data with synthetic models, can enhance model accuracy, and machine learning can predict model attributes using training data from other models. However, the complexity of infrastructure systems means that even minor changes in network connectivity can significantly impact simulation results. Resilience analysis using synthetic infrastructure models shows that while some system behaviors are preserved, the magnitude of disruptions may not be accurately represented, indicating the need for more research and validation before using synthetic models for critical infrastructure investment decisions. The framework outlined in this report represents a significant advance to infrastructure model development and could be applied to additional domains and sites. Future research will continue to streamline and validate methods to help reduce roadblocks to resilience analysis.
The Runtime for Airspace Concept Evaluation (RACE) is an open-source software architecture and framework to build configurable, highly concurrent and distributed message-based systems that offer scalable, low-latency performance on commodity hardware. RACE was used in commercial aviation applications to rapidly build systems that span several machines (including synchronized displays), interface existing hardware simulators and other live data feeds, and incorporate sophisticated visualization components such as NASA WorldWind. These RACE applications validated elements of the FAA’s System Wide Information Management (SWIM) Program, handling up to 1000 messages/sec from diverse sources (SFDPS, TFM-DATA, TAIS, ASDE-X, ITWS and local ADS) for 4,500 simultaneous flights tracked in the next-generation air transportation system’s digital backbone. We have since generalized RACE to support Open Data Integration (ODIN) applications outside aviation. Systems built with RACE/ODIN can be deployed in the field, on commodity hardware, and operate with limited or intermittent connectivity to the outside world. Our primary use case is a web-server with local/persistent data storage that runs within and only serves the stakeholder network (e.g. an incident command post). We are tailoring the RACE/ODIN system to support wildland fire management for the upcoming NASA Wildland Fire Safety Demonstration Series. RACE-ODIN is under consideration for application in the Scalable Traffic Management for Emergency Response Operations project, or STEReO, which aims to create a system that can be deployed during emergencies, to coordinate multiple elements of disaster response. Such data sources predominantly come from existing services on the internet (e.g. weather and satellite data, imported from so called "edge servers") but can also include dynamic (real-time) data from computer simulations and within the stakeholder network (such as aircraft and personnel tracking information). We will present the architecture and ODIN system demonstration incorporating local data from instrumented power-line towers, interpolated weather data and geospatial data from space-based platforms.
The Runtime for Airspace Concept Evaluation (RACE) is an open-source software architecture and framework to build configurable, highly concurrent and distributed message-based systems that offer scalable, low-latency performance on commodity hardware. RACE was used in commercial aviation applications to rapidly build systems that span several machines (including synchronized displays), interface existing hardware simulators and other live data feeds, and incorporate sophisticated visualization components such as NASA WorldWind. These RACE applications validated elements of the FAA’s System Wide Information Management (SWIM) Program, handling up to 1000 messages/sec from diverse sources (SFDPS, TFM-DATA, TAIS, ASDE-X, ITWS and local ADS) for 4,500 simultaneous flights tracked in the next-generation air transportation system’s digital backbone. We have since generalized RACE to support Open Data Integration (ODIN) applications outside aviation. Systems built with RACE/ODIN can be deployed in the field, on commodity hardware, and operate with limited or intermittent connectivity to the outside world. Our primary use case is a web-server with local/persistent data storage that runs within and only serves the stakeholder network (e.g. an incident command post). We are tailoring the RACE/ODIN system to support wildland fire management for the upcoming NASA Wildland Fire Safety Demonstration Series. RACE-ODIN is under consideration for application in the Scalable Traffic Management for Emergency Response Operations project, or STEReO, which aims to create a system that can be deployed during emergencies, to coordinate multiple elements of disaster response. Such data sources predominantly come from existing services on the internet (e.g. weather and satellite data, imported from so called "edge servers") but can also include dynamic (real-time) data from computer simulations and within the stakeholder network (such as aircraft and personnel tracking information). We will present the architecture and ODIN system demonstration incorporating local data from instrumented power-line towers, interpolated weather data and geospatial data from space-based platforms.
Traditionally, NASA Earth Science data archives have file-based storage using proprietary data file formats, such as HDF and HDF-EOS, which are optimized to support fast and efficient storage of spaceborne and model data as they are generated. The use of file-based storage essentially imposes an indexing strategy based on data dimensions. In most cases, NASA Earth Science data uses time as the primary index, leading to poor performance in accessing data in spatial dimensions. For example, producing a time series for a single spatial grid cell involves accessing a large number of data files. With exponential growth in data volume due to the ever-increasing spatial and temporal resolution of the data, using file-based archives poses significant performance and cost barriers to data discovery and access. Storing and disseminating data in proprietary data formats imposes an additional access barrier for users outside the mainstream research community. At the NASA Goddard Earth Sciences Data Information Services Center (GES DISC), we have evaluated applying the schema-on-read principle to data access and distribution. We used Apache Parquet to store geospatial data, and have exposed data through Amazon Web Services (AWS) Athena, AWS Simple Storage Service (S3), and Apache Spark. Using the schema-on-read approach allows customization of indexing spatially or temporally to suit the data access pattern. The storage of data in open formats such as Apache Parquet has widespread support in popular programming languages. A wide range of solutions for handling big data lowers the access barrier for all users. This presentation will discuss formats used for data storage, frameworks with This presentation will discuss formats used for data storage, frameworks with support for schema-on-read used for data access, and common use cases covering data usage patterns seen in a geospatial data archive.
The Basin-Scale Structural Features database provides spatial datasets of faults, fractures, folds, and earthquakes compiled from public, authoritative sources (e.g., U.S. Geological Survey and State Geological Surveys) and aggregated into derivative forms to support subsurface assessments. Recognizing that characterizing basin-scale structural features requires interpreting data that are often ambiguous or lack key information, the source data were evaluated using a knowledge-data framework and geospatial fuzzy logic method (Justman et al., 2020) to represent both measured (observed) and predicted (inferred or potential) structural features as derivative datasets. This workflow employs conceptual models for known structural features and predicted structural features, incorporating geospatial data to estimate potential, even with limited data. The aim is to aid and support an understanding of basin-scale features and identify potential gaps in data and knowledge. As of 4/30/2025, the database includes resources for nine sedimentary basins: Appalachian, Denver, U.S. Gulf Coast, Illinois, Michigan, Permian, Sacramento, San Joquin and Williston. The database is organized by basin and then data category: 1) Faults, fractures, folds, 2) Earthquakes, 3) Topographic, 4) Structural contours and isopachs, 5) Geophysical, and 6) Structural feature density assessment maps.
This project would identify a methodology and implement a living solution to map, both visually and utilizing some form of database, the complex network of stakeholders that the Disasters Program routinely interacts with to maximize efficiency and minimize confusion and overlapping effort during disaster responses. This project's solution will take into account factors such as stakeholder data production type, geospatial data maturity, geographic areas of interest, federal mandates, type of relationship, national priorities and many additional relevant attributes.
The VA Environmental Determinants of Health (EDH) Advanced Software Pipeline Framework is designed to enhance the efficiency, scalability, and security of geospatial data processing workflows. This framework integrates modern data orchestration and containerization technologies, including Prefect for workflow automation, Docker for containerization, and PostgreSQL/PostGIS for geospatial data storage and analysis. It ensures standardized, reproducible, and automated data processing, supporting VA objectives related to substance use risk assessment and recovery research. The pipeline addresses key scalability and performance challenges through horizontal and vertical scaling, high-performance computing (HPC) integration, parallel processing, task caching, and dynamic resource allocation. These optimizations improve throughput and reduce latency, allowing the system to efficiently manage large and complex datasets. Additionally, security and compliance measures—such as data encryption (SSL), Role-Based Access Control (RBAC), and adherence to GDPR and HIPAA standards—safeguard sensitive information throughout data transmission and storage. A key implementation of this framework includes the automation of shelter list geolocation workflows, ensuring that up-to-date data is readily available for VA decision-making. Lessons learned from this project include the transition from in-memory processing to incremental storage writes, improving resource management and reliability. Future enhancements aim to expand automation, integrate AI-driven anomaly detection, and incorporate high-performance computing resources. This framework provides a scalable, secure, and adaptable solution for managing geospatial datasets, reinforcing the VA’s ability to support clinical and strategic initiatives through data-driven decision-making.
This dataset provides site and endmember spectra collected during the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS) campaign. The site spectra were collected to help validate airborne hyperspectral data acquired by the National Ecological Observatory Network's aerial observation platform (NEON AOP). Endmember spectra were collected to augment existing spectral libraries with additional samples of bare surfaces and non-photosynthetic vegetation. All measurements were acquired with an Analytical Spectral Devices (ASD) FieldSpec4 Hi-Res NG (Next Generation) spectroradiometer, which records radiance at 1nm (nanometer) intervals from the ultraviolet to the short-wave infrared (350-2500 nm). The dataset includes spectra measured at meadow sites where the CHESS team also collected vegetation samples for trait analyses. The site spectra were collected with the ASD FieldSpec4 palm grip attachment using an 8° field-of-view foreoptic. Site spectra are integrated measurements of the entire surface within the foreoptic’s field of view. For site-level spectra, the sun is the illumination source. A Spectralon panel mounted on a tripod was used for instrument optimization and white reference measurements for all site spectra. Site spectra were acquired within two hours of solar noon and within 48 hours of a NEON AOP overflight. Site spectra are labeled by date, sampling area, and site number according to the naming conventions of the CHESS campaign’s data management plan. The dataset also contains endmember spectra in the following categories: photosynthetic vegetation (PV), non-photosynthetic vegetation (NPV), bare (soil/rock), and flowers. Endmember measurements were acquired using either the contact probe or the leaf clip attachments of the ASD FieldSpec4. In these configurations, the bulb inside the spectrometer provides the light source for the measurements. The spectrometer was optimized and white reference measurements were recorded using the circular white pucks attached to the contact probe and leaf clip. Because they do not rely on solar illumination, contact probe and leaf clip measurements were collected during a broader time frame than the palm grip site spectra. Some endmembers were measured at CHESS meadow sites, while others were collected within the larger sampling area or in nearby locations (e.g. Gothic Townsite) with similar characteristics. Radiance, reflectance, and metadata files are split into three subfolders according to measurement type: proximal/palm grip (prx), contact probe (cp), and leaf clip (lc). Radiance spectra are provided in ASD file format (.asd file extension). All ASD files can be opened using the provided scripts. Metadata is provided in two formats: CSV file format (no geolocation) and GEOJSON file format (includes geolocation for each spectra). The dataset includes a set of pre-processed reflectance spectra as CSV files (yyyymmdd_rfl.csv). The python scripts and jupyter notebook used to calculate reflectance spectra from the ASD radiance data is included here and was previously published at: https://doi.org/10.3334/ORNLDAAC/2446. There is also a folder of JPEG photographs corresponding to selected spectra. We include a protocol document with detailed steps for ASD FieldSpec4 assembly and operations. This data additionally contains a file level metadata (flmd.csv) and data dictionary (dd.csv) file. Geospatial information: Geospatial data for mapping measurement site locations are in the files CHESS_polygons_lai_UTM.geojson, CHESS_polygons_shrub_UTM.geojson, and CHESS_polygons_meadow_UTM.geojson in the companion geospatial package for the 2025 CHESS campaign, ‘CHESS 2025: Location data for field observations and sampling’ (Henderson et al., 2026). CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgment: This research was carried out at the Jet Propulsion Laboratory, California Institute of Technology, under a contract with the National Aeronautics and Space Administration (80NM0018D0004) and was funded by EMIT Extended Mission Phase E Science.
Explore the source record for details and available documents.
This data package contains tabular and geospatial data used to quantify and model meteoric beryllium-10 fluxes in the East River watershed, Colorado, USA. The tabular component includes calibration-site data from five glacial moraine sites and includes environmental variables used to evaluate spatial controls on meteoric 10Be delivery, including elevation, mean annual precipitation (MAP), mean snow depth, and mean snow water equivalent (SWE). These site-level data were used to compare observed fluxes with environmental gradients across the watershed and to evaluate the effects of erosion correction on flux estimates. The package also includes supporting slope and curvature values used to assess topographic inputs to the erosion analysis. A second component of the data package contains updated manuscript tables and regression outputs used to summarize the relationships between meteoric 10Be flux and environmental predictors. These tables include meteoric 10Be sample information and AMS results, site-level environmental values, site-level meteoric 10Be inventory and flux values, watershed-averaged predicted fluxes, soil bulk density measurements, fine-fraction values, soil pH measurements, and regression statistics including slope, intercept, coefficient of determination, and p-value. The regression products include both standard linear regressions and regressions in which the intercept is constrained to pass through zero, and they support the analyses presented in the companion manuscript. Together, these tabular files provide the numerical basis for the manuscript tables and the regression-based interpretation of meteoric 10Be flux variability in a snow-dominated mountain watershed. The geospatial component of the package consists of GeoTIFF raster files used to generate the map products presented in Figures 2 and 6 of the companion manuscript. These rasters represent watershed-scale spatial layers for environmental variables and regression-based predictions of meteoric 10Be flux. This dataset contains comma-separated values files (.csv), Microsoft Excel files (.xlsx), GeoTIFF raster files (.tif), and upporting metadata files, including CSV data dictionaries and readme text files (.csv, .txt). The tabular files can be opened with standard spreadsheet software, and the raster files can be viewed and analyzed in GIS software such as ArcGIS Pro or QGIS. Together, these files document the numerical and spatial datasets used to calibrate and predict meteoric 10Be delivery in the East River watershed.
This presentation discusses MODIS vegetation phenology products used in the ForWarn Early Warning System (EWS) tool for near real time regional forest disturbance detection and surveillance at regional to national scales. The ForWarn EWS is being developed by the USDA Forest Service NASA, ORNL, and USGS to aid federal and state forest health management activities. ForWarn employs multiple historical land surface phenology products that are derived from MODIS MOD13 Normalized Difference Vegetation Index (NDVI) data. The latter is temporally processed into phenology products with the Time Series Product Tool (TSPT) and the Phenological Parameter Estimation Tool (PPET) software produced at NASA Stennis Space Center. TSPT is used to effectively noise reduce, fuse, and void interpolate MODIS NDVI data. PPET employs TSPT-processed NDVI time series data as an input, outputting multiple vegetation phenology products at a 232 meter resolution for 2000 to 2011, including NDVI magnitude and day of year products for seven key points along the growing season (peak of growing season and the minima, 20%, and 80% of the peak NDVI for both the left and right side of growing season), cumulative NDVI integral products for the most active part of the growing season and sequentially across the growing season at 8 day intervals, and maximum value NDVI products composited at 24 day intervals in which each product date has 8 days of overlap between the previous and following product dates. MODIS NDVI phenology products are also used to compute nationwide NRT forest change products refreshed every 8 days. These include percent change in forest NDVI products that compare the current NDVI from USGS eMODIS products to historical MODIS MOD13 NDVI. For each date, three forest change products are produced using three different maximum value NDVI baselines (from the previous year, three previous years, and all previous years). All change products are output with a rainbow color table in which forests with the most severe NDVI decreases are assigned hot colors (yellow to red) and forests with prominent NDVI increases are assigned cold colors (blue tones). All mentioned products have been integrated as data layers into ForWarn s geospatial data viewer known as the U.S. Forest Change Assessment Viewer (FCAV). The latter is used to view and assess the context of the mentioned forest change products with respect to ancillary data layers, such as land cover, elevation, hydrologic features, climatic data, storm data, aerial disturbance surveys, fire data, and land ownership. The FCAV also includes a temporal NDVI profiler for viewing phenological change in multi-year NDVI associated with known or suspected regionally apparent forest disturbances (e.g., from fire and insects). ForWarn forest change products have been used to detect, track, and assess several biotic and abiotic regional forest disturbance events across the country, including ephemeral and longer lasting damage from storms, drought, and insects. Such change products are most effective for viewing severe disturbances affecting multiple MODIS pixels. MODIS vegetation phenology products contribute vital current information on forest conditions to the ForWarn system and this role is expected to grow as these products are refined and derivative products are added.
The Geospatial Raster Input Data for Capacity Expansion Regional Feasibility (GRIDCERF) data package is a high-resolution product to evaluate siting suitability for renewable and non-renewable power plants in the conterminous United States. GRIDCERF offers hundreds of individual suitability layers for use with both renewable and non-renewable power plant technology configurations in a harmonized format that can be easily ingested by geospatially-enabled modeling software. It also provides pre-compiled technology-specific suitability layers and allows for user customization to robustly address science objectives when evaluating varying future conditions. GRIDCERF data can be directly used with the CERF (Capacity Expansion Regional Feasibility) model to site power plants at a 1km resolution. GRIDCERF includes composite technology siting suitability raster layers for the following utility scale technology configurations. Note that, in addition to technology sub-types shown below, various cooling types are also included (recirculating, pond, once-through, recirculating-seawater, dry-hybrid, or dry) for various technologies. Biomass Conventional (with or without CCS) IGCC (with or without CCS) Coal Conventional (with or without CCS) IGCC (with or without CCS) Natural Gas Combined-cycle (CC) (with or without CCS) Turbine Geothermal Enhanced Geothermal Systems (EGS) - Class 1 through Class 5 resource potential Nuclear Gen 2 Light Water Reactor (LWR) Gen 3 Small Modular Reactor (SMR) Gen 3 AP1000 Refined Liquids Combined-cycle (CC) (with or without CCS) Turbine Solar Photovoltaic (PV) - for capacity factors in the range of 6-18% Utility-scale Concentrating Solar Power (CSP) - for capacity factors in the range of 24-46% Tower Wind (Onshore) - for capacity factors in the range of 5-50% 80m hub height 100m hub height 120m hub height 140m hub height Wind (Offshore) - for capacity factors in the range of 25-60% 100m hub height 140m hub height 160m hub height
Earth science data users almost always have an interest in utilizing geospatial data from multiple agencies. As computing capability and cloud-based infrastructures accelerate the pace at which scientific research can be done, there is a growing need to enable search, discovery, and use of multi-agency geospatial observations relevant for a common use case - without undergoing the search and discovery process in a less efficient, disparate path with each agency. NASA’s Earth Observing System Data and Information System (EOSDIS) and NOAA’s National Environmental Satellite, Data and Information Service (NESDIS) both support a wide range of Earth science disciplines’ research, operations, and applications activities. Presently, however, there are few examples of data discovery frameworks supporting an inquiry of both NASA’s and NOAA’s extensive archives of Earth observations that are equally suitable for a particular science scenario, regardless of the agency that “owns” the data. NASA and NOAA are collaborating on a data expedition platform for exploring fire weather using data products from both agencies. Users will be able to search, discover, and visualize NASA and NOAA products in one interface. Each agency will curate metadata for its respective datasets, providing for a rich search experience. The collaboration will pilot a shared search interface into these metadata datastores. Data products will be stored in the cloud in cloud-optimized format(s). These formats will allow for optimized data access and visualization to support the “data expedition”. Avenues for further development and application of this cloud-based, multi-agency data provisioning platform will also be discussed.