Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Hierarchical Data Format”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Hierarchical Data Format (HDF) Status Update

In the "Hierarchical Data Format (HDF) Status Update" talk we will give an update on the recent and upcoming HDF5 software releases, and how to upgrade HDF5-based software to the new major releases HDF5 1.10.* to assure backward and forward compatibility for the files produces with the newest versions of the HDF5 software. We will also focus on a compression feature of HDF5 and will talk about new mechanism for storing HDF5 data in Object Store.

Virtual File Driver↗

Closing the Gap between FAIR Data Repositories and Hierarchical Data Formats

Many in the scientific community, particularly in publicly funded research, are pushing to adhere to more accessible data standards to maximize the findability, accessibility, interoperability, and reusability (FAIR) of scientific data, especially with the growing prevalence of machine learning augmented research. Online FAIR data repositories, such as the Open Science Framework (OSF), help facilitate the adoption of these standards by providing frameworks for storage, access, search, APIs, and other features that create organized hubs of scientific data. However, the wider acceptance of such repositories is hindered by the lack of support of hierarchical data formats, such as Technical Data Management Streaming (TDMS) and Hierarchical Data Format 5 (HDF5), that many researchers rely on to organize their datasets. Various tools and strategies should be used to allow hierarchical data formats, FAIR data repositories, and scientific organizations to work more seamlessly together. A pilot project at Los Alamos National Laboratory (LANL) addresses the disconnect between them by integrating the OSF FAIR data repository with hierarchical data renderers, extending support for additional file types in their framework. The multifaceted interactive renderer displays a tree of metadata alongside a table and plot of the data channels in the file. This allows users to quickly and efficiently load large and complex data files directly in the OSF webapp. Users who are browsing files can quickly and intuitively see the files in the way they or their colleagues structured the hierarchical form and immediately grasp their contents. This solution helps bridge the gap between hierarchical data storage techniques and FAIR data repositories, making both of them more viable options for scientific institutions like LANL which have been put off by the lack of integration between them.

97 MATHEMATICS AND COMPUTING↗

AIRS Data Subsetting Service at the Goddard Earth Sciences (GES) DISC/DAAC

The AIRS mission, as a combination of the Atmospheric Infrared Sounder (AIRS), the Advanced Microwave Sounding Unit (AMSU) and the Humidity Sounder for Brazil (HSB), brings climate research and weather prediction into 21st century. From NASA' Aqua spacecraft, the AIRS/AMSU/HSB instruments measure humidity, temperature, cloud properties and the amounts of greenhouse gases. The AIRS also reveals land and sea- surface temperatures. Measurements from these three instruments are analyzed . jointly to filter out the effects of clouds from the IR data in order to derive clear-column air-temperature profiles and surface temperatures with high vertical resolution and accuracy. Together, they constitute an advanced operational sounding data system that have contributed to improve global modeling efforts and numerical weather prediction; enhance studies of the global energy and water cycles, the effects of greenhouse gases, and atmosphere-surface interactions; and facilitate monitoring of climate variations and trends. The high data volume generated by the AIRS/AMSU/HSB instruments and the complexity of its data format (Hierarchical Data Format, HDF) are barriers to AIRS data use. Although many researchers are interested in only a fraction of the data they receive or request, they are forced to run their algorithms on a much larger data set to extract the information of interest. In order to better server its users, the GES DISC/DAAC, provider of long-term archives and distribution services as well science support for the AIRS/AMSU/HSB data products, has developed various tools for performing channels, variables, parameter, spatial and derived products subsetting, resampling and reformatting operations. This presentation mainly describes the web-enabled subsetting services currently available at the GES DISC/DAAC that provide subsetting functions for all the Level 1B and Level 2 data products from the AIRS/AMSU/HSB instruments.

Vicente, Gilberto A.↗

Hierarchical Data Format for Earth Observing System Data Product Developer's Guide

The "Hierarchical Data Format for Earth Observing System" talk will address the best practices for creating ESDIS data products. The work presented is done in support of Data Product Developers Guide Working Group with mission "to help data product developers make data usable for end users". During the presentation, we will use some examples of NASA data products and show how to modify them to make data more usable.

Data usability↗

Hierarchical Data Format for Nuclear Data Sensitivities

The SCALE code system includes capabilities for sensitivity and uncertainty (S/U) analysis as part of its TSUNAMI code suite. The sensitivity of a quantity of interest (for example, an application’s $k_{eff}$) to nuclear data is stored as a profile in a text-based file, which is known as a sensitivity data file (SDF). The sensitivity profile can be used to calculate uncertainties, correlation coefficients, and similarity indices. One of the goals of the present work was to seek general performance improvements in the TSUNAMI code suite, starting with the TSUNAMI-IP code for calculating similarity indices. Through profiling, it was found that reading the text-based sensitivity files was a performance bottleneck in the TSUNAMI-IP code. In a typical TSUNAMI-IP calculation, an application might be compared to thousands of benchmarks, thus requiring the reading of thousands of SDFs. Reading of binary-based data is generally faster than reading text-based data. Hierarchical Data Format 5 (HDF5) is a binary-based format that also benefits from being portable, and it can be inspected with nonproprietary tools. This paper describes an HDF5-based file format that has been introduced for SDFs.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Hierarchical Data Format for Nuclear Data Sensitivities [Slides]

An HDF5-based file format was introduced for the sensitivity data calculated by TSUNAMI. The format was defined to collect the sensitivity coefficients into hyperslabs, which optimizes file reading time and therefore improves the time-to-solution for applications. In future work, this format will be extended to store sensitivity data for depletion calculations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

The Hierarchical Data Format for EOS (HDF-EOS)

HDF is a file format and a software library for data storage, management, exchange, and archiving. It is written and maintained by the National Center for Supercomputing Applications (NCSA). HDF5 has a very simple but versatile data model which is compatible with most competing formats. Through its grouping and linking mechanisms, the HDF5 data model enables complex data relationships and dependencies. HDF5 accommodates the inclusion of many common types of metadata and arbitrary types and quantities of user-defined metadata.

Ullman, Richard↗

International Satellite Cloud Climatology Project (ISCCP) Stage D1 3-Hourly Cloud Product - Revised Algorithm in Hierarchical Data Format (ISCCP_D1)

Since 1983 an international group of institutions has collected and analyzed satellite radiance measurements from up to five geostationary and two polar orbiting satellites to infer the global distribution of cloud properties and their diurnal, seasonal and interannual variations. The primary focus of the first phase of the project (1983-1995) was the elucidation of the role of clouds in the radiation budget (top of the atmosphere and surface). In the second phase of the project (1995 onwards) the analysis also concerns improving understanding of clouds in the global hydrological cycle. [Location=TROPOSPHERE] [Temporal_Coverage: Start_Date=1983-07-01; Stop_Date=] [Spatial_Coverage: Southernmost_Latitude=-90; Northernmost_Latitude=90; Westernmost_Longitude=-180; Easternmost_Longitude=180] [Data_Resolution: Latitude_Resolution=280 Km; Longitude_Resolution=280 Km; Temporal_Resolution=3 Hourly].

CLOUD LIQUID WATER PATH↗

International Satellite Cloud Climatology Project (ISCCP) Stage D2 Monthly Cloud Product - Revised Algorithm in Hierarchical Data Format (ISCCP_D2)

Since 1983 an international group of institutions has collected and analyzed satellite radiance measurements from up to five geostationary and two polar orbiting satellites to infer the global distribution of cloud properties and their diurnal, seasonal and interannual variations. The primary focus of the first phase of the project (1983-1995) was the elucidation of the role of clouds in the radiation budget (top of the atmosphere and surface). In the second phase of the project (1995 onwards) the analysis also concerns improving understanding of clouds in the global hydrological cycle. [Location=TROPOSPHERE] [Temporal_Coverage: Start_Date=1983-07-01; Stop_Date=] [Spatial_Coverage: Southernmost_Latitude=-90; Northernmost_Latitude=90; Westernmost_Longitude=-180; Easternmost_Longitude=180] [Data_Resolution: Latitude_Resolution=280 Km; Longitude_Resolution=280 Km; Temporal_Resolution=Monthly].

CLOUD TOP PRESSURE↗

Hierarchical Data Formats (HDF) Update

In this presentation, we will talk about the latest releases of HDF4 and HDF5 software and tools, new features available in HDF5, and roadmap for the HDF software. We will also solicit feedback from the users of HDF data and HDF application developers on new features and new tools. The talk will cover: Difference between 1.8 and 1.10 releases and how and when to move to the latest release Features of the recent HDF5 1.8.19, 1.10.1 and HDF 4.2.13 Overview of HDF View 3.0 and other enhancements to tools Supported compilers and systems Open discussion of new requirements and wish list of the HDF features Compression library for interoperability with h5py and Pandas and better floating-point data compression.

HDFView↗

Ocean Data from MODIS at the NASA Goddard DAAC

Terra satellite carrying the Moderate Resolution Imaging Spectroradiometer (MODIS) was successfully launched on December 18, 1999. Some of the 36 different wavelengths that MODIS samples have never before been measured from space. New ocean data products, which have not been derived on a global scale before, are made available for research to the scientific community. For example, MODIS uses a new split window in the four-micron region for the better measurement of Sea Surface Temperature (SST), and provides the unprecedented ability (683 nm band) to measure chlorophyll fluorescence. At full ocean production, more than a thousand different ocean products in three major categories (ocean color, sea surface temperature, and ocean primary production) are archived at the NASA Goddard Earth Sciences (GES) Distributed Active Archive Center (DAAC) at the rate of approx. 230GB/day. The challenge is to distribute such large volumes of data to the ocean community. It is achieved through a combination of public and restricted EOS Data Gateways, the GES DAAC Search and Order WWW interface, and an FTP site that contains samples of MODIS data. A new Search and Order WWW interface at http://acdisx.gsfc.nasa.gov/data/ developed at the GES DAAC is based on a hierarchical organization of data, will always return non-zero results. It has a very convenient geographical representation of five-minute data granule coverage for each day MODIS Data Support Team (MDST) continues the tradition of quality support at the GES DAAC for the ocean color data from the Coastal Zone Color Scanner (CZCS) and the Sea Viewing Wide Field-of-View Sensor (SeaWiFS) by providing expert assistance to users in accessing data products, information on visualization tools, documentation for data products and formats (Hierarchical Data Format-Earth Observing System (HDF-EOS)), information on the scientific content of products and metadata. Visit the MDST website at http://daac.gsfc.nasa.gov/CAMPAIGN DOCS/MODIS/index.html

Leptoukh, Gregory G.↗

Re-Organizing Earth Observation Data Storage to Support Temporal Analysis of Big Data

The Earth Observing System Data and Information System archives many datasets that are critical to understanding long-term variations in Earth science properties. Thus, some of these are large, multi-decadal datasets. Yet the challenge in long time series analysis comes less from the sheer volume than the data organization, which is typically one (or a small number of) time steps per file. The overhead of opening and inventorying complex, API-driven data formats such as Hierarchical Data Format introduces a small latency at each time step, which nonetheless adds up for datasets with O(10^6) single-timestep files. Several approaches to reorganizing the data can mitigate this overhead by an order of magnitude: pre-aggregating data along the time axis (time-chunking); storing the data in a highly distributed file system; or storing data in distributed columnar databases. Storing a second copy of the data incurs extra costs, so some selection criteria must be employed, which would be driven by expected or actual usage by the end user community, balanced against the extra cost.

data storage↗

SeaWiFS technical report series. Volume 19: Case studies for SeaWiFS calibration and validation, part 2

This document provides brief reports, or case studies, on a number of investigations and data set development activities sponsored by the Calibration and Validation Team (CVT) within the Sea-viewing Wide Field-of-view Sensor (SeaWiFS) Project. Chapter 1 is a comparison with the atmospheric correction of Coastal Zone Color Scanner (CZCS) data using two independent radiative transfer formulations. Chapter 2 is a study on lunar reflectance at the SeaWiFS wavelengths which was useful in establishing the SeaWiFS lunar gain. Chapter 3 reports the results of the first ground-based solar calibration of the SeaWiFS instrument. The experiment was repeated in the fall of 1993 after the instrument was modified to reduce stray light; the results from the second experiment will be provided in the next case studies volume. Chapter 4 is a laboratory experiment using trap detectors which may be useful tools in the calibration round-robin program. Chapter 5 is the original data format evaluation study conducted in 1992 which outlines the technical criteria used in considering three candidate formats, the hierarchical data format (HDF), the common data format (CDF), and the network CDF (netCDF). Chapter 6 summarizes the meteorological data sets accumulated during the first three years of CZCS operation which are being used for initial testing of the operational SeaWiFS algorithms and systems and would be used during a second global processing of the CZCS data set. Chapter 7 describes how near-real time surface meteorological and total ozone data required for the atmospheric correction algorithm will be retrieved and processed. Finally, Chapter 8 is a comparison of surface wind products from various operational meteorological centers and field observations. Surface winds are used in the atmospheric correction scheme to estimate glint and foam radiances.

Hooker, Stanford B.↗

Processing TES Level-2 Data

TES Level 2 Subsystem is a set of computer programs that performs functions complementary to those of the program summarized in the immediately preceding article. TES Level-2 data pertain to retrieved species (or temperature) profiles, and errors thereof. Geolocation, quality, and other data (e.g., surface characteristics for nadir observations) are also included. The subsystem processes gridded meteorological information and extracts parameters that can be interpolated to the appropriate latitude, longitude, and pressure level based on the date and time. Radiances are simulated using the aforementioned meteorological information for initial guesses, and spectroscopic-parameter tables are generated. At each step of the retrieval, a nonlinear-least-squares- solving routine is run over multiple iterations, retrieving a subset of atmospheric constituents, and error analysis is performed. Scientific TES Level-2 data products are written in a format known as Hierarchical Data Format Earth Observing System 5 (HDF-EOS 5) for public distribution.

Poosti, Sassaneh↗

Simplifying Analysis of Hierarchical HDF5 and NetCDF4 Files with Xarray-Datatree

NASA’s Earth Observing System Data and Information System (EOSDIS) contains thousands of Earth science datasets from satellites, models, and field campaigns. EOSDIS data are stored in formats that are well supported by the Earth Science community. These formats include the Hierarchical Data Format (HDF), with derivative flavors such as HDF-5 and the Network Common Data Format (NetCDF-4). The HDF specification allows for a directory-like hierarchy within a single file, known as "groups". Observational data and associated metadata within a single file can be distributed amongst multiple internal groups, which can also be nested to multiple levels. Working with datasets that have a group hierarchical structure can be difficult because of the nested structure of groups. Widely used packages, such as xarray, have data models that do not accommodate the hierarchical structure within HDF files, requiring users to traverse the file and open different HDF groups as separate, unrelated objects. Xarray-datatree is a Python package developed to solve the difficulty of traversing HDFs with a hierarchical group structure by creating a tree-like hierarchical data structure in xarray. The tree-like structure allows each group to be accessed once a DataTree object is instantiated. The migration of xarray-datatree into the xarray core library will reduce barriers to accessing Earth science data by eliminating the need to understand and traverse the specific hierarchy of a grouped HDF file.

Eni Awowale↗

Fostering Geothermal Machine Learning Success: Elevating Big Data Accessibility and Automated Data Standardization in the Geothermal Data Repository

The Department of Energy's (DOE's) Geothermal Data Repository (GDR) has implemented improvements to both its data lakes and its data standards and automated data pipelines. The GDR data lakes have reduced storage and compute-related barriers to using large geothermal datasets, enabling these large datasets to be accessed by anyone with a modern computer and internet access. More recently, the GDR has been working to further reduce barriers through streamlining the data intake process, educating users on the process and requirements, and helping users access data from the data lakes. These improvements have augmented the quantity of datasets the GDR is able to accept into its data lakes and have enabled users who are new to cloud tools to access these datasets more easily, overall increasing the accessibility of big geothermal data for use in machine learning and other projects. In addition, the GDR now has built-in data standards and pipelines for drilling data, geospatial data, and distributed acoustic sensing (DAS) data. These standardization efforts aim to enhance the real-world applicability of geothermal machine learning outcomes by improving the quality of training data. Specifically, through standardizing high-value datasets, the GDR is reducing project-specific data curation requirements, thus allowing more time for actual research. By automating this process, the burden of standardization is lifted from the user, ultimately increasing the availability of standardized data. This paper provides an update on recent improvements made to the GDR's data lakes and automated data pipelines, including: (1) streamlining the data lake intake process, (2) better educating users on the process and requirements through a new data lakes page, (3) adding data lake direct access links to GDR data lake submission pages, (4) implementing a DAS data pipeline to convert DAS data uploaded in SEG-Y format to a standardized hierarchical data format v5 (HDF5), (5) extending this pipeline to encompass data in the GDR data lake, (6) adding metadata requirements for geospatial data, (7) making user interface/user experience (UX) enhancements to the data pipelines' documentation pages, and (8) improving the GDR's data standards and pipelines pages to better guide users in ensuring that their data is standardized by the GDR's automated data pipelines. 2024 Geothermal Resources Council. All rights reserved.

accessibility↗

Subsetting and Formatting Landsat-7 LOR ETM+ and Data Products

The Landsat-7 Processing System (LPS) processes Landsat-7 Enhanced Thematic Mapper (ETM+) instrument data into large, contiguous segments called "subintervals" and stores them in Level OR (LOR) data files. The LPS processed subinterval products must be subsetted and reformatted before the Level I processing systems can ingest them. The initial full subintervals produced by the LPS are stored mainly in HDF Earth Observing System (HDF-EOS) format which is an extension to the Hierarchical Data Format (HDF). The final LOR products are stored in native HDF format. Primarily the EOS Core System (ECS) and alternately the DAAC Emergency System (DES) subset the subinterval data for the operational Landsat-7 data processing systems. The HDF and HDF-EOS application programming interfaces (APIs) can be used for extensive data subsetting and data reorganization. A stand-alone subsetter tool has been developed which is based on some of the DES code. This tool makes use of the HDF and HDFEOS APIs to perform Landsat-7 LOR product subsetting and demonstrates how HDF and HDFEOS can be used for creating various configurations of full LOR products. How these APIs can be used to efficiently subset, format, and organize Landsat-7 LOR data as demonstrated by the subsetter tool and the DES is discussed.

Reid, Michael R.↗

Cloud Optimized Data Formats

Cloud computing offers the promise of being able to analyze Big Data earth Observations at scale, by allowing scientists to deploy many nodes at once to analyze the data. However, in order to take full advantage of cloud scalability, it is often necessary to reorganize and reformat the data to enable fine-grained, parallel access to the data in Web Object Storage. NASA recently conducted a study of several formats that are optimized for analysis in the cloud: Parquet, zarr, HDF (Hierarchical Data Format) in the Cloud, and Cloud-Optimized GeoTIFF (Tagged Image File Format). They were compared against non-cloud-optimized formats, netCDF (network Common Data Form) and GeoTIFF, with criteria based both on stewardship and analysis performance.

Christopher Lynnes↗