Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hierarchical data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Re-Organizing Earth Observation Data Storage to Support Temporal Analysis of Big Data

The Earth Observing System Data and Information System archives many datasets that are critical to understanding long-term variations in Earth science properties. Thus, some of these are large, multi-decadal datasets. Yet the challenge in long time series analysis comes less from the sheer volume than the data organization, which is typically one (or a small number of) time steps per file. The overhead of opening and inventorying complex, API-driven data formats such as Hierarchical Data Format introduces a small latency at each time step, which nonetheless adds up for datasets with O(10^6) single-timestep files. Several approaches to reorganizing the data can mitigate this overhead by an order of magnitude: pre-aggregating data along the time axis (time-chunking); storing the data in a highly distributed file system; or storing data in distributed columnar databases. Storing a second copy of the data incurs extra costs, so some selection criteria must be employed, which would be driven by expected or actual usage by the end user community, balanced against the extra cost.

data storage↗

Microphone Phased Array NetCDF/HDF5 Archival Files: Application Program Interface Reference

An application program interface (API) has been developed for the creation and access of structured data files generated by microphone phased arrays utilized in aeroacoustics research. Two structured binary file formats are supported, namely NetCDF (Network Common Data Form) and HDF5 (Hierarchical Data Format) files. The API consists of a library of routines callable from C, Fortran or Matlab, with native versions of the API provided for each language. The libraries are divided into categories for file handling, file definition and initialization, data writing, data recovery, and error handling. The API is intended to provide a mechanism for generating self-describing binary files for long-term archiving of raw and processed data generated by phased array systems.

Humphreys, William M., Jr.↗

Forming Aggregations using Virtual Sharding: Lessons Learned from Simple Scalable Storage (S3)

Data aggregation is the ability to combine separate datasets to form a single new logical dataset provides users with a powerful abstraction. The advantage of an aggregate dataset is that the users are freed from having to understand, and incorporate into their workflow, knowledge about the (ad hoc) organization of the constituent datasets. However, aggregating large numbers of files can be computationally complex with data server systems performing many repetitive operations. As part of the authors work on subsetting data stored on Amazon Web Service (AWS) Simple Storage Service (S3), we developed technology to read portions of otherwise monolithic data files. This enables the formation of virtual shards for user in subsetting data stored in HDF5 (hierarchical data format, version 5) files. This same tool can be used to form aggregations that combine data stored in many HDF5 files when those files are stored on S3. The nature of the virtual sharding and the algorithm that exploits it for subsetting is such that it can also be used for aggregation with the need for many of the repetitive operations required by the per file aggregation techniques. We will present timing information that demonstrates the flexibility of this approach. However, the lessons learned is that while this is a useful result in and of itself, these very same techniques can be applied in other contexts where data are stored in services and on media other than S3. For example, this same technique can be applied to data stored on spinning disk. Pushing the envelope for S3 forced a reexamination of our data access techniques which lead to unexpected positive benefits.

Gallagher, James↗

Using Big Data Technologies with Earth Science Data in HDF5: HDF5 Scalable Solutions

HDF5 (Hierarchical Data Format 5) is open-source, high-performance software that consists of an abstract data model, library, and fileformat used for storing and managing extremely large and/or complex data collections. NASA Earth Observing System (EOS) Data and Information Systems use HDF5 as an archival format to store remote sensing data from EOS satellites. HDF5 is also used to store other types of Geoscience and Strophysical data, e.g., seismic data and data from Low-Frequency Array (LOFAR) radio telescopes. Data stored in HDF5 has reached tens of petabytes and is growing at an accelerated rate.With the growing amout of HDF5 Earth Science data to analyze and process, scientists need to adopt big data technologies including new storage paradigms such as cloud and object storage. To run models and perform data analysis they also need to utilizied efficient and diverse ways to access data, from high-performance computing's (HPC) Message Passing Interface (MPI) I/O and deep memory hierarchies (DMH) to non-HPC frameworks such as Apache Hadoop, Spark, and Drill. The HDF Group continually works to enable usage of big data technologies in HDF software.

Knox, Larry↗

Extending CF Conventions to Enhance Data FAIRness for Atmospheric Composition Observations

The Hierarchical Data Format (HDF) and Network Common Data Form (NetCDF) are data file formats created to aid users in the creation or use of scientific data. These file formats are useful for handling large data volumes and hosting extensive metadata as global, group, or variable attributes and are popular with the modeling community. HDF and NetCDF files are widely used with atmospheric remote sensing data and have been used to support measurements from numerous field campaigns, from satellite to aircraft or ground and mobile based measurements. The files from airborne field studies, however, vary greatly in terms of the file structure and the amount and content of their metadata. Information relevant to the file that can be useful to the user such as the data producer, location where data was taken, variable descriptions, or information about the instrument might not be included in the file. Recently, the Measurements of Aerosols, Clouds, and their Interactions for Earth System Models (MACIE) group started a grassroots effort to develop a CF-based template for the HDF and NetCDF files for field studies, with the aim of making the data products more interoperable and usable. This template seeks to make the files more compliant to Climate and Forecast (CF) metadata conventions and to standardize the file structure and the global and variable attributes. The template would help to ensure that HDF and NetCDF files contain adequate metadata to better support their use for research, e.g., the modeling community, and to enhance the usability and interoperability of data for research communities at large. The draft template has been applied to recent field studies for various instruments and their merge files in support of the Atmosphere Observing System (AOS) project. The details of the revised template are to be presented, as well as examples of the implementation of these requirements for merge files and lidar observation data files and issues revealed during the implementation process.

Sean Leavor↗

Multifrequency data analysis software on STARLINK

Although the STARLINK project was set up to provide image processing facilities to UK astronomers, it has grown over the last 12 years to the extent that it now provides most of the data analysis facilities for UK astronomers. One aspect of the growth of the STARLINK network is that it now has to cater for astronomers working in a diverse range of wavelengths. Since a given individual may be working with data obtained in a variety of wavelengths, it is most convenient if the data can be stored in a common format and the programs that analyze the data have a similar 'look and feel'. What is known as 'STARLINK software' is obtained from many sources: STARLINK funded programmers; astronomers; foreign projects such as AIPS; generally available shareware; and commercial sources when this proves cost effective. This means that the ideal situation of a completely integrated system cannot be realized in practice. Nevertheless, many of the major packages written by STARLINK application programmers and by astronomers do use a common data format, based on the Hierarchical Data System, so that interchange of data between packages designed separately from each other is simply a matter of using the same file names. For example, as astronomer might use KAPPA to read some optical spectra off a FITS tape, then use CCDPACK to debias and flat field the data (it is easy to set up an overnight batch job to do this if there is a lot of data), then use KAPPA to have a quick look at the data and then use Figaro to reduce the spectra. It is useful to divide data analysis packages into wavelength specific packages, or even instrument specific packages, and general purpose ones. Once the instrumental signature has been removed from some data, any appropriate general purpose package can be used to analyze te data. For example, the ASTERIX package deals with x-ray data reduction, but after dealing with all of the x-ray specific processing, an astronomer may well want to find the brightness of objects in a given frame. Since ASTERIX uses the standard STARLINK data format, the astronomer can use PHOTOM or DAOPHOT 2 to measure the brightness of the objects. Although DAOPHOT was written with optical astronomy in mind, it is useful for analyzing data from several wavelengths. The ability of DAOPHOT 2 to handle non-standard point spread functions can be especially useful in many areas of astronomy.

Allan, P. M.↗

Implementation of CCSDS Lossless Data Compression in HDF

The Earth Science Data and Information System (ESDIS) handles over one terabyte (10(exp 12) bytes) of data daily and is using the Hierarchical Data Format (EDF) for data archiving and distribution. This report provides the progress and status of our effort to alleviate bandwidth and storage burdens by first performing compression studies on various science data products and later integrating the selected compression scheme into HDF.

Pen-Shu Yeh↗

MODIS Data from the GES DISC DAAC: Moderate-Resolution Imaging Spectroradiometer (MODIS)

The Goddard Earth Sciences (GES) Distributed Active Archive Center (DAAC) is responsible for the distribution of the Level 1 data, and the higher levels of all Ocean and Atmosphere products (Land products are distributed through the Land Processes (LP) DAAC DAAC, and the Snow and Ice products are distributed though the National Snow and Ice Data Center (NSIDC) DAAC). Ocean products include sea surface temperature (SST), concentrations of chlorophyll, pigment and coccolithophores, fluorescence, absorptions, and primary productivity. Atmosphere products include aerosols, atmospheric water vapor, clouds and cloud masks, and atmospheric profiles from 20 layers. While most MODIS data products are archived in the Hierarchical Data Format-Earth Observing System (HDF-EOS 2.7) format, the ocean binned products and primary productivity products (Level 4) are in the native HDF4 format. MODIS Level 1 and 2 data are of the Swath type and are packaged in files representing five minutes of Files for Level 3 and 4 are global products at daily, weekly, monthly or yearly resolutions. Apart from the ocean binned and Level 4 products, these are in Grid type, and the maps are in the Cylindrical Equidistant projection with rectangular grid. Terra viewing (scenes of approximately 2000 by 2330 km). MODIS data have several levels of maturity. Most products are released with a provisional level of maturity and only announced as validated after rigorous testing by the MODIS Science Teams. MODIS/Terra Level 1, and all MODIS/Terra 11 micron SST products are announced as validated. At the time of this publication, the MODIS Data Support Team (MDST) is working with the Ocean Science Team toward announcing the validated status of the remainder of MODIS/Terra Ocean products. MODIS/Aqua Level 1 and cloud mask products are released with provisional maturity.

Source record↗

Airborne Spectral BRDF of Various Surface Types (Ocean, Vegetation, Snow, Desert, Wetlands, Cloud Decks, Smoke Layers) for Remote Sensing Applications

In this paper we describe measurements of the bidirectional reflectance-distribution function (BRDF) acquired over a 30-year period (1984-2014) by the National Aeronautics and Space Administration's (NASA's) Cloud Absorption Radiometer (CAR). Our BRDF database encompasses various natural surfaces that are representative of many land cover or ecosystem types found throughout the world. CAR's unique measurement geometry allows a comparison of measurements acquired from different satellite instruments with various geometrical configurations, none of which are capable of obtaining such a complete and nearly instantaneous BRDF. This database is therefore of great value in validating many satellite sensors and assessing corrections of reflectances for angular effects. These data can also be used to evaluate the ability of analytical models to reproduce the observed directional signatures, to develop BRDF models that are suitable for sub-kilometer-scale satellite observations over both homogeneous and heterogeneous landscape types, and to test future spaceborne sensors. All of these BRDF data are publicly available and accessible in hierarchical data format (http:car.gsfc.nasa.gov/).

albedo↗

Geographic Information Systems for Assessing Existing and Potential Bio-energy Resources: Their Use in Determining Land Use and Management Options which Minimize Ecological and Landscape Impacts in Rural Areas

A management construct is described which forms part of an overall landscape ecological planning model which has as a principal objective the extension of the traditional descriptive land use mapping capabilities of geographic information systems into land management realms. It is noted that geographic information systems appear to be moving to more comprehensive methods of data handling and storage, such as relational and hierarchical data management systems, and a clear need has simultaneously arisen therefore for planning assessment techniques and methodologies which can actually use such complex levels of data in a systematic, yet flexible and scenario dependent way. The descriptive of mapping method proposed broaches such issues and utilizes a current New England bioenergy scenario, stimulated by the use of hardwoods for household heating purposes established in the post oil crisis era and the increased awareness of the possible landscape and ecological ramifications of the continued increasing use of the resource.

Jackman, A. E.↗

An object-oriented data reduction system in Fortran

A data reduction system for the AAO two-degree field project is being developed using an object-oriented approach. Rather than use an object-oriented language (such as C++) the system is written in Fortran and makes extensive use of existing subroutine libraries provided by the UK Starlink project. Objects are created using the extensible N-dimensional Data Format (NDF) which itself is based on the Hierarchical Data System (HDS). The software consists of a class library, with each class corresponding to a Fortran subroutine with a standard calling sequence. The methods of the classes provide operations on NDF objects at a similar level of functionality to the applications of conventional data reduction systems. However, because they are provided as callable subroutines, they can be used as building blocks for more specialist applications. The class library is not dependent on a particular software environment thought it can be used effectively in ADAM applications. It can also be used from standalone Fortran programs. It is intended to develop a graphical user interface for use with the class library to form the 2dF data reduction system.

Bailey, J.↗

Tropospheric Emission Spectrometer Product File Readers

TES Product File Reader software extracts data from publicly available Tropospheric Emission Spectrometer (TES) HDF (Hierarchical Data Format) product data files using publicly available format specifications for scientific analysis in IDL (interactive data language). In this innovation, the software returns data fields as simple arrays for a given file. A file name is provided, and the contents are returned as simple IDL variables.

Fisher, Brendan M.↗

Processing TES Level-2 Data

TES Level 2 Subsystem is a set of computer programs that performs functions complementary to those of the program summarized in the immediately preceding article. TES Level-2 data pertain to retrieved species (or temperature) profiles, and errors thereof. Geolocation, quality, and other data (e.g., surface characteristics for nadir observations) are also included. The subsystem processes gridded meteorological information and extracts parameters that can be interpolated to the appropriate latitude, longitude, and pressure level based on the date and time. Radiances are simulated using the aforementioned meteorological information for initial guesses, and spectroscopic-parameter tables are generated. At each step of the retrieval, a nonlinear-least-squares- solving routine is run over multiple iterations, retrieving a subset of atmospheric constituents, and error analysis is performed. Scientific TES Level-2 data products are written in a format known as Hierarchical Data Format Earth Observing System 5 (HDF-EOS 5) for public distribution.

Poosti, Sassaneh↗

Software to Compare NPP HDF5 Data Files

This software was developed for the NPOESS (National Polar-orbiting Operational Environmental Satellite System) Preparatory Project (NPP) Science Data Segment. The purpose of this software is to compare HDF5 (Hierarchical Data Format) files specific to NPP and report whether the HDF5 files are identical. If the HDF5 files are different, users have the option of printing out the list of differences in the HDF5 data files. The user provides paths to two directories containing a list of HDF5 files to compare. The tool would select matching HDF5 file names from the two directories and run the comparison on each file. The user can also select from three levels of detail. Level 0 is the basic level, which simply states whether the files match or not. Level 1 is the intermediate level, which lists the differences between the files. Level 2 lists all the details regarding the comparison, such as which objects were compared, and how and where they are different. The HDF5 tool is written specifically for the NPP project. As such, it ignores certain attributes (such as creation_date, creation_ time, etc.) in the HDF5 files. This is because even though two HDF5 files could represent exactly the same granule, if they are created at different times, the creation date and time would be different. This tool is smart enough to ignore differences that are not relevant to NPP users.

Wiegand, Chiu P.↗

GMI-IPS: Python Processing Software for Aircraft Campaigns

NASA's Atmospheric Tomography Mission (ATom) seeks to understand the impact of anthropogenic air pollution on gases in the Earth's atmosphere. Four flight campaigns are being deployed on a seasonal basis to establish a continuous global-scale data set intended to improve the representation of chemically reactive gases in global atmospheric chemistry models. The Global Modeling Initiative (GMI), is creating chemical transport simulations on a global scale for each of the ATom flight campaigns. To meet the computational demands required to translate the GMI simulation data to grids associated with the flights from the ATom campaigns, the GMI ICARTT Processing Software (GMI-IPS) has been developed and is providing key functionality for data processing and analysis in this ongoing effort. The GMI-IPS is written in Python and provides computational kernels for data interpolation and visualization tasks on GMI simulation data. A key feature of the GMI-IPS, is its ability to read ICARTT files, a text-based file format for airborne instrument data, and extract the required flight information that defines regional and temporal grid parameters associated with an ATom flight. Perhaps most importantly, the GMI-IPS creates ICARTT files containing GMI simulated data, which are used in collaboration with ATom instrument teams and other modeling groups. The initial main task of the GMI-IPS is to interpolate GMI model data to the finer temporal resolution (1-10 seconds) of a given flight. The model data includes basic fields such as temperature and pressure, but the main focus of this effort is to provide species concentrations of chemical gases for ATom flights. The software, which uses parallel computation techniques for data intensive tasks, linearly interpolates each of the model fields to the time resolution of the flight. The temporally interpolated data is then saved to disk, and is used to create additional derived quantities. In order to translate the GMI model data to the spatial grid of the flight path as defined by the pressure, latitude, and longitude points at each flight time record, a weighted average is then calculated from the nearest neighbors in two dimensions (latitude, longitude). Using SciPya's Regular Grid Interpolator, interpolation functions are generated for the GMI model grid and the calculated weighted averages. The flight path points are then extracted from the ATom ICARTT instrument file, and are sent to the multi-dimensional interpolating functions to generate GMI field quantities along the spatial path of the flight. The interpolated field quantities are then written to a ICARTT data file, which is stored for further manipulation. The GMI-IPS is aware of a generic ATom ICARTT header format, containing basic information for all flight campaigns. The GMI-IPS includes logic to edit metadata for the derived field quantities, as well as modify the generic header data such as processing dates and associated instrument files. The ICARTT interpolated data is then appended to the modified header data, and the ICARTT processing is complete for the given flight and ready for collaboration. The output ICARTT data adheres to the ICARTT file format standards V1.1. The visualization component of the GMI-IPS uses Matplotlib extensively and has several functions ranging in complexity. First, it creates a model background curtain for the flight (time versus model eta levels) with the interpolated flight data superimposed on the curtain. Secondly, it creates a time-series plot of the interpolated flight data. Lastly, the visualization component creates averaged 2D model slices (longitude versus latitude) with overlaid flight track circles at key pressure levels. The GMI-IPS consists of a handful of classes and supporting functionality that have been generalized to be compatible with any ICARTT file that adheres to the base class definition. The base class represents a generic ICARTT entry, only defining a single time entry and 3D spatial positioning parameters. Other classes inherit from this base class; several classes for input ICARTT instrument files, which contain the necessary flight positioning information as a basis for data processing, as well as other classes for output ICARTT files, which contain the interpolated model data. Utility classes provide functionality for routine procedures such as: comparing field names among ICARTT files, reading ICARTT entries from a data file and storing them in data structures, and returning a reduced spatial grid based on a collection of ICARTT entries. Although the GMI-IPS is compatible with GMI model data, it can be adapted with reasonable effort for any simulation that creates Hierarchical Data Format (HDF) files. The same can be said of its adaptability to ICARTT files outside of the context of the ATom mission. The GMI-IPS contains just under 30,000 lines of code, eight classes, and a dozen drivers and utility programs. It is maintained with GIT source code management and has been used to deliver processed GMI model data for the ATom campaigns that have taken place to date.

Damon, M. R.↗

File servers, networking, and supercomputers

One of the major tasks of a supercomputer center is managing the massive amount of data generated by application codes. A data flow analysis of the San Diego Supercomputer Center is presented that illustrates the hierarchical data buffering/caching capacity requirements and the associated I/O throughput requirements needed to sustain file service and archival storage. Usage paradigms are examined for both tightly-coupled and loosely-coupled file servers linked to the supercomputer by high-speed networks.

Moore, Reagan W.↗

File servers, networking, and supercomputers

One of the major tasks of a supercomputer center is managing the massive amount of data generated by application codes. A data flow analysis of the San Diego Supercomputer Center is presented that illustrates the hierarchical data buffering/caching capacity requirements and the associated I/O throughput requirements needed to sustain file service and archival storage. Usage paradigms are examined for both tightly-coupled and loosely-coupled file servers linked to the supercomputer by high-speed networks.

Moore, Reagan W.↗

Visualization, Analysis and Subsetting Tools for EOS Aura Data Products in HDF-EOS5

Aura data products are among the first to use the new version 5 of the Hierarchical Data Format for the Earth Observing System, or HDF-EOS5. This presentation discusses the common HDF-EOS5 file layout that is adopted for most of the EOS Aura standard data products. Details of the various tools that can be used to access, visualize and subset these data will also be provided. Aura, the NASA Earth Observing System's atmospheric chemistry mission, was successfully launched July 15, 2004. The Aura spacecraft includes four instruments: the High Resolution Dynamics Limb Sounder (HIRDLS), the Microwave Limb Sounder (MLS), the Ozone Monitoring Instrument (OMI), and the Tropospheric Emission Spectrometer (TES). Data from the HIRDLS, MLS and OMI will be archived at the NASA Goddard Earth Sciences (GES) Distributed Active Archive Center (DAAC), while TES data will be archived at the NASA Langley Research Center DAAC. For more information see http://daac.gsfc.nasa.gov/.

Johnson, J.↗