Engineering PapersSearch

Engineering topics

Petrenko, Maksym

Publications and source records attributed to Petrenko, Maksym.

At least 19 records

Cloud Giovanni: Reining in Costs and Improving Performance with Analytical Data Stores Using Scalable Serverless Architecture

Giovanni is the Geospatial Interactive Online Visualization ANd aNalysis Infrastructure developed at NASA GES DISC which provides a simple and intuitive way to visualize, analyze, and access vast amounts of Earth science data. It receives large number of user requests each day for a variety of analysis and visualization services, which leads to the big data challenge of serving gradually increasing large data volumes with diverse statistical algorithms. We hereby propose a multi-dimensional accumulation method which provides fast and cost-efficient cloud analysis for diverse services including both area averaging and time averaging. This method involves the weighted volume integration over multiple variable dimensions (time and space), and is implemented in AWS using Athena providing serverless and highly scalable data analysis. Compared to the standard method, this approach dramatically reduced the computational time by order of magnitude with a minimal AWS cost incurred. For example, for a benchmark of 10-year area averaging over the 1x1 degree daily variable, the computational time was reduced from minutes to seconds, and the Athena cost is only $5 for 100,000 requests.

Zhang, Hailiang

The Value of Data and Metadata Standardization for Interoperability in Giovanni Or: Why Your Product's Metadata Causes Us Headaches!

Giovanni is a data exploration and visualization tool at the NASA Goddard Earth Sciences Data Information Services Center (GES DISC). It has been around in one form or another for more than 15 years. Giovanni calculates simple statistics and produces 22 different visualizations for more than 1600 geophysical parameters from more than 90 satellite and model products. Giovanni relies on external data format standards to ensure interoperability, including the NetCDF CF Metadata Conventions. Unfortunately, these standards were insufficient to make Giovanni's internal data representation truly simple to use. Finding and working with dimensions can be convoluted with the CF Conventions. Furthermore, the CF Conventions are silent on machine-friendly descriptive metadata such as the parameter's source product and product version. In order to simplify analyzing disparate earth science data parameters in a unified way, we developed Giovanni's internal standard. First, the format standardizes parameter dimensions and variables so they can be easily found. Second, the format adds all the machine-friendly metadata Giovanni needs to present our parameters to users in a consistent and clear manner. At a glance, users can grasp all the pertinent information about parameters both during parameter selection and after visualization.

interoperability

Giovanni in the Cloud: Earth Science Data Exploration in Amazon Web Services

Giovanni is an exploration tool at the NASA Goddard Earth Sciences Data Information Services Center (GES DISC), providing 22 analysis and visualization services for over 1600 Earth Science data variables. Owing to its popularity, Giovanni has experienced a consistent growth in overall demand, with periodic usage spikes attributed to trainings by education organizations, extensive data analysis in response to natural disasters, preparations for science meetings, etc. Furthermore, the new generation of spaceborne sensors and high resolution models have resulted in an exponential growth in data volume with data distributed across the traditional boundaries of data centers. Seamless exploration of data (without users having to worry about data center boundaries) has been a key recommendation of the GES DISC User Working Group. These factors have required new strategies for delivering acceptable performance. The cloud-based Giovanni, built on Amazon Web Services (AWS), evaluates (1) AWS native solutions to provide a scalable, serverless architecture; (2) open standards for data storage in the Cloud; (3) a cost model for operations; and (4) end-user performance. Our preliminary findings indicate that the use of serverless architecture has a potential to significantly reduce development and operational cost of Giovanni. The combination of using AWS managed services, storage of data in open standards, and schema-on-read data access strategy simplifies data access and analytics, in addition to making data more accessible to the end users of Giovanni through popular programming languages.

Giovanni

Use of Schema on Read in Earth Science Data Archives

Traditionally, NASA Earth Science data archives have file-based storage using proprietary data file formats, such as HDF and HDF-EOS, which are optimized to support fast and efficient storage of spaceborne and model data as they are generated. The use of file-based storage essentially imposes an indexing strategy based on data dimensions. In most cases, NASA Earth Science data uses time as the primary index, leading to poor performance in accessing data in spatial dimensions. For example, producing a time series for a single spatial grid cell involves accessing a large number of data files. With exponential growth in data volume due to the ever-increasing spatial and temporal resolution of the data, using file-based archives poses significant performance and cost barriers to data discovery and access. Storing and disseminating data in proprietary data formats imposes an additional access barrier for users outside the mainstream research community. At the NASA Goddard Earth Sciences Data Information Services Center (GES DISC), we have evaluated applying the schema-on-read principle to data access and distribution. We used Apache Parquet to store geospatial data, and have exposed data through Amazon Web Services (AWS) Athena, AWS Simple Storage Service (S3), and Apache Spark. Using the schema-on-read approach allows customization of indexing spatially or temporally to suit the data access pattern. The storage of data in open formats such as Apache Parquet has widespread support in popular programming languages. A wide range of solutions for handling big data lowers the access barrier for all users. This presentation will discuss formats used for data storage, frameworks with This presentation will discuss formats used for data storage, frameworks with support for schema-on-read used for data access, and common use cases covering data usage patterns seen in a geospatial data archive.

cloud applications

A-Train Datalist - A New GES DISC Service to Allow One-Stop Shopping for A-Train Data

The currently available services at the Goddard Earth Sciences Data Information Services Center (GES DISC) only allow users to select variables from a single data set at a time. Because entire variables from a data set are often displayed, user selection of variables of interest can be overwhelming. At the American Geophysical Union (AGU) 2016 Fall Meeting, GES DISC unveiled a new service called Datalist: a collection of predefined or user-defined data variables from one or more archived data sets. Our science support team has been curating Datalists and providing added value to the user community.Originally known as Afternoon Constellation, A-Train includes six currently on polar-orbiting Earth observation satellites: OCO-2, GCOM-W1, Aqua, CALIPSO, CloudSat, and Aura, which travel a few minutes apart from each other. This constellation arrangement has enabled coordinated science observations further forming comprehensive pictures of Earth weather and climate that are readily for use in crucial studies such as climate change.GES DISC Datalists are based on the software architecture of the new GES DISC website (also unveiled at the AGU 2016 Fall Meeting). The GES DISC science support team has created a Datalist to support the A-Train Data Depot (ATDD). Using pre-defined Datalist should hopefully save users significant effort in their data searches.

A-Train data ordering

GES DISC Datalist Enables Easy Data Selection For Natural Phenomena Studies

In order to investigate and assess natural hazards such as tropical storms, winter storms, volcanic eruptions, floods, and drought in a timely manner, the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) has been developing an efficient data search and access service. Called "Datalist," this service enables users to acquire their data of interest "all at once," with minimum effort. A Datalist is a virtual collection of predefined or user-defined data variables from one or more archived data sets. Datalists are more than just data. Datalists effectively provide users with a sophisticated integrated data and services package, including metadata, citation, documentation, visualization, and data-specific services (e.g., subset and OPeNDAP), all available from one-stop shopping. The predefined Datalists, created by the experienced GES DISC science support team, should save a significant amount of time that users would otherwise have to spend. The Datalist service is an extension of the new GES DISC website, which is completely data-driven. A Datalist, also known as "data bundle," is treated just as any other data set. Being a virtual collection, a Datalist requires no extra storage space.

natural hazards

New GES DISC Services Shortening the Path in Science Data Discovery

The Current GES DISC available services only allow user to select variables from a single dataset at a time and too many variables from a dataset are displayed, choice is hard. At American Geophysical Union (AGU) 2016 Fall Meeting, Goddard Earth Sciences Data Information Services Center (GES DISC) unveiled a new service: Datalist. A Datalist is a collection of predefined or user-defined data variables from one or more archived datasets. Our science support team curated predefined datalist and provided value to the user community. Imagine some novice user wants to study hurricane and typed in hurricane in the search box. The first item in the search result is GES DISC provided Hurricane Datalist. It contains scientists recommended variables from multiple datasets like TRMM, GPM, MERRA, etc. Datalist uses the same architecture as that of our new website, which also provides one-stop shopping for data, metadata, citation, documentation, visualization and other available services.We implemented Datalist with new GES DISC web architecture, one single web page that unified all user interfaces. From that webpage, users can find data by either type in keyword, or browse by category. It also provides user with a sophisticated integrated data and services package, including metadata, citation, documentation, visualization, and data-specific services, all available from one-stop shopping.

Datalist

GES DISC Datalist Enables Easy Data Selection for Natural Phenomena Studies

In order to investigate and assess natural hazards such as tropical storms, winter storms, volcanic eruptions, floods, and drought in a timely manner, the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) has been developing an efficient data search and access service. Called Datalist, this service enables users to acquire their data of interest all at once, with minimum effort. A Datalistis a virtual collection of predefined or user-defined data variables from one or more archived data sets. Datalistsare more than just data. Datalistseffectively provide users with a sophisticated integrated data and services package, including metadata, citation, documentation, visualization, and data-specific services (e.g., subset and OPeNDAP), all available from one-stop shopping. The predefined Datalists, created by the experienced GES DISC science support team, should save a significant amount of time that users would otherwise have to spend. The Datalistservice is an extension of the new GES DISC website, which is completely data-driven. A Datalist, also known as data bundle, is treated just as any other data set. Being a virtual collection, a Datalistrequires no extra storage space.

Earth events

UUI: Reusable Spatial Data Services in Unified User Interface at NASA GES DISC

Unified User Interface (UUI) is a next-generation operational data access tool that has been developed at Goddard Earth Sciences Data and Information Services Center(GES DISC) to provide a simple, unified, and intuitive one-stop shop experience for the key data services available at GES DISC, including subsetting (Simple Subset Wizard -SSW), granule file search (Mirador), plotting (Giovanni), and other legacy spatial data services. UUI has been built based on a flexible infrastructure of reusable web services self-contained building blocks that can easily be plugged into spatial applications, including third-party clients or services, to easily enable new functionality as new datasets and services become available. In this presentation, we will discuss our experience in designing UUI services based on open industry standards. We will also explain how the resulting framework can be used for a rapid development, deployment, and integration of spatial data services, facilitating efficient access and dissemination of spatial data sets.

EOSDIS

Datalist: A Value Added Service to Enable Easy Data Selection

Imagine a user wanting to study hurricane events. This could involve searching and downloading multiple data variables from multiple data sets. The currently available services from the Goddard Earth Sciences Data and Information Services Center (GES DISC) only allow the user to select one data set at a time. The GES DISC started a Data List initiative, in order to enable users to easily select multiple data variables. A Data List is a collection of predefined or user-defined data variables from one or more archived data sets. Target users of Data Lists include science teams, individual science researchers, application users, and educational users. Data Lists are more than just data. Data Lists effectively provide users with a sophisticated integrated data and services package, including metadata, citation, documentation, visualization, and data-specific services, all available from one-stop shopping. Data Lists are created based on the software architecture of the GES DISC Unified User Interface (UUI). The Data List service is completely data-driven, and a Data List is treated just as any other data set. The predefined Data Lists, created by the experienced GES DISC science support team, should save a significant amount of time that users would otherwise have to spend.

Datalist

Validation and Expected Error Estimation of Suomi-NNP VIIRS Aerosol Optical Thickness and Angstrom Exponent with AERONET

The new-generation polar-orbiting operational environmental sensor, the Visible Infrared Imaging Radiometer Suite (VIIRS) on board the Suomi National Polar-orbiting Partnership (S-NPP) satellite, provides critical daily global aerosol observations. As older satellite sensors age out, the VIIRS aerosol product will become the primary observational source for global assessments of aerosol emission and transport, aerosol meteorological and climatic effects, air quality monitoring, and public health. To prove their validity and to assess their maturity level, the VIIRS aerosol products were compared to the spatiotemporally matched Aerosol Robotic Network (AERONET)measurements. Over land, the VIIRS aerosol optical thickness (AOT) environmental data record (EDR) exhibits an overall global bias against AERONET of 0.0008 with root-mean-square error(RMSE) of the biases as 0.12. Over ocean, the mean bias of VIIRS AOT EDR is 0.02 with RMSE of the biases as 0.06.The mean bias of VIIRS Ocean Angstrom Exponent (AE) EDR is 0.12 with RMSE of the biases as 0.57. The matchups between each product and its AERONET counterpart allow estimates of expected error in each case. Increased uncertainty in the VIIRS AOT and AE products is linked to specific regions, seasons, surface characteristics, and aerosol types, suggesting opportunity for future modifications as understanding of algorithm assumptions improves. Based on the assessment, the VIIRS AOT EDR over land reached Validated maturity beginning 23 January 2013; the AOT EDR and AE EDR over ocean reached Validated maturity beginning 2 May 2012, excluding the processing error period 15 October to 27 November 2012. These findings demonstrate the integrity and usefulness of the VIIRS aerosol products that will transition from S-NPP to future polar-orbiting environmental satellites in the decades to come and become the standard global aerosol data set as the previous generations missions come to an end.

Huang, Jingfeng

Spatial Distribution of Accuracy of Aerosol Retrievals from Multiple Satellite Sensors

Remote sensing of aerosols from space has been a subject of extensive research, with multiple sensors retrieving aerosol properties globally on a daily or weekly basis. The diverse algorithms used for these retrievals operate on different types of reflected signals based on different assumptions about the underlying physical phenomena. Depending on the actual retrieval conditions and especially on the geographical location of the sensed aerosol parcels, the combination of these factors might be advantageous for one or more of the sensors and unfavorable for others, resulting in disagreements between similar aerosol parameters retrieved from different sensors. In this presentation, we will demonstrate the use of the Multi-sensor Aerosol Products Sampling System (MAPSS) to analyze and intercompare aerosol retrievals from multiple spaceborne sensors, including MODIS (on Terra and Aqua), MISR, OMI, POLDER, CALIOP, and SeaWiFS. Based on this intercomparison, we are determining geographical locations where these products provide the greatest accuracy of the retrievals and identifying the products that are the most suitable for retrieval at these locations. The analyses are performed by comparing quality-screened satellite aerosol products to available collocated ground-based aerosol observations from the Aerosol Robotic Network (AERONET) stations, during the period of 2006-2010 when all the satellite sensors were operating concurrently. Furthermore, we will discuss results of a statistical approach that is applied to the collocated data to detect and remove potential data outliers that can bias the results of the analysis.

Petrenko, Maksym

Estimation and Bias Correction of Aerosol Abundance using Data-driven Machine Learning and Remote Sensing

Air quality information is increasingly becoming a public health concern, since some of the aerosol particles pose harmful effects to peoples health. One widely available metric of aerosol abundance is the aerosol optical depth (AOD). The AOD is the integrated light extinction coefficient over a vertical atmospheric column of unit cross section, which represents the extent to which the aerosols in that vertical profile prevent the transmission of light by absorption or scattering. The comparison between the AOD measured from the ground-based Aerosol Robotic Network (AERONET) system and the satellite MODIS instruments at 550 nm shows that there is a bias between the two data products. We performed a comprehensive analysis exploring possible factors which may be contributing to the inter-instrumental bias between MODIS and AERONET. The analysis used several measured variables, including the MODIS AOD, as input in order to train a neural network in regression mode to predict the AERONET AOD values. This not only allowed us to obtain an estimate, but also allowed us to infer the optimal sets of variables that played an important role in the prediction. In addition, we applied machine learning to infer the global abundance of ground level PM2.5 from the AOD data and other ancillary satellite and meteorology products. This research is part of our goal to provide air quality information, which can also be useful for global epidemiology studies.

Malakar, Nabin K.

Multi-Satellite Synergy for Aerosol Analysis in the Asian Monsoon Region

Atmospheric aerosols represent one of the greatest uncertainties in environmental and climate research, particularly in tropical monsoon regions such as the Southeast Asian regions, where significant contributions from a variety of aerosol sources and types is complicated by unstable atmospheric dynamics. Although aerosols are now routinely retrieved from multiple satellite Sensors, in trying to answer important science questions about aerosol distribution, properties, and impacts, researchers often rely on retrievals from only one or two sensors, thereby running the risk of incurring biases due to sensor/algorithm peculiarities. We are conducting detailed studies of aerosol retrieval uncertainties from various satellite sensors (including Terra-/ Aqua-MODIS, Terra-MISR, Aura-OMI, Parasol-POLDER, SeaWiFS, and Calipso-CALIOP), based on the collocation of these data products over AERONET and other important ground stations, within the online Multi-sensor Aerosol Products Sampling System (MAPSS) framework that was developed recently. Such analyses are aimed at developing a synthesis of results that can be utilized in building reliable unified aerosol information and climate data records from multiple satellite measurements. In this presentation, we will show preliminary results of. an integrated comparative uncertainly analysis of aerosol products from multiple satellite sensors, particularly focused on the Asian Monsoon region, along with some comparisons from the African Monsoon region.

Ichoku, Charles

Effects of Data Quality on the Characterization of Aerosol Properties from Multiple Sensors

Cross-comparison of aerosol properties between ground-based and spaceborne measurements is an important validation technique that helps to investigate the uncertainties of aerosol products acquired using spaceborne sensors. However, it has been shown that even minor differences in the cross-characterization procedure may significantly impact the results of such validation. Of particular consideration is the quality assurance I quality control (QA/QC) information - an auxiliary data indicating a "confidence" level (e.g., Bad, Fair, Good, Excellent, etc.) conferred by the retrieval algorithms on the produced data. Depending on the treatment of available QA/QC information, a cross-characterization procedure has the potential of filtering out invalid data points, such as uncertain or erroneous retrievals, which tend to reduce the credibility of such comparisons. However, under certain circumstances, even high QA/QC values may not fully guarantee the quality of the data. For example, retrievals in proximity of a cloud might be particularly perplexing for an aerosol retrieval algorithm, resulting in an invalid data that, nonetheless, could be assigned a high QA/QC confidence. In this presentation, we will study the effects of several QA/QC parameters on cross-characterization of aerosol properties between the data acquired by multiple spaceborne sensors. We will utilize the Multi-sensor Aerosol Products Sampling System (MAPSS) that provides a consistent platform for multi-sensor comparison, including collocation with measurements acquired by the ground-based Aerosol Robotic Network (AERONET), The multi-sensor spaceborne data analyzed include those acquired by the Terra-MODIS, Aqua-MODIS, Terra-MISR, Aura-OMI, Parasol-POLDER, and CalipsoCALIOP satellite instruments.

Petrenko, Maksym

MODIS Aerosol Optical Depth Bias Adjustment Using Machine Learning Algorithms

To monitor the earth atmosphere and its surface changes, satellite based instruments collect continuous data. While some of the data is directly used, some others such as aerosol properties are indirectly retrieved from the observation data. While retrieved variables (RV) form very powerful products, they don't come without obstacles. Different satellite viewing geometries, calibration issues, dynamically changing atmospheric and earth surface conditions, together with complex interactions between observed entities and their environment affect them greatly. This results in random and systematic errors in the final products.

Albayrak, Arif