Engineering PapersSearch

Engineering topics

Pham, Long

Publications and source records attributed to Pham, Long.

At least 19 records

Citizen Science Twitter Data Management for Earth Science Applications

Social media data can provide useful real-time and historical information relating to the natural world, but managing this data poses challenges. Scientists at GES DISC are exploring the potential of Twitter data to augment precipitation data from the Global Precipitation Measurement (GPM) mission. However, the format of Twitter data is unconventional in the context of NASA data centers, resulting in frustration for scientists who need to work with the data. This study investigated procedures and standards needed to properly manage Twitter data to make them compatible with these data centers. After comparing databases, the study found that the MongoDB database was best suited for the storage of raw Twitter data due to its flexibility, ability to be accessed by multiple users, and querying functionality. The study used the Python package Zarr to transform processed Twitter data into a gridded format similar to that of satellite data. Each Tweet was mapped onto a time-space grid; each grid location contained information about Tweet attributes and precipitation. The study developed a pipeline for downloading, storing, and gridding Twitter data and transformed Twitter data into an understandable format for users of NASA satellite data.

Li, Rachel

Understanding Machine Learning in Earth Science: A Natural Language Processing Approach

Machine learning (ML) is being increasingly utilized in Earth science research. Benefits of ML include efficiency, reduction of human error, and ability to extract hidden patterns within data. However, the mutual lack of each other’s domain knowledge by ML and Earth science stands as a barrier to timely and effective implementation. Earth science, in particular, faces challenges in generating sample data, compared to those of traditional ML problems such as face recognition or stock predictions, where data is abundant and not lacking in ground truth, which is necessary for labeling. Earth science data are more varying in formats, such as HDF5 and image resolutions, and are not standardized across instruments, even within a given Earth science discipline. Previous studies have been done to outline the specific challenges that Earth science faces with ML, while others have focused on using existing publications to mine information efficiently. Other resources such as Scikit-Learn have developed decision trees for choosing appropriate machine learning algorithms, but application within Earth science subjects becomes much more complex. For the current study, we propose a methodology and tool that aids in implementation of ML in Earth science using natural language processing (NLP). Our work comprises three main parts: (1) analyzing existing publications related to ML and Earth science, using natural language processing: (2) extracting from the publications information on ML models subjects in Earth Science: and (3) visualizing the extracted relationships as a network graph. The resulting network graph should aid the Earth science communities in applying optimal ML algorithms and guiding data preparation through visualization of similar studies. The network graph and analysis of document similarity will be the basis of our next step, which is to develop a decision tree for selecting optimal machine learning methodologies for specified Earth science applications.

Zheng, Laura

Cloud Giovanni: Reining in Costs and Improving Performance with Analytical Data Stores Using Scalable Serverless Architecture

Giovanni is the Geospatial Interactive Online Visualization ANd aNalysis Infrastructure developed at NASA GES DISC which provides a simple and intuitive way to visualize, analyze, and access vast amounts of Earth science data. It receives large number of user requests each day for a variety of analysis and visualization services, which leads to the big data challenge of serving gradually increasing large data volumes with diverse statistical algorithms. We hereby propose a multi-dimensional accumulation method which provides fast and cost-efficient cloud analysis for diverse services including both area averaging and time averaging. This method involves the weighted volume integration over multiple variable dimensions (time and space), and is implemented in AWS using Athena providing serverless and highly scalable data analysis. Compared to the standard method, this approach dramatically reduced the computational time by order of magnitude with a minimal AWS cost incurred. For example, for a benchmark of 10-year area averaging over the 1x1 degree daily variable, the computational time was reduced from minutes to seconds, and the Athena cost is only $5 for 100,000 requests.

Zhang, Hailiang

Developing a Machine-Learning-Based Processing Framework for Twitter and Other Crowdsourced Data

Crowdsourced data streams such as Twitter and other social media are important sources of real-time and historical global information for Earth science applications. At the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), we have been exploring the Twitter data stream for its potential in augmenting the validation program of NASA's Global Precipitation Measurement (GPM) mission. To realize this potential, we need to increase the information density and enhance the quality of filtered precipitation tweets. We have implemented various components of a machine learning (ML)-based processing infrastructure for crowdsourced data that outputs, in this instance, useful and usable information derived from precipitation tweets. We have test enriched the Twitter stream with higher quality active tweets from those knowingly contributing to our effort and from existing crowdsourced programs (e.g., mPING, CoCoRaHS). We have experimented with various algorithms for processing tweets, including Naà ve Bayes, Convolutional Neural Network (CNN), Hierarchical Attention Network (HAN), and semi-supervised learning (with tri-training). Our current work focuses on (1) automated review of Earth science-related publications to determine relationships between discipline research needs and ML algorithms; (2) investigating Sequential Generative Adversarial Network (SeqGAN) for processing precipitation tweets for anomaly detection; and (3) managing crowdsourced data in a way that is compatible with existing NASA satellite data archives and using the data for ML applications. Key results include (1) network visualization of NLP-processed publications in various Earth science disciplines; (2) difference between GPM-linked, generated tweets and collected actual tweets that is small for GPM-determined light to moderate rain cases and high for GPM-determined heavy rain cases; and (3) identification of MongoDB for storing raw tweets and Zarr format for gridded tweets (compatible with GPM data). Our results have taken us a step closer to an operational ML-based tweet processing infrastructure and have already demonstrated that tweet-derived precipitation information is potentially useful for validation of Earth science satellite data.

Teng, William

Using NASA Earth Observation Data in ArcGIS

The NASA Goddard Earth Sciences Data and Information Services Center archives tens of thousands of Earth Observation (EO) parameters for land, atmosphere, and ocean. To facilitate GIS users to easily find, visualize, obtain, and analyze these EO data through, we developed an ArcGIS infrastructure with the Server, image services, Portal, and AOL. We will show how this capability supports broad GIS applications. Use cases including water management and air quality analyses will be demonstrated.

Wei, Jennifer

NASA GES DISC's Customized Services for Climatology and Meteorology

At the NASA Goddard Earth Sciences (GES) Data and Information Service Center (DISC), we have archived and distributed more than 2,400 Earth science data products, from different missions or projects containing more than 100 M data files/granules with a total volume size nearly 2 PB that broadly serve user needs in science areas such as Atmospheric Composition, Water & Energy Cycles and Climate Variability. To date, GES DISC has developed many pertinent services to facilitate the usage of data products by our research communities, represented by approximately 24,000 registered users. We are facing the big data with increasingly archival volume and data types, moreover, we also encounter increasing users' demands and the demands are more diversified. It is still a challenge for us to better understand exactly what our users' needs are, even after developing more than 70 services, including well-known online tools such as Giovanni and MERRA subsetter. In this presentation, we will try to address how we can accommodate the users' needs from two applicational user communities, Air Quality and Wind Energy, from data or service discovery to guide them properly utilize the data and services to fit their needs.

customizable services for climate and meteorology

Enriching the Twitter Stream Increasing Data Mining Yield and Quality Using Machine Learning

Social media data streams are important sources of real-time and historical global information for science applications. At the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), we are exploring the Twitter data stream for its potential in augmenting the validation program of NASA Earth science missions, specifically the Global Precipitation Measurement (GPM) mission. We have implemented a tweet processing infrastructure that outputs classified precipitation tweets. Inputs are "passive" tweets, along with a smaller number of tweets from "active" participants, i.e., those knowingly contributing to our effort. The "active" tweets, presumably of higher quality, enrich the Twitter stream. "Active" sources include data scraped from other social media (e.g., public Facebook posts) and data from existing crowdsourcing programs (e.g., mPING reports). In addition, there is likely relevant precipitation information in images and documents that are the end points of links often included in tweets. Information derived from these "active" sources could then be tweeted into the Twitter stream, thus enriching its quality. The objective of our current work is to mine these tweet­ linked images and documents, using neural networks, to increase the information content and quality related to precipitation. For images, we classified them as either precipitation-related or not. For training and validation, we used images obtained via the Google custom search API. We created two models: (1) by training a simple Convolutional Neural Network and (2) by using transfer learning principles to adapt a pre-trained object recognition model. For documents, both those linked to tweets and the tweet contents, we trained Hierarchical Attention Networks to determine precipitation occurrence, type, and intensity. For training and validation, we used a keyword-filtered tweet data set labelled with ground truth data from Dark Sky (an API to retrieve weather-related labels) and the National Severe Storms Laboratory's Multi­ Radar/Multi-Sensor (MRMS) system. Our results demonstrated the efficacy of our machine learning approaches for enriching the Twitter stream, to derive information potentially useful for validation of earth science satellite data.

Albayrak, Arif

Serving NASA GES DISC Multi-Spatiotemporal Earth Science Data to the GIS community

NASA Earth Science (ES) data is essential to a wide range of GIS research and applications. However, for many GIS users, searching, accessing, using and analyzing NASA ES data can be of a great challenge- ranging from the sheer data volumes, types of science parameters, and to the complexity of data encoding formats. As one of the twelve NASA Science Mission Directorate (SMD) Data Centers, Goddard Earth Sciences (GES) Data and Information Services Center (DISC) archives and distributes petabytes of ES parameters covering atmosphere, land, and ocean fields. Most data are multidimensional and multi-spatiotemporal in nature and are encoded in different science data formats (e.g, HDF, HDF-EOS, netCDF, GRIB, binary), which usually contain multiple variables and different metadata information. By far, GES DISC has been developing a number of services and online tools to help GIS users to easily explore our data products. In this presentation, we will describe our ArcGIS-based data accessing and visualization services and portals, which allow users directly exploring the multi-spatiotemporal ES data in ArcGIS clients without having to pre-download/import the data. The ArcGIS services are also compliant with the Open Geospatial Consortium (OGC) Web Coverage Service (WCS) and Web Map Service (WMS) protocols and can be accessed by any other WCS/WMS clients to get customized GES DISC EO data on-the-fly from such services.

Wei, Jennifer

Giovanni in the Cloud: Earth Science Data Exploration in Amazon Web Services

Giovanni is an exploration tool at the NASA Goddard Earth Sciences Data Information Services Center (GES DISC), providing 22 analysis and visualization services for over 1600 Earth Science data variables. Owing to its popularity, Giovanni has experienced a consistent growth in overall demand, with periodic usage spikes attributed to trainings by education organizations, extensive data analysis in response to natural disasters, preparations for science meetings, etc. Furthermore, the new generation of spaceborne sensors and high resolution models have resulted in an exponential growth in data volume with data distributed across the traditional boundaries of data centers. Seamless exploration of data (without users having to worry about data center boundaries) has been a key recommendation of the GES DISC User Working Group. These factors have required new strategies for delivering acceptable performance. The cloud-based Giovanni, built on Amazon Web Services (AWS), evaluates (1) AWS native solutions to provide a scalable, serverless architecture; (2) open standards for data storage in the Cloud; (3) a cost model for operations; and (4) end-user performance. Our preliminary findings indicate that the use of serverless architecture has a potential to significantly reduce development and operational cost of Giovanni. The combination of using AWS managed services, storage of data in open standards, and schema-on-read data access strategy simplifies data access and analytics, in addition to making data more accessible to the end users of Giovanni through popular programming languages.

Giovanni

Use of Schema on Read in Earth Science Data Archives

Traditionally, NASA Earth Science data archives have file-based storage using proprietary data file formats, such as HDF and HDF-EOS, which are optimized to support fast and efficient storage of spaceborne and model data as they are generated. The use of file-based storage essentially imposes an indexing strategy based on data dimensions. In most cases, NASA Earth Science data uses time as the primary index, leading to poor performance in accessing data in spatial dimensions. For example, producing a time series for a single spatial grid cell involves accessing a large number of data files. With exponential growth in data volume due to the ever-increasing spatial and temporal resolution of the data, using file-based archives poses significant performance and cost barriers to data discovery and access. Storing and disseminating data in proprietary data formats imposes an additional access barrier for users outside the mainstream research community. At the NASA Goddard Earth Sciences Data Information Services Center (GES DISC), we have evaluated applying the schema-on-read principle to data access and distribution. We used Apache Parquet to store geospatial data, and have exposed data through Amazon Web Services (AWS) Athena, AWS Simple Storage Service (S3), and Apache Spark. Using the schema-on-read approach allows customization of indexing spatially or temporally to suit the data access pattern. The storage of data in open formats such as Apache Parquet has widespread support in popular programming languages. A wide range of solutions for handling big data lowers the access barrier for all users. This presentation will discuss formats used for data storage, frameworks with This presentation will discuss formats used for data storage, frameworks with support for schema-on-read used for data access, and common use cases covering data usage patterns seen in a geospatial data archive.

cloud applications

Challenges in Obtaining and Visualizing Satellite Level 2 Data in GIS

Satellite data products are important for a wide variety of applications that can bring far-reaching benefits to the science community and the broader society. These benefits can best be achieved if the satellite data are well utilized and interpreted. Unfortunately, this is not always the case, despite the abundance and relative maturity of numerous satellite data products provided by NASA and other organizations. One way to help users better understand the satellite data is to provide data along with Images, including accurate pixel coverage area delineation, and science team recommended quality screening for individual geophysical parameters. However, there are challenges of visualizing remote sensed non-gridded products: (1) different geodetics of space-borne instruments (2) data often arranged in a long-track and a cross-track axes (3) spatially and temporally continuous data chunked into granule files: data for a portion (or all) of a satellite orbit (4) no general rule of resampling or interpolations to a grid (5) geophysical retrieval only based on pixel center location without shape information. In this presentation, we will unravel a new Goddard Earth Sciences Data and Information Services Center (GES DISC) Level 2 (L2) visualization on-demand service. The service's front end provides various visualization and data accessing capabilities, such as overlay and swipe of multiply variables and subset and download of data in different formats. The backend of the service consists of Open Geospatial Consortium (OGC) standard-compliant Web Mapping Service (WMS) and Web Coverage Service. The infrastructure allows inclusion of outside data sources served in OGC compliant protocols and allows other interoperable clients, such as ArcGIS clients, to connect to our L2 WCS/WMS.

GI

New GES DISC Services Shortening the Path in Science Data Discovery

The Current GES DISC available services only allow user to select variables from a single dataset at a time and too many variables from a dataset are displayed, choice is hard. At American Geophysical Union (AGU) 2016 Fall Meeting, Goddard Earth Sciences Data Information Services Center (GES DISC) unveiled a new service: Datalist. A Datalist is a collection of predefined or user-defined data variables from one or more archived datasets. Our science support team curated predefined datalist and provided value to the user community. Imagine some novice user wants to study hurricane and typed in hurricane in the search box. The first item in the search result is GES DISC provided Hurricane Datalist. It contains scientists recommended variables from multiple datasets like TRMM, GPM, MERRA, etc. Datalist uses the same architecture as that of our new website, which also provides one-stop shopping for data, metadata, citation, documentation, visualization and other available services.We implemented Datalist with new GES DISC web architecture, one single web page that unified all user interfaces. From that webpage, users can find data by either type in keyword, or browse by category. It also provides user with a sophisticated integrated data and services package, including metadata, citation, documentation, visualization, and data-specific services, all available from one-stop shopping.

Datalist

Analyzing and Visualizing Precipitation and Soil Moisture in ArcGIS

Precipitation and soil moisture are among the most important parameters in many land GIS (Geographic Information System) research and applications. These data are available globally from NASA GES DISC (Goddard Earth Science Data and Information Services Center) in GIS-ready format at 10-kilometer spatial resolution and 24-hour or less temporal resolutions. In this presentation, well demonstrate how rainfall and soil moisture data are used in ArcGIS to analyze and visualize spatiotemporal patterns of droughts and their impacts on natural vegetation and agriculture in different parts of the world.

precipitation

Exploring NASA GES DISC Data with Interoperable Services

Overview of NASA GES DISC (NASA Goddard Earth Science Data and Information Services Center) data with interoperable services: Open-standard and Interoperable Services Improve data discoverability, accessibility, and usability with metadata, catalogue and portal standards Achieve data, information and knowledge sharing across applications with standardized interfaces and protocols Open Geospatial Consortium (OGC) Data Services and Specifications Web Coverage Service (WCS) -- data Web Map Service (WMS) -- pictures of data Web Map Tile Service (WMTS) --- pictures of data tiles Styled Layer Descriptors (SLD) --- rendered styles.

Giovanni

Open Source GIS Connectors to NASA GES DISC Satellite Data

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) houses a suite of high spatiotemporal resolution GIS data including satellite-derived and modeled precipitation, air quality, and land surface parameter data. The data are valuable to various GIS research and applications at regional, continental, and global scales. On the other hand, many GIS users, especially those from the ArcGIS community, have difficulties in obtaining, importing, and using our data due to factors such as the variety of data products, the complexity of satellite remote sensing data, and the data encoding formats. We introduce a simple open source ArcGIS data connector that significantly simplifies the access and use of GES DISC data in ArcGIS.

user

Exploring NASA OMI Level 2 Data With Visualization

Satellite data products are important for a wide variety of applications that can bring far-reaching benefits to the science community and the broader society. These benefits can best be achieved if the satellite data are well utilized and interpreted, such as model inputs from satellite, or extreme events (such as volcano eruptions, dust storms, etc.).Unfortunately, this is not always the case, despite the abundance and relative maturity of numerous satellite data products provided by NASA and other organizations. Such obstacles may be avoided by allowing users to visualize satellite data as images, with accurate pixel-level (Level-2) information, including pixel coverage area delineation and science team recommended quality screening for individual geophysical parameters. We present a prototype service from the Goddard Earth Sciences Data and Information Services Center (GES DISC) supporting Aura OMI Level-2 Data with GIS-like capabilities. Functionality includes selecting data sources (e.g., multiple parameters under the same scene, like NO2 and SO2, or the same parameter with different aggregation methods, like NO2 in OMNO2G and OMNO2D products), user-defined area-of-interest and temporal extents, zooming, panning, overlaying, sliding, and data subsetting, reformatting, and reprojection. The system will allow any user-defined portal interface (front-end) to connect to our backend server with OGC standard-compliant Web Mapping Service (WMS) and Web Coverage Service (WCS) calls. This back-end service should greatly enhance its expandability to integrate additional outside data-map sources.

Level 2

Exploring NASA OMI Level 2 Data With Visualization

Satellite data products are important for a wide variety of applications that can bring far-reaching benefits to the science community and the broader society. These benefits can best be achieved if the satellite data are well utilized and interpreted, such as model inputs from satellite, or extreme events (such as volcano eruptions, dust storms,... etc.). Unfortunately, this is not always the case, despite the abundance and relative maturity of numerous satellite data products provided by NASA and other organizations. Such obstacles may be avoided by allowing users to visualize satellite data as "images", with accurate pixel-level (Level-2) information, including pixel coverage area delineation and science team recommended quality screening for individual geophysical parameters. We present a prototype service from the Goddard Earth Sciences Data and Information Services Center (GES DISC) supporting Aura OMI Level-2 Data with GIS-like capabilities. Functionality includes selecting data sources (e.g., multiple parameters under the same scene, like NO2 and SO2, or the same parameter with different aggregation methods, like NO2 in OMNO2G and OMNO2D products), user-defined area-of-interest and temporal extents, zooming, panning, overlaying, sliding, and data subsetting, reformatting, and reprojection. The system will allow any user-defined portal interface (front-end) to connect to our backend server with OGC standard-compliant Web Mapping Service (WMS) and Web Coverage Service (WCS) calls. This back-end service should greatly enhance its expandability to integrate additional outside data/map sources.

Aura

Processing NASA Earth Science Data on Nebula Cloud

Three applications were successfully migrated to Nebula, including S4PM, AIRS L1/L2 algorithms, and Giovanni MAPSS. Nebula has some advantages compared with local machines (e.g. performance, cost, scalability, bundling, etc.). Nebula still faces some challenges (e.g. stability, object storage, networking, etc.). Migrating applications to Nebula is feasible but time consuming. Lessons learned from our Nebula experience will benefit future Cloud Computing efforts at GES DISC.

Chen, Aijun