Engineering PapersSearch

Engineering topics

Albayrak, Arif

Publications and source records attributed to Albayrak, Arif.

Citizen Science Twitter Data Management for Earth Science Applications

Social media data can provide useful real-time and historical information relating to the natural world, but managing this data poses challenges. Scientists at GES DISC are exploring the potential of Twitter data to augment precipitation data from the Global Precipitation Measurement (GPM) mission. However, the format of Twitter data is unconventional in the context of NASA data centers, resulting in frustration for scientists who need to work with the data. This study investigated procedures and standards needed to properly manage Twitter data to make them compatible with these data centers. After comparing databases, the study found that the MongoDB database was best suited for the storage of raw Twitter data due to its flexibility, ability to be accessed by multiple users, and querying functionality. The study used the Python package Zarr to transform processed Twitter data into a gridded format similar to that of satellite data. Each Tweet was mapped onto a time-space grid; each grid location contained information about Tweet attributes and precipitation. The study developed a pipeline for downloading, storing, and gridding Twitter data and transformed Twitter data into an understandable format for users of NASA satellite data.

Li, Rachel

Understanding Machine Learning in Earth Science: A Natural Language Processing Approach

Machine learning (ML) is being increasingly utilized in Earth science research. Benefits of ML include efficiency, reduction of human error, and ability to extract hidden patterns within data. However, the mutual lack of each other’s domain knowledge by ML and Earth science stands as a barrier to timely and effective implementation. Earth science, in particular, faces challenges in generating sample data, compared to those of traditional ML problems such as face recognition or stock predictions, where data is abundant and not lacking in ground truth, which is necessary for labeling. Earth science data are more varying in formats, such as HDF5 and image resolutions, and are not standardized across instruments, even within a given Earth science discipline. Previous studies have been done to outline the specific challenges that Earth science faces with ML, while others have focused on using existing publications to mine information efficiently. Other resources such as Scikit-Learn have developed decision trees for choosing appropriate machine learning algorithms, but application within Earth science subjects becomes much more complex. For the current study, we propose a methodology and tool that aids in implementation of ML in Earth science using natural language processing (NLP). Our work comprises three main parts: (1) analyzing existing publications related to ML and Earth science, using natural language processing: (2) extracting from the publications information on ML models subjects in Earth Science: and (3) visualizing the extracted relationships as a network graph. The resulting network graph should aid the Earth science communities in applying optimal ML algorithms and guiding data preparation through visualization of similar studies. The network graph and analysis of document similarity will be the basis of our next step, which is to develop a decision tree for selecting optimal machine learning methodologies for specified Earth science applications.

Zheng, Laura

Bringing Analysis Closer to Data: Developing a Visualization Tool for L2 Earth Science Satellite Data

Earth Science satellite missions provide a unique opportunity for scientists to visualize complex and multifaceted observations projected geospatially across maps of the Earth. While visualization tools can help scientists comprehend, analyze, and share data, visualizing Level-2 Earth Sciences data poses its own specific set of challenges. Since the geospatial information in Level-2 data files is stored as independent variables, the plotting process involves matching dimensional information from latitude and longitude with a desired variable. Variables are stored in different ways across various Earth Science data file formats, which complicates the process of extracting data and plotting variables from a given file without requiring extensive user input and prerequisite familiarity with the file type variable structure. In coordination with NASA’s Goddard Earth Sciences Data Information Services Center (GES DISC), the team developed a Level-2 Earth Science data visualization tool that aims to address some of the complexities associated with plotting Level-2 data. This tool offers command-line and user interface support for file and variable selection to accommodate varying use cases and degrees of user familiarity with the structure of a given file. The visualization tool is written in Python 3 and utilizes a modular approach to facilitate continued expansion and reuse. In addressing some common complications involved in plotting Level-2 Earth Sciences data, the tool aims to help to link the process of analysis more directly with data acquisition and visualization, bringing analysis closer to data across levels of processing.

Li, Angela W.

Developing a Machine-Learning-Based Processing Framework for Twitter and Other Crowdsourced Data

Crowdsourced data streams such as Twitter and other social media are important sources of real-time and historical global information for Earth science applications. At the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), we have been exploring the Twitter data stream for its potential in augmenting the validation program of NASA's Global Precipitation Measurement (GPM) mission. To realize this potential, we need to increase the information density and enhance the quality of filtered precipitation tweets. We have implemented various components of a machine learning (ML)-based processing infrastructure for crowdsourced data that outputs, in this instance, useful and usable information derived from precipitation tweets. We have test enriched the Twitter stream with higher quality active tweets from those knowingly contributing to our effort and from existing crowdsourced programs (e.g., mPING, CoCoRaHS). We have experimented with various algorithms for processing tweets, including Naà ve Bayes, Convolutional Neural Network (CNN), Hierarchical Attention Network (HAN), and semi-supervised learning (with tri-training). Our current work focuses on (1) automated review of Earth science-related publications to determine relationships between discipline research needs and ML algorithms; (2) investigating Sequential Generative Adversarial Network (SeqGAN) for processing precipitation tweets for anomaly detection; and (3) managing crowdsourced data in a way that is compatible with existing NASA satellite data archives and using the data for ML applications. Key results include (1) network visualization of NLP-processed publications in various Earth science disciplines; (2) difference between GPM-linked, generated tweets and collected actual tweets that is small for GPM-determined light to moderate rain cases and high for GPM-determined heavy rain cases; and (3) identification of MongoDB for storing raw tweets and Zarr format for gridded tweets (compatible with GPM data). Our results have taken us a step closer to an operational ML-based tweet processing infrastructure and have already demonstrated that tweet-derived precipitation information is potentially useful for validation of Earth science satellite data.

Teng, William

Enriching the Twitter Stream Increasing Data Mining Yield and Quality Using Machine Learning

Social media data streams are important sources of real-time and historical global information for science applications. At the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), we are exploring the Twitter data stream for its potential in augmenting the validation program of NASA Earth science missions, specifically the Global Precipitation Measurement (GPM) mission. We have implemented a tweet processing infrastructure that outputs classified precipitation tweets. Inputs are "passive" tweets, along with a smaller number of tweets from "active" participants, i.e., those knowingly contributing to our effort. The "active" tweets, presumably of higher quality, enrich the Twitter stream. "Active" sources include data scraped from other social media (e.g., public Facebook posts) and data from existing crowdsourcing programs (e.g., mPING reports). In addition, there is likely relevant precipitation information in images and documents that are the end points of links often included in tweets. Information derived from these "active" sources could then be tweeted into the Twitter stream, thus enriching its quality. The objective of our current work is to mine these tweet­ linked images and documents, using neural networks, to increase the information content and quality related to precipitation. For images, we classified them as either precipitation-related or not. For training and validation, we used images obtained via the Google custom search API. We created two models: (1) by training a simple Convolutional Neural Network and (2) by using transfer learning principles to adapt a pre-trained object recognition model. For documents, both those linked to tweets and the tweet contents, we trained Hierarchical Attention Networks to determine precipitation occurrence, type, and intensity. For training and validation, we used a keyword-filtered tweet data set labelled with ground truth data from Dark Sky (an API to retrieve weather-related labels) and the National Severe Storms Laboratory's Multi­ Radar/Multi-Sensor (MRMS) system. Our results demonstrated the efficacy of our machine learning approaches for enriching the Twitter stream, to derive information potentially useful for validation of earth science satellite data.

Albayrak, Arif

Mining Twitter Data to Augment NASA GPM Validation

The Twitter data stream is an important new source of real-time and historical global information for potentially augmenting the validation program of NASA's Global Precipitation Measurement (GPM) mission. There have been other similar uses of Twitter, though mostly related to natural hazards monitoring and management. The validation of satellite precipitation estimates is challenging, because many regions lack data or access to data, especially outside of the U.S. and in remote and developing areas. The time-varying set of "precipitation" tweets can be thought of as an organic network of rain gauges, potentially providing a widespread view of precipitation occurrence. Twitter provides a large source of crowd for crowdsourcing. During a 24-hour period in the middle of the snow storm this past March in the U.S. Northeast, we collected more than 13,000 relevant precipitation tweets with exact geolocation. The overall objective of our project is to determine the extent to which processed tweets can provide additional information that improves the validation of GPM data. Though our current effort focuses on tweets and precipitation, our approach is general and applicable to other social media and other geophysical measurements. Specifically, we have developed an operational infrastructure for processing tweets, in a format suitable for analysis with GPM data; engaged with potential participants, both passive and active, to "enrich" the Twitter stream; and inter-compared "precipitation" tweet data, ground station data, and GPM retrievals. In this presentation, we detail the technical capabilities of our tweet processing infrastructure, including data abstraction, feature extraction, search engine, context-awareness, real-time processing, and high volume (big) data processing; various means for "enriching" the Twitter stream; and results of inter-comparisons. Our project should bring a new kind of visibility to Twitter and engender a new kind of appreciation of the value of Twitter by the science research communities.

validatio

Complexities in Subsetting Level 2 Data

Satellite Level 2 data presents unique challenges for tools and services. From nonlinear spatial geometry to inhomogeneous file data structure to inconsistent temporal variables to complex data variable dimensionality to multiple file formats, there are many difficulties in creating general tools for Level 2 data support. At NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), we are implementing a general Level 2 Subsetting service for Level 2 data to a user-specified spatio-temporal region of interest (ROI). In this presentation, we will unravel some of the challenges faced in creating this service and the strategies we used to surmount them.

Data acces

Framework for Processing Citizens Science Data for Applications to NASA Earth Science Missions

Citizen science (or crowdsourcing) has drawn much high-level recent and ongoing interest and support. It is poised to be applied, beyond the by-now fairly familiar use of, e.g., Twitter for natural hazards monitoring, to science research, such as augmenting the validation of NASA earth science mission data. This interest and support is seen in the 2014 National Plan for Civil Earth Observations, the 2015 White House forum on citizen science and crowdsourcing, the ongoing Senate Bill 2013 (Crowdsourcing and Citizen Science Act of 2015), the recent (August 2016) Open Geospatial Consortium (OGC) call for public participation in its newly-established Citizen Science Domain Working Group, and NASA's initiation of a new Citizen Science for Earth Systems Program (along with its first citizen science-focused solicitation for proposals). Over the past several years, we have been exploring the feasibility of extracting from the Twitter data stream useful information for application to NASA precipitation research, with both "passive" and "active" participation by the twitterers. The Twitter database, which recently passed its tenth anniversary, is potentially a rich source of real-time and historical global information for science applications. The time-varying set of "precipitation" tweets can be thought of as an organic network of rain gauges, potentially providing a widespread view of precipitation occurrence. The validation of satellite precipitation estimates is challenging, because many regions lack data or access to data, especially outside of the U.S. and in remote and developing areas. Mining the Twitter stream could augment these validation programs and, potentially, help tune existing algorithms. Our ongoing work, though exploratory, has resulted in key components for processing and managing tweets, including the capabilities to filter the Twitter stream in real time, to extract location information, to filter for exact phrases, and to plot tweet distributions. The key step is to process the "precipitation" tweets to be compatible with satellite-retrieved precipitation data. These key components for processing and managing "precipitation" tweets (and additional ones to be developed) are not limited to precipitation, nor are they limited to the Twitter social medium. Indeed, to maximize the value of our work for NASA earth science programs, these components should be generalized and be part of an overall framework for processing citizen science data for science research. In this paper, we outline such a framework.

earth science satellite data

Characterize Aerosols from MODIS MISR OMI MERRA-2: Dynamic Image Browse Perspective

Among the known atmospheric constituents, aerosols still represent the greatest uncertainty in climate research. To understand the uncertainty is to bring altogether of observational (in-situ and remote sensing) and modeling datasets and inter-compare them synergistically for a wide variety of applications that can bring far-reaching benefits to the science community and the broader society. These benefits can best be achieved if these earth science data (satellite and modeling) are well utilized and interpreted. Unfortunately, this is not always the case, despite the abundance and relative maturity of numerous satellite-borne sensors routinely measure aerosols. There is often disagreement between similar aerosol parameters retrieved from different sensors, leaving users confused as to which sensors to trust for answering important science questions about the distribution, properties, and impacts of aerosols. NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) have developed a new visualization service (NASA Level 2 Data Quality Visualization, DQViz)supporting various visualization and data accessing capabilities from satellite Level 2(MODISMISROMI) and long term assimilated aerosols from NASA Modern-Era Retrospective analysis for Research and Applications, Version 2 (MERRA-2 displaying at their own native physical-retrieved spatial resolution. Functionality will include selecting data sources (e.g., multiple parameters under the same measurement), defining area-of-interest and temporal extents, zooming, panning, overlaying, sliding, and data subsetting and reformatting.

data quality

Application of Data Cubes for Improving Detection of Water Cycle Extreme Events

As part of an ongoing NASA-funded project to remove a longstanding barrier to accessing NASA data (i.e., accessing archived time-step array data as point-time series), for the hydrology and other point-time series-oriented communities, "data cubes" are created from which time series files (aka "data rods") are generated on-the-fly and made available as Web services from the Goddard Earth Sciences Data and Information Services Center (GES DISC). Data cubes are data as archived rearranged into spatio-temporal matrices, which allow for easy access to the data, both spatially and temporally. A data cube is a specific case of the general optimal strategy of reorganizing data to match the desired means of access. The gain from such reorganization is greater the larger the data set. As a use case of our project, we are leveraging existing software to explore the application of the data cubes concept to machine learning, for the purpose of detecting water cycle extreme events, a specific case of anomaly detection, requiring time series data. We investigate the use of support vector machines (SVM) for anomaly classification. We show an example of detection of water cycle extreme events, using data from the Tropical Rainfall Measuring Mission (TRMM).

water cycle extreme events

CO2 Data Distribution and Support from the Goddard Earth Science Data and Information Services Center (GES-DISC)

This talk will describe the support and distribution of CO2 data products from OCO-2, AIRS, and ACOS, that are archived and distributed from the Goddard Earth Sciences Data and Information Services Center. We will provide a brief summary of the current online archive and distribution metrics for the OCO-2 Level 1 products and plans for the Level 2 products. We will also describe collaborative data sets and services (e.g., matchups with other sensors) and solicit feedback for potential future services.

AIRS

Estimation and Bias Correction of Aerosol Abundance using Data-driven Machine Learning and Remote Sensing

Air quality information is increasingly becoming a public health concern, since some of the aerosol particles pose harmful effects to peoples health. One widely available metric of aerosol abundance is the aerosol optical depth (AOD). The AOD is the integrated light extinction coefficient over a vertical atmospheric column of unit cross section, which represents the extent to which the aerosols in that vertical profile prevent the transmission of light by absorption or scattering. The comparison between the AOD measured from the ground-based Aerosol Robotic Network (AERONET) system and the satellite MODIS instruments at 550 nm shows that there is a bias between the two data products. We performed a comprehensive analysis exploring possible factors which may be contributing to the inter-instrumental bias between MODIS and AERONET. The analysis used several measured variables, including the MODIS AOD, as input in order to train a neural network in regression mode to predict the AERONET AOD values. This not only allowed us to obtain an estimate, but also allowed us to infer the optimal sets of variables that played an important role in the prediction. In addition, we applied machine learning to infer the global abundance of ground level PM2.5 from the AOD data and other ancillary satellite and meteorology products. This research is part of our goal to provide air quality information, which can also be useful for global epidemiology studies.

Malakar, Nabin K.

MODIS Aerosol Optical Depth Bias Adjustment Using Machine Learning Algorithms

To monitor the earth atmosphere and its surface changes, satellite based instruments collect continuous data. While some of the data is directly used, some others such as aerosol properties are indirectly retrieved from the observation data. While retrieved variables (RV) form very powerful products, they don't come without obstacles. Different satellite viewing geometries, calibration issues, dynamically changing atmospheric and earth surface conditions, together with complex interactions between observed entities and their environment affect them greatly. This results in random and systematic errors in the final products.

Albayrak, Arif