Engineering PapersSearch

Engineering topics

Li, Angela

Publications and source records attributed to Li, Angela.

Simplifying NASA Earth Science Data and Information Access Through Natural Language Processing Based Data Analysis and Visualization

NASA Earth science data collected from satellites, model assimilation, airborne missions, and field campaigns, are large, complex and evolving. Such characteristics pose great challenges for end users (e.g., Earth science and applied science users, students, citizen scientists), particularly for those who are unfamiliar with NASA's EOSDIS and thus unable to access and utilize datasets effectively. For example, a novice user may simply ask: what is the total rainfall for a flooding event in my county yesterday? For an experienced user (e.g., algorithm developer), a question can be: how did my rainfall product perform, compared to ground observations, during a flooding event? Nonetheless, with rapid information technology development such as natural language processing, it is possible to develop simplified Web interfaces and back-end processing components to handle such questions and deliver answers in terms of text, data, or graphic results directly to users.In this presentation, we describe the main challenges for end users with different levels of expertise in accessing and utilizing NASA Earth science data. Surveys reveal that most non-professional users normally do not want to download and handle raw data as well as conduct heavy-duty data processing tasks. Often they just want some simple graphics or data for various purposes. To them, simple and intuitive user interfaces are sufficient because complicated ones can be difficult and time-consuming to learn. Professionals also want such interfaces to answer many questions from datasets. One solution is to develop a natural language based search box like Google and the search results can be text, data, graphics and more. Now the challenge is, with natural language processing, can we design a system to process a scientific question typed in by a user? In this presentation, we describe our plan for such a prototype. The workflow is: 1) extract needed information (e.g., variables, spatial and temporal information, processing methods, etc.) from the input, 2) process the data in the backend, and 3) deliver the results (data or graphics) to the user.

Liu, Zhong

Integrated Analysis of Multiple User Metrics - A “Sequel”; and Introducing the Google Analytic

For decades, the Goddard Earth Sciences Data and Information Services Center (GES DISC) has archived and distributed enormous volumes of NASA Earth science data (accompanied with many developed tools and services) to various research/applications communities and the general public. Being “immersed” in the Big Data era, we have inevitably faced the challenges of our continually increasing archived data in both volume and variety, as well as enhanced user needs and demands. In recent years, we have actively analyzed different types of user metrics, such as operational distribution metrics (recording numbers of distinct users and downloaded data files, size of distributed data volume): user publication metrics (mining info from our Giovanni users’ publications): and Bugzilla metrics (collecting info from user questions or feedback from user assistance tickets). Such metrics have helped us achieve a better understanding of user needs, demands, characteristics, and behaviors, which has then helped us improve our user services. Now we will present a “Sequel” of integrated analysis of multiple metrics at the GES DISC by introducing and adding one new kind of metrics acquired via utilizing our recently implemented Google Analytic 360 suite. Several “newer” reports, e.g., “What web site features and links are the most popular (and least)?” and “What are the top 25 dataset Keyword searches?” retrieved from this new metrics set will be presented, along with the aforementioned “traditional” metrics results.

Shie, Chung-Lin

Bringing Analysis Closer to Data: Developing a Visualization Tool for L2 Earth Science Satellite Data

Earth Science satellite missions provide a unique opportunity for scientists to visualize complex and multifaceted observations projected geospatially across maps of the Earth. While visualization tools can help scientists comprehend, analyze, and share data, visualizing Level-2 Earth Sciences data poses its own specific set of challenges. Since the geospatial information in Level-2 data files is stored as independent variables, the plotting process involves matching dimensional information from latitude and longitude with a desired variable. Variables are stored in different ways across various Earth Science data file formats, which complicates the process of extracting data and plotting variables from a given file without requiring extensive user input and prerequisite familiarity with the file type variable structure. In coordination with NASA’s Goddard Earth Sciences Data Information Services Center (GES DISC), the team developed a Level-2 Earth Science data visualization tool that aims to address some of the complexities associated with plotting Level-2 data. This tool offers command-line and user interface support for file and variable selection to accommodate varying use cases and degrees of user familiarity with the structure of a given file. The visualization tool is written in Python 3 and utilizes a modular approach to facilitate continued expansion and reuse. In addressing some common complications involved in plotting Level-2 Earth Sciences data, the tool aims to help to link the process of analysis more directly with data acquisition and visualization, bringing analysis closer to data across levels of processing.

Li, Angela W.

NASA Global Satellite and Model Data Products and Services for Tropical Cyclone Research

The lack of observations over vast tropical oceans is a major challenge for tropical cyclone research. Satellite observations and model reanalysis data play an important role in filling these- gaps. Established in the mid-1980's, the Goddard Earth Sciences Data and Information Services Center (GES DISC), as one of the 12 NASA data centers, archives and distributes data from several Earth science disciplines such as precipitation, atmospheric dynamics, atmospheric composition, hydrology, including well-known NASA satellite missions (e.g. TRMM, GPM) and model assimilation projects (MERRA-2). Acquiring datasets suitable for tropical cyclone research in a large data archive is a challenge for many, especially for those who are not familiar with satellite or model data. Over the years, the GES DISC has developed user-friendly data services. For example, Giovanni is an online visualization and analysis tool, allowing users to visualize and analyze over 2000 satellite- and model-based variables with a Web browser, without downloading data and software. In this chapter, we will describe data and services at the GES DISC with emphasis on tropical cyclone research. We will also present two case studies and discuss future plans.

Liu, Zhong

MERRA-2 Data and Analytic Services at NASA GES DISC for Climate Extremes Study

NASA's climate reanalysis datasets from the Modern Era Retrospective-analysis for Research and Applications, Version 2 (MERRA-2) contains numerous long-term atmosphere, land, and ocean data products from 1980-present. MERRA-2 datasets, such as precipitation, soil moisture, and temperature, have been used widely to study extreme events. The native archived MERRA-2 data files are day-file (hourly time interval) and month-file, containing up to 125 parameters in one file. Due to the large number of data files and volumes, it is challenging for users, especially the applications research community, to handle the original hourly data files for long time periods to analyze extreme events. In this presentation, we review MERRA-2 data for studies of extreme conditions, and demonstrate analytic services at the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). One of the current operational services, 'subsetter', allows users to download only specific data of interest, i.e. data selected by parameter, region, and time period. New services are under development that will provide more 'on-the-fly' statistical calculations when downloading data; improve efficiency when accessing long time-series data. We will provide additional "How-to" resources that include step-by-step instructions on data access and usage. We have tested restructuring of day-files in an optimized data cube, which has significantly improved system performance for accessing long time-series. Overall performance is associated with cube size and structure, data compression method, and how the data are accessed. The optimized data cube structure will enable better online analytic services for statistical analysis and extreme events mining. To demonstrate the service, we use an extreme drought associated with the anomalous 2016 monsoon over southern Asia. This prototype time-series service may be augmented in the cloud infrastructure in the future.

data access

NASA GES DISC's Customized Services for Climatology and Meteorology

At the NASA Goddard Earth Sciences (GES) Data and Information Service Center (DISC), we have archived and distributed more than 2,400 Earth science data products, from different missions or projects containing more than 100 M data files/granules with a total volume size nearly 2 PB that broadly serve user needs in science areas such as Atmospheric Composition, Water & Energy Cycles and Climate Variability. To date, GES DISC has developed many pertinent services to facilitate the usage of data products by our research communities, represented by approximately 24,000 registered users. We are facing the big data with increasingly archival volume and data types, moreover, we also encounter increasing users' demands and the demands are more diversified. It is still a challenge for us to better understand exactly what our users' needs are, even after developing more than 70 services, including well-known online tools such as Giovanni and MERRA subsetter. In this presentation, we will try to address how we can accommodate the users' needs from two applicational user communities, Air Quality and Wind Energy, from data or service discovery to guide them properly utilize the data and services to fit their needs.

customizable services for climate and meteorology

Improving NASA Earth Science Data and Information Access Through Natural Language Processing Based Data Analysis and Visualization

NASA: The Research Access initiative is part of the agency's framework for increasing public access to scientific publications and digital scientific data. The initiative follows the release of White House Office of Science and Technology Policy's (OSTP) memorandum "Increasing Access to the Results of Federally Funded Research," to ensure federally funded research is available to the public within one year of publication. NASA answered the mandate by creating an agency plan entitled "NASA Plan for Increasing Access to the Results of Scientific Research" and associated policy, NPD 2230.1, Research Data and Publication Access. Principles in NASA SMD Strategic Plan for Scientific Data and Computing: Continued free and open access to scientific data for any use. Improved ease of use and discoverability. Enhanced science applications and new use cases. Incorporates best practices and "state of the art" through partnerships. Earth Data and Systems are Evolving: Increasing archive and file sizes. More complicated data structures. More user-friendly and data services. What is the future direction?

Liu, Zhong

Global Satellite Observations for Smart Cities

The smart city approach requires collection of interdisciplinary data and information from multiple sources and integration with modern technologies to provide a new and cost-effective way for researchers and decision makers to study and manage cities. In this book chapter, we introduce NASA satellite-based global and regional observations with emphasis on the hydrologic cycle (e.g., precipitation, wind, temperature, soil moisture) for smart cities. These products, consisting of both near-real-time and historical datasets, are publicly available free of charge and can be used for global and regional research and applications. Examples of using these datasets in smart cities are included. The chapter is organized as follows, first, a brief overview of NASA global satellite-based data products, followed by data services and tools, two examples of using satellite-based datasets in megacities, and finally summary and future plans.

hydrology

The Value of Data and Metadata Standardization for Interoperability in Giovanni Or: Why Your Product's Metadata Causes Us Headaches!

Giovanni is a data exploration and visualization tool at the NASA Goddard Earth Sciences Data Information Services Center (GES DISC). It has been around in one form or another for more than 15 years. Giovanni calculates simple statistics and produces 22 different visualizations for more than 1600 geophysical parameters from more than 90 satellite and model products. Giovanni relies on external data format standards to ensure interoperability, including the NetCDF CF Metadata Conventions. Unfortunately, these standards were insufficient to make Giovanni's internal data representation truly simple to use. Finding and working with dimensions can be convoluted with the CF Conventions. Furthermore, the CF Conventions are silent on machine-friendly descriptive metadata such as the parameter's source product and product version. In order to simplify analyzing disparate earth science data parameters in a unified way, we developed Giovanni's internal standard. First, the format standardizes parameter dimensions and variables so they can be easily found. Second, the format adds all the machine-friendly metadata Giovanni needs to present our parameters to users in a consistent and clear manner. At a glance, users can grasp all the pertinent information about parameters both during parameter selection and after visualization.

interoperability

A-Train Datalist - A New GES DISC Service to Allow One-Stop Shopping for A-Train Data

The currently available services at the Goddard Earth Sciences Data Information Services Center (GES DISC) only allow users to select variables from a single data set at a time. Because entire variables from a data set are often displayed, user selection of variables of interest can be overwhelming. At the American Geophysical Union (AGU) 2016 Fall Meeting, GES DISC unveiled a new service called Datalist: a collection of predefined or user-defined data variables from one or more archived data sets. Our science support team has been curating Datalists and providing added value to the user community.Originally known as Afternoon Constellation, A-Train includes six currently on polar-orbiting Earth observation satellites: OCO-2, GCOM-W1, Aqua, CALIPSO, CloudSat, and Aura, which travel a few minutes apart from each other. This constellation arrangement has enabled coordinated science observations further forming comprehensive pictures of Earth weather and climate that are readily for use in crucial studies such as climate change.GES DISC Datalists are based on the software architecture of the new GES DISC website (also unveiled at the AGU 2016 Fall Meeting). The GES DISC science support team has created a Datalist to support the A-Train Data Depot (ATDD). Using pre-defined Datalist should hopefully save users significant effort in their data searches.

A-Train data ordering

GES DISC Datalist Enables Easy Data Selection For Natural Phenomena Studies

In order to investigate and assess natural hazards such as tropical storms, winter storms, volcanic eruptions, floods, and drought in a timely manner, the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) has been developing an efficient data search and access service. Called "Datalist," this service enables users to acquire their data of interest "all at once," with minimum effort. A Datalist is a virtual collection of predefined or user-defined data variables from one or more archived data sets. Datalists are more than just data. Datalists effectively provide users with a sophisticated integrated data and services package, including metadata, citation, documentation, visualization, and data-specific services (e.g., subset and OPeNDAP), all available from one-stop shopping. The predefined Datalists, created by the experienced GES DISC science support team, should save a significant amount of time that users would otherwise have to spend. The Datalist service is an extension of the new GES DISC website, which is completely data-driven. A Datalist, also known as "data bundle," is treated just as any other data set. Being a virtual collection, a Datalist requires no extra storage space.

natural hazards

New GES DISC Services Shortening the Path in Science Data Discovery

The Current GES DISC available services only allow user to select variables from a single dataset at a time and too many variables from a dataset are displayed, choice is hard. At American Geophysical Union (AGU) 2016 Fall Meeting, Goddard Earth Sciences Data Information Services Center (GES DISC) unveiled a new service: Datalist. A Datalist is a collection of predefined or user-defined data variables from one or more archived datasets. Our science support team curated predefined datalist and provided value to the user community. Imagine some novice user wants to study hurricane and typed in hurricane in the search box. The first item in the search result is GES DISC provided Hurricane Datalist. It contains scientists recommended variables from multiple datasets like TRMM, GPM, MERRA, etc. Datalist uses the same architecture as that of our new website, which also provides one-stop shopping for data, metadata, citation, documentation, visualization and other available services.We implemented Datalist with new GES DISC web architecture, one single web page that unified all user interfaces. From that webpage, users can find data by either type in keyword, or browse by category. It also provides user with a sophisticated integrated data and services package, including metadata, citation, documentation, visualization, and data-specific services, all available from one-stop shopping.

Datalist

Datalist: A Value Added Service to Enable Easy Data Selection

Imagine a user wanting to study hurricane events. This could involve searching and downloading multiple data variables from multiple data sets. The currently available services from the Goddard Earth Sciences Data and Information Services Center (GES DISC) only allow the user to select one data set at a time. The GES DISC started a Data List initiative, in order to enable users to easily select multiple data variables. A Data List is a collection of predefined or user-defined data variables from one or more archived data sets. Target users of Data Lists include science teams, individual science researchers, application users, and educational users. Data Lists are more than just data. Data Lists effectively provide users with a sophisticated integrated data and services package, including metadata, citation, documentation, visualization, and data-specific services, all available from one-stop shopping. Data Lists are created based on the software architecture of the GES DISC Unified User Interface (UUI). The Data List service is completely data-driven, and a Data List is treated just as any other data set. The predefined Data Lists, created by the experienced GES DISC science support team, should save a significant amount of time that users would otherwise have to spend.

Datalist

Integrated Data Modeling and Simulation on the Joint Polar Satellite System Program

The Joint Polar Satellite System is a modern, large-scale, complex, multi-mission aerospace program, and presents a variety of design, testing and operational challenges due to: (1) System Scope: multi-mission coordination, role, responsibility and accountability challenges stemming from porous/ill-defined system and organizational boundaries (including foreign policy interactions) (2) Degree of Concurrency: design, implementation, integration, verification and operation occurring simultaneously, at multiple scales in the system hierarchy (3) Multi-Decadal Lifecycle: technical obsolesce, reliability and sustainment concerns, including those related to organizational and industrial base. Additionally, these systems tend to become embedded in the broader societal infrastructure, resulting in new system stakeholders with perhaps different preferences (4) Barriers to Effective Communications: process and cultural issues that emerge due to geographic dispersion and as one spans boundaries including gov./contractor, NASA/Other USG, and international relationships.

Roberts, Christopher J.