Engineering PapersSearch

SEARCH · Engineering Papers

Results for “analysis ready data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Analysis Ready Data in Analytics Optimized Data Stores for Analysis of Big Earth Data in the Cloud

Cloud computing offers the possibility of making the analysis of Big Data approachable for a wider community due to affordable access to computing power, an ecosystem of usable tools for parallel processing, and migration of many large datasets to archives in the cloud, allowing data-proximal computing. Generally, data analysis acceleration in the cloud comes from running multiple nodes in a split-combine-apply strategy. Data systems such as the Earth Observing System Data and Information System are in a position to "pre-split" the data by storing them in a data store that is optimized for data parallel computing, i.e., an Analytics-Optimized Data Store (AODS). A variety of approaches to AODS are possible, from highly scalable databases to scalable filesystems to data formats optimized for cloud access (e.g., zarr and cloud-optimized datasets), with the optimal choice dependent on both the types of analysis and the geospatial structure of the data. A key question is how much preprocessing of the data to do, both before splitting and as the first part of the apply step. Again, the geospatial structure of the data and the analysis type influence the decision, with the added complexity of the user type. Trans-disciplinary users who are not well-versed in the nuances of quality-filtering and georeferencing of remote sensing orbit/swath/scene data tend to ask for more highly processed data, relying on the data provider to make sensible decisions on preprocessing parameters. (This accounts for the popularity of "Level 3" gridded data, despite the lower spatial resolution it provides.) In this case, data can be preprocessed before the split, resulting in higher performance in the rest of the "apply" step, which can be transformative for use cases such as interactive data exploration at scale. Discipline researchers who are experienced with remote sensing data often prefer more flexibility in customizing the preprocessing data into Analysis Ready Data, resulting in more need for on-the-fly preprocessing.

Lynnes, Christopher

Data Quality Challenges for Analysis Ready Data (ARD)

Data quality plays a critical role in research and applications. The Earth Science Information Partners (ESIP) Information Quality Cluster (IQC) defines four aspects of information quality: Science, Product, Stewardship, and Services. The ESIP IQC has become internationally recognized as an authoritative and responsive resource of information and guidance to data producers and distributors on how to implement data quality standards and best practices for their science data systems, datasets, and data/metadata dissemination services. In recent years, cloud computing environments have provided scale-up capabilities such as data archives and services, enabling interdisciplinary science and applications. More value-added products are expected from data service providers, including Analysis Ready Data (ARD). ARD refers to data that has been preprocessed into a form that allows immediate analysis by the end user, processed to a minimum set of requirements and provides interoperability over time and across multiple datasets. Once a dataset has been developed from its original form to produce ARD, what quality characteristics should the derived dataset or ARD possess? Also, is it safe to assume that the quality of the ARD is consistent with the quality of the source data, or are there special attributes to an ARD that would warrant a secondary, independent quality assessment? What provenance (also called “data lineage”) information needs to be included in ARD? It is important to answer these questions, especially given the ease of use of ARD, and the consequent temptation by users to trust ARD without understanding the limitations or possible variations in quality compared to the source data. In this presentation, we will discuss data quality challenges for ARD products and services and introduce IQC for participation.

data quality

Role of CEOS Working Group on Calibration and Validation in Analysis Ready Data Products

The Committee on Earth Observation Satellites (CEOS) is leading the CEOS Analysis Ready Data for Land (CARD4L) initiative. A goal of analysis ready data products is to limit the effort needed by users to pre-process the data allowing them to concentrate on the end products. CARD4L provides a set of specifications that data providers need to meet to be considered to satisfy CARD4L. One of the working groups within CEOS, the Working Group on Calibration and Validation (WGCV) is providing a peer review process to evaluate the documentation of the data providers validation and data product accuracy assessment. The approach makes use of the expertise within WGCV to collaborate with both the CEOS Land Surface Imaging Virtual Constellation and the data providers to work towards acceptance of the...

Thome, K.

The CEOS Data Cube Portal: A User-Friendly, Open Source Software Solution for the Distribution, Exploration, Analysis, and Visualization of Analysis Ready Data

There is an urgent need to increase the capacity of developing countries to take part in the study and monitoring of their environments through remote sensing and space-based Earth observation technologies. The Open Data Cube (ODC) provides a mechanism for efficient storage and a powerful framework for processing and analyzing satellite data. While this is ideal for scientific research, the expansive feature space can also be daunting for end-users and decision-makers who simply require a solution which provides easy exploration, analysis, and visualization of Analysis Ready Data (ARD). Utilizing innovative web-design and a modular architecture, the Committee on Earth Observation Satellites (CEOS) has created a web-based user interface (UI) which harnesses the power of the ODC yet provides a simple and familiar user experience: the CEOS Data Cube (CDC). This paper presents an overview of the CDC architecture and the salient features of the UI. In order to provide adaptability, flexibility, scalability, and robustness, we leverage widely-adopted and well-supported technologies such as the Django web framework and the AWS Cloud platform. The fully-customizable source code of the UI is available at our public repository. Interested parties can download the source and build their own UIs. The UI empowers users by providing features that assist with streamlining data preparation, data processing, data visualization, and sub-setting ARD products in order to achieve a wide variety of Earth imaging objectives through an easy to use web interface.

User Interface

NASA Earth Systems Digital Twins (ESDT)

"Similarly to artificial intelligence, which is now revolutionizing many aspects of our daily lives, Earth system digital twin technologies have the potential to revolutionize the way Earth Science research will be conducted in the future, and how results and knowledge from this research will provide information to support decision making and yield impactful societal benefits. An Earth System Digital Twin or ESDT is a dynamic and interactive information system that first provides a digital replica of the past and current states of the Earth or Earth system as accurately and timely as possible; second, allows for computing forecasts of future states under nominal assumptions and based on the current replica; and third, offers the capability to investigate many hypothetical scenarios under varying impact assumptions. In other words, an ESDT provides the integrated What-Now, What-Next, and What-If pictures of the Earth or Earth system, by continuously ingesting newly observed data and by leveraging multiple interconnected models, machine learning as well advanced computing and visualization capabilities. Digital twins have been developed in engineering since 2002, but the interest in digital twins for the Earth domain is more recent and stems from the convergence of several developments: - The huge amount of diverse data that has now been collected continuously for more than 50 years, and that is becoming more and more difficult to access, understand, and utilize. - At the same time, because of climate change and its impacts the information produced by all of this data is becoming of interest to many new non-traditional users for analyzing and predicting various phenomena. - Because of advances in computational and visualization capabilities and the parallel unprecedented development of machine learning (ML), extracting relevant information from these large amounts of data and running complex models faster has become possible. As a result, it is becoming necessary and possible to build intuitive and interactive frameworks that will enable users with various skill levels and/or organizational hierarchy levels to easily access large amounts of targeted information along with the relevant tools and models (Earth system and human activity models), to support them in analyzing and visualizing this information, to help them understand interactions among models, to visualize the potential outcomes of various impacts, and to support decision or policy making. The full power of digital twins is that, through an integrated representation and standardized tools and software technologies, the same digital replica can address the needs of multiple users at various resolutions (spatial and temporal) and for various applications (science, economic, policy, etc.) – “from farmer to scientist”. With all these interests at stake, the challenges of building optimal digital twins are many and complex. The first challenge is to determine if a Digital Twin should be global or local, and multi-domain or thematic. For example, some domains such as Climate or Weather will require a global Digital Twin or Digital Twin capabilities while science areas such as Biodiversity might be more local. We can also envision that multiple thematic ESDTs, e.g., Air Quality, Wildfires, Hydrology could be federated or provide input to other ESDTs, either on a regional level or to a more global ESDT. Overall, we can imagine a future “web” of Digital Twins co-existing in a hierarchy or in a network, and capable of being connected or federated depending on the needs. This last point brings up the very important challenge of interoperability, including standards and protocols that will need to be built into these systems from the beginning. Each individual digital twin would have full flexibility in internal construction but would need standards-based interfaces (input and output) or hooks to make it compatible with others. Another challenge when building digital twins will be to decide how to organize each digital replica. Based on the applications targeted by the DT under implementation, various amounts and types of raw data, Analysis Ready Data (ARD) and information will need to be incorporated. Depending on the required latencies and needs of the users, various solutions can be considered, including Data Cubes, Data Lakes, pointers, or computing information on demand. We envision that each ESDT will choose a solution adapted to its specific objectives. Another important challenge is the type(s) of visualization that will be used, as well as the level of interactivity and refresh rate that will be required. Again, this will depend on the objectives of the ESDT, but also on the various users’ needs. In most cases, several types of visualizations and human interfaces will need to be offered depending on the projected users of that system. In parallel to the challenges highlighted above, there are also many tools and technologies that will need to be developed or improved for all types of digital twins. Among those are improved machine learning technologies, for example providing explainability, but also ML techniques for causality and providing a better integration of physics models. Additionally, reliable uncertainty quantification methods will be needed for all ESDT components, from validating data fusion and assimilation to assessing the accuracy of ML models and weighing the values of decisions supported by those systems. This presentation introduces the ESDT concept, presents several ESDT use cases, and a proposed ESDT architecture framework, as well as various technologies being developed by the Advanced Information Systems Technology (AIST) Program."

Earth Science Remote Sensing; Information Systems

Making NASA GES DISC Level 2 Data GIS Analysis Ready

There are many valuable data hosted by NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) for GIS applications in the areas of extreme weather events, climatic anomaly, and public health. However, using NASA Earth Science Data poses some challenges for GIS users. Many of these users are not experts in Earth Observation and have little knowledge about NASA's Earth science data. In the GIS community, GeoTiff is the most widely used raster format, whereas NASA's data is primarily in complex multidimensional netCDF and HDF formats. This complexity makes it difficult for GIS users, especially those who are unfamiliar with these formats. Although GIS software like ArcGIS has made progress in processing multidimensional netCDF data, certain issues still remain, particularly with level 2 data. In this study, we use TROPSpheric Monitoring instrument (TROPOMI) level 2 data as an example to demonstrate how to make such data GIS analysis ready. The process involves: 1. Creating a feature layer from the TROPOMI level 2 data. 2. Converting the feature layer to a gridded raster dataset. 3. Mosaicking gridded raster datasets into a raster dataset covering the entire desired extent. 4. Generating a symbology with a GIBS-specific style that aligns with the visual standards and requirements of GIBS. 5. Publishing image services. By performing these preprocessing and transformation steps, NASA level 2 data can be made compatible and ready for use within GIS software for various spatial analysis and visualization tasks.

Geographic Information System

Visualization and Analysis of Near Real-Time Global Cloud Composites (GCC): Integration in a Geospatial Web Mapping Application

The NASA Langley Satellite ClOud and Radiation Property retrieval System (SatCORPS) team provides tools for retrieving cloud information from operational and research meteorological imager data. LEO and GEO satellite imagers are used for fusing hourly mosaics which create the Global Cloud Composite (GCC). The GCC provides global cloud products with a low latency which can help satisfy the growing needs of the research, modeling and business communities. The Esri ArcGIS system is used to transform the GCC data into Analysis Ready Data (ARD) which can be utilized for a variety of visualization and analysis activities in support of weather diagnoses/forecasting, Earth Sciences remote sensing applications, and disaster management. Additionally, the data is enabled as ArcGIS Image Services and Open Geospatial Consortium (OGC) Web Mapping/Coverage Services to be consumed via API within a web mapping application. We will present a preview of the new interactive SatCORPS GCC web mapping application and its capabilities in allowing users to access and use GIS data within it.

Web Mapping Application

Analysis Ready Satellite Data

Analyais-Ready Data (ARD) specifications have gained rcent prominence in the field of land-related Earth Observations. The ARD label enables users to recognize data that need a minimum of preprocessing before analysis. However, the regular geolocation requirements make Level 2 data in other disciplines problematic. While Level 3 gridded data can satisfy the geolocation requirement, they often sacrifice spatial resolution and other information, such as extreme values. This talk outlines this dilemma with some potential approaches to it.

Analysis-Ready Data

Integrated Topographic Corrections Improve Forest Mapping Using Landsat Imagery

In mountainous environments, topography strongly affects the reflectance due to illumination effects and cast shadows, which introduce errors in land cover classifications. However, topographic correction is not routinely implemented in standard data pre-processing chains (e.g., Landsat Analysis Ready Data), and there is a lack of consensus whether topographic correction is necessary, and if so, how to conduct it. Furthermore, methods that correct simultaneously for atmospheric and topographic effects are becoming available, but they have not been compared directly. Our objects were to investigate (1) the effectiveness of two topographic correction approaches that integrate atmospheric and topographic correction, (2) improvements in classification accuracy when analyzing topographically corrected single-date imagery (14 July 2016 and 2 October 2016), versus a full Landsat time series from 2014 to 2016, and 3) improvements in classification accuracy when including additional terrain information (i.e., topographic slope, elevation, and aspect). We developed a physical based model and compared it with an enhanced C-correction, both of which integrate atmospheric and topographic correction. We compared classification accuracies with and without topographic correction using combinations of single-date imagery, image composites and spectral-temporal metrics generated from the full Landsat time series, and additional terrain information in the Caucasus Mountains. We found that both the enhanced C-correction and the physical model performed very well and largely eliminated the correlation (Pearson’s correlation coefficient r ranges from 0.06 to 0.24) between surface reflectance and illumination condition, but the physical model performed best (r ranges from 0.05 to 0.11). Both image composites, and spectral-temporal metrics generated from corrected imagery, resulted in significantly (p ≤ 0.05) higher classification accuracies and better forest classifications, especially for the mixed forests. Adding terrain information reduced classification error significantly, but not as much as topographic correction. In summary, topographic correction remains necessary, even when analyzing a full Landsat time series and including a digital elevation model in the classification. We recommend that topographic correction should be applied when analyzing Landsat satellite imagery in mountainous region for forest cover classification.

Atmospheric correction

Lessons Learned and Cost Analysis of Hosting a Full Stack Open Data Cube (ODC) Application on the Amazon Web Services (AWS)

The Open Data Cube (ODC) initiative, with support from the Committee on Earth Observation Satellites (CEOS) System Engineering Office (SEO) has developed a state-of-the-art suite of software tools and products to facilitate the analysis of Earth Observation data. This paper presents a short summary and cost analysis of our experience using Amazon Web Services (AWS) to host one such software product, the CEOS Data Cube (CDC) web-based User Interface (UI). In order to provide adaptability, flexibility, scalability, and robustness, we leverage widely-adopted and well-supported technologies such as the Django web framework and the AWS Cloud platform. The UI has empowered users by providing features that assist with streamlining data preparation, data processing, data visualization, and the sub-setting of Analysis Ready Data (ARD) products in order to achieve a wide variety of Earth imaging objectives.

Rizvi, Syed R.

Fostering Open Science Inclusiveness for Interdisciplinary Users of Earth Observations

The term Open Science is subject to a variety of interpretations because of a key (and useful) ambiguity in the meaning of “Open”. Open in the sense of Transparency enables more trust in science research by making the details of the scientific process visible and accessible to anyone. “Open” in the sense of Inclusiveness enables more scientists from other disciplines to participate in research in a given discipline, thus producing more interdisciplinary research. Data Systems can play a major role in enabling Open (Inclusive) Science by making it easier for users from other disciplines to work with data within a given discipline. This is challenging for Earth Observation datasets, most of which are the product of advanced instrumentation and sophisticated, specialized variable retrieval algorithms and code. Serving the “extra-disciplinary”communities begins with simple things, like accessible, readable data documentation with adequate scaffolding. But just as important is provisioning Analysis-Ready data that does not require expert pre-processing. Disciplines also often have dominant toolsets, such as R in the biomass community or GIS in many applications communities. Ensuring that EO data are easy to use in the tools favored in other communities will enable more interdisciplinary research. Ideally, interdisciplinary research also benefits from scientists with different domain expertise. Platforms and frameworks that facilitate frictionless collaboration with discipline experts, together with capacity building efforts in those external disciplines also improve the inclusiveness aspect of Open Science. In short, Open Science is at root a way of thinking about how users from diverse discipline can best access and use data and services from a particular discipline.

Christopher Lynnes

NASA’s Prototype Spectral Water Inversion Processor and Emulator (SWIPE): Towards Global Coastal and Inland Water Quality and Algal Biodiversity Monitoring

Degradation of Earth’s inland water resources due to anthropogenic perturbations and climate anomalies at both local and global scales continues to place human health at substantial risk. There is now a growing necessity to develop pragmatic approaches that allow timely and effective extrapolation of local processes, to spatially resolved global products, and to promote operational and sustainable resource policy management. This presentation will provide updates on NASA’s prototype open-source aquatic modeling platform, Spectral Water Inversion Processor and Emulator (SWIPE), which is a comprehensive, multi-faceted modeling platform for both forward and inverse modeling of diverse aquatic ecosystems from the benthos to top-of-atmosphere (TOA). SWIPE provides a cohesive application which leverages recent advancements in particle modeling, Big Data analytics, and machine learning to develop a high-fidelity synthetic training ground for sensitivity studies and algorithm development for multispectral or upcoming hyperspectral missions. Some of the prominent features of SWIPE to be discussed include: 1. Advanced hyperspectral modeling of globally diverse algal and non-algal particles using a novel two-layer coated sphere scattering model and radiative transfer modeling, 2. Massive, highly detailed synthetic spectral libraries of Analysis-Ready-Data (ARD) which include spectral libraries of particle microphysics, water biogeophysical and optical properties, as well as surface and TOA reflectances at 1 nm resolution, 3. An ensemble of pre-built analytic, machine learning, and deep learning inversion algorithms for various water quality and biodiversity related retrieval parameters and uncertainty quantification, 4. Sensor-agnostic water quality inversion at wide ranging spatial and spectral resolutions including a codebase for seamless application in the Google Earth Engine and NASA Earth Exchange (NEX) for planetary scale analysis. SWIPE will be a fully open-source platform based in python with comprehensive documentation, tutorials, and options for distributed computing on high performance computing clusters or on single, local machines. Further, we will discuss how we envision SWIPE contributing towards a global analysis of coastal and inland water quality dynamics.

top-of-atmosphere (TOA)

TPSAS-NF1676L-32493-DND

The Committee on Earth Observation Satellites (CEOS) System Engineering Office (SEO) has supported the Open Data Cube (ODC) initiative to provide a data architecture solution that has value to its global users and increases the impact of EO satellite data. ODC is an open-source platform for processing satellite data. We have developed software products and tools around the core ODC that would help users perform machine learning on EO satellite data. The recent United Nations (UN) Sustainable Development Agenda provides a shared blueprint for peace and prosperity for people and for the planet, considering our current situation and helping to create a plan. The core of this agenda is a set of seventeen Sustainable Development Goals (SDGs), which represent an urgent call for action by all countries - both developed and developing - in a global partnership. The CEOS SEO team has recently developed and released a set of innovative Jupyter notebooks addressing UN SDGs 6.6.1 (spatial extents of water-related ecosystems), 11.3.1 (ratio of land consumption rate to population growth rate), and 15.3.1 (proportion of land that is degraded over total land area). These notebooks empower users by providing features that will assist with streamlining analysis ready data retrieval, processing, and visualization. We have recently incorporated several machine learning techniques in these notebooks. In this paper, we present the lessons learned from our experience on classifying land using supervised and unsupervised machine learning techniques using ODC framework for UN SDGs. We identify the current limitations of ODC to seamlessly support machine learning techniques. We propose features that would help machine learning, specifically within the ODC framework. We propose a thematic indexing/loading of data for both unsupervised learning as well as data annotation/labeling pipeline. Currently, ODC supports machine learning by separating data-management from the analysis process. It works as a mechanism to load cubes of data. ODC does not natively support features that are vital in machine learning such as validation splits, fair/balanced sampling, establishing load size constraints, etc. We believe that our proposed features will empower users by providing features that bring machine learning techniques closed to ODC. Enhancements to ODC to better accommodate machine learning techniques can assist in fulfilling UN SDGs such as 6.3.2, 6.4.2, 6.6.1, 11.3.1, 14.1.1, 15.1.1, 15.3.1, and 15.4.2.

Syed R Rizvi