Engineering PapersSearch

NASA NTRS · 20180004248

Sherlock Data Warehouse

Abstract

Overview of NASA Ames Aviation Systems Division's Sherlock data warehouse.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Arneson, Heather M.. 2018-06-01. Sherlock Data Warehouse. https://ntrs.nasa.gov/citations/20180004248

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Evaluating the Impact of Power Outages on Occupancy Patterns During the 2021 Texas Power Crisis

Large-scale power outages, such as those caused by extreme weather events, have a big impact on human behavior. A short power outage is merely a nuisance for most, and may not change people's locations. An outage that lasts for a few hours can result in spoiled food and medical supplies, and people will have to restock spoiled items. Long outages result in temperatures outside tolerable levels in homes, and may prompt people to acquire supplies, such as generators and gas, or change location. The long outages during Winter Storm Uri in Texas resulted in millions of dollars in property damage due to freezing pipes. This level of damage is expected to result in a sharp increase in supply runs and contractor activity. In this paper, we present a tool to explore differences in visiting patterns before, during, and after power outages. It allows to compare different points of interest like medical facilities, grocery stores, hardware stores, and other types of businesses.

big data

A Spatiotemporal Indexing Approach for Efficient Processing of Big Array-Based Climate Data with MapReduce

Climate observations and model simulations are producing vast amounts of array-based spatiotemporal data. Efficient processing of these data is essential for assessing global challenges such as climate change, natural disasters, and diseases. This is challenging not only because of the large data volume, but also because of the intrinsic high-dimensional nature of geoscience data. To tackle this challenge, we propose a spatiotemporal indexing approach to efficiently manage and process big climate data with MapReduce in a highly scalable environment. Using this approach, big climate data are directly stored in a Hadoop Distributed File System in its original, native file format. A spatiotemporal index is built to bridge the logical array-based data model and the physical data layout, which enables fast data retrieval when performing spatiotemporal queries. Based on the index, a data-partitioning algorithm is applied to enable MapReduce to achieve high data locality, as well as balancing the workload. The proposed indexing approach is evaluated using the National Aeronautics and Space Administration (NASA) Modern-Era Retrospective Analysis for Research and Applications (MERRA) climate reanalysis dataset. The experimental results show that the index can significantly accelerate querying and processing (10 speedup compared to the baseline test using the same computing cluster), while keeping the index-to-data ratio small (0.0328). The applicability of the indexing approach is demonstrated by a climate anomaly detection deployed on a NASA Hadoop cluster. This approach is also able to support efficient processing of general array-based spatiotemporal data in various geoscience domains without special configuration on a Hadoop cluster.

big data

Climate Analytics as a Service

Exascale computing, big data, and cloud computing are driving the evolution of large-scale information systems toward a model of data-proximal analysis. In response, we are developing a concept of climate analytics as a service (CAaaS) that represents a convergence of data analytics and archive management. With this approach, high-performance compute-storage implemented as an analytic system is part of a dynamic archive comprising both static and computationally realized objects. It is a system whose capabilities are framed as behaviors over a static data collection, but where queries cause results to be created, not found and retrieved. Those results can be the product of a complex analysis, but, importantly, they also can be tailored responses to the simplest of requests. NASA's MERRA Analytic Service and associated Climate Data Services API provide a real-world example of climate analytics delivered as a service in this way. Our experiences reveal several advantages to this approach, not the least of which is orders-of-magnitude time reduction in the data assembly task common to many scientific workflows.

big data