Big Data and Cloud Computing Expert Panel
This talk is a brief introduction to the exploratory cloud prototypes currently underway within ESDIS and the implications of cloud technologies on existing systems.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
This talk is a brief introduction to the exploratory cloud prototypes currently underway within ESDIS and the implications of cloud technologies on existing systems.
No abstract available
No abstract available
We have implemented an updated Hierarchical Triangular Mesh (HTM) as the basis for a unified data model and an indexing scheme for geoscience data to address the variety challenge of Big Earth Data. We observe that, in the absence of variety, the volume challenge of Big Data is relatively easily addressable with parallel processing. The more important challenge in achieving optimal value with a Big Data solution for Earth Science (ES) data analysis, however, is being able to achieve good scalability with variety. With HTM unifying at least the three popular data models, i.e. Grid, Swath, and Point, used by current ES data products, data preparation time for integrative analysis of diverse datasets can be drastically reduced and better variety scaling can be achieved. In addition, since HTM is also an indexing scheme, when it is used to index all ES datasets, data placement alignment (or co-location) on the shared nothing architecture, which most Big Data systems are based on, is guaranteed and better performance is ensured. Moreover, our updated HTM encoding turns most geospatial set operations into integer interval operations, gaining further performance advantages.
Big Earth Data Initiative (BEDI) The Big Earth Data Initiative (BEDI) invests in standardizing and optimizing the collection, management and delivery of U.S. Government's civil Earth observation data to improve discovery, access use, and understanding of Earth observations by the broader user community. Complete and consistent standard metadata helps address all three goals.
Objectives of the NASA Information And Data System (NAIADS) project are to develop a prototype of a conceptually new middleware framework to modernize and significantly improve efficiency of the Earth Science data fusion, big data processing and analytics. The key components of the NAIADS include: Service Oriented Architecture (SOA) multi-lingual framework, multi-sensor coincident data Predictor, fast into-memory data Staging, multi-sensor data-Event Builder, complete data-Event streaming (a work flow with minimized IO), on-line data processing control and analytics services. The NAIADS project is leveraging CLARA framework, developed in Jefferson Lab, and integrated with the ZeroMQ messaging library. The science services are prototyped and incorporated into the system. Merging the SCIAMACHY Level-1 observations and MODIS/Terra Level-2 (Clouds and Aerosols) data products, and ECMWF re- analysis will be used for NAIADS demonstration and performance tests in compute Cloud and Cluster environments.
Climate science is a big data domain that is experiencing unprecedented growth. In our efforts to address the big data challenges of climate science, we are moving toward a notion of Climate Analytics-as-a-Service (CAaaS). CAaaS combines high-performance computing and data-proximal analytics with scalable data management, cloud computing virtualization, the notion of adaptive analytics, and a domain-harmonized API to improve the accessibility and usability of large collections of climate data. MERRA Analytic Services (MERRA/AS) provides an example of CAaaS. MERRA/AS enables MapReduce analytics over NASA's Modern-Era Retrospective Analysis for Research and Applications (MERRA) data collection. The MERRA reanalysis integrates observational data with numerical models to produce a global temporally and spatially consistent synthesis of key climate variables. The effectiveness of MERRA/AS has been demonstrated in several applications. In our experience, CAaaS is providing the agility required to meet our customers' increasing and changing data management and data analysis needs.
The 2015 Space Radiation Standing Review Panel (from here on referred to as the SRP) met for a site visit in Houston, TX on December 8 - 9, 2015. The SRP met with representatives from the Space Radiation Element and members of the Human Research Program (HRP) to review the updated research plan for the Risk of Radiation Carcinogenesis Cancer Risk. The SRP also reviewed the newly revised Evidence Reports for the Risk of Acute Radiation Syndromes Due to Solar Particle Events (SPEs) (Acute Risk), the Risk of Acute (In-flight) and Late Central Nervous System Effects from Radiation Exposure (CNS Risk), and the Risk of Cardiovascular Disease and Other Degenerative Tissue Effects from Radiation (Degen Risk), as well as a status update on these Risks. The SRP would like to commend Dr. Simonsen, Dr. Huff, Dr. Nelson, and Dr. Patel for their detailed presentations. The Space Radiation Element did a great job presenting a very large volume of material. The SRP considers it to be a strong program that is well-organized, well-coordinated and generates valuable data. The SRP commended the tissue sharing protocols, working groups, systems biology analysis, and standardization of models. In several of the discussed areas the SRP suggested improvements of the research plans in the future. These include the following: It is important that the team has expanded efforts examining immunology and inflammation as important components of the space radiation biological response. This is an overarching and important focus that is likely to apply to all aspects of the program including acute, CVD, CNS, cancer and others. Given that the area of immunology/inflammation is highly complex (and especially so as it relates to radiation), it warrants the expansion of investigators expertise in immunology and inflammation to work with the individual research projects and also the NASA Specialized Center of Research (NSCORs). Historical data on radiation injury to be entered into the Watson “big data” study must be used with caution. The general scientific issues of reproducibility, details of experimental methods and data analysis from preclinical and basic research laboratories have been raised broadly over the last few years (not specific to this work) and indicate that caution must be applied in the ways these data are used. This pertains to preclinical data and also to phase 3 clinical trials in radiation oncology and medical oncology. Of course, appropriate use and analysis of these “big-data” sets also offer the potential of pinpointing limitations and extracting remaining useful information. Emphasis should be placed on the latter possibility. A key target is risk reduction from radiation exposure. Progress of the entire space program, now moving towards the Mars mission, requires timely answers to key components of human risk, which are known to be complex. Periodic review of progress should be conducted with additional resources directed into achieving critical milestones. Turning the long red bars to yellow and green (or for some risks such as CNS possibly to grey) must be high priority. That such progress will require new science and not engineering means that it should be viewed in a knowledge-based light. The technology-based aspects of engineering issues are certainly as important, however, science and knowledge-based problems are solved in a different way than engineering. Timelines for engineering are more predictable, while for science, progress can be methodical with occasional major incremental findings that can rapidly change the rate of progress. As opportunities for rapid incremental changes arise, periodic enhancement of investment is strongly recommended to enable such new knowledge to be quickly and efficiently exploited. Collaborations and linkages with National Institute of Allergy and Infectious Diseases (NIAID), the Biomedical Advanced Research and Development Authority (BARDA) and the Department of Defense (DoD) are in place and more are encouraged, where possible, with the radiation injury and medical countermeasure studies. This could include utilizing some of their animal model testing contracts to facilitate obtaining results using common platforms. Such approach will facilitate the comparison of results among laboratories, and will facilitate and accelerate the development of medical countermeasures. It is particularly noteworthy that the NASA Space Radiation Element is reaching out to the Multidisciplinary European Low Dose Initiative (MELODI) platform coordinating low dose radiation risk research, and to other international agencies that are studying low dose radiation effects in an effort to fill the void generated by the cancelation of the Department of Energy (DOE) low dose radiation program. While NASA is working actively with NIAID and BARDA to integrate their relevant findings of radiation mitigator investigations to NASA programs, the committee notes its disappointment that the United States currently lacks a dedicated low dose radiation program with clear mechanistic orientation and aimed at the quantification and mitigation of human radiation risk on Earth. This void gives to the NASA Space Radiation Program Element special societal value, but also makes its overall design more challenging.
At the Goddard Distributed Active Archive Center, we have recently added a 35-Year record of output data from the North American Land Assimilation System (NLDAS) to the Giovanni web-based analysis and visualization tool. Giovanni (Geospatial Interactive Online Visualization ANd aNalysis Infrastructure) offers a variety of data summarization and visualization to users that operate at the data center, obviating the need for users to download and read the data themselves for exploratory data analysis. However, the NLDAS data has proven surprisingly resistant to application of the summarization algorithms. Algorithms that were perfectly happy analyzing 15 years of daily satellite data encountered limitations both at the algorithm and system level for 35 years of hourly data. Failures arose, sometimes unexpectedly, from command line overflows, memory overflows, internal buffer overflows, and time-outs, among others. These serve as an early warning sign for the problems likely to be encountered by the general user community as they try to scale up to Big Data analytics. Indeed, it is likely that more users will seek to perform remote web-based analysis precisely to avoid the issues, or the need to reprogram around them. We will discuss approaches to mitigating the limitations and the implications for data systems serving the user communities that try to scale up their current techniques to analyze Big Data.
The fields of machine learning and big data analytics have made significant advances in recent years, which has created an environment where cross-fertilization of methods and collaborations can achieve previously unattainable outcomes. The Comprehensive Digital Transformation (CDT) Machine Learning and Big Data Analytics team planned a workshop at NASA Langley in August 2016 to unite leading experts the field of machine learning and NASA scientists and engineers. The primary goal for this workshop was to assess the state-of-the-art in this field, introduce these leading experts to the aerospace and science subject matter experts, and develop opportunities for collaboration. The workshop was held over a three day-period with lectures from 15 leading experts followed by significant interactive discussions. This report provides an overview of the 15 invited lectures and a summary of the key discussion topics that arose during both formal and informal discussion sections. Four key workshop themes were identified after the closure of the workshop and are also highlighted in the report. Furthermore, several workshop attendees provided their feedback on how they are already utilizing machine learning algorithms to advance their research, new methods they learned about during the workshop, and collaboration opportunities they identified during the workshop.
NASA's EOSDIS system faces several challenges in the Big Data Era. Although volumes are large (but not unmanageably so), the variety of different data collections is daunting. That variety also brings with it a large and diverse user community. One key evolution EOSDIS is working toward is to enable more science analysis to be performed close to the data.
We investigate the impact of data placement for two Big Data technologies, Spark and SciDB, with a use case from Earth Science where data arrays are multidimensional. Simultaneously, this investigation provides an opportunity to evaluate the performance of the technologies involved. Two datastores, HDFS and Cassandra, are used with Spark for our comparison. It is found that Spark with Cassandra performs better than with HDFS, but SciDB performs better yet than Spark with either datastore. The investigation also underscores the value of having data aligned for the most common analysis scenarios in advance on a shared nothing architecture. Otherwise, repartitioning needs to be carried out on the fly, degrading overall performance.
Each day, the global air transportation industry generates a vast amount of heterogeneous data from air carriers, air traffic control providers, and secondary aviation entities handling baggage, ticketing, catering, fuel delivery, and other services. Generally, these data are stored in isolated data systems, separated from each other by significant political, regulatory, economic, and technological divides. These realities aside, integrating aviation data into a single, queryable, big data store could enable insights leading to major efficiency, safety, and cost advantages. In this paper, we describe an implemented system for combining heterogeneous air traffic management data using semantic integration techniques. The system transforms data from its original disparate source formats into a unified semantic representation within an ontology-based triple store. Our initial prototype stores only a small sliver of air traffic data covering one day of operations at a major airport. The paper also describes our analysis of difficulties ahead as we prepare to scale up data storage to accommodate successively larger quantities of data -- eventually covering all US commercial domestic flights over an extended multi-year timeframe. We review several approaches to mitigating scale-up related query performance concerns.
Each day, the global air transportation industry generates a vast amount of heterogeneous data from air carriers, air traffic control providers, and secondary aviation entities handling baggage, ticketing, catering, fuel delivery, and other services. Generally, these data are stored in isolated data systems, separated from each other by significant political, regulatory, economic, and technological divides. These realities aside, integrating aviation data into a single, queryable, big data store could enable insights leading to major efficiency, safety, and cost advantages. In this paper, we describe an implemented system for combining heterogeneous air traffic management data using semantic integration techniques. The system transforms data from its original disparate source formats into a unified semantic representation within an ontology-based triple store. Our initial prototype stores only a small sliver of air traffic data covering one day of operations at a major airport. The paper also describes our analysis of difficulties ahead as we prepare to scale up data storage to accommodate successively larger quantities of data -- eventually covering all US commercial domestic flights over an extended multi-year timeframe. We review several approaches to mitigating scale-up related query performance concerns.
Exascale computing, big data, and cloud computing are driving the evolution of large-scale information systems toward a model of data-proximal analysis. In response, we are developing a concept of climate analytics as a service (CAaaS) that represents a convergence of data analytics and archive management. With this approach, high-performance compute-storage implemented as an analytic system is part of a dynamic archive comprising both static and computationally realized objects. It is a system whose capabilities are framed as behaviors over a static data collection, but where queries cause results to be created, not found and retrieved. Those results can be the product of a complex analysis, but, importantly, they also can be tailored responses to the simplest of requests. NASA's MERRA Analytic Service and associated Climate Data Services API provide a real-world example of climate analytics delivered as a service in this way. Our experiences reveal several advantages to this approach, not the least of which is orders-of-magnitude time reduction in the data assembly task common to many scientific workflows.
Traditionally, NASA Earth Science data archives have file-based storage using proprietary data file formats, such as HDF and HDF-EOS, which are optimized to support fast and efficient storage of spaceborne and model data as they are generated. The use of file-based storage essentially imposes an indexing strategy based on data dimensions. In most cases, NASA Earth Science data uses time as the primary index, leading to poor performance in accessing data in spatial dimensions. For example, producing a time series for a single spatial grid cell involves accessing a large number of data files. With exponential growth in data volume due to the ever-increasing spatial and temporal resolution of the data, using file-based archives poses significant performance and cost barriers to data discovery and access. Storing and disseminating data in proprietary data formats imposes an additional access barrier for users outside the mainstream research community. At the NASA Goddard Earth Sciences Data Information Services Center (GES DISC), we have evaluated applying the schema-on-read principle to data access and distribution. We used Apache Parquet to store geospatial data, and have exposed data through Amazon Web Services (AWS) Athena, AWS Simple Storage Service (S3), and Apache Spark. Using the schema-on-read approach allows customization of indexing spatially or temporally to suit the data access pattern. The storage of data in open formats such as Apache Parquet has widespread support in popular programming languages. A wide range of solutions for handling big data lowers the access barrier for all users. This presentation will discuss formats used for data storage, frameworks with This presentation will discuss formats used for data storage, frameworks with support for schema-on-read used for data access, and common use cases covering data usage patterns seen in a geospatial data archive.