Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Big data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

FAIRness and Usability for Open-access Omics Data Systems

Omics data sharing is crucial to the biological research community, and the last decade or two has seen a huge rise in collaborative analysis systems, databases, and knowledge bases for omics and other systems biology data. We assessed the “FAIRness” of NASA’s GeneLab Data Systems (GLDS) along with four similar kinds of systems in the research omics data domain, using 14 FAIRness metrics. The range of overall FAIRness scores was 6-12 (out of 14), average 10.1, and standard deviation 2.4. The range of Pass ratings for the metrics was 29-79%, Partial Pass 0-21%, and Fail 7-50%. The systems we evaluated performed the best in the areas of data findability and accessibility, and worst in the area of data interoperability. Reusability of metadata, in particular, was frequently not well supported. We relate our experiences implementing semantic integration of omics data from some of the assessed systems for federated querying and retrieval functions, given their shortcomings in data interoperability. Finally, we propose two new principles that Big Data system developers, in particular, should consider for maximizing data accessibility.

Berrios, Daniel C.↗

SNPP and N20 VIIRS Thermal Emissive Bands Calibration Comparison Using the GEO-LEO Double Difference Method

The VIIRS instruments onboard the SNPP and NOAA-20 satellites have identical spatial resolutions and the same spectral bands. Similar prelaunch tests and identical on-orbit calibration algorithms established the foundation for their consistent Earth measurements. Calibration assessment and consistency comparisons are useful to maintain their performance and measurement accuracy. Simultaneous nadir overpasses (SNO) between two satellites are commonly used for a direct calibration comparison between sensors. However, there are no SNO between SNPP and NOAA20. Hence, a reference sensor or Earth measurements are normally used to bridge the comparison. As a reference, we focus on the Advanced Baseline Imager (ABI) onboard the GOES-R series spacecraft and its application to the SNPP and NOAA-20 VIIRS comparison. GOES16 and GOES17are the first two satellites of the GOES-R series and were launched on November 19, 2016, and March 12, 2018, respectively. Their operational positions are on the equator with longitudes of 75.2° West over land and 137.2° West over ocean, respectively. The ABI is the primary imaging instrument of these spacecrafts for the Earth’s weather, oceans, and environment, with observations (every 10 minutes) that provide vast data for GEO-Low Earth orbit (LEO)and LEO-LEO comparisons utilizing it as an intermediate reference sensor. VIIRS and ABI have spectrally matched bands and can have simultaneous measurements over any selected site every day. The simultaneous measurements over the same site also have various scan angles. These features provide advantages for a VIIRS-to-ABI comparison. The spectral response function difference between instruments, sites selected, and view angles will have effects on the instrument measurements. Their impacts on the calibration comparison, including the use of double differences, will be discussed. By collecting VIIRS measurements over a large range of view angles, the view angle effect will also be investigated. The collection of an ample amount of data provides an advantage for statistical analyses and potential big data applications to sensor calibration assessments. This method can also be applied to other sensor calibration comparison and performance assessments, such as GOES16 and GOES17 ABI, and Terra and Aqua MODIS.

Tiejun Chang↗

Detection of Hail Storms in Radar Imagery Using Deep Learning

In 2016, hail was responsible for 3.5 billion and 23 million dollars in damage to property and crops, respectively, making it the second costliest weather phenomenon in the United States. In an effort to improve hail-prediction techniques and reduce the societal impacts associated with hail storms, we propose a deep learning technique that leverages radar imagery for automatic detection of hail storms. The technique is applied to radar imagery from 2011 to 2016 for the contiguous United States and achieved a precision of 0.848. Hail storms are primarily detected through the visual interpretation of radar imagery (Mrozet al., 2017). With radars providing data every two minutes, the detection of hail storms has become a big data task. As a result, scientists have turned to neural networks that employ computer vision to identify hail-bearing storms (Marzbanet al., 2001). In this study, we propose a deep Convolutional Neural Network (ConvNet) to understand the spatial features and patterns of radar echoes for detecting hailstorms.

natural hazard↗

VEDA Visualization Exploration & Data Analysis

Why? - Interdisciplinary science depends on large amount of Earth science data and computational resources - Working with these datasets is non-trivial - Big data science requires advanced distributed computing knowledge What? VEDA is an open platform that brings key Earth science datasets next to open source tools for data processing, analysis, visualization, and exploration in a managed and more accessible computing environment.

Manil Maskey↗

Scalable Adaptive Graphics Environment (SAGE) Software for the Visualization of Large Data Sets on a Video Wall

The use of collaborative scientific visualization systems for the analysis, visualization, and sharing of "big data" available from new high resolution remote sensing satellite sensors or four‐dimensional numerical model simulations is propelling the wider adoption of ultra‐resolution tiled display walls interconnected by high speed networks. These systems require a globally connected and well‐integrated operating environment that provides persistent visualization and collaboration services. This abstract and subsequent presentation describes a new collaborative visualization system installed for NASA's Shortterm Prediction Research and Transition (SPoRT) program at Marshall Space Flight Center and its use for Earth science applications. The system consists of a 3 x 4 array of 1920 x 1080 pixel thin bezel video monitors mounted on a wall in a scientific collaboration lab. The monitors are physically and virtually integrated into a 14' x 7' for video display. The display of scientific data on the video wall is controlled by a single Alienware Aurora PC with a 2nd Generation Intel Core 4.1 GHz processor, 32 GB memory, and an AMD Fire Pro W600 video card with 6 mini display port connections. Six mini display‐to‐dual DVI cables are used to connect the 12 individual video monitors. The open source Scalable Adaptive Graphics Environment (SAGE) windowing and media control framework, running on top of the Ubuntu 12 Linux operating system, allows several users to simultaneously control the display and storage of high resolution still and moving graphics in a variety of formats, on tiled display walls of any size. The Ubuntu operating system supports the open source Scalable Adaptive Graphics Environment (SAGE) software which provides a common environment, or framework, enabling its users to access, display and share a variety of data‐intensive information. This information can be digital‐cinema animations, high‐resolution images, high‐definition video‐teleconferences, presentation slides, documents, spreadsheets or laptop screens. SAGE is cross‐platform, community‐driven, open‐source visualization and collaboration middleware that utilizes shared national and international cyberinfrastructure for the advancement of scientific research and education.

Jedlovec, Gary↗

An Integrated Data Analytics Platform

An Integrated Science Data Analytics Platform is an environment that enables the confluence of resources for scientific investigation. It harmonizes data, tools and computational resources which subsequently enable the research community to focus on the investigation rather than spending time on security, data preparation, management, etc. OceanWorks is a NASA technology integration project to establish a cloud-based Integrated Ocean Science Data Analytics Platform at NASA’s Physical Oceanography Distributed Active Archive Center (PO.DAAC) for big ocean science. It focuses on advancement and maturity by bringing together several NASA open-source, big data projects for parallel analytics, anomaly detection, in-situ to satellite data matchup, quality-screened data subsetting, search relevancy, and data discovery. Our communities are relying on data distributed through data centers such as the PO.DAAC, COAPS, NCAR, and many others to conduct their research. In typical investigations, scientists would engage in: search for data, evaluate the relevance of that data, download it, and then apply algorithms to identify trends. Such workflow cannot scale if the research involves a massive amount of data or multi-variate measurements. NASA’s Surface Water and Ocean Topography (SWOT) mission is expected to produce massive amount of observational data during its 3-year nominal mission. Collections like SWOT challenges all existing Earth Science data archival, distribution and analysis paradigms. In this paper, we will discuss how OceanWorks enhances the analysis of physical ocean data where the computation is done on an elastic cloud platform next to the archive to deliver fast, web-accessible services for working with oceanographic measurements.

Yang, Chaowei↗

Mining Twitter Data to Augment NASA GPM Validation

The Twitter data stream is an important new source of real-time and historical global information for potentially augmenting the validation program of NASA's Global Precipitation Measurement (GPM) mission. There have been other similar uses of Twitter, though mostly related to natural hazards monitoring and management. The validation of satellite precipitation estimates is challenging, because many regions lack data or access to data, especially outside of the U.S. and in remote and developing areas. The time-varying set of "precipitation" tweets can be thought of as an organic network of rain gauges, potentially providing a widespread view of precipitation occurrence. Twitter provides a large source of crowd for crowdsourcing. During a 24-hour period in the middle of the snow storm this past March in the U.S. Northeast, we collected more than 13,000 relevant precipitation tweets with exact geolocation. The overall objective of our project is to determine the extent to which processed tweets can provide additional information that improves the validation of GPM data. Though our current effort focuses on tweets and precipitation, our approach is general and applicable to other social media and other geophysical measurements. Specifically, we have developed an operational infrastructure for processing tweets, in a format suitable for analysis with GPM data; engaged with potential participants, both passive and active, to "enrich" the Twitter stream; and inter-compared "precipitation" tweet data, ground station data, and GPM retrievals. In this presentation, we detail the technical capabilities of our tweet processing infrastructure, including data abstraction, feature extraction, search engine, context-awareness, real-time processing, and high volume (big) data processing; various means for "enriching" the Twitter stream; and results of inter-comparisons. Our project should bring a new kind of visibility to Twitter and engender a new kind of appreciation of the value of Twitter by the science research communities.

validatio↗

Sherlock Data Warehouse

This slide deck provides an overview of the data and resources available in the Sherlock Data Warehouse. Sherlock was developed and is currently maintained by the Aviation Systems Division at NASA Ames Research Center. Sherlock contains a valuable collection of flight, air traffic management, and weather data. But Sherlock is not just a data archive. Sherlock also includes tools and resources to access, download, and visualize data, as well as resources to process the data. This overview summarizes Sherlock data sources, demonstrates data analytics and visualization with MicroStrategy, illustrates disparate data integration using the ATM Knowledge graph, and presents a machine learning use case using the Big Data system.

data warehouse↗

Achieving Fast Operational Intelligence in NASA's Deep Space Network Through Complex Event Processing

NASA’s Deep Space Network (DSN) is a complex, global project, in which the expertise of human operators remain crucial for its successful operation. To find ways to save costs in operations and to improve its services, a number of modernization efforts are underway in the DSN. One such effort is a research and technology development task at the Jet Propulsion Laboratory that is investigating the use of complex event processing (CEP) for intelligent assessment of situations, trend analysis, and advanced automation. The technology leverages the significant business intelligence (BI) and data science advancements made in the enterprise industries over the last several years. The open source big data processing engine Apache SparkTM and the high-throughput, distributed messaging system Apache Kafka form the core of the DSN Complex Event Processing (DCEP) framework. This paper discusses the system engineering perspective of why achieving efficient, lower-cost operations in the DSN is a challenging problem, how the DCEP system handles the use cases that help realize intelligent operations, and how this solution fits into the overall model of the planned DSN Follow-the- Sun Operations (FtSO).

Choi, Joshua S.↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA’s Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related ‘omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata ‘omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data re-use, resulting in 38 additional publications derived from the original 67 publication over the past four years. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA “Open Science Data Repositories (OSDR)” and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Fluorescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to “big data” from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology. Several other talks will cover these topics in this conference.

life sciences↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, there-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA's Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomatic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related 'omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata 'omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data-use, resulting in 40 enabled publications by open data. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA "Open Science Data Repositories (OSDR)" and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Flourescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to "big data" from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology.

omics↗

Introduction to NASA Goddard Workshop on Artificial Intelligence

Artificial Intelligence (AI) is a collection of advanced technologies that allows machines to think and act, both humanly and rationally, through sensing, comprehending, acting and learning. AI's foundations lie at the intersection of several traditional fields Philosophy, Mathematics, Economics, Neuroscience, Psychology and Computer Science. Although the inception of AI started in the 1950's, it has recently made a strong comeback in all aspects of society and all over the world; this is mainly due to the timely combination of increased data volumes, advanced and mature algorithms, and improvements in computing power and storage. Current AI applications include big data analytics, robotics, intelligent sensing, assisted decision making, and speech recognition just to name a few.This workshop will be investigating how AI technologies can be adapted or developed to address the following challenges: Discover events of interest and correlations in large amounts of science data; improve the outcomes of science modeling and data assimilation using improved data processing, integration, and analysis. Design advisors for mission planning and operations, including anomaly detection and spacecraft health monitoring. Develop tools for engineering support, including advanced manufacturing, orbit determination, new component design and system engineering. Customize intelligent user interfaces, including visual analytics and natural language processing.

Le Moigne, Jacqueline↗

Overview of Artificial Intelligence (AI) at NASA Goddard

Artificial Intelligence (AI) is a collection of advanced technologies that allows machines to think and act, both humanly and rationally, through sensing, comprehending, acting and learning. AI's foundations lie at the intersection of several traditional fields Philosophy, Mathematics, Economics, Neuroscience, Psychology and Computer Science. Although the inception of AI started in the 1950's, it has recently made a strong comeback in all aspects of society and all over the world; this is mainly due to the timely combination of increased data volumes, advanced and mature algorithms, and improvements in computing power and storage. Current AI applications include big data analytics, robotics, intelligent sensing, assisted decision making, and speech recognition just to name a few. During the Tour, we will show a few examples of the current AI activities at NASA Goddard.

Le Moigne, Jacqueline↗

NASA GES DISC's Customized Services for Climatology and Meteorology

At the NASA Goddard Earth Sciences (GES) Data and Information Service Center (DISC), we have archived and distributed more than 2,400 Earth science data products, from different missions or projects containing more than 100 M data files/granules with a total volume size nearly 2 PB that broadly serve user needs in science areas such as Atmospheric Composition, Water & Energy Cycles and Climate Variability. To date, GES DISC has developed many pertinent services to facilitate the usage of data products by our research communities, represented by approximately 24,000 registered users. We are facing the big data with increasingly archival volume and data types, moreover, we also encounter increasing users' demands and the demands are more diversified. It is still a challenge for us to better understand exactly what our users' needs are, even after developing more than 70 services, including well-known online tools such as Giovanni and MERRA subsetter. In this presentation, we will try to address how we can accommodate the users' needs from two applicational user communities, Air Quality and Wind Energy, from data or service discovery to guide them properly utilize the data and services to fit their needs.

customizable services for climate and meteorology↗

NASA GES DISC Giovanni: Current and Future

Giovanni (Geospatial Interactive Online Visualization and Analysis Infrastructure), developed by the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC), has established a reputation among NASA users for easy access, analysis, and visualization of NASA Earth science data. Currently, Giovanni supports over 1900 variables in eight disciplinary areas. Like any other enterprise application, Giovanni faces big data challenges, such as servicing increasingly large data volumes and more complex data types, while at the same time addressing the demands of a more diverse user community, e.g., placing requests for long-term time series from multiple spatially and temporally dense data records. I will present how Giovanni has been evolving from an on-premises, monolithic software application towards a cloud-enabled implementation to address these challenges.

Analytics↗

Smarter Earth Science Data System

The explosive growth in Earth observational data in the recent decade demands a better method of interoperability across heterogeneous systems. The Earth science data system community has mastered the art in storing large volume of observational data, but it is still unclear how this traditional method scale over time as we are entering the age of Big Data. Indexed search solutions such as Apache Solr (Smiley and Pugh, 2011) provides fast, scalable search via keyword or phases without any reasoning or inference. The modern search solutions such as Googles Knowledge Graph (Singhal, 2012) and Microsoft Bing, all utilize semantic reasoning to improve its accuracy in searches. The Earth science user community is demanding for an intelligent solution to help them finding the right data for their researches. The Ontological System for Context Artifacts and Resources (OSCAR) (Huang et al., 2012), was created in response to the DARPA Adaptive Vehicle Make (AVM) programs need for an intelligent context models management system to empower its terrain simulation subsystem. The core component of OSCAR is the Environmental Context Ontology (ECO) is built using the Semantic Web for Earth and Environmental Terminology (SWEET) (Raskin and Pan, 2005). This paper presents the current data archival methodology within a NASA Earth science data centers and discuss using semantic web to improve the way we capture and serve data to our users.

data center↗

Integrated Analysis of Multiple User Metrics - A “Sequel”; and Introducing the Google Analytic

For decades, the Goddard Earth Sciences Data and Information Services Center (GES DISC) has archived and distributed enormous volumes of NASA Earth science data (accompanied with many developed tools and services) to various research/applications communities and the general public. Being “immersed” in the Big Data era, we have inevitably faced the challenges of our continually increasing archived data in both volume and variety, as well as enhanced user needs and demands. In recent years, we have actively analyzed different types of user metrics, such as operational distribution metrics (recording numbers of distinct users and downloaded data files, size of distributed data volume): user publication metrics (mining info from our Giovanni users’ publications): and Bugzilla metrics (collecting info from user questions or feedback from user assistance tickets). Such metrics have helped us achieve a better understanding of user needs, demands, characteristics, and behaviors, which has then helped us improve our user services. Now we will present a “Sequel” of integrated analysis of multiple metrics at the GES DISC by introducing and adding one new kind of metrics acquired via utilizing our recently implemented Google Analytic 360 suite. Several “newer” reports, e.g., “What web site features and links are the most popular (and least)?” and “What are the top 25 dataset Keyword searches?” retrieved from this new metrics set will be presented, along with the aforementioned “traditional” metrics results.

Shie, Chung-Lin↗

Cloud Giovanni: Reining in Costs and Improving Performance with Analytical Data Stores Using Scalable Serverless Architecture

Giovanni is the Geospatial Interactive Online Visualization ANd aNalysis Infrastructure developed at NASA GES DISC which provides a simple and intuitive way to visualize, analyze, and access vast amounts of Earth science data. It receives large number of user requests each day for a variety of analysis and visualization services, which leads to the big data challenge of serving gradually increasing large data volumes with diverse statistical algorithms. We hereby propose a multi-dimensional accumulation method which provides fast and cost-efficient cloud analysis for diverse services including both area averaging and time averaging. This method involves the weighted volume integration over multiple variable dimensions (time and space), and is implemented in AWS using Athena providing serverless and highly scalable data analysis. Compared to the standard method, this approach dramatically reduced the computational time by order of magnitude with a minimal AWS cost incurred. For example, for a benchmark of 10-year area averaging over the 1x1 degree daily variable, the computational time was reduced from minutes to seconds, and the Athena cost is only $5 for 100,000 requests.

Zhang, Hailiang↗