Engineering PapersSearch

SEARCH · Engineering Papers

Results for “heterogeneous data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Interpretable Categorization of Heterogeneous Time Series Data

We analyze data from simulated aircraft encounters to validate and inform the development of a prototype aircraft collision avoidance system. The high-dimensional and heterogeneous time series dataset is analyzed to discover properties of near mid-air collisions (NMACs) and categorize the NMAC encounters. Domain experts use these properties to better organize and understand NMAC occurrences. Existing solutions either are not capable of handling high-dimensional and heterogeneous time series datasets or do not provide explanations that are interpretable by a domain expert. The latter is critical to the acceptance and deployment of safety-critical systems. To address this gap, we propose grammar-based decision trees along with a learning algorithm. Our approach extends decision trees with a grammar framework for classifying heterogeneous time series data. A context-free grammar is used to derive decision expressions that are interpretable, application-specific, and support heterogeneous data types. In addition to classification, we show how grammar-based decision trees can also be used for categorization, which is a combination of clustering and generating interpretable explanations for each cluster. We apply grammar-based decision trees to a simulated aircraft encounter dataset and evaluate the performance of four variants of our learning algorithm. The best algorithm is used to analyze and categorize near mid-air collisions in the aircraft encounter dataset. We describe each discovered category in detail and discuss its relevance to aircraft collision avoidance.

Drones

Agent Architecture for Aviation Data Integration System

This paper describes the proposed agent-based architecture of the Aviation Data Integration System (ADIS). ADIS is a software system that provides integrated heterogeneous data to support aviation problem-solving activities. Examples of aviation problem-solving activities include engineering troubleshooting, incident and accident investigation, routine flight operations monitoring, safety assessment, maintenance procedure debugging, and training assessment. A wide variety of information is typically referenced when engaging in these activities. Some of this information includes flight recorder data, Automatic Terminal Information Service (ATIS) reports, Jeppesen charts, weather data, air traffic control information, safety reports, and runway visual range data. Such wide-ranging information cannot be found in any single unified information source. Therefore, this information must be actively collected, assembled, and presented in a manner that supports the users problem-solving activities. This information integration task is non-trivial and presents a variety of technical challenges. ADIS has been developed to do this task and it permits integration of weather, RVR, radar data, and Jeppesen charts with flight data. ADIS has been implemented and used by several airlines FOQA teams. The initial feedback from airlines is that such a system is very useful in FOQA analysis. Based on the feedback from the initial deployment, we are developing a new version of the system that would make further progress in achieving following goals of our project.

Kulkarni, Deepak

Observations on Cost Modeling and Performance Measurement of Long Term Archives

This paper describes a prototype suite of Excel-based tools that could be used for estimating lifecycle costs for newly planned or modified long-term archival facilities. These tools may also prove valuable for monitoring the long-term performance of such facilities once operational. Cost estimation is by analogy, using statistical curve-fitting techniques across a database of comparable data activities. The database currently includes 29 operational data centers ranging from small (2 FTEs) to large (66 FTEs), and is readily expandable to include additional activities specifically involving data preservation and added value. Each comparable data center is described in terms of its staffing, throughput workload, archival and distribution requirements, levels of user service, overall complexity, degree of automation, and other data, comprising 94 distinct descriptors in all. The descriptors were developed by normalizing heterogeneous data from the various centers and mapping them into an Excel framework consistent with the OAIS reference model. A user-friendly tool is provided for generating input to and updating the comparables database. This tool can also be used to benchmark the performance (in terms of cost versus throughput) of an operational data center, and to update the staffing, cost and workload data on a periodic basis. The comparables database could thus provide a history of staffing and throughput over time, as a means of performance monitoring and providing feedback for continuous improvement. Ancillary tools are also provided for performing "what-if' cost exercises for planning purposes, and for graphical display of data and results. We provide a high-level description of the tools; present our experiences and observations on gathering the information and maintaining the database; and discuss how this tool set might be applied to long term archives.

Fontaine, Kathy

Telescience, an operational approach to science investigation

The NASA Science and Applications Information System, which is based on telescience and must provide remote interaction between information system services in space and on the ground, is discussed. An infrastructure of networked facilities and institutionally provided support services is being developed. The technologies involved with providing telescience capability are examined, including automated data management services, new data acquisition systems, user support environment for system access, and the capability to access heterogeneous data bases and computational facilities from remote locations.

Weiss, James R.

Development of an intelligent interface for adding spatial objects to a knowledge-based geographic information system

Earth Scientists lack adequate tools for quantifying complex relationships between existing data layers and studying and modeling the dynamic interactions of these data layers. There is a need for an earth systems tool to manipulate multi-layered, heterogeneous data sets that are spatially indexed, such as sensor imagery and maps, easily and intelligently in a single system. The system can access and manipulate data from multiple sensor sources, maps, and from a learned object hierarchy using an advanced knowledge-based geographical information system. A prototype Knowledge-Based Geographic Information System (KBGIS) was recently constructed. Many of the system internals are well developed, but the system lacks an adequate user interface. A methodology is described for developing an intelligent user interface and extending KBGIS to interconnect with existing NASA systems, such as imagery from the Land Analysis System (LAS), atmospheric data in Common Data Format (CDF), and visualization of complex data with the National Space Science Data Center Graphics System. This would allow NASA to quickly explore the utility of such a system, given the ability to transfer data in and out of KBGIS easily. The use and maintenance of the object hierarchies as polymorphic data types brings, to data management, a while new set of problems and issues, few of which have been explored above the prototype level.

Campbell, William J.

Understanding Earthquake Fault Systems Using QuakeSim Analysis and Data Assimilation Tools

We are using the QuakeSim environment to model interacting fault systems. One goal of QuakeSim is to prepare for the large volumes of data that spaceborne missions such as DESDynI will produce. QuakeSim has the ability to ingest distributed heterogenous data in the form of InSAR, GPS, seismicity, and fault data into various earthquake modeling applications, automating the analysis when possible. Virtual California simulates interacting faults in California. We can compare output from long time history Virtual California runs with the current state of strain and the strain history in California. In addition to spaceborne data we will begin assimilating data from UAVSAR airborne flights over the San Francisco Bay Area, the Transverse Ranges, and the Salton Trough. Results of the models are important for understanding future earthquake risk and for providing decision support following earthquakes. Improved models require this sensor web of different data sources, and a modeling environment for understanding the combined data.

earthquake simulations

The Best Educational Tool for Interdisciplinary Earth Science Giovanni

Accessing and using NASA Earth science data has commonly presented a challenge to many educators and students, due to issues such as heterogeneous data formats, complex data structures, large volumes of data storage, special programming requirements, and diverse analytical software options that often require a significant investment in time and resources, especially for novices. By facilitating data access and evaluation, as well as promoting open access to create a more level playing field for non-funded scientists, NASA Earth observation data can be more readily used for scientific discovery and societal benefits. To advance this goal, the NASA Goddard Earth Sciences (GES) Data and Information Services Center (DISC) developed the Geospatial Interactive Online Visualization ANd aNalysis Infrastructure (Giovanni). To date, Giovanni has assisted researchers around the world publish over 1300 peer-reviewed papers in a wide range of Earth science disciplines. In this presentation, we will demonstrate how easy it is to use Giovanni for the rapid creation of many different analyses of both weather and climate events.

interdisciplinary

Earth Science Data Analytics: Preparing for Extracting Knowledge from Information

Data analytics is the process of examining large amounts of data of a variety of types to uncover hidden patterns, unknown correlations and other useful information. Data analytics is a broad term that includes data analysis, as well as an understanding of the cognitive processes an analyst uses to understand problems and explore data in meaningful ways. Analytics also include data extraction, transformation, and reduction, utilizing specific tools, techniques, and methods. Turning to data science, definitions of data science sound very similar to those of data analytics (which leads to a lot of the confusion between the two). But the skills needed for both, co-analyzing large amounts of heterogeneous data, understanding and utilizing relevant tools and techniques, and subject matter expertise, although similar, serve different purposes. Data Analytics takes on a practitioners approach to applying expertise and skills to solve issues and gain subject knowledge. Data Science, is more theoretical (research in itself) in nature, providing strategic actionable insights and new innovative methodologies. Earth Science Data Analytics (ESDA) is the process of examining, preparing, reducing, and analyzing large amounts of spatial (multi-dimensional), temporal, or spectral data using a variety of data types to uncover patterns, correlations and other information, to better understand our Earth. The large variety of datasets (temporal spatial differences, data types, formats, etc.) invite the need for data analytics skills that understand the science domain, and data preparation, reduction, and analysis techniques, from a practitioners point of view. The application of these skills to ESDA is the focus of this presentation. The Earth Science Information Partners (ESIP) Federation Earth Science Data Analytics (ESDA) Cluster was created in recognition of the practical need to facilitate the co-analysis of large amounts of data and information for Earth science. Thus, from a to advance science point of view: On the continuum of ever evolving data management systems, we need to understand and develop ways that allow for the variety of data relationships to be examined, and information to be manipulated, such that knowledge can be enhanced, to facilitate science. Recognizing the importance and potential impacts of the unlimited ways to co-analyze heterogeneous datasets, now and especially in the future, one of the objectives of the ESDA cluster is to facilitate the preparation of individuals to understand and apply needed skills to Earth science data analytics. Pinpointing and communicating the needed skills and expertise is new, and not easy. Information technology is just beginning to provide the tools for advancing the analysis of heterogeneous datasets in a big way, thus, providing opportunity to discover unobvious scientific relationships, previously invisible to the science eye. And it is not easy It takes individuals, or teams of individuals, with just the right combination of skills to understand the data and develop the methods to glean knowledge out of data and information. In addition, whereas definitions of data science and big data are (more or less) available (summarized in Reference 5), Earth science data analytics is virtually ignored in the literature, (barring a few excellent sources).

data analytics

Advanced data structures for the interpretation of image and cartographic data in geo-based information systems

A growing need to usse geographic information systems (GIS) to improve the flexibility and overall performance of very large, heterogeneous data bases was examined. The Vaster structure and the Topological Grid structure were compared to test whether such hybrid structures represent an improvement in performance. The use of artificial intelligence in a geographic/earth sciences data base context is being explored. The architecture of the Knowledge Based GIS (KBGIS) has a dual object/spatial data base and a three tier hierarchial search subsystem. Quadtree Spatial Spectra (QTSS) are derived, based on the quadtree data structure, to generate and represent spatial distribution information for large volumes of spatial data.

Peuquet, D. J.

CCSDS Time-Critical Onboard Networking Service

The Consultative Committee for Space Data Systems (CCSDS) is developing recommendations for communication services onboard spacecraft. Today many different communication buses are used on spacecraft requiring software with the same basic functionality to be rewritten for each type of bus. This impacts on the application software resulting in custom software for almost every new mission. The Spacecraft Onboard Interface Services (SOIS) working group aims to provide a consistent interface to various onboard buses and sub-networks, enabling a common interface to the application software. The eventual goal is reusable software that can be easily ported to new missions and run on a range of onboard buses without substantial modification. The system engineer will then be able to select a bus based on its performance, power, etc and be confident that a particular choice of bus will not place excessive demands on software development. This paper describes the SOIS Intra-Networking Service which is designed to enable data transfer and multiplexing of a variety of internetworking protocols with a range of quality of service support, over underlying heterogeneous data links. The Intra-network service interface provides users with a common Quality of Service interface when transporting data across a variety of underlying data links. Supported Quality of Service (QoS) elements include: Priority, Resource Reservation and Retry/Redundancy. These three QoS elements combine and map into four TCONS services for onboard data communications: Best Effort, Assured, Reserved, and Guaranteed. Data to be transported is passed to the Intra-network service with a requested QoS. The requested QoS includes the type of service, priority and where appropriate, a channel identifier. The data is de-multiplexed, prioritized, and the required resources for transport are allocated. The data is then passed to the appropriate data link for transfer across the bus. The SOIS supported data links may inherently provide the quality of service support requested by the intra-network layer. In the case where the data link does not have the required level of support, the missing functionality is added by SOIS. As a result of this architecture, re-usable software applications can be designed and used across missions thereby promoting common mission operations. In addition, the protocol multiplexing function enables the blending of multiple onboard networks. This paper starts by giving an overview of the SOIS architecture in section 11, illustrating where the TCONS services fit into the overall architecture. It then describes the quality of service approach adopted, in section III. The prototyping efforts that have been going on are introduced in section JY. Finally, in section V the current status of the CCSDS recommendations is summarized.

Parkes, Steve

Machine Learning Approaches to Increasing Value of Spaceflight Omics Databases

The number of spaceflight bioscience mission opportunities is too small to allow all relevant biological and environmental parameters to be experimentally identified. Simulated spaceflight experiments in ground-based facilities (GBFs), such as clinostats, are each suitable only for particular investigations -- a rotating-wall vessel may be 'simulated microgravity' for cell differentiation (hours), but not DNA repair (seconds) -- and introduce confounding stimuli, such as motor vibration and fluid shear effects. This uncertainty over which biological mechanisms respond to a given form of simulated space radiation or gravity, as well as its side effects, limits our ability to baseline spaceflight data and validate mission science. Machine learning techniques autonomously identify relevant and interdependent factors in a data set given the set of desired metrics to be evaluated: to automatically identify related studies, compare data from related studies, or determine linkages between types of data in the same study. System-of-systems (SoS) machine learning models have the ability to deal with both sparse and heterogeneous data, such as that provided by the small and diverse number of space biosciences flight missions; however, they require appropriate user-defined metrics for any given data set. Although machine learning in bioinformatics is rapidly expanding, the need to combine spaceflight/GBF mission parameters with omics data is unique. This work characterizes the basic requirements for implementing the SoS approach through the System Map (SM) technique, a composite of a dynamic Bayesian network and Gaussian mixture model, in real-world repositories such as the GeneLab Data System and Life Sciences Data Archive. The three primary steps are metadata management for experimental description using open-source ontologies, defining similarity and consistency metrics, and generating testing and validation data sets. Such approaches to spaceflight and GBF omics data may soon enable unique insight into which measured phenomena correlate to biological mechanisms that are truly affected by spaceflight conditions; which are most likely to be confounded by other variables; and which are insufficiently characterized, significantly increasing existing and future science return from ISS and spaceflight missions.

Gentry, Diana

Trend Detection of Atmospheric Time Series: Incorporating Appropriate Uncertainty Estimates and Handling Extreme Events

This paper is aimed at atmospheric scientists without formal training in statistical theory. Its goal is to, 1) provide a critical review of the rationale for trend analysis of the time series typically encountered in the field of atmospheric chemistry; 2) describe a range of trend-detection methods; and 3) demonstrate effective means of conveying the results to a general audience. Trend detections in atmospheric chemical composition data are often challenged by a variety of sources of uncertainty, which often behave differently to other environmental phenomena such as temperature, precipitation rate, or stream flow, and may require specific methods depending on the science questions to be addressed. Some sources of uncertainty can be explicitly included in the model specification, such as autocorrelation and seasonality, but some inherent uncertainties are difficult to quantify, such as data heterogeneity and measurement uncertainty due to the combined effect of short- and long-term natural variability, instrumental stability, and aggregation of data from sparse sampling frequency. Failure to account for these uncertainties might result in an inappropriate inference of the trends and their estimation errors. On the other hand, the variation in extreme events might be interesting for different scientific questions, for example, the frequency of extremely high surface ozone events and their relevance to human health. In this study we aim to, 1) review trend detection methods for addressing different levels of data complexity in different chemical species; 2) demonstrate that the incorporation of scientifically interpretable covariates can outperform pure numerical curve fitting techniques in terms of uncertainty reduction and improved predictability; 3) illustrate the study of trends based on extreme quantiles that can provide insight beyond standard mean or median based trend estimates; and 4) present an advanced method of quantifying regional trends based on the inter-site correlations of multi-site data. All demonstrations are based on time series of observed trace gases relevant to atmospheric chemistry, but the methods can be applied to other environmental data sets.

Trace gas

Flexible Data Fusion for Air Quality Estimation and Forecasting in Google Earth Engine to support Global Health Management Needs

The assessment and forecasting of air quality around the world at high spatial and temporal resolution can be enhanced by integrating data from multiple sources including models, satellites, regulatory monitors, and low-cost sensors. Such integration is subject to numerous technical challenges, however, including heterogeneous data resolution and formatting, different levels of data availability and reliability, and computational and capacity challenges to developing data fusion tools and platforms. This presentation will provide an overview of a NASA-funded effort to develop a data fusion system within the Google Earth Engine platform which integrates these air quality data sources to produce comprehensive assessments and forecasts of key air pollutants at sub-daily and sub-city scales. The system is being developed in collaboration with city- and regional-level air quality managers, and will provide them with information to the assess and anticipate the health impacts of poor air quality, track local changes in air quality due to ongoing transportation and land use changes, and identify potential gaps in their current air quality monitoring strategies. The presentation will report advances achieved through the project, including bringing local air quality monitoring data into Google Earth Engine, quantifying uncertainties in air quality estimates and forecasts, and tailored communications tools providing integration into end-user processes to meet their needs.

Carl Malings

Flexible Data Fusion for Air Quality Estimation and Forecasting in Google Earth Engine to support Global Health Management Needs

The assessment and forecasting of air quality around the world at high spatial and temporal resolution can be enhanced by integrating data from multiple sources including models, satellites, regulatory monitors, and low-cost sensors. Such integration is subject to numerous technical challenges, however, including heterogeneous data resolution and formatting, different levels of data availability and reliability, and computational and capacity challenges to developing data fusion tools and platforms. This presentation will provide an overview of a NASA-funded effort to develop a data fusion system within the Google Earth Engine platform which integrates these air quality data sources to produce comprehensive assessments and forecasts of key air pollutants at sub-daily and sub-city scales. The system is being developed in collaboration with city- and regional-level air quality managers and will provide them with information to the assess and anticipate the health impacts of poor air quality, track local changes in air quality due to ongoing transportation and land use changes, and identify potential gaps in their current air quality monitoring strategies. The presentation will report advances achieved through the project, including bringing local air quality monitoring data into Google Earth Engine, quantifying uncertainties in air quality estimates and forecasts, and tailored communications tools providing integration into end-user processes to meet their needs.

Nathan R. Pavlovic

Logical optimization for database uniformization

Data base uniformization refers to the building of a common user interface facility to support uniform access to any or all of a collection of distributed heterogeneous data bases. Such a system should enable a user, situated anywhere along a set of distributed data bases, to access all of the information in the data bases without having to learn the various data manipulation languages. Furthermore, such a system should leave intact the component data bases, and in particular, their already existing software. A survey of various aspects of the data bases uniformization problem and a proposed solution are presented.

Grant, J.

A distributed component framework for science data product interoperability

Correlation of science results from multi-disciplinary communities is a difficult task. Traditionally data from science missions is archived in proprietary data systems that are not interoperable. The Object Oriented Data Technology (OODT) task at the Jet Propulsion Laboratory is working on building a distributed product server as part of a distributed component framework to allow heterogeneous data systems to communicate and share scientific results.

XML distributed framework science interprobability

Long-term archiving and data access: modelling and standardization

This paper reports on the multiple difficulties inherent in the long-term archiving of digital data, and in particular on the different possible causes of definitive data loss. It defines the basic principles which must be respected when creating long-term archives. Such principles concern both the archival systems and the data. The archival systems should have two primary qualities: independence of architecture with respect to technological evolution, and generic-ness, i.e., the capability of ensuring identical service for heterogeneous data. These characteristics are implicit in the Reference Model for Archival Services, currently being designed within an ISO-CCSDS framework. A system prototype has been developed at the French Space Agency (CNES) in conformance with these principles, and its main characteristics will be discussed in this paper. Moreover, the data archived should be capable of abstract representation regardless of the technology used, and should, to the extent that it is possible, be organized, structured and described with the help of existing standards. The immediate advantage of standardization is illustrated by several concrete examples. Both the positive facets and the limitations of this approach are analyzed. The advantages of developing an object-oriented data model within this contxt are then examined.

Hoc, Claude

Harnessing the Risk-Related Data Supply Chain: An Information Architecture Approach to Enriching Human System Research and Operations Knowledge

NASA's Human Research Program (HRP) and Space Life Sciences Directorate (SLSD), not unlike many NASA organizations today, struggle with the inherent inefficiencies caused by dependencies on heterogeneous data systems and silos of data and information spread across decentralized discipline domains. The capture of operational and research-based data/information (both in-flight and ground-based) in disparate IT systems impedes the extent to which that data/information can be efficiently and securely shared, analyzed, and enriched into knowledge that directly and more rapidly supports HRP's research-focused human system risk mitigation efforts and SLSD s operationally oriented risk management efforts. As a result, an integrated effort is underway to more fully understand and document how specific sets of risk-related data/information are generated and used and in what IT systems that data/information currently resides. By mapping the risk-related data flow from raw data to useable information and knowledge (think of it as the data supply chain), HRP and SLSD are building an information architecture plan to leverage their existing, shared IT infrastructure. In addition, it is important to create a centralized structured tool to represent risks including attributes such as likelihood, consequence, contributing factors, and the evidence supporting the information in all these fields. Representing the risks in this way enables reasoning about the risks, e.g. revisiting a risk assessment when a mitigation strategy is unavailable, updating a risk assessment when new information becomes available, etc. Such a system also provides a concise way to communicate the risks both within the organization as well as with collaborators. Understanding and, hence, harnessing the human system risk-related data supply chain enhances both organizations' abilities to securely collect, integrate, and share data assets that improve human system research and operations.

Buquo, Lynn