Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data analytics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

CIF Report - Information Fusion and Data Analytics for Human Lunar Exploration

This project leverages the Concept Exploration Laboratory (CEL) to collect, warehouse, and augment data relevant to human lunar exploration as a platform for NA (S&MA) to develop operational data integration techniques. The project capitalizes on 16+ years of CEL experience applied to NASA, DoD, the City of Houston, the State of Texas, and private industry. The integrated data will be utilized in the two scenarios described in a definition of concept for development of a full scale data analysis suite and storage solution, useful to all JSC organizations engaged in real time operations and safety tasks, and may be useful as pathfinders for the Digital Transformation Program.

information fusion

Analysis Ready Data in Analytics Optimized Data Stores for Analysis of Big Earth Data in the Cloud

Cloud computing offers the possibility of making the analysis of Big Data approachable for a wider community due to affordable access to computing power, an ecosystem of usable tools for parallel processing, and migration of many large datasets to archives in the cloud, allowing data-proximal computing. Generally, data analysis acceleration in the cloud comes from running multiple nodes in a split-combine-apply strategy. Data systems such as the Earth Observing System Data and Information System are in a position to "pre-split" the data by storing them in a data store that is optimized for data parallel computing, i.e., an Analytics-Optimized Data Store (AODS). A variety of approaches to AODS are possible, from highly scalable databases to scalable filesystems to data formats optimized for cloud access (e.g., zarr and cloud-optimized datasets), with the optimal choice dependent on both the types of analysis and the geospatial structure of the data. A key question is how much preprocessing of the data to do, both before splitting and as the first part of the apply step. Again, the geospatial structure of the data and the analysis type influence the decision, with the added complexity of the user type. Trans-disciplinary users who are not well-versed in the nuances of quality-filtering and georeferencing of remote sensing orbit/swath/scene data tend to ask for more highly processed data, relying on the data provider to make sensible decisions on preprocessing parameters. (This accounts for the popularity of "Level 3" gridded data, despite the lower spatial resolution it provides.) In this case, data can be preprocessed before the split, resulting in higher performance in the rest of the "apply" step, which can be transformative for use cases such as interactive data exploration at scale. Discipline researchers who are experienced with remote sensing data often prefer more flexibility in customizing the preprocessing data into Analysis Ready Data, resulting in more need for on-the-fly preprocessing.

Lynnes, Christopher

An Integrated Gate Turnaround Management Concept Leveraging Big Data Analytics for NAS Performance Improvements

"Gate Turnaround" plays a key role in the National Air Space (NAS) gate-to-gate performance by receiving aircraft when they reach their destination airport, and delivering aircraft into the NAS upon departing from the gate and subsequent takeoff. The time spent at the gate in meeting the planned departure time is influenced by many factors and often with considerable uncertainties. Uncertainties such as weather, early or late arrivals, disembarking and boarding passengers, unloading/reloading cargo, aircraft logistics/maintenance services and ground handling, traffic in ramp and movement areas for taxi-in and taxi-out, and departure queue management for takeoff are likely encountered on the daily basis. The Integrated Gate Turnaround Management (IGTM) concept is leveraging relevant historical data to support optimization of the gate operations, which include arrival, at the gate, departure based on constraints (e.g., available gates at the arrival, ground crew and equipment for the gate turnaround, and over capacity demand upon departure), and collaborative decision-making. The IGTM concept provides effective information services and decision tools to the stakeholders, such as airline dispatchers, gate agents, airport operators, ramp controllers, and air traffic control (ATC) traffic managers and ground controllers to mitigate uncertainties arising from both nominal and off-nominal airport gate operations. IGTM will provide NAS stakeholders customized decision making tools through a User Interface (UI) by leveraging historical data (Big Data), net-enabled Air Traffic Management (ATM) live data, and analytics according to dependencies among NAS parameters for the stakeholders to manage and optimize the NAS performance in the gate turnaround domain. The application will give stakeholders predictable results based on the past and current NAS performance according to selected decision trees through the UI. The predictable results are generated based on analysis of the unique airport attributes (e.g., runway, taxiway, terminal, and gate configurations and tenants), and combined statistics from past data and live data based on a specific set of ATM concept-of-operations (ConOps) and operational parameters via systems analysis using an analytic network learning model. The IGTM tool will then bound the uncertainties that arise from nominal and off-nominal operational conditions with direct assessment of the gate turnaround status and the impact of a certain operational decision on the NAS performance, and provide a set of recommended actions to optimize the NAS performance by allowing stakeholders to take mitigation actions to reduce uncertainty and time deviation of planned operational events. An IGTM prototype was developed at NASA Ames Simulation Laboratories (SimLabs) to demonstrate the benefits and applicability of the concept. A data network, using the System Wide Information Management (SWIM)-like messaging application using the ActiveMQ message service, was connected to the simulated data warehouse, scheduled flight plans, a fast-time airport simulator, and a graphic UI. A fast-time simulation was integrated with the data warehouse or Big Data/Analytics (BAI), scheduled flight plans from Aeronautical Operational Control AOC, IGTM Controller, and a UI via a SWIM-like data messaging network using the ActiveMQ message service, illustrated in Figure 1, to demonstrate selected use-cases showing the benefits of the IGTM concept on the NAS performance.

Efficent ATM systems

Information Management Platform for Data Analytics and Aggregation (IMPALA) System Design Document

The System Design document tracks the design activities that are performed to guide the integration, installation, verification, and acceptance testing of the IMPALA Platform. The inputs to the design document are derived from the activities recorded in Tasks 1 through 6 of the Statement of Work (SOW), with the proposed technical solution being the completion of Phase 1-A. With the documentation of the architecture of the IMPALA Platform and the installation steps taken, the SDD will be a living document, capturing the details about capability enhancements and system improvements to the IMPALA Platform to provide users in development of accurate and precise analytical models. The IMPALA Platform infrastructure team, data architecture team, system integration team, security management team, project manager, NASA data scientists and users are the intended audience of this document. The IMPALA Platform is an assembly of commercial-off-the-shelf (COTS) products installed on an Apache-Hadoop platform. User interface details for the COTS products will be sourced from the COTS tools vendor documentation. The SDD is a focused explanation of the inputs, design steps, and projected outcomes of every design activity for the IMPALA Platform through installation and validation.

Carnell, Andrew

Understanding the International Space Station Crew Perspective following Long-Duration Missions through Data Analytics & Visualization of Crew Feedback

The International Space Station (ISS) first became a home and research laboratory for NASA and International Partner crewmembers over 16 years ago. Each ISS mission lasts approximately 6 months and consists of three to six crewmembers. After returning to Earth, most crewmembers participate in an extensive series of 30+ debriefs intended to further understand life onboard ISS and allow crews to reflect on their experiences. Examples of debrief data collected include ISS crew feedback about sleep, dining, payload science, scheduling and time planning, health & safety, and maintenance. The Flight Crew Integration (FCI) Operational Habitability (OpsHab) team, based at Johnson Space Center (JSC), is a small group of Human Factors engineers and one stenographer that has worked collaboratively with the NASA Astronaut office and ISS Program to collect, maintain, disseminate and analyze this data. The database provides an exceptional and unique resource for understanding the "crew perspective" on long duration space missions. Data is formatted and categorized to allow for ease of search, reporting, and ultimately trending, in order to understand lessons learned, recurring issues and efficiencies gained over time. Recently, the FCI OpsHab team began collaborating with the NASA JSC Knowledge Management team to provide analytical analysis and visualization of these over 75,000 crew comments in order to better ascertain the crew's perspective on long duration spaceflight and gain insight on changes over time. In this initial phase of study, a text mining framework was used to cluster similar comments and develop measures of similarity useful for identifying relevant topics affecting crew health or performance, locating similar comments when a particular issue or item of operational interest is identified, and providing search capabilities to identify information pertinent to future spaceflight systems and processes for things like procedure development and training. In addition, the comments were scored for sentiment using a polarity scoring algorithm to identify both positive and negative comments for particular groups and clusters, allowing the team to make analytically informed decisions regarding future hardware and operating procedures. The use of polarity scoring with time series analysis was used to provide insight into how crew health and habitability is changing throughout various spaceflight increments or the station lifecycle as a whole. Finally, a visualization framework was developed to address the needs of the end users to search for and analyze comments by user, category or mission. This paper will discuss how the use of an analytical framework in conjunction with the current human interface, improved the understanding of crew perspective and shortened the time for analysis allowing for more informed decisions and rapid development of improvements. These methods are significantly optimizing the way that this valuable data can be assessed and applied to current and future spaceflight design and development. This collaboration allows the FCI OpsHab team to effectively analyze and share data in a more automated and timely fashion. Trends are no longer derived manually and can be illustrated effectively and accurately with these evolving techniques to an ever growing group of human spaceflight end users.

Bryant, Cody

Ask-The-Expert: Minimizing Human Review for Big Data Analytics Through Active Learning

In this CIF project, we worked toward semi-automating knowledge discovery from anomaly detection algorithms through the use of active learning. Active learning is an area of research within machine learning that uses an "expert in the loop" to learn from large data sets that have very few annotations or labels available, and where providing such labels is expensive. In our case, the task can be defined as the identification of safety events from flight operational data. Since traditional anomaly detection algorithms cannot differentiate between operationally relevant and irrelevant statistical anomalies, Subject Matter Experts (SMEs) have a lengthy and expensive burden of investigating every example identified by the detection algorithm, classifying and labeling them as relevant or irrelevant. Active learningidentifies the unlabeled example for which a label would most improve the classifier, asks the domain expert for a label, and repeats this process until there are no more resources (time, budget) available for labeling or a minimum required performance is reached. A positive label indicates an operationally significant safety event whereas a negative label indicates otherwise. Based on these few labels we propose to build an active learning system that utilizes the SME's time in the most effective manner by iteratively asking for labels for as few informative instances as possible. Our work was proposed to be a stepping stone toward implementation and deployment of the system with user interface to be pursued by the Aviation Operations and Safety Program (AOSP) given its interest in safety monitoring and discovery of safety incidents.

aviation safety

Cloud Giovanni: Reining in Costs and Improving Performance with Analytical Data Stores Using Scalable Serverless Architecture

Giovanni is the Geospatial Interactive Online Visualization ANd aNalysis Infrastructure developed at NASA GES DISC which provides a simple and intuitive way to visualize, analyze, and access vast amounts of Earth science data. It receives large number of user requests each day for a variety of analysis and visualization services, which leads to the big data challenge of serving gradually increasing large data volumes with diverse statistical algorithms. We hereby propose a multi-dimensional accumulation method which provides fast and cost-efficient cloud analysis for diverse services including both area averaging and time averaging. This method involves the weighted volume integration over multiple variable dimensions (time and space), and is implemented in AWS using Athena providing serverless and highly scalable data analysis. Compared to the standard method, this approach dramatically reduced the computational time by order of magnitude with a minimal AWS cost incurred. For example, for a benchmark of 10-year area averaging over the 1x1 degree daily variable, the computational time was reduced from minutes to seconds, and the Athena cost is only $5 for 100,000 requests.

Zhang, Hailiang