Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data mining”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Data Fusion and Mining Techniques to Map Water Use and Drought across Spatial and Temporal Scales

As the world’s water resources come under increasing tension due to dual stressors of climate change and population growth, accurate knowledge of water consumption through evapotranspiration (ET) over a range in spatial scales will be critical in developing adaptation strategies. Remote sensing methods for monitoring consumptive water use (e.g, ET) are becoming increasingly important, especially in areas of significant water and food insecurity. One method to estimate ET from satellite-based methods, the Atmosphere Land Exchange Inverse (ALEXI) model uses the change in mid-morning land surface temperature to estimate the partitioning of sensible and latent heat fluxes which are then used to estimate daily ET. This presentation will outline several recent enhancements to the ALEXI modeling system, with a focus on global ET and drought monitoring. Until recently, ALEXI has been limited to areas with high resolution temporal sampling of geostationary sensors. The use of geostationary sensors makes global mapping a complicated process, especially for real-time applications, as data from as many as five different sensors are required to be ingested and harmonized to create a global mosaic. However, our research team has developed a new and novel method of using twice-daily observations from polar-orbiting sensors such as MODIS and VIIRS to estimate the mid-morning rise in LST that is used to drive the energy balance estimations within ALEXI. This allows the method to be applied globally using a single sensor (in this case, initially MODIS with a planned transition to VIIRS) rather than a global compositing of all available geostationary data. Other advantages of this new method include the higher spatial resolution provided by MODIS and VIIRS and the increased sampling at high latitudes where oblique view angles limit the utility of geostationary sensors. This presentation will focus on global applications for mapping water use and drought using data mining and data fusion across spatial scales extending from 5-km to 30-m “field-scale” estimates.

Christopher Hain↗

Polynomial chaos expansions on principal geodesic Grassmannian submanifolds for surrogate modeling and uncertainty quantification

In this work we introduce a manifold learning-based surrogate modeling framework for uncertainty quantification in high-dimensional stochastic systems. Our first goal is to perform data mining on the available simulation data to identify a set of low-dimensional (latent) descriptors that efficiently parameterize the response of the high-dimensional computational model. To this end, we employ Principal Geodesic Analysis on the Grassmann manifold of the response to identify a set of disjoint principal geodesic submanifolds, of possibly different dimension, that captures the variation in the data. Since operations on the Grassmann require the data to be concentrated, we propose an adaptive algorithm based on Riemannian K-means and the minimization of the sample Fréchet variance on the Grassmann manifold to identify “local” principal geodesic submanifolds that represent different system behavior across the parameter space. Polynomial chaos expansion is then used to construct a mapping between the random input parameters and the projection of the response on these local principal geodesic submanifolds. Here, the method is demonstrated on four test cases, a toy-example that involves points on a hypersphere, a Lotka-Volterra dynamical system, a continuous-flow stirred-tank chemical reactor system, and a two-dimensional Rayleigh-Bénard convection problem.

42 ENGINEERING↗

Occupant Protection at NASA

This slide presentation reviews NASA's efforts to arrive at protection of occupants of the ORION space craft on landing. An Abbreviated Injury Scale (AIS) has been developed, it is an anatomically-based, consensus-derived, global severity scoring system that classifies each injury by body region according to its relative importance on a 6-point ordinal scale. It reviews an Operationmally Relevant Injury Scale (ORIS), a classification methodology, and shows charts that detail the results of applying this ORIS to the injury databases. One chart uses NASCAR injury classification. It discusses providing a context for the level of risk inherent in the Orion landings in terms that people understand and have a sense for. For example is the risk of injury during an Orion landing roughly the same, better or worse than: An aircraft carrier landing, a NASCAR crash, or a helicopter crash, etc? The data for NASCAR and Indy Racing league (IRL) racing crash and injury data was reviewed. The risk from the Air Force, Navy, and Army injury data was also reviewed. Past NASA and the Soyuz programs injury risks are also reviewed. The work is an attempt to formulate a recommendation to the Orion Project for an acceptable level of injury risk associated with Nominal and Off-Nominal landing cases. The presentation also discusses the data mining and use of the data to Validate NASA Operationally-Relevant Injury Scale (NORIS) / Military Operationally-Relevant Injury Scale (MORIS), developing injury risk criteria, the types of data that are required, NASCAR modeling techniques and crash data, and comparison with the Brinkley model. The development of injury risk curves for each biodynamic response parameter is discussed. One of the main outcomes of this work is to establish an accurate Automated Test Dummy (ATD) that can be used to measure human tolerances.

Somers, Jeffrey↗

Technology Alignment and Portfolio Prioritization (TAPP): Advanced Methods in Strategic Analysis, Technology Forecasting and Long Term Planning for Human Exploration and Operations, Advanced Exploration Systems and Advanced Concepts

The Advanced Concepts Office (ACO) at NASA, Marshall Space Flight Center is expanding its current technology assessment methodologies. ACO is developing a framework called TAPP that uses a variety of methods, such as association mining and rule learning from data mining, structure development using a Technological Innovation System (TIS), and social network modeling to measure structural relationships. The role of ACO is to 1) produce a broad spectrum of ideas and alternatives for a variety of NASA's missions, 2) determine mission architecture feasibility and appropriateness to NASA's strategic plans, and 3) define a project in enough detail to establish an initial baseline capable of meeting mission objectives ACO's role supports the decision­-making process associated with the maturation of concepts for traveling through, living in, and understanding space. ACO performs concept studies and technology assessments to determine the degree of alignment between mission objectives and new technologies. The first step in technology assessment is to identify the current technology maturity in terms of a technology readiness level (TRL). The second step is to determine the difficulty associated with advancing a technology from one state to the next state. NASA has used TRLs since 1970 and ACO formalized them in 1995. The DoD, ESA, Oil & Gas, and DoE have adopted TRLs as a means to assess technology maturity. However, "with the emergence of more complex systems and system of systems, it has been increasingly recognized that TRL assessments have limitations, especially when considering [the] integration of complex systems." When performing the second step in a technology assessment, NASA requires that an Advancement Degree of Difficulty (AD2) method be utilized. NASA has used and developed or used a variety of methods to perform this step: Expert Opinion or Delphi Approach, Value Engineering or Value Stream, Analytical Hierarchy Process (AHP), Technique for the Order of Prioritization by Similarity to Ideal Solution (TOPSIS), and other multi­‐criteria decision-making methods. These methods can be labor-intensive, often contain cognitive or parochial bias, and do not consider the competing prioritization between mission architectures. Strategic Decision-Making (SDM) processes cannot be properly understood unless the context of the technology is understood. This makes assessing technological change particularly challenging due to the relationships "between incumbent technology and the incumbent (innovation) system in relation to the emerging technology and the emerging innovation system." The central idea in technology dynamics is to consider all activities that contribute to the development, diffusion, and use of innovations as system functions. Bergek defines system functions within a TIS to address what is actually happening and has a direct influence on the ultimate performance of the system and technology development. ACO uses similar metrics and is expanding these metrics to account for the structure and context of the technology. At NASA technology and strategy is strongly interrelated. NASA's Strategic Space Technology Investment Plan (SSTIP) prioritizes those technologies essential to the pursuit of NASA's missions and national interests. The SSTIP is strongly coupled with NASA's Technology Roadmaps to provide investment guidance during the next four years, within a twenty-year horizon. This paper discusses the methods ACO is currently developing to better perform technology assessments while taking into consideration Strategic Alignment, Technology Forecasting, and Long Term Planning.

Funaro, Gregory V.↗

Validation of A Simulation Environment for Future Space Traffic Management

This study presents initial results of a newly developed simulation environment intendedto explore and assess future Space Traffic Management (STM) scenarios. The number ofnew Resident Space Objects (RSOs) in near-Earth orbit is expected to increase significantlywith the expected future deployment of a number of large constellations. These futurescenarios involve the addition of many tens of thousands of new RSOs, making the analysisinto their impact on collision risk extend beyond what traditional data-mining of present-dayconjunction data can reliably predict. To address this, a robust simulation environment wasdeveloped that implements a full force-model for orbit propagation, and computes continuousall-on-all conjunction statistics for arbitrarily large catalog sizes and simulation timeframes.Collision avoidance and station-keeping maneuvers can be optionally implemented based onconfigurable user inputs including physical characteristics and spacecraft meta-data (e.g.,commercial/government, owner country, etc.). Constellation build-out and de-orbit scenarioswere also implemented and modeled based on real-data analysis. Validation of the simulationresults was a critical component of the simulation development, and was done using the currentcatalog with comparisons against both public and internal NASA data-sets. The comparisonsdemonstrate that, with the appropriate settings, representative levels of conjunction ratesand probabilities can be obtained, providing confidence that the simulation tool can generatemeaningful outcomes for test scenarios. As an initial demonstration of the tool’s capabilities,year-long simulations with station-keeping were conducted to examine conjunction historiesusing both the current object catalog (5800 active satellites) and a hypothetical 60,000 objectscenario involving five potential large constellations. Output metrics include the number ofconjunction events, estimates of collision consequence (fragmentation), delta-V maneuver costs,and the probability of at least one collision occurring. The results highlight the potential thatthe simulation tool has for incorporating and running performance comparisons between, e.g.,various sets of maneuver guidelines, industry norms, and definitions of risk, with the overallobjective of providing actionable data to STM policy makers. The presentation will providean overview of the simulation development and validation efforts, as well as a discussion ofobservations gathered from the initial simulations performed.

conjunction assessment↗

Clustering Days with Similar Airport Weather Conditions

On any given day, traffic flow managers must often rely on past experience and intuition when developing traffic flow management initiatives that mitigate imbalances between the aircraft demand and the weather impacted airport capacity. The goal of this study was to build on recent efforts to apply data mining classification and clustering algorithms to vast archives of historical weather and air traffic data to identify patterns and past decisions that can ultimately inform day-of-operations decision-making. More specifically, this study identified similar weather impacted days at select U.S. airports, and analyzed the traffic management initiatives implemented on these representative days. The identification of the similar days was accomplished by applying a decision tree algorithm to the hourly Localized Aviation Model Output Statistics Program observations and the arrival delays for Newark Liberty International Airport. The branches from the trained decision tree were subsequently pruned to identify four weather conditions that resulted in medium to high delays for the arrivals scheduled to Newark in 2012. Using these weather conditions, four, daily airport-level Weather Impacted Traffic Index values were calculated using the Localized Aviation Model Output Statistics Program observations and the 2012 scheduled arrival counts from the FAAs Aviation System Performance Metric system. The four, daily Weather Impacted Traffic Index values for 2012 were subsequently clustered using an Expectation Maximization clustering algorithm, and nine unique types of weather days at Newark were identified. By far the most prominent type of day at Newark was a day associated with relatively good weather conditions, where there was little convective activity, winds were low, ceilings and visibility were high and there was little precipitation. Moderate levels of convective activity characterized the next most prominent type of day. Days with persistently high winds or low ceiling and visibility levels were relatively rare in 2012. Lastly, the frequency at which Ground Delay Programs, Ground Stops and Miles-in-Trail restrictions were implemented on each of the typical types of days at Newark were analyzed. Based on the results, it does appear as if the usage of Miles-in-Trail, Ground Delay Program and Ground Stop restrictions correlates well with the severity of the weather associated with each unique type of weather impacted day at Newark. Furthermore, the results demonstrate that it is feasible to use historical weather and air traffic archives to provide guidance on the types of traffic management restrictions to implement in response to the weather conditions impacting an airport.

traffic flow management↗

Clustering Days with Similar Airport Weather Conditions

On any given day, traffic flow managers must often rely on past experience and intuition when developing traffic flow management initiatives that mitigate imbalances between the aircraft demand and the weather impacted airport capacity. The goal of this study was to build on recent efforts to apply data mining classification and clustering algorithms to vast archives of historical weather and air traffic data to identify patterns and past decisions that can ultimately inform day-of-operations decision-making. More specifically, this study identified similar weather impacted days at select U.S. airports, and analyzed the traffic management initiatives implemented on these representative days. The identification of the similar days was accomplished by applying a decision tree algorithm to the hourly Localized Aviation Model Output Statistics Program observations and the arrival delays for Newark Liberty International Airport. The branches from the trained decision tree were subsequently pruned to identify four weather conditions that resulted in medium to high delays for the arrivals scheduled to Newark in 2012. Using these weather conditions, four, daily airport-level Weather Impacted Traffic Index values were calculated using the Localized Aviation Model Output Statistics Program observations and the 2012 scheduled arrival counts from the FAAs Aviation System Performance Metric system. The four, daily Weather Impacted Traffic Index values for 2012 were subsequently clustered using an Expectation Maximization clustering algorithm, and nine unique types of weather days at Newark were identified. By far the most prominent type of day at Newark was a day associated with relatively good weather conditions, where there was little convective activity, winds were low, ceilings and visibility were high and there was little precipitation. Moderate levels of convective activity characterized the next most prominent type of day. Days with persistently high winds or low ceiling and visibility levels were relatively rare in 2012. Lastly, the frequency at which Ground Delay Programs, Ground Stops and Miles-in-Trail restrictions were implemented on each of the typical types of days at Newark were analyzed. Based on the results, it does appear as if the usage of Miles-in-Trail, Ground Delay Program and Ground Stop restrictions correlates well with the severity of the weather associated with each unique type of weather impacted day at Newark. Furthermore, the results demonstrate that it is feasible to use historical weather and air traffic archives to provide guidance on the types of traffic management restrictions to implement in response to the weather conditions impacting an airport.

weather↗

RNAV STAR Procedural Adherence

Flight crews and air traffic controllers have reported many safety concerns regarding area navigation standard terminal arrival routes (RNAV STARs). However, our information sources to quantify these issues are limited to subjective reporting and time consuming case-by-case investigations. This work is a preliminary study into the objective performance of instrument procedures and provides a framework to track procedural concepts and assess design functionality. We created a tool and analysis methods for gauging aircraft adherence as it relates to RNAV STARs. This information is vital for comprehensive understanding of how our air traffic behaves. In this exploratory archival study, we mined the performance of 24 major US airports over the preceding three years. Overlaying radar track data on top of RNAV STAR routes provided a comparison between aircraft flight paths and the waypoint positions and altitude restrictions. NASA Ames Supercomputing resources were utilized to perform the data mining and processing. We assessed STARs by lateral transition path (full-lateral), vertical restrictions (full-lateralfull-vertical), and skipped waypoints (skips). In addition, we graphed aircraft altitudes relative to the altitude restrictions and their occurrence rates. Full-lateral adherence was generally greater than Full-lateralfull-vertical, but the difference between the rates was not always consistent. Full-lateralfull-vertical adherence medians of the 2016 procedures ranged from 0 in KDEN (Denver) to 21 in KMEM (Memphis). Waypoint skips ranged from 0 to nearly 100 for specific waypoints. Altitudes restrictions were sometimes missed by systematic amounts in 1000 ft. increments from the restriction, creating multi-modal distributions. Other times, altitude misses looked to be more normally distributed around the restriction. This tool may aid in providing acceptability metrics as well as risk assessment information.

rnav↗

Pre-Launch GOES-R Risk Reduction Activities for the Geostationary Lightning Mapper

The GOES-R Geostationary Lightning Mapper (GLM) is a new instrument planned for GOES-R that will greatly improve storm hazard nowcasting and increase warning lead time day and night. Daytime detection of lightning is a particularly significant technological advance given the fact that the solar illuminated cloud-top signal can exceed the intensity of the lightning signal by a factor of one hundred. Our approach is detailed across three broad themes which include: Data Processing Algorithm Readiness, Forecast Applications, and Radiance Data Mining. These themes address how the data will be processed and distributed, and the algorithms and models for developing, producing, and using the data products. These pre-launch risk reduction activities will accelerate the operational and research use of the GLM data once GOES-R begins on-orbit operations. The GLM will provide unprecedented capabilities for tracking thunderstorms and earlier warning of impending severe and hazardous weather threats. By providing direct information on lightning initiation, propagation, extent, and rate, the GLM will also capture the updraft dynamics and life cycle of convective storms, as well as internal ice precipitation processes. The GLM provides information directly from the heart of the thunderstorm as opposed to cloud-top only. Nowcasting applications enabled by the GLM data will expedite the warning and response time of emergency management systems, improve the dispatch of electric power utility repair crews, and improve airline routing around thunderstorms thereby improving safety and efficiency, saving fuel and reducing delays. The use of GLM data will assist the Bureau of Land Management (BLM) and the Forest Service in quickly detecting lightning ground strikes that have a high probability of causing fires. Finally, GLM data will help assess the role of thunderstorms and deep convection in global climate, and will improve regional air quality and global chemistry/climate modeling. The GLM has a robust design that benefits and improves upon its strong heritage of NASA-developed LEO predecessors, the Optical Transient Detector (OTD) and the Lightning Imaging Sensor (LIS). GLM will have a substantially larger number of pixels within the focal plane, two lens systems, and multiple Real-Time Event Processors REPS for on-board event detection and data compression to provide continuous observations of the Americas and adjacent oceans.

Goodman, S. J.↗

Using ADOPT Algorithm and Operational Data to Discover Precursors to Aviation Adverse Events

The US National Airspace System (NAS) is making its transition to the NextGen system and assuring safety is one of the top priorities in NextGen. At present, safety is managed reactively (correct after occurrence of an unsafe event). While this strategy works for current operations, it may soon become ineffective for future airspace designs and high density operations. There is a need for proactive management of safety risks by identifying hidden and "unknown" risks and evaluating the impacts on future operations. To this end, NASA Ames has developed data mining algorithms that finds anomalies and precursors (high-risk states) to safety issues in the NAS. In this paper, we describe a recently developed algorithm called ADOPT that analyzes large volumes of data and automatically identifies precursors from real world data. Precursors help in detecting safety risks early so that the operator can mitigate the risk in time. In addition, precursors also help identify causal factors and help predict the safety incident. The ADOPT algorithm scales well to large data sets and to multidimensional time series, reduce analyst time significantly, quantify multiple safety risks giving a holistic view of safety among other benefits. This paper details the algorithm and includes several case studies to demonstrate its application to discover the "known" and "unknown" safety precursors in aviation operation.

aviation safet↗

Alpha Seeding for Support Vector Machines

A key practical obstacle in applying support vector machines to many large-scale data mining tasks is that SVM's generally scale quadratically (or worse) in the number of examples or support vectors.

Support Vector Machines Alpha Seeding Data Mining↗

Applications of Anomaly Detection and Precursor Identification in Airspace Operations

As we continue to advance the U.S. National Airspace into the next generation of air traffic, we face challenges in both increase in complexity, as well as, a significant growth in traffic volume. Addressing these challenges, while maintaining the same level of safety is an important application of data mining. Because of these significant shifts in airspace design and usage there is a need to identify current and emergent safety risks along with their potential precursors. In recent years NASA has made advancements in developing scalable methods to address this effort in the Big Data paradigm. Multiple kernel anomaly detection approaches have been employed on both surveillance radar data and flight operational quality assurance data to identify operationally significant safety risks. Additionally, events have been explored with a recently developed precursor identification tool to discover states that reveal an increased probability of a safety event. These tools can be used to discover emerging safety risks that may not be currently monitored, which allows for mitigation tactics to be employed and ultimately make the overall airspace safer. This talk will discuss an overview of these methods and a discussion of the findings.

anomaly detection↗

LIB Design Module for Grid Energy System Application

We will employ a machine learning approach with intelligent data mining and database construction to analyze enormous data repositories for identifying and extracting geographic-dependent cell design specifications from publicly accessible grid-scale energy storage usage databases in an automated way at scale.

Liu, Dianying [Pacific Northwest National Laborato↗

Interactive Computer Graphics

Aerospace data analysis tools that significantly reduce the time and effort needed to analyze large-scale computational fluid dynamics simulations have emerged this year. The current approach for most postprocessing and visualization work is to explore the 3D flow simulations with one of a dozen or so interactive tools. While effective for analyzing small data sets, this approach becomes extremely time consuming when working with data sets larger than one gigabyte. An active area of research this year has been the development of data mining tools that automatically search through gigabyte data sets and extract the salient features with little or no human intervention. With these so-called feature extraction tools, engineers are spared the tedious task of manually exploring huge amounts of data to find the important flow phenomena. The software tools identify features such as vortex cores, shocks, separation and attachment lines, recirculation bubbles, and boundary layers. Some of these features can be extracted in a few seconds; others take minutes to hours on extremely large data sets. The analysis can be performed off-line in a batch process, either during or following the supercomputer simulations. These computations have to be performed only once, because the feature extraction programs search the entire data set and find every occurrence of the phenomena being sought. Because the important questions about the data are being answered automatically, interactivity is less critical than it is with traditional approaches.

Kenwright, David↗

A Local Scalable Distributed Expectation Maximization Algorithm for Large Peer-to-Peer Networks

This paper offers a local distributed algorithm for expectation maximization in large peer-to-peer environments. The algorithm can be used for a variety of well-known data mining tasks in a distributed environment such as clustering, anomaly detection, target tracking to name a few. This technology is crucial for many emerging peer-to-peer applications for bioinformatics, astronomy, social networking, sensor networks and web mining. Centralizing all or some of the data for building global models is impractical in such peer-to-peer environments because of the large number of data sources, the asynchronous nature of the peer-to-peer networks, and dynamic nature of the data/network. The distributed algorithm we have developed in this paper is provably-correct i.e. it converges to the same result compared to a similar centralized algorithm and can automatically adapt to changes to the data and the network. We show that the communication overhead of the algorithm is very low due to its local nature. This monitoring algorithm is then used as a feedback loop to sample data from the network and rebuild the model when it is outdated. We present thorough experimental results to verify our theoretical claims.

Bhaduri, Kanishka↗

The Ocean Carbon States Database: A Proof-of-Concept Application of Cluster Analysis in the Ocean Carbon Cycle

In this paper, we present a database of the basic regimes of the carbon cycle in the ocean, the 'ocean carbon states', as obtained using a data mining/pattern recognition technique in observation-based as well as model data. The goal of this study is to establish a new data analysis methodology, test it and assess its utility in providing more insights into the regional and temporal variability of the marine carbon cycle. This is important as advanced data mining techniques are becoming widely used in climate and Earth sciences and in particular in studies of the global carbon cycle, where the interaction of physical and biogeochemical drivers confounds our ability to accurately describe, understand, and predict CO2 concentrations and their changes in the major planetary carbon reservoirs. In this proof-of-concept study, we focus on using well-understood data that are based on observations, as well as model results from the NASA Goddard Institute for Space Studies (GISS) climate model. Our analysis shows that ocean carbon states are associated with the subtropical-subpolar gyre during the colder months of the year and the tropics during the warmer season in the North Atlantic basin. Conversely, in the Southern Ocean, the ocean carbon states can be associated with the subtropical and Antarctic convergence zones in the warmer season and the coastal Antarctic divergence zone in the colder season. With respect to model evaluation, we find that the GISS model reproduces the cold and warm season regimes more skillfully in the North Atlantic than in the Southern Ocean and matches the observed seasonality better than the spatial distribution of the regimes. Finally, the ocean carbon states provide useful information in the model error attribution. Model air-sea CO2 flux biases in the North Atlantic stem from wind speed and salinity biases in the subpolar region and nutrient and wind speed biases in the subtropics and tropics. Nutrient biases are shown to be most important in the Southern Ocean flux bias.

carbon cycle↗