Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DATA MINING”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Systems Development, Data Mining, and Knowledge Discovery

I worked as a NASA OSTEM virtually during the Fall 2020 term. Working in the IT division under my mentor Dr. Ali Shaykhian, our overall goal for the duration of this internship is to get a better understanding of the Visual Basic Language and how it can be used to make forms and collaboration with other workers to be more dynamics. I also worked with 3 other interns throughout this internship, using Microsoft Office to help each other to get a better understanding of how we approached our projects individually. Every week, my mentor Dr. Ali assigned me a task and a goal to finish by the end of each week. My overall project was figuring out a way to make email submissions and emails in general more dynamic for the NASA database. For example, instead of only using the same generic email for hundreds of different workers, a code can be used to send an individual email with more personalization such as each recipient’s name and personal info. I have also been assigned to create an email graphical user interface by the end of this internship. For these tasks to be made possible, I had to learn more about the capabilities of Microsoft Office, such as Macros in Microsoft Excel and Visual Basic in both Microsoft Excel and Word. Although it was challenging at first, I developed new programming skilled in a new language and actually made it possible to create this code along the way. Alongside doing the individual assignments, Dr. Ali also assigned up to replicate other intern’s work to get a better understanding of their approach and learn how to do it ourselves.

Janelisse Morales Gonzalez↗

Physics Mining of Multi-Source Data Sets

Powerful new parallel data mining algorithms can produce diagnostic and prognostic numerical models and analyses from observational data. These techniques yield higher-resolution measures than ever before of environmental parameters by fusing synoptic imagery and time-series measurements. These techniques are general and relevant to observational data, including raster, vector, and scalar, and can be applied in all Earth- and environmental science domains. Because they can be highly automated and are parallel, they scale to large spatial domains and are well suited to change and gap detection. This makes it possible to analyze spatial and temporal gaps in information, and facilitates within-mission replanning to optimize the allocation of observational resources. The basis of the innovation is the extension of a recently developed set of algorithms packaged into MineTool to multi-variate time-series data. MineTool is unique in that it automates the various steps of the data mining process, thus making it amenable to autonomous analysis of large data sets. Unlike techniques such as Artificial Neural Nets, which yield a blackbox solution, MineTool's outcome is always an analytical model in parametric form that expresses the output in terms of the input variables. This has the advantage that the derived equation can then be used to gain insight into the physical relevance and relative importance of the parameters and coefficients in the model. This is referred to as physics-mining of data. The capabilities of MineTool are extended to include both supervised and unsupervised algorithms, handle multi-type data sets, and parallelize it.

Helly, John↗

Mining Twitter Data to Augment NASA GPM Validation

The Twitter data stream is an important new source of real-time and historical global information for potentially augmenting the validation program of NASA's Global Precipitation Measurement (GPM) mission. There have been other similar uses of Twitter, though mostly related to natural hazards monitoring and management. The validation of satellite precipitation estimates is challenging, because many regions lack data or access to data, especially outside of the U.S. and in remote and developing areas. The time-varying set of "precipitation" tweets can be thought of as an organic network of rain gauges, potentially providing a widespread view of precipitation occurrence. Twitter provides a large source of crowd for crowdsourcing. During a 24-hour period in the middle of the snow storm this past March in the U.S. Northeast, we collected more than 13,000 relevant precipitation tweets with exact geolocation. The overall objective of our project is to determine the extent to which processed tweets can provide additional information that improves the validation of GPM data. Though our current effort focuses on tweets and precipitation, our approach is general and applicable to other social media and other geophysical measurements. Specifically, we have developed an operational infrastructure for processing tweets, in a format suitable for analysis with GPM data; engaged with potential participants, both passive and active, to "enrich" the Twitter stream; and inter-compared "precipitation" tweet data, ground station data, and GPM retrievals. In this presentation, we detail the technical capabilities of our tweet processing infrastructure, including data abstraction, feature extraction, search engine, context-awareness, real-time processing, and high volume (big) data processing; various means for "enriching" the Twitter stream; and results of inter-comparisons. Our project should bring a new kind of visibility to Twitter and engender a new kind of appreciation of the value of Twitter by the science research communities.

validatio↗

A Testbed Demonstration of an Intelligent Archive in a Knowledge Building System

The last decade's influx of raw data and derived geophysical parameters from several Earth observing satellites to NASA data centers has created a data-rich environment for Earth science research and applications. While advances in hardware and information management have made it possible to archive petabytes of data and distribute terabytes of data daily to a broad community of users, further progress is necessary in the transformation of data into information, and information into knowledge that can be used in particular applications in order to realize the full potential of these valuable datasets. In examining what is needed to enable this progress in the data provider environment that exists today and is expected to evolve in the next several years, we arrived at the concept of an Intelligent Archive in context of a Knowledge Building System (IA/KBS). Our prior work and associated papers investigated usage scenarios, required capabilities, system architecture, data volume issues, and supporting technologies. We identified six key capabilities of an IA/KBS: Virtual Product Generation, Significant Event Detection, Automated Data Quality Assessment, Large-Scale Data Mining, Dynamic Feedback Loop, and Data Discovery and Efficient Requesting. Among these capabilities, large-scale data mining is perceived by many in the community to be an area of technical risk. One of the main reasons for this is that standard data mining research and algorithms operate on datasets that are several orders of magnitude smaller than the actual sizes of datasets maintained by realistic earth science data archives. Therefore, we defined a test-bed activity to implement a large-scale data mining algorithm in a pseudo-operational scale environment and to examine any issues involved. The application chosen for applying the data mining algorithm is wildfire prediction over the continental U.S. This paper reports a number of observations based on our experience with this test-bed. While proof-of-concept for data mining scalability and utility has been a major goal for the research reported here, it was not the only one. The other five capabilities of an WKBS named above have been considered as well, and an assessment of the implications of our experience for these other areas will also be presented. The lessons learned through the testbed effort and presented in this paper will benefit technologists, scientists, and system operators as they consider introducing IA/KBS capabilities into production systems.

Ramapriyan, Hampapuram↗

KDD Services at the Goddard Earth Sciences Distributed Active Archive Center

NASA's Goddard Earth Sciences Distributed Active Archive Center (GES DAAC) processes, stores and distributes earth science data from a variety of remote sensing satellites. End users of the data range from instrument scientists to global change and climate researchers to federal agencies and foreign governments. Many of these users apply data mining techniques to large volumes of data (up to 1 TB) received from the GES DAAC. However, rapid advances in processing power are enabling increases in data processing that are outpacing tape drive performance and network capacity. As a result, the proportion of data that can be distributed to users continues to decrease. As mitigation, we are migrating more data mining and mining preparation activities into the data center in order to reduce the data volume that needs to be distributed and to offer the users a more useful and manageable product. This migration of activities faces a number of technical and human-factor challenges. As data reduction and mining algorithms are normally quite specific to the user's research needs, the user's algorithm must be integrated virtually unchanged into the archive environment. Also, the archive itself is busy with everyday data archive and distribution activities and cannot be dedicated to, or even impacted by, the mining activities. Therefore, we schedule KDD 'campaigns' (similar to reprocessing campaigns), during which we schedule a wholesale retrieval of specific data products, offering users the opportunity to extract information from the data being retrieved during the campaign.

Lynnes, Christopher↗

Solid Propulsion Systems, Subsystems, and Components Service Life Extension

The service life extension of solid propulsion systems, subsystems, and components will be discussed based on the service life extension of the Space Transportation System Reusable Solid Rocket Motor (RSRM) and Booster Separation Motors (BSM). The RSRM is certified for an age life of five years. In the aftermath of the Columbia accident there were a number of motors that were approaching the end of their five year service life certification. The RSRM Project initiated an assessment to determine if the service life of these motors could be extended. With the advent of the Constellation Program, a flight test was proposed that would utilize one of the RSRMs which had been returned from the launch site due to the expiration of its five year service life certification and twelve surplus Chemical Systems Division BSMs which had exceeded their eight year service life. The RSRM age life tracking philosophy which establishes when the clock starts for age life tracking will be described. The role of the following activities in service life extension will be discussed: subscale testing, accelerated aging, dissecting full scale aged hardware, static testing full scale aged motors, data mining industry data, and using the fleet leader approach. The service life certification and extension of the BSMs will also be presented.

Hundley, Nedra H.↗

Damage Detection Response Characteristics of Open Circuit Resonant (SansEC) Sensors

The capability to assess the current or future state of the health of an aircraft to improve safety, availability, and reliability while reducing maintenance costs has been a continuous goal for decades. Many companies, commercial entities, and academic institutions have become interested in Integrated Vehicle Health Management (IVHM) and a growing effort of research into "smart" vehicle sensing systems has emerged. Methods to detect damage to aircraft materials and structures have historically relied on visual inspection during pre-flight or post-flight operations by flight and ground crews. More quantitative non-destructive investigations with various instruments and sensors have traditionally been performed when the aircraft is out of operational service during major scheduled maintenance. Through the use of reliable sensors coupled with data monitoring, data mining, and data analysis techniques, the health state of a vehicle can be detected in-situ. NASA Langley Research Center (LaRC) is developing a composite aircraft skin damage detection method and system based on open circuit SansEC (Sans Electric Connection) sensor technology. Composite materials are increasingly used in modern aircraft for reducing weight, improving fuel efficiency, and enhancing the overall design, performance, and manufacturability of airborne vehicles. Materials such as fiberglass reinforced composites (FRC) and carbon-fiber-reinforced polymers (CFRP) are being used to great advantage in airframes, wings, engine nacelles, turbine blades, fairings, fuselage structures, empennage structures, control surfaces and aircraft skins. SansEC sensor technology is a new technical framework for designing, powering, and interrogating sensors to detect various types of damage in composite materials. The source cause of the in-service damage (lightning strike, impact damage, material fatigue, etc.) to the aircraft composite is not relevant. The sensor will detect damage independent of the cause. Damage in composite material is generally associated with a localized change in material permittivity and/or conductivity. These changes are sensed using SansEC. The unique electrical signatures (amplitude, frequency, bandwidth, and phase) are used for damage detection and diagnosis. An operational system and method would incorporate a SansEC sensor array on select areas of the aircraft exterior surfaces to form a "Smart skin" sensing surface. In this paper a new method and system for aircraft in-situ damage detection and diagnosis is presented. Experimental test results on seeded fault damage coupons and computational modeling simulation results are presented. NASA LaRC has demonstrated with individual sensors that SansEC sensors can be effectively used for in-situ composite damage detection of delamination, voids, fractures, and rips. Keywords: Damage Detection, Composites, Integrated Vehicle Health Monitoring (IVHM), Aviation Safety, SansEC Sensors

Dudley, Kenneth L.↗

RHSEG and Subdue: Background and Preliminary Approach for Combining these Technologies for Enhanced Image Data Analysis, Mining and Knowledge Discovery

Under a project recently selected for funding by NASA's Science Mission Directorate under the Applied Information Systems Research (AISR) program, Tilton and Cook will design and implement the integration of the Subdue graph based knowledge discovery system, developed at the University of Texas Arlington and Washington State University, with image segmentation hierarchies produced by the RHSEG software, developed at NASA GSFC, and perform pilot demonstration studies of data analysis, mining and knowledge discovery on NASA data. Subdue represents a method for discovering substructures in structural databases. Subdue is devised for general-purpose automated discovery, concept learning, and hierarchical clustering, with or without domain knowledge. Subdue was developed by Cook and her colleague, Lawrence B. Holder. For Subdue to be effective in finding patterns in imagery data, the data must be abstracted up from the pixel domain. An appropriate abstraction of imagery data is a segmentation hierarchy: a set of several segmentations of the same image at different levels of detail in which the segmentations at coarser levels of detail can be produced from simple merges of regions at finer levels of detail. The RHSEG program, a recursive approximation to a Hierarchical Segmentation approach (HSEG), can produce segmentation hierarchies quickly and effectively for a wide variety of images. RHSEG and HSEG were developed at NASA GSFC by Tilton. In this presentation we provide background on the RHSEG and Subdue technologies and present a preliminary analysis on how RHSEG and Subdue may be combined to enhance image data analysis, mining and knowledge discovery.

Tilton, James C.↗

FJET Database Project: Extract, Transform, and Load

The Data Mining & Knowledge Management team at Kennedy Space Center is providing data management services to the Frangible Joint Empirical Test (FJET) project at Langley Research Center (LARC). FJET is a project under the NASA Engineering and Safety Center (NESC). The purpose of FJET is to conduct an assessment of mild detonating fuse (MDF) frangible joints (FJs) for human spacecraft separation tasks in support of the NASA Commercial Crew Program. The Data Mining & Knowledge Management team has been tasked with creating and managing a database for the efficient storage and retrieval of FJET test data. This paper details the Extract, Transform, and Load (ETL) process as it is related to gathering FJET test data into a Microsoft SQL relational database, and making that data available to the data users. Lessons learned, procedures implemented, and programming code samples are discussed to help detail the learning experienced as the Data Mining & Knowledge Management team adapted to changing requirements and new technology while maintaining flexibility of design in various aspects of the data management project.

excel vba↗

A High-Throughput Computing Infrastructure to Generate Custom, Open Community Geothermal Datasets

The most significant challenge facing geothermal research, development, and deployment is a lack of comprehensive datasets describing the geological and economical properties of North America. Automated knowledge base construction, the process of designing algorithms to analyze text and images to programmatically build new datasets, is one possible solution to this problem. The xDD library of full-text scientific articles (https://xdd.wisc.edu) is one of the largest collections of open and controlled-access scientific documents available for knowledge base construction in the world, but it has been underutilized by experts in geothermal research. The xDD development team attributed the lack of engagement by software developers and geothermal researchers to two perceived shortcomings of the system. First, the workflow for obtaining data from xDD for local development and testing of data mining applications was unnecessarily abstruse and required significant manual intervention by xDD systems administrators. Second, although xDD already held articles from a broad cross-section of scientific literature with an emphasis on the geosciences, it did not have an explicit set of geothermal research documents that could serve as the nucleus of a geothermal data mining application. To address these issues, the Automated Data Extraction PlaTform (ADEPT) was proposed to extend the data distribution capabilities of the xDD system. The ADEPT extension added the following four key features to xDD: 1) integration of National Geothermal Data System (NGDS) documents into the xDD library to provide an explicitly geothermally-themed collection; 2) improved RESTful (i.e., https-protocol driven) web services for external partners to access xDD data for machine learning application development; 3) a web platform for end-users and xDD administrators to coordinate the development of data mining applications from the initial step of browsing available documents to the final stage of deploying a production-quality machine learning application on high-throughput computing infrastructure; and 4) the development of demonstration data mining applications to illustrate the new workflow to potential collaborators. A total of 21,674 geothermal documents from NGDS were fully ingested into the xDD library and the associated metadata is publicly available through the xDD web services; furthermore, the ADEPT web platform is now publicly accessible and fully live at https://xdd.wisc.edu/adept/.

15 GEOTHERMAL ENERGY↗

Automated Data Assimilation and Flight Planning for Multi-Platform Observation Missions

This is a progress report on an effort in which our goal is to demonstrate the effectiveness of automated data mining and planning for the daily management of Earth Science missions. Currently, data mining and machine learning technologies are being used by scientists at research labs for validating Earth science models. However, few if any of these advanced techniques are currently being integrated into daily mission operations. Consequently, there are significant gaps in the knowledge that can be derived from the models and data that are used each day for guiding mission activities. The result can be sub-optimal observation plans, lack of useful data, and wasteful use of resources. Recent advances in data mining, machine learning, and planning make it feasible to migrate these technologies into the daily mission planning cycle. We describe the design of a closed loop system for data acquisition, processing, and flight planning that integrates the results of machine learning into the flight planning process.

Oza, Nikunj↗

A Data Miner for the Information Power Grid

Grid Miner (GM) is one of the early data mining applications developed by NASA to help users obtain information from the Information Power Grid (IPG). Topics cover include: benefits of data mining, potential use of grids in data mining activities, an overview of the GM application, and a brief review of GM architecture and implementation issues. The current status of the GM system is also discussed.

Hinke, Thomas H.↗

New Directions in Space Operations Services in Support of Interplanetary Exploration

To gain access to the necessary operational processes and data in support of NASA's Lunar/Mars Exploration Initiative, new services, adequate levels of computing cycles and access to myriad forms of data must be provided to onboard spacecraft and ground based personnel/systems (earth, lunar and Martian) to enable interplanetary exploration by humans. These systems, cycles and access to vast amounts of development, test and operational data will be required to provide a new level of services not currently available to existing spacecraft, on board crews and other operational personnel. Although current voice, video and data systems in support of current space based operations has been adequate, new highly reliable and autonomous processes and services will be necessary for future space exploration activities. These services will range from the more mundane voice in LEO to voice in interplanetary travel which because of the high latencies will require new voice processes and standards. New services, like component failure predictions based on data mining of significant quantities of data, located at disparate locations, will be required. 3D or holographic representation of onboard components, systems or family members will greatly improve maintenance, operations and service restoration not to mention crew morale. Current operational systems and standards, like the Internet Protocol, will not able to provide the level of service required end to end from an end point on the Martian surface like a scientific instrument to a researcher at a university. Ground operations whether earth, lunar or Martian and in flight operations to the moon and especially to Mars will require significant autonomy that will require access to highly reliable processing capabilities, data storage based on network storage technologies. Significant processing cycles will be needed onboard but could be borrowed from other locations either ground based or onboard other spacecraft. Reliability will be a key factor with onboard and distributed backup processing an absolutely necessary requirement. Current cluster processing/Grid technologies may provide the basis for providing these services. An overview of existing services, future services that will be required and the technologies and standards required to be developed will be presented. The purpose of this paper will be to initiate a technological roadmap, albeit at a high level, of current voice, video, data and network technologies and standards (which show promise for adaptation or evolution) to what technologies and standards need to be redefined, adjusted or areas where new ones require development. The roadmap should begin the differentiation between non manned and manned processes/services where applicable. The paper will be based in part on the activities of the CCSDS Monitor and Control working group which is beginning the process of standardization of the these processes. Another element of the paper will be based on an analysis of current technologies supporting space flight processes and services at JSC, MSFC, GSFC and to a lesser extent at KSC. Work being accomplished in areas such as Grid computing, data mining and network storage at ARC, IBM and the University of Alabama at Huntsville will be researched and analyzed.

Bradford, Robert N.↗

Effects of Long Duration Spaceflight on Venous and Arterial Compliance

The visual impairment and intracranial pressure syndrome (VIIP) is a newly described space flight-associated medical condition made up of a constellation of symptoms affecting at least 34% of American astronauts who have flown International Space Station (ISS) missions. VIIP is defined primarily by visual acuity deficits and anatomical changes to eye structures, and is thought to be related to elevated intracranial pressure secondary to space flightinduced cephalad fluid shifts. Loss of visual acuity could be a significant threat to crew health and performance and may be suggestive of other adaptations with implications for years post-flight. Our primary objective is to determine whether vascular compliance is altered by space flight and whether such adaptations are related to the incidence of VIIP. In particular, we will measure ocular parameters and vascular compliance in vessels of the head and neck in astronauts who have no space flight experience, in astronauts before, during, and after space flight, and in bed rest subjects with conditions similar to space flight. Additionally, we will analyze astronaut data from the Lifetime Surveillance of Astronaut Health (LSAH) archive to determine which factors might be predictive of the development of VIIP. The project will be conducted in four separate but related parts. To understand the baseline condition of astronauts without any prior space flight experience, we will study 10 astronauts who have never flown in space by performing a comprehensive evaluation of the vasculature of the head, neck and eyes. Hemodynamic data (stroke volume and blood pressure), ocular (tonometry and ocular ultrasound), venous and arterial parameters will be acquired across a range of tilt angles (20, 10, 0, -10, -20 degrees). Vessels to be studied include the temporal, jugular, and vertebral veins and the cerebral, carotid and vertebral arteries. Ophthalmic data from the annual physical will be obtained through data sharing. To examine the relation between vascular compliance in the head and neck and the development of VIIP after a long duration space flight, we will study 10 astronauts before, during, and after long-duration ISS missions. Pre- and post-flight testing will be identical to that described above. During flight, images of the same vessels of interest will be obtained for later analysis. Ophthalmic data including VIIP scores will be obtained through data sharing from medically-required tests. To investigate the effects of age and elevated sodium intake, two potential contributors to VIIP, we will study 24 men (in two age groups: 25-35 and 45-55) during a 14 day 6deg head-down bed rest, a well-accepted analog of space flight. Standard NASA bed rest conditions will be maintained except for dietary sodium. Sodium intake will be similar to that of ISS astronauts, which is higher than consumed in previous bed rest studies. Pre- and post-bed rest testing procedures will be identical to the testing protocol described above for astronauts. Ophthalmic testing (optical coherence tomography, fundoscopy, and tonometry) will be conducted on the same day that vascular compliance measures are obtained. To identify parameters that may relate to an increase in an astronaut's susceptibility to developing VIIP, we will use data mining techniques to evaluate astronaut data obtained from the LSAH. Medical history, family history, space flight history and its related exposures, and history of high performance jet aircraft exposure will be examined for their potential relationship to ocular data. We hypothesize that the cephalad fluid shift induced by space flight will result in structural and functional adaptations in head and neck vessels leading to decreased vascular compliance and related to the development of VIIP symptoms. Further, although VIIP has not been observed in previous bed rest studies, we hypothesize that an elevated sodium intake will increase the incidence of VIIP symptoms in this space flight analog. Finally, we hypothesize that data mining analyses will reveal relationships between health history, previous exposures (including space flight and high performance aircraft), and the development of VIIP in the astronaut population.

Ribeiro, L. C.↗