Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Improve Data Mining and Knowledge Discovery Through the Use of MatLab

Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(R) (MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.

Shaykhian, Gholam Ali↗

Improve Data Mining and Knowledge Discovery through the use of MatLab

Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(TradeMark)(MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.

Shaykahian, Gholan Ali↗

Vostok Subglacial Lake: A Review of Geophysical Data Regarding Its Discovery and Topographic Setting

Vostok Subglacial Lake is the largest and best known sub-ice lake in Antarctica. The establishment of its water depth (>500 m) led to an appreciation that such environments may be habitats for life and could contain ancient records of ice sheet change, which catalyzed plans for exploration and research. Here we discuss geophysical data used to identify the lake and the likely physical, chemical, and biological processes that occur in it. The lake is more than 250 km long and around 80 km wide in one place. It lies beneath 4.2 to 3.7 km of ice and exists because background levels of geothermal heating are sufficient to warm the ice base to the pressure melting value. Seismic and gravity measurements show the lake has two distinct basins. The Vostok ice core extracted >200 m of ice accreted from the lake to the ice sheet base. Analysis of this ice has given valuable insights into the lake s biological and chemical setting. The inclination of the ice-water interface leads to differential basal melting in the north versus freezing in the south, which excites circulation and potential mixing of the water. The exact nature of circulation depends on hydrochemical properties, which are not known at this stage. The age of the subglacial lake is likely to be as old as the ice sheet (approx.14 Ma). The age of the water within the lake will be related to the age of the ice melting into it and the level of mixing. Rough estimates put that combined age as approx.1 Ma.

Siegert, Martin J.↗

Enhancing Discovery, Search, and Access of NASA Hydrological Data by Leveraging GEOSS

An ongoing NASA-funded project has removed a longstanding barrier to accessing NASA data (i.e., accessing archived time-step array data as point-time series) for selected variables of the North American and Global Land Data Assimilation Systems (NLDAS and GLDAS, respectively) and other EOSDIS (Earth Observing System Data Information System) data sets (e.g., precipitation, soil moisture). These time series (data rods) are pre-generated. Data rods Web services are accessible through the CUAHSI Hydrologic Information System (HIS) and the Goddard Earth Sciences Data and Information Services Center (GES DISC) but are not easily discoverable by users of other non-NASA data systems. The Global Earth Observation System of Systems (GEOSS) is a logical mechanism for providing access to the data rods. An ongoing GEOSS Water Services project aims to develop a distributed, global registry of water data, map, and modeling services cataloged using the standards and procedures of the Open Geospatial Consortium and the World Meteorological Organization. The ongoing data rods project has demonstrated the feasibility of leveraging the GEOSS infrastructure to help provide access to time series of model grid information or grids of information over a geographical domain for a particular time interval. A recently-begun, related NASA-funded ACCESS-GEOSS project expands on these prior efforts. Current work is focused on both improving the performance of the generation of on-the-fly (OTF) data rods and the Web interfaces from which users can easily discover, search, and access NASA data.

GEOSS↗

Global Change Master Directory (GCMD) Keyword Management Process and Lifecycle

The Global Change Master Directory (GCMD) keywords are a hierarchical set of controlled vocabulary covering the Earth science disciplines that have been evolving for over 25 years. The process for how these keywords have been curated and reviewed has also evolved. This presentation will convey the process for reviewing and approving the GCMD keywords, including fast track and yearly reviews through the ESDIS Standards Office, and how Earth science users can influence keyword additions and modifications. The presentation will also highlight how keywords facilitate the discovery of EOSDIS data and services and how other organizations are using the GCMD keywords.

EOSDIS↗

Previously Unrecognized Large Lunar Impact Basins Revealed by Topographic Data

The discovery of a large population of apparently buried impact craters on Mars, revealed as Quasi- Circular Depressions (QCDs) in Mars Orbiting Laser Altimeter (MOLA) data [1,2,3] and as Circular Thin Areas (CTAs) [4] in crustal thickness model data [5] leads to the obvious question: are there unrecognized impact features on the Moon and other bodies in the solar system? Early analysis of Clementine topography revealed several large impact basins not previously known [6,7], so the answer certainly is "Yes." How large a population of previously undetected impact basins, their size frequency distribution, and how much these added craters and basins will change ideas about the early cratering history and Late Heavy Bombardment on the Moon remains to be determined. Lunar Orbiter Laser Altimeter (LOLA) data [8] will be able to address these issues. As a prelude, we searched the state-of-the-art global topographic grid for the Moon, the Unified Lunar Control Net (ULCN) [9] for evidence of large impact features not previously recognized by photogeologic mapping, as summarized by Wilhelms [lo].

Frey, Herbert V.↗

Direct Manipulation in Virtual Reality

Virtual Reality interfaces offer several advantages for scientific visualization such as the ability to perceive three-dimensional data structures in a natural way. The focus of this chapter is direct manipulation, the ability for a user in virtual reality to control objects in the virtual environment in a direct and natural way, much as objects are manipulated in the real world. Direct manipulation provides many advantages for the exploration of complex, multi-dimensional data sets, by allowing the investigator the ability to intuitively explore the data environment. Because direct manipulation is essentially a control interface, it is better suited for the exploration and analysis of a data set than for the publishing or communication of features found in that data set. Thus direct manipulation is most relevant to the analysis of complex data that fills a volume of three-dimensional space, such as a fluid flow data set. Direct manipulation allows the intuitive exploration of that data, which facilitates the discovery of data features that would be difficult to find using more conventional visualization methods. Using a direct manipulation interface in virtual reality, an investigator can, for example, move a data probe about in space, watching the results and getting a sense of how the data varies within its spatial volume.

Bryson, Steve↗

Gullies on Mars and Constraints Imposed by Mars Global Surveyor Data

The discovery of geologically recent gully features on Mars has spawned a wide variety of proposed theories of their origin including water versus carbon dioxide based erosion and shallow versus deep fluid sources. To test the validity of such gully formation mechanisms, data from the Mars Global Surveyor spacecraft has been analyzed to uncover trends in the dimensional and physical properties of the gullies and their surrounding terrain. Over 100 Mars Orbiter Camera (MOC) images containing clear evidence of gully landforms, distributed in the southern mid and high latitudes, have been analyzed in combination with Mars Orbiter Laser Altimeter (MOLA) and Thermal Emission Spectrometer (TES) data to provide quantitative measurements of numerous gully characteristics. Parameters measured include apparent source depth and distribution, vertical and horizontal dimensions, slopes, compass orientations, and factors controlling present-day climatic conditions.

Heldmann, J. L.↗

Properties of cirrus from multispectral AVHRR imagery data

The discovery that the 11 and 12 microns window channels of AVHRR could be used to detect and even characterize the properties of cirrus stimulated the present study which reexamines the general multispectral approach for retrieving cirrus cloud top temperature and emissivity. The generalized multispectral approach described compliments the CO2 slicing method used by Wylie and the bispectral threshold methods used by Minnis et al. While the results shown were for 11 and 12 micron radiances, better definition of the cloud top temperature is probably obtainable using 3.7 micron radiances in combination with the 11 and 12 micron radiances. During the day reflection of solar radiation at 3.7 micron by low level water clouds makes the analysis untenable. At night, at least with the NOAA-9 AVHRR, instrument noise in the 3.7 micron channel also makes the analysis untenable. The identification of semitransparent systems using 3.7 micron radiances has been noted elsewhere.

Coakley, James A., Jr.↗

Data Science and the Knowledge Discovery Adventure

This talk will cover the important steps involved in the data science and knowledge discovery process: • Initial fact gathering (interview domain experts, review reports, articles, state-of-the-art) • Identify the problem (prediction, classification, statistical analysis, etc.) • Survey supporting data sources • Understand the data (numerical, categorical, text, sampling rate, data quality issues, etc.) • Selecting relevant features and sources • Acquire the data (set up agreements with the data stewards, APIs to download, etc.) • Merge data sources (temporal, spatial, common key, other ontologies...) • Feature Engineering (non linear domain knowledge or physics-based relationships) • Build data processing pipeline (may need to tap into data stream, develop parallel processing algorithm, federated learning etc.) • Build model and test (tune hyper-parameters, cross validation.) • Analyze/Validate results (do the results make sense. Does it answer the original question). • Deploy/Publish (Monitor and assess benefits)

Data science↗

RHSEG and Subdue: Background and Preliminary Approach for Combining these Technologies for Enhanced Image Data Analysis, Mining and Knowledge Discovery

Under a project recently selected for funding by NASA's Science Mission Directorate under the Applied Information Systems Research (AISR) program, Tilton and Cook will design and implement the integration of the Subdue graph based knowledge discovery system, developed at the University of Texas Arlington and Washington State University, with image segmentation hierarchies produced by the RHSEG software, developed at NASA GSFC, and perform pilot demonstration studies of data analysis, mining and knowledge discovery on NASA data. Subdue represents a method for discovering substructures in structural databases. Subdue is devised for general-purpose automated discovery, concept learning, and hierarchical clustering, with or without domain knowledge. Subdue was developed by Cook and her colleague, Lawrence B. Holder. For Subdue to be effective in finding patterns in imagery data, the data must be abstracted up from the pixel domain. An appropriate abstraction of imagery data is a segmentation hierarchy: a set of several segmentations of the same image at different levels of detail in which the segmentations at coarser levels of detail can be produced from simple merges of regions at finer levels of detail. The RHSEG program, a recursive approximation to a Hierarchical Segmentation approach (HSEG), can produce segmentation hierarchies quickly and effectively for a wide variety of images. RHSEG and HSEG were developed at NASA GSFC by Tilton. In this presentation we provide background on the RHSEG and Subdue technologies and present a preliminary analysis on how RHSEG and Subdue may be combined to enhance image data analysis, mining and knowledge discovery.

Tilton, James C.↗

Systems Development, Data Mining, and Knowledge Discovery

The primary role of the Technical Integration Office is to provide technical solutions and services to different branches at KSC (Kennedy Space Center) and NASA program customers. The Technical Integration Office helps support KSC's operational needs by providing services such as digital connectivity, data center services, modelling and simulation tools, and communication video services. To learn the necessary technology and processes for my internship, I am working on two projects: learning C# (C Sharp programming language) with SQL and developing requirements for a PX (Communication and Public Engagement) inventory management system. To learn how to efficiently program with C#, my mentor assigned me to complete a sports informatics application that would let users discover facts and rules about various sports. The sports informatics application comes with search capabilities, report generating features, rule lists that users can modify, and diagrams for various sport strategies. To further build upon this project, I also developed a sport simulation game with the application. Once I begin more SQL-based projects, I will have the opportunity to learn how to manage databases and link SQL servers with C# programs. To develop requirements for the inventory management system, I have met with PX representatives and toured their storage facilities to see how they organize and store their items and equipment. I will also be meeting with representatives from the budget office to find out what information must be in a system budget report. The main components the system must have are customer request management, a search feature for items and equipment, report generation capabilities, and automated system warnings when item quantities reach or go below administrator-specified threshold levels. I have drafted questions and shall statements that will ultimately become part of the inventory management system requirements document.

Espinosa, Gabriel↗

The Science Discovery Engine: Connecting Heterogeneous Scientific Data and Information

Transformative science often occurs at the boundaries of different disciplines. Making interdisciplinary science data, software and documentation discoverable and accessible is essential to enabling transformative science. However, connecting this diverse and heterogeneous information is often a challenge due to several factors including the dispersed and sometimes isolated nature of data and the semantic differences between topical areas. NASA’s Science Discovery Engine (SDE) has developed several approaches to tackling these challenges. The SDE is a unified, insightful search experience that enables discovery of NASA’s open science data across five topical areas: astrophysics, biological and physical sciences, Earth science, heliophysics and planetary science. In this presentation, we will discuss our efforts to develop a systematic scientific curation workflow to integrate diverse content into a single search environment. We will also share lessons learned from our work to create a metadata crosswalk across the five disciplines.

Kaylin Bugbee↗

Integrating Multi-agency Data Products in a Cloud-based Platform for Streamlined Discovery, Visualization, and Use

Earth science data users almost always have an interest in utilizing geospatial data from multiple agencies. As computing capability and cloud-based infrastructures accelerate the pace at which scientific research can be done, there is a growing need to enable search, discovery, and use of multi-agency geospatial observations relevant for a common use case - without undergoing the search and discovery process in a less efficient, disparate path with each agency. NASA’s Earth Observing System Data and Information System (EOSDIS) and NOAA’s National Environmental Satellite, Data and Information Service (NESDIS) both support a wide range of Earth science disciplines’ research, operations, and applications activities. Presently, however, there are few examples of data discovery frameworks supporting an inquiry of both NASA’s and NOAA’s extensive archives of Earth observations that are equally suitable for a particular science scenario, regardless of the agency that “owns” the data. NASA and NOAA are collaborating on a data expedition platform for exploring fire weather using data products from both agencies. Users will be able to search, discover, and visualize NASA and NOAA products in one interface. Each agency will curate metadata for its respective datasets, providing for a rich search experience. The collaboration will pilot a shared search interface into these metadata datastores. Data products will be stored in the cloud in cloud-optimized format(s). These formats will allow for optimized data access and visualization to support the “data expedition”. Avenues for further development and application of this cloud-based, multi-agency data provisioning platform will also be discussed.

cloud-based technology↗