Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data discovery”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Qualitative Future Safety Risk Identification an Update

The purpose of this report is to document the results of a high-level qualitative study that was conducted to identify future aviation safety risks and to assess the potential impacts to the National Airspace System (NAS) of NASA Aviation Safety research on these risks. Multiple external sources (for example, the National Transportation Safety Board, the Flight Safety Foundation, the National Research Council, and the Joint Planning and Development Office) were used to develop a compilation of future safety issues risks, also referred to as future tall poles. The primary criterion used to identify the most critical future safety risk issues was that the issue must be cited in several of these sources as a safety area of concern. The tall poles in future safety risk, in no particular order of importance, are as follows: Runway Safety, Loss of Control In Flight, Icing Ice Detection, Loss of Separation, Near Midair Collision Human Fatigue, Increasing Complexity and Reliance on Automation, Vulnerability Discovery, Data Sharing and Dissemination, and Enhanced Survivability in the Event of an Accident.

future aviation safety risk

Carbon Cycle Science Data and Services at the Goddard Earth Sciences Data Information and Services Center (GES DISC)

The Goddard Earth Sciences Data Information and Services Center (GES DISC) archives and distributes a number of observational and model carbon cycle science data sets. We also provide services that facilitate data discovery, intercomparison, and visualization of these heterogeneous datasets for both research and applications users, such as subsetting, format conversion, How-To documentation, and the Help Desk.

Hearty, Thomas

EOSDIS STAC Briefing

The Spatio-Temporal Asset Catalog (STAC) specification provides a common language to describe a range of geospatial information, so it can more easily be indexed and discovered. A 'spatiotemporal asset' is any file that represents information about the earth captured in a certain space and time. STAC provides a standards-based, web and cloud friendly cataloging specification that has seen a large amount of adoption in the cloud geospatial information system space. This presentation will demonstrate how NASA EOSDIS leverages STAC technology to provide value-added features to both our data discovery and transformation services. Finally, we will propose a way to improve our federated discovery capabilities using STAC.

NASA EOSDIS STAC Cloud catalog metadata

NASA GeneLab: Open Science for Life in Space

The NASA GeneLab project capitalizes on multi-omic technologies to maximize the return on spaceflight experiments. To do this, GeneLab maintains a publicly accessible database (GLDS) that houses spaceflight and spaceflight relevant multi-omics data and collaborates with NASA principal investigators and projects to generate additional omics data. GeneLab houses more than 350 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, animal and microbial experiments, with a growing number of these having been produced by the GeneLab Sequencing Lab. The GLDS contains rich metadata about each experiment and has integrated radiation dosimetry data from experiments flown on the Space Shuttle, International Space Station, and Free Flying spacecrafts. With the increasing amount and complexity of omics data being generated, GeneLab utilizes community-defined, common models for metadata and terminology so that omics data and results are discoverable and reliably reproducible. GeneLab uses the ISA-Tab specification and semantic model for organizing and representing omics metadata. In addition to metadata standards, data files must be open-source file or common exchange formats to ensure accessibility and usability by all users. To ease data ingestion and transfer, the web-based submission tool allows PIs a user-friendly user interface to curate, organize, and publish their space relevant omics data. In the more recent years, data curation and submission portal has incorporated the FAIR principles making data findable, accessible, interoperable, and reusable. To increase reusability of data, GeneLab has implemented an effort to present processed data in the GLDS in addition to the raw omics data. The processed data will enable interpretation of the data by a larger group of students, scientists and the general public. Standard pipelines for the transformation of raw data into visualizations were developed by four GeneLab Analysis Working Groups (animals, plants, microbes, multi-omics) comprised of over 200 scientists from NASA, industry, and academia. To explore the data, the GLDS provides users various tools for data analysis, collaborative workspace for file storage and sharing, and a visualization portal. The analysis platform built using the Galaxy toolshed provides access to a broad variety of users including those with limited bioinformatics experience and students to learn how to analyze spaceflight omics data. The visualization portal takes GeneLab one step closer to data democratization by removing all bioinformatics requisites to interpret transcriptomics data hosted in the repository. To train the next generation of scientists, NASA offers training programs such as GeneLab 4 High School (GL4HS) and GeneLab 4 Universities. NLM Curation at a Scale Workshop 2022 | NASA GeneLab (GL4U) to teach students bioinformatics and computational biology methods to analyze omics data. Discoveries made using GeneLab have begun and will continue to deepen our understanding of biology, advance the field of genomics, and help to discover cures for diseases, create better diagnostic tools, and ultimately allow astronauts to better withstand the rigors of long-duration spaceflight.

GeneLab

Improving GES Disc Data Search and Discovery Through AI Metadata Augmentation

NASA’s Goddard Earth Science (GES) Data and Information Services Center (DISC) is one of twelve data centers in NASA's Science Mission Directorate (SMD), providing vital earth science data to a diverse user base. To enhance the discoverability of this data, GES DISC employs a keyword search system, which leverages scientific keywords embedded in dataset metadata. However, the evolving nature of scientific applications of our data necessitates regular review and augmentation of these keywords. To address this, we developed a service to automatically predict missing science keywords in the metadata. This service constructs a knowledge graph from the latest GES DISC metadata within NASA’s Common Metadata Repository (CMR). Using an open-source library, we trained a machine learning model to predict absent science keywords in the metadata. Our preliminary results indicate that the model has high levels of accuracy at predicting science keywords in the dataset metadata when exposed to data not included in its training. These predicted keywords were then evaluated by GES DISC data curation scientists and compared against other AI tools for metadata augmentation. We aim to enhance the overall usability and accessibility of NASA’s earth science data by implementing this tool in our data curation processes.

Kendall Gilbert

Improve Data Mining and Knowledge Discovery Through the Use of MatLab

Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(R) (MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.

Shaykhian, Gholam Ali

Improve Data Mining and Knowledge Discovery through the use of MatLab

Data mining is widely used to mine business, engineering, and scientific data. Data mining uses pattern based queries, searches, or other analyses of one or more electronic databases/datasets in order to discover or locate a predictive pattern or anomaly indicative of system failure, criminal or terrorist activity, etc. There are various algorithms, techniques and methods used to mine data; including neural networks, genetic algorithms, decision trees, nearest neighbor method, rule induction association analysis, slice and dice, segmentation, and clustering. These algorithms, techniques and methods used to detect patterns in a dataset, have been used in the development of numerous open source and commercially available products and technology for data mining. Data mining is best realized when latent information in a large quantity of data stored is discovered. No one technique solves all data mining problems; challenges are to select algorithms or methods appropriate to strengthen data/text mining and trending within given datasets. In recent years, throughout industry, academia and government agencies, thousands of data systems have been designed and tailored to serve specific engineering and business needs. Many of these systems use databases with relational algebra and structured query language to categorize and retrieve data. In these systems, data analyses are limited and require prior explicit knowledge of metadata and database relations; lacking exploratory data mining and discoveries of latent information. This presentation introduces MatLab(TradeMark)(MATrix LABoratory), an engineering and scientific data analyses tool to perform data mining. MatLab was originally intended to perform purely numerical calculations (a glorified calculator). Now, in addition to having hundreds of mathematical functions, it is a programming language with hundreds built in standard functions and numerous available toolboxes. MatLab's ease of data processing, visualization and its enormous availability of built in functionalities and toolboxes make it suitable to perform numerical computations and simulations as well as a data mining tool. Engineers and scientists can take advantage of the readily available functions/toolboxes to gain wider insight in their perspective data mining experiments.

Shaykahian, Gholan Ali

Vostok Subglacial Lake: A Review of Geophysical Data Regarding Its Discovery and Topographic Setting

Vostok Subglacial Lake is the largest and best known sub-ice lake in Antarctica. The establishment of its water depth (>500 m) led to an appreciation that such environments may be habitats for life and could contain ancient records of ice sheet change, which catalyzed plans for exploration and research. Here we discuss geophysical data used to identify the lake and the likely physical, chemical, and biological processes that occur in it. The lake is more than 250 km long and around 80 km wide in one place. It lies beneath 4.2 to 3.7 km of ice and exists because background levels of geothermal heating are sufficient to warm the ice base to the pressure melting value. Seismic and gravity measurements show the lake has two distinct basins. The Vostok ice core extracted >200 m of ice accreted from the lake to the ice sheet base. Analysis of this ice has given valuable insights into the lake s biological and chemical setting. The inclination of the ice-water interface leads to differential basal melting in the north versus freezing in the south, which excites circulation and potential mixing of the water. The exact nature of circulation depends on hydrochemical properties, which are not known at this stage. The age of the subglacial lake is likely to be as old as the ice sheet (approx.14 Ma). The age of the water within the lake will be related to the age of the ice melting into it and the level of mixing. Rough estimates put that combined age as approx.1 Ma.

Siegert, Martin J.

Enhancing Discovery, Search, and Access of NASA Hydrological Data by Leveraging GEOSS

An ongoing NASA-funded project has removed a longstanding barrier to accessing NASA data (i.e., accessing archived time-step array data as point-time series) for selected variables of the North American and Global Land Data Assimilation Systems (NLDAS and GLDAS, respectively) and other EOSDIS (Earth Observing System Data Information System) data sets (e.g., precipitation, soil moisture). These time series (data rods) are pre-generated. Data rods Web services are accessible through the CUAHSI Hydrologic Information System (HIS) and the Goddard Earth Sciences Data and Information Services Center (GES DISC) but are not easily discoverable by users of other non-NASA data systems. The Global Earth Observation System of Systems (GEOSS) is a logical mechanism for providing access to the data rods. An ongoing GEOSS Water Services project aims to develop a distributed, global registry of water data, map, and modeling services cataloged using the standards and procedures of the Open Geospatial Consortium and the World Meteorological Organization. The ongoing data rods project has demonstrated the feasibility of leveraging the GEOSS infrastructure to help provide access to time series of model grid information or grids of information over a geographical domain for a particular time interval. A recently-begun, related NASA-funded ACCESS-GEOSS project expands on these prior efforts. Current work is focused on both improving the performance of the generation of on-the-fly (OTF) data rods and the Web interfaces from which users can easily discover, search, and access NASA data.

GEOSS

Global Change Master Directory (GCMD) Keyword Management Process and Lifecycle

The Global Change Master Directory (GCMD) keywords are a hierarchical set of controlled vocabulary covering the Earth science disciplines that have been evolving for over 25 years. The process for how these keywords have been curated and reviewed has also evolved. This presentation will convey the process for reviewing and approving the GCMD keywords, including fast track and yearly reviews through the ESDIS Standards Office, and how Earth science users can influence keyword additions and modifications. The presentation will also highlight how keywords facilitate the discovery of EOSDIS data and services and how other organizations are using the GCMD keywords.

EOSDIS

Previously Unrecognized Large Lunar Impact Basins Revealed by Topographic Data

The discovery of a large population of apparently buried impact craters on Mars, revealed as Quasi- Circular Depressions (QCDs) in Mars Orbiting Laser Altimeter (MOLA) data [1,2,3] and as Circular Thin Areas (CTAs) [4] in crustal thickness model data [5] leads to the obvious question: are there unrecognized impact features on the Moon and other bodies in the solar system? Early analysis of Clementine topography revealed several large impact basins not previously known [6,7], so the answer certainly is "Yes." How large a population of previously undetected impact basins, their size frequency distribution, and how much these added craters and basins will change ideas about the early cratering history and Late Heavy Bombardment on the Moon remains to be determined. Lunar Orbiter Laser Altimeter (LOLA) data [8] will be able to address these issues. As a prelude, we searched the state-of-the-art global topographic grid for the Moon, the Unified Lunar Control Net (ULCN) [9] for evidence of large impact features not previously recognized by photogeologic mapping, as summarized by Wilhelms [lo].

Frey, Herbert V.

Direct Manipulation in Virtual Reality

Virtual Reality interfaces offer several advantages for scientific visualization such as the ability to perceive three-dimensional data structures in a natural way. The focus of this chapter is direct manipulation, the ability for a user in virtual reality to control objects in the virtual environment in a direct and natural way, much as objects are manipulated in the real world. Direct manipulation provides many advantages for the exploration of complex, multi-dimensional data sets, by allowing the investigator the ability to intuitively explore the data environment. Because direct manipulation is essentially a control interface, it is better suited for the exploration and analysis of a data set than for the publishing or communication of features found in that data set. Thus direct manipulation is most relevant to the analysis of complex data that fills a volume of three-dimensional space, such as a fluid flow data set. Direct manipulation allows the intuitive exploration of that data, which facilitates the discovery of data features that would be difficult to find using more conventional visualization methods. Using a direct manipulation interface in virtual reality, an investigator can, for example, move a data probe about in space, watching the results and getting a sense of how the data varies within its spatial volume.

Bryson, Steve

Gullies on Mars and Constraints Imposed by Mars Global Surveyor Data

The discovery of geologically recent gully features on Mars has spawned a wide variety of proposed theories of their origin including water versus carbon dioxide based erosion and shallow versus deep fluid sources. To test the validity of such gully formation mechanisms, data from the Mars Global Surveyor spacecraft has been analyzed to uncover trends in the dimensional and physical properties of the gullies and their surrounding terrain. Over 100 Mars Orbiter Camera (MOC) images containing clear evidence of gully landforms, distributed in the southern mid and high latitudes, have been analyzed in combination with Mars Orbiter Laser Altimeter (MOLA) and Thermal Emission Spectrometer (TES) data to provide quantitative measurements of numerous gully characteristics. Parameters measured include apparent source depth and distribution, vertical and horizontal dimensions, slopes, compass orientations, and factors controlling present-day climatic conditions.

Heldmann, J. L.

Properties of cirrus from multispectral AVHRR imagery data

The discovery that the 11 and 12 microns window channels of AVHRR could be used to detect and even characterize the properties of cirrus stimulated the present study which reexamines the general multispectral approach for retrieving cirrus cloud top temperature and emissivity. The generalized multispectral approach described compliments the CO2 slicing method used by Wylie and the bispectral threshold methods used by Minnis et al. While the results shown were for 11 and 12 micron radiances, better definition of the cloud top temperature is probably obtainable using 3.7 micron radiances in combination with the 11 and 12 micron radiances. During the day reflection of solar radiation at 3.7 micron by low level water clouds makes the analysis untenable. At night, at least with the NOAA-9 AVHRR, instrument noise in the 3.7 micron channel also makes the analysis untenable. The identification of semitransparent systems using 3.7 micron radiances has been noted elsewhere.

Coakley, James A., Jr.

Data Science and the Knowledge Discovery Adventure

This talk will cover the important steps involved in the data science and knowledge discovery process: • Initial fact gathering (interview domain experts, review reports, articles, state-of-the-art) • Identify the problem (prediction, classification, statistical analysis, etc.) • Survey supporting data sources • Understand the data (numerical, categorical, text, sampling rate, data quality issues, etc.) • Selecting relevant features and sources • Acquire the data (set up agreements with the data stewards, APIs to download, etc.) • Merge data sources (temporal, spatial, common key, other ontologies...) • Feature Engineering (non linear domain knowledge or physics-based relationships) • Build data processing pipeline (may need to tap into data stream, develop parallel processing algorithm, federated learning etc.) • Build model and test (tune hyper-parameters, cross validation.) • Analyze/Validate results (do the results make sense. Does it answer the original question). • Deploy/Publish (Monitor and assess benefits)

Data science