Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Metadata and Buckets in the Smart Object, Dumb Archive (SODA) Model

We present the Smart Object, Dumb Archive (SODA) model for digital libraries (DLs), and discuss the role of metadata in SODA. The premise of the SODA model is to "push down" many of the functionalities generally associated with archives into the data objects themselves. Thus the data objects become "smarter", and the archives "dumber". In the SODA model, archives become primarily set managers, and the objects themselves negotiate and handle presentation, enforce terms and conditions, and perform data content management. Buckets are our implementation of smart objects, and da is our reference implementation for dumb archives. We also present our approach to metadata translation for buckets.

Nelson, Michael L.

SPASE, Metadata, and the Heliophysics Virtual Observatories

To provide data search and access capability in the field of Heliophysics (the study of the Sun and its effects on the Solar System, especially the Earth) a number of Virtual Observatories (VO) have been established both via direct funding from the U.S. National Aeronautics and Space Administration (NASA) and through other funding agencies in the U.S. and worldwide. At least 15 systems can be labeled as Virtual Observatories in the Heliophysics community, 9 of them funded by NASA. The problem is that different metadata and data search approaches are used by these VO's and a search for data relevant to a particular research question can involve consulting with multiple VO's - needing to learn a different approach for finding and acquiring data for each. The Space Physics Archive Search and Extract (SPASE) project is intended to provide a common data model for Heliophysics data and therefore a common set of metadata for searches of the VO's. The SPASE Data Model has been developed through the common efforts of the Heliophysics Data and Model Consortium (HDMC) representatives over a number of years. We currently have released Version 2.1 of the Data Model. The advantages and disadvantages of the Data Model will be discussed along with the plans for the future. Recent changes requested by new members of the SPASE community indicate some of the directions for further development.

Thieman, James

Benchmarking CME Arrival Time and Impact: Progress on Metadata, Metrics, and Events

Accurate forecasting of the arrival time and subsequent geomagnetic impacts of coronal mass ejections (CMEs) at Earth is an important objective for space weather forecasting agencies. Recently, the CME Arrival and Impact working team has made significant progress toward defining communit yagreed metrics and validation methods to assess the current state of CME modeling capabilities. This will allow the community to quantify our current capabilities and track progress in models over time. First, it is crucial that the community focuses on the collection of the necessary metadata for transparency and reproducibility of results. Concerning CME arrival and impact we have identified six different metadata types: 3D CME measurement, model description, model input, CME (non)arrival observation, model output data, and metrics and validation methods. Second, the working team has also identified a validation time period, where all events within the following two periods will be considered: 1 January 2011 to 31 December 2012 and January 2015 to 31 December 2015. Those two periods amount to a total of about 100 hit events at Earth and a large amount of misses. Considering a time period will remove any bias in selecting events and the event set will represent a sample set that will not be biased by user selection. Lastly, we have defined the basic metrics and skill scores that the CME Arrival and Impact working team will focus on.

Verbeke, C.

Big Earth Data Initiative: Metadata Improvement: Case Studies

Big Earth Data Initiative (BEDI) The Big Earth Data Initiative (BEDI) invests in standardizing and optimizing the collection, management and delivery of U.S. Government's civil Earth observation data to improve discovery, access use, and understanding of Earth observations by the broader user community. Complete and consistent standard metadata helps address all three goals.

BEDI

ES2Vec: Earth Science Metadata Suggestions and Analogical Reasoning

As the volume of text-based Earth science research grows, it is increasingly possible to discover latent relationships in the literature. However, traditional methodologies are restricted by limited computational capabilities and intractable problem spaces. Advancements in natural language processing (NLP) have allowed us to use an extensive Earth science corpus to create a domain-specific word vector model, Es2Vec, which we have used to surface latent relationships between Earth science concepts and generate improved keyword tags. Earth science metadata keyword assignment is a challenging problem. Dataset curators select appropriate keywords from the Global Change Master Directory (GCMD) set of keywords. The keywords an are integral part of the search and discovery of these datasets. Hence, the selection of keywords is crucial to increasing the discoverability of datasets. Utilizing machine learning techniques, we provide users with automated keyword suggestions to complement manual selection. We trained a machine learning model that leverages the semantic embedding ability of Word2Vec models to process abstracts and suggest relevant keywords. A user interface tool we built to assist data curators in the assignment of such keywords is also described.

word vectors

Development of an Improved Spatial Metadata Simplification Algorithm

The National Aeronautics and Space Administration's (NASA) Atmospheric Science Data Center (ASDC) at NASA Langley Research Center in Hampton, VA provides atmospheric science data products and services to the science community, including enhanced search and subsetting capabilities for numerous Earth Science datasets. The ASDC is the official Distributed Active Archive Center (DAAC) of record for the Tropospheric Emissions: Monitoring of Pollution (TEMPO) instrument. TEMPO is situated on a geostationary satellite positioned at a longitude near the center of the conterminous United States and focused on North America, making hourly swaths of its field of regard from east to west. Spatial metadata is an essential component for the discovery and distribution of Earth Science data. The simplified polygonal boundaries representing the archived data files ensure that any granule can be identified quickly and accurately by a geospatial query. Historically the Douglas-Peucker algorithm has been used for polygon simplification; however, due to the nature of the algorithm, a buffer must be added to the polygon before simplification to ensure pivotal points are not removed by the algorithm. This adds in additional error to the polygon simplification. ASDC’s goal is to test other methods of polyline simplification, such as Visvalingan-Whyatt and Opheim simplification alongside of Douglas-Peucker and different buffering methods, to produce less error during polygon simplification of TEMPO data swaths, and special spatial query geometries such as EPA non-attainment regions, and geopolitical boundaries.

Spatial Metadata

Use of Spatial Metadata Simplification for TEMPO

The National Aeronautics and Space Administration's (NASA) Atmospheric Science Data Center (ASDC) at NASA Langley Research Center in Hampton, VA provides atmospheric science data products and services to the science community, including enhanced search and subsetting capabilities for numerous Earth Science datasets. The ASDC is the official Distributed Active Archive Center (DAAC) of record for the Tropospheric Emissions: Monitoring of Pollution (TEMPO) instrument. TEMPO is situated on a geostationary satellite positioned at a longitude near the center of the conterminous United States and focused on North America, making hourly swaths of its field of regard from east to west. Spatial metadata is an essential component for the discovery and distribution of Earth Science data. The simplified polygonal boundaries representing the archived data files ensure that any granule can be identified quickly and accurately by a geospatial query. Historically the Douglas-Peucker algorithm has been used for polygon simplification; however, due to the nature of the algorithm, a buffer must be added to the polygon before simplification to ensure pivotal points are not removed by the algorithm. This adds in additional error to the polygon simplification. ASDC’s goal is to test other methods of polyline simplification, such as Visvalingan-Whyatt and Opheim simplification alongside of Douglas-Peucker and different buffering methods, to produce less error during polygon simplification of TEMPO data swaths, and special spatial query geometries such as EPA non-attainment regions, and geopolitical boundaries.

Spatial Metadata

Avian Activity Classification Using Recurrent Networks to Fuse Videos with Metadata on Imbalanced Datasets

Activity classification plays a crucial role in various real-life scenarios involving both humans and animals. There is an increasing need for precise activity classification focused on avian-solar interactions, as the usage of solar energy facilities, such as photovoltaic array power stations, has been observed to impact bird species richness, behavior, and activity. However, there has been no work to develop an automated system to monitor and classify these avian-solar interactions. All current methods rely on human observers, which is time and human resources costly and subject to errors related to searcher efficiency. With the recent success of Deep Learning models in activity classification problems, this paper develops a recurrent neural network-based model to automatically classify six avian activities around solar energy facilities. Our proposed model integrates critical feature engineering metadata with video frame data, enabling improved learning and more accurate activity classification. Furthermore, we address the challenge of data imbalance during training and demonstrate the efficacy of our model in detecting and classifying different activities within video tracks. Additionally, we analyze the saliency/backpropagation map of the trained proposed model and validate its decision-making rationale.

Avian activity classification; bidirectional LSTM;

Microbial community data from throughfall exclusion experiment: Metadata, SI, community composition, LefSe, and FunGuilR data tables from PARCHED Panama tropical forest soils, 2024-2025

Soil contains more carbon (C) than terrestrial vegetation and the atmosphere combined, with some of the largest terrestrial C stocks in tropical rainforests. Soil microbes decompose organic matter, playing a vital role in the storage or loss of soil C. With climate change, drought conditions are predicted to increase in many tropical regions, including both chronic drying and extended drought, potentially influencing these processes. This project explored the effects of chronic and seasonal drying on soil microbial communities across four distinct tropical forests in a long-term drying experiment. We investigated the effects of a chronic drying manipulation on soil microbial community abundance and variation across different forests and seasons. We also compared findings with previously published data from these forests after short-term drying. This project used soils from a long-term drying experiment established in 2018 across four seasonal lowland forests in Panama. Soils were collected from 0 – 10 cm depths during three seasonal periods in control and drying plots in 2024 and 2025 from a total of 32 plots (n = 4 per forest per treatment). The forests varied in baseline rainfall and soil fertility. We calculated alpha and beta diversity indices and compared taxonomic community composition. We found significant biogeographic variation in microbial diversity and taxonomy, with significant differences across the forests and significant effects of the drying treatment. Metadata and sample IDs are within Metadata_16S.csv and Metadata_ITS.csv. Relative abundance tables of every sample at every season are shown in the Excel workbooks 16S Relative Abundance.xlsx and ITS Relative Abundance.xlsx. They are then also shown in CSV files by each taxonomic level. Linear discriminant analysis effect size (LefSe) tables are shown for the full 16S and ITS datasets (n = 96), subsets for every site at every season (n = 8), and then for the forests with each plot merged by season (n = 8). FunGuildR data table of ITS data is uploaded.

Bacteria

Data Cards for Standardized Metadata Across DOE-Aligned Data Initiatives: Toward Transparent, Interoperable, and Governed Dataset Documentation

As data-intensive research, advanced computing, and artificial intelligence become increasingly central to scientific and operational workflows, the need for consistent, transparent, and machine-actionable documentation has grown correspondingly. Multiple DOE-aligned communities—including Office of Science, Genesis Mission, American Science Cloud (AmSC), National Nuclear Security Administration (NNSA) stewardship and governance, and related cross-laboratory collaborations—have independently developed metadata practices to support discovery, access, reuse, repository deposit, and compliance.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

DMTN-227: The Consolidated Database of Image Metadata

This document proposes a specification for what the content of an image metadata database should be, how it should be managed, how it could be implemented, and how it might be extended. A phased strategy for bringing it to production is also proposed.

79 ASTRONOMY AND ASTROPHYSICS

Logic programming and metadata specifications

Artificial intelligence (AI) ideas and techniques are critical to the development of intelligent information systems that will be used to collect, manipulate, and retrieve the vast amounts of space data produced by 'Missions to Planet Earth.' Natural language processing, inference, and expert systems are at the core of this space application of AI. This paper presents logic programming as an AI tool that can support inference (the ability to draw conclusions from a set of complicated and interrelated facts). It reports on the use of logic programming in the study of metadata specifications for a small problem domain of airborne sensors, and the dataset characteristics and pointers that are needed for data access.

Lopez, Antonio M., Jr.

The VIS-AD data model: Integrating metadata and polymorphic display with a scientific programming language

The VIS-AD data model integrates metadata about the precision of values, including missing data indicators and the way that arrays sample continuous functions, with the data objects of a scientific programming language. The data objects of this data model form a lattice, ordered by the precision with which they approximate mathematical objects. We define a similar lattice of displays and study visualization processes as functions from data lattices to display lattices. Such functions can be applied to visualize data objects of all data types and are thus polymorphic.

Hibbard, William L.

Textural-Contextual Labeling and Metadata Generation for Remote Sensing Applications

Despite the extensive research and the advent of several new information technologies in the last three decades, machine labeling of ground categories using remotely sensed data has not become a routine process. Considerable amount of human intervention is needed to achieve a level of acceptable labeling accuracy. A number of fundamental reasons may explain why machine labeling has not become automatic. In addition, there may be shortcomings in the methodology for labeling ground categories. The spatial information of a pixel, whether textural or contextual, relates a pixel to its surroundings. This information should be utilized to improve the performance of machine labeling of ground categories. Landsat-4 Thematic Mapper (TM) data taken in July 1982 over an area in the vicinity of Washington, D.C. are used in this study. On-line texture extraction by neural networks may not be the most efficient way to incorporate textural information into the labeling process. Texture features are pre-computed from cooccurrence matrices and then combined with a pixel's spectral and contextual information as the input to a neural network. The improvement in labeling accuracy with spatial information included is significant. The prospect of automatic generation of metadata consisting of ground categories, textural and contextual information is discussed.

Kiang, Richard K.

Data Archival and Retrieval Enhancement (DARE) Metadata Modeling and Its User Interface

The Defense Nuclear Agency (DNA) has acquired terabytes of valuable data which need to be archived and effectively distributed to the entire nuclear weapons effects community and others...This paper describes the DARE (Data Archival and Retrieval Enhancement) metadata model and explains how it is used as a source for generating HyperText Markup Language (HTML)or Standard Generalized Markup Language (SGML) documents for access through web browsers such as Netscape.

The Defense Nuclear Agency DNA DARE Data Archival

Availability of Previously Unprocessed ALSEP Raw Instrument Data, Derivative Data, and Metadata Products

In year 2010, 440 original data archival tapes for the Apollo Lunar Science Experiment Package (ALSEP) experiments were found at the Washington National Records Center. These tapes hold raw instrument data received from the Moon for all the ALSEP instruments for the period of April through June 1975. We have recently completed extraction of binary files from these tapes, and we have delivered them to the NASA Space Science Data Cordinated Archive (NSSDCA). We are currently processing the raw data into higher order data products in file formats more readily usable by contemporary researchers. These data products will fill a number of gaps in the current ALSEP data collection at NSSDCA. In addition, we have estabilished a digital, searcheable archive of ALSEP document and metadata as part of the web portal of the Lunar and Planetary Institute. It currently holds approx. 700 documents totaling approx. 40,000 pages

ALSEP

Storage of Physical Sample Metadata in the Astrobiology Habitable Environments Database (AHED)

The National Aeronautics and Space Administration has begun an effort to store, curate, and publish information about physical samples collected and analyzed in conjunction with NASA-funded astrobiology research. Astrobiology is a multidisciplinary area of scientific research being conducted by collaborating teams of biologists, chemists, geologists, atmospheric scientists, oceanographers, astrophysicists, astronomers, and other specialists. Astrobiology studies the origin, evolution, and distribution of life in the Universe. NASA uses the results of astrobiology research to focus its future missions on targets of opportunity for the discovery of life off Earth. Astrobiology researchers conduct both field-based and laboratory-based research, during which physical samples are collected, processed, and catalogued. The cataloguing practices employed by different teams of astrobiologists vary widely, and there are no specific standards available to guide the collection and recording of astrobiology sample data. The disparity in data collection approaches and the lack of a centralized sample repository makes it difficult for astrobiology teams to share data and benefit from resultant synergies.To facilitate data sharing within the astrobiology community, NASA is developing a prototype database the Astrobiology Habitable Environments Database (AHED) and an associated set of data collection templates. The database will store information about samples, along with associated measurements and analyses, including information about biological cultures enriched or isolated from samples, and the results of analyses performed on the samples (e.g., via spectrography, microscopy, etc.). In addition, the system will store contextual information about field sites where samples were collected, the instruments or equipment used for analysis, and people and institutions involved in their collection. AHED is being implemented on top of Open Data Repository's Data Publisher [1], an open source software platform for the publication of scientific datasets. The data collection templates under development represent an initial attempt to propose a set of metadata for capture and storage within AHED. The design of these templates is being conducted by a consolidated group of astrobiologists from active research teams at NASA Ames Research Center, assisted by data science and software engineering specialists. These initial templates must be vetted with the broader astrobiology community through a defined process to ensure that they meet community needs. Each template captures a different type of data collection record. For each template, we are developing a list of fields to be captured, including a set of required entry fields, a set of recommended but optional fields, and a set of discretionary fields. A datatype selected from a variety of text and numeric types is specified for each field. Included is a 'choice' type that restricts user input to an enumerated list of values. Many of the fields and field values capture information of particular interest to the astrobiology community, and are intended to facilitate search and retrieval of relevant data across multiple datasets.

Keller, Rich