Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “system metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Beyond microbial abundance: metadata integration enhances disease prediction in human microbiome studies

Multiple studies have highlighted the interaction of the human microbiome with physiological systems such as the gut, immune, liver, and skin, via key axes. Advances in sequencing technologies and high-performance computing have enabled the analysis of large-scale metagenomic data, facilitating the use of machine learning to predict disease likelihood from microbiome profiles. However, challenges such as compositionality, high dimensionality, sparsity, and limited sample sizes have hindered the development of actionable models. One strategy to improve these models is by incorporating key metadata from both the human host and sample collection/processing protocols. This remains challenging due to sparsity and inconsistency in metadata annotation and availability. In this paper, we introduce a machine learning-based pipeline for predicting human disease states by integrating host and protocol metadata with microbiome abundance profiles from 68 different studies, processed through a consistent pipeline. Our findings indicate that metadata can enhance machine learning predictions, particularly at higher taxonomic ranks like Kingdom and Phylum, though this effect diminishes at lower ranks. Our study leverages a large collection of microbiome datasets comprising 11,208 samples, therefore enhancing the robustness and statistical confidence of our findings. This work is a critical step toward utilizing microbiome and metadata for predicting diseases such as gastrointestinal infections, diabetes, cancer, and neurological disorders.

Mathematics and Computing↗

Enabling Space Biological Knowledge Discovery Through Image and Video Data Sharing

Increased biomedical risks and challenges associated with deep space missions and experiments (cis-Lunar, Mars transit/surface) require new knowledge discovery and development of novel ecosystems. Supporting distant and long-duration missions and experiments requires biological data (from yeast, microbes, fruit flies, C. elegans, plants, crops, rodents, humans) be findable, accessible, interoperable, reusable (FAIR), and maximally open-access. As data-intensive, bioinformatic, meta-analytical, and computer-assisted approaches continue to be a centerpiece of modern research, the NASA Biological and Physical Sciences division is expanding its Open Science capabilities beyond NASA GeneLab. The NASA Ames Life Sciences Data Archive (ALSDA) is a repository which is responsible for collecting and access to space biological imagery and video, alongside tabular and environmental data. In this presentation, we will discuss strategies dealing with archiving, curating, and accessibility of images from very distinct imaging modalities (e.g., micro-computed tomography, magnetic resonance imaging, photographic images of plants, fluorescence microscopy, behavioral videos, etc.). There are two main challenges: 1. Open-source data storage and 2. Metadata related to the imagery-video. Both have been solved by leveraging two existing open-source systems. For data storage, ALSDA is utilizing components through the Open Microscopy Environment (OME), which can read most imaging proprietary formats and display on a web interface complex multidimensional images (Z stack, multi-channel, temporal, spectral). Most technical metadata from imaging modalities are captured seamlessly. For metadata capturing experimental details, ALSDA (like GeneLab) uses the ISA-Tab specification which relies on the ISA data model to order and classify metadata. The ISA data model uses a tree structure with three files to capture the metadata: The top layer is the Investigations file, the second layer is the Study file(s), and the last layer is the Assay file(s). We believe such an approach may be useful for other types of image research data from other investigators in the AGU community.

imaging↗

A machine learning method of modern urban building energy modeling: A case study of Chicago

Urban-scale building energy modeling is vital for urban planning. However, it can be challenging to assimilate reliable non-geometry building data for urban-scale modeling without extensive investment. Here, this study introduces a novel approach to developing modern urban-scale building energy stock data using geographic information systems and machine learning algorithms without necessarily requiring pre-supplied non-geometric metadata. The proposed framework integrates building footprint and height data to estimate gross floor areas, and matches each building to a pool of candidate records from ComStock or ResStock—filtered to the same county and ranked by geometric similarity—demonstrate a proof-of-concept case study in Chicago for predicting energy use intensity (EUI) using scalable datasets. The model achieved a mean bias error (MBE) of 0.08 kWh/m² and root mean square error (RMSE) of 14.84 kWh/m² under full metadata input for EUI prediction. With only location inputs, the model captured 69.2 % of EUI within predicted ranges. These results demonstrate the model’s potential to support early-stage urban planning, identify candidates for energy-efficient retrofits. By removing the dependency on detailed pre-surveys or extensive building metadata, the approach overcomes a key barrier in traditional urban-scale building energy modeling, illustrating a pathway toward broader and more cost-effective application, though further multi-city validation and improved treatment of pre-1925 buildings are needed.

Energy Use Intensity↗

Analysis of Conservation Voltage Reduction under Inverter-Based VAR-Support [Slides]

Conservation voltage reduction (CVR) is a common technique used by utilities to strategically reduce demand during peak periods. As penetration levels of distributed generation (DG) continue to rise and advanced inverter capabilities become more common, it is unclear how the effectiveness of CVR will be impacted and how CVR interacts with advanced inverter functions. In this work, we investigated the mutual impacts of CVR and DG from photovoltaic (PV) systems (with and without autonomous Volt-VAR enabled). The analysis was conducted on an actual utility dataset, including a feeder model, measurement data from smart meters and intelligent reclosers, and metadata for more than 30 CVR events triggered by the utility over the year. The installed capacity of the modeled PV systems represented 66% of peak load, but reached instantaneous penetrations reached up to 2.5x the load consumption over the year. While the objectives of CVR and autonomous Volt-VAR are opposed to one another, this study found that their interactions were mostly inconsequential since the CVR events occurred when total PV output was low.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Applying Waveform Correlation and Waveform Template Metadata to Aftershocks in the Middle East to Reduce Analyst Workload

Organizations that monitor for underground nuclear explosive tests are interested in techniques that automatically characterize recurring events such as aftershocks to reduce the human analyst effort required to produce high-quality event bulletins. Waveform correlation is a technique that is effective in finding similar waveforms from repeating seismic events. In this study, we apply waveform correlation in combination with template event metadata to two aftershock sequences in the Middle East to seek corroborating detections from multiple stations in the International Monitoring System of the Preparatory Commission for the Comprehensive Nuclear-Test-Ban Treaty Organization. We use waveform templates from stations that are within regional distance of aftershock sequences to detect subsequent events, then use template event metadata to discover what stations are likely to record corroborating arrival waveforms for recurring aftershock events at the same location, and develop additional waveform templates to seek corroborating detections. We evaluate the results with the goal of determining whether applying the method to aftershock events will improve the choice of waveform correlation detections that lead to bulletin-worthy events and reduction of analyst effort.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Plant Physiology, Alder and Willow Species, Seward Peninsula, Alaska, 2019

Photosynthetic parameters (Vcmax and Jmax), CO2 response (ACi) curves and dark respiration (Rdark) measurements of tundra shrub species. Data were collected at the NGEE-Arctic Kougarok and Teller study sites on the Seward Peninsula, AK, USA in July 2019. Four species were measured using LI-COR LI-6400 portable gas exchange systems: Alnus viridis ssp. fruticosa, Salix richardsonii and Salix pulchra. The data package files include data files, metadata files, dGPS locations and the complete instrument output for all gas exchange measurements in .csv format. See the related data packages for foliar trait data (leaf mass per area, leaf nitrogen concentration), and leaf spectral reflectance data collected simultaneously on the same shrubs (NGA210, NGA212).The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗

CEDS Differential Privacy (CEDSDP) v0.1

A Python package that provides differentially private queries optimized for energy systems' data. It may be used to publish queries such as clustering, averaging, metadata inference, etc. that are useful for a variety of grid-related analytics, including cyberattack detection.

Peisert, Sean↗

Pan-Arctic daily streamflow at 0.5-degree resolution

Data was derived using daily runoff from models participating in the Coupled Model Intercomparison Project, version 6 (CMIP6) that have been routed through the Model for Scale Adaptive River Transport (MOSART) routing model to represent streamflow at individual streamflow gage locations. Data from 750 streamflow gage locations are represented, which are indexed by gage id according to the metadata.csv file. Streamflow gage location sites were based on gage records downloaded from the United States Geological Survey (USGS), National Water Data (HYDAT) Archive from Canada, and the State Hydrological Institute of Russia. Data are available from 11 Earth System Models (ESMs) in total at a daily timestep with a duration of 1920-2099 (1850-2014 for E3SM only). The spatial domain of the modeled data is 0.5 degree resolution. Earth System Models and Units of mean daily streamflow are presented in cubic meters per second (cms). Data files consist of 11 Earth System Model mean daily streamflow output per gage location in .csv format, 1 metadata file in .csv format and 1 Readme file in .docx format (13 files in total). These data were used to benchmark CMIP6 modeled representations of streamflow against gage records.

54 ENVIRONMENTAL SCIENCES↗

DPADL: An Action Language for Data Processing Domains

This paper presents DPADL (Data Processing Action Description Language), a language for describing planning domains that involve data processing. DPADL is a declarative object-oriented language that supports constraints and embedded Java code, object creation and copying, explicit inputs and outputs for actions, and metadata descriptions of existing and desired data. DPADL is supported by the IMAGEbot system, which will provide automation for an ecosystem forecasting system called TOPS.

Golden, Keith↗

A Domain Description Language for Data Processing

We discuss an application of planning to data processing, a planning problem which poses unique challenges for domain description languages. We discuss these challenges and why the current PDDL standard does not meet them. We discuss DPADL (Data Processing Action Description Language), a language for describing planning domains that involve data processing. DPADL is a declarative, object-oriented language that supports constraints and embedded Java code, object creation and copying, explicit inputs and outputs for actions, and metadata descriptions of existing and desired data. DPADL is supported by the IMAGEbot system, which we are using to provide automation for an ecological forecasting application. We compare DPADL to PDDL and discuss changes that could be made to PDDL to make it more suitable for representing planning domains that involve data processing actions.

Golden, Keith↗

Collaborative Data Curation to Support the Multi-Mission Algorithm and Analysis Platform (MAAP)

Upcoming space-borne missions will offer unprecedented data about Earth but will also feature exponentially high data volumes. These high data volumes will change the way the scientific community works with data and will also create a unique need for improved data sharing and collaboration. NASA and ESA are working together to address these issues by collaboratively developing the Multi-Mission Algorithm and Analysis Platform (MAAP) to improve the understanding of global aboveground terrestrial carbon dynamics. The MAAP will support ESA’s BIOMASS mission, NASA’s GEDI mission and NASA/ISRO’s NISAR mission. The MAAP will be developed in two phases: a pilot phase and a full production phase. The pilot phase will demonstrate collaboration and basic capabilities. The pilot phase will focus on biomass relevant airborne and field campaign data. Two NASA teams are supporting the development of the MAAP. The MAAP engineering team is responsible for the development, maintenance and operations of the MAAP system while the MAAP data team ensures the ongoing quality of the data, metadata and other information provided in the MAAP. The MAAP data team also supports the ingest and archive of identified data to the MAAP platform. This poster describes the use case development process for the pilot MAAP and the data curated in support of those use cases. Additionally, this presentation will outline the pilot MAAP data ingest process and metadata curation effort along with efforts to ensure interoperability between ESA and NASA data and metadata.

Bugbee, Kaylin↗

SetGo: Metadata Readiness for Scientific AI Datasets

Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse. The Readiness Engine for Data Integration (REDI) addresses computational readiness, but no corresponding tool evaluates whether a dataset’s metadata are sufficiently complete, governed, and standards-compliant for publication and agent-based consumption. Existing FAIR assessors operate only on published repository records, and no single system covers FAIR compliance, licensing, provenance, governance, reproducibility, and catalog readiness together. We present SetGo, an open-source Python toolkit that assesses and repairs metadata readiness across these six dimensions before a dataset is published or archived. Applied to four scientific corpora, SetGo surfaces deficiencies that general-purpose tools do not detect: ERA5 climate metadata scores 4% on ACDD 1.3 compliance; materials datasets fail OPTIMADE species-definition requirements; and PDB-derived proteomics data carries licensing terms incompatible with standard SPDX identifiers. Guided enrichment raises overall FAIR scores from 52–57% to 81–91%, and a single setgo publish command pushes to Hugging Face Hub, CKAN, or OpenMetadata with ML Commons Croissant 1.0 metadata sidecars. To support interactive and automated workflows, SetGo integrates with coding agents powered by large language models (LLMs) through a /setgo skill that enables natural-language execution of the full assess–enrich–publish loop, with user involvement limited to supplying missing metadata values.

Wilkinson, Sean [ORNL] (ORCID:0000000214437479)↗

System and Method for Providing a Climate Data Persistence Service

A system, method and computer-readable storage devices for providing a climate data persistence service. A system configured to provide the service can include a climate data server that performs data and metadata storage and management functions for climate data objects, a compute-storage platform that provides the resources needed to support a climate data server, provisioning software that allows climate data server instances to be deployed as virtual climate data servers in a cloud computing environment, and a service interface, wherein persistence service capabilities are invoked by software applications running on a client device. The climate data objects can be in various formats, such as International Organization for Standards (ISO) Open Archival Information System (OAIS) Reference Model Submission Information Packages, Archive Information Packages, and Dissemination Information Packages. The climate data server can enable scalable, federated storage, management, discovery, and access, and can be tailored for particular use cases.

Schnase, John L.↗

Intelligent Systems Technologies to Assist in Utilization of Earth Observation Data

With the launch of several Earth observing satellites over the last decade, we are now in a data rich environment. From NASA's Earth Observing System (EOS) satellites alone, we are accumulating more than 3 TB per day of raw data and derived geophysical parameters. The data products are being distributed to a large user community comprising scientific researchers, educators and operational government agencies. Notable progress has been made in the last decade in facilitating access to data. However, to realize the full potential of the growing archives of valuable scientific data, further progress is necessary in the transformation of data into information, and information into knowledge that can be used in particular applications. Sponsored by NASA s Intelligent Systems Project within the Computing, Information and Communication Technology (CICT) Program, a conceptual architecture study has been conducted to examine ideas to improve data utilization through the addition of intelligence into the archives in the context of an overall knowledge building system. Potential Intelligent Archive concepts include: 1) Mining archived data holdings using Intelligent Data Understanding algorithms to improve metadata to facilitate data access and usability; 2) Building intelligence about transformations on data, information, knowledge, and accompanying services involved in a scientific enterprise; 3) Recognizing the value of results, indexing and formatting them for easy access, and delivering them to concerned individuals; 4) Interacting as a cooperative node in a web of distributed systems to perform knowledge building (i.e., the transformations from data to information to knowledge) instead of just data pipelining; and 5) Being aware of other nodes in the knowledge building system, participating in open systems interfaces and protocols for virtualization, and collaborative interoperability. This paper presents some of these concepts and identifies issues to be addressed by research in future intelligent systems technology.

Ramapriyan, Hampapuram K.↗

Post-fire soil respiration in late growing season (2023 and 2024), Kougarok Fire Complex, Seward Peninsula, Alaska

Field soil respiration data collected in 2023 and 2024 from burned and unburned tussock tundra sites in the Kougarok Fire Complex, near Nome, on the Seward Peninsula of Alaska. Specifically, we measured soil properties and late-growing season CO2 fluxes in patches of unique plant functional types (forbs, shrubs, and graminoids) across two years in tundra recovering from repeated wildfires over the decade. The goal was to identify the main drivers of soil respiration in Arctic tundra underlain by discontinuous permafrost that is recovering from two recent, repeated wildfires that differed in fire age and number of times burned, thereby resulting in different levels of vegetation and subsurface property changes (i.e., successional trajectories). There are five files in *.csv format with one data file and four data description files including data, dictionary, methods, terminology, and file-level metadata. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), is a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research. The NGEE Arctic Phase 3 project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

Santos, Fernanda [ORNL] (ORCID:0000000191555623)↗

IDN Update

To present on recent developments of the IDN including the migration of Earth Dataset collection and services to the Common Metadata Repository (CMR), development of dMMT and IDN Search Portal. Discuss working with the other CEOS agencies to get agreements to make decisions on future interoperability approaches for IDN and for the WGISS Connected Assets systems approaches for IDN and for the WGISS Connected Assets systems.

Committee on Earth Observation Satellites (CEOS)↗

Integrating Ideas for International Data Collaborations Through The Committee on Earth Observation Satellites (CEOS) International Directory Network (IDN)

The capabilities of the International Directory Network's (IDN) version MD9.5, along with a new version of the metadata authoring tool, "docBUILDER", will be presented during the Technology and Services Subgroup session of the Working Group on Information Systems and Services (WGISS). Feedback provided through the international community has proven instrumental in positively influencing the direction of the IDN s development. The international community was instrumental in encouraging support for using the IS0 international character set that is now available through the directory. Supporting metadata descriptions in additional languages encourages extended use of the IDN. Temporal and spatial attributes often prove pivotal in the search for data. Prior to the new software release, the IDN s geospatial and temporal searches suffered from browser incompatibilities and often resulted in unreliable performance for users attempting to initiate a spatial search using a map based on aging Java applet technology. The IDN now offers an integrated Google map and date search that replaces that technology. In addition, one of the most defining characteristics in the search for data relates to the temporal and spatial resolution of the data. The ability to refine the search for data sets meeting defined resolution requirements is now possible. Data set authors are encouraged to indicate the precise resolution values for their data sets and subsequently bin these into one of the pre-selected resolution ranges. New metadata authoring tools have been well received. In response to requests for a standalone metadata authoring tool, a new shareable software package called "docBUILDER solo" will soon be released to the public. This tool permits researchers to document their data during experiments and observational periods in the field. interoperability has been enhanced through the use of the Open Archives Initiative s (OAI) Protocol for Metadata Harvesting (PMH). Harvesting of XML content through OAI-MPH has been successfully tested with several organizations. The protocol appears to be a prime candidate for sharing metadata throughout the international community. Data services for visualizing and analyzing data have become valuable assets in facilitating the use of data. Data providers are offering many of their data-related services through the directory. The IDN plans to develop a service-based architecture to further promote the use of web services. During the IDN Task Team session, ideas for further enhancements will be discussed.

Olsen, Lola M.↗

SPRUCE: Peat Core Sample Collection Metadata, Marcell Experimental Forest, Minnesota, August 2024

This data set contains metadata associated with peat core samples collected from the Spruce and Peatland Responses Under Changing Environments (SPRUCE) experiment in August 2024. This sample metadata contains no analytical results and is a reference for analytical datasets. To ensure accessibility and discoverability, each sample was assigned an International Generic Sample Number (IGSN), a persistent identifier, using System for Earth and Extraterrestrial Sample Registration (SESAR). These samples were used for downstream analysis by multiple teams of researchers the results of which will be reported separately. This dataset contains one data file in comma separate (.csv) format. Additional metadata are provided: one data dictionary and a file-level metadata file in comma separate (.csv) format and a user guide in PDF (*.pdf) format. An aliquot of most samples is stored at Oak Ridge National Laboratory and may be available for further analysis. Access this collection event on SESAR https://doi.org/10.58052/IEJ9B00VQ. To inquire about obtaining archived samples for analysis, reach out using the Contact Sample Owner form located on the bottom of the landing page in SESAR. Note: Only dried and ground material from C Cores are available for new analysis.

Birkebak, Joshua [ORNL] (ORCID:0009000955611494)↗