Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “system metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Metadata Schemas and Ontologies for Building Energy Applications: A Critical Review and Use Case Analysis

With the increasing digitalization of processes throughout the lifecycle of buildings, data exchanged between stakeholders and between building systems has grown significantly. However, a lack of semantic interoperability between data in different systems is still prevalent, hindering the development of applications that can be reused across buildings and limiting the scalability of innovative solutions. Semantics refers to the description of the meaning of the data in a way that can be consistently understood by applications. Recently, several competing initiatives have been developing metadata schemas and ontologies to express this semantic information for different applications in the building domain. This paper systematically reviews these schemas and conducts an analysis of five of them to evaluate their applicability to three high-value use cases for building operations: energy audits, automated fault detection and diagnostics and optimal control. The survey finds 40 schemas published in the last 10 years but but their actual use in industry is difficult to estimate. Among the five selected ontologies, several gaps are highlighted in relation to the three use cases. Recommendations for the future include better harmonization of these initiatives, more centralized repositories and search engines for these schemas as well as better industry engagement to facilitate their adoption.

Smart Building, Sematic, Metadata, Ontology, Data ↗

Genesis Data Card Schema, Template and Supporting Tools

Genesis Data Cards provide a standardized template and schema for documenting scientific datasets in support of discovery, access, interoperability, reusability, governed use, and AI usability. This release of the Genesis Data Card repository includes a versioned Markdown template, a LinkML schema with generated Pydantic and JSON artifacts, schema documentation, and example completed data cards. Validation tooling is provided to ensure that completed data cards conform to the schema prior to submission. Accompanying documentation for the structured metadata is provided as a Field Reference Guide. The schema and accompanying template provided in this repository address the call for actionable context that enables humans and AI systems to find, access, interpret, cite, and reuse data, and, when appropriate, integrate it into AI and machine learning workflows. The data card is intended to serve as a common metadata artifact intended to support standardized, cross-program dataset documentation across Department of Energy (DOE)-aligned efforts, including but not limited to Genesis Mission-related implementations, the Office of Science, National Nuclear Security Administration (NNSA), and Advanced Simulation and Computing (ASC) data governance and stewardship initiatives.

data card↗

Control systems and data management for high-power laser facilities

The next generation of high-power lasers enables repetition of experiments at orders of magnitude higher frequency than what was possible using the prior generation. Facilities requiring human intervention between laser repetitions need to adapt in order to keep pace with the new laser technology. A distributed networked control system can enable laboratory-wide automation and feedback control loops. These higher-repetition-rate experiments will create enormous quantities of data. A consistent approach to managing data can increase data accessibility, reduce repetitive data-software development and mitigate poorly organized metadata. An opportunity arises to share knowledge of improvements to control and data infrastructure currently being undertaken. We compare platforms and approaches to state-of-the-art control systems and data management at high-power laser facilities, and we illustrate these topics with case studies from our community

47 OTHER INSTRUMENTATION↗

Development of a Digital Twin for Hydrogen Dispersion and Safety Assessment in an Electrolyzer Based Hydrogen Production Facility

Digital twin models are virtual representations of physical systems that use real-time data to simulate and optimize performance. This study presents the development and initial implementation of a digital twin (DT) for the electrolyzer-based hydrogen production facility at NREL's Advanced Research on Integrated Energy Systems (ARIES), focused on enhancing safety and optimizing sensor placement through physics-based simulations and metadata integration. The DT incorporates detailed facility-specific information, including component layout, leak locations, and controlled release parameters, to model hydrogen dispersion under varying environmental conditions. Using steady-state computational fluid dynamics (CFD) simulations informed by real meteorological data, such as wind speed, direction, and vertical wind profiles, the DT enables visualization of hydrogen plume behavior and spatial concentration distributions. Comparative analysis between high and low wind speed scenarios illustrates the significant influence of wind dynamics on plume shape and extent, with horizontal momentum dominating dispersion at higher speeds, while buoyancy effects become more prominent under low wind conditions. These simulations generate a rich dataset embedded within the DT, allowing users to assess potential leak outcomes and identify optimal sensor locations based on concentration thresholds. The model supports scenario-based analysis to guide safety strategies and equipment deployment for open-area hydrogen infrastructure. The digital twin thus serves as a dynamic platform for virtual prototyping, providing predictive insight into hydrogen behavior and enhancing risk-informed decision-making. This initial phase establishes a validated foundation for future integration of transient, uncontrolled leak scenarios and real-time sensor feedback, positioning the DT as a critical tool for safety design, operational planning, and adaptive monitoring in hydrogen systems. Overall, the approach demonstrates the value of combining environmental data with digital simulations to inform safer and more efficient deployment of hydrogen technologies.

08 HYDROGEN↗

Development of a Digital Twin for Hydrogen Dispersion and Safety Assessment in an Electrolyzer-Based Hydrogen Production Facility: Preprint

Digital twin models are virtual representations of physical systems that use real-time data to simulate and optimize performance. This study presents the development and initial implementation of a digital twin (DT) for the electrolyzer-based hydrogen production facility at the National Renewable Energy Laboratory (NREL)'s Advanced Research on Integrated Energy Systems (ARIES), focused on enhancing safety and optimizing sensor placement through physics-based simulations and metadata integration. The DT incorporates detailed facility-specific information, including component layout, leak locations, and controlled release parameters, to model hydrogen dispersion under varying environmental conditions. Using steady-state computational fluid dynamics (CFD) simulations informed by real meteorological data, such as wind speed, direction, and vertical wind profiles, the DT enables visualization of hydrogen plume behavior and spatial concentration distributions. Comparative analysis between high and low wind speed scenarios illustrates the significant influence of wind dynamics on plume shape and extent, with horizontal momentum dominating dispersion at higher speeds, while buoyancy effects become more prominent under low wind conditions. These simulations generate a rich dataset embedded within the DT, allowing users to assess potential leak outcomes and identify optimal sensor locations based on concentration thresholds. The model supports scenario-based analysis to guide safety strategies and equipment deployment for open-area hydrogen infrastructure. The digital twin thus serves as a dynamic platform for virtual prototyping, providing predictive insight into hydrogen behavior and enhancing risk-informed decision-making. This initial phase establishes a validated foundation for future integration of transient, uncontrolled leak scenarios and real-time sensor feedback, positioning the DT as a critical tool for safety design, operational planning, and adaptive monitoring in hydrogen systems. Overall, the approach demonstrates the value of combining environmental data with digital simulations to inform safer and more efficient deployment of hydrogen technologies.

08 HYDROGEN↗

Frontier Job-Centric Telemetry Dataset

Comprehensive analysis of high-performance computing (HPC) systems requires linking workload execution to system behavior. This kind of analysis is vital for diagnosing performance issues, managing capacity, detecting anomalous workloads, and understanding how applications interact with system hardware. This job-centric telemetry dataset unifies scheduler job records with node-level measurements, enabling direct association between workloads and their corresponding power, thermal, and performance characteristics. It contains sanitized, scheduler related metadata for 152,400 individual jobs that ran on the Frontier supercomputer and ended on selected days throughout 2024 and 2025, a subpopulation of ~6.8% of the total number of allocated jobs with non-zero run time on the system over that same period. Each is linked with files that contain telemetry time series records of the power utilization and temperature behavior of its allocated nodes and their processors during the run time of the job. Where available, a portion of the job files also contain network performance time series. Jobs are sampled from select days that reflect normal levels of user activity and possess job size distributions with large numbers of leadership class jobs (>20% of Frontier nodes). Jobs in this dataset attempt to best represent successful user workflows.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Clustering-Based Predictive Analytics to Improve Scientific Data Discovery

Given the sheer volume of scientific data archived within the data-intensive projects at the US Department of Energy's Oak Ridge National Laboratory, finding precisely what data we are looking for may not be a trivial task; conversely, we may also miss a more prominent data product. To address such issues, we propose improving the data discovery system and using data analytics methods to comprehend what specific users might be interested in based on their physiological state, search patterns, and past data usage history. This work's primary goal is to prune the complexity, increase the visibility of popular data products, and direct users toward the data that best meet their needs. The proposed algorithm constructs a user profile based on the user's explicit or implicit interactions with the system, such as items they are currently looking at on-site and the key metadata mappings related to the data set. The pattern is then used to build a training data set, which will help find relevant data to recommend to the user.

Devarakonda, Ranjeet↗

Produced Water DNA Database (PW-DNA): Utilizing KBase to generate an environmental specific curated molecular database

The deep subsurface is estimated to host the majority of Earth’s microbial biomass yet remains one of the most challenging environments to access and study. One common approach to investigate these microbial communities is through the analysis of produced water from subsurface reservoirs, where researchers can assess water and gas chemistry along with molecular (DNA/RNA) sequence data. Advances in high-throughput sequencing have greatly expanded our understanding of these environments and their biotechnological potential. However, further progress requires large-scale, integrative meta-analyses across diverse datasets. To address this need, we developed the Produced Water-DNA (PW-DNA) Database, a curated, publicly available resource that consolidates microbial DNA/RNA sequences, geochemical data, and relevant metadata from in situ hydrocarbon environments such as coal beds, oil reservoirs, and natural gas systems. The PW-DNA database delivers three core benefits to the research community: (1) it improves data sharing by linking environmental microbial datasets with corresponding geochemical parameters, enabling more robust filtering and analysis; (2) it connects with complementary research databases to promote broader dissemination and interoperability; and (3) it supports technological innovation by serving as a resource for identifying microbial trends and exploring genetic potential. While individual studies have highlighted basin-specific microbial communities and functional redundancy in biogeochemical cycling, a comprehensive, system-wide perspective is needed to better understand connectivity and novelty across subsurface ecosystems. By designing the PW-DNA in the KBase platform, we provide a reproducible, visual framework for integrating large-scale genomic and geochemical data, enabling researchers to perform more informed analyses and experimental design. Ultimately, this resource enhances the ability to identify, characterize, and interpret microbial functions across diverse subsurface environments, thereby accelerating discovery in subsurface microbiology and biotechnology.

59 BASIC BIOLOGICAL SCIENCES↗

System and method of managing large data files

Disclosed are systems and software that provide a high-performance, extensible file format and web API for remote data access and a visual interface for data viewing, query, and analysis. The described system can support storage of raw spectroscopic data such as neural recording data, MSI data, metadata, and derived analyses in a single, self-describing format that may be compatible by a large range of analysis software.

Bowen, Benjamin P.↗

WHONDRS laboratory time series moisture manipulative experiment from soil core layers across eastern contiguous US: time series aerobic respiration, geochemistry, and aggregates

This dataset supports a broader study examining the effects of wetting and drying on soil layers across the eastern contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata. Samples were collected as part of a collaboration between WHONDRS (Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems; https://whondrs.pnnl.gov) and MONet (Molecular Observation Network; https://www.emsl.pnnl.gov/monet). The field samples (soil cores) were labeled as MEL_##_COR and subsequent subsamples begin with MEL_##. Additional subsamples were taken for the laboratory experiment and were labeled as EL_##. The labels from the MEL field samples and the EL subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EL_01 is a subsample from MEL_01). See the critical details section below for more details on sample naming and experimental design.For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions.This dataset is comprised of (1) a folder containing environmental context photos; (2) file-level metadata; (3) data dictionary; (4) field metadata; (5) readme; (6) international generic sample number (IGSN) mapping file; and (7) a subfolder with soil sample data from field samples and the incubation experiment. The sample data subfolder contains (1) effect size; (2) gravimetric moisture from field samples and incubation experiment; (3) respiration rates, raw dissolved oxygen values, and plots; (4) specific conductance, pH, and temperature from the incubation; (5) soil aggregates; (6) a summary containing median values of each data type for each treatment (wet and dry) in the incubation; (7) a summary containing averages for each data type of each soil layer; and (8) methods codes. All files are .csv, .pdf, .jpeg, or .jpg.

54 ENVIRONMENTAL SCIENCES↗

NEPATEC v2.0: Standardized Metadata and Text Corpus of National Environmental Policy Act Documents

The National Environmental Policy Act of 1969, as amended (NEPA), is a major environmental law in the United States, requiring Federal agencies to consider and document potential environmental impacts before deciding on a proposed action. Modernization of NEPA and permitting processes faces significant challenges due to the lack of standardized formats and interoperable systems for organizing and sharing NEPA-related information across agencies. Much of the information gathered during NEPA reviews is written into documents such as categorical exclusions, environmental assessments, and environmental impact statements, then filed in predominately independent agency file stores that may or may not be publicly accessible. The application of metadata and data standards, such as those recommended by the Council on Environmental Quality (CEQ), to NEPA documents offers a shared vocabulary and structure for key entities like projects, processes, and documents that can streamline information exchange and enhance collaboration across systems. In this work, we publicly release NEPATEC2.0, an expanded corpus of NEPA documents with associated metadata. NEPATEC2.0 encompasses approximately 120,000 documents from 60,000 projects prepared by more than 60 different agencies. Modeled to align with CEQ metadata standards, NEPATEC2.0 promotes consistency in environmental reviews and supports the ongoing effort to modernize permitting technologies by facilitating more transparent, efficient, and data-driven decision-making. Importantly, NEPATEC2.0 demonstrates the possibilities and limitations of large language model-based prompting to extract information from NEPA documents at scale.

54 ENVIRONMENTAL SCIENCES↗

PoliMOR

PoliMOR is a scalable, automated, and customizable policy engine framework for multi-tiered parallel file systems. It is composed of single-purpose agents that handle tasks such as gathering file metadata, making policy decisions, and then executing actions based on those policies. These agents are designed to communicate using distributed message queues, allowing the number of individual agents to be scaled up as needed. PoliMOR automates the data management tasks by precluding the need for admin intervention. The agents in PoliMOR can be customized to integrate any utilities/tools that perform tasks like metadata scanning and data placement management.

Brumgard, Christopher↗

Remote sensing images, DEM, and point clouds associated with “Accuracy evaluation of cost-effective 3D reconstruction approaches for hydrobiogeochemical processes in non-perennial stream riverbeds”

This data package is associated with the publication “Accuracy evaluation of cost-effective 3D reconstruction approaches for hydrobiogeochemical processes in non-perennial stream riverbeds” published in Frontiers in Environmental Science, Environmental Informatics and Remote Sensing (Bao et al., 2026; doi: 10.3389/fenvs.2026.1725258). This data package includes the drone photos for a section of Umtanum Creek in Washington, Unted States. The photos were used to reconstruct the 3-dimensional (3D) digital elevation model (DEM) of the riverbed for the investigated stream section. The reconstruction results from four approaches are provided: (1) unoccupied aerial vehicle (UAV, colloquially known as drone) imagery-based Structure-from-Motion (SfM), (2) a machine learning-based 3D reconstruction model, Visual Geometry Grounded Deep Structure from Motion (VGGSfM), (3) Visual Geometry Grounded Transformer for long sequence of images (VGGT-Long), and (4) handheld smartphone LiDAR scanning. The ground truth measurements by tripod-mounted optical level kit and ground control points GPS locations for evaluating the accuracy of the four reconstruction approaches are also provided in this data package. A preliminary version of this data package was published in October 2025 at the time of manuscript submission. It was updated in March 2026, at the time of manuscript acceptance, to include additional metadata (this readme, data dictionary, and file level metadata). The data did not change. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. In addition to a readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. This dataset is comprised of (1) 8 folders; (2) the detailed flight configuration html files; (3) field metadata; (4) a readme; (5) a data dictionary; and (6) file-level metadata. The folders “2024_10_18_d01” and “2024_10_18_d02” contain the original drone photos for the two drone flights (d01 and d02) on October 18, 2024. The reconstruction results from each of the approaches are in the folders called “ODM_SfM”, “VGGSfM”, “VGGTLong”, and “LiDAR”. The ground truth measurements are in the folder called “optical_level_kit”. Lastly, results comparing the different approaches are in the folder called “comparisons”. All files are .csv, .html, .jpg, .obj, .txt, and .npy. For information on using the .obj and .npy files, see the readme files within the same folder as the files.

54 ENVIRONMENTAL SCIENCES↗

Multispectral UAV imagery of experimental freshwater wetlands under 5 ppt saltwater intrusion, Louisiana, 2023 and 2024

Multispectral imagery was collected using an unmanned aerial vehicle (UAV) to evaluate how freshwater vegetation responds to short-term simulated saltwater intrusion events. The purpose of this data collection was to understand how plant health changes in response to acute salinity exposure, which is increasingly relevant in coastal wetland ecosystems facing sea level rise and storm surge events, such as in coastal Louisiana. Three experimental saltwater intrusions were conducted at a salinity of approximately 5 parts per thousand (ppt) for durations of 6-days, 10-days, and 17-days. UAV flights occurred both before and after each treatment. The resulting imagery was processed using Pix4DMapper software to georeference the images and generate orthomosaics. The multispectral sensor used in this study captures reflectance in five bands: blue, green, red, red-edge, and near-infrared. The uploaded data consist of georeferenced .tif orthomosaics for each spectral band, which are compatible with GIS software for vegetation analysis. This imagery can be utilized in investigations into vegetation stress, remote sensing of freshwater wetland ecosystems, and modeling of plant response to environmental changes.

EARTH SCIENCE > BIOSPHERE > ECOSYSTEMS↗

(U)Vendor Sanitization Requirements – SIEMENS Control

This document defines the software requirements for the sanitization of the IPC and or the controller. The system is intended to manage part programs and associated data (e.g. simulations, probing data) by capturing metadata, securely sanitizing files, logging operations, backing up data to a defined path, deploying a user defined program to a controller, and restoring part programs and associated data when required.

42 ENGINEERING↗

Brick Schema Standardized Plug Load Control Strategies for Load Reduction: Preprint

Plug loads comprise a significant percentage of commercial building energy consumption. Applying intelligent controls to turn off plug loads when unused can provide dynamic load reduction and flexibility, which are key traits of grid-interactive efficient buildings. This capability is important for equitable decarbonization as it can enable disadvantaged communities to electrify buildings without costly upgrades to electrical infrastructure. In this work, we present the effectiveness of various control strategies along with the operational lessons that informed their design. During a three-year period, we operated over 600 smart outlets in 12 university office buildings. The attached plug loads consisted primarily of printers, TVs, water dispensers, and copiers. After recording baseline power measurements for one year, we designed plug load control (PLC) strategies for each plug load type, use, and for different risk tolerance levels because PLC can potentially be disruptive to daily work. We used the Brick Schema to facilitate the management of plug load locations and other metadata. For advanced controls, we integrated the smart plugs with heating, ventilation, and air conditioning (HVAC) systems through the campus building automation system. We found static schedules to be the least disruptive and most predictable for occupants, resulting in 38% and 66% energy savings in two studies. For printers, print server-triggered PLC produced 86% savings, the highest of all strategies with minimal occupant impact. Scheduling of water dispensers and digital signage TVs produced 49% and 70% savings respectively with opportunities to improve performance with the use of HVAC occupancy data.

ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATION↗

In Situ Soil Moisture and Thaw Depth Measurements Coincident with Airborne SAR Data Collections, Seward Peninsula, Alaska, 2022

The in-situ soil moisture and thaw depth measurements provided in this dataset were collected coincident with airborne overflights of L-band synthetic aperture radar (SAR) instruments at the Teller, Kougarok, and Council study sites on the Seward Peninsula, Alaska. Overflights occurred on August 19, 2022. Soil moisture data at Teller and Kougarok was collected on August 19, and at Council on August 20. Thaw depth, soil pits, and any additional measurements were recorded on August 20 and 21. Field measurements and flights were conducted during the summer of 2022 as a collaboration between the National Aeronautics and Space Administration (NASA) Arctic-Boreal Vulnerability Experiment (ABoVE) Project’s Airborne SAR Campaign and the Next-Generation Ecosystem Experiments (NGEE) Arctic Project. This dataset includes a data file (*.csv), a data dictionary (*_dd.csv) and a file-level metadata (*_flmd.csv). The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy’s Office of Biological and Environmental Research. The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska. Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy’s Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

EARTH SCIENCE > LAND SURFACE > FROZEN GROUND↗

Organic layer thickness and carbon concentration in burned and unburned sites, Seward Peninsula, AK, 2022

Measurements associated with organic layer samples collected from naturally burned (1971, 2002, 2015, 2019) and unburned sites at the Kougarok Fire Complex, Seward Peninsula, AK, 2022. Here, a discontinuous permafrost underlies an arctic tundra ecosystem. Measurements include elemental carbon and nitrogen concentrations and stocks, organic layer thickness, and thaw depth. There are five files in *.csv format with one data file and four data description files including data dictionary, methods, terminology, and file-level metadata. The Next-Generation Ecosystem Experiments: Arctic (NGEE Arctic), was a research effort to reduce uncertainty in Earth System Models by developing a predictive understanding of carbon-rich Arctic ecosystems and feedbacks to climate. NGEE Arctic was supported by the Department of Energy's Office of Biological and Environmental Research.The NGEE Arctic project had two field research sites: 1) located within the Arctic polygonal tundra coastal region on the Barrow Environmental Observatory (BEO) and the North Slope near Utqiagvik (Barrow), Alaska and 2) multiple areas on the discontinuous permafrost region of the Seward Peninsula north of Nome, Alaska.Through observations, experiments, and synthesis with existing datasets, NGEE Arctic provided an enhanced knowledge base for multi-scale modeling and contributed to improved process representation at global pan-Arctic scales within the Department of Energy's Earth system Model (the Energy Exascale Earth System Model, or E3SM), and specifically within the E3SM Land Model component (ELM).

54 ENVIRONMENTAL SCIENCES↗