Engineering PapersSearch

SEARCH · Engineering Papers

Results for “schema”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

COMPASS-FME Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) Experiment Tree Inventory

This is the tree inventory (diameter, species, and live/dead status) data from the Terrestrial Ecosystem Manipulation to Probe the Effects of Storm Treatments (TEMPEST) experimental site. This manipulative, ecosystem-scale TEMPEST experiment is part of the COMPASS-FME (Coastal Observations, Mechanisms, and Predictions Across Systems and Scales: Field Measurements and Experiments; see https://compass.pnnl.gov/FME/COMPASSFME) project. It addresses the potential for freshwater and estuarine-water disturbance events to alter tree function, species composition, and ecosystem processes in a deciduous coastal forest in eastern Maryland, USA. The experiment uses a large-unit (2000 m2), un-replicated experimental design, with three 50 m × 40 m plots serving as control, freshwater, and estuarine-water treatments.This dataset includes:- An overall dataset README file.- The tree inventory data in both "wide" and "long" forms. These contain the same information but are structured differently, with the former more useful for human viewers and the latter more amenable for programmatic analyses.- A key to the species/genus codes used, which follow the U.S. Department of Agriculture's PLANTS schema (https://plants.usda.gov/).- A copy of the R code used to generate the wide- and long-form data files.All files are comma-separated value (CSV) and no special software is required to read them.

54 ENVIRONMENTAL SCIENCES

Lab Homes

This dataset includes processed data from the Lab Homes (LH) Test Facility located on the PNNL campus in Richland, WA. This a set of 2 identical homes that allow for the side-by-side comparison/performance evaluation of different technologies under the same weather at any given time. The dataset spans December 6, 2021 to December 27, 2021 and represents a series of tests performed; calibration, set-point excitation, pre-heating, free-floating and warm up. The measurements correspond to whole building electrical power, HVAC energy use, water heating, appliances and lighting, as well as space temperatures, space humidity, window glass surface temperatures, through glass solar radiation, and meterological data from an onsite meteorological weather station. In addition to the measurements, a metadata .json file, a .ttl file to visualize the data as per BRICK schema, and a detailed .pdf description of the dataset are also provided.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Electricity Baseline 2022

The Electricity Baseline (2022) is a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data and was created using the ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0). The Python package used the "ELCI_2022" model configuration to set the facility and generation data sources and years that were used to create this life cycle inventory, which were taken from publicly accessible datasets and automatically curated into a local data store. An archive of the data stores used in this model is available online: https://doi.org/10.18141/2569193. This model is presented in GreenDelta's openLCA schema v2 JSON-LD format (https://greendelta.github.io/olca-schema/).

Electricity; LCA; data inventory

Electricity Baseline 2021

The Electricity Baseline (2021) is a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data and was created using the ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0). The Python package used the "ELCI_2021" model configuration to set the facility and generation data sources and years that were used to create this life cycle inventory, which were taken from publicly accessible datasets and automatically curated into a local data store. An archive of the data stores used in this model is available online: https://doi.org/10.18141/2569576. This model is presented in GreenDelta's openLCA schema v2 JSON-LD format (https://greendelta.github.io/olca-schema/).

Electricity; LCA; LCI; Life Cycle

Electricity Baseline 2020

The Electricity Baseline (2020) is a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data and was created using the ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0). The Python package used the "ELCI_2020" model configuration to set the facility and generation data sources and years that were used to create this life cycle inventory, which were taken from publicly accessible datasets and automatically curated into a local data store. An archive of the data stores used in this model is available online: https://doi.org/10.18141/2569605. This model is presented in GreenDelta's openLCA schema v2 JSON-LD format (https://greendelta.github.io/olca-schema/).

Electricity; LCA; LCI; data inventory

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris

AI in Science Communication

Generative AI has brought innovations across multiple fields, offering great tools for enhanced communication and efficiency. This project focused on developing a custom AI chatbot using OpenAI's Chat GPT (GPT-4o) to support the Fermilab communications team. An analysis identified Chat GPT as the optimal choice, leading to the adoption of its team version and the implementation of a real-time JSON schema for website scanning. Four distinct personas were created to tailor responses to specific audiences, and Fermilab's published content was uploaded to ensure tone consistency. The training involved iterative prompt trials, resulting in a responsive and effective communication assistant. Initial evaluations indicate that the custom GPT shows promise.

Valle, Diego

CO2-Locate: A Dynamic Database and Tool for Accessing National Oil and Gas Well Data to Inform Carbon Storage Projects

The CO2-Locate Database is a growing compilation of publicly available wellbore resources that have been merged based on common attributes across data sources with an attribute schema developed to be consistent across disparate resources, reduce data gaps, and eliminate record redundancy. The first version of CO2-Locate has been published to Energy Data eXchange (EDX) and includes the integrated public wells dataset as well as additional geospatial summary layers of key wellbore characteristics to protect proprietary resources. Additionally, the CO2-Locate database has been deployed into a web application, enabling easy access, data filtering capabilities, and visualization of U.S. wellbore infrastructure by stakeholders to inform injection site selection and risk assessments.

Dyer, Alec S. [NETL Site Support Contractor, Natio

Metadata Standards for the NSE: Core Fields

This standard presents a core set of metadata fields required for each managed digital object within the Nuclear Security Enterprise (NSE). Metadata standardization is a critical enabler for two primary objectives: 1) effectively sharing data, documents, and other digital objects between NSE sites; and 2) supporting digital engineering through the digital thread at the object level. Standardization is necessary for both schemas and vocabularies, meaning that both field standards and value standards must be specified. This document serves as a foundational field standard, recommending a core set of fields that should be uniformly required for all managed digital objects within the NSE.

99 GENERAL AND MISCELLANEOUS

SLIA Reference Architecture Models

The SLIA Reference Architecture Models project, sponsored by the DOE CESER Energy CyberSense Program (Oct 2024–Sep 2025), advanced LLNL’s PySCES simulation tool to better support CyTRICS Prioritization and Initial Risk Assessment (PIRA) reference architectures. Key achievements include enhancements to the PySCES transmission substation facility model, expanded asset coverage, and enhancements to the PySCES code base. Software improvements reduced code complexity, migrated PySCES to Python version 3.11, introduced an object-oriented design, and added a schema database for easier updates and validation. New features support device criticality assessments and a more precise parametric simulation mode. Remaining gaps include model validation, workflow limitations, Monte Carlo convergence issues, full device criticality metric implementation, model fidelity, and general software improvements. Continued development is recommended to address these gaps and fully align PySCES with CyTRICS PIRA requirements.

97 MATHEMATICS AND COMPUTING

What Are Ontologies and When Should They Be Used?

Data without description is at best unusable, and at worst, misused. If we do not understand the assumptions and meaning of our data, we are unable to confidently use it. Data today is largely described within a database’s schema, detailing structure and primitive datatypes as part of a relational model, but if we require assurance some data value can be correctly evaluated alongside others beyond the immediate systems in which they are defined, a more portable, richer semantics is needed. Ontologies define knowledge unambiguously across systems and establish the means to reason upon said knowledge using logical inference. They model neutral domains of information rather than data definitions from software or databases that would only serve to enrich a single system’s idiosyncrasies. In this paper, we take a casual stance to explore what ontologies are, how they are built, why they are useful, and when they should be used.

97 MATHEMATICS AND COMPUTING

Instantiation of the Damara Tern Platform for Advanced Materials and Manufacturing Technologies (AMMT) Program Collaborative Data Management

This work package focused on deploying an instance of the Damara Tern platform to support AMMT collaborative research activities. The objectives were to provide selected AMMT collaborators with access to a shared environment for capturing operations, trackables, and associated metadata, and to implement data entry functionalities that reflect site-specific procedures. Key activities included creating configurable, schema-driven entry forms and validating the data collection process. The report details the deployment process, the platform infrastructure, and the implemented data entry workflows, providing a reference for end users and establishing a foundation for future production-scale deployments.

36 MATERIALS SCIENCE

OEDI—Solar Grid Integration Data and Analytics Library

As a part of the Open Energy Data Initiative, this effort aims to develop and demonstrate novel distribution state estimation, control optimization, and transient analysis as well as provide access to data, data integration, and mapping information. More specifically, the focus of the effort will be on physics-based distribution system state estimation, hybrid (physics-based and machine learning) distribution optimal power flow, and event detection/analysis for solar integration and analytics. This work will enable reproducible, robust, replicable, and generalizable R&D in simulation and emulation of solar system integration. These test models and datasets will provide an integrated library for developing and testing power system operation technologies. To make the library user-friendly, this project will provide data curation tools such as data translators, mapping scripts and APIs, database schemas and metadata, interfaces and user dashboard, source code for the reference algorithms, description of the use-cases/scenarios, and comprehensive information on all the assumptions.

14 SOLAR ENERGY

Vulcan-Forge: Architecture and Design of a Multi-Modal Forensic Analysis Plugin for CALDERA

Forge and VULCAN together describe an open-architecture cybersecurity analysis ecosystem that unifies forensic artifact processing, detection engineering, and vulnerability intelligence within integrated platforms. Forge operates as a plugin for MITRE CALDERA, ingesting diverse evidence formats—including EVTX, PCAP/PCAPNG, CSV, JSON, YAML, XML, binaries, and archives—to construct a unified artifact graph enriched with severity scoring, TLP classification, and audit trails. It provides subsystems for artifact parsing, streaming structured-data visualization, NetworkMiner-based packet inspection, PE/.NET binary analysis, and LLM-assisted triage and rule generation, with outputs validated against CCCS-YARA and pySigma schemas. VULCAN complements this by serving as a cybersecurity analyst platform that integrates a Neo4j knowledge graph, Qdrant vector retrieval, SSVC-based triage, and a local LLM to deliver CVE intelligence and forensic analysis through a multi-source ingest pipeline drawing from NVD, CISA KEV, EPSS, MITRE ATT&CK, and CAPEC. Together, they bridge structured threat intelligence with automated forensic analysis and detection workflows.

97 MATHEMATICS AND COMPUTING

Metadata Standards for the NSE: Extended Field Standards

This standard presents a set of optional metadata fields for managed digital objects within the Nuclear Security Enterprise (NSE) and provides a deeper look at data representation in metadata by looking at the representation of 1) Records Management required metadata, and 2) common representations of technical/scientific data. Metadata standardization is a critical enabler for effectively sharing data, documents, and other digital objects between NSE sites, and for tracing the digital thread at the object level. Standardization is necessary for both schemas and vocabularies, meaning that both field standards and value standards must be specified. This document serves as a complementary field standard, recommending an optional set of fields that should be uniformly built for all managed digital objects within the NSE. This document specifically focuses on extending the shared discovery layer defined in the first white paper by introducing additional descriptive and data representation fields that improve cross-site search and interpretation.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

Genesis Mission Data cards

As data-intensive research and artificial intelligence become central to DOE mission science, the need for machine-actionable dataset documentation has grown accordingly. However, many DOE-aligned communities, including the Office of Science, NNSA, and cross-laboratory collaborations, have developed independent metadata practices. This fragmentation creates friction for discovery, federation, and reuse across programs. To address these challenges, this talk introduces the Genesis Data Card: a shared metadata artifact developed in collaboration with a broad DOE community (Jefferson Lab and the National Lab of the Rockies, Oak Ridge, Sandia, Idaho, Berkeley, and Los Alamos). The Genesis Data Card aims to standardize dataset documentation across DOE-aligned initiatives while remaining extensible to discipline-specific needs. This talk will describe the data card template and the supporting code to validate completed data cards, using a companion LinkML schema. I'll walk through the design decisions behind the template, its alignment with existing standards, its treatment of sensitivity and governance metadata, and the phased roadmap toward lifecycle-integrated "xCards" that support autonomous discovery and reuse. The talk closes with current gaps, ongoing work, and how others can contribute datasets and feedback to the shared repository.

McSpadden, Helen [Thomas Jefferson National Accele

C-HER Metadata Overview: Approach, Standards, and Rigor for the Centralized Health and Exposomic Resource

The Centralized Health and Exposomic Resource (C-HER) unifies environmental, demographic, geographic, and health-related data for exposomic research. The source data differ in format, geographic coverage, time period, resolution, terminology, and documentation. We use a common metadata framework to describe those differences and to record how each data resource has been processed, documented, and ingested. This document relates only to the C-HER metadata framework. It explains the information that is recorded for each resource, the standards used to organize that information, the conditions for metadata completeness, and the relationship between metadata and quality review. It is intended for those who need to understand what C-HER metadata communicates and how it supports appropriate use of the data. It is not an implementation specification or procedure. It does not document the database schema, source code, deployment configuration, transformation algorithms, or dataset-specific QA/QC thresholds. Those materials are maintained separately.

MacFarland, Midgie [ORNL] (ORCID:0009000807354078)

Knowledge Graph for End-to-End Traceability of an Integrated Human-Earth System Model

Integrated human-Earth system models inform energy-water-land system dynamics and policies, yet their results are difficult to trace through input-data, model structure, scenario configurations, and solved outputs. Because this information is siloed across disconnected artifacts, process-based IAMs have historically lacked a unified, queryable representation. Such lack of traceability prevents researchers from systematically isolating the multi-sector drivers of complex outcomes (such as tracing water-scarcity results back to distant energy-system dynamics) or conducting holistic uncertainty attribution across hundreds of interacting parameters. To address this concern, our work documents the software engineering process of a knowledge graph that unifies these four layers for the Global Change Analysis Model (GCAM-USA_Reference scenario, GCAM v9.1). The graph was built as a relational property graph in DuckDB from the run’s own artifacts: the input-preparation dependency map (gcamdata chunk map), the model’s XML input files, the run configuration, and the results database (BaseX), successfully mapping the model’s declared structure. The resulting graph comprises 204,321 nodes and 1,687,814 edges across 16 node types and 15 edge types, with approximately 16.3 million time-series values stored separately to maintain structural efficiency. To ensure representation fidelity, every edge carries an epistemic-status annotation recording the warrant for the relationship (structural, provenance, dependency, or model-derived), and a machine-readable provenance ledger classifying the origin of every schema element. Evaluation against a fixed five-benchmark suite with locked baselines reports zero structural orphans, zero dangling edge endpoints, and 100% of output-producing technologies traceable to raw input files. Two interactive interfaces present the graph, including a serverless browser application built on DuckDB-Wasm. By establishing the first end-to-end provenance framework for an IAM, this work enables researchers and scientists to systematically audit complex policy scenarios, debug model structures, and trace policy-relevant outputs to their data origins in real time.

Artifical Intelligence