Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data systems standards”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Data from: “Enabling FAIR data in Earth and environmental science with community-centric (meta)data reporting formats”

This dataset contains supplementary information for a manuscript describing the ESS-DIVE (Environmental Systems Science Data Infrastructure for a Virtual Ecosystem) data repository's community data and metadata reporting formats. The purpose of creating the ESS-DIVE reporting formats was to provide guidelines for formatting some of the diverse data types that can be found in the ESS-DIVE repository. The 6 teams of community partners who developed the reporting formats included scientists and engineers from across the Department of Energy National Lab network. Additionally, during the development process, 247 individuals representing 128 institutions provided input on the formats. The primary files in this dataset are 10 data and metadata crosswalk for ESS-DIVE’s reporting formats (all files ending in _crosswalk.csv). The crosswalks compare elements used in each of the reporting formats to other related standards and data resources (e.g., repositories, datasets, data systems). This dataset also contains additional files recommended by ESS-DIVE’s file-level metadata reporting format. Each data file has an associated dictionary (files ending in _dd.csv) which provide a brief description of each standard or data resource consulted in the data reporting format development process. The flmd.csv file describes each file contained within the dataset.

54 ENVIRONMENTAL SCIENCES↗

The Monarch Initiative in 2024: an analytic platform integrating phenotypes, genes and diseases across species

Abstract Bridging the gap between genetic variations, environmental determinants, and phenotypic outcomes is critical for supporting clinical diagnosis and understanding mechanisms of diseases. It requires integrating open data at a global scale. The Monarch Initiative advances these goals by developing open ontologies, semantic data models, and knowledge graphs for translational research. The Monarch App is an integrated platform combining data about genes, phenotypes, and diseases across species. Monarch's APIs enable access to carefully curated datasets and advanced analysis tools that support the understanding and diagnosis of disease for diverse applications such as variant prioritization, deep phenotyping, and patient profile-matching. We have migrated our system into a scalable, cloud-based infrastructure; simplified Monarch's data ingestion and knowledge graph integration systems; enhanced data mapping and integration standards; and developed a new user interface with novel search and graph navigation features. Furthermore, we advanced Monarch's analytic tools by developing a customized plugin for OpenAI’s ChatGPT to increase the reliability of its responses about phenotypic data, allowing us to interrogate the knowledge in the Monarch graph using state-of-the-art Large Language Models. The resources of the Monarch Initiative can be found at monarchinitiative.org and its corresponding code repository at github.com/monarch-initiative/monarch-app.

60 APPLIED LIFE SCIENCES↗

Suppressing Quantum Circuit Errors Due to System Variability

We present a quantum circuit optimization technique that takes into account the variability in error rates that is inherent across present-day noisy quantum computing platforms. This method can be run after qubit routing or postcompilation and consists of computing isomorphic subgraphs to input circuits and scoring each using heuristic cost functions derived from system calibration data. Using an independent standard algorithmic test suite, we show that it is possible to recover on average nearly 40% of missing fidelity using better qubit selection via efficient to compute cost functions. We demonstrate additional performance gains by considering qubit placement over multiple quantum processors. The overhead from these tools is minimal with respect to other compilation steps, such as qubit routing, as the number of qubits increases. As such, our method can be used to find qubit mappings for problems at the scale of quantum advantage and beyond.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

PowerSystemsData Specification [SWR-25-92]

This repository defines a standardized data format for representing power system datasets, with a focus on supporting time-series visualizations and 3D visualization tools. By providing a consistent and extensible structure, this format aims to streamline the development and interoperability of visualization codebases in the power systems domain. NOODLES is a cross-platform/device/tool protocol for collaborative visualization. NOODLES was Developed at the National Renewable Energy Laboratory (NREL) as a capability of the Insight Center https://www.nrel.gov/computational-science/insight-center.html

Brunhart-Lupo, Nicholas [National Renewable Energy↗

Electrical Energy Storage Data Submission Guidelines, Version 3

The knowledge of long-term health and reliability of energy storage systems is still unknown, yet these systems are proliferating and are expected increasingly to assist in the maintenance of grid reliability. Understanding long-term reliability and performance characteristics to the degree of knowledge similar to that of traditional utility assets requires operational data. This guideline is intended to inform numerous stakeholders on what data are needed for given functions, how to prescribe access to those data and the considerations impacting data architecture design, as well as provide these stakeholders insight into the data and data systems necessary to ensure storage can meet growing expectations in a safe and cost-efficient manner. Understanding data needs, the systems required, relevant standards, and user needs early in a project conception aids greatly in ensuring that a project ultimately performs to expectations.

25 ENERGY STORAGE↗

A Power Application Developer’s Guide to the Common Information Model: An Introduction for Power Systems Engineers and Application Developers – CIM17v40

A key issue in creating the next generation of energy management system (EMS) and advanced distribution management system (ADMS) platforms will be the ability to represent and exchange power system network model data in a consistent manner. To this end, the Common Information Model (CIM) stands out as the only standardized vocabulary (or ontology) for defining power system network models and asset data in a comprehensive, consistent manner across the generation-transmission-distribution boundary. The CIM is freely available to use and extend. The CIM is maintained by the UCAiug (informally known as the CIM User’s Group) under an Apache 2.0 license. The CIM Users Group collaborates with the IEC and other standards communities for the development of technical and informative specifications. Although portions of the information model are referred to by the corresponding IEC standards naming, it is not necessary to purchase any of the IEC standards to use the CIM information model. This document provides a roadmap for power system engineers and application developers not familiar with semantic modeling to start using the CIM for modeling, simulation, optimization, and development of advanced power applications. The key classes needed for defining power system topology and equipment are explained systematically. Key focus areas include modeling of lines, transformers, generators, switching equipment, loads, and distributed energy resources (DERs).

97 MATHEMATICS AND COMPUTING↗

Lifetime fatigue response due to wake steering on a pair of utility-scale wind turbines

Quantifying the impacts on turbine loads during wind farm control is an important consideration in assessing power production benefits. Wake steering controls aim to improve total wind farm performance by coordinating the control actions of individual turbines, wherein an upstream turbine is intentionally yawed at an offset angle from the measured wind direction. Consequently, this redirects its wake for improved power production and potentially reduces fatigue loads of the downstream turbines. This paper studies the lifetime fatigue loads associated with wake steering by using utility-scale wind turbine experimental data to conduct an analysis on a pair of wind turbines. This study was part of a large field experiment in which a group of five GE 1.5-MW SLE CWE turbines were selected as targets for conducting wake steering research. A standard loads instrumentation package and data acquisition system were installed on two turbines within the cluster to measure turbine fundamental loads. The time-series databases were used to calculate loads statistics as well as short-term and lifetime damage equivalent loads. Fatigue calculations followed the guidance in Annex H of the International Electrotechnical Commission standard 61400-1, Edition 4. Lifetime fatigue calculation results are presented in this analysis; three methods of assessing lifetime fatigue were used to determine percent difference for blade root moments, main shaft moments, main shaft torque, and tower top torque. For all three fatigue treatments, some components’ lifetime fatigue increases for the controlled turbine; however, the downwind turbine experienced a reduction in lifetime fatigue and combined effect for the turbine pair results in a reduction of fatigue when wake steering controls are applied.

17 WIND ENERGY↗

Photovoltaic Analysis and Response Support (PARS) Platform for Solar Situational Awareness and Resiliency Services

The project's primary objective is to develop a digital-twin based Photovoltaic (PV) Analysis and Response Support (PARS) platform, which aims to provide real-time situational awareness and optimal response plans. This platform is designed to enhance the performance of hybrid PV systems, making them competitive with or even superior to conventional generation resources. The PARS platform enabled the project team to develop and evaluate an extensive suite of grid support functionalities for the hybrid PV systems to enhance grid performance, across key areas including visibility, dispatchability, security, resilience, and reliability. Given the global push toward achieving 100% clean energy by 2035, there is a significant increase in the integration of inverter-based resources (IBRs) throughout the energy grid. Effectively managing the inherent variability and uncertainty associated with IBRs is crucial for ensuring cost-effectiveness, reliability, and security in both the main grid and islanded microgrids. Constrained to a limited array of IEEE test systems or standard feeder models, traditional IBR modeling struggles to assimilate new field data, accurately reflect system dynamics, and adapt to the evolving energy landscape. In our project, we embraced a Digital Twin (DT) strategy for crafting the PARS platform. A digital twin acts as a precise virtual counterpart of a physical system, built on historical data and continuously honed with real-time insights. This enables the high-fidelity DT to accurately mirror current system operations and forecast future scenarios. Consequently, the PARS platform becomes an ideal environment for testing and refining monitoring, control, power, and energy management algorithms designed to boost hybrid PV system performance. The defining feature of the PARS platform, distinguishing it from other advanced simulation tools, is its exceptional adaptability. This is achieved by employing actual network topologies and utilizing real-time field data for fine-tuning and calibration, ensuring a close emulation of real-world conditions. The project deliverables include: 1) High-fidelity IBR models and tools for real-time parameterization, utilizing real-time field measurements to refine IBR models for enhanced accuracy and performance; 2) Grid-forming and Grid-following capabilities to deliver resilience services, including blackstart, voltage and frequency support, cold-load pick-up, power reserves, and three-phase load balancing across grid-connected and microgrid settings; 3) Machine learning-based forecasting tools and methods for generating synthetic data and topologies, creating diverse and realistic simulation environments for evaluating varied operational scenarios; 4) Advanced microgrid power and energy management algorithms for optimizing the integration and operation of PV, storage, and demand response resources within both feeder and community scales. The power grid data sets are provided by four utility companies in North Carolina and the New York Power Administration. Acting as industry advisors, our industry partners communicated stakeholder needs and regulatory standards to the research teams, aiding technology transfer by incorporating the developed methodologies into their daily operations. This collaboration ensures that the PARS platform, functioning as a power system digital twin, enhances our understanding of IBR dynamic behaviors and enables the development and evaluation of IBR control functions that match or exceed the capabilities of conventional synchronous generators.

14 SOLAR ENERGY↗

ESS-DIVE Reporting Format for Location Metadata

The ESS-DIVE location metadata reporting format provides instructions and templates for reporting a minimum set of metadata for discrete point locations in geographic space represented by x, y, and z coordinates. This format was created based on a need for earth and environmental science researchers to more consistently provide metadata about locations where they conduct studies. To create the format, we incorporated elements from ESS-DIVE’s community reporting formats as well as 12 additional data standards or other data resources (e.g., databases, data systems, or repositories). In the template, we ask researchers to indicate unique locations using Location IDs and indicate hierarchies of locations through parent location IDs. We also provide additional optional fields for researchers to indicate how they measured the point location and the date and time that the location was first used as a research siteThis dataset contains support documentation for the reporting format (README.md and instructions.md), a terminology guide (guide.md), a crosswalk indicating how this reporting format relates to existing standards and data resources (Location_metadata_crosswalk.csv), a data dictionary (dd.csv), file-level metadata (flmd.csv), and the location metadata templates in both CSV (Location_metadata_template.csv) and Excel formats (Location_metadata_template.xlsx).

54 ENVIRONMENTAL SCIENCES↗

Resilience Metrics for Solar Photovoltaics

This workshop presentation proposes the development of solar photovoltaic (PV) system resilience metrics and a methodology and framework for evaluation of PV resilience metrics. PV resilience metrics are needed to establish a consistent basis for reporting, evaluation, and data collection by industry, evaluate performance of PV systems that have been subject to natural hazards, correlating resilience to system attributes, and predicting resilience for any PV system. PV resilience metrics can guide improved system design, standards, and insurance coverage. Establishing consistent metrics can foster data collection on impacts of natural hazards on PV systems.

14 SOLAR ENERGY↗

Physics-Informed Gaussian Process Regression for States Estimation and Forecasting in Power Grids

Real-time state estimation and forecasting are critical for the efficient operation of power grids. In this paper, a physics-informed Gaussian process regression (PhI-GPR) method is presented and used for forecasting and estimating the phase angle, angular speed, and wind mechanical power of a three-generator power grid system using sparse measurements. In standard data-driven Gaussian process regression (GPR), parameterized models for the prior statistics are fit by maximizing the marginal likelihood of observed data. In the PhI-GPR method, we propose to compute the prior statistics offline by solving stochastic differential equations (SDEs) governing the power grid dynamics. The short-term forecast of a power grid system dominated by wind generation is complicated by the stochastic nature of the wind and the resulting uncertainty in wind mechanical power. Here, we assume that the power grid dynamics are governed by swing equations, with the wind mechanical power fluctuating randomly in time. We solve these equations for the mean and covariances of the power grid states using the Monte Carlo simulation method. We demonstrate that the proposed PhI-GPR method can accurately forecast and estimate observed and unobserved states. For the considered problem, PhI-GPR has computational advantages over the ensemble Kalman filter (EnKF) method: In PhI-GPR, ensembles are computed offline and independently of the data acquisition process, whereas for EnFK, ensembles are computed online with data acquisition, rendering real-time forecast more challenging. We also demonstrate that the PhI-GPR forecast is more accurate than the EnKF forecast when the random mechanical wind power is non-Markovian. In contrast, the two methods produce similar forecasts for the Markovian mechanical wind power. For observed states, we show that PhI-GPR provides a forecast comparable to the standard data-driven GPR; both forecasts are significantly more accurate than the autoregressive integrated moving average (ARIMA) forecast. We also show that the ARIMA forecast is more sensitive to observation frequency and measurement errors than the PhI-GPR forecast.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Bayesian learning with Gaussian processes for low-dimensional representations of time-dependent nonlinear systems

This work presents a data-driven method for learning low-dimensional time-dependent physics-based surrogate models whose predictions are endowed with uncertainty estimates. We use the operator inference approach to model reduction that poses the problem of learning low-dimensional model terms as a regression of state space data and corresponding time derivatives by minimizing the residual of reduced system equations. Standard operator inference models perform well with accurate training data that are dense in time, but producing stable and accurate models when the state data are noisy and/or sparse in time remains a challenge. Another challenge is the lack of uncertainty estimation for the predictions from the operator inference models. Our approach addresses these challenges by incorporating Gaussian process surrogates into the operator inference framework to (1) probabilistically describe uncertainties in the state predictions and (2) procure analytical time derivative estimates with quantified uncertainties. The formulation leads to a generalized least-squares regression and, ultimately, reduced-order models that are described probabilistically with a closed-form expression for the posterior distribution of the operators. The resulting probabilistic surrogate model propagates uncertainties from the observed state data to reduced-order predictions. Furthermore, we demonstrate the method is effective for constructing low-dimensional models of two nonlinear partial differential equations representing a compressible flow and a nonlinear diffusion–reaction process, as well as for estimating the parameters of a low-dimensional system of nonlinear ordinary differential equations representing compartmental models in epidemiology.

Data-driven model reduction↗

Critical review and analysis of hydrogen safety data collection tools

The wider adoption of hydrogen in multiple sectors of the economy requires that safety and risk issues be rigorously investigated. Quantitative Risk Assessment (QRA) is an important tool for enabling safe deployment of hydrogen fueling stations and is increasingly embedded in the permitting process. QRA requires reliability data, and currently hydrogen QRA is limited by the lack of hydrogen specific reliability data, thereby hindering the development of necessary safety codes and standards [1]. Four tools have been identified that collect hydrogen system safety data: H2Tools Lessons Learned, Hydrogen Incidents and Accidents Database (HIAD), National Renewable Energy Lab's (NREL) Composite Data Products (CDPs), and the Center for Hydrogen Safety (CHS) Equipment and Component Failure Rate Data Submission Form. This work critically reviews and analyzes these tools for their quality and usability in QRA. It is determined that these tools lay a good foundation, however, the data collected by these tools needs improvement for use in QRA. Areas in which these tools can be improved are highlighted, and can be used to develop a path towards adequate reliability data collection for hydrogen systems.

08 HYDROGEN↗

Detection of Stealthy False Data Injection Attacks in Unobservable Distribution Networks

In this paper, a composite scheme is proposed for detecting stealthy data manipulation attacks on distribution system which is unobservable with standard least squares based state estimators. This technique has three stages where the process of data imputation, voltage phasor estimation and the bad data detection are carried out in a systematic manner. The proposed approach is then integrated with moving target defense strategies which perturbs the network parameters to reveal stealthy false data injection attacks. The proposed approach is tested is validated on a three-phase, unbalanced 37-node distribution system and its results are presented. It is shown that the proposed approach has the ability to accurately detect the presence of FDI attacks using limited measurements (i.e., the test system is unobservable).

Rajasekaran, James K.↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

Extraction and Analysis of Time Series Data from Building Automation Systems Using Large Language Models

Semantic schemas like Haystack 4, Brick and ASHRAE standard 223 enable the structured, standardized, and machine-readable representation of building data, facilitating interoperability, data integration, and advanced analytics. However, extracting information from these models requires specialized expertise in SPARQL and other programming languages, skills that are not commonly found among building professionals. Recent advancements in Large Language Models (LLMs), such as ChatGPT, enable the construction of queries using natural language, making it easier for individuals to interact with these systems in a manner that resembles everyday speech. However, these methods have not yet been tested on building semantic ontologies. This paper introduces a novel workflow and tool for enabling users to ask questions about a specific building's data, using natural language and receive answers automatically generated by GPT-4o. Our approach integrates semantic ontologies with advanced LLM capabilities to automate three critical steps: (1) generating SPARQL queries to retrieve time series references from ontological models, (2) extracting the corresponding time series data from the Building Automation System, and (3) performing computations and visualizations tailored to the user's query. The proposed method simplifies access to BAS data, allowing both domain experts and non-specialists to conduct sophisticated analyses without needing extensive technical knowledge of semantic web technologies. By demonstrating this pipeline, we facilitate more accessible and scalable data-driven decision-making in building operations and management.

Mulayim, Ozan Baris↗

NETL Well Integrity Workshop: Identifying Well Integrity Research Needs for Subsurface Energy Infrastructure

Wells are a critical component of subsurface energy infrastructure. Ensuring the integrity of wells as engineered pathways for the safe extraction, injection, and storage of fluids in the subsurface is key to maximizing the effectiveness and resilience of that infrastructure. Addressing well integrity issues in a technically robust manner that promotes environmental sustainability and social equity is also an important focus of the United States (U.S.) Department of Energy’s (DOE) Office of Fossil Energy and Carbon Management. Industry best practices, regulatory standards, modern monitoring data acquisition and control systems, and decades of research and development have dramatically improved the performance and reliability of wells for hydrocarbon extraction and underground injection in the oil and gas industry. Yet, important innovation is required to improve and ensure well integrity performance in engineered geologic systems where operational environments (fluid composition, temperature, pressure, and/or stress conditions) and long functional life cycles of well systems present unique challenges. Additionally, work is needed to understand and manage the long-term integrity and risks associated with legacy wells—especially those located adjacent to and presenting hazards for new subsurface activity.

02 PETROLEUM↗