Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data systems standards”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Ontologizing health systems data at scale: making translational discovery a reality

Common data models solve many challenges of standardizing electronic health record (EHR) data but are unable to semantically integrate all of the resources needed for deep phenotyping. Open Biological and Biomedical Ontology (OBO) Foundry ontologies provide computable representations of biological knowledge and enable the integration of heterogeneous data. However, mapping EHR data to OBO ontologies requires significant manual curation and domain expertise. We introduce OMOP2OBO, an algorithm for mapping Observational Medical Outcomes Partnership (OMOP) vocabularies to OBO ontologies. Using OMOP2OBO, we produced mappings for 92,367 conditions, 8611 drug ingredients, and 10,673 measurement results, which covered 68–99% of concepts used in clinical practice when examined across 24 hospitals. When used to phenotype rare disease patients, the mappings helped systematically identify undiagnosed patients who might benefit from genetic testing. By aligning OMOP vocabularies to OBO ontologies our algorithm presents new opportunities to advance EHR-based deep phenotyping.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

BrickQA: Bridging the Semantic Gap in Building Operations with Dynamic Graph Exploration

While standardized ontologies like the Brick schema address data heterogeneity in Building Automation Systems (BAS), accessing this semantic data remains a challenge as domain experts often lack the expertise to formulate complex SPARQL queries. To bridge this gap, we present BrickQA, a Large Language Model (LLM)-based framework that translates natural language into executable SPARQL queries through structured query decomposition, dynamic schema exploration, and inline validation. BrickQA utilizes an iterative reasoning agent to actively navigate graph topology through dynamic exploration actions without requiring exhaustive context injection or model fine-tuning. This approach effectively mitigates hallucinations, particularly in large-scale building knowledge graphs. Empirical evaluation on BuildingQA, a standardized benchmark, demonstrates that BrickQA significantly outperforms ReAct baselines, delivering a 0.291–0.355 absolute F1 improvement while achieving 3 × –12.7 × higher token cost-efficiency. Beyond these metrics, the framework maintains structural fidelity across heterogeneous buildings and remains resilient to ambiguous queries without requiring site-specific fine-tuning. Furthermore, a case study on operational analytics validates the framework’s capability to handle temporal and aggregation constraints, effectively transforming abstract semantic models into actionable facility management insights.1

Ko, Yun-Dam↗

Correlation approach for quality assurance of additive manufactured parts based on optical metrology

Surface topography and surface finish are two significant factors for evaluating the quality of products in additive manufacturing (AM). AM parts are fabricated layer by layer, which is quite different from traditional formative or subtractive methods. Despite rapid progress in additive manufacturing and associated optical metrology for quality control and in-situ monitoring, limited research has been conducted to investigate the reliability of 3D surface measurement data. The surface topologies scanned by multiple optical systems demonstrated significant differences due to varying sampling mechanisms, resolutions, system noises, etc. The 3D datasets should be trustworthy in order to extract parameters for quality assurance or feedback control from 3D surface measurements. In this paper, we set up new standards to evaluate the reliability of 3D surface measurement data and analyze the variation in the topographical profile. In this study, two non-contact optical methods based on Focus Variation Microscopy (FVM) and Structured Light System (SLS) were adopted to measure the surface topography of the target components. The two optical metrology systems generated two entirely different point cloud datasets. Statistical methods were applied to test the difference between the data obtained from the two systems. By using a data analytics approach for comparison, it was found that the surface roughness estimated from the point cloud data sets of FVM and SLS has no significant difference, though the point cloud data sets were completely different. This paper provides standard validation approach to evaluate the plausibility of metrology data from in-situ real-time surface analysis for process planning of AM.

36 MATERIALS SCIENCE↗

Common Information Model for Electromagnetic Transients (CIM for EMT): CRADA 533 [Abstract only}

The control and protection functions of inverter-based resources (IBR) have raised concerns with bulk system reliability. Most of the current interest lies with solar photovoltaic generation but increasing amounts of storage would pose the same risks. Newer North American Electric Reliability Corporation (NERC) guidelines call for electromagnetic transient (EMT) studies of IBR and recommend that transmission operators collect distributed energy resource (DER) data to support such modeling. IEEE Standard P2800.1 is defining tests for model parameterization, so good model data should become available from inverter vendors. (EMT studies also apply to large power transformer reliability, and transformer vendors can provide EMT models.) Utilities don’t currently have the rest of the bulk system represented for EMT studies at large scale. An International Electrotechnical Commission (IEC) standard Common Information Model (CIM) provides a way of supporting these detailed models from physical asset data, e.g., conductors, towers, transformer data sheets, control block diagrams, while avoiding software vendor lock-in. CIM-for-EMT, with proposed schema extensions and open-source converters, provides a way to exchange EMT data between organizations and tools. This project leverages Office of Electricity (OE) funding of CIM-for-EMT code base through the GridAPPS-DTM project, and of GridPACKTM (parallelized transmission solver), for interoperability testing in CIM-for-EMT. The project also leverages partner PGSTech investments in EMTP® interoperability with CIM.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Data for Autonomous Transportation Awareness: Data Exchange Use Cases, Standards, and Barriers

This report examines the critical data exchanges between automated vehicle (AV) service providers and the cities and municipalities they serve. It assists municipal authorities in navigating the often complex and real-time digital data exchanges needed to support AV mobility services, with emphasis in three areas: (1) critical safety data for broad-area situational awareness of hazards typically associated emergency dispatch or roadway work zones; (2) performance metrics of AV services that inform the quantity, quality, spatial extents, and impact on the roadway network; and (3) regulatory and policy information, particularly dynamic information that governs how AV services interact with the roadway network, with emphasis on curb space. The report reviews existing practices and emerging protocols and standards and identifies key gaps to address moving forward.

33 ADVANCED PROPULSION SYSTEMS↗

The U.S. Agrivoltaic Shading Tool: A National-Scale Interface for Modeling Light and Shade Patterns in Ten Common Agrivoltaic Configurations

Agrivoltaic systems are dual-use configurations that co-locate agriculture and photovoltaic (PV) infrastructure and require careful design to balance crop performance and energy generation. A critical element of agrivoltaic design is the spatial and temporal distribution of irradiance and shade within and around PV arrays. To support research, planning, and stakeholder decision-making, we introduce the U.S. Agrivoltaic Shading Tool, a novel web-based application that delivers high-resolution irradiance and photosynthetically active radiation (PAR) modeling for ten standardized PV configurations across the conterminous United States. The tool leverages the National Laboratory of the Rockies (NLR) System Advisor Model (SAM) to perform detailed irradiance simulations, using meteorological data from the National Solar Radiation Database (NSRDB). Outputs include seasonal, monthly, weekly, and diurnal patterns of available sunlight, amount of shade, irradiance, and PAR at ground level within agrivoltaic system footprints. For a user's selected location, these results are visualized through interactive visualizations, heatmaps, and time-series plots, designed to be accessible to both technical and non-technical users. In addition to facilitating rapid spatial exploration of agrivoltaic light environments, the tool will offer seamless integration with the InSPIRE Agrivoltaics Design and Analysis Model (ADAM). This optional workflow will allow users to port selected site and configuration parameters into a more advanced modeling environment for further customization of structural layouts, crop-system compatibility, power generation, and technoeconomic performance. Finally, to promote open science, the entire dataset will be hosted and available for open access through the OpenEI platform. By standardizing and disseminating high-quality irradiance data and design tools, the U.S. Agrivoltaic Shading Tool supports a wide range of users, including researchers, landowners, energy developers, and policymakers, in evaluating the agronomic and energetic feasibility of agrivoltaic systems across the United States.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Electrical Energy Storage Data Submission Guidelines, Version 2

Energy storage technologies are positioned to play a substantial role in power delivery systems. They have the potential to serve as an effective new resource to maintain reliability and allow for increased penetration of renewable energy. However, because of their relative infancy, there is a lack of knowledge about how these resources truly operate over time. A data analysis can help ascertain the operational and performance characteristics of these emerging technologies. Rigorous testing and a data analysis are important for all stakeholders to ensure a safe, reliable system that performs predictably on a macro level. Standardizing testing and analysis approaches to verify the performance of energy storage devices, equipment, and systems when integrating them into the grid will improve the understanding and benefit of energy storage over time from technical and economic vantage points. Demonstrating the life-cycle value and capabilities of energy storage systems begins with the data that the provider supplies for the analysis. After a review of energy storage data received from several providers, some of these data have clearly shown to be inconsistent and incomplete, raising the question of their efficacy for a robust analysis. This report reviews and proposes general guidelines, such as sampling rates and data points, that providers must supply for a robust data analysis to take place. Consistent guidelines are the basis of a proper protocol and ensuing standards to (1) reduce the time that it takes for data to reach those who are providing the analysis; (2) allow them to better understand the energy storage installations; and (3) enable them to provide a high-quality analysis of the installations. The report is intended to serve as a starting point for what data points should be provided when monitoring. Readers are encouraged to use the guidance in the report to develop specifications for new systems, as well as enhance current efforts to ensure optimal storage performance. As battery technologies continue to advance and the industry expands, the report will be updated to remain current.

25 ENERGY STORAGE↗

Unified architecture for data-driven metadata tagging of building automation systems

This article presents a Unified Architecture (UA) for automated point tagging of Building Automation System (BAS) data, based on a combination of data-driven approaches. Advanced energy analytics applications—including fault detection and diagnostics and supervisory control—have emerged as a significant opportunity for improving the performance of our built environment. Effective application of these analytics depends on harnessing structured data from the various building control and monitoring systems, but typical BAS implementations do not employ any standardized metadata schema. While standards such as Project Haystack and Brick Schema have been developed to address this issue, the process of structuring the data, i.e., tagging the points to apply a standard metadata schema, has, to date, been a manual process. This process is typically costly, labor-intensive, and error-prone. In this work we address this gap by proposing a UA that automates the process of point tagging by leveraging the data accessible through connection to the BAS, including time-series data and the raw point names. The UA intertwines supervised classification and unsupervised clustering techniques from machine learning and leverages both their deterministic and probabilistic outputs to inform the point tagging process. Furthermore, we extend the UA to embed additional input and output data-processing modules that are designed to address the challenges associated with the real-time deployment of this automation solution. We test the UA on two datasets for real-life buildings: (i) commercial retail buildings and (ii) office buildings from the National Renewable Energy Laboratory (NREL) campus. We report the proposed methodology correctly applied 85–90% and 70–75% of the tags in each of these test scenarios, respectively for two significantly different building types used for testing UA's fully-functional prototype. The proposed UA, therefore, offers promising approach for automatically tagging BAS data as it reaches close to 90% accuracy. Further building upon this framework to algorithmically identify the equipment type and their relationships is an apt future research direction to pursue.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Generation and validation of comprehensive synthetic weather histories using auto-regressive moving-average models

As energy system design moves to more complex methods of optimization including machine learning there is a significant need for more weather data than is available. One method to solve this is using synthetic data models such as the auto-regressive moving-average (ARMA) model which has been frequently utilized to create such data. This paper looks at extending the ARMA algorithm to generate solar components through the use of clearsky detrending, maintaining vector relationships and by leveraging physical relationships. The method for the creation of entirely synthetic weather data files including key weather variables for energy system analysis is presented. Furthermore, a detailed comparison of energy system simulations utilizing both real and synthetic data is made using NREL’s System Advisor Model. Whilst good agreement is made for the solar variables, and other weather variables, ARMA methods often fail to capture the standard deviation and skew of annual weather distributions. Vector-ARMA is shown to maintain correlations between variables and thus generate data sets that perform similarly in energy system design. Here, it is finally shown that the ARMA method fails to preserve day-today correlations in weather variables and thus over-predicts optimal energy storage by 21% for a residential solar application.

42 ENGINEERING↗

A Privacy-Preserving Strategy for the Trust Layer of the Energy Grid of Things Distributed Energy Resource Management System

Emergent from the shadows of the traditional grid flaws, the Smart Grid (SG) idea was born and led by government mandates toward cleaner energy production. The SG represents the next generation of electricity distribution systems that subsume recent technological innovations. It uses digital communication between its components and entities to attain more automation, self-sufficiency, and reliability. Unfortunately, this relatively new concept is not flawless; the intrinsic reliance on increased digital communication spreads open attack paths for adversaries. Therefore, finding solutions that address information exchange vulnerabilities has become imperative. The Energy Grid of Things (EGoT) is Portland State University’s (PSU’s) implementation of a Distributed Energy Resource Management System (DERMS). The EGoT DERMS requires access to customers’ information to achieve operational objectives. The system’s access to customers’ information needs to be restricted such that it does not violate customers’ privacy. Applying privacy protection models such as K-anonymity to EGoT DERMS sub-components safeguards that privacy. This thesis work proposes a strategy to ensure communication in the EGoT DERMS is privacy-preserving and secure. Specifically, it provides an approach to applying the Mondrian Algorithm to ensure data within the system excludes Personally Identifiable Information (PII) and provides means for securing the communication according to industry standards (IEEE 2030.5). Results suggest that the generalization hierarchy derived for the EGoT DERMS exhibits an Identical Generalization Hierarchy structure. Guarantees of sameness manifested in the test feeder topology would not hold in real-world scenarios. This thesis work proposes a strategy to ensure communication in the EGoT DERMS is privacy-preserving and secure. Specifically, it provides an approach to applying the Mondrian Algorithm to ensure data within the system excludes Personally Identifiable Information (PII) and provides means for securing the communication according to industry standards (IEEE 2030.5). Results suggest that the generalization hierarchy derived for the EGoT DERMS exhibits an Identical Generalization Hierarchy structure. Guarantees of sameness manifested in the test feeder topology would not hold in real-world scenarios.

Alsiad, Mohammed↗

Baseline Climate Variables for Earth System Modelling

The Baseline Climate Variables for Earth System Modelling (ESM-BCVs) are defined as a list of 135 variables which have high utility for the evaluation and exploitation of climate simulations. The list reflects the most frequently used elements of the Coupled Model Intercomparison Project Phase 6 (CMIP6) archive. Successive phases of CMIP have supported strong results in science and substantially influence international climate policy formulation. This paper responds to both interest in exploiting CMIP data standards in a broader range of climate modelling activities and a need to achieve greater clarity about the significance and intention of variables in the CMIP Data Request. As Earth system modelling archives grow in scale and complexity, there are emerging problems associated with weak standardisation at the variable collection level. That is, there are good standards covering how specific variables should be archived, but this paper fills a gap in the standardisation of which variables should be archived. The ESM-BCV list is intended as a resource for ESM intercomparison projects (MIPs) developing requests to enable greater consistency among MIPs and as a reference for modelling centres to enhance consistency within MIPs. Provisional planning for the CMIP7 Data Request exploits the ESM-BCVs as a core element. The baseline variable list includes 98 variables which have modest or minor data volume footprints and could be generated systematically when simulations are produced and archived for exploitation by the World Climate Research Programme (WCRP) community. A further 35 variables are classed as “high volume” and are only suitable for production when the resource implications are justified.

Juckes, Martin [University of Oxford (United Kingd↗

Impact of duration and missing data on the long-term photovoltaic degradation rate estimation

Accurate quantification of photovoltaic (PV) system degradation rate (R D ) is essential for lifetime yield predictions. Although R D is a critical parameter, its estimation lacks a standardized methodology that can be applied on outdoor field data. The purpose of this paper is to investigate the impact of time period duration and missing data on R D by analyzing the performance of different techniques applied to synthetic PV system data at different linear R D patterns and known noise conditions. The analysis includes the application of different techniques to a 10-year synthetic dataset of a crystalline Silicon PV system, with emulated degradation levels and imputed missing data. Here, the analysis demonstrated that the accuracy of ordinary least squares (OLS), year-on-year (YOY), autoregressive integrated moving average (ARIMA) and robust principal component analysis (RPCA) techniques is affected by the evaluation duration with all techniques converging to lower R D deviations over the 10-year evaluation, apart from RPCA at high degradation levels. Moreover, the estimated R D is strongly affected by the amount of missing data. Filtering out the corrupted data yielded more accurate R D results for all techniques. It is proven that the application of a change-point detection stage is necessary and guidelines for accurate R D estimation are provided.

14 SOLAR ENERGY↗

Measured and projected beam backgrounds in the Belle II experiment at the SuperKEKB collider

The Belle II experiment at the SuperKEKB electron–positron collider aims to collect an unprecedented data set of 50 ab -1 to study CP -violation in the B -meson system and to search for Physics beyond the Standard Model. SuperKEKB is already the world’s highest-luminosity collider. In order to collect the planned data set within approximately one decade, the target is to reach a peak luminosity of 6 x 10 35 cm −2 s −1 by further increasing the beam currents and reducing the beam size at the interaction point by squeezing the betatron function down to β $^*_y$ = 0.3 mm. To ensure detector longevity and maintain good reconstruction performance, beam backgrounds must remain well controlled. We report on current background rates in Belle II and compare these against simulation. We find that a number of recent refinements have significantly improved the background simulation accuracy. Finally, we estimate the safety margins going forward. We predict that backgrounds should remain high but acceptable until a luminosity of at least 2.8 x 10 35 cm −2 s −1 is reached for β $^*_y$ = 0.6 mm. At this point, the most vulnerable Belle II detectors, the Time-of-Propagation (TOP) particle identification system and the Central Drift Chamber (CDC), have predicted background hit rates from single-beam and luminosity backgrounds that add up to approximately half of the maximum acceptable rates.

Detector background↗

Tsdat: An Open-Source Data Standardization Framework for Marine Energy and Beyond: Preprint

Many organizations are tasked with the collection and processing of large quantities of data from various measurement devices. Data reported from these sources are often not interoperable with datasets and software used by analysts and other organizations in the same field, introducing barriers for collaboration on large-scale projects. This poses a particular problem for cross-device comparisons and machine learning applications. To address these challenges, the open source Time-Series Data Pipelines (Tsdat) software was developed by a joint collaboration between Pacific Northwest National Laboratory, the National Renewable Energy Laboratory, and Sandia National Laboratories to facilitate collaboration and accelerate advancements in the Marine Energy domain through the development of an open-source ecosystem of tools. This paper will describe the Tsdat software and the data standards within which the framework operates. A beta version of the framework has been released and is currently being used by several projects in marine energy, wind energy, and building energy systems.

big data↗

Path Forward: Materials Data Modernization for ASME Codes and Standards in the Artificial Intelligence Era

Development of the ASME Materials Properties Database was initiated in the early 2010s to support the ASME Codes and Standards. As information technologies advance at an accelerated pace with the artificial intelligence era on the horizon, the ASME Materials Properties Database must be further modernized from a database to a knowledgebase to ride the wave of digital information revolution and effectively support the ASME Codes and Standards in the new era.This paper is intended to provide an overview of the ASME Materials Properties Database and discuss a roadmap for its future development to facilitate understanding of and participation from different sectors of the Codes and Standards community. It first reviews the basic concepts of data, information, knowledge, database, and database system; as well as the pros and cons in different types of data management, and then discusses the path forward for a desired evolution of the database into a self-explanatory and machine-readable knowledgebase that is consistent with human cognitive processes for the Codes and Standards development and furthermore provides resources for data processing and analysis to reach an eventual goal of streamlining the Codes and Standards development from the initial inquiry, throughout data submission, analysis, …, to Codes and Standards rule establishment for final publication.

Ren, Weiju↗

Making a Water Data System Responsive to Information Needs of Decision Makers

Evidence-based environmental management requires data that are sufficient, accessible, useful and used. A mismatch between data, data systems, and data needs for decision making can result in inefficient and inequitable capital investments, resource allocations, environmental protection, hazard mitigation, and quality of life. In this paper, we examine the relationship between data and decision making in environmental management, with a focus on water management. We focus on the concept of decision-driven data systems —data systems that incorporate an assessment of decision-makers' data needs into their design. The aim of the research was to examine the process of translating data into effective decision making by engaging stakeholders in the development of a water data system. Using California's legislative mandate for state agencies to integrate existing water and other environmental data as a case study, we developed and applied a participatory approach to inform data-system design and identify unmet data needs. Using workshops and focused stakeholder meetings, we developed 20 diverse use cases to assess data sources, availability, characteristics, gaps, and other attributes of data used for representative decisions. Federal and state agencies made up about 90% of the data sources, and could readily adapt to a federated data system, our recommended model for the state. The remaining 10% of more-specialized data, central to important decisions across multiple use cases, would require additional investment or incentives to achieve data consistency, interoperability, and compatibility with a federated system. Based on this assessment, we propose a typology of different types of data limitations and gaps described by stakeholders. We also propose technical, governance, and stakeholder engagement evaluation criteria to guide planning and building environmental data systems. Data-system governance involving both producers and users of data was seen as essential to achieving workable standards, stable funding, convenient data availability, resilience to institutional change, and long-term buy-in by stakeholders. Our work provides a replicable lesson for using decision-maker and stakeholder engagement to shape the design of an environmental data system, and inform a technical design that addresses both user and producer needs.

Cantor, Alida↗