Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “system metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

A Data Science and Machine Learning Platform Supporting Large Particle Accelerator Control and Diagnostics Applications Final Report: SBIR Initial Phase II DE-SC0022583

The Machine Learning Data Platform (MLDP) is a product providing full-stack support for data science, Machine Learning, and Artificial Intelligence (ML/AI) applications at particle accelerator and large experimental physics facilities. It supports ML/AI applications from front-end, high-speed acquisition of heterogeneous, time-series data, through data archiving and management, to back-end analysis. The MLDP embodies a “data-science ready” platform for data analysis and ML/AI applications in diagnosis, modelling, control, and optimization of these facilities. It provides data scientists and applications a consistent, datacentric interface to archive data standardizing implementation and deployment of ML/AI algorithms to different operations configurations within the same facility, or between facilities. Being an open-source, public-domain project, the MLDP is intended for broadest possible impact by increasing accessibility and minimizing the required expertise for installation and operation. The MLDP can also be deployed at user facilities for experimental data collection, archiving, and analysis. It is capable of acquisition and archiving of heterogeneous data from experimental equipment (e.g., images, arrays, structures, etc.) along with system hardware configurations (e.g., scalars, tables), control system process variables, and any metadata required for provenance. Thus, the MLDP can manage experimental data through its entire lifecycle, from acquisition and archiving, through analysis and investigation, to release and final publication.

43 PARTICLE ACCELERATORS↗

Towards Next-Generation Urban Decision Support Systems through AI-Powered Construction of Scientific Ontology Using Large Language Models—A Case in Optimizing Intermodal Freight Transportation

The incorporation of Artificial Intelligence (AI) models into various optimization systems is on the rise. However, addressing complex urban and environmental management challenges often demands deep expertise in domain science and informatics. This expertise is essential for deriving data and simulation-driven insights that support informed decision-making. In this context, we investigate the potential of leveraging the pre-trained Large Language Models (LLMs) to create knowledge representations for supporting operations research. By adopting ChatGPT-4 API as the reasoning core, we outline an applied workflow that encompasses natural language processing, Methontology-based prompt tuning, and Generative Pre-trained Transformer (GPT), to automate the construction of scenario-based ontologies using existing research articles and technical manuals of urban datasets and simulations. From these ontologies, knowledge graphs can be derived using widely adopted formats and protocols, guiding various tasks towards data-informed decision support. The performance of our methodology is evaluated through a comparative analysis that contrasts our AI-generated ontology with the widely recognized pizza ontology, commonly used in tutorials for popular ontology software. We conclude with a real-world case study on optimizing the complex system of multi-modal freight transportation. Our approach advances urban decision support systems by enhancing data and metadata modeling, improving data integration and simulation coupling, and guiding the development of decision support strategies and essential software components.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

DAISY Variant and Tether Tests, Admirality Inlet, WA

Acoustic data and metadata from Drifting Acoustic Instrumentation SYstem (DAISY) testing in Admiralty Inlet (connecting Puget Sound to the Strait of San Juan de Fuca) in July 2022. Tests focused on occurrences of flow noise for three hydrophone package variants and on the potential for alternative tether materials.

16 TIDAL AND WAVE POWER↗

DAISY Acoustic Measurements in Agate Pass, WA

Acoustic data and metadata from Drifting Acoustic Instrumentation SYstem (DAISY) testing in Agate Pass (separating the north end of Bainbridge Island and the Kitsap Peninsula in Puget Sound), WA in April 2022. The goal was to characterize radiated noise from a cross-flow turbine deployed from a moored vessel. As discussed in the accompanying report, sound produced by the turbine was below the ambient nose floor at the surveyed ranges.

16 TIDAL AND WAVE POWER↗

WHONDRS River Corridor Surface Water Metabolites and Geochemistry from Global Sites

This dataset supports a broader study examining the character of organic matter that may be delivered to subsurface sediments via hydrologic exchange. To implement the global survey, free stream sampling kits were provided to interested volunteers throughout the world. Samples were collected with minimal constraints in terms of location, but following strict protocols, and shipped for metabolomic analysis via Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS). In addition, basic geochemistry analyses (e.g., dissolved organic matter concentration) were conducted, standardized photos of each field system were taken, and extensive metadata were captured. Sampling began in 2018 and is ongoing as of 2025. This dataset is comprised of one folders of field photos, one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data, and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) international generic sample number (IGSN) mapping file; (6) field protocol; and (7) a subfolder with sample data. The sample data subfolder contains (1) surface water dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) methods codes; (3) surface water FTICR methods; and (4) a subfolder of 12 Tesla (12T) FTICR-MS data. This folder contains three subfolders, one containing the.xml files, one containing the CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .xml, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, or .png. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About.

Biogeochemistry↗

AI and ML Applications for PV Reliability and System Performance

This poster discusses AI and ML topics in PV reliability and system performance. In particular, automated metadata extraction and QA for fielded solar installations is covered for the PV Fleets Project. Additionally, statistical learning topics for the PVInsight Project are addressed, as well as development of the PV Validation Hub.

algorithm↗

SysCaps (Language Interfaces for Simulation Surrogates of Complex Systems) [SWR-24-97]

You've found the official code repository for the paper "SysCaps: Language Interfaces for Simulation Surrogates of Complex Systems," presented at the Foundation Models for Science: Progress, Opportunities, and Challenges workshop at NeurIPS 2024. Our paper conjectures that interfaces (both text templates as well as conversational) makes interacting with simulation surrogate models for complex systems more intuitive and accessible for both non-experts and experts. "System captions", or SysCaps, are text-based descriptions of systems based on information contained in simulation metadata. Our paper's goal is to train multimodal regression models that take text inputs (SysCaps) and timeseries inputs (exogenous system conditions such as hourly weather) and regress timeseries simulation outputs (e.g. hourly building energy consumption). The experiments in our paper with building and wind farm simulators, which can be reproduced using this codebase, aim to help us understand whether a) accurate regression in this setting is possible and b) if so, how well can we do it. Paper: https://arxiv.org/abs/2405.19653

Emami, Patrick↗

Explainable discrepancy checker and diagnosis for digital Twin-based supervisory control system

By virtually representing a physical object and process, a digital twin (DT) enables optimal autonomous operations by combining classical and novel frameworks in sensors, state predictions, and multi-input/multi-output systems. A DT’s values depend on how well models estimate quantities of interest and on how uncertainty is handled. Moreover, DTs often combine physics-based and data-driven models with mixed fidelities, where classical uncertainty quantification (UQ) struggles with many sources of uncertainty and real-time constraints. Here, this work presents a UQ-based discrepancy checking and diagnosis tool for a DT-based supervisory control system. The tool is developed using metadata from an automated DT development process to learn correlations between sources of uncertainties and outcomes. During operation, it compares predictions with measurements, attributes discrepancies to dominant sources, and recommends parameter and configuration updates. We verify the workflow on a synthetic temperature-control problem and deploy it on a virtual Thermal Energy Delivery System, reducing mismatch and improving control robustness.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Uncovering I/O demands on HPC platforms: Peeking under the hood of Santos Dumont

High-Performance Computing (HPC) platforms are required to solve the most diverse large-scale scientific problems in various research areas, such as biology, chemistry, physics, and health sciences. Researchers use a multitude of scientific softwares, which have different requirements. These include input and output operations, which directly impact performance due to the existing difference in processing and data access speeds. Thus, supercomputers must efficiently handle mixed workload when storing data from the applications. Understanding the set of applications and their performance running in a supercomputer is paramount to understanding the storage system's usage, pinpointing possible bottlenecks, and guiding optimization techniques. This research proposes a methodology and visualization tool to evaluate a supercomputer's data storage infrastructure's performance, taking into account the diverse workload and demands of the system over a long period of operation. As a study case, we focus on the Santos Dumont supercomputer, identifying inefficient usage, problematic performance factors, and providing guidelines on how to tackle those issues.

97 MATHEMATICS AND COMPUTING↗

Integrating AEAD Ciphers into Software-Defined-Storage Systems

The use of software-defined storage (SDS) systems to store sensitive data is becoming increasingly prevalent. However, these systems primarily implement security measures to ensure the confidentiality and availability of stored data, with limited consideration for the protection of its integrity. This paper outlines why this is a harmful development, as well as how integrity-protecting measures can be included into SDS systems. To demonstrate the practical challenges and opportunities of such measures, we integrated "authenticated encryption with associated data" (AEAD) ciphers into the widely used SDS system Ceph, specifically, into its block storage interface, to secure the integrity of stored data and metadata. Ultimately, we identify the characteristics that an SDS system should possess to adopt our methodology.

Mohren, David [University of New Brunswick, Canada↗

Towards Semantic Search in Building Sensor Data

This paper presents a search engine system for sensor time series data and metadata in the context of building management. It takes natural language queries as input, retrieves sensor time series data, ranks them with respect to their relevance to a given query, and visualizes the time series as search results. In addition, the system allows users to interact with the search results: they can define events of interest in the visualized results and search across sensor data for similar events, i.e., the search by example scheme. Quantitative evaluations and user studies demonstrate the value of this system for managing building sensor data.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Data about data – when, why and how metadata can support the digital plant

A structured approach for recording data quality and contextual information about how and why a signal exists – i.e. metadata – is central to interpret and use sensor data correctly. This is becoming increasingly important with the global trend with data-driven applications such as digital twins and AI-models. But a structured metadata collection and organization of sensor data is not routine in most plants, which can result in lost information and missed opportunities to make use of the investments made in the data collection. Therefore, the IWA task group on Metadata Collection and Organization in wastewater resource recovery systems (MetaCO) was initiated in 2020 and recently delivered the IWA scientific and technical report number 31. The report gives and in-depth description about metadata in water resources recovery facilities (WRRFs) and is available as open access at IWA publishing. The report is the outcome of the collaboration between more than 80 water professionals with the intention to serve WRRF data users with a guide on how to structure and make use of metadata throughout the data pipeline in order to maximize the value of sensor data.

Alferes, Janelcy [VITO, Belgium]↗

An overview of data tools for representing and managing building information and performance data

Building information modeling (BIM) has been widely adopted for representing and exchanging building data across disciplines during building design and construction. However, BIM's use in the building operation phase is limited. With the increasing deployment of low-cost sensors and meters, as well as affordable digital storage and computing technologies, growing volumes of data have been collected from buildings, their energy services systems, and occupants. Such data are crucial to help decision makers understand what, how, and when energy is consumed in buildings—a critical step to improving building performance for energy efficiency, demand flexibility, and resilience. However, practical analyses and use of the collected data are very limited due to various reasons, including poor data quality, ad-hoc representation of data, and lack of data science skills. To unlock value from building data, there is a strong need for a toolchain to curate and represent building information and performance data in common standardized terminologies and schemas, to enable interoperability between tools and applications. This study selected and reviewed 24 data tools based on common use cases of data across the building life cycle, from design to construction, commissioning, operation, and retrofits. The selected data tools are grouped into three categories: (1) data dictionary or terminology, (2) data ontology and schemas, and (3) data platforms. The data are grouped into ten typologies covering most types of data collected in buildings. This study resulted in five main findings: (1) most data representation tools can represent their intended data typologies well, such as Green Button for smart meter data and Brick schema for metadata of sensors in buildings and HVAC systems, but none of the tools cover all ten types of data; (2) there is a need for data schemas to represent the basis of design data and metadata of occupant data; (3) standard terminologies such as those defined in BEDES are only adopted in a few data tools; (4) integrating data across various stages in the building life cycle remains a challenge; and (5) most data tools were developed and maintained by different parties for different purposes, their flexibility and interoperability can be improved to support broader use cases. Finally, recommendations for future research on building data tools are provided for the data and buildings community based on the FAIR principles to make data Findable, Accessible, Interoperable, and Reusable.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Development of a Discrepancy Checker for the Digital Twin in a Supervisory Control System for a Thermal Energy Delivery System

Defined as a virtual representation of a physical object, process, or service, and used to support real-world decision-making, a digital twin (DT) can be utilized to combine classical and novel frameworks in sensors, state predictions, and multi-input/multi-output systems, and to enable optimal autonomous operations. However, a DT’s usefulness largely depends on its ability to adequately mirror the state of its physical counterpart, and this adequacy should be reflected by the level of uncertainty in the underlying simulation models when estimating and predicting quantities of interest (QOIs). Moreover, simulation models in a DT may involve multiple fidelities of representations—ranging from physics-based models to data-driven ones—but classical uncertainty quantification (UQ) methods struggle to handle numerous uncertainty sources, nor are they designed for real-time applications. This work presents a UQ-based discrepancy checking and diagnosis tool for a DT-based supervisory control system applied to a thermal energy delivery system (TEDS) at Idaho National Laboratory. The discrepancy checker was developed using metadata from an automated DT development process, and these metadata included different combinations of physical model forms and model parameters, training data and hyperparameters for surrogate models, and design parameters for supervisory control systems. Next, correlations between the uncertainty results and the metadata were established and then applied to the DT operations. The discrepancy checker evaluates the discrepancies between model predictions from virtual and sensor measurements and backtraces them to the corresponding major sources of uncertainty. The discrepancy checker showed reasonable performance in detecting discrepancies and diagnosing sources of uncertainty in testing scenarios.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Titan on the Red (advanced AI / ML system) [Slides]

Born out of a need to make LANL Weapons Program archival material available to scientists and engineers. Material is the end result of decades of consolidation of mini-libraries and mini archives at LANL. Latest consolidation brought together LANL’s digital archives and physical archives. This houses over 75 years of nuclear weapons research, designs, procedures, videos, photos, and other reports. For the Titan on the Red machine learning project, the system must be able to automatically extract metadata from digitized documents, perform natural-language search, enforce security classification and NTK protocols, be expandable to ingest various data stores (Online Vault, PDMLink, shared drives, SharePoint, etc.), and utilize commercial, public domain, and LANL ontologies.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

DAISY: A Rapid Approach to Evaluating Marine Energy Converter Sound (Final Technical Report)

This project’s objective was to improve the quality of acoustic information about marine energy converters that could be collected from groups of drifting hydrophones, while reducing the costs of deployment and data analysis. This was achieved through technology development addressing four focus areas: (1) minimizing flow-noise and self-noise, (2) integrating metadata streams into a single data acquisition system, (3) developing post-processing routines to facilitate rapid data review, and (4) enabling objective identification of marine energy converter sound against a backdrop of ambient noise using time-delay-of-arrival localization.

16 TIDAL AND WAVE POWER↗

Evaluation of Best Practices in Mitigating Startup Costs on Leadership-Class Supercomputers

Supercomputers at Department of Energy (DOE) National Laboratories face a widening range of workloads, from traditional modeling and simulation to Artificial Intelligence model training or complex multi-stage workflows, and beyond. At DOE Leadership Computing Facilities like the Oak Ridge Leadership Computing Facility (OLCF), these workloads demand concurrent access to large portions of the supercomputer’s resources. Launching a job across massive supercomputers is challenging from the start; the file system struggles with a large backlog of metadata requests as tens of thousands of processes read thousands of the same files, and the compute job cannot start until this is completed. There are multiple existing approaches to calm this metadata storm, ranging from vendor-developed tools like sbcast to National Laboratory-developed tools like Spindle and Copper. In this paper, we benchmark and discuss three common approaches to improving compute job launch latencies on Frontier: Slurm’s sbcast tool, Spindle, and Copper. We evaluate these tools by measuring the launch latencies of four workloads: OSU Microbenchmark’s osu_init, Pynamic, Python import mpi4py, and Python import torch. We provide discussion of the results, highlighting data that meet expectations and that do not meet expectations.

Hagerty, Nick [ORNL] (ORCID:0000000330014414)↗