Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “system metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

TomoPyUI : a user-friendly tool for rapid tomography alignment and reconstruction

The management and processing of synchrotron and neutron computed tomography data can be a complex, labor-intensive and unstructured process. Users devote substantial time to both manually processing their data ( i.e. organizing data/metadata, applying image filters etc. ) and waiting for the computation of iterative alignment and reconstruction algorithms to finish. In this work, we present a solution to these problems: TomoPyUI , a user interface for the well known tomography data processing package TomoPy . This highly visual Python software package guides the user through the tomography processing pipeline from data import, preprocessing, alignment and finally to 3D volume reconstruction. The TomoPyUI systematic intermediate data and metadata storage system improves organization, and the inspection and manipulation tools (built within the application) help to avoid interrupted workflows. Notably, TomoPyUI operates entirely within a Jupyter environment. Herein, we provide a summary of these key features of TomoPyUI , along with an overview of the tomography processing pipeline, a discussion of the landscape of existing tomography processing software and the purpose of TomoPyUI , and a demonstration of its capabilities for real tomography data collected at SSRL beamline 6-2c.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

The 2021 Blind PVPMC Modeling Intercomparison

This document provides the instructions for participating in the 2021 blind photovoltaic (PV) modeling intercomparison organized by the PV Performance Modeling Collaborative (PVPMC). It describes the system configurations, metadata, and other information necessary for the modeling exercise. The practical details of the validation datasets are also described. The datasets were published online in open access in April 2023, after completing the analysis of the results.

14 SOLAR ENERGY↗

Distributed Real-time Plume Monitoring for Deep Sea Mineral Extraction​

In the emerging industry of deep-sea mining for minerals and deposits (e.g. polymetallic nodules for nickel, cobalt, copper, and manganese), more data is required to understand the effects of sediment plume generation and predict the distribution of disturbed sediment. There are two main sources of plume generation, the first being at the active mining site where the “collector” directly removes the top layer of the sea floor. The other is the “midwater plume” consisting of unwanted sediment that was collected during extraction that is pumped back into the aphotic zone. The vast majority of plume generation is caused by the collector, causing detrimental and long-lasting impacts on seafloor ecosystems due to the lack of wave activity or strong currents at the sea floor. Therefore, it is crucial to invest in the infrastructure to support the study and constant monitoring over a large area of the sea floor where plume generation is present. Due to the limited number of usable channels and power requirements, current subsea wireless communications technologies are not well suited to instrumenting the large areas of the sea floor needed to monitor plume migration. The scope of this effort is to transition experimental demonstrations of high-bandwidth, full-duplex scalable underwater laser communications to the seafloor in an open ocean environment. Specifically tackling challenges associated with the dynamic nature of the subsea world, including but not limited to, deployment logistics, sustainability, and range. The goal is to enable the internet of underwater things for deep sea industries by broadening the capabilities of subsea communications. By using high-precision laser transmitters, many of the challenges current subsea optical systems face can be circumvented, such as power consumption, interference, and bandwidth limitations. This approach lends itself to wireless interlinking multi-node networks, in series or parallel, facilitating the implementation of a wide array of sensor types. This interlinking allows all the data gathered from the network to be processed through a single hardline uplink to the surface, lowering the complexity required for near real-time data processing. Additionally, the laser control systems produce metadata that can be used to help characterize the water column between the nodes. Combining data from various sensors such as turbidity, temperature, current velocity with metadata such as beam attenuation and deflection can produce a high-resolution model of sea floor conditions around an active mining zone. The resulting near real-time model can be used to optimize location and flow rate of the mining operation to minimize and quantify the environmental impact.

Mons, Ishan↗

Application-Driven Creation of Building Metadata Models with Semantic Sufficiency

Semantic metadata models such as Brick, RealEstateCore, Project Haystack, and BOT promise to simplify and lower the cost of developing software for smart buildings, enabling the widespread deployment of energy efficiency applications. However, creating these models remains a challenge. Despite recent advances in creating models from existing digital representations like point labels and architectural models, there is still no feedback mechanism to ensure that the human input to these methods results in a model that can actually support the desired software. In this paper, we introduce the notion of semantic sufficiency, a practical principle for semantic metadata model creation that asserts that a model is "finished" when it contains the metadata necessary to support a given set of applications. To support semantic sufficiency, we design a standard representation for capturing application metadata requirements and a templating system for generating common metadata model components with limited user input. We then construct an iterative model creation workflow that integrates metadata requirements to direct the model creation effort, and present several novel optimizations that increase the model utility while minimizing the effort by a human operator. These new abstractions for model creation and validation lower model development costs and ensure the utility of the resulting model, thus facilitating the adoption of intelligent building applications.

applications↗

Machine Learning for Automated Metadata Assignment in Buildings: Cooperative Research and Development (Final Report, CRADA Number CRD-18-00767)

RealTerm Energy and NREL have identified a shared vision to evaluate opportunities to facilitate the organization and assignment of metadata to building control system (BCS) data via industry-informed machine learning (ML). Manual metadata assignment is labor intensive and costly, slowing down any Energy Management and Information System (EMIS) deployment in the building space. This project aims to develop methodologies to accurately assign this metadata and significantly decrease the level of effort associated with deploying EMIS. The objective of this project is to identify/design methodologies to assign metadata to HVAC control points automatically. The identified methodologies will be programmed in analytics algorithms so they can ingest a list of points and produce a detailed tagging following the Haystack classification nomenclature. To validate the efficacy of each methodology, tagging results will be compared utilizing a list of points extracted from RealTerm's building database-as well as data extracted from the NREL campus via the Intelligent Campus program-enabling testing against large datasets with real world challenges. The developed methodologies may leverage building manager/operator input on a limited basis to add context to the classifying algorithms. The partnership aims to advance global efforts in areas related to the DOE missions through improving operational performance of commercial buildings. It is well documented that buildings fall out of commission after they are occupied, wasting significant energy and incurring associated costs simply due to poor operational performance. Emerging EMIS technologies that perform continuous commissioning help to address this issue, yet integration of these systems can be labor intensive both for the technology vendor and the building owner/operator. This project will enable more efficient and cost-effective analytics for buildings, enabling improvement in building operations at lower cost points.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

ESS-DIVE Reporting Format for Dataset Package Metadata

ESS-DIVE’s (Environmental Systems Science Data Infrastructure for a Virtual Ecosystem) dataset metadata reporting format is intended to compile information about a dataset (e.g., title, description, funding sources) that can enable reuse of data submitted to the ESS-DIVE data repository. The files contained in this dataset include instructions (dataset_metadata_guide.md and README.md) that can be used to understand the types of metadata ESS-DIVE collects. The data dictionary (dd.csv) follows ESS-DIVE’s file-level metadata reporting format and includes brief descriptions about each element of the dataset metadata reporting format. This dataset also includes a terminology crosswalk (dataset_metadata_crosswalk.csv) that shows how ESS-DIVE’s metadata reporting format maps onto other existing metadata standards and reporting formats.Data contributors to ESS-DIVE can provide this metadata by manual entry using a web form or programmatically via ESS-DIVE’s API (Application Programming Interface). A metadata template (dataset_metadata_template.docx or dataset_metadata_template.pdf) can be used to collaboratively compile metadata before providing it to ESS-DIVE.Since being incorporated into ESS-DIVE’s data submission user interface, ESS-DIVE’s dataset metadata reporting format, has enabled features like automated metadata quality checks, and dissemination of ESS-DIVE datasets onto other data platforms including Google Dataset Search and DataCite.

54 ENVIRONMENTAL SCIENCES↗

Metadata for a systematic description of signal data

This chapter aims to provide a comprehensive overview of metadata types that may be useful during system design, optimization, and automation. Metadata are grouped into three main categories: (a) metadata describing signal generation, (b) metadata describing signal quality, and (c) contextual information in the form of annotations. Each of these categories is introduced and explained in three separate sections. Importantly, this chapter mainly answers what is considered metadata. To a lesser degree, recommendations are made regarding the selection of metadata for long-term storage. Chapter 4 will explain where and how to store metadata. Chapters 5 and 6 explain how to collect certain metadata through dedicated sensor validation tests (Chapter 5) or algorithmic analysis (Chapter 6).

Alferes, Janelcy↗

Optimizing Metadata Exchange: Leveraging DAOS for ADIOS Metadata I/O

In HPC I/O middleware like the Adaptable I/O System (ADIOS) often mediates data transfers between applications. The metadata I/O generated by such systems often presents significant scaling and performance limitations. This work seeks improvement opportunities for metadata I/O by leveraging the DAOS storage systems, a recent storage system solution deployed on high-end systems such as the Aurora supercomputer. We investigate the tradeoffs and the design space for integrating I/O engines for the ADIOS middleware based on the different storage mechanisms supported by DAOS. We present a new DAOS-Array-ChunkSize-aligned engine which provides up to 2.3× improved performance than when using the existing DAOS-POSIX interface, without requiring any application modifications.

Venkatesh, Ranjan Sarpangala↗

Groundwater and river water elevations and temperature from 2017 to 2022 across Meander Z in the East River Watershed, Colorado

This dataset includes groundwater and river water elevations and temperature data collected in the East River watershed located in the Upper Colorado River Basin. The data were collected in order to investigate the coupling between hydrology and biogeochemical processes in the floodplain. Data was collected at ten groundwater locations in Meander Z (MZ), located just upstream of the confluence with Brush Creek and two river locations directly adjacent to Meander Z from 2017-2019. From 2019-2022, data was collected at five groundwater locations in Meander Z. Note that location names, not location identifiers (IDs), are used in the related publication Dewey et al. (2022). Both location IDs and names are included in data files. Files in this dataset include the main data files for each location zipped into a single folder (waterlevel_data.zip), an installation methods file describing sensor installation (InstallationMethods.csv), a file containing field metadata including GPS (Global Positioning System) coordinates and ground surface elevations (transducers_locations.csv). This dataset also includes a file-level metadata (flmd.csv) file that lists each file contained in the dataset with associated metadata and a data dictionary (dd.csv) file that contains column/row headers used throughout the files along with a definition, units, and data type. This dataset conforms to the ESS-DIVE hydrological reporting format. 2026-04-27 Update: The river water elevation data files (ER-MZR1.csv and ER-MZR2.csv) were corrected. The data for these two locations were inadvertently swapped in the original published data. This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

WorkJournalMaker (WJMaker) v0.5

The software generates and maintains daily work journal entries in text format, via a web browser. The journal entries are saved in a structured directory file tree on the system running the software. The software also incorporates a database so that it can track the location of files in the file system and various other metadata. The software allows the users to access their journal entries either through the browser or as discrete text files, facilitating sharing and open science. Additionally, to assist with the yearly PMP process, this tool connects to LLM APIs to provide summarization of the journal entries on a month-by-month or weekly basis. The advantage over similar technologies such as Apple Notes (extremely popular for notetaking) is that the instant software does not force the user to stay inside the Apple ecosystem, since it allows for export of the user's text files. This facilitates open science, so that researchers who use the tool can easily transfer their research notes to any other system. The WebJournalMaker repository is here: https://github.com/lbnl-science-it/WorkJournalMaker The WebJournalMaker repository is forked from the JournalSummarizer: https://github.com/tyfong-lbl/JournalSummarizer and builds on its code. I wrote the code for both of these software repos, using generative AI.

Fong, Timothy [Lawrence Berkeley National Laborato↗

ESS-DIVE Unoccupied Aerial Systems (UAS) Reporting Format v1

Here we present documentation of the ESS-DIVE reporting format for Unoccupied Aerial System (UAS) data and metadata. This reporting format provides guidance to data contributors on how to store data to maximize their discoverability, facilitate their efficient reuse, and add value to individual datasets. For data users, the reporting format will better allow data repositories to optimize data search and extraction, and more readily integrate similar data into harmonized synthesis products. The reporting format provides templates and guidance for the reporting of metadata for UAS experimental campaigns, individual flights, platform and sensor description. To improve data access and discoverability, the reporting format proposes a data description scheme of Levels based on the degree of processing, where Level 0 includes raw data, through to Level 3 being derived data end products. A range of examples of data types for each Level are given, with suggested file naming schemes. The reporting format presented here is intended to form a foundation for future development that will accommodate new UAS technologies and approaches to data access and use in the future. The reporting format documentation is maintained and updated on the ESS-DIVE Community Space GitHub at https://github.com/ess-dive-community/essdive-uas. This data package is the first published version of this reporting format, and comprises a zip file of the complete content of https://github.com/ess-dive-community/essdive-uas v1.0. The zip contains the reporting format description, instructions and variable definitions in GitHub markdown language (*.md) and metadata templates in csv format. The reporting format is designed to be compatible with other ESS-DIVE formats, and it is specifically recommended that this reporting format be used in conjunction with the File-level metadata (FLMD) and comma separated values (csv) reporting formats for submission to the ESS-DIVE repository.

54 ENVIRONMENTAL SCIENCES↗

Datum: A Scientific Metadata Catalog

The data catalog market is currently flooded with a myriad of different products, but none serve the scientific community well. There are cloud-native tools like Databricks, Snowflake,to on-premise solutions like Collibra and Datahub. The common failing of all these tools however, is their inability to serve the scientific data community directly. Most catalogs are targeted towards financial, health, or user data - not sensor or scientific domain data. They also prioritize integrations that often don’t exist or are just starting to be used in the scientific realm - all while ignoring common scientific tools and file types. Datum is a catalog which targets the scientific data directly, including the tools and networks in which those tools are used. We work with the producers and consumers of the data where they are, targeting cloud and on-premise with a focus on classified networks. Datum is an Erlang/Elixir application. Technical Features Note: The features listed below are still under development and may change, slightly, upon final delivery of the product. File Formats - Datum has the ability to read additional metadata and provides processing pipelines for the following file formats: Plain Text, PDF, LaTeX, HTML, Open Document Format (.odt), XML, CSV/TSV (and other standard delimiters), OpenDocument Database and Spreadsheets, Geo-Referenced TIFF, Common Data Format, HDF/HDF5, LabView TDMS, Excel, DeltaTables, Parquet, Apache Iceberg, Apache Hudi and many others. Metadata Collection - Scanners for the local and networked file systems and cloud storage providers. Network integration with common databases such as MSSQL and MySQL. User Plugin System - Users are able to provide either file processing, metadata extraction, or sampling plugins in the programming language of their choice. Authentication/Authorization -: OIDC integration, SCIM provisioning and EntraID integration out of the box. Full user and group management system with a “least privilege” operating mode. Governance - Customizable data governance platform; dictate and enforce required metadata, enforce data embargos, and enforce user agreements and NDAs before data access. Ability to create health checks on data, rejecting abandoned or poorly curated data and automatically removing it from the search index. Ability for users to submit corrections. Search - Semantic search is a first class citizen. No licenses to expensive, external software required. Integrated use of vectors and vector-based search allows for AI agent integration at all levels of operation. Metadata Model - Display and control data’s lineage and connections to other data and data directories. Data is modeled after a filesystem - an organization instantly recognizable and navigable by most any user. CLI and SDK - Ships with a Command Line Interface (CLI) tool and with a fully-featured Python SDK. This allows for rapid and programmatic use of Datum by every level of user. Minimal Infrastructure - Datum ships as a single executable file and can be run on any operating system and most CPU architectures. Datum has no reliance on external databases, search indexing tools, or other outside services - and it runs equally well on edge computing devices, cloud services, or in a clustered HPC environment.

darrington, john↗

ScholarGuard

The ScholarGuard framework aims to address the gap in archiving and preservation efforts for scholarly artifacts beyond traditional research papers, such as software source code, datasets, presentation slides, workflows, protocols, videos, and more. It introduces a prototype system designed to automatically track researchers' outputs across various scholarly productivity portals on the open web, including platforms like GitHub, Slideshare, Figshare, and Wikipedia. The system detects the availability of new scholarly artifacts and applies modern web archiving technology to create a durable archival record, including high-level metadata for each artifact. This metadata is displayed within the system, linking both to the live version and the archived version of the resource, ensuring long-term accessibility and preservation of diverse research outputs. The software serves as a critical tool for preserving the broader spectrum of scholarly contributions, facilitating visibility, searchability, and long-term access to research artifacts beyond the traditional scope of journal publications.

Balakireva, Lyudmila↗

Assessing Machine Learning as a Tool to Explain Variance in Deployed Photovoltaic (PV) System Degradation

Degradation remains a large uncertainty in forecasting production for PV plants, creating significant risk for developers and financiers. This study aims to quantify the distribution and drivers of degradation across 10,000 PV systems deployed for distributed or utility generation by training a machine learning model to predict year-over-year degradation rates from metadata characteristics. A combination of K-Means clustering and random forest regressor were found to associate multiple metadata features as potential drivers of degradation, including module characteristics, system design, and climate features. From this, it is inferred that if machine learning is able to find complex patterns between metadata features and system performance loss, such methods can be employed to help developers and financiers make data-informed decisions when estimating long-term energy production forecasts in financial models.

Dunn, Jimmy C.↗

A three-year dataset supporting research on building energy management and occupancy analytics

Abstract This paper presents the curation of a monitored dataset from an office building constructed in 2015 in Berkeley, California. The dataset includes whole-building and end-use energy consumption, HVAC system operating conditions, indoor and outdoor environmental parameters, as well as occupant counts. The data were collected during a period of three years from more than 300 sensors and meters on two office floors (each 2,325 m 2 ) of the building. A three-step data curation strategy is applied to transform the raw data into research-grade data: (1) cleaning the raw data to detect and adjust the outlier values and fill the data gaps; (2) creating the metadata model of the building systems and data points using the Brick schema; and (3) representing the metadata of the dataset using a semantic JSON schema. This dataset can be used in various applications—building energy benchmarking, load shape analysis, energy prediction, occupancy prediction and analytics, and HVAC controls—to improve the understanding and efficiency of building operations for reducing energy use, energy costs, and carbon emissions.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Scaling Ensembles of Data-Intensive Quantum Chemical Calculations for Millions of Molecules

Deep learning models are efficient computational tools that can accelerate the inverse design of molecules with desired functional properties by generating predictions at a fraction of the time required by traditional quantum chemical approaches. To ensure that a model maintains accuracy and transferability across broad regions of the chemical space explored during the inverse design, it must be trained on massively large volumes of simulation data. This requires running large-scale ensemble quantum chemical calculations on high-performance computing (HPC) systems for data collection. However, the efficient execution of such large ensemble calculations and the management of large volumes of output data require tools that can judiciously utilize computational resources and manage metadata overhead on the file system. Therefore, we present a high-performance, scalable, ensemble management framework for performing data-intensive quantum chemical electronic structure calculations for organic molecules. This framework provides abstractions to plug different ab initio, first principles, and first principles-based semi-empirical methods and executes them efficiently at large scale on HPC systems. It dynamically distributes tasks to resources and uses tiered storage for managing large collections of files. We employed this framework to process over ten million organic molecules and generate open-source datasets that provide UV-vis absorption spectra by running time-dependent density-functional tight-binding calculations. It is the largest database containing molecular optical spectra that were simulated with quantum chemical methods in a consistent manner.

Mehta, Kshitij↗

A Unified User-Friendly Instrument Control and Data Acquisition System for the ORNL SANS Instrument Suite

In an effort to upgrade and provide a unified and improved instrument control and data acquisition system for the Oak Ridge National Laboratory (ORNL) small-angle neutron scattering (SANS) instrument suite—biological small-angle neutron scattering instrument (Bio-SANS), the extended q-range small-angle neutron scattering diffractometer (EQ-SANS), the general-purpose small-angle neutron scattering diffractometer (GP-SANS)—beamline scientists and developers teamed up and worked closely together to design and develop a new system. We began with an in-depth analysis of user needs and requirements, covering all perspectives of control and data acquisition based on previous usage data and user feedback. Our design and implementation were guided by the principles from the latest user experience and design research and based on effective practices from our previous projects. In this article, we share details of our design process as well as prominent features of the new instrument control and data acquisition system. The new system provides a sophisticated Q-Range Planner to help scientists and users plan and execute instrument configurations easily and efficiently. The system also provides different user operation interfaces, such as wizard-type tool Panel Scan, a Scripting Tool based on Python Language, and Table Scan, all of which are tailored to different user needs. The new system further captures all the metadata to enable post-experiment data reduction and possibly automatic reduction and provides users with enhanced live displays and additional feedback at the run time. We hope our results will serve as a good example for developing a user-friendly instrument control and data acquisition system at large user facilities.

47 OTHER INSTRUMENTATION↗