Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “standardized data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

A change language for ontologies and knowledge graphs

Ontologies and knowledge graphs (KGs) are general-purpose computable representations of some domain, such as human anatomy, and are frequently a crucial part of modern information systems. Most of these structures change over time, incorporating new knowledge or information that was previously missing. Managing these changes is a challenge, both in terms of communicating changes to users and providing mechanisms to make it easier for multiple stakeholders to contribute. To fill that need, we have created KGCL, the Knowledge Graph Change Language (https://github.com/INCATools/kgcl), a standard data model for describing changes to KGs and ontologies at a high level, and an accompanying human-readable Controlled Natural Language (CNL). This language serves two purposes: a curator can use it to request desired changes, and it can also be used to describe changes that have already happened, corresponding to the concepts of “apply patch” and “diff” commonly used for managing changes in text documents and computer programs. Another key feature of KGCL is that descriptions are at a high enough level to be useful and understood by a variety of stakeholders—e.g. ontology edits can be specified by commands like “add synonym ‘arm’ to ‘forelimb’” or “move ‘Parkinson disease’ under ‘neurodegenerative disease’.” We have also built a suite of tools for managing ontology changes. These include an automated agent that integrates with and monitors GitHub ontology repositories and applies any requested changes and a new component in the BioPortal ontology resource that allows users to make change requests directly from within the BioPortal user interface. Overall, the KGCL data model, its CNL, and associated tooling allow for easier management and processing of changes associated with the development of ontologies and KGs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

The Case for Strong Scaling in Deep Learning: Training Large 3D CNNs with Hybrid Parallelism

Here, we present scalable hybrid-parallel algorithms for training large-scale 3D convolutional neural networks. Deep learning-based emerging scientific workflows often require model training with large, high-dimensional samples, which can make training much more costly and even infeasible due to excessive memory usage. We solve these challenges by extensively applying hybrid parallelism throughout the end-to-end training pipeline, including both computations and I/O. Our hybrid-parallel algorithm extends the standard data parallelism with spatial parallelism, which partitions a single sample in the spatial domain, realizing strong scaling beyond the mini-batch dimension with a larger aggregated memory capacity. We evaluate our proposed training algorithms with two challenging 3D CNNs, CosmoFlow and 3D U-Net. Our comprehensive performance studies show that good weak and strong scaling can be achieved for both networks using up to 2K GPUs. More importantly, we enable training of CosmoFlow with much larger samples than previously possible, realizing an order-of-magnitude improvement in prediction accuracy.

97 MATHEMATICS AND COMPUTING↗

Characterization and differentiation of aluminum powders used in improvised explosive devices. Part 2: Micromorphometric method refinement and preliminary statistical analysis

Abstract Aluminum (Al) powder is commonly encountered in improvised explosive devices (IEDs) as a metallic fuel due to its availability and low cost. Although available commercially in powder form, amateur bomb‐makers also produce their own Al powder via simple methods found online. In order to provide investigative leads and forensic intelligence, it is important to evaluate not only the composition of homemade devices, but also to distinguish between the various forms of Al powder they contain. To achieve this goal, a method using automated microscopy in combination with statistical techniques has been demonstrated to have the potential to provide source discrimination and investigative leads in source attribution of Al powders in IEDs. The present research refined this method and investigated 59 industrially and amateurly produced Al powder sources with seven subsamples per source using two traditional linear discriminant analyses (LDA), one with a standard data split for training and testing, and another using leave‐one‐out cross‐validation. Averaging the classification accuracies for the two LDA‐based analyses, LDA has the ability to correctly classify 59.26%, 83.35%, and 80.69% of the samples based on their powder source, type, and production method, respectively. This classification accuracy represents a 3407%, 317%, and 61.38% increase in accuracy from random class assignment, respectively. Further, in most instances of incorrect data attribution to a particular source, the subsample has been misidentified with another sample of the same powder type or production method.

Ommen, Danica M.↗

Long-Term Stability of Mullite and Magnesia-encased Ag|Ag + Reference Electrodes in Molten MgCl 2 -KCl-NaCl

In this study, the stability of Ag|Ag + molten chloride reference electrodes was studied using mullite and magnesia tubes as reference electrode membrane materials in molten MgCl 2 -KCl-NaCl for continuous immersion up to 31 d. Two electrochemical methods were used to characterize the long-term performance of the reference electrodes at 500 °C in an inert atmosphere using a Mg rod as the working electrode: open circuit potential (OCP) and onset of Mg(II) reduction as measured by cyclic voltammetry (CV). There was only a 3% difference between the average OCP measured versus the mullite and MgO reference electrodes. Potential drift was higher for mullite (−1.34 mV d −1 ) than MgO (0.22 mV d −1 ). But both drift rates were very low compared to the OCP data standard deviation, indicating that no significant drift was measured for us to 31 d.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Neurodata Extensions Catalog (NDX Catalog) v1.0.0

The Neurodata Extensions Catalog (NDX Catalog) is a novel web archive and associated software tools that enable users to easily create, search for, test, review, and install extensions to the Neurodata Without Borders (NWB) data standard for neurophysiology. The Catalog consists of several software repositories that can be found at https://github.com/nwb-extensions.

Tritt, Andrew↗

AqNWB v0.1

This is a C ++ acquisition API for neurophysiological data in the NeurodataWithoutBorders (NWB) data standard.

Prince, Stephanie↗

Common Electric Power Transmission System Model JSON Schema Specification

The Common Electric Power Transmission System Model (CTM) is an intuitive, extensible, language-agnostic, and error-resistant specification of electric power network components parameter names and units, and relation between components, intended for use by the research community developing new computational methods for power systems operations and simulation. Power system datasets following the CTM specification can be read as dictionaries and manipulated in that form in most programming languages (e.g., Python, Julia, C++). This standard data structure in CTM makes it easy to work in multiple power systems domains (e.g., economic operation, reliability assessment, electricity markets, stability assessment, etc.) without requiring conversions between use-case-specific file formats with information loss in the process. This repository specifies CTM as a JSON Schema, provides documentation, derivate (code-generated) implementations of CTM, and example data and usage of the schema for important use cases.

Aravena Solis, Ignacio↗

PowerSystemsData Specification [SWR-25-92]

This repository defines a standardized data format for representing power system datasets, with a focus on supporting time-series visualizations and 3D visualization tools. By providing a consistent and extensible structure, this format aims to streamline the development and interoperability of visualization codebases in the power systems domain. NOODLES is a cross-platform/device/tool protocol for collaborative visualization. NOODLES was Developed at the National Renewable Energy Laboratory (NREL) as a capability of the Insight Center https://www.nrel.gov/computational-science/insight-center.html

Brunhart-Lupo, Nicholas [National Renewable Energy↗

ESS-DIVE Reporting Format for File-level Metadata

The ESS-DIVE reporting format for file-level metadata (FLMD) provides granular information at the data file level to describe the contents, scope, and structure of the data file to enable comparison of data files within a data package. The FLMD are fully consistent with and augment the metadata collected at the data package level. We developed the FLMD template based on a review of a small number of existing FLMD in use at other agencies and repositories with valuable input from the Environmental Systems Science (ESS) Community. Also included is a template for a CSV Data Dictionary where users can provide file-level information about the contents of a CSV data file (e.g., define column names, provide units). Files are in .csv, .xlsx, and .md. Templates are in both .csv and .xlsx (open with e.g. Microsoft Excel, LibreOffice, or Google Sheets). Open the .md files by downloading and using a text editor (e.g. Notepad or TextEdit). Though we provide Excel templates for the file-level metadata reporting format, our instructions encourage users to 'Save the FLMD template as a CSV following the CSV Reporting Format guidance'. In addition, we developed the ESS-DIVE File Level Metadata Extractor which is a lightweight python script that can extract some FLMD fields following the recommended FLMD format and structure.

54 ENVIRONMENTAL SCIENCES↗

ESS-DIVE Reporting Format for Comma-separated Values (CSV) File Structure

The ESS-DIVE reporting format for Comma-separated Values (CSV) file structure is based on a combination of existing guidelines and recommendations including some found within the Earth Science Community with valuable input from the Environmental Systems Science (ESS) Community. The CSV reporting format is designed to promote interoperability and machine-readability of CSV data files while also facilitating the collection of some file-level metadata content. Tabular data in the form of rows and columns should be archived in its simplest form, and we recommend submitting these tabular data following the ESS-DIVE reporting format for generic comma-separated values (CSV) text format files. In general, the CSV file format is more likely accessible by future systems when compared to a proprietary format and CSV files are preferred because this format is easier to exchange between different programs increasing the interoperability of a data file. By defining the reporting format and providing guidelines for how to structure CSV files and some field content within, this can increase the machine-readability of the data file for extracting, compiling, and comparing the data across files and systems.Data package files are in .csv, .png, and .md. Open the .csv with e.g. Microsoft Excel, LibreOffice, or Google Sheets. Open the .md files by downloading and using a text editor (e.g., notepad or TextEdit). Open the .png in e.g. a web browser, photo viewer/editor, or Google Drive.

54 ENVIRONMENTAL SCIENCES↗

PEAR Objective 3.2

NTESS and its aviation security industry partners developed the Open Threat Assessment Platform (OTAP), an open-architecture platform project that provides a common set of software interfaces and data standards, for Transportation Security Administration airport screening.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Establishing Validation Sites for Field-Level Emissions Quantification from Grain Sorghum in Southern Great Plains

Establish publicly available open-source, high-resolution datasets of greenhouse gas (GHG) emissions, meteorology, soil, and vegetation to support testing and validation of emerging biofuel production monitoring technologies. The project team deployed latest technologies to collect “Gold Standard” data for field- level fluxes of CO 2 , N 2 O, CH 4 and H 2 O in grain sorghum production systems in Texas, Oklahoma, and Kansas. The Oklahoma State University led SMARTFARM was initiated on Oct 1, 2020, due to COVID-19 pandemic delay. We installed closed-path dual laser Eddy Covariance (EC) flux measurement system along with a well-established open-path EC system for simultaneously measuring field-level fluxes of CO 2 , N 2 O, CH 4 , and H2O. The eddy flux data are combined with sub-acre soil moisture and temperature measurements, plant and soil analysis data (nutrient, composition, carbon, growth, and phenology), spatial surveys of soil EC, and pH. Grain sorghum crop was managed including planting and harvesting by producers. All teams recruited postdocs and students for the project and submitted purchase orders for field equipment. The team has developed a draft impact sheet and finalized the advisory board for the project.

09 BIOMASS FUELS↗

Winter Storm Elliott: Energy Emergency Alert Lessons Learned

Power blackouts in the United States in 2022 highlighted the need for better communication and planning among power grid stakeholders. This need was further demonstrated at a 2021 workshop held by the Grid Reliability Initiative Utilizing Data Standardization (GRIDS) Project. The GRIDS Project was established to understand and address the gaps in communication between the transmission and distribution operators. With the ongoing energy transition including electrification, distributed energy resource penetration, and so on, this project aims for improved energy emergency communications at the transmission-distribution interface. In 2022, Winter Storm Elliott affected 60% of the United States population (Weather Underground 2022), resulting in rolling blackouts in many areas of the eastern United States (Knoxville Utilities Board [KUB] 2023). Information presented here is based on the GRIDS workshop findings and subsequent research and interviews with stakeholders. After a brief summary of the effects of Winter Storm Elliott, this report covers the standard energy emergency communications processes currently in use, how the storm affected power management operations for the Tennessee Valley Authority (TVA) and its local utility companies, and related experiences in other states. The report concludes with lessons learned and recommendations for improvements and a path forward.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Automating Detection and Diagnosis of Faults, Failures, and Underperformance in PV Plants

The project developed hybrid physics-based and machine-learning methods for near-real-time detection of balance-of-system faults (e.g., string, combiner, and tracker outages) in utility-scale Photovoltaic plants, achieving over 50% true positive rates with under 10% false positives and significantly reducing engineering setup time. In the extended phase, the scope expanded to plant-level underperformance analysis and industry benchmarking through the SUPER.epri.com platform. SUPER standardizes data processing and performance metrics across more than 9 GWac and 120+ plants, enabling robust comparisons and insights into loss rates, inverter downtime, and capacity degradation.

14 SOLAR ENERGY↗

California - Leosphere Windcube 866 (120), Humboldt / Reviewed Data

The purpose of this dataset is to provide filtered, averaged lidar data and standardize the data format of various data streams from the buoy into NetCDF. The attached Lidar Buoy Data Dictionary provides further details on the various instruments mounted on the buoys, parameters measured by each instrument, and the frequency of data collection.

17 WIND ENERGY↗

Buoy - California - Wind Sentinel (120), Humboldt - Processed Data

The purpose of the dataset is to provide preliminary filtered, averaged lidar data and standardize the data format of various datastreams from the buoy into NetCDF. The attached Lidar Buoy Data Dictionary provides further details on the various instruments mounted on the buoys, parameters measured by each instrument, and the frequency of data collection.

17 WIND ENERGY↗

Buoy - California - Wind Sentinel (130), Morro Bay - Processed Data

The purpose of the dataset is to provide preliminary filtered, averaged buoy data and standardize the data format of various data streams from the buoy into NetCDF. The attached Lidar Buoy Data Dictionary provides further details on the various instruments mounted on the buoys, parameters measured by each instrument, and the frequency of data collection.

17 WIND ENERGY↗