Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “standardized data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Understanding Biases in Sample Preparation Techniques for Coupled Scanning Electron Microscopy and MAMA PuO 2 Morphological Analysis

In this project, the scanning electron microscopy (SEM) sampling method used during the statistical design study (SDS) was investigated to determine if any sampling biases were present in the analyzed data. Using standard particle size distribution powders from the National Institute of Standards and Technology (NIST 1984 standard reference material) with the origin wet dispersion method, it was determined that a bias to smaller particles was present. This was supported by theoretical calculations using Stokes’ law to determine the settling rate of spherical particles of roughly the same size and mass as those found in the SDS. Based on the theoretical calculations, it was determined that the settling rate for each of the 76 powder sets in the SDS could be unique based on specific particle shape and mass distributions, making a universal correction factor/formula not applicable. Therefore, priority shifted to developing an improved wet dispersion method that significantly reduced the particle settling rate for all particle size and shapes. This was achieved by replacing the original solvent (isopropyl alcohol) with a heavy liquid (lithium heteropolytungstates), which dramatically slowed the settling rate and allowed for the capture of a suitable homogeneous aliquot. SEM imaging and Morphological Analysis for Material Attribution (MAMA) software analysis were conducted on the NIST standard, and the SEM/MAMA data were compared to data captured by a dynamic image analysis particle size analyzer. The resulting data confirmed that the new wet dispersion method does indeed deliver an improved representative aliquot to the SEM stub. For instance, in the NIST certificate, the average particle size is ~17.1 µm ± 2.2 µm with a normal distribution. The initial wet dispersion method resulted in a drastically reduced average particle size of 6.1 µm in addition to a non-representative heavy bi-modal distribution whereas the improved LST wet dispersion method resulting in an average particle size that was much closer to the NIST certificate (12.7 µm) with a similar normal distribution. Although the improved method was still short of the NIST certificate average, atomic force microscopy analysis determined that the resulting ~20-25% reduction in size was due to particles sinking into the carbon sticky tape used for SEM imaging. It is believed that that this bias can be calibrated in a much more predicable manner than the original settling rate bias. In addition, the matching normal distribution curves between the NIST certificate and the heavy liquid method indicate a much-improved representative aliquot has been sampled and imaged. A surrogate CeO 2 powder was used to reflect PuO 2 more accurately and to aid in implementing radiological controls and shielding. The resulting data sets from the SEM/MAMA method and the particle size analyzer give almost identical average particle sizes and particle distribution statistics. Future work will re-analyze several select runs from the SDS to determine if morphological signatures can be found with the improved sampling method.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

LTE Electrolyzer Data Collection

The goal for NREL is to collect, develop and publish performance metrics relative to low temperature electrolyzer installations. This will be done through the development of: Secure storage solution to house the collection of data from multiple projects Standardization of data to be collected and analyzed. This will be done using data templates developed with the help of partners involved with electrolyzer installations. Analysis that produces metrics of interest for all stakeholders Aggregation of results from multiple projects to view industry progress as a whole Publication of aggregated results in the form of composite data products (CDPs) Collaboration with Idaho National Lab and their work with high temperature electrolyzer installations will enable efficient use of storage and analysis tools.

data↗

Supply Chain Risk Management: Data Structuring

Supply chain risk management (SCRM) is an area of research that addresses both logistics concepts to maximize efficiency, reliability, and revenue as well as risk features, such as potential weak points, break points, and vulnerabilities within the supply chain. SCRM is used to find risks introduced at each node in a supply chain and how these risks can impact a company’s products, individuals, customers, and reputation. SCRM is a relatively new field, so standardized processes including data structuring are not fully documented. This paper explains the importance of a standard data structuring methodology and how it can enhance current SCRM efforts. Data ingest, structuring, and analysis are predominantly managed by humans. Automating some of the less complex steps can positively impact SCRM by allowing human analysts to focus on more strategic analyses. Types of data to be collected and structured are collected via publicly available information related to hardware, software, and corporate entities. After the data has been collected, the information is formatted in a specific manner, conforming to a schema, to allow for more effective and efficient ingest for further analysis. This paper outlines data structures used by Pacific Northwest National Laboratory for SCRM research and analysis purposes. These structures have been used for hundreds of analyses and have been successful in developing a common baseline. Data structuring is one of the first steps in data standardization, which will further mature and enhance the SCRM research area.

supply chain risk management, data structuring, re↗

DOE BSSD Performance Management Metrics Report Q1

Microbes play key roles in our biosphere, from driving global nutrient cycling to impacting plant, animal and human health and disease. Complex data from microbial genomes, proteins, and metabolites provide a window into these tiny engines that drive life on our planet. Yet these data are dispersed among researchers’ laboratories and various repositories, making it difficult to access. This calls for new ways of managing data, improving data interoperability, advancing community standards, and creating an infrastructure where data are shared efficiently. We have built the National Microbiome Data Collaborative (NMDC) to advance how scientists create, use, and reuse data to redefine the way we understand and harness the power of microbes. The vision of the National Microbiome Data Collaborative (NMDC) is to drive a microbiome data sharing network connecting data, people, and ideas to advance microbiome innovation and discovery. The NMDC was launched in 2019 and brought together DOE National Laboratories to collaborate across resources, capabilities, and expertise. The NMDC team was strategically assembled to include software developers, microbial researchers, metadata experts, and multi-omics specialists. The diversity of the NMDC team reflects the inherently interdisciplinary nature of microbiome science, and we leverage the strengths of the DOE National Laboratory system. Towards BER’s goal of advancing an iterative systems biology approach to the understanding of microbial genomes, the NMDC serves as a foundation for infrastructure, data standards, and community building. Together with the flagship DOE User Facilities, the Joint Genome Institute (JGI) and the Environmental Molecular Sciences Laboratory (EMSL), we are developing core capabilities in metadata standards for environmental descriptors and sample handling and processing; standardized bioinformatic workflows; an interface for data search and access; and robust community engagement activities. The NMDC production platform supports long-term data infrastructure and community building for BER’s bioenergy and environmental research goals. Our approach leverages lessons learned and an ambitious framework for collaborative, interdisciplinary data infrastructure to support microbiome research. The NMDC supports data, information, and knowledge access through three defined software tools – the Submission Portal, NMDC EDGE, and the Data Portal – driven by community needs. Herein, we describe the value proposition for the microbiome research community, our overarching strategy, and challenges and opportunities for developing the NMDC as both an infrastructure and community engagement program.

59 BASIC BIOLOGICAL SCIENCES↗

Sample IEEE123 Bus system for OEDI SI

Time series load and PV data from an IEEE123 bus system. An example electrical system, named the OEDI SI feeder, is used to test the workflow in a co-simulation. The system used is the IEEE123 test system, which is a well studied test system (see link below to IEEE PES Test Feeder), but some modifications were made to it to add some solar power modules and measurements on the system. The aim of this project is to create an easy-to-use platform where various types of analytics can be performed on a wide range of electrical grid datasets. The aim is to establish an open-source library of algorithms that universities, national labs and other developers can contribute to which can be used on both open-source and proprietary grid data to improve the analysis of electrical distribution systems for the grid modeling community. OEDI Systems Integration (SI) is a grid algorithms and data analytics API created to standardize how data is sent between different modules that are run as part of a co-simulation. The readme file included in the S3 bucket provides information about the directory structure and how to use the algorithms. The sensors.json file is used to define the measurement locations.

123 bus↗

PandAna: A Python Analysis Framework for Scalable High Performance Computing in High Energy Physics

Modern experiments in high energy physics analyze millions of events recorded in particle detectors to select the events of interest and make measurements of physics parameters. These data can often be stored as tabular data in files with detector information and reconstructed quantities. Current techniques for event selection in these files lack the scalability needed for high performance computing environments. We describe our work to develop a high energy physics analysis framework suitable for high performance computing. This new framework utilizes modern tools for reading files and implicit data parallelism. Framework users analyze tabular data using standard, easy-to-use data analysis techniques in Python while the framework handles the file manipulations and parallelism without the user needing advanced experience in parallel programming. In future versions, we hope to provide a framework that can be utilized on a personal computer or a high performance computing cluster with little change to the user code.

Groh, Micah↗

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration↗

Reference Correlations for the Density and Viscosity of Molten Alkali and Alkaline Earth Fluoride Salts

While there is a significant body of literature pertaining to thermophysical property measurements of molten salts, there is often a wide degree of variability among independent measurements of the same compounds. As such, the scientific community benefits greatly from an unbiased, independent assessment of duplicate datasets, so that reference correlations which describe these thermophysical properties as functions of temperature can be determined and then commonly used by researchers, scientists, and engineers. With regard to molten fluoride compounds, a significant time has elapsed since density and viscosity reference correlations have been determined; Janz conducted the most recent effort, in 1988, to provide reference correlations for the densities and viscosities of molten fluoride compounds via the National Standard Reference Data System coordinated by the National Bureau of Standards. Since then, new data have been published for molten fluoride compounds, and a new precedent has surfaced for putting forth reference correlations that involve fitting to multiple primary datasets. In this work, reference correlations are put forth for molten alkali and alkaline earth fluoride compounds in an effort to provide updated, improved correlations for general use. For molten alkali fluoride densities, estimated uncertainties with a 95% confidence interval are summarized as follows: LiF (0.63%), NaF (0.48%), KF (0.76%), RbF (0.93%), and CsF (0.75%). For molten alkaline earth fluoride densities, an estimated uncertainty was not able to be quantified for BeF 2 because of limited data; however, estimated uncertainties with a 95% confidence interval are summarized as follows for the remaining alkaline earth fluorides: MgF 2 (1.5%), CaF 2 (0.92%), SrF 2 (1.6%), and BaF 2 (0.23%). For molten alkali fluoride viscosities, uncertainty was not able to be quantified for RbF and CsF because of limited data; however, estimated uncertainties with a 95% confidence interval are summarized as follows for the remaining alkali fluorides: LiF (4.4%), NaF (3.0%), and KF (4.0%). For molten alkaline earth fluoride viscosities, limited consistent data resulted in the recommendation of single datasets (from literature) that are deemed to be the most trustworthy based on the quality of the underlying experimental studies.

Birri, A. [Oak Ridge National Laboratory (ORNL), O↗

Cloud-based Testbed for Adaptive Under-Frequency Load Shedding with High DER Penetration

Increasing penetration of distributed energy resources and behind-the-meter renewables may soon disrupt the efficacy of critical protection schemes, such as under-frequency load shedding (UFLS). Improved data exchange and coordination across the transmission-distribution boundary will be required to maintain reliability of bulk electric system. Standards-based data integration platforms using agreed-upon semantic vocabularies, such as the Common Information Model, will be key to enabling adaptive protection schemes requiring synthesized data from both the bulk power system and behind-the-meter resources. This paper introduces a cloud-based open-source data integration environment and UFLS clustering algorithm being developed to enable adaptive relay coordination between transmission and distribution utilities in the state of Vermont.

Anderson, Alexander A.↗

Operation and Maintenance of PV Systems: Data Science, Analysis, and Standards

This effort improves the effectiveness and reduce uncertainty in O&M cost through four primary objectives/tasks: 1) institutionalize standards for reliability and availability reporting for large PV power plants; 2) bridge systemic O&M knowledge gaps around important topics affecting O&M; 3) characterize systemic failure modes and patterns and accelerate O&M experiential learning cycles using field data; and 4) establish a baseline understanding of UPVS O&M cost drivers. Key results of this effort include publication of IEC standards, published topical papers on O&M topics, training, and characterize field data for climate- and service-related patterns (additional details below). Integrating these results serves to reduce performance risk and facilitate improvement in the way solar projects are operated and maintained. Results are well received and two publications are among the most successful SETO publications at NREL ("Model of Operation and Maintenance Costs for Photovoltaic Systems with over 40,000 downloads and "Best Practices in Operation and Maintenance of PV Systems, 3rd Ed." with over 90,000 downloads).

14 SOLAR ENERGY↗

A practical approach to using the Genomic Standards Consortium MIxS reporting standard for comparative genomics and metagenomics

Comparative analysis of (meta)genomes necessitates aggregation, integration, and synthesis of well-annotated data using standards. The Genomic Standards Consortium (GSC) collaborates with the research community to develop and maintain the Minimal Information about any (x) Sequence (MIxS) reporting standard for genomic data. To facilitate use of the GSC’s MIxS reporting standard, we provide a description of the structure and terminology, how to navigate ontologies for required terms in MIxS, and demonstrate practical usage through a soil metagenome example.

standards, metadata, genome, metagenome, schema, v↗

A use-case-driven approach for demonstrating the added value of digitalisation in wind energy

Digitalisation is one of the key drivers for reducing the costs and risks of wind energy. When considering whether to embark on a digitalisation initiative, two key questions arise. The first is what business or operational opportunities might feasibly be addressed and the second is which of the many potential aspects of digitalisation are relevant to those opportunities. In this work, we show how these questions can be answered with a use-case-driven approach, based around a survey aiming to collect and collate the main "pain points" (or everyday challenges) of people in the wind energy sector. Although the relatively low number of participants of the survey (46) means that the results should only be used indicatively, it is still possible to make some general recommendations for priorities for digitalisation efforts in the wind energy sector. Firstly, digitalisation efforts should focus both on supporting people carrying out cross-lifecycle tasks, in particular sharing data, managing data, undertaking general data analyses and accessing data. Tools to do this should deal with varying data formats and naming conventions, make metadata more accessible, define data and metadata standards, make more data publicly available and improve the quality of data. Secondly, efforts should also focus on supporting people in the wind farm operational phase, in particular with failure detection, fault diagnosis, failure rate modelling and predictive maintenance. Solutions to do this should focus on accessible and validated tools for fault detection, cloud or other data pipeline solutions for SCADA data and tools for exhaustive data documentation. Finally, digitalisation efforts should focus on better communicating and helping people become aware of existing solutions and tools, as well as on helping people to exert a stronger influence on possible solutions.

17 WIND ENERGY↗

TripleGraph

RDF triplestores are great tools for online graph analytic processing (i.e., graph pattern query processing), but they do not provide graph mining capabilities (e.g., PageRank, connected-component analysis, node eccentricity, etc.). The software title “TripleGraph” is a graph analysis toolkit, which uses an RDF triplestore as its backend for creating, manipulating, mining, and programming with large scale property graphs. It allows users to run various graph mining algorithms easily. User can import edgelist-formatted (homogeneous graph) or JSON-formatted graph (property graph) into the RDF triplestore using the provided tool and perform various analysis such as (1) Node/edge retrieval and manipulation, (2) Pathfinding between two given nodes, (3) Running graph mining algorithms (PageRank/Personalized PageRank, Single Source Shortest Path/Multi-Source Shortest Path, Connected Component, Node Eccentricity, Peer Pressure Clustering). It supports standard graph data format and works with a standard SPARQL endpoint like Jena Fuseki. It allows users to perform online graph analytic processing and graph mining on the same platform (a triplestore).

Sangkeun, MattLee↗

Foliar isotopic and elemental biochemistry and leaf mass per area (LMA), Parque Natural Metropolitano, Panama, 2022

Foliar N and C isotope composition, bulk elemental N and C composition, and leaf mass per area (LMA) for leaves sampled from the Parque Natural Metropolitano (PNM), Panama, from January to April 2022 are presented. Measurements were made on leaves from 17 different tree, shrub and liana species, from sunlit canopy and understory locations on 8 vertical profiles. Full sample details and leaf area index (LAI) from the vertical profiles at PNM are also included in the data package. The aim of this measurement campaign was two-fold: to improve our understanding of the vertical variation in leaf-level water use efficiency, and to improve models which can predict leaf traits from leaf contact spectral measurements. Biochemistry data and sample metadata are presented in .csv files. The original isotopic data report, which includes details about standards and data accuracy, is provided in .xlsx format. Data and metadata meet the ESS-DIVE reporting format requirements for file level metadata (FLMD) and comma separated values (csv). The protocol details are provided as pdf documents. In addition to foliar biochemistry data reported here these samples were also used for measurement of leaf gas exchange, and leaf optical properties. These data can be linked using the unique sample ID and are provided in separate data packages (NGT0192 and NGT0193).

54 ENVIRONMENTAL SCIENCES↗

Ka-Band ARM Zenith Radar Corrections (KAZRCOR, KAZRCFRCOR) Value-Added Products

The Ka-Band Atmospheric Radiation Measurement (ARM) Zenith Radar Corrections (KAZRCOR) and KA-Band ARM Zenith Radar CF-Radial, Corrected (KAZRCFRCOR) value-added products (VAPs) perform several corrections to the ingested KAZR moments and also create a significant detection mask for each radar mode. The VAPs compute gaseous attenuation as a function of time and radial distance from the radar antenna, based on ambient meteorological observations, and correct observed reflectivities for that effect. Mean Doppler velocities are dealiased to correct velocities whose magnitudes exceed the radar’s Nyquist velocity. Input KAZR data fields are passed through to the KAZRCOR or KAZRCFRCOR output files, in their native time and range coordinates. Complementary corrected reflectivity and velocity fields are provided, along with a mask of significant detections and a number of data quality flags. This report covers the KAZRCOR VAP as applied to the original KAZR radars and the upgraded KAZR2 radars. Originally, prior to late 2019, there were two separate code bases for the different radar versions. Following the harmonization of KAZR and KAZR2 data formats in 2019, only a single code base is required. The new combined KAZR and KAZR2 code base is called the KAZRCFRCOR VAP. The ‘cfr portion of the VAP name refers to the Radial Climate and Forecasting data format standards that are used in these data sets. Throughout this report, references to ‘KAZRCOR’ should be taken to apply to the ‘KAZRCFRCOR’ VAP as well unless there is an explicit statement to the contrary.

54 ENVIRONMENTAL SCIENCES↗

Transfer Learning for HVAC System Fault Detection

Faults in HVAC systems degrade thermal comfort and energy efficiency in buildings and have received significant attention from the research community, with data driven methods gaining in popularity. Yet the lack of labeled data, such as normal versus faulty operational status, has slowed the application of machine learning to HVAC systems. In addition, for any particular building, there may be an insufficient number of observed faults over a reasonable amount of time for training. To overcome these challenges, we present a transfer methodology for a novel Bayesian classifier designed to distinguish between normal operations and faulty operations. The key is to train this classifier on a building with a large amount of sensor and fault data (for example, via simulation or standard test data) then transfer the classifier to a new building using a small amount of normal operations data from the new building. We demonstrate a proof-of-concept for transferring a classifier between architecturally similar buildings in different climates and show few samples are required to maintain classification precision and recall.

transfer learning, Building HVAC, Bayesian framewo↗