Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Rotational Equivariance for Object Classification Using xView

With the recent addition of large, curated and labeled data sets to the remote sensing discipline, deep learning models have largely surpassed the performance of classical techniques. These deep models, typically Convolutional Neural Networks, are invariant to translation through the use of successive convolution layers which are themselves equivariant to translation. Further, the combination of multiple convolution and pooling layers means that in practice, the model is also approximately invariant to translation. However, until recently these models could only approach rotational invariance through data augmentation. Here we propose using a new model formulation which achieves rotational equaivariance without data augmentation for overhead imagery classification. We utilize the popular xView data set to compare the rotational equivariance formalization against a regular CNN and CNN with rotational data augmentation for the task of image classification.

Bynum, Lucius EJ↗

Smart Methane Emission Detection System Development (Final Report)

Working with the Department of Energy’s National Energy Technology Laboratory, Southwest Research Institute® (SwRI®) developed a system to identify methane leaks reliably, accurately, and autonomously at critical midstream sections of the natural gas distribution network in real- time for the purpose of mitigating methane emissions using Optical Gas Imaging (OGI) cameras. SwRI’s Smart Leak Detection – Methane (SLED/M) adds a high degree of automation to the process of methane leak detection to minimize sources of human error, minimize response time to a leak event, and maximize midstream visibility. Furthermore, SwRI has been working towards integrating Quantitative OGI (QOGI) capabilities into this existing technology. By leveraging Deep Learning, SwRI now has the capability to estimate fugitive emission leak rates quickly and reliably, which allows operators to detect emissions, quantify leak rate, prioritize repairs, and validate the repairs in a single instrument. The next generation QOGI technology leverages the same cameras used in Leak Detection and Repair (LDAR) programs, with improvements in safety and speed for traditional quantification-based repairs, ultimately leading to less overhead cost for the operators. The goals for this research were to develop two types of models with the following goals: 1. Run in real-time on the edge (≥ 12 Hz) 2. Classification: Achieve less than 5% false positive detection 3. Classification: Achieve ≥ 95% methane plume detection rate 4. Regression: achieve ≤ 10 standard cubic feet per hour (scfh) prediction > 70% of the time In order to achieve these results, multiple infrared (IR) and other sensors were investigated in tandem with the midwave IR (MWIR) OGI to provide additional information to train the underlying models. Information on atmospheric conditions including humidity, temperature, pressure, and solar radiation was provided by a weather station. Several machine learning and deep learning architectures and methods, including looking at quantized classification networks and regressions networks, were explored. As further data was collected, curated, and labeled, it allowed for more refined regressive networks to be adequately trained, leading to better insight into the true flow rates being observed. An important valuable deliverable of this research effort was the development of an advanced network which underwent multiple iterations capable of giving a continuous output. The current network has a predicted mean average percentage error (MAPE) of 12.3% just outside our target goal of 10.00%, but an accuracy of 97.78% at ±50 scfh, well within the overall goal for the Department of Energy (DOE) program. Upon closer inspection, it was observed that more than 10% of datapoints contributing to the MAPE predictions were the result of low flow rate predictions and are beyond the sensitivity of instrument measurement as a result of normal operational variation and noise.

03 NATURAL GAS↗

Thermal Performance of Spandrel Assemblies in Glazed Wall Systems: Laboratory Test Design – Challenges and Test Results

Accurate thermal performance calculation procedures for opaque spandrel areas in curtain wall and window wall systems are essential for rating systems when comparing spandrel systems. However, there is a lack of consensus in thermal modeling needed for accurately characterizing heat transfer through spandrel assemblies due to the complex arrangement of materials and structural components. Several studies indicate that conventional 2D thermal simulations may overestimate R-values by 30% compared to physical testing and 3D simulations. Detailed simulations and well-curated laboratory test data are necessary to build confidence in simulation models, which will later be used to develop correlations to improve widely used conventional 2D thermal simulations. This study aims to experimentally test heat transfer through various spandrel assemblies to validate 3D simulation models. Also, the challenges of conducting a thorough testing design along with the solutions would be documented. The team developed a design for testing spandrel assemblies, making appropriate modifications to the existing heat, air, and moisture (HAM) chamber to accommodate the testing needs. Two moveable baffles were designed and fabricated to guide airflow direction parallel to the test article surface. The data acquisition capabilities in the chamber were upgraded to add more than two hundred sensors to the climate and indoor side of the chamber. The goal is to provide a quality dataset for validating complex 3D modeling simulations, which will be used to develop improved thermal simulation techniques that more accurately represent the thermal behavior of spandrel assemblies and their integration within the building envelope. This paper will summarize the results for the boundary conditions of the testing and the temperature variation across different locations of the spandrel assemblies.

Kunwar, Niraj [ORNL] (ORCID:0000000263457652)↗

Compilation of Experimental Yield Data for Spontaneous Fission of 252 Cf

We present a comprehensive compilation and curation of experimental fission yield (FY) data for the spontaneous fission of 252 Cf, extracted from the EXFOR database. The compilation follows a structured methodology developed for prior compilations of neutron-induced fission yields, and incorporates both independent (IFY) and cumulative (CFY) yields. A total of 62 datasets were reviewed, with entries spanning from 1955 to 2021. A significant portion of the literature reports pre-neutron emission yields, which were excluded from the present compilation due to limitations in format compatibility. Each accepted dataset was processed into a standardized JSON format, including metadata, uncertainties, and bibliographic references. Where available, decay radiation information was used to update the FY data using the latest ENSDF evaluations; 237 data points were corrected accordingly. These corrections are fully traceable and preserve original values. The result is a curated dataset suitable for use in nuclear data evaluations. This work is part of an ongoing effort to modernize the handling of FY data and provide evaluators with high-quality, machine-readable experimental inputs

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Natural Language Processing for Text Based Event Extraction: Identifying Events of Interest Related to Worldwide State-Sponsored Civil Nuclear Power

Beginning in FY20, SRNL was funded by the National Nuclear Security Administration’s Office of Defense Nuclear Non-Proliferation Research and Development to develop a prototype natural language processing/natural language understating machine learning-based modeling and analysis pipeline to extract and forecast events of interest from massive open data sources. The working hypothesis within the approach is that contextual shifts in key words and phrases act as indicators of events of interest over time. Therefore, by identifying points in time where contextual shifts occur, events of interest can be extracted along with explicit and implicit connections of entities and activities. The development of the preliminary prototype pipeline proved successful, meriting further testing of the pipeline on more broad topical domains and in a worldwide data environment. Therefore, SRNL, in collaboration with the Sanghani Center for Artificial Intelligence and Data Analytics at Virginia Tech, have continued development with a test case of identifying events of interest related to worldwide state-sponsored civil nuclear power in open data sources. In the first year of this follow-on effort, the team has curated domain-specific data corpuses using an automated scheme and applied the modeling and analysis pipeline. This robust, focused, and efficient approach consists of an ensemble of analyses applied to time dependent word embedding models that are trained on the data corpuses. In this report, the team has demonstrated the capability of the existing pipeline (as development has continued in parallel) by exploring several specific case-studies centered around Rosatom’s international activities regarding the planning, construction, operation, and/or shutdown of nuclear reactors. A basic timeline events has been generated by manually cataloging known “milestone” events that have occurred at reactors in Turkey, Finland, Hungary, and Egypt and compared with the output of the modeling pipeline. In this approach, the team has characterized the lead time using the prototype pipeline, as well as the ability to capture relevant information, which proved 100% successful. A deep dive example of the Akkuyu reactor (Turkey) is presented that shows the breadth of information that can be captured using the approach. In this case study, events were extracted pertaining to the planning/construction of Akkuyu including protests from the population, information campaigns in response to the protests, forged regulatory documents and lawsuits, budgetary/shareholder information, geopolitical tensions, and the various construction milestones. This has demonstrated the pipeline’s utility as a research aid or real-time event extraction tool, where summary-level information and detailed text extractions from millions of articles or Tweets across long time periods can be generated with significantly less effort than current techniques.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Dataset Repository for Investigating Suicide Risk Using Social and Environmental Determinants of Health

Suicide is frequently modeled as a function of genetics and environment, where the latter refers to factors other than direct biological consequences, such as air quality, financial level, social connectivity, transportation and food access, and homelessness status. According to the World Health Organization, clean air, a stable climate, adequate water, sanitation and hygiene, safe chemical use, radiation protection, healthy and safe workplaces, sound agricultural practices, health-supportive cities and built environments, and a preserved natural environment are all prerequisites for good health. Understanding the relationships between these determinants and mental health outcomes requires standardized data that can be included in healthcare programs and health outcome models. There is a wealth of publicly available data on social and environmental factors provided by various US organizations that can benefit the design of health care systems and public health interventions, as well as improve our comprehension of factors that impact health. Such information would not only help improve the understanding of individual and community risk but also identify new risk factors that have not previously been therapeutically targeted, especially in terms of their impact on mental health. However, curating and standardizing such datasets is challenging because they are often recorded at numerous geographical and temporal resolutions and with varying spatial and temporal granularities. To address this challenge, we launched an endeavor in conjunction with the Veterans Health Administration to collect publicly available socioeconomic and environmental determinants of health statistics in the US. In this manuscript, we describe a social and environmental determinants of health (SEDH) datasets repository, data curation documentation, and a pipeline framework for data generation; This effort started in 2020, when we began constructing a scalable pipeline to automate the download, extraction, preparation, analysis, and production of datasets. These datasets have been made available to the VHA and may be shared upon agreement with collaborating organizations.

60 APPLIED LIFE SCIENCES↗

SG50 Data-format Requirement Document for an Automatically Readable, Comprehensive and Curated Experimental Reaction Database MEDUSA

This report constitutes the requirement document that guides the development of the experimental reaction database, MEDUSAL (Machine-readable Experimental Data User App & Library), created by OECD/NEA/WPEC SubGroup 50. Experimental reaction data are usually stored in the EXFOR library in EXFOR format. With MEDUSAL, the WPEC sub-group 50 wants to go beyond the EXFOR format and database to generate a library that is (a) automatically readable, (b) comprehensive, and (c) curated.

Nuclear Criticality Safety Program (NCSP)↗

GraphAide: Advanced Graph-Assisted Query and Reasoning System

Curating knowledge from multiple siloed sources that contain both structured and unstructured data is a major challenge in many real-world applications. Pattern matching and querying represent fundamental tasks in modern data analytics that leverage this curated knowledge. The development of such applications necessitates overcoming several research challenges, including data extraction, named entity recognition, data modeling, and designing query interfaces. Moreover, the explainability of these functionalities is critical for their broader adoption. The emergence of Large Language Models (LLMs) has accelerated the development lifecycle of new capabilities. Nonetheless, there is an ongoing need for domain-specific tools tailored to user activities. The creation of digital assistants has gained considerable traction in recent years, with LLMs offering a promising avenue to develop such assistants utilizing domain-specific knowledge and assumptions. In this context, we introduce an advanced query and reasoning system, GraphAide, which constructs a knowledge graph (KG) from diverse sources and allows to query and reason over the resulting KG. GraphAide harnesses both the KG and LLMs to rapidly develop domain-specific digital assistants. It integrates design patterns from retrieval augmented generation (RAG) and the semantic web to create an agentic LLM application. GraphAide underscores the potential for streamlined and efficient development of specialized digital assistants, thereby enhancing their applicability across various domains.

Purohit, Sumit [BATTELLE (PACIFIC NW LAB)] (ORCID:↗

SG50 Data-format Specifications Document for the Automatically Readable, Comprehensive, and Curated Experimental Reaction Database MEDUSAL

The aim of this document is to lay out a first draft of the specifications for the MEDUSAL database (Machine-readable Experimental Data User App & Library) that is being described by OECD/NEA/WPEC SG-50. The EXFOR database (Otuka et al., 2014) has a format that is based on code-value pairs, and a significant portion of the information in the EXFOR entry is contained in free text sections. Several high-level requirements for the MEDUSAL database, as laid out in the Use Cases and Requirements Working Paper (citation), relate to the definition of the specifications

Nuclear Criticality Safety Program (NCSP)↗

Correlating real-world incidents with vessel traffic off the coast of Hawaii, 2017–2020

Abstract Objectives Because of the high-risk nature of emergencies and illegal activities at sea, it is critical that algorithms designed to detect anomalies from maritime traffic data be robust. However, there exist no publicly available maritime traffic data sets with real-world expert-labeled anomalies. As a result, most anomaly detection algorithms for maritime traffic are validated without ground truth. Data description We introduce the HawaiiCoast_GT data set, the first ever publicly available automatic identification system (AIS) data set with a large corresponding set of true anomalous incidents. This data set—cleaned and curated from raw Bureau of Ocean Energy Management (BOEM) and National Oceanic and Atmospheric Administration (NOAA) automatic identification system (AIS) data—covers Hawaii’s coastal waters for four years (2017–2020) and contains 88,749,176 AIS points for a total of 2622 unique vessels. This includes 208 labeled tracks corresponding to 154 rigorously documented real-world incidents.

99 GENERAL AND MISCELLANEOUS↗

ECP Software Technology Capability Assessment Report

The Exascale Computing Project Software Technology (ECP ST) focus area represents the key bridge between Exascale systems and the scientists developing applications that will run on those platforms. ECP offers a unique opportunity to build a coherent set of software (often referred to as the "software stack") that will allow application developers to maximize their ability to write highly parallel applications, targeting multiple Exascale architectures with runtime environments that will provide high performance and resilience. But applications are only useful if they can provide scientific insight, and the unprecedented data produced by these applications require a complete analysis work ow that includes new technology to scalably collect, reduce, organize, curate, and analyze the data into actionable decisions. This requires approaching scientific computing in a holistic manner, encompassing the entire user workflow - from conception of a problem, setting up the problem with validated inputs, performing high-fidelity simulations, to the application of uncertainty quantification to the final analysis. The software stack plan defined here aims to address all of these needs by extending current technologies to Exascale where possible, by performing the research required to conceive of new approaches necessary to address unique problems where current approaches will not suffice, and by deploying high-quality and robust software products on the platforms developed in the Exascale systems project. The ECP ST portfolio has established a set of interdependent projects that will allow for the research, development, and delivery of a comprehensive software stack,

97 MATHEMATICS AND COMPUTING↗

The Unified Phenotype Ontology : a framework for cross-species integrative phenomics

Phenotypic data are critical for understanding biological mechanisms and consequences of genomic variation, and are pivotal for clinical use cases such as disease diagnostics and treatment development. For over a century, vast quantities of phenotype data have been collected in many different contexts covering a variety of organisms. The emerging field of phenomics focuses on integrating and interpreting these data to inform biological hypotheses. A major impediment in phenomics is the wide range of distinct and disconnected approaches to recording the observable characteristics of an organism. Phenotype data are collected and curated using free text, single terms or combinations of terms, using multiple vocabularies, terminologies, or ontologies. Integrating these heterogeneous and often siloed data enables the application of biological knowledge both within and across species. Existing integration efforts are typically limited to mappings between pairs of terminologies; a generic knowledge representation that captures the full range of cross-species phenomics data is much needed. We have developed the Unified Phenotype Ontology (uPheno) framework, a community effort to provide an integration layer over domain-specific phenotype ontologies, as a single, unified, logical representation. uPheno comprises (1) a system for consistent computational definition of phenotype terms using ontology design patterns, maintained as a community library; (2) a hierarchical vocabulary of species-neutral phenotype terms under which their species-specific counterparts are grouped; and (3) mapping tables between species-specific ontologies. This harmonized representation supports use cases such as cross-species integration of genotype-phenotype associations from different organisms and cross-species informed variant prioritization.

59 BASIC BIOLOGICAL SCIENCES↗

Electricity Baseline 2022

The Electricity Baseline (2022) is a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data and was created using the ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0). The Python package used the "ELCI_2022" model configuration to set the facility and generation data sources and years that were used to create this life cycle inventory, which were taken from publicly accessible datasets and automatically curated into a local data store. An archive of the data stores used in this model is available online: https://doi.org/10.18141/2569193. This model is presented in GreenDelta's openLCA schema v2 JSON-LD format (https://greendelta.github.io/olca-schema/).

Electricity; LCA; data inventory↗

Electricity Baseline 2021

The Electricity Baseline (2021) is a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data and was created using the ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0). The Python package used the "ELCI_2021" model configuration to set the facility and generation data sources and years that were used to create this life cycle inventory, which were taken from publicly accessible datasets and automatically curated into a local data store. An archive of the data stores used in this model is available online: https://doi.org/10.18141/2569576. This model is presented in GreenDelta's openLCA schema v2 JSON-LD format (https://greendelta.github.io/olca-schema/).

Electricity; LCA; LCI; Life Cycle↗

Electricity Baseline 2020

The Electricity Baseline (2020) is a regionalized life cycle inventory model of U.S. electricity generation, consumption, and distribution using standardized facility and generation data and was created using the ElectricityLCI v2 Python package (https://github.com/USEPA/ElectricityLCI/tree/v2.0). The Python package used the "ELCI_2020" model configuration to set the facility and generation data sources and years that were used to create this life cycle inventory, which were taken from publicly accessible datasets and automatically curated into a local data store. An archive of the data stores used in this model is available online: https://doi.org/10.18141/2569605. This model is presented in GreenDelta's openLCA schema v2 JSON-LD format (https://greendelta.github.io/olca-schema/).

Electricity; LCA; LCI; data inventory↗

Multiple Cases of Bacterial Sequence Erroneously Incorporated Into Publicly Available Chloroplast Genomes

Public sequencing databases are invaluable resources to biological researchers, but assessing data veracity as well as the curation and maintenance of such large collections of data can be challenging. Genomes of eukaryotic organelles, such as chloroplasts and other plastids, are particularly susceptible to assembly errors and misrepresentations in these databases due to their close evolutionary relationships with bacteria, which may co-occur within the same environment, as can be the case when sequencing plants. Here, based on sequence similarities with bacterial genomes, we identified several suspicious chloroplast assemblies present in the National Institutes of Health (NIH) Reference Sequence (RefSeq) collection. Investigations into these chloroplast assemblies reveal examples of erroneous integration of bacterial sequences into chloroplast ribosomal RNA (rRNA) loci, often within the rRNA genes, presumably due to the high similarity between plastid and bacterial rRNAs. The bacterial lineages identified within the examined chloroplasts as the most likely source of contamination are either known associates of plants, or co-occur in the same environmental niches as the examined plants. Modifications to the methods used to process untargeted ‘raw’ shotgun sequencing data from whole genome sequencing efforts, such as the identification and removal of bacterial reads prior to plastome assembly, could eliminate similar errors in the future.

59 BASIC BIOLOGICAL SCIENCES↗

Spatial and Temporal Characterization of Activity in Public Space, 2019–2020

The data reported here characterize spatial and temporal variation in the ratio of short-to-long-duration visits in public places (i.e., points of interest) in the United States for each week between January 2019 and December 2020. The underlying data on anonymized and aggregated foot traffic to public places is curated by SafeGraph, a geospatial data provider. In this work, we report the estimated number and duration of “short” (i.e., <4 hours) and “long” (i.e., >4 hours) visits to public places at the US census block group level. Long visits are shown to be a good proxy for workers based on formal economic data. We propose that short visits are more likely to represent nonobligate activities: people visiting a public place for leisure, shopping, entertainment, or civic or cultural engagement. Our work constructs a ratio of short to long visits, which can be used to inform population estimates for nonworker use of public space. These data may be useful for understanding how people’s use of public space has changed during the COVID-19 pandemic and, more generally, for understanding activity patterns in public.

99 GENERAL AND MISCELLANEOUS↗

In-Water Data Acquisition Tool Supports Four Marine Energy Projects

The NREL-developed Modular Ocean Data Acquisition (MODAQ) system is built from open-source hardware and software. The tool can both collect and store data and even share curated information through the cloud. Marine energy developers can work with NREL's engineers to design their own customized MODAQ and capture high-quality field data to help monitor and improve their technology designs. For example, users could assess how much power their device produces at sea, analyze the durability of a specific device component, or even control their device from a desk halfway around the world.

data acquisition↗