Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Data Curation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

EXSCLAIM!

Due to recent improvements in image resolution and acquisition speed, materials microscopy is experiencing an explosion of published imaging data. The standard publication format, while sufficient for traditional data ingestion scenarios where a select number of images can be critically examined and curated manually, is not conducive tolarge-scale data aggregation or analysis, hindering data sharing and reuse. Most images in publications are presented as components of a larger figure with their explicit context buried in the main body or caption text, so even if aggregated, collections of images with weak or no digitized contextual labels have limited value. To solve the problem of curating labeled microscopy data from literature, we introduce the EXSCLAIM! Python toolkit for the automatic EXtraction, Separation, and Caption-based natural Language Annotation of IMages from scientific literature. The software is implemented through a three part pipeline: the JournalScraper, which searches the web and downloads figures and captions based on a user provided query, the CaptionDistributor, which separates caption text based on the subfigure each portion of the caption refers to, and the FigueSeparator, which separates figures into component subfigures and extracts other visual information. Also included is a Django user interface for exploring the resulting dataset.

CHAN, MARIA↗

Crowdsourcing biocuration: The Community Assessment of Community Annotation with Ontologies (CACAO)

Experimental data about gene functions curated from the primary literature have enormous value for research scientists in understanding biology. Using the Gene Ontology (GO), manual curation by experts has provided an important resource for studying gene function, especially within model organisms. Unprecedented expansion of the scientific literature and validation of the predicted proteins have increased both data value and the challenges of keeping pace. Capturing literature-based functional annotations is limited by the ability of biocurators to handle the massive and rapidly growing scientific literature. Within the community-oriented wiki framework for GO annotation called the Gene Ontology Normal Usage Tracking System (GONUTS), we describe an approach to expand biocuration through crowdsourcing with undergraduates. This multiplies the number of high-quality annotations in international databases, enriches our coverage of the literature on normal gene function, and pushes the field in new directions. From an intercollegiate competition judged by experienced biocurators, Community Assessment of Community Annotation with Ontologies (CACAO), we have contributed nearly 5,000 literature-based annotations. Many of those annotations are to organisms not currently well-represented within GO. Over a 10-year history, our community contributors have spurred changes to the ontology not traditionally covered by professional biocurators. The CACAO principle of relying on community members to participate in and shape the future of biocuration in GO is a powerful and scalable model used to promote the scientific enterprise. It also provides undergraduate students with a unique and enriching introduction to critical reading of primary literature and acquisition of marketable skills.

59 BASIC BIOLOGICAL SCIENCES↗

OBO Foundry in 2021: operationalizing open data principles to evaluate ontologies

Biological ontologies are used to organize, curate and interpret the vast quantities of data arising from biological experiments. While this works well when using a single ontology, integrating multiple ontologies can be problematic, as they are developed independently, which can lead to incompatibilities. The Open Biological and Biomedical Ontologies (OBO) Foundry was created to address this by facilitating the development, harmonization, application and sharing of ontologies, guided by a set of overarching principles. One challenge in reaching these goals was that the OBO principles were not originally encoded in a precise fashion, and interpretation was subjective. Here, we show how we have addressed this by formally encoding the OBO principles as operational rules and implementing a suite of automated validation checks and a dashboard for objectively evaluating each ontology’s compliance with each principle. This entailed a substantial effort to curate metadata across all ontologies and to coordinate with individual stakeholders. We have applied these checks across the full OBO suite of ontologies, revealing areas where individual ontologies require changes to conform to our principles. Our work demonstrates how a sizable, federated community can be organized and evaluated on objective criteria that help improve overall quality and interoperability, which is vital for the sustenance of the OBO project and towards the overall goals of making data Findable, Accessible, Interoperable, and Reusable (FAIR).

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Evaluation of the PhaseNet Model Applied to the IMS Seismic Network

Producing a complete and accurate set of signal detections is essential for automatically building and characterizing seismic events of interest for nuclear explosion monitoring. Signal detection algorithms have been an area of research for decades, but still produce large quantities of false detections and misidentify real signals that must be detected to produce a complete global catalog of events of interest. Deep learning methods have shown promising capabilities in effectively characterizing seismic signals for complex tasks such as identifying phase arrival times. We use the PhaseNet model, a UNet-based Neural Network, trained on local distance data from northern California to predict seismic arrivals on data from the International Monitoring System (IMS) global network. We use an analyst-curated bulletin generated from this data set to compare the performance of PhaseNet to that of the Short-Term Average/Long-Term Average (STA/LTA) algorithm. We find that PhaseNet has the potential of outperforming traditional processing methods and recommend the training of a new model with the IMS data to achieve optimal performance.

58 GEOSCIENCES↗

Tripal, a community update after 10 years of supporting open source, standards-based genetic, genomic and breeding databases

Abstract Online, open access databases for biological knowledge serve as central repositories for research communities to store, find and analyze integrated, multi-disciplinary datasets. With increasing volumes, complexity and the need to integrate genomic, transcriptomic, metabolomic, proteomic, phenomic and environmental data, community databases face tremendous challenges in ongoing maintenance, expansion and upgrades. A common infrastructure framework using community standards shared by many databases can reduce development burden, provide interoperability, ensure use of common standards and support long-term sustainability. Tripal is a mature, open source platform built to meet this need. With ongoing improvement since its first release in 2009, Tripal provides full functionality for searching, browsing, loading and curating numerous types of data and is a primary technology powering at least 31 publicly available databases spanning plants, animals and human data, primarily storing genomics, genetics and breeding data. Tripal software development is managed by a shared, inclusive governance structure including both project management and advisory teams. Here, we report on the most important and innovative aspects of Tripal after 11 years development, including integration of diverse types of biological data, successful collaborative projects across member databases, and support for implementing FAIR principles.

59 BASIC BIOLOGICAL SCIENCES↗

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration↗

The INTERSECT Open Federated Architecture for the Laboratory of the Future

A federated instrument-to-edge-to-center architecture is needed to autonomously collect, transfer, store, process, curate, and archive scientific data and reduce human-in-the-loop needs with (a) common interfaces to leverage community and custom software, (b) pluggability to permit adaptable solutions, reuse, and digital twins, and (c) an open standard to enable adoption by science facilities world-wide. The Selfdriven Experiments for Science/Interconnected Science Ecosystem (INTERSECT) Open Architecture enables science breakthroughs using intelligent networked systems, instruments and facilities with autonomous experiments, “self-driving” laboratories, smart manufacturing and artificial intelligence (AI) driven design, discovery and evaluation. It creates an open federated architecture for the laboratory of the future using a novel approach, consisting of (1) science use case design patterns, (2) a system of systems architecture, and (3) a microservice architecture.

Engelmann, Christian↗

Decoding substrate specificity determining factors in glycosyltransferase-B enzymes – insights from machine learning models

Substrate specificity is an essential characteristic of any enzyme's function and an understanding of the factors that determine this specificity is crucial for enzyme engineering. Unlike the structure of an enzyme which is directly impacted by its sequence, substrate specificity as an enzyme attribute involves a rather indirect relationship with sequence as it also depends on structural aspects that dictate substrate accessibility and active site dynamics. In this study, we explore the performance of classifier-based machine learning models trained on curated sequence and structural data for a class of glycosyltransferases (GTs), namely GT-Bs, to understand their substrate specificity determining factors. GTs enable the transfer of sugar moieties to other biomolecules such as oligosaccharides or proteins and are found in all kingdoms of life. In plants, GTs participate in the biosynthesis of plant cell wall biopolymers (e.g.: hemicelluloses and pectins) and are an integral part of the enzymatic machinery that enables the storage of carbon and energy as plant biomass. To elucidate the substrate specificity of uncharacterized GT-Bs, we constructed multi-label machine learning models (Support Vector Classifier, K-Nearest Neighbors, Gaussian Naïve-Bayes, Random Forest) that incorporate both sequence and structural features. These models achieve good predictive accuracies on test datasets. However, despite our use of structural information, we highlight that there is further scope for improvement in training these models to draw interpretable relationships between sequence, structure and substrate specificity determining motifs in GT-Bs.

97 MATHEMATICS AND COMPUTING↗

Multi-strain analysis of Pseudomonas putida reveals the metabolic and genetic diversity of the species

Pseudomonas putida is a gram-negative bacterial species increasingly utilized in biotechnology due to its robust growth, ability to degrade aromatic compounds, solvent tolerance, and genetic tractability. In this study, we report a comprehensive multi-strain analysis of 164 P. putida strains based on the reconstruction of a pan-putida metabolic network and the formulation of strain-specific genome-scale metabolic models (GEMs). We performed whole-genome sequencing and hybrid assembly for 40 strains, contributing a ~8% increase to the available genomic data for P. putida . Furthermore, high-throughput phenotypic profiling using the Biolog phenotype microarray system for 24 strains on 190 unique carbon sources, along with 15 aromatic compounds not present on Biolog plates, yielded 4,920 unique strain-phenotype measurements. These data were leveraged to curate GEMs for 24 representative strains, including a refined model for strain KT2440, which comprised 1,480 genes and 2,191 metabolites, achieving a prediction accuracy of 91.2% in carbon utilization. Systematic comparison of genomes and GEMs revealed both conserved core pathways and significant allelic and functional divergence across strains, highlighting strain-specific variation in aromatic degradation. While pathways for protocatechuate and phenylacetate degradation were widely conserved, metabolic capabilities for compounds such as ferulate, phenol, and cresols varied markedly, suggesting adaptation to distinct ecological niches. Alleleome analysis of enzymes, such as PcaI and PcaJ, revealed distinct, functionally similar clades, indicating possible convergent evolution or horizontal gene transfer. These results provide computable resources and informative models for selecting P. putida strains with desired traits for biomanufacturing and bioremediation and offer insights into the evolution and phylogeny of the P. putida species.

aromatics utilization↗

DETECTING FIRE WITH MACHINE LEARNING-ENABLED VISUAL MONITORING FOR NUCLEAR POWER PLANT ENVIRONMENTS

Nuclear power plants are experiencing significant cost challenges to remain competitive with other energy-generation utilities. Unlike other industries, the cost of operation and maintenance activities is mostly attributed to workforce costs. To mitigate this, nuclear power plant stakeholders are increasingly interested in the development and deployment of machine learning methods to potentially automate or augment manually intensive tasks to reduce costs, especially for monitoring activities. One monitoring function that is visually demanding and that can occur frequently to meet the requirements of a fire protection program is visually monitoring an area for fire occurrence. Currently, fire watch activities consist of a worker physically stationed at a given location with the sole responsibility of observing a given area to ensure a fire is detected and mitigated promptly. This effort focused on the development and evaluation of a suitable deep convolutional neural network to classify individual video frames at a sub-second frequency for the occurrence of “fire” and “no fire” in varying industrial environments similar to nuclear power plants. It is believed that a trained neural network model could be integrated with existing facility video surveillance camera feeds to generate alerts when fire inferences occur in individual frames captured at sub-second temporal resolutions. Extensive effort was dedicated to identifying and curating suitable imagery training data representing varying environments and scene settings with and without flame features to maximize generalization in nuclear power plant environments. The data collection effort resulted in the aggregation of a large, labeled image library exceeding 12,000 images to support model training for diverse industrial environments. A deep neural network model incorporating parallel multi-scale capabilities was developed and trained to support accurate image-based detection of flame incidents of varying sizes and spectral feature properties within heterogeneous scenes. Analysis results show that the trained model can achieve high inference accuracy despite heterogeneous scene environments and components. Testing accuracy exceeded 95.0 percent with very low false positive and false negative inferences.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Curation of Ground-Truth Validated Benchmarking Datasets for Fault Detection & Diagnostics Tools

Fault detection and diagnostics (FDD) analytical tools for heating, ventilation and air conditioning (HVAC) systems represent one of the most active areas of smart building technology development. A diversity of techniques is used for FDD analytics, spanning physical models, black box, and rule-based approaches, and researchers continuously strive to develop improved algorithms. With FDD algorithm numbers now in the hundreds, there is a need for performance evaluation of these algorithms in order to assess improvements, improve costeffectiveness, and to prioritize investment in the further development of these technologies. A persistent challenge of FDD advance has been the lack of common datasets to benchmark the performance accuracy of FDD algorithms. This paper summarizes the successful curation of HVAC operational data, paired with validated ground-truth information regarding the presence and absence of faults. The current dataset, consisting of both simulation and experimental data, will evolve to include a larger set of HVAC systems with the objective of creating the largest publicly available dataset to be used by FDD developers, users, and researchers to compare and contrast performance accuracy across FDD algorithms, helping to drive improvements that will spur greater market adoption of FDD tools. Furthermore, in order to avoid previously observed issues with contributed datasets and ensure high quality and consistency of future submissions, the development of data validation and ground-truth assessment protocol is detailed in this study.

Casillas, Armando↗

A Semi-Automated Approach for Curating a Glossary of Key Terms for Open-Source Data Queries

In FY20, the Savannah River National Laboratory (SRNL) was funded by the National Nuclear Security Administration’s Office of Defense Nuclear Non-Proliferation Research and Development (NA-22) to build a machine learning based modeling pipeline that could extract proliferation events of interest from open text-based data sources. As a test case, the research team targeted the identification/fusion of events and indicators that fissile core fabrication would be executed at the Savannah River Site prior to its official announcement in May of 2018. The demonstration prototype proved successful by applying natural language processing and graph theoretical techniques to identify contextual shifts in key words and phrases that acted as indicators that pit production would be carried out at the Savannah River Site up to two years prior to the official announcement.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

An overview of data tools for representing and managing building information and performance data

Building information modeling (BIM) has been widely adopted for representing and exchanging building data across disciplines during building design and construction. However, BIM's use in the building operation phase is limited. With the increasing deployment of low-cost sensors and meters, as well as affordable digital storage and computing technologies, growing volumes of data have been collected from buildings, their energy services systems, and occupants. Such data are crucial to help decision makers understand what, how, and when energy is consumed in buildings—a critical step to improving building performance for energy efficiency, demand flexibility, and resilience. However, practical analyses and use of the collected data are very limited due to various reasons, including poor data quality, ad-hoc representation of data, and lack of data science skills. To unlock value from building data, there is a strong need for a toolchain to curate and represent building information and performance data in common standardized terminologies and schemas, to enable interoperability between tools and applications. This study selected and reviewed 24 data tools based on common use cases of data across the building life cycle, from design to construction, commissioning, operation, and retrofits. The selected data tools are grouped into three categories: (1) data dictionary or terminology, (2) data ontology and schemas, and (3) data platforms. The data are grouped into ten typologies covering most types of data collected in buildings. This study resulted in five main findings: (1) most data representation tools can represent their intended data typologies well, such as Green Button for smart meter data and Brick schema for metadata of sensors in buildings and HVAC systems, but none of the tools cover all ten types of data; (2) there is a need for data schemas to represent the basis of design data and metadata of occupant data; (3) standard terminologies such as those defined in BEDES are only adopted in a few data tools; (4) integrating data across various stages in the building life cycle remains a challenge; and (5) most data tools were developed and maintained by different parties for different purposes, their flexibility and interoperability can be improved to support broader use cases. Finally, recommendations for future research on building data tools are provided for the data and buildings community based on the FAIR principles to make data Findable, Accessible, Interoperable, and Reusable.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Using online tools at the Bovine Genome Database to manually annotate genes in the new reference genome

With the availability of a new highly contiguous Bos taurus reference genome assembly (ARS-UCD1.2), it is the opportune time to upgrade the bovine gene set by seeking input from researchers. Furthermore, advances in graphical genome annotation tools now make it possible for researchers to leverage sequence data generated with the latest technologies to collaboratively curate genes. For many years the Bovine Genome Database (BGD) has provided tools such as the APOLLO genome annotation editor to support manual bovine gene curation. The goal of this paper is to explain the reasoning behind the decisions made in the manual gene curation process while providing examples using the existing BGD tools. We will describe the sources of gene annotation evidence provided at the BGD, including RNAseq and Iso-Seq data. We will also explain how to interpret various data visualizations when curating gene models, and will demonstrate the value of manual gene annotation. The process described here can be applied to manual gene curation for other species with similar tools. With a better understanding of manual gene annotation, researchers will be encouraged to edit gene models and contribute to the enhancement of livestock gene sets.

59 BASIC BIOLOGICAL SCIENCES↗

Data Archive and Portal (DAP) Platform for Solid Phase Processing Technologies

The scale and speed of data generated by modern scientific experiments have constantly challenged the research community to store, curate, manage and optimally use it to drive scientific discoveries. In this work, we have developed a data archive and portal (DAP) platform including analytics capabilities to collect, curate, and manage data and metadata stream for solid phase processing (SPP) techniques. We successfully hosted around ~347K files of data related to processing parameters, microscopic images, and spectroscopic data related to solid phase processing. The DAP platform for SPP will establish an enduring capability to support machine learning and grow collaboration at the intersection of materials science and data science.

36 MATERIALS SCIENCE↗

Carbon Storage Core Characterization Efforts at NETL

The multi-scale Computed Tomography (CT) and core flow facility in the Geocharacterization Laboratory at NETL, Morgantown yields porosity, permeability, and fracture properties of rock core samples obtained from the subsurface while maintaining the integrity of the sample. Additionally, geophysical bulk rock properties are analyzed with the laboratory’s GeoTEK multi-sensor core logger in a comparable fashion to downhole methods. NETL researchers collaborate with stakeholders within the carbon storage, oil and gas, and critical minerals sectors. Since 2017, over 1.88 miles of core have been analyzed within the laboratory and all data is publicly available through the Technical Report Series (TRS) on the Energy Data eXchange (EDX). Additionally, the website, RokBase, was curated to extrapolate and visualize the high-resolution data from field operations. The characterization of the Lively Grove #1 (LG#1) well provides a case study into the full capabilities of the Geocharacterization Laboratory. During the comprehensive study of LG#1 ~1-2 mm in diameter, vertical to bedding, cylindrical structures were identified throughout the St. Peter Formation. These structures are pervasive throughout the St. Peter Formation at depth and are characterized as the trace fossil, Skolithos.

Isom, Shelby L↗

Machine learning analysis of RB-TnSeq fitness data predicts functional gene modules in Pseudomonas putida KT2440

ABSTRACT There is growing interest in engineering Pseudomonas putida KT2440 as a microbial chassis for the conversion of renewable and waste-based feedstocks, and metabolic engineering of P. putida relies on the understanding of the functional relationships between genes. In this work, independent component analysis (ICA) was applied to a compendium of existing fitness data from randomly barcoded transposon insertion sequencing (RB-TnSeq) of P. putida KT2440 grown in 179 unique experimental conditions. ICA identified 84 independent groups of genes, which we call fModules (“functional modules”), where gene members displayed shared functional influence in a specific cellular process. This machine learning-based approach both successfully recapitulated previously characterized functional relationships and established hitherto unknown associations between genes. Selected gene members from fModules for hydroxycinnamate metabolism and stress resistance, acetyl coenzyme A assimilation, and nitrogen metabolism were validated with engineered mutants of P. putida . Additionally, functional gene clusters from ICA of RB-TnSeq data sets were compared with regulatory gene clusters from prior ICA of RNAseq data sets to draw connections between gene regulation and function. Because ICA profiles the functional role of several distinct gene networks simultaneously, it can reduce the time required to annotate gene function relative to manual curation of RB-TnSeq data sets. IMPORTANCE This study demonstrates a rapid, automated approach for elucidating functional modules within complex genetic networks. While Pseudomonas putida randomly barcoded transposon insertion sequencing data were used as a proof of concept, this approach is applicable to any organism with existing functional genomics data sets and may serve as a useful tool for many valuable applications, such as guiding metabolic engineering efforts in other microbes or understanding functional relationships between virulence-associated genes in pathogenic microbes. Furthermore, this work demonstrates that comparison of data obtained from independent component analysis of transcriptomics and gene fitness datasets can elucidate regulatory-functional relationships between genes, which may have utility in a variety of applications, such as metabolic modeling, strain engineering, or identification of antimicrobial drug targets.

09 BIOMASS FUELS↗

Challenges for monitoring and data analytics in a leadership public data repository

The availability and disposition of data has assumed increasing importance in large-scale computational science. Data repositories are evolving to meet new classes of requirements: compliance with government access guidelines, support for reproducibility of experimental results, and long-term availability of data products. The Constellation public data repository at the Oak Ridge Leadership Computing Facility faces these issues while being situated in one of the most productive data centers in the world. While monitoring and operational data analysis are ingrained in the operation of the OLCF’s large-scale high performance computing platforms, data repositories do not have this history of support. Problems faced by Constellation range from data size (over 7 petabytes in current holdings) to analytic complexity (detailed curation is both absolutely necessary for many data sets and absolutely impossible for humans to accomplish in any practical manner) to deployment environment (OLCF storage resources are oriented toward the needs of the compute platforms). In this paper we describe some of the challenges for collecting monitoring and analytic data from a leadership public data repository. We also discuss various strategies we are pursuing in order to address these challenges, from manual data collection to plans for introducing machine learning-based curatorial techniques.

Widener, Patrick [ORNL] (ORCID:0000000258820816)↗