Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Dataset Annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Livewire: Automatic Annotations

Diogenes processes datasets to provide data quality metrics for the Livewire platform and creates standardized data dictionaries from data annotations. Diogenes needs data annotations that clearly outline thenformat and organization of the data. It also relies on the type, class, and unit of each data piece for comprehensive analysis, which it cannot determine independently. The Annotation Tool significantly reduces the time needed to create annotations for Diogenes by generating data annotations with the correct formatting and content. It also employs machine learning and hard-coded models to automatically annotate data class, quality type, and data units.

33 - ADVANCED PROPULSION SYSTEMS↗

Annotation of DOM metabolomes with an ultrahigh resolution mass spectrometry molecular formula library

Current approaches to analyzing metabolomic data often rely on matching MS/MS fragmentation data to sparse libraries or databases. This approach results in limited identification of features, often with less than 10% of the dataset being annotated. A complementary approach is to assign molecular formula to features based on accurate mass measurements, but the platforms commonly used for metabolomics do not have the needed accuracy or resolving power to do this robustly, particularly for larger molecules. Using our newly modified analysis tool, CoreMS, we generated a library of molecular formula from pooled samples analyzed with LC-21T FT-ICR MS. This library successfully annotated approximately 53.2% of features identified from the exometabolome of marine diatom Phaeodactylum tricornutum – a nearly ten-fold increase over the 5.9% annotation rate achieved using a conventional MS/MS library matching approach. Using this FT-ICR MS library approach, we were able to differentiate differences in the exometabolome of P. tricornutum in iron replete and iron limited conditions, with 668 metabolites being differentially expressed (p < 0.05, 2 x intensity difference) under these conditions. The traditional MS/MS fragmentation-based annotation approach only annotated 61 of these metabolites, while our novel pipeline annotated 450 metabolites and revealed 12 metabolites that were significantly more abundant under low iron conditions. Our results demonstrate the utility of ultrahigh resolution mass spectrometry for generating more comprehensive and confident molecular annotations.

21T-FTICR-MS, CoreMS↗

Self-Supervised Cloud Classification

Abstract Low-level marine clouds play a pivotal role in Earth’s weather and climate through their interactions with radiation, heat and moisture transport, and the hydrological cycle. These interactions depend on a range of dynamical and microphysical processes that result in a broad diversity of cloud types and spatial structures, and a comprehensive understanding of cloud morphology is critical for continued improvement of our atmospheric modeling and prediction capabilities moving forward. Deep learning has recently accelerated our ability to study clouds using satellite remote sensing, and machine learning classifiers have enabled detailed studies of cloud morphology. A major limitation of deep learning approaches to this problem, however, is the large number of hand-labeled samples that are required for training. This work applies a recently developed self-supervised learning scheme to train a deep convolutional neural network (CNN) to map marine cloud imagery to vector embeddings that capture information about mesoscale cloud morphology and can be used for satellite image classification. The model is evaluated against existing cloud classification datasets and several use cases are demonstrated, including training cloud classifiers with very few labeled samples, interrogation of the CNN’s learned internal feature representations, cross-instrument application, and resilience against sensor calibration drift and changing scene brightness. The self-supervised approach learns meaningful internal representations of cloud structures and achieves comparable classification accuracy to supervised deep learning methods without the expense of creating large hand-annotated training datasets. Significance Statement Marine clouds heavily influence Earth’s weather and climate, and improved understanding of marine clouds is required to improve our atmospheric modeling capabilities and physical understanding of the atmosphere. Recently, deep learning has emerged as a powerful research tool that can be used to identify and study specific marine cloud types in the vast number of images collected by Earth-observing satellites. While powerful, these approaches require hand-labeling of training data, which is prohibitively time intensive. This study evaluates a recently developed self-supervised deep learning method that does not require human-labeled training data for processing images of clouds. We show that the trained algorithm performs competitively with algorithms trained on hand-labeled data for image classification tasks. We also discuss potential downstream uses and demonstrate some exciting features of the approach including application to multiple satellite instruments, resilience against changing image brightness, and its learned internal representations of cloud types. The self-supervised technique removes one of the major hurdles for applying deep learning to very large atmospheric datasets.

54 ENVIRONMENTAL SCIENCES↗

PlantSegNet: 3D point cloud instance segmentation of nearby plant organs with identical semantics

In this study, we introduce PlantSegNet, a novel neural network model for instance segmentation of nearby objects with similar geometric structures. Our work addresses the challenges of instance segmentation of plant point clouds, including the difficulty of annotating and labeling point clouds, the loss of local structural information in neural network components, and the generation of large numbers of incorrect small clusters due to poor choices of the loss function. One of the key contributions of our approach is a digital twin of sorghum, i.e., a procedural sorghum model, which was used to generate point clouds of sorghum fields. This allowed us to create a large-scale, annotated, synthetic dataset of sorghum plants that we used to train our PlantSegNet model. We demonstrated the effectiveness of our method in segmenting instances of sorghum leaves grown in outdoor field settings. To the best of our knowledge, this is the first study to address this specific instance segmentation problem for plants grown in such a setting. We compared our proposed method with other state-of-the-art methods for indoor settings, including SGPN and TreePartNet, on both synthetic and real data. Furthermore, our results show that PlantSegNet outperforms these methods regarding accuracy, robustness, and efficiency.

97 MATHEMATICS AND COMPUTING↗

Identifying Disinformation Using Rhetorical Devices in Natural Language Models

Foreign disinformation campaigns are strategically organized, extended efforts using disinformation – false or misleading information deliberately placed by an adversary – to achieve some goal. Disinformation campaigns pose severe threats to our nation’s security by misinforming decision makers and negatively influencing their actions when they are operating on limited amounts of evidence. Current efforts rely on subject matter experts to manually identify disinformation, or on computers and traditional natural language processing algorithms to identify patterns in data to calculate the probability that something is disinformation or not. While both have their merits and successes, subject matter experts are unable to keep up with the high volumes of global information and traditional natural language algorithms do not do well in identifying “why” something is disinformation or not. Our hypothesis is that we can identify disinformation by looking at the way someone speaks, in the rhetorical devices they use. We have curated and annotated a dataset designed for multiple natural language processing tasks, but specifically useful for disinformation detection algorithms.

97 MATHEMATICS AND COMPUTING↗

2025 TEM Workshop

The TEM Data Management Workshop will take place on August 26 from 9 a.m. to 12 p.m. MT, and will be held virtually on TEAMS. The primary goal of this workshop is to engage NSUF users and stakeholders in discussions about the data needs for the utilization of AI and ML in the analysis of TEM data. Key topics to be covered include data storage, data sharing, data tagging, metadata inclusion, standardized data formats, data augmentation, and annotated training datasets. Additionally, the workshop will provide valuable insights into resources such as the Nuclear Research Data System (NRDS) for data storage and sharing, as well as open-source codes for data analysis.

Bachhav, Mukesh↗

Dataset: Breaking the barrier of human-annotated training data for machine-learning-aided plant research using aerial imagery

This dataset supports the implementation described in the manuscript "Breaking the Barrier of Human-Annotated Training Data for Machine-Learning-Aided Biological Research Using Aerial Imagery." It comprises UAV aerial imagery used to execute the code available at https://github.com/pixelvar79/GAN-Flowering-Detection-paper. For detailed information on dataset usage and instructions for implementing the code to reproduce the study, please refer to the GitHub repository.

generative and adversarial learning↗

Ecological Trait-Based Digital Categorization of Microbial Genomes for Denitrification Potential

Microorganisms encode proteins that function in the transformations of useful and harmful nitrogenous compounds in the global nitrogen cycle. The major transformations in the nitrogen cycle are nitrogen fixation, nitrification, denitrification, anaerobic ammonium oxidation, and ammonification. The focus of this report is the complex biogeochemical process of denitrification, which, in the complete form, consists of a series of four enzyme-catalyzed reduction reactions that transforms nitrate to nitrogen gas. Denitrification is a microbial strain-level ecological trait (characteristic), and denitrification potential (functional performance) can be inferred from trait rules that rely on the presence or absence of genes for denitrifying enzymes in microbial genomes. Despite the global significance of denitrification and associated large-scale genomic and scholarly data sources, there is lack of datasets and interactive computational tools for investigating microbial genomes according to denitrification trait rules. Therefore, our goal is to categorize archaeal and bacterial genomes by denitrification potential based on denitrification traits defined by rules of enzyme involvement in the denitrification reduction steps. We report the integration of datasets on genome, taxonomic lineage, ecosystem, and denitrifying enzymes to provide data investigations context for the denitrification potential of microbial strains. We constructed an ecosystem and taxonomic annotated denitrification potential dataset of 62,624 microbial genomes (866 archaea and 61,758 bacteria) that encode at least one of the twelve denitrifying enzymes in the four-step canonical denitrification pathway. Our four-digit binary-coding scheme categorized the microbial genomes to one of sixteen denitrification traits including complete denitrification traits assigned to 3280 genomes from 260 bacteria genera. The bacterial strains with complete denitrification potential pattern included Arcobacteraceae strains isolated or detected in diverse ecosystems including aquatic, human, plant, and Mollusca (shellfish). The dataset on microbial denitrification potential and associated interactive data investigations tools can serve as research resources for understanding the biochemical, molecular, and physiological aspects of microbial denitrification, among others. The microbial denitrification data resources produced in our research can also be useful for identifying microbial strains for synthetic denitrifying communities.

59 BASIC BIOLOGICAL SCIENCES↗

Strategies for community-sourced biocuration in bioinformatics: a case study on MIBiG 4.0

Biocuration is essential to transform molecular sequence data into standardized, machine-readable resources. Such curated datasets enable comparative analysis, predictive modeling, and data integration across bioinformatics platforms. While professional biocuration is resource-intensive and usually limited to institutional settings, community-driven approaches can mobilize large-scale annotation of specialized datasets and are more resilient to disruptions in scientific funding. Here, we present a model for community-powered curation applied to the Minimum Information about a Biosynthetic Gene Cluster (MIBiG) repository. Through a framework of workflows for metadata capture, annotation validation, and contributor coordination, the MIBiG 4.0 initiative recruited 267 scientists across 178 institutions from 33 countries, volunteering an estimated 4000 h of work. These efforts expanded the MIBiG repository by 22% and enhanced its usability in downstream molecular data analyses in comparative genomic analyses, natural product discovery, and machine learning applications. We provide strategies and actionable lessons for adopting this model, supporting the sustainability of curated bioinformatics resources central to nucleic acid research and related fields.

biocuration↗

CholecTriplet2021: A benchmark challenge for surgical action triplet recognition

Context-aware decision support in the operating room can foster surgical safety and efficiency by leveraging real-time feedback from surgical workflow analysis. Most existing works recognize surgical activities at a coarse-grained level, such as phases, steps or events, leaving out fine-grained interaction details about the surgical activity; yet those are needed for more helpful AI assistance in the operating room. Recognizing surgical actions as triplets of ‹ instrument, verb, target › combination delivers more comprehensive details about the activities taking place in surgical videos. This paper presents CholecTriplet2021: an endoscopic vision challenge organized at MICCAI 2021 for the recognition of surgical action triplets in laparoscopic videos. Here, the challenge granted private access to the large-scale CholecT50 dataset, which is annotated with action triplet information. In this paper, we present the challenge setup and the assessment of the state-of-the-art deep learning methods proposed by the participants during the challenge. Here, a total of 4 baseline methods from the challenge organizers and 19 new deep learning algorithms from the competing teams are presented to recognize surgical action triplets directly from surgical videos, achieving mean average precision (mAP) ranging from 4.2% to 38.1%. This study also analyzes the significance of the results obtained by the presented approaches, performs a thorough methodological comparison between them, in-depth result analysis, and proposes a novel ensemble method for enhanced recognition. Our analysis shows that surgical workflow analysis is not yet solved, and also highlights interesting directions for future research on fine-grained surgical activity recognition which is of utmost importance for the development of AI in surgery.

60 APPLIED LIFE SCIENCES↗

Rapid Automated Annotation and Analysis of N-Glycan Mass Spectrometry Imaging Data Sets Using NGlycDB in METASPACE

Imaging N-glycan spatial distribution in tissues using mass spectrometry imaging (MSI) is emerging as a promising tool in biological and clinical applications. However, there is currently no high throughput tool for visualization and molecular annotation of N-glycans in MSI data, which significantly slows down data processing and hampers the applicability of this approach. In this work, we present how METASPACE, an open-source cloud engine for molecular annotation of MSI data, can be used to automatically annotate, visualize, analyze, and interpret high-resolution mass spectrometry-based spatial N-glycomics data. METASPACE is an emerging tool in spatial metabolomics, but the lack of compatible glycan databases has limited its application for comprehensive N-glycan annotations from MSI datasets. We created NGlycDB, a public database of N-glycans, by adapting available glycan databases. We demonstrate the applicability of NGlycDB in METASPACE by analyzing MALDI-MSI data from formalin-fixed paraffin-embedded (FFPE) human kidney and mouse lung tissue sections. We added NGlycDB to METASPACE for public use, thus facilitating applications of MSI in glycobiology.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Synthetic data-driven deep learning for label-free autonomous atomic force microscopy

Atomic force microscopy (AFM) is a widely used tool for nanoscale characterization across materials science, energy research, and biology. However, its adoption in high-throughput materials discovery and statistically driven studies remains limited by a strong dependence on expert operator input and by the scarcity of annotated experimental AFM datasets needed to enable data-driven automation. Here, we introduce SimuScan, a synthetic-data–driven framework that enables reliable AFM feature identification, segmentation, and targeted imaging without requiring large manually labeled experimental datasets. SimuScan generates tunable, high-fidelity synthetic AFM images of defined morphologies while incorporating realistic experimental artifacts, including tip–sample convolution, noise, flattening distortions, and surface debris. These datasets are shown to support scalable, label-free training of modern deep learning models for AFM analysis. When integrated into data-driven AFM workflows, SimuScan-trained models can locate and analyze nanoscale structures across large datasets and guide targeted follow-up imaging. We validate this approach on nanostructured surfaces, DNA assemblies, and bacterial cells, demonstrating robust generalization across diverse sample types with minimal operator intervention. More broadly, this work establishes a general strategy for generating explicitly conditioned, task-relevant synthetic data to improve the reliability of downstream models in autonomous microscopy.

Millan-Solsona, Ruben [Oak Ridge National Laborato↗

FaceBase 3: analytical tools and FAIR resources for craniofacial and dental research

ABSTRACT The FaceBase Consortium was established by the National Institute of Dental and Craniofacial Research in 2009 as a ‘big data’ resource for the craniofacial research community. Over the past decade, researchers have deposited hundreds of annotated and curated datasets on both normal and disordered craniofacial development in FaceBase, all freely available to the research community on the FaceBase Hub website. The Hub has developed numerous visualization and analysis tools designed to promote integration of multidisciplinary data while remaining dedicated to the FAIR principles of data management (findability, accessibility, interoperability and reusability) and providing a faceted search infrastructure for locating desired data efficiently. Summaries of the datasets generated by the FaceBase projects from 2014 to 2019 are provided here. FaceBase 3 now welcomes contributions of data on craniofacial and dental development in humans, model organisms and cell lines. Collectively, the FaceBase Consortium, along with other NIH-supported data resources, provide a continuously growing, dynamic and current resource for the scientific community while improving data reproducibility and fulfilling data sharing requirements.

60 APPLIED LIFE SCIENCES↗

Predicting metabolic modules in incomplete bacterial genomes with MetaPathPredict

The reconstruction of complete microbial metabolic pathways using ‘omics data from environmental samples remains challenging. Computational pipelines for pathway reconstruction that utilize machine learning methods to predict the presence or absence of KEGG modules in incomplete genomes are lacking. Here, we present MetaPathPredict, a software tool that incorporates machine learning models to predict the presence of complete KEGG modules within bacterial genomic datasets. Using gene annotation data and information from the KEGG module database, MetaPathPredict employs deep learning models to predict the presence of KEGG modules in a genome. MetaPathPredict can be used as a command line tool or as a Python module, and both options are designed to be run locally or on a compute cluster. Benchmarks show that MetaPathPredict makes robust predictions of KEGG module presence within highly incomplete genomes.

59 BASIC BIOLOGICAL SCIENCES↗

Integration of Simulated and Real Distributed Acoustic Sensor Measurements to Develop AI Enhanced Intelligent Sensing Operations

Distributed acoustic sensors (DAS) have shown remarkable success in monitoring critical civil, energy and transportation assets over long distances and under harsh environmental conditions. Integration of artificial intelligence (AI) technologies with DAS can extend their operational capabilities from conventional tasks (e.g., vibration frequency detection) to more advanced intelligent tasks like anomaly detection. However, these AI technologies often require large labeled/ annotated DAS measurements datasets of the assets in both normal and anomalous operating conditions. This data acquisition task can be both time consuming and costly. This work develops a generative adversarial network based unsupervised domain adaptation framework to build intelligent DAS operational capabilities.

Venketeswaran, Abhishek↗

Expanded genetic variation (SNP) and phenomics (image based) dataset for Populus trichocarpa

The image dataset consists of 11,791 images representing 1,219 genotypes of Populus trichocarpa undergoing in planta regeneration. Genotypes were imaged with a median of four weekly timepoints and a median of two replicates each. A representative and diverse subset of 249 images was annotated using the IDEAS annotation interface (ideas.eecs.oregonstate.edu) and these annotated images were used to train a deep semantic segmentation model (PSPNet), which was deployed for inference over the entire dataset. Annotated classes include specific stages of regeneration (callus and shoot) in addition to unregenerated plant material and background. Statistics of relative tissue area were extracted and used for downstream genetic association mapping in a genome-wide association study. The SNP dataset consists of over 40 million single-nucleotide polymorphisms across 1,323 wild accessions of Populus trichocarpa

09 BIOMASS FUELS↗

Data for "Genetics of flooding tolerance in an F2 Miscanthus sacchariflorus ssp. lutarioriparius × M. sinensis population"

This dataset contains all data and supplementary materials from "Genetics of flooding tolerance in an F2 Miscanthus sacchariflorus ssp. lutarioriparius × M. sinensis population". 1. The dataset S1 table contains the raw phenotypic data collected during the experiment. 2. The dataset S2 table contains the LSmean values for the 24 traits studied. 3. The dataset S3 table contains the TASSEL GBSv2 map, marker information, and genotype data used for mapping. 4. The dataset S4 table contains information on candidate genes found in each of the QTL intervals. 5. The dataset S5 table contains the GO annotations and KEGG enrichment analyses for those candidate genes. 6. The dataset S6 table contains information on the sequences used to classify AP2 ERF transcription factors. 7. The dataset S7 table contains information on AP2 ERF orthologs between Miscanthus and rice based on synteny. 8. Supplementary file 1 contains the ANOVA results using the raw phenotypic data collected from protocol "A". 9. Supplementary file 2 contains the ANOVA results using the raw phenotypic data collected from protocol "B". 10. Supplementary file 3 contains notes on the comparison of SNP calling methods. 11. Supplementary file 4 is a script for analyzing candidate genes found in QTL intervals.

Miscanthus, flood, partial submergence, complete s↗

PyCMG-based Simulation of Volumetric Concrete Microstructure

Concrete is a complex, heterogeneous material with a microstructure composed of aggregates, cement paste, and pores spanning multiple length scales. Understanding this microstructure is critical for advancing the performance, durability, and modeling of concrete-based systems. While experimental imaging such as X-ray computed tomography (XCT) provides valuable insights, generating large datasets with detailed ground truth annotations is both costly and labor-intensive due to challenges in segmenting similar phases, such as aggregates and cement paste, that often share similar attenuation properties. To address this, we developed a pipeline to simulate realistic 3D concrete microstructures using the open-source Python package PyCMG. This simulation effort focuses on generating high-fidelity, annotated microstructures that can serve as training or benchmarking datasets for image analysis, segmentation algorithms, and machine learning models, particularly in scenarios where experimental data is scarce.

Ziabari, Amir [Oak Ridge National Laboratory; ORNL↗