Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Dataset Annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Automated Bacterial Identification and Morphological Feature Analysis in Low‐Dose Cryo‐EM Using YOLOv11

Bacteria rapidly adapt to environmental cues through morphological and ultrastructural changes that correlate with physiology and behavior. Cryogenic transmission electron microscopy (cryo‐TEM) can capture these phenotypic changes in near‐native, vitrified states, but manual analysis of low‐dose micrographs is labor intensive and limits throughput. Here, we present an end‐to‐end workflow that combines low‐dose cryo‐TEM imaging with a YOLOv11‐based instance‐segmentation model to automatically identify bacteria and quantify key structural features directly from the micrographs. This workflow enables (i) robust bacterial localization and counting from low‐magnification atlas/montage images, (ii) automated measurements of cell‐envelope (outer–inner membrane) thickness and anisotropy from higher‐magnification views, and (iii) detection and quantification of bacteria–flagella interactions, including overlap length and curvature metrics for interacting versus noninteracting flagella. Using Pantoea sp. YR343 grown under distinct media conditions, we show that the automated measurements agree with manual annotations while substantially reducing analysis time. Together, these tools provide a practical framework for scalable bacterial identification and quantitative phenotyping in low‐dose cryo‐TEM datasets and establish a foundation for extending cryo‐TEM image analysis toward higher‐throughput studies of microbial heterogeneity and biointerfaces.

YOLOv11↗

Metagenome-assembled-genomes recovered from the Arctic drift expedition MOSAiC

The Multidisciplinary Observatory for Study of the Arctic Climate (MOSAiC) expedition consisted of a year-long drifting survey of the Central Arctic Ocean. The ecosystems component of MOSAiC included the sampling of molecular data, with metagenomes collected from a diverse range of environments. The generation of metagenome-assembled-genomes (MAGs) from metagenomes are a starting point for genome-resolved analyses. This dataset presents a catalogue of MAGs recovered from a set of 73 samples from MOSAiC, including 2407 prokaryotic and 56 eukaryotic MAGs, as well as annotations of a near complete eukaryotic MAG using the Joint Genome Institute (JGI) annotation pipeline. The metagenomic samples are from the surface ocean, chlorophyll maximum, mesopelagic and bathypelagic, within leads and under-ice ocean, as well as melt ponds, ice ridges, and first- and second-year sea ice. This set of MAGs can be used to benchmark microbial biodiversity in the Central Arctic Ocean, compare individual strains across space and time, and to study changes in Arctic microbial communities from the winter to summer, at a genomic level.

59 BASIC BIOLOGICAL SCIENCES↗

Reverse engineering environmental metatranscriptomes clarifies best practices for eukaryotic assembly

Abstract Background Diverse communities of microbial eukaryotes in the global ocean provide a variety of essential ecosystem services, from primary production and carbon flow through trophic transfer to cooperation via symbioses. Increasingly, these communities are being understood through the lens of omics tools, which enable high-throughput processing of diverse communities. Metatranscriptomics offers an understanding of near real-time gene expression in microbial eukaryotic communities, providing a window into community metabolic activity. Results Here we present a workflow for eukaryotic metatranscriptome assembly, and validate the ability of the pipeline to recapitulate real and manufactured eukaryotic community-level expression data. We also include an open-source tool for simulating environmental metatranscriptomes for testing and validation purposes. We reanalyze previously published metatranscriptomic datasets using our metatranscriptome analysis approach. Conclusion We determined that a multi-assembler approach improves eukaryotic metatranscriptome assembly based on recapitulated taxonomic and functional annotations from an in-silico mock community. The systematic validation of metatranscriptome assembly and annotation methods provided here is a necessary step to assess the fidelity of our community composition measurements and functional content assignments from eukaryotic metatranscriptomes.

Krinos, Arianna I. (ORCID:0000000197678392)↗

Efforts and Innovations to Promote Data Sharing and Data Accessibility in the EGS Collab Experiments

Large-scale scientific research programs such as the EGS Collab experiments and the FORGE program play extremely important roles in advancing geothermal technologies. Such efforts involve a large number of researchers from multiple institutions, last multiple years, and generate large, complex datasets. The value of the research efforts is only realized when the datasets are used by a large community of researchers in the decades to come. A challenge is that due to the complexity of the data, it could require a user to devote a serious effort into understanding the data before the data can be effectively utilized. In the EGS Collab experiments, the team has devoted remarkable efforts and developed many innovative solutions to make the data more accessible to broader team members and future users. This paper documents the experience gained and lessons learned by the EGS Collab team in disseminating the data in the most informative and inspiring forms to maximize the value of this precious dataset. "Accessibility"in the title does not only mean making the raw data available for download. Particularly, we want to emphasize the importance of organizing, annotating, and presenting the data in ways to make it easy to digest by prospective consumers of the data.

circulation↗

Constructing Self-Labeled Materials Imaging Datasets from Open Access Scientific Journals with EXSCLAIM!

Due to recent improvements in image resolution and acquisition speeds, materials microscopy is experiencing an explosion in imaging data. Yet, despite the volume of images generated, the overall accessibility landscape is highly fragmented, as researchers who do release images to the public, often only do so as snapshots of their larger private dataset in context of scientific journal publications. The effort to automatically consolidate images and descriptive information from web-based platforms has garnered broad attention from the computer vision, language technologies, and chemistry/materials informatics communities. However, these methods are problematic for scientific figures because over 30% of figures are compound in nature, and it is the individual images themselves, paired with relevant context, that are necessary for construction a proper labeled dataset. To this end, we outline in this paper the design of a software pipeline for the automatic EXtraction, Separation, and Caption-based natural Language Annotation of IMages from scientific figures (EXSCLAIM!). Successful consolidation of materials imaging across literature sources will enhance navigation and searchability of materials microscopy images for both novice and experienced researchers, as well as establish the framework necessary for users to search by images, text, or some combination of both.

36 MATERIALS SCIENCE↗

Hyaloscypha finlandica Metabolome Repository

This repository provides the curated data tables, manuscript figure and table exports, dependency records, and workflow scripts supporting an integrated comparative genomics and untargeted LC-MS/MS metabolomics analysis of Hyaloscypha finlandica strain PMI 746, a root-associated dark septate endophyte of poplar. The repository includes genome-mining summaries from antiSMASH, FunBGCeX, BGC-Prophet, and BiG-SCAPE; processed metabolomics inputs; metabolite annotation evidence; statistical outputs; and publication-facing figures and tables. Raw LC-MS/MS spectra, full genome/protein downloads, and large generated tool outputs are referenced through public archive/accession records and are not stored in Git.

59 BASIC BIOLOGICAL SCIENCES↗

What you get is not always what you see—pitfalls in solar array assessment using overhead imagery

Effective integration planning for small, distributed solar photovoltaic (PV) arrays into electric power grids requires access to high quality data: the location and power capacity of individual solar PV arrays. Unfortunately, national databases of small-scale solar PV do not exist; those that do are limited in their spatial resolution, typically aggregated up to state or national levels. While several promising approaches for solar PV detection have been published, strategies for evaluating the performance of these models are often highly heterogeneous from study to study. The resulting comparison of these methods for practical applications for energy assessments becomes challenging and may imply that the reported performance evaluations overly optimistic. The heterogeneity comes in many forms, each of which we explore in this work: the degree of diversity of the locations and sensors (e.g. different satellites, aerial photography) from which the training and validation data originate, the validation of ground truth (manual annotation of imagery vs known solar PV locations), the level of spatial aggregation (e.g. array-level vs regional estimates), and inconsistencies in the training and validation datasets (e.g. different datasets are used for each study and those data are not always made accessible). For each, we discuss emerging practices from the literature to address them or suggest directions of future research. As part of our investigation, we evaluate solar PV identification performance in two large regions: the entire state of Connecticut and the city of San Diego, CA. In Connecticut, we also use 33,114 known parcel-level solar PV installations from Berkeley Lab’s Tracking the Sun dataset to evaluate parcel-level performance and evaluate capacity estimates using 169 municipalities. We also make our code (which we call SolarMapper), pre-trained models, training data, and predictions publicly available and provide a web portal for interactively inspecting each prediction that was made. Here our findings suggest that traditional performance evaluation of the automated identification of solar PV from satellite imagery may be optimistic due to common limitations in the validation process. The takeaways from this work are intended to inform and catalyze the large-scale practical application of automated solar PV assessment techniques by energy researchers and professionals.

14 SOLAR ENERGY↗

Comparative mitogenomics of kingdom Fungi – evolutionary insights and metagenomic applications

Mitochondria are essential components of eukaryotic cells, responsible for ATP production through oxidative phosphorylation. Despite their biological importance, unique challenges have hindered the adoption of automated mitochondrial genome (mitogenome) annotation methods, obstructing mitochondrial comparative genomics in a broad evolutionary context. Using Fungi as a study system and a Joint Genome Institute (JGI) annotated high-quality reference set, we observed broad patterns of mitochondrial evolution across the kingdom. We found that the median fungal mitogenome size is 58 kb and identified exceptionally large examples over 1 Mb in Pezizomycetes. All 14 expected oxidative phosphorylation protein-coding genes, plus rps3, were generally conserved. We found evidence of major evolutionary transitions within the Ascomycota, including the transfer of mitochondrially encoded atp8 and atp9 to the nuclear genomes across the Pezizomycotina and shifts in mitogenome tRNA patterns across the kingdom. We found substantial concordance between mitochondrial and nuclear evolution, enabling us to document 3131 total fungal mitogenomes from JGI-derived metagenomic datasets. We also identified 6467 total undeclared mitogenomes embedded in Genbank fungal nuclear assemblies. We provide interactive tools for mitogenome analysis through the JGI MycoCosm platform. Collectively, this work generated nearly 10 000 new fungal mitogenome annotations, providing a foundation and resources for future exploration of comparative fungal mitogenomics.

Ahrendt, Steven R. [USDOE Joint Genome Institute (↗

Database of virus genomes from ultra-deep sequencing of wastewater

Researchers at University of Missouri have conducted ultra-deep RNA sequencing of viral concentrates from wastewater (1 billion Illumina reads per sample). The resulting dataset spans 321 samples collected weekly from 11 cities between 2023-2025. As part of a tri-lab collaboration, scientists at LLNL and LANL cleaned, assembled, and annotated this metagenomic data, identifying nearly 200,000 viral genomes. Careful data curation resulted in a database containing 21,015 high-quality, near-complete viral genomes from wastewater. This database contains viruses predicted to infect a range of hosts including bacteria (most common viruses), plants (most abundant viruses), and vertebrates (rarest viruses). There are also numerous novel viruses that could not be well identified and whose host(s) are unknown. Just 7% of all genomes in the wastewater virus database had genus-level matches in the public NCBI database, and 17% matched to a recently created metagenomic virus database at that level (metaVR). The database will provide baseline information about viruses in wastewater that may be used to additional identify novel viruses during ongoing monitoring

Allen, Jonathan [Lawrence Livermore National Labor↗

Conditional Filamentation of Paraburkholderia elongata 5N - Data Container

Overview of Dataset This narrative contains assemblies for all Paraburkholderia discussed in Karasz et al. 2022, where their phosphate-solubilizing activity was measured. Assemblies were downloaded from NCBI and annotated with Prokka. These genomes were created in many Narratives and collected here to be more accessible Note: The following naming of assemblies does not correspond with current taxonomic names for the following: Paraburkholderia 5N = Paraburkholderia elongata 5NT Paraburkholderia 1N = Paraburkholderia solitsuga 1NT Paraburkholderia vancouverensis = Paraburkholderia madseniana RL16

59 BASIC BIOLOGICAL SCIENCES↗

Integrative analysis of the 3D genome and epigenome in mouse embryonic tissues

While a rich set of putative cis-regulatory sequences involved in mouse fetal development have been annotated recently on the basis of chromatin accessibility and histone modification patterns, delineating their role in developmentally regulated gene expression continues to be challenging. To fill this gap, here we mapped chromatin contacts between gene promoters and distal sequences across the genome in seven mouse fetal tissues and across six developmental stages of the forebrain. We identified 248,620 long-range chromatin interactions centered at 14,138 protein-coding genes and characterized their tissue-to-tissue variations and developmental dynamics. Integrative analysis of the interactome with previous epigenome and transcriptome datasets from the same tissues revealed a strong correlation between the chromatin contacts and chromatin state at distal enhancers, as well as gene expression patterns at predicted target genes. We predicted target genes of 15,098 candidate enhancers and used them to annotate target genes of homologous candidate enhancers in the human genome that harbor risk variants of human diseases. We present evidence that schizophrenia and other adult disease risk variants are frequently found in fetal enhancers, providing support for the hypothesis of fetal origins of adult diseases.

59 BASIC BIOLOGICAL SCIENCES↗

A global soil plasmidome resource unveils functional and ecological roles of plasmids in soil microbiomes

Plasmids play significant roles in microbial adaptation to ecosystems, yet their dynamics remain poorly understood due to identification challenges. We present the Global Soil Plasmidome Resource (GSPR), a comprehensive dataset of 98,728 plasmid sequences amassed from 6860 terrestrial microbial communities and isolates. We explore this resource through various computational approaches, including phylogenetic diversity analysis, host prediction, and extensive functional annotation, to understand the contribution of plasmids to the genetic and functional diversity in soil, correlating these findings with sample type, as well as the soil habitat they were retrieved from. Our analysis reveals insights into plasmid-encoded functions such as effector modules, quorum sensing, and stress resistance, which may contribute to their persistence and microbial adaptation in soil. Furthermore, CRISPR analysis suggests a prevalent role of these elements related to intra-plasmid competition. By contrasting plasmids from cultivated and uncultivated organisms, we identify important functions that expand existing knowledge of plasmid roles in these habitats. This study represents a notable step forward in elucidating plasmid diversity and function within soil microbiomes and establishes a foundational framework for exploring their roles in natural environments.

Fiamenghi, Mateus B↗

Dataset for Leveraging CryoEM and AI-Driven Morphological Feature Analysis for Insights on Bacterial Structures

This repository hosts an AI-assisted image segmentation and analysis pipeline for Pantoea sp. YR343 cryo-electron microscopy (cryoEM) datasets. The workflow automates membrane thickness measurements, flagella detection, and field-of-view (FOV) screening from low-dose, high-resolution cryoEM micrographs eliminating the need for slow manual annotation. By integrating deep-learning based segmentation (YOLOv11) with quantitative post-processing, this toolkit provides a scalable and reproducible way to study bacterial morphology under hydrated, near-native conditions. The GitHub repository for AI-based tools for cryoEM bacteria ultrastructures can be found here: https://github.com/Sireesiru/Cryo-EM-Ultrastructures/tree/main

60 APPLIED LIFE SCIENCES↗

Improving an Acoustic Vehicle Detector Using an Iterative Self-Supervision Procedure

In many non-canonical data science scenarios, obtaining, detecting, attributing, and annotating enough high-quality training data is the primary barrier to developing highly effective models. Moreover, in many problems that are not sufficiently defined or constrained, manually developing a training dataset can often overlook interesting phenomena that should be included. To this end, we have developed and demonstrated an iterative self-supervised learning procedure, whereby models are successfully trained and applied to new data to extract new training examples that are added to the corpus of training data. Successive generations of classifiers are then trained on this augmented corpus. Using low-frequency acoustic data collected by a network of infrasound sensors deployed around the High Flux Isotope Reactor and Radiochemical Engineering Development Center at Oak Ridge National Laboratory, we test the viability of our proposed approach to develop a powerful classifier with the goal of identifying vehicles from continuously streamed data and differentiating these from other sources of noise such as tools, people, airplanes, and wind. Using a small collection of exhaustively manually labeled data, we test several implementation details of the procedure and demonstrate its success regardless of the fidelity of the initial model used to seed the iterative procedure. Finally, we demonstrate the method’s ability to update a model to accommodate changes in the data-generating distribution encountered during long-term persistent data collection.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Plant sulfate transporter protein sequences for phylogenetic analysis

Sulfur is an essential macronutrient that supports plant growth, development, and responses to environmental stress. Sulfate is the predominant inorganic form of sulfur in soils, and its uptake by roots and translocation to shoots are facilitated by the sulfate transporter (SULTR) family of proteins. Although the first plant SULTR gene was identified nearly three decades ago, several subfamily members, particularly those in the expansive and angiosperm-specific SULTR3 group, remain poorly characterized. To support comprehensive phylogenetic and sequence-based analyses, we compiled a curated dataset of 262 SULTR protein sequences from 22 plant species spanning the evolutionary breadth of land plants. This collection includes representatives from two basal lineages, two early-divergent angiosperms, six monocots, and ten dicots. All sequences were extracted from genome assemblies available in Phytozome v13 (Joint Genome Institute) and manually curated, with cross-referencing to additional databases such as NCBI when needed. This dataset provides a valuable resource for reconstructing the evolutionary history of the SULTR family, with particular emphasis on the diversification of SULTR3 transporters in flowering plants. This resource may also support functional annotation, comparative genomics, and structural modeling of sulfate transport proteins.

CBI↗

Data and scripts associated with a manuscript investigating dissolved organic matter and microbial community linkages across seven globally distributed rivers

This data package is associated with the publication “Meta-metabolome ecology reveals that geochemistry and microbial functional potential are linked to organic matter development across seven rivers” submitted to Science of the Total Environment. This data package includes the data necessary to replicate the analyses presented within the manuscript to investigate dissolved organic matter (DOM) development across broad spatial distances and within divergent biomes. Specifically, we included the Fourier transform ion cyclotron mass spectrometry (FTICR-MS) data, geochemistry data, annotated metagenomic data, and results from ecological null modeling analyses in this data package. Additionally, we included the scripts necessary to generate the figures from the manuscript. Complete metagenomic data associated with this data package can be found at the National Center for Biotechnology (NCBI) under Bioproject PRJNA946291. This dataset consists of (1) four folders; (2) a file-level metadata (flmd) file; (3) a data dictionary (dd) file; (4) a factor sheet describing samples; and (5) a readme. The FTICR Data folder contains (1) the processed Fourier transform ion cyclotron mass spectrometry (FTICR-MS) data; (2) a transformation-weighted characteristics dendrogram generated from the FTICR-MS data; and (3) the script used to generate all FTICR-MS related figures. The Geochemical Data folder contains (1) the single geochemistry data file and (2) the R script responsible for generating associated figures. The Metagenomic Data folder contains (1) annotation information across different levels; (2) carbohydrate active enzyme (CAZyme) information from the dbCAN database (Yin et al., 2012); (3) phylogenetic tree data (FASTAs, alignments, and tree file); and (4) the scripts necessary to analyze all of these data and generate figures. The Null Modeling Data folder contains (1) data generated during null modeling for each river and all rivers combined and (2) the R scripts necessary to process the data. All files are .csv, .pdf, .tsv, .tre, .faa, .afa, .tree, or .R.

54 ENVIRONMENTAL SCIENCES↗

Philympics 2021: Prophage Predictions Perplex Programs

Most bacterial genomes contain integrated bacteriophages—prophages—in various states of decay. Many are active and able to excise from the genome and replicate, while others are cryptic prophages, remnants of their former selves. Over the last two decades, many computational tools have been developed to identify the prophage components of bacterial genomes, and it is a particularly active area for the application of machine learning approaches. However, progress is hindered and comparisons thwarted because there are no manually curated bacterial genomes that can be used to test new prophage prediction algorithms. Here, we present a library of gold-standard bacterial genome annotations that include manually curated prophage annotations, and a computational framework to compare the predictions from different algorithms. We use this suite to compare all extant stand-alone prophage prediction algorithms to identify their strengths and weaknesses. We provide a FAIR dataset for prophage identification, and demonstrate the accuracy, precision, recall, and f 1 score from the analysis of seven different algorithms for the prediction of prophages. We discuss caveats and concerns in this analysis and how those concerns may be mitigated.

Roach, Michael J.↗

Automated Quantification of Wind Turbine Blade Leading Edge Erosion from Field Images

Wind turbine blade leading edge erosion is a major source of power production loss and early detection benefits optimization of repair strategies. Two machine learning (ML) models are developed and evaluated for automated quantification of the areal extent, morphology and nature (deep, shallow) of damage from field images. The supervised ML model employs convolutional neural networks (CNN) and learns features (specific types of damage) present in an annotated set of training images. The unsupervised approach aggregates pixel intensity thresholding with calculation of pixel-by-pixel shadow ratio (PTS) to independently identify features within images. The models are developed and tested using a dataset of 140 field images. The images sample across a range of blade orientation, aspect ratio, lighting and resolution. Each model (CNN v PTS) is applied to quantify the percent area of the visible blade that is damaged and classifies the damage into deep or shallow using only the images as input. Both models successfully identify approximately 65% of total damage area in the independent images, and both perform better at quantifying deep damage. The CNN is more successful at identifying shallow damage and exhibits better performance when applied to the images after they are preprocessed to a common blade orientation.

Aird, Jeanie A.↗