Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Harnessing the predicted maize pan-interactome for putative gene function prediction and prioritization of candidate genes for important traits

Abstract The recent assembly and annotation of the 26 maize nested association mapping population founder inbreds have enabled large-scale pan-genomic comparative studies. These studies have expanded our understanding of agronomically important traits by integrating pan-transcriptomic data with trait-specific gene candidates from previous association mapping results. In contrast to the availability of pan-transcriptomic data, obtaining reliable protein–protein interaction (PPI) data has remained a challenge due to its high cost and complexity. We generated predicted PPI networks for each of the 26 genomes using the established STRING database. The individual genome-interactomes were then integrated to generate core- and pan-interactomes. We deployed the PPI clustering algorithm ClusterONE to identify numerous PPI clusters that were functionally annotated using gene ontology (GO) functional enrichment, demonstrating a diverse range of enriched GO terms across different clusters. Additional cluster annotations were generated by integrating gene coexpression data and gene description annotations, providing additional useful information. We show that the functionally annotated PPI clusters establish a useful framework for protein function prediction and prioritization of candidate genes of interest. Our study not only provides a comprehensive resource of predicted PPI networks for 26 maize genomes but also offers annotated interactome clusters for predicting protein functions and prioritizing gene candidates. The source code for the Python implementation of the analysis workflow and a standalone web application for accessing the analysis results are available at https://github.com/eporetsky/PanPPI.

Genetics & Heredity↗

Improve Learning from Crowds via Generative Augmentation

Crowdsourcing provides an efficient label collection schema for supervised machine learning. However, to control annotation cost, each instance in the crowdsourced data is typically annotated by a small number of annotators. This creates a sparsity issue and limits the quality of machine learning models trained on such data. In this paper, we study how to handle sparsity in crowdsourced data using data augmentation. Specifically, we propose to directly learn a classifier by augmenting the raw sparse annotations. We implement two principles of high-quality augmentation using Generative Adversarial Networks: 1) the generated annotations should follow the distribution of authentic ones, which is measured by a discriminator; 2) the generated annotations should have high mutual information with the ground-truth labels, which is measured by an auxiliary network. Extensive experiments and comparisons against an array of state-of-the-art learning from crowds methods on three real-world datasets proved the effectiveness of our data augmentation framework. It shows the potential of our algorithm for low-budget crowdsourcing in general.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Improvement of eukaryotic protein predictions from soil metagenomes

During the last decades, metagenomics has highlighted the diversity of microorganisms from environmental or host-associated samples. Most metagenomics public repositories use annotation pipelines tailored for prokaryotes regardless of the taxonomic origin of contigs. Consequently, eukaryotic contigs with intrinsically different gene features, are not optimally annotated. Using a bioinformatics pipeline, we have filtered 7.9 billion contigs from 6,872 soil metagenomes in the JGI’s IMG/M database to identify eukaryotic contigs. We have re-annotated genes using eukaryote-tailored methods, yielding 8 million eukaryotic proteins and over 300,000 orphan proteins lacking homology in public databases. Comparing the gene predictions we made with initial JGI ones on the same contigs, we confirmed our pipeline improves eukaryotic proteins completeness and contiguity in soil metagenomes. The improved quality of eukaryotic proteins combined with a more comprehensive assignment method yielded more reliable taxonomic annotation. This dataset of eukaryotic soil proteins with improved completeness, quality and taxonomic annotation reliability is of interest for any scientist aiming at studying the composition, biological functions and gene flux in soil communities involving eukaryotes.

54 ENVIRONMENTAL SCIENCES↗

EXSCLAIM!: Harnessing materials science literature for self-labeled microscopy datasets

This work introduces the EXSCLAIM! toolkit for the automatic extraction, separation, and caption-based natural language annotation of images from scientific literature. EXSCLAIM! is used to show how rule-based natural language processing and image recognition can be leveraged to construct an electron microscopy data set containing thousands of keyword-annotated nanostructure images. Moreover, it is demonstrated how a combination of statistical topic modeling and semantic word similarity comparisons can be used to increase the number and variety of keyword annotations on top of the standard annotations from EXSCLAIM! With large-scale imaging datasets constructed from scientific literature, users are well positioned to train neural networks for classification and recognition tasks specific to microscopy-tasks often otherwise inhibited by a lack of sufficient annotated training data.

36 MATERIALS SCIENCE↗

Rapid Adaptation of Chemical Named Entity Recognition Using Few-Shot Learning and LLM Distillation

Named entity recognition (NER) has been widely used in chemical text mining for the automatic identification and extraction of chemical entities. However, existing chemical NER systems primarily focus on scenarios with abundant training data, requiring significant human effort on annotations. This poses challenges for applications in the chemical field, such as catalysis, where many advancements have traditionally relied on trial-and-error investigations and incremental adjustment of variables. This hinders catalysis science and technology progress in addressing emerging energy and environmental crises. In this work, we propose a few-shot NER model that can quickly adapt to extract new types of chemical entities by using only a limited number of annotated examples. Our model employs a metric-learning approach to transfer entity similarity knowledge from high-resource chemical domains (with abundant annotations) to enable effective entity recognition in low-resource specialized domains (limited annotation). We validate the effectiveness of our model on a few-shot chemical NER benchmark built based on six existing chemical NER data sets. Experiments show that the proposed few-shot NER model can achieve reasonable performance with only 5 examples per entity type and shows consistent improvement as the number of examples increases. Furthermore, we demonstrate how the proposed model can be trained with large language model (LLM) annotated data, opening a new pathway for rapid adaptation of NER systems. Furthermore, our approach leverages the knowledge broadness of large language models for chemistry while distilling this knowledge into a lightweight model suitable for efficient and in-house use.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Missing microbial eukaryotes and misleading meta-omic conclusions

Meta-omics is commonly used for large-scale analyses of microbial eukaryotes, including species or taxonomic group distribution mapping, gene catalog construction, and inference on the functional roles and activities of microbial eukaryotes in situ. Here, we explore the potential pitfalls of common approaches to taxonomic annotation of protistan meta-omic datasets. We re-analyze three environmental datasets at three levels of taxonomic hierarchy in order to illustrate the crucial importance of database completeness and curation in enabling accurate environmental interpretation. We show that taxonomic membership of sequence clusters estimates community composition more accurately than returning exact sequence labels, and overlap between clusters can address database shortcomings. Clustering approaches can be applied to diverse environments while continuing to exploit the wealth of annotation data collated in databases, and selecting and evaluating these databases is a critical part of correctly annotating protistan taxonomy in environmental datasets. We argue that ongoing curation of genetic resources is crucial in accurately annotating protists in in situ meta-omic datasets. Moreover, we propose that precise taxonomic annotation of meta-omic data is a clustering problem rather than a feasible alignment problem.

59 BASIC BIOLOGICAL SCIENCES↗

GenomeDepot: data management system for microbial comparative genomics

Summary GenomeDepot is an open-source web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of websites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, Basic Local Alignment Search Tool (BLAST) search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools. Availability and implementation GenomeDepot is open source and distributed under the GNU General Public License via GitHub (https://github.com/aekazakov/genome-depot). GenomeDepot is implemented in Python and was tested in Ubuntu Linux. Full installation instructions and documentation are available at https://aekazakov.github.io/genome-depot/. GenomeDepot demo server is freely accessible at https://iseq.lbl.gov/demogd/.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

GenomeDepot v1.0

GenomeDepot is a web-based platform for annotation, management, and comparative analysis of microbial genomic sequences and associated data including ortholog families, protein domains, operons, regulatory interactions, strain taxonomy, and sample metadata. GenomeDepot supports rapid creation of web-sites for user-defined genome collections that include bioinformatic tools for interactive genome browsing, BLAST search, annotation search, comparative genomic neighborhood visualization, and sequence download. Gene function annotations are generated by a customizable annotation pipeline. The pipeline runs annotation tools in Conda environments and can be easily extended with additional user-specified tools.

Kazakov, Alexey [Lawrence Berkeley National Labora↗

KBase Narrative - Porphyromonadaceae sp. W3.11 genome

Narratives for The phenotype and genotype of fermentative prokaryotes This is the Narrative for Porphyromonadaceae sp. W3.11. A complementary Narrative for Lachnospiraceae sp. C1.1 is available here. This is the Narrative for Lachnospiraceae sp. C1.1. A complementary Narrative for Porphyromonadaceae sp. W3.11 is available here. Background and Isolation This Narrative and its complementary Narrative contain assembly and annotation of two bacterial isolates that were isolated by our laboratory from the rumen of a Holstein heifer. All procedures with animals have been approved by University of California Davis’s Institutional Animal Care and Use Committee. Rumen contents were collected through a rumen fistula and strained through two layers of cheesecloth into a bottle. The bottle was sealed to exclude air and maintained at 39°C. Contents were brought to the laboratory and bubbled under O2-free CO2 within 15 min. At the laboratory, serial dilutions were made with anaerobic dilution solution for Lachnospiraceae sp. C1.1 and propionibacterium diluent for Porphyromonadaceae sp. W3.11 (table S2). Aliquots (0.1 ml) of each dilution were injected into anaerobic bottle plates (1) containing 9 ml of LH medium (table S2). After incubation at 37°C for 7 days, isolated colonies were picked. Lachnospiraceae sp. C1.1 was picked from a bottle inoculated with a 104 dilution of rumen contents, and Porphyromonadaceae sp. W3.11 was picked from a bottle inoculated with a 103 dilution. After initial isolation, these organisms were purified by growing on anaerobic roll tubes (2) and picking isolated colonies. We performed de novo sequencing of Lachnospiraceae sp. C1.1 and Porphyromonadaceae sp. W3.11. Aliquots of liquid culture (9 and 1.5 ml, respectively) were collected by syringe and centrifuged (21,000g for 10 min at 4°C). Cell pellets were submitted to Molecular Research LP for DNA extraction, library preparation, and sequencing. After resuspending pellets in 180 µl of ATL buffer (Qiagen), DNA was extracted using the MagAttract HMW DNA Kit (Qiagen). DNA was eluted in 100 µl of AE buffer (Qiagen) and then cleaned using the DNEasy PowerClean Pro Cleanup Kit (Qiagen). DNA was then sheared using the Covaris g-TUBE (Covaris). Sequencing libraries were prepared using the SMRTbell Express Template Prep Kit 2.0 (Pacific Biosciences) and 1500 ng of the sheared and purified DNA. The SMRTbell libraries were size-selected (>6 Kb) using a BluePippin instrument (Sage Science) and 0.75% agarose gel. Libraries were then sequenced using the PacBio Sequel II (Pacific Biosciences) platform and a 30-hour movie time. Narrative Summary In these Narratives, we filtered low-quality reads using Trimmomatic (v0.36), assembled filtered reads with SPAdes (v3.15.3), and then checked completeness and contamination of the assembled genomes with CheckM (v1.0.18). Statistics for sequencing and assembly are in table S3. Using the assembled contigs (genomes), we called genes and annotated them. Protein-coding genes were called using Prodigal (v2.6.3) (3) locally or using KBase via RASTtk (v1.073), with identical results. Genes were annotated with KO IDs using KAAS (4). They were further annotated with pfam and TIGRFAM IDs using KBase and the Annotate Domains in a Genome app. We classified putative genes for hydrogenases using HydDB. Genes for 16S ribosomal RNA (rRNA) were called using RASTtk (v1.073) in KBase. The contigs (genomes) were analyzed to determine whether they belonged to new species. Taxonomy was assigned using GTDB-Tk (v1.7.0) in KBase. The identity of 16S rRNA genes to other organisms was found using EzBioCloud (5). Values of digital DNA-DNA hybridization (dDDH) were found with Type (Strain) Genome Server (6). These analyses suggest that Lachnospiraceae sp. C1.1 and Porphyromonadaceae sp. W3.11 represent novel species or genera. GTDB-Tk assigned Lachnospiracae sp. C1.1 to family Lachnospiraceae and genus NK4A144, which contains no type strains. It assigned Porphyromonadaceae sp. W3.11 to Porphyromonadaceae and genus Porphyromonas_A. Values of 16S rRNA identity and dDDH with respect to type strains were low (table S4). Although more phenotypic data are needed, available evidence supports assignment of genomes to new species or genera. Related publication Hackmann TJ, Zhang B. The phenotype and genotype of fermentative prokaryotes. Sci Adv. 2023 Sep 29;9(39):eadg8687. doi: 10.1126/sciadv.adg8687. Epub 2023 Sep 27. PMID: 37756392; PMCID: PMC10530074.

Hackmann, Timothy↗

KBase Narrative - Lachnospiraceae sp. C1.1 genome

Narratives for The phenotype and genotype of fermentative prokaryotes This is the Narrative for Porphyromonadaceae sp. W3.11. A complementary Narrative for Lachnospiraceae sp. C1.1 is available here. This is the Narrative for Lachnospiraceae sp. C1.1. A complementary Narrative for Porphyromonadaceae sp. W3.11 is available here. Background and Isolation This Narrative and its complementary Narrative contain assembly and annotation of two bacterial isolates that were isolated by our laboratory from the rumen of a Holstein heifer. All procedures with animals have been approved by University of California Davis’s Institutional Animal Care and Use Committee. Rumen contents were collected through a rumen fistula and strained through two layers of cheesecloth into a bottle. The bottle was sealed to exclude air and maintained at 39°C. Contents were brought to the laboratory and bubbled under O2-free CO2 within 15 min. At the laboratory, serial dilutions were made with anaerobic dilution solution for Lachnospiraceae sp. C1.1 and propionibacterium diluent for Porphyromonadaceae sp. W3.11 (table S2). Aliquots (0.1 ml) of each dilution were injected into anaerobic bottle plates (1) containing 9 ml of LH medium (table S2). After incubation at 37°C for 7 days, isolated colonies were picked. Lachnospiraceae sp. C1.1 was picked from a bottle inoculated with a 104 dilution of rumen contents, and Porphyromonadaceae sp. W3.11 was picked from a bottle inoculated with a 103 dilution. After initial isolation, these organisms were purified by growing on anaerobic roll tubes (2) and picking isolated colonies. We performed de novo sequencing of Lachnospiraceae sp. C1.1 and Porphyromonadaceae sp. W3.11. Aliquots of liquid culture (9 and 1.5 ml, respectively) were collected by syringe and centrifuged (21,000g for 10 min at 4°C). Cell pellets were submitted to Molecular Research LP for DNA extraction, library preparation, and sequencing. After resuspending pellets in 180 µl of ATL buffer (Qiagen), DNA was extracted using the MagAttract HMW DNA Kit (Qiagen). DNA was eluted in 100 µl of AE buffer (Qiagen) and then cleaned using the DNEasy PowerClean Pro Cleanup Kit (Qiagen). DNA was then sheared using the Covaris g-TUBE (Covaris). Sequencing libraries were prepared using the SMRTbell Express Template Prep Kit 2.0 (Pacific Biosciences) and 1500 ng of the sheared and purified DNA. The SMRTbell libraries were size-selected (>6 Kb) using a BluePippin instrument (Sage Science) and 0.75% agarose gel. Libraries were then sequenced using the PacBio Sequel II (Pacific Biosciences) platform and a 30-hour movie time. Narrative Summary In these Narratives, we filtered low-quality reads using Trimmomatic (v0.36), assembled filtered reads with SPAdes (v3.15.3), and then checked completeness and contamination of the assembled genomes with CheckM (v1.0.18). Statistics for sequencing and assembly are in table S3. Using the assembled contigs (genomes), we called genes and annotated them. Protein-coding genes were called using Prodigal (v2.6.3) (3) locally or using KBase via RASTtk (v1.073), with identical results. Genes were annotated with KO IDs using KAAS (4). They were further annotated with pfam and TIGRFAM IDs using KBase and the Annotate Domains in a Genome app. We classified putative genes for hydrogenases using HydDB. Genes for 16S ribosomal RNA (rRNA) were called using RASTtk (v1.073) in KBase. The contigs (genomes) were analyzed to determine whether they belonged to new species. Taxonomy was assigned using GTDB-Tk (v1.7.0) in KBase. The identity of 16S rRNA genes to other organisms was found using EzBioCloud (5). Values of digital DNA-DNA hybridization (dDDH) were found with Type (Strain) Genome Server (6). These analyses suggest that Lachnospiraceae sp. C1.1 and Porphyromonadaceae sp. W3.11 represent novel species or genera. GTDB-Tk assigned Lachnospiracae sp. C1.1 to family Lachnospiraceae and genus NK4A144, which contains no type strains. It assigned Porphyromonadaceae sp. W3.11 to Porphyromonadaceae and genus Porphyromonas_A. Values of 16S rRNA identity and dDDH with respect to type strains were low (table S4). Although more phenotypic data are needed, available evidence supports assignment of genomes to new species or genera. Related publication Hackmann TJ, Zhang B. The phenotype and genotype of fermentative prokaryotes. Sci Adv. 2023 Sep 29;9(39):eadg8687. doi: 10.1126/sciadv.adg8687. Epub 2023 Sep 27. PMID: 37756392; PMCID: PMC10530074.

Hackmann, Timothy↗

Introducing Molecular Hypernetworks for Discovery in Multidimensional Metabolomics Data

Orthogonal separations of data from high-resolution mass spectrometry can provide insight into sample composition and address challenges of complete annotation of molecules in untargeted metabolomics. “Molecular networks” (MNs), as used in the Global Natural Products Social Molecular Networking platform, are a prominent strategy for exploring and visualizing molecular relationships and improving annotation. MNs are mathematical graphs showing the relationships between measured multidimensional data features. MNs also show promise for using network science algorithms to automatically identify targets for annotation candidates and to dereplicate features associated with a single molecular identity. Here, this paper introduces “molecular hypernetworks” (MHNs) as more complex MN models able to natively represent multiway relationships among observations. Compared to MNs, MHNs can more parsimoniously represent the inherent complexity present among groups of observations, initially supporting improved exploratory data analysis and visualization. MHNs also promise to increase confidence in annotation propagation, for both human and analytical processing. We first illustrate MHNs with simple examples, and build them from liquid chromatography- and ion mobility spectrometry-separated MS data. We then describe a method to construct MHNs directly from existing MNs as their “clique reconstructions”, demonstrating their utility by comparing examples of previously published graph-based MNs to their respective MHNs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Molecular Hypernetworks for Exploration of Multi-Dimensional Metabolomics Data (Chyper)

Orthogonal separations of data from high-resolution mass spectrometry can provide insight into sample composition and help address the challenge of complete annotation of molecules in untargeted metabolomics. “Molecular networks” (MNs), as used, for example, in the Global Natural Products Social Molecular Networking platform, are an increasingly popular computational strategy for exploring and visualizing molecular relationships and improving annotation. MNs use graph representations to show the relationships between measured multidimensional data features. MNs also show promise for using network science algorithms to automatically identify targets for annotation candidates and to dereplicate features associated to a single molecular identity. However, more advanced methods may better represent the complexity present in samples. Our work aims to increase confidence in annotation propagation by extending molecular network methods to include “molecular hypernetworks” (MHNs), able to natively represent multiway relationships among observations supporting both human and analytical processing. In this paper we first introduce MHNs illustrated with simple examples, and demonstrate how to build them from liquid chromatography- and ion mobility spectrometry- separated MS data. We then describe a method to construct MHNs directly from existing MNs as their “clique reconstructions”, demonstrating their utility by comparing examples of previously published graph-based MNs to their respective MHNs.

59 BASIC BIOLOGICAL SCIENCES↗

Improving microstructures segmentation via pretraining with synthetic data

Image analysis of material microstructures through microscopy is an integral capability in the field of materials science. The topological and chemical information obtained through microscopy allow us to draw vital connections between material microstructures, properties, and processing. While scanning electron microscopy (SEM) is able to yield a considerable wealth of information interpretable by the intuition of experts, there has been considerable interest in using machine learning, convolutional neural networks (CNNs) in particular, for such image analysis task. Training CNNs for an image analysis task requires a large annotated dataset. However, in many materials science applications, obtaining a large annotated dataset is cost and labor intensive. In this work, we study the use of synthetic data to enlarge the available annotated experimental data of uranium oxide. We utilize a modified Potts model to simulate uranium oxide particles with morphologies similar to those observed experimentally. We then leverage an image-to-image translation model to synthesize the simulated particles as if they are acquired with SEM. Through this process, we obtain pairs of particle images and their corresponding SEM representations, which corresponds to pairs of annotations and images. Unlike previous works, we leverage synthetic data for pretraining a CNN model prior, and finetune that model further with experimental data. We experimentally demonstrate that using synthetic data as incremental learning process benefits the overall performance compared to training a model on combined synthetic and experimental data.

36 MATERIALS SCIENCE↗

Open access repository-scale propagated nearest neighbor suspect spectral library for untargeted metabolomics

Despite the increasing availability of tandem mass spectrometry (MS/MS) community spectral libraries for untargeted metabolomics over the past decade, the majority of acquired MS/MS spectra remain uninterpreted. To further aid in interpreting unannotated spectra, we created a nearest neighbor suspect spectral library, consisting of 87,916 annotated MS/MS spectra derived from hundreds of millions of MS/MS spectra originating from published untargeted metabolomics experiments. Entries in this library, or “suspects,” were derived from unannotated spectra that could be linked in a molecular network to an annotated spectrum. Annotations were propagated to unknowns based on structural relationships to reference molecules using MS/MS-based spectrum alignment. We demonstrate the broad relevance of the nearest neighbor suspect spectral library through representative examples of propagation-based annotation of acylcarnitines, bacterial and plant natural products, and drug metabolism. Our results also highlight how the library can help to better understand an Alzheimer’s brain phenotype. The nearest neighbor suspect spectral library is openly available for download or for data analysis through the GNPS platform to help investigators hypothesize candidate structures for unknown MS/MS spectra in untargeted metabolomics data.

59 BASIC BIOLOGICAL SCIENCES↗

efam: an e xpanded, metaproteome-supported HMM profile database of viral protein fam ilies

Viruses infect, reprogram and kill microbes, leading to profound ecosystem consequences, from elemental cycling in oceans and soils to microbiome-modulated diseases in plants and animals. Although metagenomic datasets are increasingly available, identifying viruses in them is challenging due to poor representation and annotation of viral sequences in databases. Here, we establish efam, an expanded collection of Hidden Markov Model (HMM) profiles that represent viral protein families conservatively identified from the Global Ocean Virome 2.0 dataset. This resulted in 240 311 HMM profiles, each with at least 2 protein sequences, making efam >7-fold larger than the next largest, pan-ecosystem viral HMM profile database. Adjusting the criteria for viral contig confidence from ‘conservative’ to ‘eXtremely Conservative’ resulted in 37 841 HMM profiles in our efam-XC database. To assess the value of this resource, we integrated efam-XC into VirSorter viral discovery software to discover viruses from less-studied, ecologically distinct oxygen minimum zone (OMZ) marine habitats. This expanded database led to an increase in viruses recovered from every tested OMZ virome by ~24% on average (up to ~42%) and especially improved the recovery of often-missed shorter contigs (<5 kb). Additionally, to help elucidate lesser-known viral protein functions, we annotated the profiles using multiple databases from the DRAM pipeline and virion-associated metaproteomic data, which doubled the number of annotations obtainable by standard, single-database annotation approaches. Together, these marine resources (efam and efam-XC) are provided as searchable, compressed HMM databases that will be updated bi-annually to help maximize viral sequence discovery and study from any ecosystem.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

KBase Narrative - Genomic and environmental controls on Castellaniella biogeography in an anthropogenically disturbed site

Genome assemblies were imported into KBase using the Batch Import Assembly from Staging Area (v1.0.57) function. All assemblies were annotated using the Annotated Multiple Microbial Assemblies with RASTtk - v1.073 tool. Annotated genomes were grouped into sets using the Add Genomes to GenomeSet - v1.7.6 function. Individual annotated genomes can be found both below and in the Data menu to the left. Taxonomy was assigned using the Classify Microbes with GTDB-Tk-v1.7.0 tool. The results of this analysis are shown below. Analysis of the Castellaniella pangenome was performed using the Compute Pangenome (v0.0.7) tool. Using the same method, we also computed the ORR-specific and non-ORR Castellaniella pangenomes. All pangenome results (including the presence/absence matrix) can be found below.

Szink, Elizabeth↗

Expanded genetic variation (SNP) and phenomics (image based) dataset for Populus trichocarpa

The image dataset consists of 11,791 images representing 1,219 genotypes of Populus trichocarpa undergoing in planta regeneration. Genotypes were imaged with a median of four weekly timepoints and a median of two replicates each. A representative and diverse subset of 249 images was annotated using the IDEAS annotation interface (ideas.eecs.oregonstate.edu) and these annotated images were used to train a deep semantic segmentation model (PSPNet), which was deployed for inference over the entire dataset. Annotated classes include specific stages of regeneration (callus and shoot) in addition to unregenerated plant material and background. Statistics of relative tissue area were extracted and used for downstream genetic association mapping in a genome-wide association study. The SNP dataset consists of over 40 million single-nucleotide polymorphisms across 1,323 wild accessions of Populus trichocarpa

09 BIOMASS FUELS↗

MaizeMine: A Data Mining Warehouse for the Maize Genetics and Genomics Database

MaizeMine is the data mining resource of the Maize Genetics and Genome Database (MaizeGDB; http://maizemine.maizegdb.org). It enables researchers to create and export customized annotation datasets that can be merged with their own research data for use in downstream analyses. MaizeMine uses the InterMine data warehousing system to integrate genomic sequences and gene annotations from the Zea mays B73 RefGen_v3 and B73 RefGen_v4 genome assemblies, Gene Ontology annotations, single nucleotide polymorphisms, protein annotations, homologs, pathways, and precomputed gene expression levels based on RNA-seq data from the Z. mays B73 Gene Expression Atlas. MaizeMine also provides database cross references between genes of alternative gene sets from Gramene and NCBI RefSeq. MaizeMine includes several search tools, including a keyword search, built-in template queries with intuitive search menus, and a QueryBuilder tool for creating custom queries. The Genomic Regions search tool executes queries based on lists of genome coordinates, and supports both the B73 RefGen_v3 and B73 RefGen_v4 assemblies. The List tool allows you to upload identifiers to create custom lists, perform set operations such as unions and intersections, and execute template queries with lists. When used with gene identifiers, the List tool automatically provides gene set enrichment for Gene Ontology (GO) and pathways, with a choice of statistical parameters and background gene sets. With the ability to save query outputs as lists that can be input to new queries, MaizeMine provides limitless possibilities for data integration and meta-analysis.

59 BASIC BIOLOGICAL SCIENCES↗