Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

The reference genome for the northeastern Pacific bull kelp, Nereocystis luetkeana

Bull kelp, Nereocystis luetkeana, is a northeastern Pacific kelp with broad distribution from Alaska to central California. Its population declines have caused severe concerns in northern California, the Salish Sea in Washington, and recently in some populations in Oregon. Despite bull kelp's accumulated ecological and physiological studies, an assembled and annotated genomic reference was still unavailable. Here, we report the complete and annotated genome of Nereocystis luetkeana, produced by the California Conservation Genomics Project (CCGP), which aims to reveal genomic diversity patterns across California by sequencing the complete genomes of approximately 150 carefully selected species. The genome was assembled into 1562 scaffolds with 449.82 Mb, 80x of coverage and 22 952 gene models. BUSCO assembly showed a completeness score of 72% for the stramenopiles gene set. The mitochondria and chloroplast genome sequences have 37 Kb and 131 Mb, respectively. The orthology analysis between 10 Phaeophycean genomes showed 1065 expanded and 286 unique orthogroups for this species. Pairwise comparisons showed 542 orthogroups present only in N. luetkeana and M. pyrifera, another large-body kelp. The enrichment analysis of these orthogroups showed important functions related to central metabolism and signaling due to ATPases enrichment in these two species. This genome assembly will provide an essential resource for the ecology, evolution, conservation, and breeding of bull kelp.

California Conservation Genomics Project—CCGP↗

Predictive Models of Genetic Redundancy in Arabidopsis thaliana

Abstract Genetic redundancy refers to a situation where an individual with a loss-of-function mutation in one gene (single mutant) does not show an apparent phenotype until one or more paralogs are also knocked out (double/higher-order mutant). Previous studies have identified some characteristics common among redundant gene pairs, but a predictive model of genetic redundancy incorporating a wide variety of features derived from accumulating omics and mutant phenotype data is yet to be established. In addition, the relative importance of these features for genetic redundancy remains largely unclear. Here, we establish machine learning models for predicting whether a gene pair is likely redundant or not in the model plant Arabidopsis thaliana based on six feature categories: functional annotations, evolutionary conservation including duplication patterns and mechanisms, epigenetic marks, protein properties including posttranslational modifications, gene expression, and gene network properties. The definition of redundancy, data transformations, feature subsets, and machine learning algorithms used significantly affected model performance based on holdout, testing phenotype data. Among the most important features in predicting gene pairs as redundant were having a paralog(s) from recent duplication events, annotation as a transcription factor, downregulation during stress conditions, and having similar expression patterns under stress conditions. We also explored the potential reasons underlying mispredictions and limitations of our studies. This genetic redundancy model sheds light on characteristics that may contribute to long-term maintenance of paralogs, and will ultimately allow for more targeted generation of functionally informative double mutants, advancing functional genomic studies.

59 BASIC BIOLOGICAL SCIENCES↗

DRAM for distilling microbial metabolism to automate the curation of microbiome function

Abstract Microbial and viral communities transform the chemistry of Earth's ecosystems, yet the specific reactions catalyzed by these biological engines are hard to decode due to the absence of a scalable, metabolically resolved, annotation software. Here, we present DRAM (Distilled and Refined Annotation of Metabolism), a framework to translate the deluge of microbiome-based genomic information into a catalog of microbial traits. To demonstrate the applicability of DRAM across metabolically diverse genomes, we evaluated DRAM performance on a defined, in silico soil community and previously published human gut metagenomes. We show that DRAM accurately assigned microbial contributions to geochemical cycles and automated the partitioning of gut microbial carbohydrate metabolism at substrate levels. DRAM-v, the viral mode of DRAM, established rules to identify virally-encoded auxiliary metabolic genes (AMGs), resulting in the metabolic categorization of thousands of putative AMGs from soils and guts. Together DRAM and DRAM-v provide critical metabolic profiling capabilities that decipher mechanisms underpinning microbiome function.

59 BASIC BIOLOGICAL SCIENCES↗

IMG/VR v4: an expanded database of uncultivated virus genomes within a framework of extensive functional, taxonomic, and ecological metadata

Viruses are widely recognized as critical members of all microbiomes. Metagenomics enables large-scale exploration of the global virosphere, progressively revealing the extensive genomic diversity of viruses on Earth and highlighting the myriad of ways by which viruses impact biological processes. IMG/VR provides access to the largest collection of viral sequences obtained from (meta)genomes, along with functional annotation and rich metadata. A web interface enables users to efficiently browse and search viruses based on genome features and/or sequence similarity. Here, for this work, we present the fourth version of IMG/VR, composed of >15 million virus genomes and genome fragments, a ≈6-fold increase in size compared to the previous version. These clustered into 8.7 million viral operational taxonomic units, including 231 408 with at least one high-quality representative. Viral sequences in IMG/VR are now systematically identified from genomes, metagenomes, and metatranscriptomes using a new detection approach (geNomad), and IMG standard annotation are complemented with genome quality estimation using CheckV, taxonomic classification reflecting the latest taxonomic standards, and microbial host taxonomy prediction. IMG/VR v4 is available at https://img.jgi.doe.gov/vr, and the underlying data are available to download at https://genome.jgi.doe.gov/portal/IMG_VR.

59 BASIC BIOLOGICAL SCIENCES↗

CFM-ID 4.0 – a web server for accurate MS-based metabolite identification

The CFM-ID 4.0 web server (https://cfmid.wishartlab.com) is an online tool for predicting, annotating and interpreting tandem mass (MS/MS) spectra of small molecules. It is specifically designed to assist researchers pursuing studies in metabolomics, exposomics and analytical chemistry. More specifically, CFM-ID 4.0 supports the: 1) prediction of electrospray ionization quadrupole time-of-flight tandem mass spectra (ESI-QTOF-MS/MS) for small molecules over multiple collision energies (10 eV, 20 eV, and 40 eV); 2) annotation of ESI-QTOF-MS/MS spectra given the structure of the compound; and 3) identification of a small molecule that generated a given ESI-QTOF-MS/MS spectrum at one or more collision energies. The CFM-ID 4.0 web server makes use of a substantially improved MS fragmentation algorithm, a much larger database of experimental and in silico predicted MS/MS spectra and improved scoring methods to offer more accurate MS/MS spectral prediction and MS/MS-based compound identification. Compared to earlier versions of CFM-ID, this new version has an MS/MS spectral prediction performance that is ~22% better and a compound identification accuracy that is ~35% better on a standard (CASMI 2016) testing dataset. CFM-ID 4.0 also features a neutral loss function that allows users to identify similar or substituent compounds where no match can be found using CFM-ID’s regular MS/MS-to-compound identification utility. Finally, the CFM-ID 4.0 web server now offers a much more refined user interface that is easier to use, supports molecular formula identification (from MS/MS data), provides more interactively viewable data (including proposed fragment ion structures) and displays MS mirror plots for comparing predicted with observed MS/MS spectra. These improvements should make CFM-ID 4.0 much more useful to the community and should make small molecule identification much easier, faster, and more accurate.

59 BASIC BIOLOGICAL SCIENCES↗

merlin , an improved framework for the reconstruction of high-quality genome-scale metabolic models

Abstract Genome-scale metabolic models have been recognised as useful tools for better understanding living organisms’ metabolism. merlin (https://www.merlin-sysbio.org/) is an open-source and user-friendly resource that hastens the models’ reconstruction process, conjugating manual and automatic procedures, while leveraging the user's expertise with a curation-oriented graphical interface. An updated and redesigned version of merlin is herein presented. Since 2015, several features have been implemented in merlin, along with deep changes in the software architecture, operational flow, and graphical interface. The current version (4.0) includes the implementation of novel algorithms and third-party tools for genome functional annotation, draft assembly, model refinement, and curation. Such updates increased the user base, resulting in multiple published works, including genome metabolic (re-)annotations and model reconstructions of multiple (lower and higher) eukaryotes and prokaryotes. merlin version 4.0 is the only tool able to perform template based and de novo draft reconstructions, while achieving competitive performance compared to state-of-the art tools both for well and less-studied organisms.

Capela, João (ORCID:0000000212352922)↗

Conserved unique peptide patterns (CUPP) online platform 2.0: implementation of +1000 JGI fungal genomes

Carbohydrate-processing enzymes, CAZymes, are classified into families based on sequence and three-dimensional fold. Because many CAZyme families contain members of diverse molecular function (different EC-numbers), sophisticated tools are required to further delineate these enzymes. Such delineation is provided by the peptide-based clustering method CUPP, Conserved Unique Peptide Patterns. CUPP operates synergistically with the CAZy family/subfamily categorizations to allow systematic exploration of CAZymes by defining small protein groups with shared sequence motifs. The updated CUPP library contains 21,930 of such motif groups including 3,842,628 proteins. The new implementation of the CUPP-webserver, https://cupp.info/, now includes all published fungal and algal genomes from the Joint Genome Institute (JGI), genome resources MycoCosm and PhycoCosm, dynamically subdivided into motif groups of CAZymes. This allows users to browse the JGI portals for specific predicted functions or specific protein families from genome sequences. Thus, a genome can be searched for proteins having specific characteristics. All JGI proteins have a hyperlink to a summary page which links to the predicted gene splicing including which regions have RNA support. The new CUPP implementation also includes an update of the annotation algorithm that uses only a fourth of the RAM while enabling multi-threading, providing an annotation speed below 1 ms/protein.

59 BASIC BIOLOGICAL SCIENCES↗

JGI Plant Gene Atlas: an updateable transcriptome resource to improve functional gene descriptions across the plant kingdom

Abstract Gene functional descriptions offer a crucial line of evidence for candidate genes underlying trait variation. Conversely, plant responses to environmental cues represent important resources to decipher gene function and subsequently provide molecular targets for plant improvement through gene editing. However, biological roles of large proportions of genes across the plant phylogeny are poorly annotated. Here we describe the Joint Genome Institute (JGI) Plant Gene Atlas, an updateable data resource consisting of transcript abundance assays spanning 18 diverse species. To integrate across these diverse genotypes, we analyzed expression profiles, built gene clusters that exhibited tissue/condition specific expression, and tested for transcriptional response to environmental queues. We discovered extensive phylogenetically constrained and condition-specific expression profiles for genes without any previously documented functional annotation. Such conserved expression patterns and tightly co-expressed gene clusters let us assign expression derived additional biological information to 64 495 genes with otherwise unknown functions. The ever-expanding Gene Atlas resource is available at JGI Plant Gene Atlas (https://plantgeneatlas.jgi.doe.gov) and Phytozome (https://phytozome.jgi.doe.gov/), providing bulk access to data and user-specified queries of gene sets. Combined, these web interfaces let users access differentially expressed genes, track orthologs across the Gene Atlas plants, graphically represent co-expressed genes, and visualize gene ontology and pathway enrichments.

59 BASIC BIOLOGICAL SCIENCES↗

The Gene Ontology knowledgebase in 2026

Abstract The Gene Ontology (GO) knowledgebase (https://geneontology.org) is a comprehensive resource describing the functions of genes. The GO knowledgebase is regularly updated and improved. We describe here the major updates that have been made in the past 3 years. The ontology and annotations have been expanded and revised, particularly in several areas of biology: cellular metabolism, multi-organism interactions (e.g. host-pathogen), extracellular matrix proteins, chromatin remodeling (e.g. the “histone code”), and noncoding RNA functions. We have released version 2 of a comprehensive set of integrated, reviewed annotations for human genes, which we call the “functionome.” We have also dramatically increased the number of GO-CAM models, with over 1500 models of metabolic and signaling pathways, primarily in human, mouse, budding and fission yeast, and fruit fly. Finally, we discuss our current recommendations and future prospects of AI in the use and development of GO.

Aleksander, Suzi A (ORCID:0000000167872901)↗

MolViewSpec: a Mol* extension for describing and sharing molecular visualizations

Data visualization is a pivotal component of a structural biologist’s arsenal. The Mol* Viewer makes molecular visualizations available to broader audiences via most web browsers. While Mol* provides a wide range of functionality, it has a steep learning curve and is only available via a JavaScript interface. To enhance the accessibility and usability of web-based molecular visualization, we introduce MolViewSpec (molstar.org/mol-view-spec), a standardized approach for defining molecular visualizations that decouples the definition of complex molecular scenes from their rendering. Scene definition can include references to commonly used structural, volumetric, and annotation data formats together with a description of how the data should be visualized and paired with optional annotations specifying colors, labels, measurements, and custom 3D geometries. Developed as an open standard, this solution paves the way for broader interoperability and support across different programming languages and molecular viewers, enabling more streamlined, standardized, and reproducible visual molecular analyses. MolViewSpec is freely available as a Mol* extension and a standalone Python package.

Midlik, Adam [European Bioinformatics Institute (U↗

CasCollect: targeted assembly of CRISPR-associated operons from high-throughput sequencing data

Abstract CRISPR arrays and CRISPR-associated (Cas) proteins comprise a widespread adaptive immune system in bacteria and archaea. These systems function as a defense against exogenous parasitic mobile genetic elements that include bacteriophages, plasmids and foreign nucleic acids. With the continuous spread of antibiotic resistance, knowledge of pathogen susceptibility to bacteriophage therapy is becoming more critical. Additionally, gene-editing applications would benefit from the discovery of new cas genes with favorable properties. While next-generation sequencing has produced staggering quantities of data, transitioning from raw sequencing reads to the identification of CRISPR/Cas systems has remained challenging. This is especially true for metagenomic data, which has the highest potential for identifying novel cas genes. We report a comprehensive computational pipeline, CasCollect, for the targeted assembly and annotation of cas genes and CRISPR arrays—even isolated arrays—from raw sequencing reads. Benchmarking our targeted assembly pipeline demonstrates significantly improved timing by almost two orders of magnitude compared with conventional assembly and annotation, while retaining the ability to detect CRISPR arrays and cas genes. CasCollect is a highly versatile pipeline and can be used for targeted assembly of any specialty gene set, reconfigurable for user provided Hidden Markov Models and/or reference nucleotide sequences.

Podlevsky, Joshua D.↗

Signature analysis of high-throughput transcriptomics screening data for mechanistic inference and chemical grouping

Abstract High-throughput transcriptomics (HTTr) uses gene expression profiling to characterize the biological activity of chemicals in in vitro cell-based test systems. As an extension of a previous study testing 44 chemicals, HTTr was used to screen an additional 1,751 unique chemicals from the EPA’s ToxCast collection in MCF7 cells using 8 concentrations and an exposure duration of 6 h. We hypothesized that concentration-response modeling of signature scores could be used to identify putative molecular targets and cluster chemicals with similar bioactivity. Clustering and enrichment analyses were conducted based on signature catalog annotations and ToxPrint chemotypes to facilitate molecular target prediction and grouping of chemicals with similar bioactivity profiles. Enrichment analysis based on signature catalog annotation identified known mechanisms of action (MeOAs) associated with well-studied chemicals and generated putative MeOAs for other active chemicals. Chemicals with predicted MeOAs included those targeting estrogen receptor (ER), glucocorticoid receptor (GR), retinoic acid receptor (RAR), the NRF2/KEAP/ARE pathway, AP-1 activation, and others. Using reference chemicals for ER modulation, the study demonstrated that HTTr in MCF7 cells was able to stratify chemicals in terms of agonist potency, distinguish ER agonists from antagonists, and cluster chemicals with similar activities as predicted by the ToxCast ER Pathway model. Uniform manifold approximation and projection (UMAP) embedding of signature-level results identified novel ER modulators with no ToxCast ER Pathway model predictions. Finally, UMAP combined with ToxPrint chemotype enrichment was used to explore the biological activity of structurally related chemicals. The study demonstrates that HTTr can be used to inform chemical risk assessment by determining in vitro points of departure, predicting chemicals’ MeOA and grouping chemicals with similar bioactivity profiles.

Toxicology↗

FUSARIUM-ID v.3.0: An Updated, Downloadable Resource for Fusarium Species Identification

Species within Fusarium are of global agricultural, medical, and food/feed safety concern and have been extensively characterized. However, accurate identification of species is challenging and usually requires DNA sequence data. FUSARIUM-ID ( http://isolate.fusariumdb.org/blast.php ) is a publicly available database designed to support the identification of Fusarium species using sequences of multiple phylogenetically informative loci, especially the highly informative ~680-bp 5' portion of the translation elongation factor 1-alpha (TEF1) gene that has been adopted as the primary barcoding locus in the genus. However, FUSARIUM-ID v.1.0 and 2.0 had several limitations, including inconsistent metadata annotation for the archived sequences and poor representation of some species complexes and marker loci. Here, we present FUSARIUM-ID v.3.0, which provides the following improvements: (i) additional and updated annotation of metadata for isolates associated with each sequence, (ii) expanded taxon representation in the TEF1 sequence database, (iii) availability of the sequence database as a downloadable file to enable local BLAST queries, and (iv) a tutorial file for users to perform local BLAST searches using either freely available software, such as SequenceServer, BLAST+ executable in the command line, and Galaxy, or the proprietary Geneious software. FUSARIUM-ID will be updated on a regular basis by archiving sequences of TEF1 and other loci from newly identified species and greater in-depth sampling of currently recognized species.

Plant Sciences↗

Philympics 2021: Prophage Predictions Perplex Programs

Most bacterial genomes contain integrated bacteriophages—prophages—in various states of decay. Many are active and able to excise from the genome and replicate, while others are cryptic prophages, remnants of their former selves. Over the last two decades, many computational tools have been developed to identify the prophage components of bacterial genomes, and it is a particularly active area for the application of machine learning approaches. However, progress is hindered and comparisons thwarted because there are no manually curated bacterial genomes that can be used to test new prophage prediction algorithms. Here, we present a library of gold-standard bacterial genome annotations that include manually curated prophage annotations, and a computational framework to compare the predictions from different algorithms. We use this suite to compare all extant stand-alone prophage prediction algorithms to identify their strengths and weaknesses. We provide a FAIR dataset for prophage identification, and demonstrate the accuracy, precision, recall, and f 1 score from the analysis of seven different algorithms for the prediction of prophages. We discuss caveats and concerns in this analysis and how those concerns may be mitigated.

Roach, Michael J.↗

DeepAndes: A Self-Supervised Vision Foundation Model for Multispectral Remote Sensing Imagery of the Andes

By mapping sites at large scales usingremotely sensed data, archaeologists can generate unique insights into long-term demographic trends, interregional social networks, and human adaptations in the past. Remote sensing surveys complement field-based approaches, and their reach can be especially great when combined with deep learning and computer vision techniques. However, conventional supervised deep learning methods face challenges in annotating fine-grained archaeological features at scale. In addition, while recent vision foundation models have shown remarkable success in learning large-scale remote sensing data with minimal annotations, most off-the-shelf solutions are designed for RGB images rather than multispectral satellite imagery, such as the eight-band data used in our study. In this article, we introduce DeepAndes, a transformer-based vision foundation model trained on three million multispectral satellite images, specifically tailored for Andean archaeology. DeepAndes incorporates a customized DINOv2 self-supervised learning algorithm optimized for eight-band multispectral imagery, marking the first foundation model designed explicitly for the Andes region. We evaluate its image understanding performance through imbalanced image classification, image instance retrieval, and pixel-level semantic segmentation tasks. Our experiments show that DeepAndes achieves superior F1 scores, mean average precision, and Dice scores in few-shot learning scenarios, significantly outperforming models trained from scratch or pretrained on smaller datasets. This underscores the effectiveness of large-scale self-supervised pretraining in archaeological remote sensing.

Guo, Junlin [Vanderbilt Univ., Nashville, TN (Unit↗

ASDFL: An adaptive super‐pixel discriminative feature‐selective learning for vehicle matching

Abstract There are a large number of cameras in modern transportation system that capture numerous vehicle images continuously. Therefore, automatic analysis of these vehicle images is helpful for traffic flow management, criminal investigations and vehicle inspections. Vehicle matching, which aims to determine whether two input images depict an identical vehicle, is one of the core tasks in vehicle analysis. Recent relevant studies have focused on local feature extraction instead of global extraction, since local details can provide crucial cues to distinguish between cars. However, these methods do not select local features; that is, they do not assign weights to local features. Therefore, in this research, we systematically study the vehicle matching task, and present a novel annotation‐free local‐based deep learning method called Adaptive super‐pixel discriminative feature‐selective learning (ASDFL) to address this issue. In ASDFL, vehicle images are segmented into clusters of super‐pixels of similar size by considering the location and colour similarities of pixels without using any component‐level annotation. These super‐pixels are deemed to be the virtual components of vehicles. Moreover, a convolutional neural network is used to extract the deep features of these virtual components. Thereafter, an instance‐specific mask generation module driven by the extracted global features is enhanced to produce a mask to select the most distinctive virtual components of each vehicle image pair in the feature space. Finally, the vehicle matching task is accomplished by classifying the selected virtual component features of each imaged vehicle pair. Extensive experiments on two popular vehicle identification benchmarks demonstrate that our method is 1.57% and 0.8% more accurate than the previous baselines in a vehicle matching task on the VeRi and VehicleID datasets, respectively, which demonstrates the effectiveness of our method.

Qin, Rong↗

A haplotype‐resolved reference genome of Quercus alba sheds light on the evolutionary history of oaks

Summary White oak ( Quercus alba ) is an abundant forest tree species across eastern North America that is ecologically, culturally, and economically important. We report the first haplotype‐resolved chromosome‐scale genome assembly of Q. alba and conduct comparative analyses of genome structure and gene content against other published Fagaceae genomes. We investigate the genetic diversity of this widespread species and the phylogenetic relationships among oaks using whole genome data. Despite strongly conserved chromosome synteny and genome size across Quercus , certain gene families have undergone rapid changes in size, including defense genes. Unbiased annotation of resistance (R) genes across oaks revealed that the overall number of R genes is similar across species – as are the chromosomal locations of R gene clusters – but, gene number within clusters is more labile. We found that Q. alba has high genetic diversity, much of which predates its divergence from other oaks and likely impacts divergence time estimations. Our phylogenetic results highlight widespread phylogenetic discordance across the genus. The white oak genome represents a major new resource for studying genome diversity and evolution in Quercus . Additionally, we show that unbiased gene annotation is key to accurately assessing R gene evolution in Quercus .

Larson, Drew A. [Department of Biology Indiana Uni↗

Morphogene-assisted transformation of Sorghum bicolor allows more efficient genome editing

Sorghum bicolor (L.) Moench, the fifth most important cereal worldwide, is a multi-use crop for feed, food, forage and fuel. To enhance the sorghum and other important crop plants, establishing gene function is essential for their improvement. For sorghum, identifying genes associated with its notable abiotic stress tolerances requires a detailed molecular understanding of the genes associated with those traits. The limits of this knowledge became evident from our earlier in-depth sorghum transcriptome study showing that over 40% of its transcriptome had not been annotated. Here, we describe a full spectrum of tools to engineer, edit, annotate and characterize sorghum’s genes. Efforts to develop those tools began with a morphogene-assisted transformation (MAT) method that led to accelerated transformation times, nearly half the time required with classical callus-based, non-MAT approaches. These efforts also led to expanded numbers of amenable genotypes, including several not previously transformed or historically recalcitrant. Another transformation advance, termed altruistic, involved introducing a gene of interest in a separate Agrobacterium strain from the one with morphogenes, leading to plants with the gene of interest but without morphogenes. The MAT approach was also successfully used to edit a target exemplary gene, phytoene desaturase. To identify single-copy transformed plants, we adapted a high-throughput technique and also developed a novel method to determine transgene independent integration. These efforts led to an efficient method to determine gene function, expediting research in numerous genotypes of this widely grown, multi-use crop.

54 ENVIRONMENTAL SCIENCES↗