Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Assessing the Impact of Measurement Precision on Metabolite Identification Probability in Multidimensional Mass Spectrometry-Based, Reference-Free Metabolomics

Identification of compounds with minimal ambiguity remains a central challenge in mass spectrometry-based metabolomics. Conventional compound identification relies on comparing analytical signatures (e.g., mass-to-charge ratio, collision cross section, tandem mass spectra) against reference data obtained from measurements of authentic chemical standards. The breadth of annotatable compounds using this approach is necessarily limited by availability of authentic standards, analytical throughput, and resolving power of the separations that underly the measurements. The maturation of computational methods, both theory-driven and artificial intelligence/machine learning-based, for prediction of various molecular properties relevant to multidimensional mass spectrometry measurements has opened the door to a new “reference-free” paradigm of compound annotation. Through augmenting existing reference data for molecular properties with computational predictions, the universe of identifiable chemical species can be expanded significantly beyond its current limits. An unexplored aspect of this novel approach is understanding how to gauge confidence in resulting annotations, especially as the compound search space is expanded. Intuitively, the confidence of a compound annotation is related to the inherent discriminatory power of the molecular properties used for identification, as well as the precision with which the properties are measured or predicted. In this work, we characterize this relationship between measurement precision and identification probability in a systematic and quantitative fashion for a defined region of chemical space that includes organic small molecule metabolites. Importantly, this work establishes a framework for conducting metabolite identification probability analysis that enables others to quantify this relationship for their own compounds and properties of interest.

Metabolite Identification↗

ULTRA-effective labeling of tandem repeats in genomic sequence

In the age of long read sequencing, genomics researchers now have access to accurate repetitive DNA sequence (including satellites) that, due to the limitations of short read-sequencing, could previously be observed only as unmappable fragments. Tools that annotate repetitive sequence are now more important than ever, so that we can better understand newly uncovered repetitive sequences, and also so that we can mitigate errors in bioinformatic software caused by those repetitive sequences. To that end, we introduce the 1.0 release of our tool for identifying and annotating locally repetitive sequence, ULTRA Locates Tandemly Repetitive Areas (ULTRA). ULTRA is fast enough to use as part of an efficient annotation pipeline, produces state-of-the-art reliable coverage of repetitive regions containing many mutations, and provides interpretable statistics and labels for repetitive regions.

59 BASIC BIOLOGICAL SCIENCES↗

MS2Planner: improved fragmentation spectra coverage in untargeted mass spectrometry by iterative optimized data acquisition

Motivation: Untargeted mass spectrometry experiments enable the profiling of metabolites in complex biological samples. The collected fragmentation spectra are the metabolite’s fingerprints that are used for molecule identification and discovery. Two main mass spectrometry strategies exist for the collection of fragmentation spectra: data-dependent acquisition (DDA) and data-independent acquisition (DIA). In the DIA strategy, all the metabolites ions in predefined mass-to-charge ratio ranges are co-isolated and co-fragmented, resulting in multiplexed fragmentation spectra that are challenging to annotate. In contrast, in the DDA strategy, fragmentation spectra are dynamically and specifically collected for the most abundant ions observed, causing redundancy and sub-optimal fragmentation spectra collection. Yet, DDA results in less multiplexed fragmentation spectra that can be readily annotated. Results: We introduce the MS2Planner workflow, an Iterative Optimized Data Acquisition strategy that optimizes the number of high-quality fragmentation spectra over multiple experimental acquisitions using topological sorting. Our results showed that MS2Planner increases the annotation rate by 38.6% and is 62.5% more sensitive and 9.4% more specific compared to DDA. Availability and implementation MS2Planner code is available at https://github.com/mohimanilab/MS2Planner. The generation of the inclusion list from MS2Planner was performed with python scripts available at https://github.com/lfnothias/IODA_MS.

47 OTHER INSTRUMENTATION↗

An expectation–maximization framework for comprehensive prediction of isoform-specific functions

Advances in RNA sequencing technologies have achieved an unprecedented accuracy in the quantification of mRNA isoforms, but our knowledge of isoform-specific functions has lagged behind. There is a need to understand the functional consequences of differential splicing, which could be supported by the generation of accurate and comprehensive isoform-specific gene ontology annotations. We present isoform interpretation, a method that uses expectation–maximization to infer isoform-specific functions based on the relationship between sequence and functional isoform similarity. We predicted isoform-specific functional annotations for 85 617 isoforms of 17 900 protein-coding human genes spanning a range of 17 430 distinct gene ontology terms. Comparison with a gold-standard corpus of manually annotated human isoform functions showed that isoform interpretation significantly outperforms state-of-the-art competing methods. We provide experimental evidence that functionally related isoforms predicted by isoform interpretation show a higher degree of domain sharing and expression correlation than functionally related genes. We also show that isoform sequence similarity correlates better with inferred isoform function than with gene-level function.

59 BASIC BIOLOGICAL SCIENCES↗

Identification of candidate host-specificity genes in Exserohilum turcicum using comparative genomics and transcriptomics

Abstract Exserohilum turcicum causes northern corn leaf blight and sorghum leaf blight. While the same species cause disease in both crops, the strains are host-specific. Here, we report the sequence and de novo annotated assemblies of one sorghum- and one maize-specific E. turcicum strain. The strains were sequenced using the PacBio Sequel II system. The total genome length for both assemblies was between 44 and 45 Mb with N50 of ∼2.5 Mb. Ninety-eight percent of the Benchmarking Universal Single-Copy Orthologs (BUSCO) for both assemblies had complete status. The estimated number of genes was 11,762 and 12,029 in the sorghum- and maize-specific isolates, respectively. Funannotate, EffectorP, SignalP, and transcriptome data were used to create functional annotation of each genome. The whole-genome comparison identified ten large-scale inversions and three translocations between the maize- and sorghum-specific strains, along with homologous genes and gene duplications. RNA was sequenced from the maize- and sorghum-specific isolate 10 days post-inoculation in maize and sorghum and from axenic cultures. Gene expression data from planta and axenic growth experiments were compared for each strain. Candidate host-specificity genes were identified by combining results from whole-genome comparison, synteny analysis, gene annotations, and transcriptome data. Overall, this study identified several candidate host-specificity genes that provide insights into E. turcicum interaction with its hosts.

Krone, Mara J. (ORCID:0000000159006624)↗

Accelerating Biological Insight for Understudied Genes

Synopsis The rapid expansion of genome sequence data is increasing the discovery of protein-coding genes across all domains of life. Annotating these genes with reliable functional information is necessary to understand evolution, to define the full biochemical space accessed by nature, and to identify target genes for biotechnology improvements. The majority of proteins are annotated based on sequence conservation with no specific biological, biochemical, genetic, or cellular function identified. Recent technical advances throughout the biological sciences enable experimental research on these understudied protein-coding genes in a broader collection of species. However, scientists have incentives and biases to continue focusing on well documented genes within their preferred model organism. This perspective suggests a research model that seeks to break historic silos of research bias by enabling interdisciplinary teams to accelerate biological functional annotation. We propose an initiative to develop coordinated projects of collaborating evolutionary biologists, cell biologists, geneticists, and biochemists that will focus on subsets of target genes in multiple model organisms. Concurrent analysis in multiple organisms takes advantage of evolutionary divergence and selection, which causes individual species to be better suited as experimental models for specific genes. Most importantly, multisystem approaches would encourage transdisciplinary critical thinking and hypothesis testing that is inherently slow in current biological research.

Zoology↗

Co-locality to co-functionality: Eukaryotic gene neighborhoods as a resource for function

Diverging from the classic paradigm of random gene order in eukaryotes, gene proximity can be leveraged to systematically identify functionally related gene neighborhoods in eukaryotes, utilizing techniques pioneered in bacteria. Current methods of identifying gene neighborhoods typically rely on sequence similarity to characterized gene products. However, this approach is not robust for non model organisms like algae, which are evolutionarily distant from well-characterized model organisms. Here, we utilize a comparative genomic approach to identify evolutionarily conserved Proximal Orthologous Gene (POG) pairs conserved across at least two taxonomic classes of green algae. A total of 317 gene neighborhoods were identified. In some cases, gene proximity appears to have been conserved since before the streptophyte-chlorophyte split, 1,000 million years ago. Using functional inferences derived from reconstructed evolutionary relationships, we identified several novel functional clusters. A putative mycosporine-like amino acid (MAA), "sunscreen", neighborhood contains genes similar to either vertebrate or cyanobacterial pathways, suggesting a novel mosaic biosynthetic pathway in green algae. One of two putative arsenic-detoxification neighborhoods includes an organoarsenical transporter (ArsJ), a glyceraldehyde 3-phosphate dehydrogenase-like gene, homologs of which are involved in arsenic detoxification in bacteria, and a novel algal-specific phosphoglycerate kinase-like gene (PGK). Mutants of the ArsJ-like transporter and PGK-like genes in Chlamydomonas reinhardtii were found to be sensitive to arsenate, providing experimental support for the role of these identified neighbors in resistance to arsenate. Potential evolutionary origins of neighborhoods are discussed, and updated annotations for formerly poorly annotated genes are presented, highlighting the potential of this strategy for functional annotation.

59 BASIC BIOLOGICAL SCIENCES↗

ECO: the Evidence and Conclusion Ontology, an update for 2022

The Evidence and Conclusion Ontology (ECO) is a community resource that provides an ontology of terms used to capture the type of evidence that supports biomedical annotations and assertions. Consistent capture of evidence information with ECO allows tracking of annotation provenance, establishment of quality control measures, and evidence-based data mining. ECO is in use by dozens of data repositories and resources with both specific and general areas of focus. ECO is continually being expanded and enhanced in response to user requests as well as our aim to adhere to community best-practices for ontology development. The ECO support team engages in multiple collaborations with other ontologies and annotating groups. Here we report on recent updates to the ECO ontology itself as well as associated resources that are available through this project. ECO project products are freely available for download from the project website (https://evidenceontology.org/) and GitHub (https://github.com/evidenceontology/evidenceontology). ECO is released into the public domain under a CC0 1.0 Universal license.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative mitogenomics of kingdom Fungi – evolutionary insights and metagenomic applications

Mitochondria are essential components of eukaryotic cells, responsible for ATP production through oxidative phosphorylation. Despite their biological importance, unique challenges have hindered the adoption of automated mitochondrial genome (mitogenome) annotation methods, obstructing mitochondrial comparative genomics in a broad evolutionary context. Using Fungi as a study system and a Joint Genome Institute (JGI) annotated high-quality reference set, we observed broad patterns of mitochondrial evolution across the kingdom. We found that the median fungal mitogenome size is 58 kb and identified exceptionally large examples over 1 Mb in Pezizomycetes. All 14 expected oxidative phosphorylation protein-coding genes, plus rps3, were generally conserved. We found evidence of major evolutionary transitions within the Ascomycota, including the transfer of mitochondrially encoded atp8 and atp9 to the nuclear genomes across the Pezizomycotina and shifts in mitogenome tRNA patterns across the kingdom. We found substantial concordance between mitochondrial and nuclear evolution, enabling us to document 3131 total fungal mitogenomes from JGI-derived metagenomic datasets. We also identified 6467 total undeclared mitogenomes embedded in Genbank fungal nuclear assemblies. We provide interactive tools for mitogenome analysis through the JGI MycoCosm platform. Collectively, this work generated nearly 10 000 new fungal mitogenome annotations, providing a foundation and resources for future exploration of comparative fungal mitogenomics.

Ahrendt, Steven R. [USDOE Joint Genome Institute (↗

The Chlamydomonas Genome Project, version 6: reference assemblies for mating type plus and minus strains reveal extensive structural mutation in the laboratory

Five versions of the Chlamydomonas reinhardtii reference genome have been produced over the last two decades. Here we present version 6, bringing significant advances in assembly quality and structural annotations. PacBio-based chromosome-level assemblies for two laboratory strains, CC-503 and CC-4532, provide resources for the plus and minus mating type alleles. We corrected major misassemblies in previous versions and validated our assemblies via linkage analyses. Contiguity increased over ten-fold and >80% of filled gaps are within genes. We used Iso-Seq and deep RNA-seq datasets to improve structural annotations, and updated gene symbols and textual annotation of functionally characterized genes via extensive manual curation. We discovered that the cell wall-less classical reference strain CC-503 exhibits genomic instability potentially caused by deletion of the helicase RECQ3, with major structural mutations identified that affect >100 genes. We therefore present the CC-4532 assembly as the primary reference, although this strain also carries unique structural mutations and is experiencing rapid proliferation of a Gypsy retrotransposon. We expect all laboratory strains to harbor gene-disrupting mutations, which should be considered when interpreting and comparing experimental results. Collectively, the resources presented here herald a new era of Chlamydomonas genomics and will provide the foundation for continued research in this important reference organism.

59 BASIC BIOLOGICAL SCIENCES↗

Global Explainability of A Deep Abstaining Classifier for Cancer Pathology Reports

We present a global explainability method to characterize sources of errors in a real-world multitask deep abstaining classifier (DAC), in the context of cancer histology prediction. Our multitask classifier, currently deployed for automated annotation of cancer pathology reports from NCI-SEER registries, was trained and evaluated on 1.04 million hand-annotated samples and makes simultaneous predictions of cancer site, subsite, histology, laterality, and behavior for each report. The DAC framework enables the model to abstain on ambiguous reports and confusing classes to achieve the target accuracy on the retained (non-abstained) samples, but at the cost of decreased coverage. Requiring 97% accuracy on the histology task caused our model to retain only 22% of all samples, mostly the less ambiguous and common classes. Local explainability with the GradInp technique provided a computationally efficient way of obtaining contextual reasoning for hundreds of thousands of individual predictions. Our method, involving dimensionality reduction of approximately 13000 aggregated local explanations (ALE), offers a tractable path to true global explainability. It enabled identification of sources of errors in histology classification, globally, as hierarchical complexity among classes, label noise, insufficient information, and conflicting evidence. This suggests several strategies for iterative improvement of our DAC, including well-designed exclusion criteria, focused annotation, and reduced penalties for errors involving hierarchically related classes.

59 BASIC BIOLOGICAL SCIENCES↗

Combining GWAS and population genomic analyses to characterize coevolution in a legume‐rhizobia symbiosis

Abstract The mutualism between legumes and rhizobia is clearly the product of past coevolution. However, the nature of ongoing evolution between these partners is less clear. To characterize the nature of recent coevolution between legumes and rhizobia, we used population genomic analysis to characterize selection on functionally annotated symbiosis genes as well as on symbiosis gene candidates identified through a two‐species association analysis. For the association analysis, we inoculated each of 202 accessions of the legume host Medicago truncatula with a community of 88 Sinorhizobia (Ensifer) meliloti strains. Multistrain inoculation, which better reflects the ecological reality of rhizobial selection in nature than single‐strain inoculation, allows strains to compete for nodulation opportunities and host resources and for hosts to preferentially form nodules and provide resources to some strains. We found extensive host by symbiont, that is, genotype‐by‐genotype, effects on rhizobial fitness and some annotated rhizobial genes bear signatures of recent positive selection. However, neither genes responsible for this variation nor annotated host symbiosis genes are enriched for signatures of either positive or balancing selection. This result suggests that stabilizing selection dominates selection acting on symbiotic traits and that variation in these traits is under mutation‐selection balance. Consistent with the lack of positive selection acting on host genes, we found that among‐host variation in growth was similar whether plants were grown with rhizobia or N‐fertilizer, suggesting that the symbiosis may not be a major driver of variation in plant growth in multistrain contexts.

59 BASIC BIOLOGICAL SCIENCES↗

Information theory and machine learning illuminate large‐scale metabolomic responses of Brachypodium distachyon to environmental change

SUMMARY Plant responses to environmental change are mediated via changes in cellular metabolomes. However, <5% of signals obtained from liquid chromatography tandem mass spectrometry (LC‐MS/MS) can be identified, limiting our understanding of how metabolomes change under biotic/abiotic stress. To address this challenge, we performed untargeted LC‐MS/MS of leaves, roots, and other organs of Brachypodium distachyon (Poaceae) under 17 organ–condition combinations, including copper deficiency, heat stress, low phosphate, and arbuscular mycorrhizal symbiosis. We found that both leaf and root metabolomes were significantly affected by the growth medium. Leaf metabolomes were more diverse than root metabolomes, but the latter were more specialized and more responsive to environmental change. We found that 1 week of copper deficiency shielded the root, but not the leaf metabolome, from perturbation due to heat stress. Machine learning (ML)‐based analysis annotated approximately 81% of the fragmented peaks versus approximately 6% using spectral matches alone. We performed one of the most extensive validations of ML‐based peak annotations in plants using thousands of authentic standards, and analyzed approximately 37% of the annotated peaks based on these assessments. Analyzing responsiveness of each predicted metabolite class to environmental change revealed significant perturbations of glycerophospholipids, sphingolipids, and flavonoids. Co‐accumulation analysis further identified condition‐specific biomarkers. To make these results accessible, we developed a visualization platform on the Bio‐Analytic Resource for Plant Biology website ( https://bar.utoronto.ca/efp_brachypodium_metabolites/cgi‐bin/efpWeb.cgi ), where perturbed metabolite classes can be readily visualized. Overall, our study illustrates how emerging chemoinformatic methods can be applied to reveal novel insights into the dynamic plant metabolome and stress adaptation.

59 BASIC BIOLOGICAL SCIENCES↗

ShadeLab/PAPER_Howe_2023_switchgrass_MetaT

The raw data (metagenomes and metatranscriptomes) for this study are available in the Joint Genomes Institute Genome Portal (https://genome.jgi.doe.gov/portal/ Project ID 503249) with projects designated by year and product type. The MAG genomes analyzed in this paper are available on NCBI, as bioproject PRJNA800073. Plants and microorganisms form beneficial associations. Understanding plant-microbe interactions will inform microbiome management to enhance crop productivity and resilience to stress. Here, we apply a genome-centric approach to identify ecologically important leaf microbiome members on field-grown switchgrass and miscanthus and to quantify their activities for switchgrass over two growing seasons. We integrate metagenome and metatranscriptome sequencing from 192 leaf samples collected over representative time points in crop phenology. We curated 40 medium- and high-quality metagenome-assembled-genomes (MAGs) and focused analysis on seasonal transcript recruitment to them. Classes represented by these focal MAGs (Actinomycetia, Alpha- and Gamma- Proteobacteria, and Bacteroidota) were active and had increases in transcripts for short-chain dehydrogenase, molybdopterin oxidoreductase, and polyketide cyclase in the late season. The majority of MAGs had activated stress-associated pathways, including trehalose metabolism, indole acetic acid degradation, betaine biosynthesis, and reactive oxygen species degradation, suggesting direct engagement with the host environment. We also detected seasonally activated biosynthetic pathways for terpenes (carotenoids and isoprenoids) and for various non-ribosomal peptide pathways that were poorly annotated. Overall, this study overcame laboratory and bioinformatic challenges associated with field-based leaf metatranscriptome analysis to inform both general and likely specialized activities of these phyllosphere populations. These activities collectively support that leaf-associated bacterial populations are seasonally dynamic, responsive to host cues, and interactively engage in feedback with the plant. This analysis represented quality filtering of metagenomes and metatranscriptomes (data-preparation folder), metagenome assemblies (metagenome-assembly folder) and metagenome-assembled genome binning, curation, refinement, annotation (mag-evaluation folder). Abundances of sequencing libraries were calculated based on reads mapped (mapping folder). Additionallly, analysis of our annotated results are also included (analysis folder).

Howe, Adina↗

Filling gaps in bacterial catabolic pathways with computation and high-throughput genetics

To discover novel catabolic enzymes and transporters, we combined high-throughput genetic data from 29 bacteria with an automated tool to find gaps in their catabolic pathways. GapMind for carbon sources automatically annotates the uptake and catabolism of 62 compounds in bacterial and archaeal genomes. For the compounds that are utilized by the 29 bacteria, we systematically examined the gaps in GapMind’s predicted pathways, and we used the mutant fitness data to find additional genes that were involved in their utilization. We identified novel pathways or enzymes for the utilization of glucosamine, citrulline, myo-inositol, lactose, and phenylacetate, and we annotated 299 diverged enzymes and transporters. We also curated 125 proteins from published reports. For the 29 bacteria with genetic data, GapMind finds high-confidence paths for 85% of utilized carbon sources. In diverse bacteria and archaea, 38% of utilized carbon sources have high-confidence paths, which was improved from 27% by incorporating the fitness-based annotations and our curation. GapMind for carbon sources is available as a web server ( http://papers.genomics.lbl.gov/carbon ) and takes just 30 seconds for the typical genome.

59 BASIC BIOLOGICAL SCIENCES↗

Individual Wave Detection and Tracking within a Rotating Detonation Engine through Computer Vision Object Detection applied to High-Speed Images

Known for their simplistic design and continuous detonation, rotating detonation engines (RDEs) constitute a majority of current pressure gain combustion (PGC) research efforts. Experimental RDE operation times have been continuously extended through the use of rig cooling techniques. As the window of observable behavior is expanded, and as the technology matures toward eventual integration within gas turbines, monitoring techniques must evolve to better match industrial diagnostics. High-speed image analysis techniques prove useful to capture and evaluate the unsteady detonation behavior within the RDE. Traditional image analysis techniques, however, require extensive processing times which prohibit simultaneous monitoring. To better address this problem, a computer vision object detection methodology is proposed to quickly detect individual detonation waves within a single down-axis image. Detonation waves are detected in individual images by the implemented computer vision method You Only Look Once (YOLO) object detection network. In order to detect detonation waves, the network must first be trained using RDE images of interest, for which each required phase of network development is outlined. Detection of waves is improved through proper treatment of the collected image set, variation of Intersection over Union (IoU) and confidence thresholding, and through a parametric study of annotation dimensions. Each detected wave is described by its location and rotational direction, and locations are tracked to calculate wave velocity across each frame, leading to a timestep resolution of 20 µs. Wave velocities are also calculated through a series of frames, leading to a suitable average velocity estimation using as few as 10 frames. Uncertainty analysis accounting for variation in camera framerate, pixel width and annotation centroid locations estimates a total uncertainty of ±4.3% for velocity calculations, using the smallest annotation boxes. This method offers great reductions in processing times, as a step toward real-time monitoring of detonation waves within an RDE. Improving on previous studies, this technique is impartial to wave modes not included in the original training set and calculates wave velocities independent of high-speed pressure data. The ability to isolate waves within predicted bounding boxes will likely facilitate analysis of pixel intensity variation as an estimation of wave strength in future work.

Johnson, Kristyn↗

Dataset for the Danczak et al., 2025 manuscript about bacterial-fungal interactions

We generated genome-resolved multiomics data from a series of metagenomic and metatranscriptomic sequencing. Specifically, we acquired, functionally annotated, and taxonomically classified both bacterial and eukaryotic metagenome assembled genomes (MAGs). For bacterial MAGs, we assembled eukaryotic float metagenomic sequencing data from JGI using MEGAHIT, binned and refined MAGs using MetaWRAP and dRep, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using GTDB-tk. For eukaryotic MAGs, we first identified potentially eukaryotic contigs from a coassembly of eukaryotic float metagenomic sequencing data from JGI using EukRep and Whokaryote, binned MAGs using MetaBAT2, functionally annotated MAGs using eggNOG mapper, and assigned taxonomy using Eukulele. Bulk metatranscriptomic reads were mapped to bacterial MAGs and polyA-metatranscriptomic read were mapped to eukaryotic MAGs using bbmap.

Danczak, Robert E. [Pacific Northwest National Lab↗

Hyaloscypha finlandica Metabolome Repository

This repository provides the curated data tables, manuscript figure and table exports, dependency records, and workflow scripts supporting an integrated comparative genomics and untargeted LC-MS/MS metabolomics analysis of Hyaloscypha finlandica strain PMI 746, a root-associated dark septate endophyte of poplar. The repository includes genome-mining summaries from antiSMASH, FunBGCeX, BGC-Prophet, and BiG-SCAPE; processed metabolomics inputs; metabolite annotation evidence; statistical outputs; and publication-facing figures and tables. Raw LC-MS/MS spectra, full genome/protein downloads, and large generated tool outputs are referenced through public archive/accession records and are not stored in Git.

59 BASIC BIOLOGICAL SCIENCES↗