Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

A Deep Exon Cryptic Splice Site Promotes Aberrant Intron Retention in a Von Willebrand Disease Patient

A translationally silent single nucleotide mutation in exon 44 (E44) of the von Willebrand factor (VWF) gene is associated with inefficient removal of intron 44 in a von Willebrand disease (VWD) patient. This intron retention (IR) event was previously attributed to reordered E44 secondary structure that sequesters the normal splice donor site. We propose an alternative mechanism: the mutation introduces a cryptic splice donor site that interferes with the function of the annotated site to favor IR. We evaluated both models using minigene splicing reporters engineered to vary in secondary structure and/or cryptic splice site content. Analysis of splicing efficiency in transfected K562 cells suggested that the mutation-generated cryptic splice site in E44 was sufficient to induce substantial IR. Mutations predicted to vary secondary structure at the annotated site also had modest effects on IR and shifted the balance of residual splicing between the cryptic site and annotated site, supporting competition among the sites. Further studies demonstrated that introduction of cryptic splice donor motifs at other positions in E44 did not promote IR, indicating that interference with the annotated site is context dependent. We conclude that mutant deep exon splice sites can interfere with proper splicing by inducing IR.

60 APPLIED LIFE SCIENCES↗

The state of algal genome quality and diversity

The genomic era of biology has created unprecedented opportunity to study how life works, including understanding evolutionary principles, bioprospecting for novel antibiotics, and genetic manipulation of bioeconomically-relevant species. While genome sequencing and genomic analysis was previously restricted to large-scale projects and consortia, sequencing has become democratized. However, as genomics has become commonplace, the cataloging of sequenced organisms has become challenging, and standardized practices for sequencing, assembly, and annotation have not been adopted. This is equally true in research fields such as algal biology, despite the growing importance of algae in a bio-based economy. Here, in this study, we provide a comprehensive review of the state of eukaryotic algal genomics, explore the quality of algal genome assemblies, and identify the biases and gaps in the current species distribution to inform the development of future genome projects. Overall, we find a trend of declining quality of genomic resources, including a reduction in assembly quality, gene annotation quality, and genome completeness. Potential solutions to improve genome quality include widespread utilization of long read and scaffolding technologies, implementation of standards for assembly quality, evidence-based gene annotation and requisite publication, and support of continued development of gene prediction and genome assessment software.

59 BASIC BIOLOGICAL SCIENCES↗

The Ontologies Community of Practice: A CGIAR Initiative for Big Data in Agrifood Systems

Heterogeneous and multidisciplinary data generated by research on sustainable global agriculture and agrifood systems requires quality data labeling or annotation in order to be interoperable. As recommended by the FAIR principles, data, labels, and metadata must use controlled vocabularies and ontologies that are popular in the knowledge domain and commonly used by the community. Despite the existence of robust ontologies in the Life Sciences, there is currently no comprehensive full set of ontologies recommended for data annotation across agricultural research disciplines. In this paper, we discuss the added value of the Ontologies Community of Practice (CoP) of the CGIAR Platform for Big Data in Agriculture for harnessing relevant expertise in ontology development and identifying innovative solutions that support quality data annotation. The Ontologies CoP stimulates knowledge sharing among stakeholders, such as researchers, data managers, domain experts, experts in ontology design, and platform development teams.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

A Step-by-Step Protocol from METASPACE to Biological Interpretation

Mass spectrometry imaging (MSI) represents an exceptional tool for exploring complex biological systems spatially at the molecular level. However, due to its multidimensional nature and large-scale data output, it presents considerable challenges when it comes to extracting meaningful biological insights. Recent advancements, such as the METASPACE platform, have enabled researchers to efficiently process, annotate, and interpret MSI datasets by leveraging machine learning and cloud-based infrastructure. In this tutorial, we present a detailed and user-friendly R-pipeline designed to help METASPACE users navigate untargeted metabolomic annotations and transform them into practical insights about their biological systems. By combining METASPACE annotations with rapid R-based screening, this workflow not only streamlined the analytical process but also enhanced the understanding of spatial molecular distribution, especially for complex systems. Here, this easy-to-follow approach has the potential for applications in diagnostics, drug discovery, environmental and ecological processes, and more. We envision this pipeline to be particularly useful for newcomers to the field of MSI and

Moreno Pedraza, Abigail↗

Text Mining for Process–Structure–Properties Relationships in Metals

With the advent of large language models (LLMs), the vast unstructured text within millions of academic papers is increasingly accessible for materials discovery—although significant challenges remain. While LLMs offer promising few- and zero-shot learning capabilities, particularly valuable in the materials domain where expert annotations are scarce, general-purpose LLMs often fail to address key materials-specific queries without further adaptation. To bridge this gap, fine-tuning LLMs on human-labeled data is essential for effective structured knowledge extraction (Liu in The Importance of Human-Labeled Data in the Era of LLMs, 2023). Here, in this study, we introduce a novel annotation schema designed to extract generic process–structure–properties relationships from scientific literature. We demonstrate the utility of this approach using a dataset of 128 abstracts, with annotations drawn from two distinct domains: high-temperature materials (Domain I) and uncertainty quantification in simulating materials microstructure (Domain II). Initially, we developed a conditional random field (CRF) model based on MatBERT—a domain-specific BERT variant—and evaluated its performance on Domain I. Subsequently, we compared this model with a fine-tuned LLM (GPT-4o from OpenAI) under identical conditions. Our results indicate that fine-tuning LLMs can significantly improve entity extraction performance over the BERT-CRF baseline on Domain I. However, when additional examples from Domain II were incorporated, the performance of the BERT-CRF model became comparable to that of the GPT-4o model. These findings underscore the potential of our schema for structured knowledge extraction and highlight the complementary strengths of both modeling approaches.

Materials science↗

Assessing the Impact of Measurement Precision on Metabolite Identification Probability in Multidimensional Mass Spectrometry-Based, Reference-Free Metabolomics

Identification of compounds with minimal ambiguity remains a central challenge in mass spectrometry-based metabolomics. Conventional compound identification relies on comparing analytical signatures (e.g., mass-to-charge ratio, collision cross section, tandem mass spectra) against reference data obtained from measurements of authentic chemical standards. The breadth of annotatable compounds using this approach is necessarily limited by availability of authentic standards, analytical throughput, and resolving power of the separations that underly the measurements. The maturation of computational methods, both theory-driven and artificial intelligence/machine learning-based, for prediction of various molecular properties relevant to multidimensional mass spectrometry measurements has opened the door to a new “reference-free” paradigm of compound annotation. Through augmenting existing reference data for molecular properties with computational predictions, the universe of identifiable chemical species can be expanded significantly beyond its current limits. An unexplored aspect of this novel approach is understanding how to gauge confidence in resulting annotations, especially as the compound search space is expanded. Intuitively, the confidence of a compound annotation is related to the inherent discriminatory power of the molecular properties used for identification, as well as the precision with which the properties are measured or predicted. In this work, we characterize this relationship between measurement precision and identification probability in a systematic and quantitative fashion for a defined region of chemical space that includes organic small molecule metabolites. Importantly, this work establishes a framework for conducting metabolite identification probability analysis that enables others to quantify this relationship for their own compounds and properties of interest.

Metabolite Identification↗

ULTRA-effective labeling of tandem repeats in genomic sequence

In the age of long read sequencing, genomics researchers now have access to accurate repetitive DNA sequence (including satellites) that, due to the limitations of short read-sequencing, could previously be observed only as unmappable fragments. Tools that annotate repetitive sequence are now more important than ever, so that we can better understand newly uncovered repetitive sequences, and also so that we can mitigate errors in bioinformatic software caused by those repetitive sequences. To that end, we introduce the 1.0 release of our tool for identifying and annotating locally repetitive sequence, ULTRA Locates Tandemly Repetitive Areas (ULTRA). ULTRA is fast enough to use as part of an efficient annotation pipeline, produces state-of-the-art reliable coverage of repetitive regions containing many mutations, and provides interpretable statistics and labels for repetitive regions.

59 BASIC BIOLOGICAL SCIENCES↗

MS2Planner: improved fragmentation spectra coverage in untargeted mass spectrometry by iterative optimized data acquisition

Motivation: Untargeted mass spectrometry experiments enable the profiling of metabolites in complex biological samples. The collected fragmentation spectra are the metabolite’s fingerprints that are used for molecule identification and discovery. Two main mass spectrometry strategies exist for the collection of fragmentation spectra: data-dependent acquisition (DDA) and data-independent acquisition (DIA). In the DIA strategy, all the metabolites ions in predefined mass-to-charge ratio ranges are co-isolated and co-fragmented, resulting in multiplexed fragmentation spectra that are challenging to annotate. In contrast, in the DDA strategy, fragmentation spectra are dynamically and specifically collected for the most abundant ions observed, causing redundancy and sub-optimal fragmentation spectra collection. Yet, DDA results in less multiplexed fragmentation spectra that can be readily annotated. Results: We introduce the MS2Planner workflow, an Iterative Optimized Data Acquisition strategy that optimizes the number of high-quality fragmentation spectra over multiple experimental acquisitions using topological sorting. Our results showed that MS2Planner increases the annotation rate by 38.6% and is 62.5% more sensitive and 9.4% more specific compared to DDA. Availability and implementation MS2Planner code is available at https://github.com/mohimanilab/MS2Planner. The generation of the inclusion list from MS2Planner was performed with python scripts available at https://github.com/lfnothias/IODA_MS.

47 OTHER INSTRUMENTATION↗

An expectation–maximization framework for comprehensive prediction of isoform-specific functions

Advances in RNA sequencing technologies have achieved an unprecedented accuracy in the quantification of mRNA isoforms, but our knowledge of isoform-specific functions has lagged behind. There is a need to understand the functional consequences of differential splicing, which could be supported by the generation of accurate and comprehensive isoform-specific gene ontology annotations. We present isoform interpretation, a method that uses expectation–maximization to infer isoform-specific functions based on the relationship between sequence and functional isoform similarity. We predicted isoform-specific functional annotations for 85 617 isoforms of 17 900 protein-coding human genes spanning a range of 17 430 distinct gene ontology terms. Comparison with a gold-standard corpus of manually annotated human isoform functions showed that isoform interpretation significantly outperforms state-of-the-art competing methods. We provide experimental evidence that functionally related isoforms predicted by isoform interpretation show a higher degree of domain sharing and expression correlation than functionally related genes. We also show that isoform sequence similarity correlates better with inferred isoform function than with gene-level function.

59 BASIC BIOLOGICAL SCIENCES↗

Identification of candidate host-specificity genes in Exserohilum turcicum using comparative genomics and transcriptomics

Abstract Exserohilum turcicum causes northern corn leaf blight and sorghum leaf blight. While the same species cause disease in both crops, the strains are host-specific. Here, we report the sequence and de novo annotated assemblies of one sorghum- and one maize-specific E. turcicum strain. The strains were sequenced using the PacBio Sequel II system. The total genome length for both assemblies was between 44 and 45 Mb with N50 of ∼2.5 Mb. Ninety-eight percent of the Benchmarking Universal Single-Copy Orthologs (BUSCO) for both assemblies had complete status. The estimated number of genes was 11,762 and 12,029 in the sorghum- and maize-specific isolates, respectively. Funannotate, EffectorP, SignalP, and transcriptome data were used to create functional annotation of each genome. The whole-genome comparison identified ten large-scale inversions and three translocations between the maize- and sorghum-specific strains, along with homologous genes and gene duplications. RNA was sequenced from the maize- and sorghum-specific isolate 10 days post-inoculation in maize and sorghum and from axenic cultures. Gene expression data from planta and axenic growth experiments were compared for each strain. Candidate host-specificity genes were identified by combining results from whole-genome comparison, synteny analysis, gene annotations, and transcriptome data. Overall, this study identified several candidate host-specificity genes that provide insights into E. turcicum interaction with its hosts.

Krone, Mara J. (ORCID:0000000159006624)↗

Accelerating Biological Insight for Understudied Genes

Synopsis The rapid expansion of genome sequence data is increasing the discovery of protein-coding genes across all domains of life. Annotating these genes with reliable functional information is necessary to understand evolution, to define the full biochemical space accessed by nature, and to identify target genes for biotechnology improvements. The majority of proteins are annotated based on sequence conservation with no specific biological, biochemical, genetic, or cellular function identified. Recent technical advances throughout the biological sciences enable experimental research on these understudied protein-coding genes in a broader collection of species. However, scientists have incentives and biases to continue focusing on well documented genes within their preferred model organism. This perspective suggests a research model that seeks to break historic silos of research bias by enabling interdisciplinary teams to accelerate biological functional annotation. We propose an initiative to develop coordinated projects of collaborating evolutionary biologists, cell biologists, geneticists, and biochemists that will focus on subsets of target genes in multiple model organisms. Concurrent analysis in multiple organisms takes advantage of evolutionary divergence and selection, which causes individual species to be better suited as experimental models for specific genes. Most importantly, multisystem approaches would encourage transdisciplinary critical thinking and hypothesis testing that is inherently slow in current biological research.

Zoology↗

Co-locality to co-functionality: Eukaryotic gene neighborhoods as a resource for function

Diverging from the classic paradigm of random gene order in eukaryotes, gene proximity can be leveraged to systematically identify functionally related gene neighborhoods in eukaryotes, utilizing techniques pioneered in bacteria. Current methods of identifying gene neighborhoods typically rely on sequence similarity to characterized gene products. However, this approach is not robust for non model organisms like algae, which are evolutionarily distant from well-characterized model organisms. Here, we utilize a comparative genomic approach to identify evolutionarily conserved Proximal Orthologous Gene (POG) pairs conserved across at least two taxonomic classes of green algae. A total of 317 gene neighborhoods were identified. In some cases, gene proximity appears to have been conserved since before the streptophyte-chlorophyte split, 1,000 million years ago. Using functional inferences derived from reconstructed evolutionary relationships, we identified several novel functional clusters. A putative mycosporine-like amino acid (MAA), "sunscreen", neighborhood contains genes similar to either vertebrate or cyanobacterial pathways, suggesting a novel mosaic biosynthetic pathway in green algae. One of two putative arsenic-detoxification neighborhoods includes an organoarsenical transporter (ArsJ), a glyceraldehyde 3-phosphate dehydrogenase-like gene, homologs of which are involved in arsenic detoxification in bacteria, and a novel algal-specific phosphoglycerate kinase-like gene (PGK). Mutants of the ArsJ-like transporter and PGK-like genes in Chlamydomonas reinhardtii were found to be sensitive to arsenate, providing experimental support for the role of these identified neighbors in resistance to arsenate. Potential evolutionary origins of neighborhoods are discussed, and updated annotations for formerly poorly annotated genes are presented, highlighting the potential of this strategy for functional annotation.

59 BASIC BIOLOGICAL SCIENCES↗

ECO: the Evidence and Conclusion Ontology, an update for 2022

The Evidence and Conclusion Ontology (ECO) is a community resource that provides an ontology of terms used to capture the type of evidence that supports biomedical annotations and assertions. Consistent capture of evidence information with ECO allows tracking of annotation provenance, establishment of quality control measures, and evidence-based data mining. ECO is in use by dozens of data repositories and resources with both specific and general areas of focus. ECO is continually being expanded and enhanced in response to user requests as well as our aim to adhere to community best-practices for ontology development. The ECO support team engages in multiple collaborations with other ontologies and annotating groups. Here we report on recent updates to the ECO ontology itself as well as associated resources that are available through this project. ECO project products are freely available for download from the project website (https://evidenceontology.org/) and GitHub (https://github.com/evidenceontology/evidenceontology). ECO is released into the public domain under a CC0 1.0 Universal license.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative mitogenomics of kingdom Fungi – evolutionary insights and metagenomic applications

Mitochondria are essential components of eukaryotic cells, responsible for ATP production through oxidative phosphorylation. Despite their biological importance, unique challenges have hindered the adoption of automated mitochondrial genome (mitogenome) annotation methods, obstructing mitochondrial comparative genomics in a broad evolutionary context. Using Fungi as a study system and a Joint Genome Institute (JGI) annotated high-quality reference set, we observed broad patterns of mitochondrial evolution across the kingdom. We found that the median fungal mitogenome size is 58 kb and identified exceptionally large examples over 1 Mb in Pezizomycetes. All 14 expected oxidative phosphorylation protein-coding genes, plus rps3, were generally conserved. We found evidence of major evolutionary transitions within the Ascomycota, including the transfer of mitochondrially encoded atp8 and atp9 to the nuclear genomes across the Pezizomycotina and shifts in mitogenome tRNA patterns across the kingdom. We found substantial concordance between mitochondrial and nuclear evolution, enabling us to document 3131 total fungal mitogenomes from JGI-derived metagenomic datasets. We also identified 6467 total undeclared mitogenomes embedded in Genbank fungal nuclear assemblies. We provide interactive tools for mitogenome analysis through the JGI MycoCosm platform. Collectively, this work generated nearly 10 000 new fungal mitogenome annotations, providing a foundation and resources for future exploration of comparative fungal mitogenomics.

Ahrendt, Steven R. [USDOE Joint Genome Institute (↗

The Chlamydomonas Genome Project, version 6: reference assemblies for mating type plus and minus strains reveal extensive structural mutation in the laboratory

Five versions of the Chlamydomonas reinhardtii reference genome have been produced over the last two decades. Here we present version 6, bringing significant advances in assembly quality and structural annotations. PacBio-based chromosome-level assemblies for two laboratory strains, CC-503 and CC-4532, provide resources for the plus and minus mating type alleles. We corrected major misassemblies in previous versions and validated our assemblies via linkage analyses. Contiguity increased over ten-fold and >80% of filled gaps are within genes. We used Iso-Seq and deep RNA-seq datasets to improve structural annotations, and updated gene symbols and textual annotation of functionally characterized genes via extensive manual curation. We discovered that the cell wall-less classical reference strain CC-503 exhibits genomic instability potentially caused by deletion of the helicase RECQ3, with major structural mutations identified that affect >100 genes. We therefore present the CC-4532 assembly as the primary reference, although this strain also carries unique structural mutations and is experiencing rapid proliferation of a Gypsy retrotransposon. We expect all laboratory strains to harbor gene-disrupting mutations, which should be considered when interpreting and comparing experimental results. Collectively, the resources presented here herald a new era of Chlamydomonas genomics and will provide the foundation for continued research in this important reference organism.

59 BASIC BIOLOGICAL SCIENCES↗

Global Explainability of A Deep Abstaining Classifier for Cancer Pathology Reports

We present a global explainability method to characterize sources of errors in a real-world multitask deep abstaining classifier (DAC), in the context of cancer histology prediction. Our multitask classifier, currently deployed for automated annotation of cancer pathology reports from NCI-SEER registries, was trained and evaluated on 1.04 million hand-annotated samples and makes simultaneous predictions of cancer site, subsite, histology, laterality, and behavior for each report. The DAC framework enables the model to abstain on ambiguous reports and confusing classes to achieve the target accuracy on the retained (non-abstained) samples, but at the cost of decreased coverage. Requiring 97% accuracy on the histology task caused our model to retain only 22% of all samples, mostly the less ambiguous and common classes. Local explainability with the GradInp technique provided a computationally efficient way of obtaining contextual reasoning for hundreds of thousands of individual predictions. Our method, involving dimensionality reduction of approximately 13000 aggregated local explanations (ALE), offers a tractable path to true global explainability. It enabled identification of sources of errors in histology classification, globally, as hierarchical complexity among classes, label noise, insufficient information, and conflicting evidence. This suggests several strategies for iterative improvement of our DAC, including well-designed exclusion criteria, focused annotation, and reduced penalties for errors involving hierarchically related classes.

59 BASIC BIOLOGICAL SCIENCES↗

Combining GWAS and population genomic analyses to characterize coevolution in a legume‐rhizobia symbiosis

Abstract The mutualism between legumes and rhizobia is clearly the product of past coevolution. However, the nature of ongoing evolution between these partners is less clear. To characterize the nature of recent coevolution between legumes and rhizobia, we used population genomic analysis to characterize selection on functionally annotated symbiosis genes as well as on symbiosis gene candidates identified through a two‐species association analysis. For the association analysis, we inoculated each of 202 accessions of the legume host Medicago truncatula with a community of 88 Sinorhizobia (Ensifer) meliloti strains. Multistrain inoculation, which better reflects the ecological reality of rhizobial selection in nature than single‐strain inoculation, allows strains to compete for nodulation opportunities and host resources and for hosts to preferentially form nodules and provide resources to some strains. We found extensive host by symbiont, that is, genotype‐by‐genotype, effects on rhizobial fitness and some annotated rhizobial genes bear signatures of recent positive selection. However, neither genes responsible for this variation nor annotated host symbiosis genes are enriched for signatures of either positive or balancing selection. This result suggests that stabilizing selection dominates selection acting on symbiotic traits and that variation in these traits is under mutation‐selection balance. Consistent with the lack of positive selection acting on host genes, we found that among‐host variation in growth was similar whether plants were grown with rhizobia or N‐fertilizer, suggesting that the symbiosis may not be a major driver of variation in plant growth in multistrain contexts.

59 BASIC BIOLOGICAL SCIENCES↗

Information theory and machine learning illuminate large‐scale metabolomic responses of Brachypodium distachyon to environmental change

SUMMARY Plant responses to environmental change are mediated via changes in cellular metabolomes. However, <5% of signals obtained from liquid chromatography tandem mass spectrometry (LC‐MS/MS) can be identified, limiting our understanding of how metabolomes change under biotic/abiotic stress. To address this challenge, we performed untargeted LC‐MS/MS of leaves, roots, and other organs of Brachypodium distachyon (Poaceae) under 17 organ–condition combinations, including copper deficiency, heat stress, low phosphate, and arbuscular mycorrhizal symbiosis. We found that both leaf and root metabolomes were significantly affected by the growth medium. Leaf metabolomes were more diverse than root metabolomes, but the latter were more specialized and more responsive to environmental change. We found that 1 week of copper deficiency shielded the root, but not the leaf metabolome, from perturbation due to heat stress. Machine learning (ML)‐based analysis annotated approximately 81% of the fragmented peaks versus approximately 6% using spectral matches alone. We performed one of the most extensive validations of ML‐based peak annotations in plants using thousands of authentic standards, and analyzed approximately 37% of the annotated peaks based on these assessments. Analyzing responsiveness of each predicted metabolite class to environmental change revealed significant perturbations of glycerophospholipids, sphingolipids, and flavonoids. Co‐accumulation analysis further identified condition‐specific biomarkers. To make these results accessible, we developed a visualization platform on the Bio‐Analytic Resource for Plant Biology website ( https://bar.utoronto.ca/efp_brachypodium_metabolites/cgi‐bin/efpWeb.cgi ), where perturbed metabolite classes can be readily visualized. Overall, our study illustrates how emerging chemoinformatic methods can be applied to reveal novel insights into the dynamic plant metabolome and stress adaptation.

59 BASIC BIOLOGICAL SCIENCES↗