Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20

High-throughput identification of novel heat tolerance genes via genome-wide pooled mutant screens in the model green alga Chlamydomonas reinhardtii

Different high temperatures adversely affect crop and algal yields with various responses in photosynthetic cells. The list of genes required for thermotolerance remains elusive. Additionally, it is unclear how carbon source availability affects heat responses in plants and algae. Here, we utilized the insertional, indexed, genome-saturating mutant library of the unicellular, eukaryotic green alga Chlamydomonas reinhardtii to perform genome-wide, quantitative, pooled screens under moderate (35°C) or acute (40°C) high temperatures with or without organic carbon sources. We identified heat-sensitive mutants based on quantitative growth rates and identified putative heat tolerance genes (HTGs). By triangulating HTGs with heat-induced transcripts or proteins in wildtype cultures and MapMan functional annotations, we presented a high/medium-confidence list of 933 Chlamydomonas genes with putative roles in heat tolerance. Triangulated HTGs include those with known thermotolerance roles and novel genes with little or no functional annotation. About 50% of these high-confidence HTGs in Chlamydomonas have orthologs in green lineage organisms, including crop species. Arabidopsis thaliana mutants deficient in the ortholog of a high-confidence Chlamydomonas HTG were also heat sensitive. This work expands our knowledge of heat responses in photosynthetic cells and provides engineering targets to improve thermotolerance in algae and crops.

59 BASIC BIOLOGICAL SCIENCES↗

A large sequenced mutant library – valuable reverse genetic resource that covers 98% of sorghum genes

SUMMARY Mutant populations are crucial for functional genomics and discovering novel traits for crop breeding. Sorghum , a drought and heat‐tolerant C4 species, requires a vast, large‐scale, annotated, and sequenced mutant resource to enhance crop improvement through functional genomics research. Here, we report a sorghum large‐scale sequenced mutant population with 9.5 million ethyl methane sulfonate (EMS)‐induced mutations that covered 98% of sorghum's annotated genes using inbred line BTx623. Remarkably, a total of 610 320 mutations within the promoter and enhancer regions of 18 000 and 11 790 genes, respectively, can be leveraged for novel research of cis ‐regulatory elements. A comparison of the distribution of mutations in the large‐scale mutant library and sorghum association panel (SAP) provides insights into the influence of selection. EMS‐induced mutations appeared to be random across different regions of the genome without significant enrichment in different sections of a gene, including the 5′ UTR, gene body, and 3′‐UTR. In contrast, there were low variation density in the coding and UTR regions in the SAP. Based on the K a / K s value, the mutant library (~1) experienced little selection, unlike the SAP (0.40), which has been strongly selected through breeding. All mutation data are publicly searchable through SorbMutDB ( https://www.depts.ttu.edu/igcast/sorbmutdb.php ) and SorghumBase ( https://sorghumbase.org/ ). This current large‐scale sequence‐indexed sorghum mutant population is a crucial resource that enriched the sorghum gene pool with novel diversity and a highly valuable tool for the Poaceae family, that will advance plant biology research and crop breeding.

59 BASIC BIOLOGICAL SCIENCES↗

The protein structurome of Orthornavirae and its dark matter

Metatranscriptomics is uncovering more and more diverse families of viruses with RNA genomes comprising the viral kingdom Orthornavirae in the realm Riboviria. Thorough protein annotation and comparison are essential to get insights into the functions of viral proteins and virus evolution. In addition to sequence- and hmm profile-based methods, protein structure comparison adds a powerful tool to uncover protein functions and relationships. We constructed an Orthornavirae “structurome” consisting of already annotated as well as unannotated (“dark matter”) proteins and domains encoded in viral genomes. We used protein structure modeling and similarity searches to illuminate the remaining dark matter in hundreds of thousands of orthornavirus genomes. The vast majority of the dark matter domains showed either “generic” folds, such as single α-helices, or no high confidence structure predictions. Nevertheless, a variety of lineage-specific globular domains that were new either to orthornaviruses in general or to particular virus families were identified within the proteomic dark matter of orthornaviruses, including several predicted nucleic acid-binding domains and nucleases. In addition, we identified a case of exaptation of a cellular nucleoside monophosphate kinase as an RNA-binding protein in several virus families. Notwithstanding the continuing discovery of numerous orthornaviruses, it appears that all the protein domains conserved in large groups of viruses have already been identified. The rest of the viral proteome seems to be dominated by poorly structured domains including intrinsically disordered ones that likely mediate specific virus-host interactions.

59 BASIC BIOLOGICAL SCIENCES↗

High Performance Approximate Computing

This code repository contains the implementation of the "High-Performance Approximate Computing" (HPAC) toolkit. The toolkit allows you to approximate your own C/C++. The developer uses "pragma's" to annotate code regions as approximate. The compiler extensions lower these pragmas to either compiletime approximate techniques or runtime approximation techniques. At execution time, the implemented runtime system decides which annotated regions it should approximate. HPAC also provides a set of script utilities. The utilities perform a grid search within approximation parameters and performance. The user can analyze the raw data to identify optimal approximation techniques for the application.

Parasyris, Konstantinos↗

FrESCO

The National Cancer Institute (NCI) monitors population level cancer trends as part of its Surveillance, Epidemiology, and End Results (SEER) program. This program consists of state or regional level cancer registries which collect, analyze, and annotate cancer pathology reports. From these annotated pathology reports, each individual registry aggregates cancer phenotype information and summary statistics about cancer prevalence to facilitate population level monitoring of cancer incidence. Extracting cancer phenotype from these reports is a labor intensive task, requiring specialized knowledge about the reports and cancer. Automating this information extraction process from cancer pathology reports has the potential to improve not only the quality of the data by extracting information in a consistent manner across registries, but to improve the quality of patient outcomes by reducing the time to assimilate new data and enabling time-sensitive applications such as precision medicine. Here we present FrESCO: Framework for Exploring Scalable Computational Oncology, a modular deep-learning natural language processing (NLP) library for extracting pathology information from clinical text documents.

Spannaus, Adam [Oak Ridge National Lab. (ORNL), Oa↗

BiGEST

Natural products have provided a rich reservoir of beneficial compounds in public health including antibiotics, therapeutics, and immunosuppressants. These natural products are synthesized by enzymes encoded by Biosynthetic Gene Clusters (BGCs), clusters of co-localized biosynthetic genes. Computational detection of BGCs has become a crucial step in natural product discovery. While this process has been facilitated in bacterial and fungal organisms thanks to the currently available tools (e.g., antiSMASH), a large spectrum of eukaryotic organisms have been neglected by these existing tools due to the scarcity and incompleteness of genome annotation resources. Here, we introduce Biosynthetic Gene cluster Extensive Search Tool (BiGEST) to provide an extensive annotation-free search for BGCs in diverse eukaryotic organisms. As a result, BiGEST uncovers eukaryotic BGCs that could be undetected by other BGC detection tools.

Adriani, Lisa↗

DiMER

SAND2025-04145O DiMER is a Python based tool that helps researchers understand the functions of genes by searching through multiple biological databases. It takes user-provided data and scans various databases to find the best matches for gene functions, generating a clear summary of results. DiMER identifies the most relevant functional annotations and improves upon previous annotations by replacing instances of "unknown protein function" with more accurate descriptions. DiMER requires minimal setup. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Mageeney, Catherine [Sandia National Lab. (SNL-CA)↗

pixelvar79/ESGAN-Flowering-Detection-paper

Machine learning (ML) can accelerate biological research. However, the adoption of such tools to facilitate phenotyping based on sensor data has been limited by (i) the need for a large amount of human-annotated training data for each context in which the tool is used and (ii) phenotypes varying across contexts defined in terms of genetics and environment. This is a major bottleneck because acquiring training data is generally costly and time-consuming. This study demonstrates how a ML approach can address these challenges by minimizing the amount of human supervision needed for tool building. A case study was performed to compare ML approaches that examine images collected by an uncrewed aerial vehicle to determine the presence/absence of panicles (i.e. “heading”) across thousands of field plots containing genetically diverse breeding populations of 2 Miscanthus species. Automated analysis of aerial imagery enabled the identification of heading approximately 9 times faster than in-field visual inspection by humans. Leveraging an Efficiently Supervised Generative Adversarial Network (ESGAN) learning strategy reduced the requirement for human-annotated data by 1 to 2 orders of magnitude compared to traditional, fully supervised learning approaches. The ESGAN model learned the salient features of the data set by using thousands of unlabeled images to inform the discriminative ability of a classifier so that it required minimal human-labeled training data. This method can accelerate the phenotyping of heading date as a measure of flowering time in Miscanthus across diverse contexts (e.g. in multistate trials) and opens avenues to promote the broad adoption of ML tools.

Varela, Sebastian↗

Deep active learning for classifying cancer pathology reports

Abstract Background Automated text classification has many important applications in the clinical setting; however, obtaining labelled data for training machine learning and deep learning models is often difficult and expensive. Active learning techniques may mitigate this challenge by reducing the amount of labelled data required to effectively train a model. In this study, we analyze the effectiveness of 11 active learning algorithms on classifying subsite and histology from cancer pathology reports using a Convolutional Neural Network as the text classification model. Results We compare the performance of each active learning strategy using two differently sized datasets and two different classification tasks. Our results show that on all tasks and dataset sizes, all active learning strategies except diversity-sampling strategies outperformed random sampling, i.e., no active learning. On our large dataset (15K initial labelled samples, adding 15K additional labelled samples each iteration of active learning), there was no clear winner between the different active learning strategies. On our small dataset (1K initial labelled samples, adding 1K additional labelled samples each iteration of active learning), marginal and ratio uncertainty sampling performed better than all other active learning techniques. We found that compared to random sampling, active learning strongly helps performance on rare classes by focusing on underrepresented classes. Conclusions Active learning can save annotation cost by helping human annotators efficiently and intelligently select which samples to label. Our results show that a dataset constructed using effective active learning techniques requires less than half the amount of labelled data to achieve the same performance as a dataset constructed using random sampling.

59 BASIC BIOLOGICAL SCIENCES↗

Asc-Seurat: analytical single-cell Seurat-based web application

Abstract Background Single-cell RNA sequencing (scRNA-seq) has revolutionized the study of transcriptomes, arising as a powerful tool for discovering and characterizing cell types and their developmental trajectories. However, scRNA-seq analysis is complex, requiring a continuous, iterative process to refine the data and uncover relevant biological information. A diversity of tools has been developed to address the multiple aspects of scRNA-seq data analysis. However, an easy-to-use web application capable of conducting all critical steps of scRNA-seq data analysis is still lacking. Summary We present Asc-Seurat, a feature-rich workbench, providing an user-friendly and easy-to-install web application encapsulating tools for an all-encompassing and fluid scRNA-seq data analysis. Asc-Seurat implements functions from the Seurat package for quality control, clustering, and genes differential expression. In addition, Asc-Seurat provides a pseudotime module containing dozens of models for the trajectory inference and a functional annotation module that allows recovering gene annotation and detecting gene ontology enriched terms. We showcase Asc-Seurat’s capabilities by analyzing a peripheral blood mononuclear cell dataset. Conclusions Asc-Seurat is a comprehensive workbench providing an accessible graphical interface for scRNA-seq analysis by biologists. Asc-Seurat significantly reduces the time and effort required to analyze and interpret the information in scRNA-seq datasets.

60 APPLIED LIFE SCIENCES↗

Reverse engineering environmental metatranscriptomes clarifies best practices for eukaryotic assembly

Abstract Background Diverse communities of microbial eukaryotes in the global ocean provide a variety of essential ecosystem services, from primary production and carbon flow through trophic transfer to cooperation via symbioses. Increasingly, these communities are being understood through the lens of omics tools, which enable high-throughput processing of diverse communities. Metatranscriptomics offers an understanding of near real-time gene expression in microbial eukaryotic communities, providing a window into community metabolic activity. Results Here we present a workflow for eukaryotic metatranscriptome assembly, and validate the ability of the pipeline to recapitulate real and manufactured eukaryotic community-level expression data. We also include an open-source tool for simulating environmental metatranscriptomes for testing and validation purposes. We reanalyze previously published metatranscriptomic datasets using our metatranscriptome analysis approach. Conclusion We determined that a multi-assembler approach improves eukaryotic metatranscriptome assembly based on recapitulated taxonomic and functional annotations from an in-silico mock community. The systematic validation of metatranscriptome assembly and annotation methods provided here is a necessary step to assess the fidelity of our community composition measurements and functional content assignments from eukaryotic metatranscriptomes.

Krinos, Arianna I. (ORCID:0000000197678392)↗

Pickaxe: a Python library for the prediction of novel metabolic reactions

Abstract Background Biochemical reaction prediction tools leverage enzymatic promiscuity rules to generate reaction networks containing novel compounds and reactions. The resulting reaction networks can be used for multiple applications such as designing novel biosynthetic pathways and annotating untargeted metabolomics data. It is vital for these tools to provide a robust, user-friendly method to generate networks for a given application. However, existing tools lack the flexibility to easily generate networks that are tailor-fit for a user’s application due to lack of exhaustive reaction rules, restriction to pre-computed networks, and difficulty in using the software due to lack of documentation. Results Here we present Pickaxe, an open-source, flexible software that provides a user-friendly method to generate novel reaction networks. This software iteratively applies reaction rules to a set of metabolites to generate novel reactions. Users can select rules from the prepackaged JN1224min ruleset, derived from MetaCyc, or define their own custom rules. Additionally, filters are provided which allow for the pruning of a network on-the-fly based on compound and reaction properties. The filters include chemical similarity to target molecules, metabolomics, thermodynamics, and reaction feasibility filters. Example applications are given to highlight the capabilities of Pickaxe: the expansion of common biological databases with novel reactions, the generation of industrially useful chemicals from a yeast metabolome database, and the annotation of untargeted metabolomics peaks from an E. coli dataset. Conclusion Pickaxe predicts novel metabolic reactions and compounds, which can be used for a variety of applications. This software is open-source and available as part of the MINE Database python package ( https://pypi.org/project/minedatabase/ ) or on GitHub ( https://github.com/tyo-nu/MINE-Database ). Documentation and examples can be found on Read the Docs ( https://mine-database.readthedocs.io/en/latest/ ). Through its documentation, pre-packaged features, and customizable nature, Pickaxe allows users to generate novel reaction networks tailored to their application.

59 BASIC BIOLOGICAL SCIENCES↗

An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models

Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.

59 BASIC BIOLOGICAL SCIENCES↗

An evaluation of GPT models for phenotype concept recognition

Clinical deep phenotyping and phenotype annotation play a critical role in both the diagnosis of patients with rare disorders as well as in building computationally-tractable knowledge in the rare disorders field. These processes rely on using ontology concepts, often from the Human Phenotype Ontology, in conjunction with a phenotype concept recognition task (supported usually by machine learning methods) to curate patient profiles or existing scientific literature. With the significant shift in the use of large language models (LLMs) for most NLP tasks, we examine the performance of the latest Generative Pre-trained Transformer (GPT) models underpinning ChatGPT as a foundation for the tasks of clinical phenotyping and phenotype annotation. The experimental setup of the study included seven prompts of various levels of specificity, two GPT models (gpt-3.5-turbo and gpt-4.0) and two established gold standard corpora for phenotype recognition, one consisting of publication abstracts and the other clinical observations. The best run, using in-context learning, achieved 0.58 document-level F1 score on publication abstracts and 0.75 document-level F1 score on clinical observations, as well as a mention-level F1 score of 0.7, which surpasses the current best in class tool. Without in-context learning, however, performance is significantly below the existing approaches. Our experiments show that gpt-4.0 surpasses the state of the art performance if the task is constrained to a subset of the target ontology where there is prior knowledge of the terms that are expected to be matched. While the results are promising, the non-deterministic nature of the outcomes, the high cost and the lack of concordance between different runs using the same prompt and input make the use of these LLMs challenging for this particular task.

59 BASIC BIOLOGICAL SCIENCES↗

AlloSHP: deconvoluting single homeologous polymorphism for phylogenetic analysis of allopolyploids

Background The genomic and evolutionary study of allopolyploid organisms involves multiple copies of homeologous chromosomes, making their assembly, annotation, and phylogenetic analysis challenging. Bioinformatics tools and protocols have been developed to study polyploid genomes, but sometimes require the assembly of their genomes, or at least the genes, limiting their use. Results We have developed AlloSHP, a command-line tool for detecting and extracting single homeologous polymorphisms (SHPs) from the subgenomes of allopolyploid species. This tool integrates three main algorithms, WGA, VCF2ALIGNMENT and VCF2SYNTENY, and allows the detection of SHPs for the study of diploid-polyploid complexes with available diploid progenitor genomes, without assembling and annotating the genomes of the allopolyploids under study. AlloSHP has been validated on three diploid-polyploid plant complexes, Brachypodium, Brassica, and Triticum-Aegilops, and a set of synthetic hybrid yeasts and their progenitors of the genus Saccharomyces. The results and congruent phylogenies obtained from the four datasets demonstrate the potential of AlloSHP for the evolutionary analysis of allopolyploids with a wide range of ploidy and genome sizes. Conclusions AlloSHP combines the strategies of simultaneous mapping against multiple reference genomes and syntenic alignment of these genomes to call SHPs, using as input data a single VCF file and the reference genomes of the known or closest extant diploid progenitor species. This novel approach provides a valuable tool for the evolutionary study of allopolyploid species, both at the interspecific and intraspecific levels, allowing the simultaneous analysis of a large number of accessions and avoiding the complex process of assembling polyploid genomes.

Allopolyploids↗

METABOLIC: high-throughput profiling of microbial genomes for functional traits, metabolism, biogeochemistry, and community-scale functional networks

Background Advances in microbiome science are being driven in large part due to our ability to study and infer microbial ecology from genomes reconstructed from mixed microbial communities using metagenomics and single-cell genomics. Such omics-based techniques allow us to read genomic blueprints of microorganisms, decipher their functional capacities and activities, and reconstruct their roles in biogeochemical processes. Currently available tools for analyses of genomic data can annotate and depict metabolic functions to some extent; however, no standardized approaches are currently available for the comprehensive characterization of metabolic predictions, metabolite exchanges, microbial interactions, and microbial contributions to biogeochemical cycling. Results We present METABOLIC (METabolic And BiogeOchemistry anaLyses In miCrobes), a scalable software to advance microbial ecology and biogeochemistry studies using genomes at the resolution of individual organisms and/or microbial communities. The genome-scale workflow includes annotation of microbial genomes, motif validation of biochemically validated conserved protein residues, metabolic pathway analyses, and calculation of contributions to individual biogeochemical transformations and cycles. The community-scale workflow supplements genome-scale analyses with determination of genome abundance in the microbiome, potential microbial metabolic handoffs and metabolite exchange, reconstruction of functional networks, and determination of microbial contributions to biogeochemical cycles. METABOLIC can take input genomes from isolates, metagenome-assembled genomes, or single-cell genomes. Results are presented in the form of tables for metabolism and a variety of visualizations including biogeochemical cycling potential, representation of sequential metabolic transformations, community-scale microbial functional networks using a newly defined metric “MW-score” (metabolic weight score), and metabolic Sankey diagrams. METABOLIC takes ~ 3 h with 40 CPU threads to process ~ 100 genomes and corresponding metagenomic reads within which the most compute-demanding part of hmmsearch takes ~ 45 min, while it takes ~ 5 h to complete hmmsearch for ~ 3600 genomes. Tests of accuracy, robustness, and consistency suggest METABOLIC provides better performance compared to other software and online servers. To highlight the utility and versatility of METABOLIC, we demonstrate its capabilities on diverse metagenomic datasets from the marine subsurface, terrestrial subsurface, meadow soil, deep sea, freshwater lakes, wastewater, and the human gut. Conclusion METABOLIC enables the consistent and reproducible study of microbial community ecology and biogeochemistry using a foundation of genome-informed microbial metabolism, and will advance the integration of uncultivated organisms into metabolic and biogeochemical models. METABOLIC is written in Perl and R and is freely available under GPLv3 at https://github.com/AnantharamanLab/METABOLIC.

59 BASIC BIOLOGICAL SCIENCES↗

Data for "Genetics of flooding tolerance in an F2 Miscanthus sacchariflorus ssp. lutarioriparius × M. sinensis population"

This dataset contains all data and supplementary materials from "Genetics of flooding tolerance in an F2 Miscanthus sacchariflorus ssp. lutarioriparius × M. sinensis population". 1. The dataset S1 table contains the raw phenotypic data collected during the experiment. 2. The dataset S2 table contains the LSmean values for the 24 traits studied. 3. The dataset S3 table contains the TASSEL GBSv2 map, marker information, and genotype data used for mapping. 4. The dataset S4 table contains information on candidate genes found in each of the QTL intervals. 5. The dataset S5 table contains the GO annotations and KEGG enrichment analyses for those candidate genes. 6. The dataset S6 table contains information on the sequences used to classify AP2 ERF transcription factors. 7. The dataset S7 table contains information on AP2 ERF orthologs between Miscanthus and rice based on synteny. 8. Supplementary file 1 contains the ANOVA results using the raw phenotypic data collected from protocol "A". 9. Supplementary file 2 contains the ANOVA results using the raw phenotypic data collected from protocol "B". 10. Supplementary file 3 contains notes on the comparison of SNP calling methods. 11. Supplementary file 4 is a script for analyzing candidate genes found in QTL intervals.

Miscanthus, flood, partial submergence, complete s↗

Layer-wise Imaging Dataset from Powder Bed Additive Manufacturing Processes for Machine Learning Applications (Peregrine v2021-03)

This dataset contains layer-wise powder bed images from three different powder bed printing technologies – laser powder bed fusion, electron beam powder bed fusion, and binder jetting. This dataset was collected and annotated using the internally-developed Peregrine software tool and is designed primarily to facilitate research into anomaly defect detection using image segmentation or similar techniques. A total of 20 layers are provided for each printing technology, with each layer of data consisting of one or more calibrated images and an annotation file containing pixel-wise ground truth labels. The ground truths were labeled by domain experts, typically printer technicians. Data in this release were collected at Oak Ridge National Laboratory between 2016 and 2020 and were compiled in March 2021.

36 MATERIALS SCIENCE↗