Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gene prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

TGCM: (T)rait, (G)ene, and (C)rop Growth (M)odel Directed Targeted Gene Characterization in Sorghum (Final Technical Report)

Understanding which genes control important crop traits could help scientists develop better bioenergy and food crops more efficiently. However, plant genomes contain tens of thousands of genes, and testing each one individually is expensive and time-consuming. This project developed computational tools to predict which genes are most likely to matter, allowing researchers to focus their efforts where they will have the greatest impact. This project developed and validated integrated approaches combining machine learning, quantitative genetics, and crop growth modeling to improve the efficiency of functional gene characterization in sorghum (Sorghum bicolor), a critical bioenergy and food security crop. The research addressed a fundamental challenge in plant biology: the majority of genes in plant genomes lack experimentally validated functions, making it difficult to prioritize which genes to study using resource-intensive reverse genetics approaches.

60 APPLIED LIFE SCIENCES↗

Draft genome sequences of strains CBS6241 and CBS6242 of the basidiomycetous yeast Filobasidium floriforme

The Tremellomycetes are a species-rich group within the basidiomycete fungi; however, most analyses of this group to date have focused on pathogenic Cryptococcus species within the order Tremellales. Recent genome-assisted studies of other Tremellomycetes have identified interesting features with respect to biotechnological applications as well as the evolution of genes involved in mating and sexual development. Here, we report genome sequences of two strains of Filobasidium floriforme, a species from the order Filobasidiales, which branches basally to the Tremellales, Trichosporonales, and Holtermanniales. The assembled genomes of strains CBS6241 and CBS6242 are 27.4 Mb and 26.4 Mb in size, respectively, with 8314 and 7695 predicted protein-coding genes. Overall sequence identity at nucleic acid level between the strains is 97%. Among the predicted genes are pheromone precursor and pheromone receptor genes as well as two genes encoding homedomain (HD) transcription factors, which are predicted to be part of the mating type (MAT) locus. Sequence analysis indicates that CBS6241 and CBS6242 carry different alleles for both the pheromone/receptor genes as well as the HD transcription factors. Orthology inference identified 1482 orthogroups exclusively found in F. floriforme, some of which were involved in carbohydrate transport and metabolism. Subsequent CAZyme repertoire characterization identified 267 and 247 enzymes for CBS6241 and CBS6242, respectively, the second highest number of CAZymes among the analyzed Tremellomycete species. In addition, F. floriforme contains five CAZymes absent in other species and several plant-cell-wall degrading CAZymes with the highest copy number in Tremellomycota, indicating the biotechnological potential of this species.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluation of an ASFV RNA Helicase Gene A859L for Virus Replication and Swine Virulence

African swine fever virus (ASFV) is producing a devastating pandemic that, since 2007, has spread to a contiguous geographical area from central Europe to Asia. In July 2021, ASFV was detected in the Dominican Republic, the first report of the disease in the Americas in more than 40 years. ASFV is a large, highly complex virus harboring a large dsDNA genome that encodes for more than 150 genes. The majority of these genes have not been functionally characterized. Bioinformatics analysis predicts that ASFV gene A859L encodes for an RNA helicase, although its function has not yet been experimentally assessed. Here, we evaluated the role of the A859L gene during virus replication in cell cultures and during infection in swine. For that purpose, a recombinant virus (ASFV-G-ΔA859L) harboring a deletion of the A859L gene was developed using the highly virulent ASFV Georgia (ASFV-G) isolate as a template. Recombinant ASFV-G-ΔA859L replicates in swine macrophage cultures as efficiently as the parental virus ASFV-G, demonstrating that the A859L gene is non-essential for ASFV replication. Experimental infection of domestic pigs demonstrated that ASFV-G-ΔA859L replicates as efficiently and induces a clinical disease indistinguishable from that caused by the parental ASFV-G. These studies conclude that the predicted RNA helicase gene A859L is not involved in the processes of virus replication or disease production in swine.

59 BASIC BIOLOGICAL SCIENCES↗

Codon Optimization Improves the Prediction of Xylose Metabolism from Gene Content in Budding Yeasts

Xylose is the second most abundant monomeric sugar in plant biomass. Consequently, xylose catabolism is an ecologically important trait for saprotrophic organisms, as well as a fundamentally important trait for industries that hope to convert plant mass to renewable fuels and other bioproducts using microbial metabolism. Although common across fungi, xylose catabolism is rare within Saccharomycotina, the subphylum that contains most industrially relevant fermentative yeast species. The genomes of several yeasts unable to consume xylose have been previously reported to contain the full set of genes in the XYL pathway, suggesting the absence of a gene–trait correlation for xylose metabolism. Here, we measured growth on xylose and systematically identified XYL pathway orthologs across the genomes of 332 budding yeast species. Although the XYL pathway coevolved with xylose metabolism, we found that pathway presence only predicted xylose catabolism about half of the time, demonstrating that a complete XYL pathway is necessary, but not sufficient, for xylose catabolism. We also found that XYL1 copy number was positively correlated, after phylogenetic correction, with xylose utilization. We then quantified codon usage bias of XYL genes and found that XYL3 codon optimization was significantly higher, after phylogenetic correction, in species able to consume xylose. Finally, we showed that codon optimization of XYL2 was positively correlated, after phylogenetic correction, with growth rates in xylose medium. We conclude that gene content alone is a weak predictor of xylose metabolism and that using codon optimization enhances the prediction of xylose metabolism from yeast genome sequence data.

59 BASIC BIOLOGICAL SCIENCES↗

Hypergraph Models of Biological Networks to Identify Genes Critical to Pathogenic Viral Response

Motivation: Representing biological networks as graphs is a powerful approach to reveal underlying patterns, signatures, and critical components from high-throughput biomolecular data. However, graphs do not natively capture the multi-way relationships present among genes and proteins in biological systems such as protein complexes, metabolic reactions, and signal transduction pathways. Hypergraphs are generalizations of graphs that naturally model multi-way interactions in data, and we therefore seek to understand how they can more faithfully identify, and potentially predict, complex relationships in genomic expression data sets. Results: We compiled a novel data set of transcriptional host response to pathogenic viral infections and formulated relationships between genes as a hypergraph where hyperedges are differentially expressed genes and vertices represent conditions. We find that hypergraph betweenness centrality is a superior method for identification of genes important to viral response when compared with graph centrality. Our results demonstrate the utility of using hypergraphs to represent complex biological systems, and highlight potentially interesting biological results about host response to highly pathogenic viruses.

systems biology, hypergraph, viral infection, biol↗

Alterations in soil pH emerge as a key driver of the impact of global change on soil microbial nitrogen cycling: Evidence from a global meta-analysis

Soil nitrogen (N) cycling is critical to the productivity of terrestrial ecosystems. However, the impact of global change factors (GCFs) on the microbial mediators of N cycling pathways has yet to be synthesized, and it also remains unclear whether the response of the abundance of N-cycling genes can predict changes in their corresponding processes. We synthesized 8322 paired observations of soil microorganisms related to N cycling from field experiments in which GCFs (climate change and nutrient addition) were manipulated. We found that the abundance of soil microbes and most N-cycling genes were resistant to elevated CO 2 , experimental warming and water addition/reduction; however, N addition and the combination of N addition with other GCFs significantly increased the abundance of ammonia oxidizer bacteria (amoA-AOB). The results indicated that in steady-state (natural) conditions, the main factors driving the global abundance of soil bacteria, archaea and N-cycling genes varied in terms of the contributions of climatic and edaphic factors. However, upon manipulation of GCFs, the induced change in soil pH was the most essential factor associated with changes in the abundance of soil microbes and N-cycling genes. Notably, the changes in ammonia-oxidizing archaea (amoA-AOA) and amoA-AOB genes, in addition to genes involved in denitrification (nirS and nirK), were significantly correlated with the rates of their corresponding processes, but GCF-induced shifts in the potential nitrification rate (PNR) were explained well by changes in the abundance of the amoA-AOB gene under GCFs. In conclusion, our study highlights how ongoing GCFs impact the abundance of soil microbes and N-cycling genes, which might have a profound impact on terrestrial N cycling. Our field-based results provide new insights into the drivers of the abundance of soil microbes and N-cycling genes.

54 ENVIRONMENTAL SCIENCES↗

Rewiring the specificity of extracytoplasmic function sigma factors

Significance Bacterial phenotypes require the concerted expression of multiple genes, usually coordinated by a transcriptional regulator. Although the functions of many genes in sequenced bacterial genomes can be inferred, the regulatory networks that coordinate their expression are only known in a few model systems. Using a bioinformatic and experimental approach, we solve the DNA-specificity code of extracytoplasmic function sigma factors (ECF σs), a major class of bacterial regulators. We develop and use a high-stringency pipeline to predict the genes regulated by 67% of ECF σs in >10,000 species, providing a comprehensive look at the role of a broadly distributed family of gene regulatory proteins. This conceptual and computational framework is potentially applicable to other bacterial regulators.

59 BASIC BIOLOGICAL SCIENCES↗

Multi-Omics integration can be used to rescue metabolic information for some of the dark region of the Pseudomonas putida proteome

In every omics experiment, genes or their products are identified for which even state of the art tools are unable to assign a function. In the biotechnology chassis organism Pseudomonas putida, these proteins of unknown function make up 14% of the proteome. This missing information can bias analyses since these proteins can carry out functions which impact the engineering of organisms. As a consequence of predicting protein function across all organisms, function prediction tools generally fail to use all of the types of data available for any specific organism, including protein and transcript expression information. Additionally, the release of Alphafold predictions for all Uniprot proteins provides a novel opportunity for leveraging structural information. We constructed a bespoke machine learning model to predict the function of recalcitrant proteins of unknown function in Pseudomonas putida based on these sources of data, which annotated 1079 terms to 213 proteins. Among the predicted functions supplied by the model, we found evidence for a significant overrepresentation of nitrogen metabolism and macromolecule processing proteins. These findings were corroborated by manual analyses of selected proteins which identified, among others, a functionally unannotated operon that likely encodes a branch of the shikimate pathway.

60 APPLIED LIFE SCIENCES↗

Prediction of condition-specific regulatory genes using machine learning

Recent advances in genomic technologies have generated data on large-scale protein–DNA interactions and open chromatin regions for many eukaryotic species. How to identify condition-specific functions of transcription factors using these data has become a major challenge in genomic research. To solve this problem, we have developed a method called ConSReg, which provides a novel approach to integrate regulatory genomic data into predictive machine learning models of key regulatory genes. Using Arabidopsis as a model system, we tested our approach to identify regulatory genes in data sets from single cell gene expression and from abiotic stress treatments. Our results showed that ConSReg accurately predicted transcription factors that regulate differentially expressed genes with an average auROC of 0.84, which is 23.5–25% better than enrichment-based approaches. To further validate the performance of ConSReg, we analyzed an independent data set related to plant nitrogen responses. ConSReg provided better rankings of the correct transcription factors in 61.7% of cases, which is three times better than other plant tools. We applied ConSReg to Arabidopsis single cell RNA-seq data, successfully identifying candidate regulatory genes that control cell wall formation. Our methods provide a new approach to define candidate regulatory genes using integrated genomic data in plants.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Predictive Models of Genetic Redundancy in Arabidopsis thaliana

Abstract Genetic redundancy refers to a situation where an individual with a loss-of-function mutation in one gene (single mutant) does not show an apparent phenotype until one or more paralogs are also knocked out (double/higher-order mutant). Previous studies have identified some characteristics common among redundant gene pairs, but a predictive model of genetic redundancy incorporating a wide variety of features derived from accumulating omics and mutant phenotype data is yet to be established. In addition, the relative importance of these features for genetic redundancy remains largely unclear. Here, we establish machine learning models for predicting whether a gene pair is likely redundant or not in the model plant Arabidopsis thaliana based on six feature categories: functional annotations, evolutionary conservation including duplication patterns and mechanisms, epigenetic marks, protein properties including posttranslational modifications, gene expression, and gene network properties. The definition of redundancy, data transformations, feature subsets, and machine learning algorithms used significantly affected model performance based on holdout, testing phenotype data. Among the most important features in predicting gene pairs as redundant were having a paralog(s) from recent duplication events, annotation as a transcription factor, downregulation during stress conditions, and having similar expression patterns under stress conditions. We also explored the potential reasons underlying mispredictions and limitations of our studies. This genetic redundancy model sheds light on characteristics that may contribute to long-term maintenance of paralogs, and will ultimately allow for more targeted generation of functionally informative double mutants, advancing functional genomic studies.

59 BASIC BIOLOGICAL SCIENCES↗

MultiPhATE2: code for functional annotation and comparison of phage genomes

To address a need for improved tools for annotation and comparative genomics of bacteriophage genomes, we developed multiPhATE2. As an extension of multiPhATE, a functional annotation code released previously, multiPhATE2 performs gene finding using multiple algorithms, compares the results of the algorithms, performs functional annotation of coding sequences, and incorporates additional search algorithms and databases to extend the search space of the original code. MultiPhATE2 performs gene matching among sets of closely related bacteriophage genomes, and uses multiprocessing to speed computations. MultiPhATE2 can be re-started at multiple points within the workflow to allow the user to examine intermediate results and adjust the subsequent computations accordingly. In addition, multiPhATE2 accommodates custom gene calls and sequence databases, again adding flexibility. MultiPhATE2 was implemented in Python 3.7 and runs as a command-line code under Linux or MAC operating systems. Full documentation is provided as a README file and a Wiki website.

59 BASIC BIOLOGICAL SCIENCES↗

Using online tools at the Bovine Genome Database to manually annotate genes in the new reference genome

With the availability of a new highly contiguous Bos taurus reference genome assembly (ARS-UCD1.2), it is the opportune time to upgrade the bovine gene set by seeking input from researchers. Furthermore, advances in graphical genome annotation tools now make it possible for researchers to leverage sequence data generated with the latest technologies to collaboratively curate genes. For many years the Bovine Genome Database (BGD) has provided tools such as the APOLLO genome annotation editor to support manual bovine gene curation. The goal of this paper is to explain the reasoning behind the decisions made in the manual gene curation process while providing examples using the existing BGD tools. We will describe the sources of gene annotation evidence provided at the BGD, including RNAseq and Iso-Seq data. We will also explain how to interpret various data visualizations when curating gene models, and will demonstrate the value of manual gene annotation. The process described here can be applied to manual gene curation for other species with similar tools. With a better understanding of manual gene annotation, researchers will be encouraged to edit gene models and contribute to the enhancement of livestock gene sets.

59 BASIC BIOLOGICAL SCIENCES↗

Nonparallel transcriptional divergence during parallel adaptation

Abstract How underlying mechanisms bias evolution toward predictable outcomes remains an area of active debate. In this study, we leveraged phenotypic plasticity and parallel adaptation across independent lineages of Trinidadian guppies ( Poecilia reticulata ) to assess the predictability of gene expression evolution during parallel adaptation. Trinidadian guppies have repeatedly and independently adapted to high‐ and low‐predation environments in the wild. We combined this natural experiment with a laboratory breeding design to attribute transcriptional variation to the genetic influences of population of origin and developmental plasticity in response to rearing with or without predators. We observed substantial gene expression plasticity, as well as the evolution of expression plasticity itself, across populations. Genes exhibiting expression plasticity within populations were more likely to also differ in expression between populations, with the direction of population differences more likely to be opposite those of plasticity. While we found more overlap than expected by chance in genes differentially expressed between high‐ and low‐predation populations from distinct evolutionary lineages, the majority of differentially expressed genes were not shared between lineages. Our data suggest alternative transcriptional configurations associated with shared phenotypes, highlighting a role for transcriptional flexibility in the parallel phenotypic evolution of a species known for rapid adaptation.

Fischer, Eva K.↗

Tandem repeats in giant archaeal Borg elements undergo rapid evolution and create new intrinsically disordered regions in proteins

Borgs are huge, linear extrachromosomal elements associated with anaerobic methane-oxidizing archaea. Striking features of Borg genomes are pervasive tandem direct repeat (TR) regions. Here, we present six new Borg genomes and investigate the characteristics of TRs in all ten complete Borg genomes. We find that TR regions are rapidly evolving, recently formed, arise independently, and are virtually absent in host Methanoperedens genomes. Flanking partial repeats and A-enriched character constrain the TR formation mechanism. TRs can be in intergenic regions, where they might serve as regulatory RNAs, or in open reading frames (ORFs). TRs in ORFs are under very strong selective pressure, leading to perfect amino acid TRs (aaTRs) that are commonly intrinsically disordered regions. Proteins with aaTRs are often extracellular or membrane proteins, and functionally similar or homologous proteins often have aaTRs composed of the same amino acids. We propose that Borg aaTR-proteins functionally diversify Methanoperedens and all TRs are crucial for specific Borg–host associations and possibly cospeciation.

59 BASIC BIOLOGICAL SCIENCES↗

Antimicrobial resistance and genomic characterization of Salmonella Dublin isolates in cattle from the United States

Salmonella enterica subspecies enterica serotype Dublin is a host-adapted serotype in cattle, associated with enteritis and systemic disease. The primary clinical manifestation of Salmonella Dublin infection in cattle, especially calves, is respiratory disease. While rare in humans, it can cause severe illness, including bacteremia, with hospitalization and death. In the United States, S . Dublin has become one of the most multidrug-resistant serotypes. The objective of this study was to characterize S . Dublin isolates from sick cattle by analyzing phenotypic and genotypic antimicrobial resistance (AMR) profiles, the presence of plasmids, and phylogenetic relationships. S . Dublin isolates (n = 140) were selected from submissions to the NVSL for Salmonella serotyping (2014–2017) from 21 states. Isolates were tested for susceptibility against 14 class-representative antimicrobial drugs. Resistance profiles were determined using the ABRicate with Resfinder and NCBI databases, AMRFinder and PointFinder. Plasmids were detected using ABRicate with PlasmidFinder. Phylogeny was determined using vSNP. We found 98% of the isolates were resistant to more than 4 antimicrobials. Only 1 isolate was pan-susceptible and had no predicted AMR genes. All S . Dublin isolates were susceptible to azithromycin and meropenem. They showed 96% resistance to sulfonamides, 97% to tetracyclines, 95% to aminoglycosides and 85% to beta-lactams. The most common AMR genes were: sulf2 and tetA (98.6%), aph(6)-Id (97.9%), aph(3’’)-Ib, (97.1%), floR (94.3%), and blaCMY-2 (85.7%). All quinolone resistant isolates presented mutations in gyr A. Ten plasmid types were identified among all isolates with IncA/C2, IncX1, and IncFII(S) being the most frequent. The S . Dublin isolates show low genomic genetic diversity. This study provided antimicrobial susceptibility and genomic insight into S . Dublin clinical isolates from cattle in the U.S. Further sequence analysis integrating food and human origin S . Dublin isolates may provide valuable insight on increased virulence observed in humans.

60 APPLIED LIFE SCIENCES↗

The multifaceted role of c-di-AMP signaling in the regulation of Porphyromonas gingivalis lipopolysaccharide structure and function

This study unveils the intricate functional association between cyclic di-3’,5’-adenylic acid (c-di-AMP) signaling, cellular bioenergetics, and the regulation of lipopolysaccharide (LPS) profile in Porphyromonas gingivalis, a Gram-negative obligate anaerobe considered as a keystone pathogen involved in the pathogenesis of chronic periodontitis. Previous research has identified variations in P. gingivalis LPS profile as a major virulence factor, yet the underlying mechanism of its modulation has remained elusive. We employed a comprehensive methodological approach, combining two mutants exhibiting varying levels of c-di-AMP compared to the wild type, alongside an optimized analytical methodology that combines conventional mass spectrometry techniques with a novel approach known as FLAT n . We demonstrate that c-di-AMP acts as a metabolic nexus, connecting bioenergetic status to nuanced shifts in fatty acid and glycosyl profiles within P. gingivalis LPS. Notably, the predicted regulator gene cdaR, serving as a potent regulator of c-di-AMP synthesis, was found essential for producing N-acetylgalactosamine and an unidentified glycolipid class associated with the LPS profile. The multifaceted roles of c-di-AMP in bacterial physiology are underscored, emphasizing its significance in orchestrating adaptive responses to stimuli. Furthermore, our findings illuminate the significance of LPS variations and c-di-AMP signaling in determining the biological activities and immunostimulatory potential of P. gingivalis LPS, promoting a pathoadaptive strategy. The study expands the understanding of c-di-AMP pathways in Gram-negative species, laying a foundation for future investigations into the mechanisms governing variations in LPS structure at the molecular level and their implications for host-pathogen interactions.

59 BASIC BIOLOGICAL SCIENCES↗

Untargeted GC-MS Metabolic Profiling of Anaerobic Gut Fungi Reveals Putative Terpenoids and Strain-Specific Metabolites

Background/Objectives: Anaerobic gut fungi (Neocallimastigomycota) are biotechnologically relevant, lignocellulose-degrading microbes with under-explored biosynthetic potential for secondary metabolites. Untargeted metabolomic profiling with gas chromatography–mass spectrometry (GC-MS) was applied to two gut fungal strains, Anaeromyces robustus and Caecomyces churrovis, to establish a foundational metabolomic dataset to identify metabolites and provide insights into gut fungal metabolic capabilities. Methods: Gut fungi were cultured anaerobically in rumen-fluid-based media with a soluble substrate (cellobiose), and metabolites were extracted using the Metabolite, Protein, and Lipid Extraction (MPLEx) method, enabling metabolomic and proteomic analysis from the same cell samples. Samples were derivatized and analyzed via GC-MS, followed by compound identification by spectral matching to reference databases, molecular networking, and statistical analyses. Results: Distinct metabolites were identified between A. robustus and C. churrovis, including 2,3-dihydroxyisovaleric acid produced by A. robustus and maltotriitol, maltotriose, and melibiose produced by C. churrovis. C. churrovis may polymerize maltotriose to form an extracellular polysaccharide, like pullulan. GC-MS profiling potentially captured sufficiently volatile products of proteomically detected, putative non-ribosomal peptide synthetases and polyketide synthases of A. robustus and C. churrovis. The triterpene squalene and triterpenoid tetrahymanol were putatively identified in A. robustus and C. churrovis. Their conserved, predicted biosynthetic genes—squalene synthase and squalene tetrahymanol cyclase—were identified in A. robustus, C. churrovis, and other anaerobic gut fungal genera. Conclusions: This study provides a foundational, untargeted metabolomic dataset to unmask gut fungal metabolic pathways and biosynthetic potential and to prioritize future efforts for compound isolation and identification.

Biochemistry & Molecular Biology↗