Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gene prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Genome-based analysis for the bioactive potential of Streptomyces yeochonensis CN732, an acidophilic filamentous soil actinobacterium

Background: Acidophilic members of the genus Streptomyces can be a good source for novel secondary metabolites and degradative enzymes of biopolymers. In this study, a genome-based approach on Streptomyces yeochonensis CN732, a representative neutrotolerant acidophilic streptomycete, was employed to examine the biosynthetic as well as enzymatic potential, and also presence of any genetic tools for adaptation in acidic environment. Results: A high quality draft genome (7.8Mb) of S. yeochonensis CN732 was obtained with a G+C content of 73.53% and 6549 protein coding genes. The in silico analysis predicted presence of multiple biosynthetic gene clusters (BGCs), which showed similarity with those for antimicrobial, anticancer or antiparasitic compounds. However, the low levels of similarity with known BGCs for most cases suggested novelty of the metabolites from those predicted gene clusters. The production of various novel metabolites was also confirmed from the combined high performance liquid chromatography-mass spectrometry analysis. Through comparative genome analysis with related Streptomyces species, genes specific to strain CN732 and also those specific to neutrotolerant acidophilic species could be identified, which showed that genes for metabolism in diverse environment were enriched among acidophilic species. In addition, the presence of strain specific genes for carbohydrate active enzymes (CAZyme) along with many other singletons indicated uniqueness of the genetic makeup of strain CN732. The presence of cysteine transpeptidases (sortases) among the BGCs was also observed from this study, which implies their putative roles in the biosynthesis of secondary metabolites. Conclusions: This study highlights the bioactive potential of strain CN732, an acidophilic streptomycete with regard to secondary metabolite production and biodegradation potential using genomics based approach. The comparative genome analysis revealed genes specific to CN732 and also those among acidophilic species, which could give some insights into the adaptation of microbial life in acidic environment.

59 BASIC BIOLOGICAL SCIENCES↗

Bayesian filtering for model predictive control of stochastic gene expression in single cells

This study describes a method for controlling the production of protein in individual cells using stochastic models of gene expression. By combining modern microscopy platforms with optogenetic gene expression, experimentalists are able to accurately apply light to individual cells, which can induce protein production. Here we use a finite state projection based stochastic model of gene expression, along with Bayesian state estimation to control protein copy numbers within individual cells. We compare this method to previous methods that use population based approaches. We also demonstrate the ability of this control strategy to ameliorate discrepancies between the predictions of a deterministic model and stochastic switching system.

59 BASIC BIOLOGICAL SCIENCES↗

Gene-informed decomposition model predicts lower soil carbon loss due to persistent microbial adaptation to warming

Abstract Soil microbial respiration is an important source of uncertainty in projecting future climate and carbon (C) cycle feedbacks. However, its feedbacks to climate warming and underlying microbial mechanisms are still poorly understood. Here we show that the temperature sensitivity of soil microbial respiration ( Q 10 ) in a temperate grassland ecosystem persistently decreases by 12.0 ± 3.7% across 7 years of warming. Also, the shifts of microbial communities play critical roles in regulating thermal adaptation of soil respiration. Incorporating microbial functional gene abundance data into a microbially-enabled ecosystem model significantly improves the modeling performance of soil microbial respiration by 5–19%, and reduces model parametric uncertainty by 55–71%. In addition, modeling analyses show that the microbial thermal adaptation can lead to considerably less heterotrophic respiration (11.6 ± 7.5%), and hence less soil C loss. If such microbially mediated dampening effects occur generally across different spatial and temporal scales, the potential positive feedback of soil microbial respiration in response to climate warming may be less than previously predicted.

54 ENVIRONMENTAL SCIENCES↗

Predictive CRISPR-mediated gene downregulation for enhanced production of sustainable aviation fuel precursor in Pseudomonas putida

CRISPR interference (CRISPRi) has emerged as a valuable tool for redirecting metabolic flux to enhance bioproduction. However, its application is often constrained by two challenges: (i) rationally identifying effective gene targets for downregulation and (ii) efficiently constructing multiplexed CRISPRi systems. In this study, we address both challenges by integrating a computational prioritization tool with a versatile assembly method for building multiplexed CRISPRi systems. FluxRETAP (Flux-Reaction Target Prioritization) accurately identified gene targets whose knockdown led to substantial increase of isoprenol titers in Pseudomonas putida KT2440, outperforming a conventional non-computational, pathway-guided target selection. The highest isoprenol titer of nearly 1.5 g/L was achieved by knocking down PP_4118 (a gene encoding α-ketoglutarate dehydrogenase). The use of VAMMPIRE (Versatile Assembly Method for MultiPlexing CRISPRi-mediated downREgulation) enabled accurate assembly of CRISPRi constructs containing up to five sgRNA arrays, reducing context dependency and achieving uniform, position-independent gene downregulation. The integration of FluxRETAP and VAMMPIRE has the potential to advance metabolic engineering by rapidly identifying CRISPRi-mediated knockdowns and knockdown combinations that enhance bioproduction titers, with potential applicability to other microbial systems.

CRISPR interference↗

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES↗

Enhanced Co-Expression Extrapolation (COXEN) Gene Selection Method for Building Anti-Cancer Drug Response Prediction Models

The co-expression extrapolation (COXEN) method has been successfully used in multiple studies to select genes for predicting the response of tumor cells to a specific drug treatment. Here, we enhance the COXEN method to select genes that are predictive of the efficacies of multiple drugs for building general drug response prediction models that are not specific to a particular drug. The enhanced COXEN method first ranks the genes according to their prediction power for each individual drug and then takes a union of top predictive genes of all the drugs, among which the algorithm further selects genes whose co-expression patterns are well preserved between cancer cases for building prediction models. We apply the proposed method on benchmark in vitro drug screening datasets and compare the performance of prediction models built based on the genes selected by the enhanced COXEN method to that of models built on genes selected by the original COXEN method and randomly picked genes. Models built with the enhanced COXEN method always present a statistically significantly improved prediction performance (adjusted p-value ≤ 0.05). Our results demonstrate the enhanced COXEN method can dramatically increase the power of gene expression data for predicting drug response.

60 APPLIED LIFE SCIENCES↗

Engineering of increased L-Threonine production in bacteria by combinatorial cloning and machine learning

The goal of this study is to develop a general strategy for bacterial engineering using an integrated synthetic biology and machine learning (ML) approach. This strategy was developed in the context of increasing L-threonine production in Escherichia coli ATCC 21277. A set of 16 genes was initially selected based on metabolic pathway relevance to threonine biosynthesis and used for combinatorial cloning to construct a set of 385 strains to generate training data (i.e., a range of L-threonine titers linked to each of the specific gene combinations). Hybrid (regression/classification) deep learning (DL) models were developed and used to predict additional gene combinations in subsequent rounds of combinatorial cloning for increased L-threonine production based on the training data. As a result, E. coli strains built after just three rounds of iterative combinatorial cloning and model prediction generated higher L-threonine titers (from 2.7 g/L to 8.4 g/L) than those of patented L-threonine strains being used as controls (4-5 g/L). Interesting combinations of genes in L-threonine production included deletions of the tdh, metL, dapA, and dhaM genes as well as overexpression of the pntAB, ppc, and aspC genes. Mechanistic analysis of the metabolic system constraints for the best performing constructs offers ways to improve the models by adjusting weights for specific gene combinations. Graph theory analysis of pairwise gene modifications and corresponding levels of L-threonine production also suggests additional rules that can be incorporated into future ML models.

60 APPLIED LIFE SCIENCES↗

Chromosome assembled and annotated genome sequence of Aspergillus flavus NRRL 3357

Abstract Aspergillus flavus is an opportunistic pathogen of crops, including peanuts and maize, and is the second leading cause of aspergillosis in immunocompromised patients. A. flavus is also a major producer of the mycotoxin, aflatoxin, a potent carcinogen, which results in significant crop losses annually. The A. flavus isolate NRRL 3357 was originally isolated from peanut and has been used as a model organism for understanding the regulation and production of secondary metabolites, such as aflatoxin. A draft genome of NRRL 3357 was previously constructed, enabling the development of molecular tools and for understanding population biology of this particular species. Here, we describe an updated, near complete, telomere-to-telomere assembly and re-annotation of the eight chromosomes of A. flavus NRRL 3357 genome, accomplished via long-read PacBio and Oxford Nanopore technologies combined with Illumina short-read sequencing. A total of 13,715 protein-coding genes were predicted. Using RNA-seq data, a significant improvement was achieved in predicted 5’ and 3’ untranslated regions, which were incorporated into the new gene models.

59 BASIC BIOLOGICAL SCIENCES↗

Gene expression for biodosimetry and effect prediction purposes: promises, pitfalls and future directions – key session ConRad 2021

In a nuclear or radiological event, an early diagnostic or prognostic tool is needed to distinguish unexposed from low- and highly exposed individuals with the latter requiring early and intensive medical care. Radiation-induced gene expression (GE) changes observed within hours and days after irradiation have shown potential to serve as biomarkers for either dose reconstruction (retrospective dosimetry) or the prediction of consecutively occurring acute or chronic health effects. The advantage of GE markers lies in their capability for early (1–3 days after irradiation), high-throughput, and point-of-care (POC) diagnosis required for the prediction of the acute radiation syndrome (ARS).

61 RADIATION PROTECTION AND DOSIMETRY↗

High School Citizen Scientists Use AI/ML to Predict Intra-Ocular Pressure From Gene Expression Data for Spaceflown Mice

Artificial Intelligence (AI) and Machine Learning (ML) have increasingly become pivotal in biological and biomedical research, largely due to the culture of open data sharing and its associated benefits. The methodologies inherent in AI/ML are particularly adept at identifying and forecasting biological phenotypes from the vast amounts of data generated by next-generation sequencing technologies. These techniques offer substantial promise for advancing research in space biosciences and for the development of automated systems for monitoring space health. Nevertheless, there are crucial aspects to consider when training, validating, and testing machine learning models in both biological research and clinical contexts. It is essential that Open Science principles, including data sharing and the availability of open-source code, are complemented by high-quality, publicly accessible training resources. These resources should focus on best practices and include modules based on real-world scientific cases and data to ensure that future AI/ML practitioners gain practical experience with genuine problems. Addressing this knowledge gap, we have designed, developed, and delivered both interactive and self-paced training programs for citizen scientists worldwide, enabling them to utilize AI/ML for space biology research. This initiative was made possible through generous funding from a Transformation to Open Science Training grant. The interactive training sessions, conducted this summer, utilized AI/ML techniques to analyze data from the Open Science Data Repository, specifically targeting the effects of spaceflight on ocular structure and function. The dataset OSD-583, from the Rodent Research 9 mission, provides experimental data detailing the ocular responses of mice subjected to a 35-day spaceflight, compared with ground control counterparts. Using OSD-583 as observational data, our summer training participants applied AI/ML methods to predict intraocular pressure from RNA-seq data and identify the genes most predictive of the observed responses. Further analysis through pathway enrichment and gene set enrichment revealed that these genes are involved in molecular and cellular processes contributing to retinal degeneration.

James Casaletto↗

Efficient chito–oligosaccharide utilization requires two TonB–dependent transporters and one hexosaminidase in Cellvibrio japonicus

Chitin utilization by microbes plays a significant role in biosphere carbon and nitrogen cycling, and studying the microbial approaches used to degrade chitin will facilitate our understanding of bacterial strategies to degrade a broad range of recalcitrant polysaccharides. The early stages of chitin depolymerization by the bacterium Cellvibrio japonicus have been characterized and are dependent on one chitin-specific lytic polysaccharide monooxygenase and non-redundant glycoside hydrolases from the family GH18 to generate chito-oligosaccharides for entry into metabolism. Here, we describe the mechanisms for the latter stages of chitin utilization by C. japonicus with an emphasis on the fate of chito-oligosaccharides. Here, our systems biology approach combined transcriptomics and bacterial genetics using ecologically relevant substrates to determine the essential mechanisms for chito-oligosaccharide transport and catabolism in Cellvibrio japonicus. Using RNAseq analysis we found a coordinated expression of genes that encode polysaccharide-degrading enzymes. Mutational analysis determined that the hex20B gene product, predicted to encode a hexosaminidase, was required for efficient utilization of chito-oligosaccharides. Furthermore, two gene loci (CJA_0353 and CJA_1157), which encode putative TonB-dependent transporters, were also essential for chito-oligosaccharides utilization. This study further develops our model of C. japonicus chitin metabolism and may be predictive for other environmentally or industrially important bacteria.

59 BASIC BIOLOGICAL SCIENCES↗

Metagenomic Alterations in Gut Microbiota Precede and Predict Onset of Colitis in the IL10 Gene-Deficient Murine Model

SUMMARY. Currently, predictive markers for the development andcourse of inflammatory bowel diseases (IBD) are not available. This study supports the notion that gut microbiome metagenomic profiles could be developed into a useful tool to assess risk and manage human IBD. BACKGROUND & AIMS: Inflammatory bowel diseases (IBD) are chronic inflammatory disorders where predictive bio-markers for the disease development and clinical course are sorely needed for development of prevention and early intervention strategies that can be implemented to improve clinical outcomes. Since gut microbiome alterations can reflect and/or contribute to impending host health changes, we examined whether gut microbiota metagenomic profiles would provide more robust measures for predicting disease outcomes in colitis-prone hosts. METHODS: Using the interleukin (IL) 10 gene-deficient (IL10KO) murine model where early life dysbiosis from antibiotic (cefoperozone [CPZ]) treated dams vertically transferred to pups increases risk for colitis later in life, we investigated temporal metagenomic profiles in the gut microbiota of post-weaning offspring and determined their relationship to even-tual clinical outcomes. RESULTS: Compared to controls, offspring acquiring maternalCPZ-induced dysbiosis exhibited a restructuring of intestinal microbial membership in both bacteriome and mycobiome that was associated with alterations in specific functional subsystems. Furthermore, among IL10 KO offspring fromCPZ-treated dams, several functional subsystems, particularly nitrogen metabolism, diverged between mice that developed spontaneous colitis (CPZ-colitis) versus those that did not (CPZ-no-colitis) at a time point prior to eventual clinical outcome. CONCLUSIONS: Our findings provide support that functional metagenomic profiling of gut microbes has potential and promise meriting further study for development of tools to assess risk and manage human IBD.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic and transcriptomic characterization of carbohydrate-active enzymes in the anaerobic fungus Neocallimastix cameroonii var. constans

Anaerobic gut fungi effectively degrade lignocellulose in the guts of large herbivores, but there remain a limited number of isolated, publicly available, and sequenced strains that impede our understanding of the role of anaerobic fungi within microbial communities. We isolated and characterized a new fungal isolate, Neocallimastix cameroonii var. constans, providing a transcriptomic and genomic understanding of its ability to degrade diverse carbohydrates. This anaerobic fungal strain was stably cultivated for multiple years in vitro among members of an initial enrichment microbial community derived from goat feces, and it demonstrated the ability to pair with other microbial members, namely, archaeal methanogens to produce methane from lignocellulose. Genomic analysis revealed a higher number of predicted carbohydrate-active enzymes encoded in the N. cameroonii var. constans genome compared to most other sequenced anaerobic fungi. The carbohydrate-active enzyme profile for this isolate contained 660 glycoside hydrolases, 160 carbohydrate esterases, 194 glycosyltransferases, and 85 polysaccharide lyases. Differential gene expression analysis showed the upregulation of thousands of genes (including predicted carbohydrate-active enzymes) when N. cameroonii var. constans was grown on lignocellulose (reed canary grass) compared to less complex substrates, such as cellulose (filter paper), cellobiose, and glucose. AlphaFold was used to predict functions of transcriptionally active yet poorly annotated genes, revealing feruloyl esterases that likely play an important role in lignocellulose degradation by anaerobic fungi. The combination of this strain's genomic and transcriptomic characterization, omics-informed structural prediction, and robustness in microbial co-culture make it a well-suited platform to conduct future investigations into bioprocessing and enzyme discovery.

CAZymes↗

Genome Sequence and Analysis of the Flavinogenic Yeast Candida membranifaciens IST 626

The ascomycetous yeast Candida membranifaciens has been isolated from diverse habitats, including humans, insects, and environmental sources, exhibiting a remarkable ability to use different carbon sources that include pentoses, melibiose, and inulin. In this study, we isolated four C. membranifaciens strains from soil and investigated their potential to overproduce riboflavin. C. membranifaciens IST 626 was found to produce the highest concentrations of riboflavin. The volumetric production of this vitamin was higher when C. membranifaciens IST 626 cells were cultured in a commercial medium without iron and when xylose was the available carbon source compared to the same basal medium with glucose. Supplementation of the growth medium with 2 g/L glycine favored the metabolization of xylose, leading to biomass increase and consequent enhancement of riboflavin volumetric production that reached 120 mg/L after 216 h of cultivation. To gain new insights into the molecular basis of riboflavin production and carbon source utilization in this species, the first annotated genome sequence of C. membranifaciens is reported in this article, as well as the result of a comparative genomic analysis with other relevant yeast species. A total of 5619 genes were predicted to be present in C. membranifaciens IST 626 genome sequence (11.5 Mbp). Among them are genes involved in riboflavin biosynthesis, iron homeostasis, and sugar uptake and metabolism. This work put forward C. membranifaciens IST 626 as a riboflavin overproducer and provides valuable molecular data for future development of superior producing strains capable of using the wide range of carbon sources, which is a characteristic trait of the species.

59 BASIC BIOLOGICAL SCIENCES↗

Comparative transcriptomics provides insights into molecular mechanisms of zinc tolerance in the ectomycorrhizal fungus Suillus luteus

Zinc (Zn) is a major soil contaminant and high Zn levels can disrupt growth, survival, and reproduction of fungi. Some fungal species evolved Zn tolerance through cell processes mitigating Zn toxicity, although the genes and detailed mechanisms underlying mycorrhizal fungal Zn tolerance remain unexplored. To fill this gap in knowledge, we investigated the gene expression of Zn tolerance in the ectomycorrhizal fungus Suillus luteus. We found that Zn tolerance in this species is mainly a constitutive trait that can also be environmentally dependent. Zinc tolerance in S. luteus is associated with differences in the expression of genes involved in metal exclusion and immobilization, as well as recognition and mitigation of metal-induced oxidative stress. Differentially expressed genes were predicted to be involved in transmembrane transport, metal chelation, oxidoreductase activity, and signal transduction. Some of these genes were previously reported as candidates for S. luteus Zn tolerance, while others are reported here for the first time. Our results contribute to understanding the mechanisms of fungal metal tolerance and pave the way for further research on the role of fungal metal tolerance in mycorrhizal associations.

59 BASIC BIOLOGICAL SCIENCES↗

Removal of primary nutrient degrading members severely reduces growth of soil microbial communities even when additional degraders are present

Understanding how microorganisms within a soil community interact to support collective respiration and growth remains challenging. Here we used a model substrate, chitin, and a Model Soil Consortium, MSC-2, to investigate how individual members of a microbial community contribute to decomposition and community growth. While MSC-2 can grow using chitin as the sole carbon source, we do not yet know how the growth kinetics or final biomass yields of MSC-2 vary when certain chitin degraders, or other important members, are absent. To characterize specific roles within this representative community, we carried out experiments leaving out members of MSC-2 and measuring biomass yields and CO2 production. We chose two members to iteratively leave out (referred to by genus name): Streptomyces, as it is predicted via gene expression analysis to be a major chitin degrader in the community, and Rhodococcus as it is predicted via species co-abundance analysis to interact with several other members. Our results showed that when MSC-2 lacked Streptomyces, growth and respiration of the community was severely reduced. Removal of either Streptomyces or Rhodococcus led to major changes in abundance for several other species, pointing to a comprehensive shifting of the microbial community when important members are removed as well as alterations in the metabolic profile, especially when Streptomyces was removed. These results show that when keystone, chitin degrading members are removed, other members, even those with the potential to degrade chitin, do not fill the same metabolic niche to promote community growth. In addition, highly connected members may be removed with similar or even increased levels of growth and respiration. Our findings are critical to a better understanding of soil microbiology, specifically in how communities maintain activity when biotic or abiotic factors lead to changes in biodiversity in soil systems.

McClure, Ryan S↗

Removal of primary nutrient degraders reduces growth of soil microbial communities with genomic redundancy

Understanding how microorganisms within a soil community interact to support collective respiration and growth remains challenging. Here, we used a model substrate, chitin, and a synthetic Model Soil Consortium (MSC-2) to investigate how individual members of a microbial community contribute to decomposition and community growth. While MSC-2 can grow using chitin as the sole carbon source, we do not yet know how the growth kinetics or final biomass yields of MSC-2 vary when certain chitin degraders, or other important members, are absent. To characterize specific roles within this synthetic community, we carried out experiments leaving out members of MSC-2 and measuring biomass yields and CO 2 production. We chose two members to iteratively leave out (referred to by genus name): Streptomyces, as it is predicted via gene expression analysis to be a major chitin degrader in the community, and Rhodococcus as it is predicted via species co-abundance analysis to interact with several other members. Our results showed that when MSC-2 lacked Streptomyces, growth and respiration of the community was severely reduced. Removal of either Streptomyces or Rhodococcus led to major changes in abundance for several other species, pointing to a comprehensive shifting of the microbial community when important members are removed, as well as alterations in the metabolic profile, especially when Streptomyces was lacking. These results show that when keystone, chitin degrading members are removed, other members, even those with the potential to degrade chitin, do not fill the same metabolic niche to promote community growth. In addition, highly connected members may be removed with similar or even increased levels of growth and respiration. Our findings are critical to a better understanding of soil microbiology, specifically in how communities maintain activity when biotic or abiotic factors lead to changes in biodiversity in soil systems.

59 BASIC BIOLOGICAL SCIENCES↗

Coupling Metabolic Source Isotopic Pair Labeling and Genome Wide Association for Metabolite and Gene Annotation in Plants (Final Technical Report)

In this project, we applied our labeling pipeline to Arabidopsis and sorghum by feeding tissues with isotopically labeled versions of commercially available amino acids to identify all metabolite features that incorporate the label. In sorghum, we fed five accessions, sampled across the diversity of sorghum, to identify the precursor-of-origin for metabolites that vary between accessions as well as those that may be missing from a single reference genotype. This provided us with precursor-of-origin annotation for thousands of unknown metabolites. We then used GWA to map genes responsible for the synthesis of precursor-of-origin classified metabolites. For sorghum leaf and root ducible metabolites, we performed untargeted metabolomics on leaf and root tissues from 300 diverse genotyped sorghum inbred lines. The amino acid precursor-of-origin metabolite library were then used to identify the corresponding metabolites in the GWA data sets and to identify novel gene-metabolite associations. Finally, we utilized existing and newly generated sequenced EMS mutants of sorghum to validate the predicted gene-metabolite relationships that our labelling analysis identified. In parallel, we conducted similar feeding experiments in Arabidopsis to categorize metabolites based on precursor-of-origin, identify those that vary across our existing Arabidopsis metabolite GWA dataset, and identify genes required for the synthesis of each metabolite. To provide an independent test of gene annotation and pathway involvement, we tested the GWA gene-metabolite associations in Arabidopsis by analyzing the metabolic phenotypes of gene knockouts. Genes of particular interest from both sorghum and Arabidopsis were studied in detail by directly measuring the activity of the corresponding enzymes following heterologous expression. In summary, this work classified as-yet-unknown amino acid-derived metabolites and identified genes involved in their production generated through “omics” technologies. This information was used to validate gene function and identify new metabolism in Arabidopsis and sorghum.

09 BIOMASS FUELS↗