Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Genes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Data for FUN-PROSE: A Deep Learning Approach to Predict Condition-Specific Gene Expression in Fungi

mRNA levels of all genes in a genome is a critical piece of information defining the overall state of the cell in a given environmental condition. Being able to reconstruct such condition-specific expression in fungal genomes is particularly important to metabolically engineer these organisms to produce desired chemicals in industrially scalable conditions. Most previous deep learning approaches focused on predicting the average expression levels of a gene based on its promoter sequence, ignoring its variation across different conditions. Here we present FUN-PROSE—a deep learning model trained to predict differential expression of individual genes across various conditions using their promoter sequences and expression levels of all transcription factors. We train and test our model on three fungal species and get the correlation between predicted and observed condition-specific gene expression as high as 0.85. We then interpret our model to extract promoter sequence motifs responsible for variable expression of individual genes. We also carried out input feature importance analysis to connect individual transcription factors to their gene targets. A sizeable fraction of both sequence motifs and TF-gene interactions learned by our model agree with previously known biological information, while the rest corresponds to either novel biological facts or indirect correlations.

Genomics↗

Population‐level gene expression can repeatedly link genes to functions in maize

SUMMARY Transcriptome‐wide association studies (TWAS) can provide single gene resolution for candidate genes in plants, complementing genome‐wide association studies (GWAS) but efforts in plants have been met with, at best, mixed success. We generated expression data from 693 maize genotypes, measured in a common field experiment, sampled over a 2‐h period to minimize diurnal and environmental effects, using full‐length RNA‐seq to maximize the accurate estimation of transcript abundance. TWAS could identify roughly 10 times as many genes likely to play a role in flowering time regulation as GWAS conducted data from the same experiment. TWAS using mature leaf tissue identified known true‐positive flowering time genes known to act in the shoot apical meristem, and trait data from a new environment enabled the identification of additional flowering time genes without the need for new expression data. eQTL analysis of TWAS‐tagged genes identified at least one additional known maize flowering time gene through trans ‐eQTL interactions. Collectively these results suggest the gene expression resource described here can link genes to functions across different plant phenotypes expressed in a range of tissues and scored in different experiments.

Torres‐Rodríguez, J. Vladimir↗

Optimized CRISPR Interference System for Investigating Pseudomonas alloputida Genes Involved in Rhizosphere Microbiome Assembly

Pseudomonas alloputida KT2440 (formerly P. putida) has become both a well-known chassis organism for synthetic biology and a model organism for rhizosphere colonization. Here, we describe a CRISPR interference (CRISPRi) system in KT2440 for exploring microbe–microbe interactions in the rhizosphere and for use in industrial systems. Our CRISPRi system features three different promoter systems (XylS/P m , LacI/P lac , and AraC/P BAD ) and a dCas9 codon-optimized for Pseudomonads, all located on a mini-Tn7-based transposon that inserts into a neutral site in the genome. It also includes a suite of pSEVA-derived sgRNA expression vectors, where the expression is driven by synthetic promoters varying in strength. We compare the three promoter systems in terms of how well they can precisely modulate gene expression, and we discuss the impact of environmental factors, such as media choice, on the success of CRISPRi. We demonstrate that CRISPRi is functional in bacteria colonizing the rhizosphere, with repression of essential genes leading to a 10–100-fold reduction in P. alloputida cells per root. Finally, we show that CRISPRi can be used to modulate microbe–microbe interactions. When the gene pvdH is repressed and P. alloputida is unable to produce pyoverdine, it loses its ability to inhibit other microbes in vitro. Furthermore, our design is amendable for future CRISPRi-seq studies and in multispecies microbial communities, with the different promoter systems providing a means to control the level of gene expression in many different environments.

Bacteria↗

Data for Promoter Deletion in the Soybean Compact Mutant Leads to Overexpression of a Gene with Homology to the C20-Gibberellin 2-Oxidase Family

Height is a critical component of plant architecture, significantly affecting crop yield. The genetic basis of this trait in soybean remains unclear. In this study, we report the characterization of the Compact mutant of soybean, which has short internodes. The candidate gene was mapped to chromosome 17, and the interval containing the causative mutation was further delineated using biparental mapping. Whole-genome sequencing of the mutant revealed an 8.7 kb deletion in the promoter of the Glyma.17g145200 gene, which encodes a member of the class III gibberellin (GA) 2-oxidases. The mutation has a dominant effect, likely via increased expression of the GA 2-oxidase transcript observed in green tissue, as a result of the deletion in the promoter of Glyma.17g145200. We further demonstrate that levels of GA precursors are altered in the Compact mutant, supporting a role in GA metabolism, and that the mutant phenotype can be rescued with exogenous GA3. We also determined that overexpression of Glyma.17g145200 in Arabidopsis results in dwarfed plants. Thus, gain of promoter activity in the Compact mutant leads to a short internode phenotype in soybean through altered metabolism of gibberellin precursors. These results provide an example of how structural variation can control an important crop trait and a role for Glyma.17g145200 in soybean architecture, with potential implications for increasing crop yield.

Biomass Analytics↗

Mining Thermophile Photosynthesis Genes: A Synthetic Operon Expressing Chloroflexota Species Reaction Center Genes in Rhodobacter sphaeroides

Photosynthesis is the foundation of the vast majority of life systems, and is therefore the most important bioenergetic process on earth. The greatest diversity of photosynthetic systems is found in microorganisms. However, our understanding of the biophysical and biochemical processes that transduce light into chemical energy is derived from a relatively small subset of proteins from microbes that are amenable to cultivation, in contrast to the huge number of predicted proteins that catalyze the initial photochemical reactions deposited in databases, such as from metagenomics. We describe the use of a Rhodobacter sphaeroides laboratory strain for the expression of heterologous photosynthesis genes to demonstrate the feasibility of mining this resource, focusing on hot spring Chloroflexota gene sequences. Using a synthetic operon of genes, we produced a photochemically active complex of reaction center proteins in our biological system. We also present bioinformatic analyses of anoxygenic type II reaction center sequences from metagenomic samples collected from hot (42–90 °C) springs available through the JGI IMG database, to generate a resource of diverse sequences that are potentially adapted to photosynthesis at such temperatures. These data provide a view into the natural diversity of anoxygenic photosynthesis, through a lens focused on high-temperature environments. The approach we took to express such genes can be applied for potential biotechnology purposes as well as for studies of fundamental catalytic properties of these heretofore inaccessible protein complexes.

Chloroflexota↗

Omics-Based Comparison of Fungal Virulence Genes, Biosynthetic Gene Clusters, and Small Molecules in Penicillium expansum and Penicillium chrysogenum

Penicillium expansum is a ubiquitous pathogenic fungus that causes blue mold decay of apple fruit postharvest, and another member of the genus, Penicillium chrysogenum, is a well-studied saprophyte valued for antibiotic and small molecule production. While these two fungi have been investigated individually, a recent discovery revealed that P. chrysogenum can block P. expansum-mediated decay of apple fruit. To shed light on this observation, we conducted a comparative genomic, transcriptomic, and metabolomic study of two P. chrysogenum (404 and 413) and two P. expansum (Pe21 and R19) isolates. Global transcriptional and metabolomic outputs were disparate between the species, nearly identical for P. chrysogenum isolates, and different between P. expansum isolates. Further, the two P. chrysogenum genomes revealed secondary metabolite gene clusters that varied widely from P. expansum. This included the absence of an intact patulin gene cluster in P. chrysogenum, which corroborates the metabolomic data regarding its inability to produce patulin. Additionally, a core subset of P. expansum virulence gene homologues were identified in P. chrysogenum and were similarly transcriptionally regulated in vitro. Molecules with varying biological activities, and phytohormone-like compounds were detected for the first time in P. expansum while antibiotics like penicillin G and other biologically active molecules were discovered in P. chrysogenum culture supernatants. Our findings provide a solid omics-based foundation of small molecule production in these two fungal species with implications in postharvest context and expand the current knowledge of the Penicillium-derived chemical repertoire for broader fundamental and practical applications.

Bartholomew, Holly P. (ORCID:0000000292726399)↗

Gene network centrality analysis identifies key regulators coordinating day-night metabolic transitions in Synechococcus elongatus PCC 7942 despite limited accuracy in predicting direct regulator-gene interactions

Synechococcus elongatus PCC 7942 is a model organism for studying circadian regulation and bioproduction, where precise temporal control of metabolism significantly impacts photosynthetic efficiency and CO 2 -to-bioproduct conversion. Despite extensive research on core clock components, our understanding of the broader regulatory network orchestrating genome-wide metabolic transitions remains incomplete. We address this gap by applying machine learning tools and network analysis to investigate the transcriptional architecture governing circadian-controlled gene expression. While our approach showed moderate accuracy in predicting individual transcription factor-gene interactions - a common challenge with real expression data - network-level topological analysis successfully revealed the organizational principles of circadian regulation. Our analysis identified distinct regulatory modules coordinating day-night metabolic transitions, with photosynthesis and carbon/nitrogen metabolism controlled by day-phase regulators, while nighttime modules orchestrate glycogen mobilization and redox metabolism. Through network centrality analysis, we identified potentially significant but previously understudied transcriptional regulators: HimA as a putative DNA architecture regulator, and TetR and SrrB as potential coordinators of nighttime metabolism, working alongside established global regulators RpaA and RpaB. This work demonstrates how network-level analysis can extract biologically meaningful insights despite limitations in predicting direct regulatory interactions. The regulatory principles uncovered here advance our understanding of how cyanobacteria coordinate complex metabolic transitions and may inform metabolic engineering strategies for enhanced photosynthetic bioproduction from CO 2 .

59 BASIC BIOLOGICAL SCIENCES↗

Data for Expression of a Bacterial Trehalose 6-Phosphate Synthase Gene otsA in Camelina sativa Seeds Promotes the Channelling of Carbon Towards Oil Accumulation

Improving seed oil yield is essential for developing Camelina sativa as a sustainable biofuel crop. Fatty acid synthesis depends on the production of acetyl-CoA from photosynthetically derived sugars. Trehalose 6-phosphate (T6P), a proxy for sucrose availability, can link sugar status to plant growth and development. Synthesised by trehalose 6-phosphate synthase (TPS) from UDP-glucose and glucose-6-phosphate, T6P plays a regulatory role in metabolism. Our previous studies on Arabidopsis transgenic lines constitutively expressing the E. coli otsA (encoding TPS) showed increased T6P levels and seed triacylglycerol, along with stunted growth. In the present study we express otsA in camelina under the control of a seed-specific Phaseolin promoter. Seeds of the resulting transgenic lines accumulated high levels of T6P, and a 15%–20% increase in total fatty acids and triacylglycerol compared to wild-type. Molecular analysis showed the transgenic seeds had reduced SnRK1 activity, elevated WRI1 protein levels, and increased the levels of WRI1 and its target genes, along with enhanced rates of fatty acid synthesis that increased seed weights relative to wild type. Notably, the increase in oil did not affect seed protein levels but did reduce the soluble metabolite fraction. Crucially, seed-specific expression of otsA mitigated the growth defects associated with constitutive otsA expression, and the transgenic lines showed normal seed development and germination. These findings demonstrate that targeted T6P modulation via seed-specific otsA expression is an effective metabolic engineering strategy to boost oil production in camelina and potentially in other oilseed crops and bioenergy crops such as energycane, sorghum and miscanthus.

Lipids↗

Stage-resolved gene regulatory network analysis reveals developmental reprogramming and genes with robust stem-preferred expression in sorghum

Sorghum bicolor is a deep-rooted, heat- and drought-tolerant crop that thrives on marginal lands and is increasingly valued for its applications in biofuel, bioenergy, and biopolymer production. The sorghum stem, which can reach 4–5 m in length, serves as the primary reservoir of both lignocellulosic biomass and soluble sugars, making it a promising bioenergy feedstock. Although recent advances in genetic, genomic, and transcriptomic resources have improved our understanding of sorghum biology, comprehensive genome-wide analyses of functional dynamics across diverse organ types and developmental stages remain limited. In particular, candidate genes with stem preferred expression pattern or their associated cis-regulatory elements, which may program key stem-related functions and enable organ- or tissue-specific engineering, have not yet been identified.

59 BASIC BIOLOGICAL SCIENCES↗

Data for An Orphan Gene BOOSTER Enhances Photosynthetic Efficiency and Plant Productivity

Seeds of Col-0 wild type, sig6 T-DNA mutants (CS877785, ABRC), PRL-1-OE, and sig6 T-DNA mutants transfected with PRL-1 (sig6::PRL-1) were planted in 1/2 MS media. Seedlings growth including chlorophyll development defects were investigated across the genotypes. Four-days-old-post-light exposure seedlings were harvested and performed RNAseq analysis with four biological replicates.

Biomass Analytics↗

Synthetic overlapping genes stabilize genetic systems

Overlapping genes—wherein two different proteins are translated from alternative reading frames of the same DNA sequence—provide a means to stabilize an engineered gene by directly linking its evolutionary fate with that of an overlapping gene. However, creating overlapping gene pairs is challenging, as it requires redesigning both protein products to accommodate overlap constraints. Here, we present a new “overlapping, alternate-frame insertion” (OAFI) method for creating synthetic overlapping genes by inserting an “inner” gene, encoded in an alternate frame, into a flexible region of an “outer” gene. Using OAFI, we create new overlapping gene pairs of genetic reporters and bacterial toxins within an antibiotic resistance gene. We show that both the inner and outer genes retain function despite redesign, with translation of the inner gene influenced by its overlap position in the outer gene. Importantly, we show that, despite these inner gene sequences not contributing to outer gene function, selection for the outer gene alters the permitted inactivating mutations in the inner gene, and that overlapping toxins can restrict horizontal gene transfer of the antibiotic resistance gene. Overall, OAFI offers a versatile tool for synthetic biology, expanding the applications of overlapping genes in gene stabilization and biocontainment.

Biological and medical sciences↗

Comparison of gene expression in the skin tissue of gray, humpback, and fin whales

Analyses of gene expression in the skin of several species of whales identified genes that are differentially expressed in association with environmental factors, suggesting that skin transcriptomics may provide a valuable tool for assessing physiological responses in marine mammals. Previous work exploring differing levels of gene expression has focused on odontocetes with comparatively limited investigation of skin gene expression has been explored in mysticetes. Here, we describe the identity of genes expressed in skin tissue of three species of baleen whales to establish a baseline of gene expression and compare gene identity and expression patterns across species. We also evaluate sex-specific differences in skin gene expression through a comparison of expression levels between males and females in gray and humpback whales. A total of 16 skin tissue samples were collected from free-ranging gray, humpback and fin whales off the central Oregon coast in the eastern North Pacific. Comparison of the expressed genes in the humpback and gray whale skin tissue to the blue whale reference database identified enriched gene ontology terms in the skin tissue of each species, suggesting genes over-represented in the whale skin related to cell epithelial development, regulation of gene expression and cell maintenance . Comparison of gene expression between male and female samples revealed sex-specific differences in gray and humpback whales. A differential gene expression analysis identified several x-linked genes that have been previously identified and show gene expression differences in male and female cetaceans, such as ZFX, DDX3X and USP9X. Establishing baseline skin gene expression profiles for these three baleen whale species sampled off the Oregon coast provides a foundation for linking transcriptome variation with physiological condition and environment.

Sremba, Angela↗

Small Signaling Peptides in Sorghum bicolor : Integrating Phylogeny and Gene Expression to Characterize Roles in Stem Development

Small signaling peptides (SSPs) are important regulators of plant growth, development, and responses to biotic and abiotic stress, yet their role in the C4 grass Sorghum bicolor is largely uncharacterized. To help fill this knowledge gap, 219 sorghum genes that encode SSPs were identified based on SSP sequences previously identified in Arabidopsis thaliana, Zea mays, Oryza sativa, Triticum aestivum , and Brachypodium distachyon . The 219 sorghum SSP-encoding genes were assigned to 19 gene families, analyzed for the presence of motifs, and aligned with genes that encode SSPs in other plants using phylogenetic analysis. Sorghum genes in 12 of the 19 SSP gene families had not been previously characterized. Expression of the 219 SSP-encoding genes in sorghum organs, during stem development, and in stem tissues and cell types revealed distinct spatial, temporal, and developmental patterns of expression. Genes associated with the SbCEP and SbRGF families were preferentially expressed in roots, whereas SbEPF genes were expressed in stem epidermal and pith parenchyma cells and panicles. The expression of genes during bioenergy sorghum stem growth and development was investigated because stems account for ~80% of harvested biomass and serve as conduits for water and nutrient transport between leaves and roots. During stem development, 28 SSP genes in several families ( CLE, EPF, CEP, GASS, PSY, ES, PSK, CAPE, POE ) were expressed at higher levels in zones of cell proliferation. For example, the TDIF homologs SbCLE41 and SbCLE42 were expressed at high levels in nascent stem nodes where they may regulate vascular bundle cambial activity and cell differentiation. A different set of 15 genes in the CIF, POE, CAPE, PSY, CEP, RALF , and CLE families were expressed at higher levels in zones of stem tissue differentiation highlighted by elevated expression of five SbRALFR s in the stem nodal plexus. Cell type–specific expression of many sorghum genes that encode SSPs was observed in fully elongated internodes indicating gene expression is regulated with high spatial resolution. Overall, the results provide a foundation of information for analysis of SSP function in sorghum that can be integrated with knowledge of sorghum gene regulatory networks to modulate traits important for production of sorghum crops.

bioenergy sorghum↗

A compendium of human gene functions derived from evolutionary modelling

A comprehensive, computable representation of the functional repertoire of all macromolecules encoded within the human genome is a foundational resource for biology and biomedical research. The Gene Ontology Consortium has been working towards this goal by generating a structured body of information about gene functions, which now includes experimental findings reported in more than 175,000 publications for human genes and genes in experimentally tractable model organisms 1,2 . Here, we describe the results of a large, international effort to integrate all of these findings to create a representation of human gene functions that is as complete and accurate as possible. Specifically, we apply an expert-curated, explicit evolutionary modelling approach to all human protein-coding genes. This approach integrates available experimental information across families of related genes into models that reconstruct the gain and loss of functional characteristics over evolutionary time. The models and the resulting set of 68,667 integrated gene functions cover approximately 82% of human protein-coding genes. The functional repertoire reveals a marked preponderance of molecular regulatory functions, and the models provide insights into the evolutionary origins of human gene functions. We show that our set of descriptions of functions can improve the widely used genomic technique of Gene Ontology enrichment analysis. The experimental evidence for each functional characteristic is recorded, thereby enabling the scientific community to help review and improve the resource, which we have made publicly available.

59 BASIC BIOLOGICAL SCIENCES↗

The relationship between gene traits and transcription in soil microbial communities varies by environmental stimulus

Codon and nucleotide frequencies are known to relate to the rate of gene transcription, yet how these traits shape transcriptional profiles of soil microbial communities remains unclear. Here we test the prediction that functional genes with high codon optimization and energetically lower cost nucleotides (i.e., nucleotides requiring less adenosine triphosphate (ATP) for synthesis) have higher transcriptional expression in a soil microbial community. In laboratory incubations, we subjected an agricultural soil to two separate short-term environmental changes: labile carbon (glucose) addition or a sudden 30-min increase in temperature from 20 °C to 60 °C. Using the total genomic codon frequencies to predict preferred codon usage for each taxon, we then estimated codon optimization for each transcript. On the community level, we found a higher average level of codon optimization after the addition of glucose. Synonymous nucleotide composition in the transcript pool also shifted towards energetically cheaper nucleotides, favoring uracil (U) over adenine (A) and cytosine (C) over guanine (G). Similarly, we found that encoded amino acid usage shifted towards energetically cheaper amino acids in response to labile carbon. In contrast, in communities responding to heat shock, there were no significant differences in the averaged gene traits of expressed transcripts. We used metagenome-assembled-genomes to further examine the ability of gene traits to predict transcriptional responses within and between taxa. We found that traits of individual genes could not reliably predict the level of transcription of a gene within or between taxa—highlighting the limits of this approach. However, we did find that when traits were averaged across several related genes, codon optimization was able to predict levels of transcription in metabolic pathways associated with growth and nutrient uptake in response to glucose. Similar relationships were not observed in response to heat, or for functions associated with stress—such as genes associated with sporulation or heat shock. These results demonstrate that gene traits, such as codon usage, nucleotide selection, and amino acid selection, relate to the transcriptional expression of genes in soil microbial communities and suggests that these relationships may be dependent on both gene function and the specific type of environmental stimuli.

Biological and medical sciences↗

Transcriptome-wide association analysis identifies candidate susceptibility genes for prostate-specific antigen levels in men without prostate cancer

Deciphering the genetic basis of prostate-specific antigen (PSA) levels may improve their utility for prostate cancer (PCa) screening. Using genome-wide association study (GWAS) summary statistics from 95,768 PCa-free men, we conducted a transcriptome-wide association study (TWAS) to examine impacts of genetically predicted gene expression on PSA. Analyses identified 41 statistically significant (p < 0.05/12,192 = 4.10 × 10 –6 ) associations in whole blood and 39 statistically significant (p < 0.05/13,844 = 3.61 × 10 –6 ) associations in prostate tissue, with 18 genes associated in both tissues. Cross-tissue analyses identified 155 statistically significantly (p < 0.05/22,249 = 2.25 × 10 –6 ) genes. Out of 173 unique PSA-associated genes across analyses, we replicated 151 (87.3%) in a TWAS of 209,318 PCa-free individuals from the Million Veteran Program. Based on conditional analyses, we found 20 genes (11 single tissue, nine cross-tissue) that were associated with PSA levels in the discovery TWAS that were not attributable to a lead variant from a GWAS. Ten of these 20 genes replicated, and two of the replicated genes had colocalization probability of >0.5: CCNA2 and HIST1H2BN. Six of the 20 identified genes are not known to impact PCa risk. Fine-mapping based on whole blood and prostate tissue revealed five protein-coding genes with evidence of causal relationships with PSA levels. Of these five genes, four exhibited evidence of colocalization and one was conditionally independent of previous GWAS findings. These results yield hypotheses that should be further explored to improve understanding of genetic factors underlying PSA levels.

60 APPLIED LIFE SCIENCES↗

Unique trajectory of gene family evolution from genomic analysis of nearly all known species in an ancient yeast lineage

Gene gains and losses are a major driver of genome evolution; their precise characterization can provide insights into the origin and diversification of major lineages. Here, we examined gene family evolution of 1154 genomes from nearly all known species in the medically and technologically important yeast subphylum Saccharomycotina. We found that yeast gene family evolution differs from that of plants, animals, and filamentous ascomycetes, and is characterized by smaller overall gene numbers yet larger gene family sizes for a given gene number. Faster-evolving lineages (FELs) in yeasts experienced significantly higher rates of gene losses—commensurate with a narrowing of metabolic niche breadth—but higher speciation rates than their slower-evolving sister lineages (SELs). Gene families most often lost are those involved in mRNA splicing, carbohydrate metabolism, and cell division and are likely associated with intron loss, metabolic breadth, and non-canonical cell cycle processes. Our results highlight the significant role of gene family contractions in the evolution of yeast metabolism, genome function, and speciation, and suggest that gene family evolutionary trajectories have differed markedly across major eukaryotic lineages.

Comparative Genomics↗