Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Genome”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A metagenomic perspective on the microbial prokaryotic genome census

Following 30 years of sequencing, we assessed the phylogenetic diversity (PD) of >1.5 million microbial genomes in public databases, including metagenome-assembled genomes (MAGs) of uncultivated microbes. As compared to the vast diversity uncovered by metagenomic sequences, cultivated taxa account for a modest portion of the overall diversity, 9.73% in bacteria and 6.55% in archaea, while MAGs contribute 48.54% and 57.05%, respectively. Therefore, a substantial fraction of bacterial (41.73%) and archaeal PD (36.39%) still lacks any genomic representation. This unrepresented diversity manifests primarily at lower taxonomic ranks, exemplified by 134,966 species identified in 18,087 metagenomic samples. Our study exposes diversity hotspots in freshwater, marine subsurface, sediment, soil, and other environments, whereas human samples yielded minimal novelty within the context of existing datasets. These results offer a roadmap for future genome recovery efforts, delineating uncaptured taxa in underexplored environments and underscoring the necessity for renewed isolation and sequencing.

59 BASIC BIOLOGICAL SCIENCES↗

Discovery of additional ancient genome duplications in yeasts

Whole-genome duplication (WGD) has had profound macroevolutionary impacts on diverse lineages, preceding adaptive radiations in vertebrates, teleost fish, and angiosperms. In contrast to the many known ancient WGDs in animals, and especially plants, we are aware of evidence for only four WGDs in fungi. The oldest of these occurred ∼100 million years ago (mya) and is shared by ∼60 extant Saccharomycetales species, including the baker’s yeast Saccharomyces cerevisiae. Notably, this is the only known ancient WGD event in the yeast subphylum Saccharomycotina. The dearth of ancient WGD events in fungi remains a mystery. Some studies have suggested that fungal lineages that experience chromosome and genome duplication quickly go extinct, leaving no trace in the genomic record, while others contend that the lack of known WGDs is due to an absence of data. Under the second hypothesis, additional sampling and deeper sequencing of fungal genomes should lead to the discovery of more WGD events. Coupling hundreds of recently published genomes from nearly every described Saccharomycotina species, with three additional long-read assemblies, we discovered three novel WGD events. Although the functions of retained duplicate genes originating from these events are broad, they bear similarities to the well-known WGD that occurred in the Saccharomycetales. In conclusion, our results suggest that WGD may be a more common evolutionary force in fungi than previously believed.

convergent evolution↗

Comparative genomic analysis of thermophilic fungi reveals convergent evolutionary adaptations and gene losses

Thermophily is a trait scattered across the fungal tree of life, with its highest prevalence within three fungal families (Chaetomiaceae, Thermoascaceae, and Trichocomaceae), as well as some members of the phylum Mucoromycota. We examined 37 thermophilic and thermotolerant species and 42 mesophilic species for this study and identified thermophily as the ancestral state of all three prominent families of thermophilic fungi. Thermophilic fungal genomes were found to encode various thermostable enzymes, including carbohydrate-active enzymes such as endoxylanases, which are useful for many industrial applications. At the same time, the overall gene counts, especially in gene families responsible for microbial defense such as secondary metabolism, are reduced in thermophiles compared to mesophiles. We also found a reduction in the core genome size of thermophiles in both the Chaetomiaceae family and the Eurotiomycetes class. The Gene Ontology terms lost in thermophilic fungi include primary metabolism, transporters, UV response, and O-methyltransferases. Comparative genomics analysis also revealed higher GC content in the third base of codons (GC3) and a lower effective number of codons in fungal thermophiles than in both thermotolerant and mesophilic fungi. Furthermore, using the Support Vector Machine classifier, we identified several Pfam domains capable of discriminating between genomes of thermophiles and mesophiles with 94% accuracy. Using AlphaFold2 to predict protein structures of endoxylanases (GH10), we built a similarity network based on the structures. We found that the number of disulfide bonds appears important for protein structure, and the network clusters based on protein structures correlate with the optimal activity temperature. Thus, comparative genomics offers new insights into the biology, adaptation, and evolutionary history of thermophilic fungi while providing a parts list for bioengineering applications.

59 BASIC BIOLOGICAL SCIENCES↗

Mechanism-guided engineering of a minimal biological particle for genome editing

The widespread application of genome editing to treat and cure disease requires the delivery of genome editors into the nucleus of target cells. Enveloped delivery vehicles (EDVs) are engineered virally derived particles capable of packaging and delivering CRISPR-Cas9 ribonucleoproteins (RNPs). However, the presence of lentiviral genome encapsulation and replication proteins in EDVs has obscured the underlying delivery mechanism and precluded particle optimization. Here, we show that Cas9 RNP nuclear delivery is independent of the native lentiviral capsid structure. Instead, EDV-mediated genome editing activity corresponds directly to the number of nuclear localization sequences on the Cas9 enzyme. EDV structural analysis using cryo-electron tomography and small molecule inhibitors guided the removal of ~80% of viral residues, creating a minimal EDV (miniEDV) that retains full RNP delivery capability. MiniEDVs are 25% smaller yet package equivalent amounts of Cas9 RNPs relative to the original EDVs and demonstrated increased editing in cell lines and therapeutically relevant primary human T cells. These results show that virally derived particles can be streamlined to create efficacious genome editing delivery vehicles with simpler production and manufacturing.

59 BASIC BIOLOGICAL SCIENCES↗

A single genomic region controls primocane fruiting in tetraploid blackberry

The fresh-market blackberry ( Rubus subgenus Rubus ) industry has expanded dramatically in the past 2 decades, driven in part by improved cultivars. Introgression of the primocane-fruiting (PF; annual flowering) trait into elite germplasm has enabled dual cropping in a single year, season extension, and cultivation in tropical and subtropical regions. Despite its economic performance, the genetic basis of PF is not well understood. It has been proposed that the PF trait is controlled by a major recessive locus, but its genomic location is unclear. Here, a genome-wide association study (GWAS) of 365 tetraploid blackberry genotypes identified a single genomic region on chromosome Ra03 (∼33 Mb) strongly associated with PF. Genetic linkage analysis in a biparental population confirmed that the same interval (32–35 Mb) was linked to the PF phenotype. Ten putative candidate genes were identified in this region. Allele mining using whole-genome resequencing of 17 genotypes highlighted 2 high-priority candidates: a CCCH-type zinc finger gene and an ubiquitin-specific protease gene. Use of an improved Rubus argutus “Hillquist” genome annotation (v1.2) enabled refined variant interpretation, including identification of regulatory 3′ UTR polymorphisms in the zinc finger homolog. Two diagnostic KASP markers (PF1 and PF2), designed from the most significant GWAS SNPs, predicted the PF phenotype with over 96% accuracy in a validation panel of 494 tetraploid blackberries from multiple breeding programs. Together, these results provide the first high-resolution mapping of the PF locus in blackberry, identify candidate genes for flowering regulation in Rubus , and deliver diagnostic markers that can be immediately deployed in breeding programs.

GWAS↗

Optimizing genomic prediction for complex traits via investigating multiple factors in switchgrass

Genomic prediction has accelerated breeding processes and provided mechanistic insights into the genetic bases of complex traits. To further optimize genomic prediction, we assess the impact of genome assemblies, genotyping approaches, variant types, allelic complexities, polyploidy levels, and population structures on the prediction of 20 complex traits in switchgrass (Panicum virgatum L.), a perennial biofuel feedstock. Surprisingly, short read-based genome assembly performs comparably to or even better than long read-based assembly. Due to higher gene coverage, exome capture and multi-allelic variants outperform genotyping-by-sequencing and bi-allelic variants, respectively. Tetraploid models show higher prediction accuracy than octoploid models for most traits, likely due to the greater genetic distances among tetraploids. Depending on the trait in question, different types of variants need to be integrated for optimal predictions. Furthermore, our study provides insights into the factors influencing genomic prediction outcomes, guiding best practices for future studies and for improving agronomic traits in switchgrass and other species through selective breeding.

60 APPLIED LIFE SCIENCES↗

BSMV-mediated genome editing exhibits host-specific heritability: germline transmission in barley and somatic edits in Nicotiana benthamiana

Plant RNA virus–mediated guide RNA (gRNA) delivery represents a transformative advance in genome editing technologies. Unlike conventional transformation methods that rely on labor-intensive tissue culture and regeneration for each individual gRNA delivery, viral vectors can rapidly and systemically transmit gRNAs into pre-established Cas-expressing plants, providing an accelerated route for functional genomics and trait discovery directly in planta . However, key design parameters, including subgenomic promoter choice, transcript architecture, and their effects on viral fitness and editing outcomes, remain to be elucidated for most viral platforms. We developed five Barley stripe mosaic virus (BSMV) vectors, each with distinct subgenomic promoter elements to drive single gRNA expression. These were initially evaluated in Cas9-expressing transgenic Nicotiana benthamiana plants targeting the Phytoene desaturase ( PDS ) gene to compare their editing efficiencies. Single gRNAs expressed under the duplicated γb subgenomic promoter or when fused directly to the γb genome achieved the highest mutation frequencies (up to 90% at 60 days post-inoculation), whereas β1- and β2-driven sgRNAs produced delayed and reduced editing. Thus, promoter selection critically determines gRNA accumulation and the efficacy of BSMV-mediated genome editing. The top-performing design was then applied to Cas9-expressing barley ( Hordeum vulgare ) targeting HvCMF7 (conferring green-white variegation) and HvGW2.1 (impacts grain width and weight). BSMV spread systemically throughout barley, inducing somatic and heritable mutations at frequencies up to 100%, with virus-free edited progeny. In contrast, despite robust somatic editing in N. benthamiana, no heritable mutations were detected indicating species-dependent limitations in germline transmission. Our systematic comparison of subgenomic promoter architectures establishes clear design principles for optimizing viral vector–mediated delivery. Promoter choice and transcript structure critically shape editing efficiency and viral stability. The host-specific boundary for germline editing, defined by efficient heritable editing in barley but not N. benthamiana , highlights where BSMV offers advantages and where alternative vectors or hybrid strategies are required, guiding rational platform selection for diverse crop species and applications. Collectively, these findings establish BSMV as a promising next-generation vector for rapid, tissue culture–free, and transformation-independent genome editing in cereals and other recalcitrant monocots.

barley↗

The Near-Gapless Penicillium fuscoglaucum Genome Enables the Discovery of Lifestyle Features as an Emerging Post-Harvest Phytopathogen

Penicillium spp. occupy many diverse biological niches that include plant pathogens, opportunistic human pathogens, saprophytes, indoor air contaminants, and those selected specifically for industrial applications to produce secondary metabolites and lifesaving antibiotics. Recent phylogenetic studies have established Penicillium fuscoglaucum as a synonym for Penicillium commune, which is an indoor air contaminant and toxin producer and can infect apple fruit during storage. During routine culturing on selective media in the lab, we obtained an isolate of P. fuscoglaucum Pf_T2 and sequenced its genome. The Pf_T2 genome is far superior to available genomic resources for the species. Our assembly exhibits a length of 35.1 Mb, a BUSCO score of 97.9% complete, and consists of five scaffolds/contigs representing the four expected chromosomes. It was determined that the Pf_T2 genome was colinear with a type specimen P. fuscoglaucum and contained a lineage-specific, intact cyclopiazonic acid (CPA) gene cluster. For comparison, a highly virulent postharvest apple pathogen, P. expansum strain TDL 12.1, was included and showed a similar growth pattern in culture to our Pf_T2 isolate but was far more aggressive in apple fruit than P. fuscoglaucum. The genome of Pf_T2 serves as a major improvement over existing resources, has superior annotation, and can inform forthcoming omics-based work and functional genetic studies to probe secondary metabolite production and disparities in aggressiveness during apple fruit decay.

59 BASIC BIOLOGICAL SCIENCES↗

Adaptive gene loss in the common bean pan-genome during range expansion and domestication

The common bean ( Phaseolus vulgaris L.) is a crucial legume crop and an ideal evolutionary model to study adaptive diversity in wild and domesticated populations. Here, we present a common bean pan-genome based on five high-quality genomes and whole-genome reads representing 339 genotypes. It reveals ~234 Mb of additional sequences containing 6,905 protein-coding genes missing from the reference, constituting 49% of all presence/absence variants (PAVs). More non-synonymous mutations are found in PAVs than core genes, probably reflecting the lower effective population size of PAVs and fitness advantages due to the purging effect of gene loss. Our results suggest pan-genome shrinkage occurred during wild range expansion. Selection signatures provide evidence that partial or complete gene loss was a key adaptive genetic change in common bean populations with major implications for plant adaptation. The pan-genome is a valuable resource for food legume research and breeding for climate change mitigation and sustainable agriculture.

59 BASIC BIOLOGICAL SCIENCES↗

Whole-genome demography of COVID-19 virus during its pandemic period and on “panvalent” vaccine design

With over 16 million submitted genomic sequences, the SARS-CoV-2 (SC2) virus, the cause of the most recent worldwide COVID-19 pandemic, has become the most sequenced genome of all known viruses, revealing, for example, a vast number of expanding viral lineages. Since the pandemic phase appears to be over, we performed a retrospective re-examination of the demographic grouping pattern and their genomic characteristics during the entire pandemic period up to the peak of the last pandemic wave. For our study, we extracted from the NCBI only unique viral sequences and converted each sequence data to a relational vector, indicating the presence/absence of each variational event compared to a “reference” sequence. Our study revealed several genomic features that are unexpected or different from those of previous studies. For example, approximately 44,000 variants with unique sequences emerged during the pandemic period; they group into only four major viral-genomic groups and each has a set of mostly unique highly-conserved variant-genotypes (HCVGs); and a small set from the first (“ancestral”) group was inherited by the three (“descendant”) groups, suggesting that HCVGs in the next group may be predictable from the current group(s). Such a concept may be potentially important in designing “panvalent” vaccines against the current and future waves of viral infections.

60 APPLIED LIFE SCIENCES↗

Directed evolution expands CRISPR–Cas12a genome-editing capacity

CRISPR-Cas12a enzymes are versatile RNA-guided genome-editing tools with applications encompassing viral diagnosis, agriculture, and human therapeutics. However, their dependence on a 5'-TTTV-3' protospacer adjacent motif (PAM) next to DNA target sequences restricts Cas12a's gene targeting capability to only ∼1% of a typical genome. To mitigate this constraint, we used a bacterial-based directed evolution assay combined with rational engineering to identify variants of Lachnospiraceae bacterium Cas12a with expanded PAM recognition. The resulting Cas12a variants use a range of noncanonical PAMs while retaining recognition of the canonical 5'-TTTV-3' PAM. In particular, biochemical and cell-based assays show that the variant Flex-Cas12a utilizes 5'-NYHV-3' PAMs that expand DNA recognition sites to ∼25% of the human genome. With enhanced targeting versatility, Flex-Cas12a unlocks access to previously inaccessible genomic loci, providing new opportunities for both therapeutic and agricultural genome engineering.

Ma, Enbo↗

Leptothrix ochracea genomes reveal potential for mixotrophic growth on Fe(II) and organic carbon

ABSTRACT Leptothrix ochracea creates distinctive iron-mineralized mats that carpet streams and wetlands. Easily recognized by its iron-mineralized sheaths, L. ochracea was one of the first microorganisms described in the 1800s. Yet it has never been isolated and does not have a complete genome sequence available, so key questions about its physiology remain unresolved. It is debated whether iron oxidation can be used for energy or growth and if L. ochracea is an autotroph, heterotroph, or mixotroph. To address these issues, we sampled L. ochracea -rich mats from three of its typical environments (a stream, wetlands, and a drainage channel) and reconstructed nine high-quality genomes of L. ochracea from metagenomes. These genomes contain iron oxidase genes cyc2 and mtoA, showing that L. ochracea has the potential to conserve energy from iron oxidation. Sox genes confer potential to oxidize sulfur for energy. There are genes for both carbon fixation (RuBisCO) and utilization of sugars and organic acids (acetate, lactate, and formate). In silico stoichiometric metabolic models further demonstrated the potential for growth using sugars and organic acids. Metatranscriptomes showed a high expression of genes for iron oxidation; aerobic respiration; and utilization of lactate, acetate, and sugars, as well as RuBisCO, supporting mixotrophic growth in the environment. In summary, our results suggest that L. ochracea has substantial metabolic flexibility. It is adapted to iron-rich, organic carbon-containing wetland niches, where it can thrive as a mixotrophic iron oxidizer by utilizing both iron oxidation and organics for energy generation and both inorganic and organic carbon for cell and sheath production. IMPORTANCE Winogradsky's observations of L. ochracea led him to propose autotrophic iron oxidation as a new microbial metabolism, following his work on autotrophic sulfur-oxidizers. While much culture-based research has ensued, isolation proved elusive, so most work on L. ochracea has been based in the environment and in microcosms. Meanwhile, the autotrophic Gallionella became the model for freshwater microbial iron oxidation, while heterotrophic and mixotrophic iron oxidation is not well-studied. Ecological studies have shown that Leptothrix overtakes Gallionella when dissolved organic carbon content increases, demonstrating distinct niches. This study presents the first near-complete genomes of L. ochracea , which share some features with autotrophic iron oxidizers, while also incorporating heterotrophic metabolisms. These genome, metabolic modeling, and transcriptome results give us a detailed metabolic picture of how the organism may combine lithoautotrophy with organoheterotrophy to promote Fe oxidation and C cycling and drive many biogeochemical processes resulting from microbial growth and iron oxyhydroxide formation in wetlands.

59 BASIC BIOLOGICAL SCIENCES↗

CRISPRi functional genomics in bacteria and its application to medical and industrial research

SUMMARY Functional genomics is the use of systematic gene perturbation approaches to determine the contributions of genes under conditions of interest. Although functional genomic strategies have been used in bacteria for decades, recent studies have taken advantage of CRISPR (clustered regularly interspaced short palindromic repeats) technologies, such as CRISPRi (CRISPR interference), that are capable of precisely modulating expression of all genes in the genome. Here, we discuss and review the use of CRISPRi and related technologies for bacterial functional genomics. We discuss the strengths and weaknesses of CRISPRi as well as design considerations for CRISPRi genetic screens. We also review examples of how CRISPRi screens have defined relevant genetic targets for medical and industrial applications. Finally, we outline a few of the many possible directions that could be pursued using CRISPR-based functional genomics in bacteria. Our view is that the most exciting screens and discoveries are yet to come.

Microbiology↗

Expanded genome and proteome reallocation in a novel, robust Bacillus coagulans strain capable of utilizing pentose and hexose sugars

Bacillus coagulans, a Gram-positive thermophilic bacterium, is recognized for its probiotic properties and recent development as a microbial cell factory. Despite its importance for biotechnological applications, the current understanding of B. coagulans’ robustness is limited, especially for undomesticated strains. To fill this knowledge gap, we characterized the metabolic capability and performed functional genomics and systems analysis of a novel, robust strain, B. coagulans B-768. Genome sequencing revealed that B-768 has the largest B. coagulans genome known to date (3.94 Mbp), about 0.63 Mbp larger than the average genome of sequenced B. coagulans strains, with expanded carbohydrate metabolism and mobilome. Functional genomics identified a well-equipped genetic portfolio for utilizing a wide range of C5 (xylose, arabinose), C6 (glucose, mannose, galactose), and C12 (cellobiose) sugars present in biomass hydrolysates, which was validated experimentally. For growth on individual xylose and glucose, the dominant sugars in biomass hydrolysates, B-768 exhibited distinct phenotypes and proteome profiles. Faster growth and glucose uptake rates resulted in lactate overflow metabolism, which makes B. coagulans a lactate overproducer; however, slower growth and xylose uptake diminished overflow metabolism due to the high energy demand for sugar assimilation. Carbohydrate Transport and Metabolism (COG-G), Translation (COG-J), and Energy Conversion and Production (COG-C) made up 60%–65% of the measured proteomes but were allocated differently when growing on xylose and glucose. The trade-off in proteome reallocation, with high investment in COG-C over COG-G, explains the xylose growth phenotype with significant upregulation of xylose metabolism, pyruvate metabolism, and tricarboxylic acid (TCA) cycle. Strain B-768 tolerates and effectively utilizes inhibitory biomass hydrolysates containing mixed sugars and exhibits hierarchical sugar utilization with glucose as the preferential substrate.

carbohydrate metabolism↗

Hyporheic zone, river, and groundwater metagenome resolved genomes and rpS3 genes in East River Watershed, Colorado USA Summer 2020, 2021

Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from water filter collected across 8 locations along the East River Watershed, CO, and 1 nearby groundwater well. The purpose was to look for connectivity and similarities across the network and to see the impact of the groundwater. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed community composition and strain similarities between the sites and we also compared it to previous metagenomic studies within the watershed looking at floodplain (Matheus Carnevali et al. 2021) and hillslope (Lavy et al. 2019) microbiomes. Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from filters across 8 locations during August 2020 and July 2021. This resulted in 32 samples. The groundwater sample was sequenced at UC Berkley's QB3. The other 31 samples were sequenced at University of Maryland. Metagenomes were assembled using four autobinners and the best bins were selected using dasTool. The genomes were dereplicated at 95% with dRep and the subset of winning genomes were manually curated based on visual inspection of taxonomic profile, GC content, coverage, and a set of 51 bacterial single copy genes (BSCG), and 38 archaeal signal copy genes (ASCG). The dataset includes a zip file of 311 genomes (HZ_River_SW_MAGS_Dereplicated_95.zip). The dataset additionally includes a zipped file of ribosomal protein small subunit 3 (rpS3) proteins from the hyporheic zone and river data (rpS3_Proteins_HZ_River.zip), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a location metadata file (locations.csv). This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

DNA↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

Conjugation-based genome engineering enables rapid prototyping and bioproduction in non-model bacteria

Abstract Non-model bacteria offer unique metabolic capabilities for sustainable bioproduction, yet their limited genetic accessibility hinders systematic strain development. Here we present conjugation-based serine recombinase-assisted genome engineering (cSAGE), a broad-host-range platform that enables predictable, iterative genomic integration in transformation-resistant bacteria. cSAGE combines conjugative DNA delivery, standardized low-copy vectors, orthogonal recombinases, and modular genetic parts to support rapid pathway assembly and cross-host benchmarking. Using purple nonsulfur bacteria as a testbed, we integrate promoter engineering, multi-payload genome modification, and genome-scale metabolic modeling to empirically evaluate host-dependent pathway performance. Applying this workflow, we identify strain-specific differences in photosynthetic conversion of lignin-derived p -coumarate to the thermoplastic precursor p -vinylphenol. By enabling genome engineering and functional comparison across diverse bacteria using a single plasmid system, cSAGE provides a general framework for non-model strain prototyping and biotransformation discovery.

Guzman, Michael S. [Department of Chemical Enginee↗

Genomic and transcriptomic characterization of carbohydrate-active enzymes in the anaerobic fungus Neocallimastix cameroonii var. constans

Anaerobic gut fungi effectively degrade lignocellulose in the guts of large herbivores, but there remain a limited number of isolated, publicly available, and sequenced strains that impede our understanding of the role of anaerobic fungi within microbial communities. We isolated and characterized a new fungal isolate, Neocallimastix cameroonii var. constans, providing a transcriptomic and genomic understanding of its ability to degrade diverse carbohydrates. This anaerobic fungal strain was stably cultivated for multiple years in vitro among members of an initial enrichment microbial community derived from goat feces, and it demonstrated the ability to pair with other microbial members, namely, archaeal methanogens to produce methane from lignocellulose. Genomic analysis revealed a higher number of predicted carbohydrate-active enzymes encoded in the N. cameroonii var. constans genome compared to most other sequenced anaerobic fungi. The carbohydrate-active enzyme profile for this isolate contained 660 glycoside hydrolases, 160 carbohydrate esterases, 194 glycosyltransferases, and 85 polysaccharide lyases. Differential gene expression analysis showed the upregulation of thousands of genes (including predicted carbohydrate-active enzymes) when N. cameroonii var. constans was grown on lignocellulose (reed canary grass) compared to less complex substrates, such as cellulose (filter paper), cellobiose, and glucose. AlphaFold was used to predict functions of transcriptionally active yet poorly annotated genes, revealing feruloyl esterases that likely play an important role in lignocellulose degradation by anaerobic fungi. The combination of this strain's genomic and transcriptomic characterization, omics-informed structural prediction, and robustness in microbial co-culture make it a well-suited platform to conduct future investigations into bioprocessing and enzyme discovery.

CAZymes↗