Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Sequence Homology”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Bottom-Up Simulation, Reconstruction, and Quantification of Macromolecule Sequences from Experimental Polymerizations

Motivated by the canonical sequence–structure–function paradigm, tools to characterize chemical patterning in natural biomacromolecules, from proteins to nucleic acids, have grown exponentially in recent years. However, analogous strategies for synthetic macromolecules remain in nascent stages, complicated by sequence polydispersity and analytical limitations. To address this, we have developed a comprehensive and open-source Python package, PRISM (polymer rate insights and sequence modeling), an end-to-end workflow that provides a path from experimental kinetics measurements to quantitative and qualitative metrics for describing chemical patterning in stochastic polymers. First, a numerical integration strategy was constructed to simulate and fit experimental data from reversible addition–fragmentation chain transfer (RAFT) polymerization kinetics, enabling the facile estimation of relevant reactivity ratios. These ratios were then used in a mechanism-specific stochastic kinetic simulation strategy to simulate sequence ensembles corresponding to model systems spanning experimental copolymers, classes of statistical polymers (e.g., alternating, block, and gradient), and multiblock copolymers. Lastly, inspired by sequence homology metrics from bioinformatics, we introduce visualization strategies and quantitative metrics to facilitate comparisons of different sequence ensembles. As the sequence–structure–function paradigm becomes increasingly central in de novo design of synthetic macromolecules, this toolkit provides a first step toward accurate and representative sequence description and featurization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

PRIME: An evaluation framework for protein representation inference and generalization in viral mutation space

Background Protein language models (PLMs) have revolutionized protein fitness prediction, yet their application to rapidly evolving viral pathogens is often confounded by extreme sequence homology. This homology leads to “data leakage” in standard random validation splits, yielding inflated performance metrics that fail to translate into real-world biosurveillance utility. Results We present Protein Representation Inference for Mutation Evaluation (PRIME), a framework that integrates domain-specific fine-tuning with a rigorous position-stratified validation protocol to evaluate viral threats. Using a dataset of 347,432 SARS-CoV-2 receptor binding domain (RBD) sequences, we demonstrate that while random training data split yields deceptive R 2 values (> 0.90), they fail to generalize to novel mutational sites. By benchmarking models up to 650 M parameters, we show that domain-specific fine-tuning of the ESM-C 600 M model with correctly stratified data provides an initial demonstration of predictive signal for binding affinity and expression at unseen mutational sites of binding affinity and expression on unseen sites (R 2 ~0.23), a significant advancement over base foundation models which exhibit no predictive power (R 2 <0). PRIME’s embedding-based clustering identified 3.03% of bat coronavirus sequences as candidates for further experimental prioritization based on their functional similarity to human-infective strains in embedding space, offering a perspective complementary to traditional phylogenetic methods. Conclusion PRIME establishes a new benchmark for the application of PLMs in pathogen surveillance. Our findings demonstrate that state-of-the-art models and fine-tuning, when paired with stratified validation, provide biologically meaningful insights into pathogen evolution and zoonotic risk.

59 BASIC BIOLOGICAL SCIENCES

An alternative pocket for binding the N‐degrons by the UBR1 and UBR2 ubiquitin E3 ligases

The UBR family of ubiquitin ligases binds to N-termini of their targets (known as N-degron) to induce their ubiquitination and degradation via a conserved domain known as UBR-box. UBR1 and UBR2 share the highest sequence homology among the family, and substantial structural studies were previously performed for substrate binding by the UBR-boxes of UBR1 and UBR2. Here, we describe a new pocket in the UBR-boxes of UBR1 and UBR2 for binding the second residues of N-degrons through determining five co-crystal structures of the UBR-boxes with various N-degron peptides. Together with binding affinities measured by fluorescence polarization, we show that the two highly homologous UBR-boxes can interact with the second residue of an N-degron differently. In addition, the UBR-boxes undergo different conformational changes when binding N-degrons. Furthermore, we demonstrate that the sidechain of the third amino acid of an N-degron has no contribution to binding the UBR-boxes. These findings represent a new conceptual advancement for the UBR E3 ligases and the new insights described here can be leveraged for developing their selective ligands for research and potential therapies.

N-end rule

Evolutionary trajectory of transcription factors and selection of targets for metabolic engineering

Transcription factors (TFs) provide potentially powerful tools for plant metabolic engineering as they often control multiple genes in a metabolic pathway. However, selecting the best TF for a particular pathway has been challenging, and the selection often relies significantly on phylogenetic relationships. Here, we offer examples where evolutionary relationships have facilitated the selection of the suitable TFs, alongside situations where such relationships are misleading from the perspective of metabolic engineering. We argue that the evolutionary trajectory of a particular TF might be a better indicator than protein sequence homology alone in helping decide the best targets for plant metabolic engineering efforts. This article is part of the theme issue ‘The evolution of plant metabolism’.

Life Sciences & Biomedicine - Other Topics

Genetics of Flooding Tolerance in an F 2 Miscanthus sacchariflorus ssp. lutarioriparius × M. sinensis Population

Miscanthus is a warm-season, perennial grass cultivated as a feedstock for bioenergy and bioproducts. M. sacchariflorus ssp. lutarioriparius has high yield potential and is well-adapted to seasonal flooding, but little is known about the genetics of this adaptation. We conducted a quantitative trait locus (QTL) analysis on a population of 332 diploid Miscanthus ×giganteus (Mxg) F2s derived from an initial cross between diploid M. sacchariflorus ssp. lutarioriparius ‘PF30022’ and diploid M. sinensis ‘PMS-014’, followed by intermating 50 F 1 s. Using tanks in a greenhouse to assess the effects of partial submergence on actively growing plants, we compared an aerobic soil control to a 6-week flood treatment. The study's primary objectives were to (1) identify QTL for flooding tolerance in Miscanthus , (2) identify candidate genes and (3) compare ethylene response factors in Miscanthus with those in rice and Arabidopsis , sorghum and maize for binding site sequence homology and synteny, especially those associated with flooding tolerance. In total, 10 QTL and 66 candidate genes for partial submergence tolerance were identified (including many for ethylene signalling), a first report for Miscanthus . Notably, none of the Miscanthus candidates were orthologs of rice Sub1A, SK1 or SK2 , yet the ‘PF30022’ parent exhibited a snorkeling phenotype, indicating convergent evolution. This study will facilitate breeding of climate-resiliant Mxg.

abiotic stress tolerance

Molecular and structural characterization of a Bacillus cereus strain producing an anthrax-like capsule

Bacillus cereus is a ubiquitous Gram-positive, spore-forming, rod-shaped saprophytic bacterium, occasionally reported to cause food-borne illnesses. However, instances of B. cereus strains harboring anthrax toxin and capsule genes have elevated certain strains as formidable pathogens and biothreats. This study focuses on the genomic analysis and the structural characterization of capsular material produced by the virulent B. cereus PATH2418 strain, isolated from the wound of a traumatic open fracture patient. The genome was sequenced using Nanopore MinION sequencing, revealing a chromosome of 5,270,283 bp and three plasmids. One plasmid, pATH1, was found to encode an operon for the biosynthesis of a bacterial capsule. This operon had sequence homology to the Bacillus anthracis capBCADE operon, which encodes the poly-γ-D-glutamate (PDGA) capsule. The capsule production in B. cereus PATH2418 was influenced by temperature and CO 2 levels. Structural analysis of the capsular material using a combined approach of nuclear magnetic resonance (NMR) and high-performance liquid chromatography (HPLC) techniques confirmed the presence of a high-molecular-weight poly-γ-glutamate capsule, with an enantiomeric composition of approximately 67% D-glutamic acid and 33% L-glutamic acid, matching that of B. anthracis.

Bacillus cereus

Spatial proteomics reveals signal sequence characteristics correlated with localization in cyanobacteria

Abstract Cyanobacteria have an inner and outer cell membrane enclosing the periplasm and cell wall and an additional set of internal membranes (called the thylakoid membranes) enclosing the thylakoid lumen. The periplasm and thylakoid lumen have unique proteomes, but the mechanisms regulating protein sorting to these locations have remained elusive. Here, proximity-based proteomics using the engineered peroxidase APEX2 was performed in the cyanobacteria Synechococcus sp. PCC 7002 to profile the proteomes of the cytoplasm, thylakoid lumen, and the periplasm and outer membrane (P-OM). Our analyses revealed specific roles for the thylakoid lumen in photosynthesis and energy generation, as well as roles for the periplasm in metabolite transport and binding, cell motility, and cell wall maintenance. Forty proteins localized to both the thylakoid lumen and the P-OM; however, their biological functions remain unclear. We also analyzed the correlation between signal sequence characteristics and differential protein localization to either the thylakoid lumen or the P-OM. In PCC 7002, as well as Synechocystis sp. PCC 6803 and Nostoc sp. PCC 7120, thylakoid lumen proteins translocated across membranes via the Secretory (Sec) system possessed more hydrophobic and alpha-helical signal sequence H-regions than P-OM proteins. The signal sequences of homologous proteins in Gloeobacter violaceus PCC 7421, a cyanobacterial species with a combined thylakoid lumen and periplasmic space, did not exhibit such differences. Therefore, the pattern of increased H-region hydrophobicity and alpha helix content is specific to cyanobacteria with a separate thylakoid lumen space and likely contributes to proper protein sorting between the thylakoid lumen and periplasm.

Plant Sciences

Revisiting synthetic lethality of Gcn5-related N-acetyltransferase (GNAT) family mutations in Haloferax volcanii

ABSTRACT Lysine acetylation is a post-translational modification that occurs in all domains of life, highlighting its evolutionary significance. Previous genome comparison identified three Gcn5-related N-acetyltransferase (GNAT) family members as lysine acetyltransferase homologs (Pat1, Pat2, and Elp3) and two deacetylase homologs (Sir2 and HdaI) in the halophilic archaeonHaloferax volcanii, withelp3andpat2proposed as a synthetic lethal gene pair. Here, we advance these findings by performing single and double mutagenesis ofelp3with thepat1andpat2lysine acetyltransferase gene homologs. Genome sequencing and PCR screens of these strains reveal successful generation of Δelp3,Δpat1Δelp3, and Δpat2Δelp3mutant strains. Although these mutant strains exhibited a reduced growth rate compared to the parent, they remained viable. Overall, this study provides genetic evidence thatelp3andpat2, while impacting cell growth, are not a synthetic lethal gene pair as previously reported. IMPORTANCE Here, we reveal by whole-genome sequencing that the GNAT family gene homologselp3andpat2can be deleted in the sameHaloferax volcaniistrain. Beyond the targeted deletions, minimal differences between the parent and Δelp3Δpat2mutant were observed, suggesting that suppressor mutations are not responsible for our ability to generate this double mutant strain. Elp3 and Pat2, thus, may not share as close a functional relationship as implied by earlier study. Our finding is significant as Elp3 is thought to function in acetylation in tRNA modification, while Pat2 likely functions in the lysine acetylation of proteins.

Microbiology

Functional role of myosin-binding protein H in thick filaments of developing vertebrate fast-twitch skeletal muscle

Myosin-binding protein H (MyBP-H) is a component of the vertebrate skeletal muscle sarcomere with sequence and domain homology to myosin-binding protein C (MyBP-C). Whereas skeletal muscle isoforms of MyBP-C (fMyBP-C, sMyBP-C) modulate muscle contractility via interactions with actin thin filaments and myosin motors within the muscle sarcomere “C-zone,” MyBP-H has no known function. This is in part due to MyBP-H having limited expression in adult fast-twitch muscle and no known involvement in muscle disease. Quantitative proteomics reported here reveal that MyBP-H is highly expressed in prenatal rat fast-twitch muscles and larval zebrafish, suggesting a conserved role in muscle development and prompting studies to define its function. We take advantage of the genetic control of the zebrafish model and a combination of structural, functional, and biophysical techniques to interrogate the role of MyBP-H. Transgenic, FLAG-tagged MyBP-H or fMyBP-C both localize to the C-zones in larval myofibers, whereas genetic depletion of endogenous MyBP-H or fMyBP-C leads to increased accumulation of the other, suggesting competition for C-zone binding sites. Does MyBP-H modulate contractility in the C-zone? Globular domains critical to MyBP-C’s modulatory functions are absent from MyBP-H, suggesting that MyBP-H may be functionally silent. However, our results suggest an active role. In vitro motility experiments indicate MyBP-H shares MyBP-C’s capacity as a molecular “brake.” These results provide new insights and raise questions about the role of the C-zone during muscle development.

59 BASIC BIOLOGICAL SCIENCES

NEAR: Neural Embeddings for Amino acid Relationships

Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present NEAR, a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR’s ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript we focus on a more straightforward analysis of the model’s value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely-used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail.

59 BASIC BIOLOGICAL SCIENCES

Biophysical and biochemical evidence for the role of acetate kinases (AckAs) in an acetogenic pathway in pathogenic spirochetes

Unraveling the metabolism of Treponema pallidum is a key component to understanding the pathogenesis of the human disease that it causes, syphilis. For decades, it was assumed that glucose was the sole carbon/energy source for this parasitic spirochete. But the lack of citric-acid-cycle enzymes suggested that alternative sources could be utilized, especially in microaerophilic host environments where glycolysis should not be robust. Recent bioinformatic, biophysical, and biochemical evidence supports the existence of an acetogenic energy-conservation pathway in T . pallidum and related treponemal species. In this hypothetical pathway, exogenous D-lactate can be utilized by the bacterium as an alternative energy source. Herein, we examined the final enzyme in this pathway, acetate kinase (named TP0476), which ostensibly catalyzes the generation of ATP from ADP and acetyl-phosphate. We found that TP0476 was able to carry out this reaction, but the protein was not suitable for biophysical and structural characterization. We thus performed additional studies on the homologous enzyme (75% amino-acid sequence identity) from the oral pathogen Treponema vincentii , TV0924. This protein also exhibited acetate kinase activity, and it was amenable to structural and biophysical studies. We established that the enzyme exists as a dimer in solution, and then determined its crystal structure at a resolution of 1.36 Å, showing that the protein has a similar fold to other known acetate kinases. Mutation of residues in the putative active site drastically altered its enzymatic activity. A second crystal structure of TV0924 in the presence of AMP (at 1.3 Å resolution) provided insight into the binding of one of the enzyme’s substrates. On balance, this evidence strongly supported the roles of TP0476 and TV0924 as acetate kinases, reinforcing the hypothesis of an acetogenic pathway in pathogenic treponemes.

Deka, Ranjit K.

Identification of candidate host-specificity genes in Exserohilum turcicum using comparative genomics and transcriptomics

Abstract Exserohilum turcicum causes northern corn leaf blight and sorghum leaf blight. While the same species cause disease in both crops, the strains are host-specific. Here, we report the sequence and de novo annotated assemblies of one sorghum- and one maize-specific E. turcicum strain. The strains were sequenced using the PacBio Sequel II system. The total genome length for both assemblies was between 44 and 45 Mb with N50 of ∼2.5 Mb. Ninety-eight percent of the Benchmarking Universal Single-Copy Orthologs (BUSCO) for both assemblies had complete status. The estimated number of genes was 11,762 and 12,029 in the sorghum- and maize-specific isolates, respectively. Funannotate, EffectorP, SignalP, and transcriptome data were used to create functional annotation of each genome. The whole-genome comparison identified ten large-scale inversions and three translocations between the maize- and sorghum-specific strains, along with homologous genes and gene duplications. RNA was sequenced from the maize- and sorghum-specific isolate 10 days post-inoculation in maize and sorghum and from axenic cultures. Gene expression data from planta and axenic growth experiments were compared for each strain. Candidate host-specificity genes were identified by combining results from whole-genome comparison, synteny analysis, gene annotations, and transcriptome data. Overall, this study identified several candidate host-specificity genes that provide insights into E. turcicum interaction with its hosts.

Krone, Mara J. (ORCID:0000000159006624)

Functional diversification within the heme-binding split-barrel family

Due to neofunctionalization, a single fold can be identified in multiple proteins that have distinct molecular functions. Depending on the time that has passed since gene duplication and the number of mutations, the sequence similarity between functionally divergent proteins can be relatively high, eroding the value of sequence similarity as the sole tool for accurately annotating the function of uncharacterized homologs. Here, we combine bioinformatic approaches with targeted experimentation to reveal a large multifunctional family of putative enzymatic and nonenzymatic proteins involved in heme metabolism. This family (homolog of HugZ (HOZ)) is embedded in the “FMN-binding split barrel” superfamily and contains separate groups of proteins from prokaryotes, plants, and algae, which bind heme and either catalyze its degradation or function as nonenzymatic heme sensors. In prokaryotes these proteins are often involved in iron assimilation, whereas several plant and algal homologs are predicted to degrade heme in the plastid or regulate heme biosynthesis. In the plant Arabidopsis thaliana, which contains two HOZ subfamilies that can degrade heme in vitro (HOZ1 and HOZ2), disruption of AtHOZ1 (AT3G03890) or AtHOZ2A (AT1G51560) causes developmental delays, pointing to important biological roles in the plastid. In the tree Populus trichocarpa, a recent duplication event of a HOZ1 ancestor has resulted in localization of a paralog to the cytosol. Structural characterization of this cytosolic paralog and comparison to published homologous structures suggests conservation of heme-binding sites. This study unifies our understanding of the sequence-structure-function relationships within this multilineage family of heme-binding proteins and presents new molecular players in plant and bacterial heme metabolism.

59 BASIC BIOLOGICAL SCIENCES

Microbial species and intraspecies units exist and are maintained by ecological cohesiveness coupled to high homologous recombination

Abstract Recent genomic analyses have revealed that microbial communities are predominantly composed of persistent, sequence-discrete species and intraspecies units (genomovars), but the mechanisms that create and maintain these units remain unclear. By analyzing closely-related isolate genomes from the same or related samples and identifying recent recombination events using a novel bioinformatics methodology, we show that high ecological cohesiveness coupled to frequent-enough and unbiased (i.e., not selection-driven) horizontal gene flow, mediated by homologous recombination, often underlie these diversity patterns. Ecological cohesiveness was inferred based on greater similarity in temporal abundance patterns of genomes of the same vs. different units, and recombination was shown to affect all sizable segments of the genome (i.e., be genome-wide) and have two times or greater impact on sequence evolution than point mutations. These results were observed in bothSalinibacter ruber, an environmental halophilic organism, andEscherichia coli, the model gut-associated organism and an opportunistic pathogen, indicating that they may be more broadly applicable to the microbial world. Therefore, our results represent a departure compared to previous models of microbial speciation that invoke either ecology or recombination, but not necessarily their synergistic effect, and answer an important question for microbiology: what a species and a subspecies are.

Science & Technology - Other Topics

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach because they provide statistically calibrated E-values; however, their gap behavior is fixed once a profile is trained, despite biological evidence that insertion/deletion tolerance varies across flexible loops and intrinsically disordered regions. We present CAHS (Context-Aware Homology Search), a lightweight query-time adapter for pHMM search that incorporates learned and biologically motivated signals without changing HMMER's downstream search pipeline or its calibrated E-value reporting. Given a query sequence, CAHS computes per-residue representations from a protein language model and a disorder predictor, maps these to profile coordinates, and modulates only match-state transition rows (gap-open and gap-extension probabilities) while preserving Plan7 constraints. We comprehensively evaluate CAHS across six structurally diverse protein families and multi-domain architectures against a 570k-sequence target corpus. CAHS expands detection capability, retrieving thousands of additional remote homologs at relaxed thresholds by maintaining alignment quality through flexible regions. For multi-domain proteins, context-aware modulation resolves 94% of fragmented alignments. Crucially, CAHS preserves hit-set invariance at stringent operating points (E<10-10), demonstrating increased statistical confidence without inflating false positives. Furthermore, sharper statistical distinction between homologs and background noise during early filter stages yields up to a 3.87× acceleration in end-to-end wall-clock time on high-performance computing clusters. Overall, CAHS illustrates a practical AI-for-science design pattern: augmenting a trusted probabilistic model with query-specific learned signals to improve interpretable, reproducible inference in data-rich biology.

Bhattaram, Swethasree [Georgia Institute of Techno

Transcriptomic and functional analyses uncover a conserved effector driving genotype-dependent virulence in the Sphaerulina musiva-Populus trichocarpa interaction

The introduction of invasive microbes compromises the structure, biodiversity, and function of naïve ecosystems. Sphaerulina musiva, a hemibiotrophic pathogen that causes leaf spot and stem cankers in Populus species, exemplifies an invasive fungal pathogen spread by human activities. However, the genetic mechanisms of pathogenicity and virulence are poorly understood, impeding mitigation strategies. We utilized RNA sequencing to identify fungal effectors linked to stem canker formation, informing the development of future strategies for effective disease management. Our analysis revealed 70 genes differentially expressed at 2 weeks and 110 genes at 3 weeks between inoculated trees and controls. Notably, the gene with the highest expression at 2 weeks and the second highest at 3 weeks was homologous to Extracellular protein 2 (Ecp2). Complementary genome-wide association studies linked sequence polymorphisms in this locus to phenotypic variation in disease severity. Infiltration of S. musiva Ecp2 into Populus trichocarpa leaves induced necrosis in susceptible genotypes. Gene disruption using a CRISPR-Cas9 RNP system resulted in a genotype-dependent reduction of stem canker and disease severity. Tracing the evolutionary history of this effector across the fungal kingdom, we uncovered clade-specific gene-family expansions and orthologs in new species. These findings raise questions about the function and adaptive significance of these gene families in fungal lifestyles. Our study provides the first tractable target for breeding resistant poplar genotypes, addressing the challenges of managing S. musiva and uncovering mechanisms that drive its virulence, and provides deeper insights into the evolutionary dynamics of a conserved small-secreted protein with a diversity of functions.

Sondreli, Kelsey L [Oregon State University]

The HIGH CHLOROPHYLL FLUORESCENCE 244 homolog CrHCF244 is required for psbA (D1) translation in Chlamydomonas reinhardtii

Translation of psbA, the chloroplast gene that encodes the D1 subunit of PSII, is important for both PSII biogenesis and repair. The translation of psbA transcripts in the chloroplast is under the control of nuclear gene products. Using a forward genetic screen and whole-genome sequencing of the alga Chlamydomonas reinhardtii , we found a mutant defective in PSII activity and mapped the causative gene to be the homolog of Arabidopsis HIGH CHLOROPHYLL FLUORESCENCE 244 (HCF244) , namely CrHCF244 . We then demonstrated that CrHCF244 is required for psbA translation in the alga, consistent with the function of HCF244 in Arabidopsis, and found that AtHCF244 also partially complemented the algal mutant. These results experimentally support the functional conservation of the homologs in green algae and land plants. Intriguingly, the CrHCF244 mutant also exhibited a relatively high rate of suppressor mutants, pointing to the presence of alternative factor(s)/pathway(s) for D1 translational control. The establishment of CrHCF244 as a psbA translation factor in C. reinhardti i shows the similarities in psbA translation regulation in algae and plants. The future identification of the alternative factor(s) in this alga will provide insights on psbA translation in plants.

Arabidopsis

Xylanolytic metabolism is regulated by coordination of transcription factors XynR and XylR in extremely thermophilic Caldicellulosiruptorales

ABSTRACT Global transcription factors (TFs) control metabolic processes in bacteria to efficiently utilize available carbon. The orderCaldicellulosiruptoraleshas drawn interest due to the ability of its members to degrade components of lignocellulosic biomass. Regulatory reconstruction ofAnaerocellum (f. Caldicellulosiruptor) besciiidentified two major global transcription factors for xylan utilization, XynR and XylR, and the corresponding putative transcription factor binding sites. Recombinant versions of XynR (LacI family) and XylR (ROK family) were subjected to fluorescence polarization (FP) and biolayer interferometry (BLI) analysis to confirm the predicted binding sites. Four XynR sites and two XylR sites were validated, accounting for 20 of 26 genes regulated by XynR and six of seven genes regulated by XylR. Bioinformatic analysis of the individual genes controlled by the two regulators showed an inter-dependent scheme for xylan conversion; the transport of xylooligosaccharides (XOS) is dependent on XylR, while enzymes responsible for hydrolysis are controlled by both regulators. For xylose catabolism by the xylose isomerase-xylulose kinase pathway, regulation is also split, with XylR controlling xylose isomerase and XynR controlling xylokinase. The XynR/XylR regulator pair withinA. besciiis conserved in all sequenced species ofCaldicellulosiruptorales, suggesting similarities in regulating linear xylan conversion. In other xylanolytic thermophiles, XylR homologs control xylan degradation, compared to just 6 out of 26 genes forA. bescii. These results show that two separate regulatory schemes (dual repression) are coordinated byA. besciito effectively regulate the hemicellulose inventory and xylan catabolism. IMPORTANCE To take full advantage of extreme thermophiles as platform metabolic engineering microorganisms, the tools for genetic manipulation must be further developed, and strategies that exploit a better understanding of metabolic regulation need to be discerned.Anaerocellum bescii, the most studied of the extremely thermophilic fermentative anaerobic bacteria that can utilize microcrystalline cellulose, can degrade microcrystalline cellulose and hemicellulose and has been metabolically engineered to convert the resulting sugars to products such as ethanol and acetone. For xylan, in particular, two major global transcription factors (TFs), XynR and XylR, play a role in sugar metabolism, although their predicted regulatory interdependence from bioinformatics analysis has not been elucidated experimentally. Here, fluorescence polarization (FP) and biolayer interferometry (BLI) were used to explore this issue to support metabolic engineering efforts aimed at improving carbohydrate processing to industrial chemicals.

Biotechnology & Applied Microbiology