Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

UnigeneFinder: An Automated Pipeline for Gene Calling From Transcriptome Assemblies Without a Reference Genome

ABSTRACT For most species, transcriptome data are much more readily available than genome data. Without a reference genome, gene calling is cumbersome and inaccurate because of the high degree of redundancy in de novo transcriptome assemblies. To simplify and increase the accuracy of de novo transcriptome assembly in the absence of a reference genome, we developed UnigeneFinder. Combining several clustering methods, UnigeneFinder substantially reduces the redundancy typical of raw transcriptome assemblies. This pipeline offers an effective solution to the problem of inflated transcript numbers, achieving a closer representation of the actual underlying genome. UnigeneFinder performs comparably or better, compared with existing tools, on plant species with varying genome complexities. UnigeneFinder is the only available transcriptome redundancy solution that fully automates the generation of primary transcript, coding region, and protein sequences, analogous to those available for high‐quality reference genomes. These features, coupled with the pipeline’s cross‐platform implementation, focus on automation, and an accessible, user‐friendly interface, make UnigeneFinder a useful tool for many downstream sequence‐based analyses in nonmodel organisms lacking a reference genome, including differential gene expression analysis, accurate ortholog identification, functional enrichments, and evolutionary analyses. UnigeneFinder also runs efficiently both on high‐performance computing (HPC) systems and personal computers, further reducing barriers to use.

Xue, Bo [Plant Resilience Institute Michigan State↗

How hydrophobicity, side chains, and salt affect the dimensions of disordered proteins

Abstract Despite the generally accepted role of the hydrophobic effect as the driving force for folding, many intrinsically disordered proteins (IDPs), including those with hydrophobic content typical of foldable proteins, behave nearly as self‐avoiding random walks (SARWs) under physiological conditions. Here, we tested how temperature and ionic conditions influence the dimensions of the N‐terminal domain of pertactin (PNt), an IDP with an amino acid composition typical of folded proteins. While PNt contracts somewhat with temperature, it nevertheless remains expanded over 10–58°C, with a Flory exponent, ν , >0.50. Both low and high ionic strength also produce contraction in PNt, but this contraction is mitigated by reducing charge segregation. With 46% glycine and low hydrophobicity, the reduced form of snow flea anti‐freeze protein (red‐sfAFP) is unaffected by temperature and ionic strength and persists as a near‐SARW, ν ~ 0.54, arguing that the thermal contraction of PNt is due to stronger interactions between hydrophobic side chains. Additionally, red‐sfAFP is a proxy for the polypeptide backbone, which has been thought to collapse in water. Increasing the glycine segregation in red‐sfAFP had minimal effect on ν . Water remained a good solvent even with 21 consecutive glycine residues ( ν > 0.5), and red‐sfAFP variants lacked stable backbone hydrogen bonds according to hydrogen exchange. Similarly, changing glycine segregation has little impact on ν in other glycine‐rich proteins. These findings underscore the generality that many disordered states can be expanded and unstructured, and that the hydrophobic effect alone is insufficient to drive significant chain collapse for typical protein sequences.

59 BASIC BIOLOGICAL SCIENCES↗

Different chemical scaffolds bind to L-phe site in Mycobacterium tuberculosis Phe-tRNA synthetase

Tuberculosis (TB), caused by Mycobacterium tuberculosis (Mt), is one of the deadliest infectious diseases. The rise of multidrug-resistant strains represents a major public health threat, requiring new therapeutic options. Bacterial aminoacyl-tRNA synthetases (aaRS) have been shown to be highly promising drug targets, including for TB treatment. These enzymes play an essential role in translating the DNA gene code into protein sequence by attaching specific amino acid to their cognate tRNAs. They have multiple binding sites that can be targeted for inhibitor discovery: amino acid binding pocket, ATP binding pocket, tRNA binding site and an editing domain. Recently we reported several high-resolution structures of M. tuberculosis phenylalanyl-tRNA synthetase (MtPheRS) complexed with tRNA Phe and either L-Phe or a nonhydrolyzable phenylalanine adenylate analog. Here, in this study, using Nucleic Magnetic Resonance (NMR) and Surface Plasmon Resonance (SPR) we identified fragments that bind to MtPheRS and we determined crystal structures of their complexes with MtPheRS/tRNA Phe . All the binders interact with the L-Phe amino acid binding site. The analysis of interactions of the new compounds combined with adenylate analog structure provides insights for the rational design of antituberculosis drugs. The 3 ' arm of the tRNA Phe in all the structures was disordered with exception of one complex with D-735 compound. In this structure the 3' CCA end of the acceptor stem is observed in the editing domain of MtPheRS providing insights regarding the post-transfer editing activity of class II aaRS.

Gade, Priyanka [Univ. of Chicago, IL (United State↗

Chromosome-level genome assembly of Quercus variabilis provides insights into the molecular mechanism of cork thickness

Quercus variabilis is a deciduous woody species with high ecological and economic value and is a major source of cork in East Asia. Cork from thick softwood sheets have higher commercial value than those from thin sheets. It is extremely difficult to genetically improve Q. variabilis to produce high quality softwood due to the lack of genomic information. Here, we present a high-quality chromosomal genome assembly for Q. variabilis with length of 791,89 Mb and 54,606 predicted genes. Comparative analysis of protein sequences of Q. variabilis with 11 other species revealed that specific and expanded gene families were significantly enriched in the "fatty acid biosynthesis" pathway in Q. variabilis, which may contribute to the formation of its unique cork. Additionally, based on weighted correlation network analysis of time-course (i.e., five important developmental ages) gene expression data in thick-cork versus thin-cork genotypes of Q. variabilis, we identified one co-expression gene module associated with the thick-cork trait. Within this co-expression gene module, 10 hub genes were associated with suberin biosynthesis. Furthermore, we identified a total of 198 suberin biosynthesis-related new candidate genes that were up-regulated in trees with a thick cork layer relative to those with a thin cork layer. Also, we found that some genes related to cell expansion and cell division were highly expressed in trees with a thick cork layer. Collectively, our results revealed that two metabolic pathways (i.e., suberin biosynthesis, fatty acid biosynthesis), along with other genes involved in cell expansion, cell division, and transcriptional regulation, were associated with the thick-cork trait in Q. variabilis, providing insights into the molecular basis of cork development and knowledge for informing genetic improvement of cork thickness in Q. variabilis and closely related species.

59 BASIC BIOLOGICAL SCIENCES↗

An in silico method to assess antibody fragment polyreactivity

Antibodies are essential biological research tools and important therapeutic agents, but some exhibit non-specific binding to off-target proteins and other biomolecules. Such polyreactive antibodies compromise screening pipelines, lead to incorrect and irreproducible experimental results, and are generally intractable for clinical development. Here, we design a set of experiments using a diverse naïve synthetic camelid antibody fragment (nanobody) library to enable machine learning models to accurately assess polyreactivity from protein sequence (AUC > 0.8). Moreover, our models provide quantitative scoring metrics that predict the effect of amino acid substitutions on polyreactivity. We experimentally test our models’ performance on three independent nanobody scaffolds, where over 90% of predicted substitutions successfully reduced polyreactivity. Importantly, the models allow us to diminish the polyreactivity of an angiotensin II type I receptor antagonist nanobody, without compromising its functional properties. We provide a companion web-server that offers a straightforward means of predicting polyreactivity and polyreactivity-reducing mutations for any given nanobody sequence.

60 APPLIED LIFE SCIENCES↗

Detecting macroevolutionary genotype–phenotype associations using error-corrected rates of protein convergence

On macroevolutionary timescales, extensive mutations and phylogenetic uncertainty mask the signals of genotype–phenotype associations underlying convergent evolution. To overcome this problem, we extended the widely used framework of non-synonymous to synonymous substitution rate ratios and developed the novel metric ω C , which measures the error-corrected convergence rate of protein evolution. While ω C distinguishes natural selection from genetic noise and phylogenetic errors in simulation and real examples, its accuracy allows an exploratory genome-wide search of adaptive molecular convergence without phenotypic hypothesis or candidate genes. Using gene expression data, we explored over 20 million branch combinations in vertebrate genes and identified the joint convergence of expression patterns and protein sequences with amino acid substitutions in functionally important sites, providing hypotheses on undiscovered phenotypes. We further extended our method with a heuristic algorithm to detect highly repetitive convergence among computationally non-trivial higher-order phylogenetic combinations. Our approach allows bidirectional searches for genotype–phenotype associations, even in lineages that diverged for hundreds of millions of years.

59 BASIC BIOLOGICAL SCIENCES↗

Three mutations repurpose a plant karrikin receptor to a strigolactone receptor

Significance Parasitic plants like witchweed cause huge losses in crop yield in Africa. A key part to the success of witchweed is to start its life cycle upon sensing small molecules called strigolactones, which are exuded from roots of host plants into the soil. Witchweed sense host-derived strigolactones through receptors called HTLs. It is thought that the evolutionary origin of HTLs is a receptor called KAI2 in nonparasitic plants, which can respond to different small molecules such as karrikins. By making three changes in the protein sequence of KAI2, this hybrid receptor can now sense both strigolactones and karrikins. These results help in understanding how receptors can evolve to sense different signals and can lead to solutions for combating pesky witchweed.

Arellano-Saab, Amir↗

An LCO-responsive homolog of NODULE INCEPTION positively regulates lateral root formation in Populus sp.

Abstract The transcription factor NODULE INCEPTION (NIN) has been studied extensively for its multiple roles in root nodule symbiosis within plants of the nitrogen-fixing clade (NFC) that associate with soil bacteria, such as rhizobia and Frankia. However, NIN homologs are present in plants outside the NFC, suggesting a role in other developmental processes. Here, we show that the biofuel crop Populus sp., which is not part of the NFC, contains eight copies of NIN with diversified protein sequence and expression patterns. Lipo-chitooligosaccharides (LCOs) are produced by rhizobia and a wide range of fungi, including mycorrhizal ones, and act as symbiotic signals that promote lateral root formation. RNAseq analysis of Populus sp. treated with purified LCO showed induction of the PtNIN2 subfamily. Moreover, the expression of PtNIN2b correlated with the formation of lateral roots and was suppressed by cytokinin treatment. Constitutive expression of PtNIN2b overcame the inhibition of lateral root development by cytokinin under high nitrate conditions. Lateral root induction in response to LCOs likely represents an ancestral function of NIN retained and repurposed in nodulating plants, as we demonstrate that the role of NIN in LCO-induced root branching is conserved in both Populus sp. and legumes. We further established a visual marker of LCO perception in Populus sp. roots, the putative sulfotransferase PtSS1 that can be used to study symbiotic interactions with the bacterial and fungal symbionts of Populus sp.

Irving, Thomas B. (ORCID:0000000330404543)↗

Reekeekee- and roodoodooviruses, two different Microviridae clades constituted by the smallest DNA phages

Small circular single-stranded DNA viruses of the Microviridae family are both prevalent and diverse in all ecosystems. They usually harbor a genome between 4.3 and 6.3 kb, with a microvirus recently isolated from a marine Alphaproteobacteria being the smallest known genome of a DNA phage (4.248 kb). A subfamily, Amoyvirinae, has been proposed to classify this virus and other related small Alphaproteobacteria-infecting phages. Here, we report the discovery, in meta-omics data sets from various aquatic ecosystems, of sixteen complete microvirus genomes significantly smaller (2.991–3.692 kb) than known ones. Phylogenetic analysis reveals that these sixteen genomes represent two related, yet distinct and diverse, novel groups of microviruses—amoyviruses being their closest known relatives. We propose that these small microviruses are members of two tentatively named subfamilies Reekeekeevirinae and Roodoodoovirinae. As known microvirus genomes encode many overlapping and overprinted genes that are not identified by gene prediction software, we developed a new methodology to identify all genes based on protein conservation, amino acid composition, and selection pressure estimations. Surprisingly, only four to five genes could be identified per genome, with the number of overprinted genes lower than that in phiX174. These small genomes thus tend to have both a lower number of genes and a shorter length for each gene, leaving no place for variable gene regions that could harbor overprinted genes. Even more surprisingly, these two Microviridae groups had specific and different gene content, and major differences in their conserved protein sequences, highlighting that these two related groups of small genome microviruses use very different strategies to fulfill their lifecycle with such a small number of genes. The discovery of these genomes and the detailed prediction and annotation of their genome content expand our understanding of ssDNA phages in nature and are further evidence that these viruses have explored a wide range of possibilities during their long evolution.

59 BASIC BIOLOGICAL SCIENCES↗

Evolutionary trajectory of transcription factors and selection of targets for metabolic engineering

Transcription factors (TFs) provide potentially powerful tools for plant metabolic engineering as they often control multiple genes in a metabolic pathway. However, selecting the best TF for a particular pathway has been challenging, and the selection often relies significantly on phylogenetic relationships. Here, we offer examples where evolutionary relationships have facilitated the selection of the suitable TFs, alongside situations where such relationships are misleading from the perspective of metabolic engineering. We argue that the evolutionary trajectory of a particular TF might be a better indicator than protein sequence homology alone in helping decide the best targets for plant metabolic engineering efforts. This article is part of the theme issue ‘The evolution of plant metabolism’.

Life Sciences & Biomedicine - Other Topics↗

An HMM approach expands the landscape of sesquiterpene cyclases across the kingdom Fungi

Sesquiterpene cyclases (STC) catalyse the cyclization of the C15 molecule farnesyl diphosphate into a vast variety of mono- or polycyclic hydrocarbons and, for a few enzymes, oxygenated structures, with diverse stereogenic centres. The huge diversity in sesquiterpene skeleton structures in nature is primarily the result of the type of cyclization driven by the STC. Despite the phenomenal impact of fungal sesquiterpenes on the ecology of fungi and their potentials for applications, the fungal sesquiterpenome is largely untapped. The identification of fungal STC is generally based on protein sequence similarity with characterized enzymes. This approach has improved our knowledge on STC in a few fungal species, but it has limited success for the discovery of distant sequences. Besides, the tools based on secondary metabolite biosynthesis gene clusters have shown poor performance for terpene cyclases. Here, we used four sets of sequences of fungal STC that catalyse four types of cyclization, and specific amino acid motives to identify phylogenetically related sequences in the genomes of basidiomycetes fungi from the order Polyporales. We validated that four STC genes newly identified from the genome sequence of Leiotrametes menziesii, each classified in a different phylogenetic clade, catalysed a predicted cyclization of farnesyl diphosphate. We built HMM models and searched STC genes in 656 fungal genomes genomes. We identified 5605 STC genes, which were classified in one of the four clades and had a predicted cyclization mechanism. We noticed that the HMM models were more accurate for the prediction of the type of cyclization catalysed by basidiomycete STC than for ascomycete STC.

59 BASIC BIOLOGICAL SCIENCES↗

Exabiome: Advancing Microbial Science through Exascale Computing

The Exabiome project seeks to improve the understanding of microbiomes through the development of methods for accelerating metagenomic science using exascale computing. This article gives an overview of scientific impact of the three components of the project: metagenome assembly, protein family detection, and comparative analysis of metagenomes. Exabiome developed MetaHipMer, the only metagenome assembler capable of scaling to full exascale systems. MetaHipMer has enabled ground-breaking assemblies on the Frontier supercomputer, with many scientific benefits, such as the discovery of rare species and viral genomes. To investigate protein families, Exabiome developed two exascale tools, PASTIS and HipMCL. Together, these can utilize exascale resources to understand the functional diversity of billions of dark matter proteins and novel protein families. For comparative analysis, Exabiome developed kmerprof, a tool that can be used to compare huge metagenomes for many different scientific purposes, for example, grouping human microbiomes according to body location.

59 BASIC BIOLOGICAL SCIENCES↗

ProteinTuneRL

ProteinTuneRL is a framework designed to harness the power of reinforcement learning for advanced protein design. The project enables fine-tuning of generative models to explore and optimize protein sequences with tailored structural and functional properties.

Landajuela Larma, Mikel [Lawrence Livermore Nation↗

Activation Domain Hunter (ADhunter) v2.0

ADhunter is a software program that enables accurate identification and quantification of transcriptional activation domains. Unlike previous software, ADhunter uses protein representations from a pre-trained protein language model, model ensembling, and a training dataset from a diverse sampling of protein sequence space for state-of-the-art performance. These advantages enable improved perception of transcriptional activation domains across sequence space that can be used for mapping natural genetic circuits and engineering synthetic genetic circuits. In particular, ADhunter enables fine-tuned control of gene expression through synthetic transcription factors that can be used for complex control of cellular programs.

Waldburger, Lucas [Lawrence Berkeley National Labo↗

PRIME: Protein Representation Inference for Mutation Evaluation

Protein language machine learning models built upon existing ESM-2 model developed by Evolutionary Scale (evolutionaryscale.ai) and an in-house protein language model based on the BERT model developed by Google. The code also includes model training scripts and saved checkpoints from our own training using publicly available SARS-CoV-2 protein sequences.

Gibson, Kaetlyn [Los Alamos National Lab]↗

High Throughput expression and characterization of laccases in Saccharomyces cerevisiae

Laccases are oxidative enzymes containing 4 conserved copper heteroatoms. Laccases catalyze cleavage of bonds in lignin using radical chemistry, yet their exact specificity for bonds (such as the β-O-4 or C-C) in lignin remains unknown and may vary with the diversity of laccases across fungi, plants and bacteria. Bond specificity may perhaps even vary for the same enzyme across different reaction conditions. Determining these differences has been difficult due to the fact that heterologous expression of soluble, active laccases has proven difficult. Here we describe the successful heterologous expression of functional laccases in two strains of Saccharomyces cerevisiae, including one we genetically modified with CRISPR. We phylogenically map the enzymes that we successfully expressed, compared to those that did not express. We also describe differences protein sequence differences and pH and temperature profiles and their ability to functionally express, leading to a potential future screening platform for directed evolution of laccases and other ligninolytic enzymes such as peroxidases.

59 BASIC BIOLOGICAL SCIENCES↗

Biosynthesis of bioprivileged, linear molecules via novel carboligase reactions

Over the award period, we made progress on the three aims. We screened twenty-five carboligases for activity coupling twenty-one possible -keto acids (Aim 1). The carboligases were selected across a diverse set of protein sequences. Using Q-Exactive UHPLC-MS, we tested a total of 210 coupled products per enzyme and generated a dataset of 5250 enzyme-substrate activity relationships. We identified multiple enzymes that had activity for synthesizing suberic acid and heptanoic acid (Aim 2). We built a random forest model for predicting the activity of each enzyme toward substrates on which it was not tested using the data from Aim 1. Finally, we evaluated growth defects that occurred due to expression of different carboligases in E. coli (Aim 3). We were able to identify specific metabolites and putative pathways that, when supplemented in the media, recovered the growth defect associated with the presence of specific carboligases. We are in the process of publishing two manuscript describing the methods for high-throughput screening of enzyme promiscuity, using machine learning to predict activity on untested substrates, and enzyme activity data we collected. This project has produced enabling data for biosynthesis of a range of new-to-nature compounds to support biomanufacturing.

60 APPLIED LIFE SCIENCES↗