Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Orthogonal glycolytic pathway enables directed evolution of noncanonical cofactor oxidase

Abstract Noncanonical cofactor biomimetics (NCBs) such as nicotinamide mononucleotide (NMN + ) provide enhanced scalability for biomanufacturing. However, engineering enzymes to accept NCBs is difficult. Here, we establish a growth selection platform to evolve enzymes to utilize NMN + -based reducing power. This is based on an orthogonal, NMN + -dependent glycolytic pathway in Escherichia coli which can be coupled to any reciprocal enzyme to recycle the ensuing reduced NMN + . With a throughput of >10 6 variants per iteration, the growth selection discovers a Lactobacillus pentosus NADH oxidase variant with ~10-fold increase in NMNH catalytic efficiency and enhanced activity for other NCBs. Molecular modeling and experimental validation suggest that instead of directly contacting NCBs, the mutations optimize the enzyme’s global conformational dynamics to resemble the WT with the native cofactor bound. Restoring the enzyme’s access to catalytically competent conformation states via deep navigation of protein sequence space with high-throughput evolution provides a universal route to engineer NCB-dependent enzymes.

59 BASIC BIOLOGICAL SCIENCES↗

Mycobacterium tuberculosis Phe-tRNA synthetase: structural insights into tRNA recognition and aminoacylation

Abstract Tuberculosis, caused by Mycobacterium tuberculosis, responsible for ∼1.5 million fatalities in 2018, is the deadliest infectious disease. Global spread of multidrug resistant strains is a public health threat, requiring new treatments. Aminoacyl-tRNA synthetases are plausible candidates as potential drug targets, because they play an essential role in translating the DNA code into protein sequence by attaching a specific amino acid to their cognate tRNAs. We report structures of M. tuberculosis Phe-tRNA synthetase complexed with an unmodified tRNAPhe transcript and either L-Phe or a nonhydrolyzable phenylalanine adenylate analog. High-resolution models reveal details of two modes of tRNA interaction with the enzyme: an initial recognition via indirect readout of anticodon stem-loop and aminoacylation ready state involving interactions of the 3′ end of tRNAPhe with the adenylate site. For the first time, we observe the protein gate controlling access to the active site and detailed geometry of the acyl donor and tRNA acceptor consistent with accepted mechanism. We biochemically validated the inhibitory potency of the adenylate analog and provide the most complete view of the Phe-tRNA synthetase/tRNAPhe system to date. The presented topography of amino adenylate-binding and editing sites at different stages of tRNA binding to the enzyme provide insights for the rational design of anti-tuberculosis drugs.

59 BASIC BIOLOGICAL SCIENCES↗

A molecular timescale of eukaryote evolution and the rise of complex multicellular life

BACKGROUND: The pattern and timing of the rise in complex multicellular life during Earth's history has not been established. Great disparity persists between the pattern suggested by the fossil record and that estimated by molecular clocks, especially for plants, animals, fungi, and the deepest branches of the eukaryote tree. Here, we used all available protein sequence data and molecular clock methods to place constraints on the increase in complexity through time. RESULTS: Our phylogenetic analyses revealed that (i) animals are more closely related to fungi than to plants, (ii) red algae are closer to plants than to animals or fungi, (iii) choanoflagellates are closer to animals than to fungi or plants, (iv) diplomonads, euglenozoans, and alveolates each are basal to plants+animals+fungi, and (v) diplomonads are basal to other eukaryotes (including alveolates and euglenozoans). Divergence times were estimated from global and local clock methods using 20-188 proteins per node, with data treated separately (multigene) and concatenated (supergene). Different time estimation methods yielded similar results (within 5%): vertebrate-arthropod (964 million years ago, Ma), Cnidaria-Bilateria (1,298 Ma), Porifera-Eumetozoa (1,351 Ma), Pyrenomycetes-Plectomycetes (551 Ma), Candida-Saccharomyces (723 Ma), Hemiascomycetes-filamentous Ascomycota (982 Ma), Basidiomycota-Ascomycota (968 Ma), Mucorales-Basidiomycota (947 Ma), Fungi-Animalia (1,513 Ma), mosses-vascular plants (707 Ma), Chlorophyta-Tracheophyta (968 Ma), Rhodophyta-Chlorophyta+Embryophyta (1,428 Ma), Plantae-Animalia (1,609 Ma), Alveolata-plants+animals+fungi (1,973 Ma), Euglenozoa-plants+animals+fungi (1,961 Ma), and Giardia-plants+animals+fungi (2,309 Ma). By extrapolation, mitochondria arose approximately 2300-1800 Ma and plastids arose 1600-1500 Ma. Estimates of the maximum number of cell types of common ancestors, combined with divergence times, showed an increase from two cell types at 2500 Ma to approximately 10 types at 1500 Ma and 50 cell types at approximately 1000 Ma. CONCLUSIONS: The results suggest that oxygen levels in the environment, and the ability of eukaryotes to extract energy from oxygen, as well as produce oxygen, were key factors in the rise of complex multicellular life. Mitochondria and organisms with more than 2-3 cell types appeared soon after the initial increase in oxygen levels at 2300 Ma. The addition of plastids at 1500 Ma, allowing eukaryotes to produce oxygen, preceded the major rise in complexity.

Evolution, Molecular↗

SPARC: Structural properties associated with residue constraints

SPARC facilitates the generation of plausible hypotheses regarding underlying biochemical mechanisms by structurally characterizing protein sequence constraints. Such constraints appear as residues co-conserved in functionally related subgroups, as subtle pairwise correlations (i.e., direct couplings), and as correlations among these sequence features or with structural features. SPARC performs three types of analyses. First, based on pairwise sequence correlations, it estimates the biological relevance of alternative conformations and of homomeric contacts, as illustrated here for death domains. Second, it estimates the statistical significance of the correspondence between directly coupled residue pairs and interactions at heterodimeric interfaces. Third, given molecular dynamics simulated structures, it characterizes interactions among constrained residues or between such residues and ligands that: (a) are stably maintained during the simulation; (b) undergo correlated formation and/or disruption of interactions with other constrained residues; or (c) switch between alternative interactions. We illustrate this for two homohexameric complexes: the bacterial enhancer binding protein (bEBP) NtrC1, which activates transcription by remodeling RNA polymerase (RNAP) containing σ 54 , and for DnaB helicase, which opens DNA at the bacterial replication fork. Based on the NtrC1 analysis, we hypothesize possible mechanisms for inhibiting ATP hydrolysis until ADP is released from an adjacent subunit and for coupling ATP hydrolysis to restructuring of σ 54 binding loops. Based on the DnaB analysis, we hypothesize that DnaB ‘grabs’ ssDNA by flipping every fourth base and inserting it into cavities between subunits and that flipping of a DnaB-specific glutamine residue triggers ATP hydrolysis.

97 MATHEMATICS AND COMPUTING↗

The MHC Associated Peptide Proteomics assay is a useful tool for the non-clinical assessment of immunogenicity

The propensity of therapeutic proteins to elicit an immune response, poses a significant challenge in clinical development and safety of the patients. Assessment of immunogenicity is crucial to predict potential adverse events and design safer biologics. In this study, we employed MHC Associated Peptide Proteomics (MAPPS) to comprehensively evaluate the immunogenic potential of re-engineered variants of immunogenic FVIIa analog (Vatreptacog Alfa). Our finding revealed the correlation between the protein sequence affinity for MHCII and the number of peptides identified in a MAPPS assay and this further correlates with the reduced T-cell responses. Moreover, MAPPS enable the identification of “relevant” T cell epitopes and may contribute to the development of biologics with lower immunogenic potential.

59 BASIC BIOLOGICAL SCIENCES↗

ATCUN-like Copper Site in βB2-Crystallin Plays a Protective Role in Cataract-Associated Aggregation

Cataract is the leading cause of blindness worldwide, and it is caused by crystallin damage and aggregation. Senile cataractous lenses have relatively high levels of metals, while some metal ions can directly induce the aggregation of human γ-crystallins. Here, for this work, we evaluated the impact of divalent metal ions in the aggregation of human βB2-crystallin, one of the most abundant crystallins in the lens. Turbidity assays showed that Pb 2+ , Hg 2+ , Cu 2+ , and Zn 2+ ions induce the aggregation of βB2-crystallin. Metal-induced aggregation is partially reverted by a chelating agent, indicating the formation of metal-bridged species. Our study focused on the mechanism of copper-induced aggregation of βB2-crystallin, finding that it involves metal-bridging, disulfide-bridging, and loss of protein stability. Circular dichroism and electron paramagnetic resonance (EPR) revealed the presence of at least three Cu 2+ binding sites in βB2-crystallin, one of them with spectroscopic features typical for Cu 2+ bound to an amino-terminal copper and nickel (ATCUN) binding motif, which is found in Cu transport proteins. The ATCUN-like Cu binding site is located at the unstructured N-terminus of βB2-crystallin, and it could be modeled by a peptide with the first six residues in the protein sequence (NH 2 -ASDHQF-). Isothermal titration calorimetry indicates a nanomolar Cu 2+ binding affinity for the ATCUN-like site. An N-truncated form of βB2-crystallin is more susceptible to Cu-induced aggregation and is less thermally stable, indicating a protective role for the ATCUN-like site. EPR and X-ray absorption spectroscopy studies reveal the presence of a copper redox active site in βB2-crystallin that is associated with metal-induced aggregation and formation of disulfide-bridged oligomers. Our study demonstrates metal-induced aggregation of βB2-crystallin and the presence of putative copper binding sites in the protein. Whether the copper-transport ATCUN-like site in βB2-crystallin plays a functional/protective role or constitutes a vestige from its evolution as a lens structural protein remains to be elucidated.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An intron within the 16S ribosomal RNA gene of the archaeon Pyrobaculum aerophilum

The 16S rRNA genes of Pyrobaculum aerophilum and Pyrobaculum islandicum were amplified by the polymerase chain reaction, and the resulting products were sequenced directly. The two organisms are closely related by this measure (over 98% similar). However, they differ in that the (lone) 16S rRNA gene of Pyrobaculum aerophilum contains a 713-bp intron not seen in the corresponding gene of Pyrobaculum islandicum. To our knowledge, this is the only intron so far reported in the small subunit rRNA gene of a prokaryote. Upon excision the intron is circularized. A secondary structure model of the intron-containing rRNA suggests a splicing mechanism of the same type as that invoked for the tRNA introns of the Archaea and Eucarya and 23S rRNAs of the Archaea. The intron contains an open reading frame whose protein translation shows no certain homology with any known protein sequence.

NASA Discipline Exobiology↗

Engineered Membrane Vesicle Production via oprF or oprI Deletion Has Distinct Phenotypic Effects in Pseudomonas putida -putative knockouts table

Table S1, putative gene knockout targets in P. putida KT2440 to enhance vesiculation; Table S2, protein sequence identity of OmpA from E. coli K12 to P. putida KT2440 genes; Table S3, strains utilized in this study and corresponding construction details; Table S4, oligonucleotides utilized in this study; Table S5, plasmids utilized in this study; Table S6, sequences for mNeonGreen, tags, and codon-optimized genes; Figure S1, particle count per gCDW for KT2440 and knockout strains corresponding to data presented in Figure 1B; Figure S2, OD600 measurements of extracted MVs from KT2440 and knockout strains; Figure S3, particle count per gCDW for WT, ΔPP_4669, and ΔPP_1502; Figure S4, particle count per gCDW for KT2440 and knockout strains corresponding to data presented in Figure 3C; Figure S5, sizes of MVs corresponding to particle counts in Figure S4; Figure S6, particle count per gCDW for KT2440 grown on 20 mM glucose alone or 20 mM glucose plus 12.5 mM p-coumarate and 12.5 mM ferulate; Figure S7 and Figure S8, principal component analysis of the cellular fractions; Figure S9, heatmap of outer membrane proteins with differential abundance; and Figure S10, mNeonGreen (mNG) fluorescence signal for the cellular fraction and the extracellular fraction

hypervesiculation↗

End-to-end learning of multiple sequence alignments with differentiable Smith–Waterman

Abstract Motivation Multiple sequence alignments (MSAs) of homologous sequences contain information on structural and functional constraints and their evolutionary histories. Despite their importance for many downstream tasks, such as structure prediction, MSA generation is often treated as a separate pre-processing step, without any guidance from the application it will be used for. Results Here, we implement a smooth and differentiable version of the Smith–Waterman pairwise alignment algorithm that enables jointly learning an MSA and a downstream machine learning system in an end-to-end fashion. To demonstrate its utility, we introduce SMURF (Smooth Markov Unaligned Random Field), a new method that jointly learns an alignment and the parameters of a Markov Random Field for unsupervised contact prediction. We find that SMURF learns MSAs that mildly improve contact prediction on a diverse set of protein and RNA families. As a proof of concept, we demonstrate that by connecting our differentiable alignment module to AlphaFold2 and maximizing predicted confidence, we can learn MSAs that improve structure predictions over the initial MSAs. Interestingly, the alignments that improve AlphaFold predictions are self-inconsistent and can be viewed as adversarial. This work highlights the potential of differentiable dynamic programming to improve neural network pipelines that rely on an alignment and the potential dangers of optimizing predictions of protein sequences with methods that are not fully understood. Availability and implementation Our code and examples are available at: https://github.com/spetti/SMURF. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

A Multiplexed Quantitative Analysis of Germline Single Amino Acid Variants by Targeted Proteomics in Nondepleted Human Plasma

Single amino acid variants (SAAVs) in protein sequences are often a direct result of single-nucleotide polymorphisms (SNPs). Certain germline SAAVs have shown biological relevance in different disease conditions but lack precise quantification in circulation, which could hinder functional investigations and progress in biomarker development. Here, we have developed a multiplexed liquid chromatography-selected reaction monitoring (LC-SRM) assay that monitors 5 wild-type and variant peptide pairs (Complement Factor B: CFB-R32Q/R32W, Clusterin: CLU-N317H, Fetuin B: FETUB-K360R, and Kininogen: KNG1-L212P) in nondepleted human plasma. The assay was optimized for imprecision, linearity, stability, and calibration assessments with CVs of under 20%. The wild-type and variant peptide pairs were characterized in a set of healthy individual plasma samples. These target identifications were also validated by SNP genotyping with more than 99% accuracy. For all protein targets, we observed significantly lower concentrations of WT species in the presence variant peptides. In CFB, the concentration of R32Q was significantly lower than its counterpart R32W variant and WT species. Furthermore, our results distinguished phenotypes of homozygosity and heterozygosity of the SAAV presence through direct concentration level characterization. These findings provide some insights into how SAAVs affect quantitative assessments of target peptides. The assay demonstrates a platform for proteogenomic analyses with potential applications in both research and clinical settings.

genetics↗

Phylogenomics and the first higher taxonomy of Placozoa, an ancient and enigmatic animal phylum

Placozoa is an ancient phylum of extraordinarily unusual animals: miniscule, ameboid creatures that lack most fundamental animal features. Despite high genetic diversity, only recently have the second and third species been named. While prior genomic studies suffer from incomplete placozoan taxon sampling, we more than double the count with protein sequences from seven key genomes and produce the first nuclear phylogenomic reconstruction of all major placozoan lineages. This leads us to the first complete Linnaean taxonomic classification of Placozoa, over a century after its discovery: This may be the only time in the 21st century when an entire higher taxonomy for a whole animal phylum is formalized. Our classification establishes 2 new classes, 4 new orders, 3 new families, 1 new genus, and 1 new species, namely classes Polyplacotomia and Uniplacotomia; orders Polyplacotomea, Trichoplacea, Cladhexea, and Hoilungea; families Polyplacotomidae, Cladtertiidae, and Hoilungidae; and genus Cladtertia with species Cladtertia collaboinventa, nov. Our likelihood and gene content tree topologies refine the relationships determined in previous studies. Adding morphological data into our phylogenomic matrices suggests sponges (Porifera) as the sister to other animals, indicating that modest data addition shifts this node away from comb jellies (Ctenophora). Furthermore, by adding the first genomic protein data of the exceptionally distinct and branching Polyplacotoma mediterranea , we solidify its position as sister to all other placozoans; a divergence we estimate to be over 400 million years old. Yet even this deep split sits on a long branch to other animals, suggesting a bottleneck event followed by diversification. Ancestral state reconstructions indicate large shifts in gene content within Placozoa, with Hoilungia hongkongensis and its closest relatives having the most unique genetics.

54 ENVIRONMENTAL SCIENCES↗

UniKP: a unified framework for the prediction of enzyme kinetic parameters

Prediction of enzyme kinetic parameters is essential for designing and optimizing enzymes for various biotechnological and industrial applications, but the limited performance of current prediction tools on diverse tasks hinders their practical applications. Here, we introduce UniKP, a unified framework based on pretrained language models for the prediction of enzyme kinetic parameters, including enzyme turnover number (k cat ), Michaelis constant (K m ), and catalytic efficiency (k cat / K m ), from protein sequences and substrate structures. A two-layer framework derived from UniKP (EF-UniKP) has also been proposed to allow robust k cat prediction in considering environmental factors, including pH and temperature. In addition, four representative re-weighting methods are systematically explored to successfully reduce the prediction error in high-value prediction tasks. We have demonstrated the application of UniKP and EF-UniKP in several enzyme discovery and directed evolution tasks, leading to the identification of new enzymes and enzyme mutants with higher activity. UniKP is a valuable tool for deciphering the mechanisms of enzyme kinetics and enables novel insights into enzyme engineering and their industrial applications.

59 BASIC BIOLOGICAL SCIENCES↗

UnigeneFinder: An Automated Pipeline for Gene Calling From Transcriptome Assemblies Without a Reference Genome

ABSTRACT For most species, transcriptome data are much more readily available than genome data. Without a reference genome, gene calling is cumbersome and inaccurate because of the high degree of redundancy in de novo transcriptome assemblies. To simplify and increase the accuracy of de novo transcriptome assembly in the absence of a reference genome, we developed UnigeneFinder. Combining several clustering methods, UnigeneFinder substantially reduces the redundancy typical of raw transcriptome assemblies. This pipeline offers an effective solution to the problem of inflated transcript numbers, achieving a closer representation of the actual underlying genome. UnigeneFinder performs comparably or better, compared with existing tools, on plant species with varying genome complexities. UnigeneFinder is the only available transcriptome redundancy solution that fully automates the generation of primary transcript, coding region, and protein sequences, analogous to those available for high‐quality reference genomes. These features, coupled with the pipeline’s cross‐platform implementation, focus on automation, and an accessible, user‐friendly interface, make UnigeneFinder a useful tool for many downstream sequence‐based analyses in nonmodel organisms lacking a reference genome, including differential gene expression analysis, accurate ortholog identification, functional enrichments, and evolutionary analyses. UnigeneFinder also runs efficiently both on high‐performance computing (HPC) systems and personal computers, further reducing barriers to use.

Xue, Bo [Plant Resilience Institute Michigan State↗

How hydrophobicity, side chains, and salt affect the dimensions of disordered proteins

Abstract Despite the generally accepted role of the hydrophobic effect as the driving force for folding, many intrinsically disordered proteins (IDPs), including those with hydrophobic content typical of foldable proteins, behave nearly as self‐avoiding random walks (SARWs) under physiological conditions. Here, we tested how temperature and ionic conditions influence the dimensions of the N‐terminal domain of pertactin (PNt), an IDP with an amino acid composition typical of folded proteins. While PNt contracts somewhat with temperature, it nevertheless remains expanded over 10–58°C, with a Flory exponent, ν , >0.50. Both low and high ionic strength also produce contraction in PNt, but this contraction is mitigated by reducing charge segregation. With 46% glycine and low hydrophobicity, the reduced form of snow flea anti‐freeze protein (red‐sfAFP) is unaffected by temperature and ionic strength and persists as a near‐SARW, ν ~ 0.54, arguing that the thermal contraction of PNt is due to stronger interactions between hydrophobic side chains. Additionally, red‐sfAFP is a proxy for the polypeptide backbone, which has been thought to collapse in water. Increasing the glycine segregation in red‐sfAFP had minimal effect on ν . Water remained a good solvent even with 21 consecutive glycine residues ( ν > 0.5), and red‐sfAFP variants lacked stable backbone hydrogen bonds according to hydrogen exchange. Similarly, changing glycine segregation has little impact on ν in other glycine‐rich proteins. These findings underscore the generality that many disordered states can be expanded and unstructured, and that the hydrophobic effect alone is insufficient to drive significant chain collapse for typical protein sequences.

59 BASIC BIOLOGICAL SCIENCES↗

Different chemical scaffolds bind to L-phe site in Mycobacterium tuberculosis Phe-tRNA synthetase

Tuberculosis (TB), caused by Mycobacterium tuberculosis (Mt), is one of the deadliest infectious diseases. The rise of multidrug-resistant strains represents a major public health threat, requiring new therapeutic options. Bacterial aminoacyl-tRNA synthetases (aaRS) have been shown to be highly promising drug targets, including for TB treatment. These enzymes play an essential role in translating the DNA gene code into protein sequence by attaching specific amino acid to their cognate tRNAs. They have multiple binding sites that can be targeted for inhibitor discovery: amino acid binding pocket, ATP binding pocket, tRNA binding site and an editing domain. Recently we reported several high-resolution structures of M. tuberculosis phenylalanyl-tRNA synthetase (MtPheRS) complexed with tRNA Phe and either L-Phe or a nonhydrolyzable phenylalanine adenylate analog. Here, in this study, using Nucleic Magnetic Resonance (NMR) and Surface Plasmon Resonance (SPR) we identified fragments that bind to MtPheRS and we determined crystal structures of their complexes with MtPheRS/tRNA Phe . All the binders interact with the L-Phe amino acid binding site. The analysis of interactions of the new compounds combined with adenylate analog structure provides insights for the rational design of antituberculosis drugs. The 3 ' arm of the tRNA Phe in all the structures was disordered with exception of one complex with D-735 compound. In this structure the 3' CCA end of the acceptor stem is observed in the editing domain of MtPheRS providing insights regarding the post-transfer editing activity of class II aaRS.

Gade, Priyanka [Univ. of Chicago, IL (United State↗

Chromosome-level genome assembly of Quercus variabilis provides insights into the molecular mechanism of cork thickness

Quercus variabilis is a deciduous woody species with high ecological and economic value and is a major source of cork in East Asia. Cork from thick softwood sheets have higher commercial value than those from thin sheets. It is extremely difficult to genetically improve Q. variabilis to produce high quality softwood due to the lack of genomic information. Here, we present a high-quality chromosomal genome assembly for Q. variabilis with length of 791,89 Mb and 54,606 predicted genes. Comparative analysis of protein sequences of Q. variabilis with 11 other species revealed that specific and expanded gene families were significantly enriched in the "fatty acid biosynthesis" pathway in Q. variabilis, which may contribute to the formation of its unique cork. Additionally, based on weighted correlation network analysis of time-course (i.e., five important developmental ages) gene expression data in thick-cork versus thin-cork genotypes of Q. variabilis, we identified one co-expression gene module associated with the thick-cork trait. Within this co-expression gene module, 10 hub genes were associated with suberin biosynthesis. Furthermore, we identified a total of 198 suberin biosynthesis-related new candidate genes that were up-regulated in trees with a thick cork layer relative to those with a thin cork layer. Also, we found that some genes related to cell expansion and cell division were highly expressed in trees with a thick cork layer. Collectively, our results revealed that two metabolic pathways (i.e., suberin biosynthesis, fatty acid biosynthesis), along with other genes involved in cell expansion, cell division, and transcriptional regulation, were associated with the thick-cork trait in Q. variabilis, providing insights into the molecular basis of cork development and knowledge for informing genetic improvement of cork thickness in Q. variabilis and closely related species.

59 BASIC BIOLOGICAL SCIENCES↗

An in silico method to assess antibody fragment polyreactivity

Antibodies are essential biological research tools and important therapeutic agents, but some exhibit non-specific binding to off-target proteins and other biomolecules. Such polyreactive antibodies compromise screening pipelines, lead to incorrect and irreproducible experimental results, and are generally intractable for clinical development. Here, we design a set of experiments using a diverse naïve synthetic camelid antibody fragment (nanobody) library to enable machine learning models to accurately assess polyreactivity from protein sequence (AUC > 0.8). Moreover, our models provide quantitative scoring metrics that predict the effect of amino acid substitutions on polyreactivity. We experimentally test our models’ performance on three independent nanobody scaffolds, where over 90% of predicted substitutions successfully reduced polyreactivity. Importantly, the models allow us to diminish the polyreactivity of an angiotensin II type I receptor antagonist nanobody, without compromising its functional properties. We provide a companion web-server that offers a straightforward means of predicting polyreactivity and polyreactivity-reducing mutations for any given nanobody sequence.

60 APPLIED LIFE SCIENCES↗

Detecting macroevolutionary genotype–phenotype associations using error-corrected rates of protein convergence

On macroevolutionary timescales, extensive mutations and phylogenetic uncertainty mask the signals of genotype–phenotype associations underlying convergent evolution. To overcome this problem, we extended the widely used framework of non-synonymous to synonymous substitution rate ratios and developed the novel metric ω C , which measures the error-corrected convergence rate of protein evolution. While ω C distinguishes natural selection from genetic noise and phylogenetic errors in simulation and real examples, its accuracy allows an exploratory genome-wide search of adaptive molecular convergence without phenotypic hypothesis or candidate genes. Using gene expression data, we explored over 20 million branch combinations in vertebrate genes and identified the joint convergence of expression patterns and protein sequences with amino acid substitutions in functionally important sites, providing hypotheses on undiscovered phenotypes. We further extended our method with a heuristic algorithm to detect highly repetitive convergence among computationally non-trivial higher-order phylogenetic combinations. Our approach allows bidirectional searches for genotype–phenotype associations, even in lineages that diverged for hundreds of millions of years.

59 BASIC BIOLOGICAL SCIENCES↗