Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Phaseolus vulgaris SUT1.1 is a high affinity sucrose–proton co–transporter

Plant sucrose transporters are required for phloem loading, and therefore are essential for plant growth and development. In common beans (Phaseolus vulgaris) there are only two sucrose transporters functionally characterized. Through a previous RNA-seq study, we identified a putative sucrose transporter in common bean, which we hypothesize to function in import of sucrose into plant cells. In silico analysis revealed that PvSUT1.1 is a putative sucrose-proton co-transporter distinct from other characterized sucrose transporters in common bean indicating that this is a previously undescribed transporter protein in beans. Further analysis revealed that PvSUT1.1 shares high protein sequence homology to the phloem loader Arabidopsis SUC2; both have 12 transmembrane domains, a typical characteristic of plant sucrose transporters. Heterologous expression in yeast further showed PvSUT1.1 to be functional and it imported sucrose into yeast cells with a K m of 0.7 mM sucrose. Import of sucrose through PvSUT1.1 is also pH-dependent with highest uptake at pH 4.0, and activity is lost in the presence of the uncoupler carbonyl cyanide 3-chlorophenylhydrazone. Consistent with identification of PvSUT1.1 as a Type I transporter, PvSUT1.1 also transports esculin. Finally, PvSUT1.1 showed expression in multiple tissues and the protein was localized to the plasma membrane. The results show that PvSUT1.1 is a sucrose transporter that is probably involved in the uptake of sucrose into source and sink cells. The potential role of PvSUT1.1 in leaf phloem loading of sucrose in common beans and its importance in heat tolerance of reproductive tissues are further discussed.

59 BASIC BIOLOGICAL SCIENCES↗

The MHC Associated Peptide Proteomics assay is a useful tool for the non-clinical assessment of immunogenicity

The propensity of therapeutic proteins to elicit an immune response, poses a significant challenge in clinical development and safety of the patients. Assessment of immunogenicity is crucial to predict potential adverse events and design safer biologics. In this study, we employed MHC Associated Peptide Proteomics (MAPPS) to comprehensively evaluate the immunogenic potential of re-engineered variants of immunogenic FVIIa analog (Vatreptacog Alfa). Our finding revealed the correlation between the protein sequence affinity for MHCII and the number of peptides identified in a MAPPS assay and this further correlates with the reduced T-cell responses. Moreover, MAPPS enable the identification of “relevant” T cell epitopes and may contribute to the development of biologics with lower immunogenic potential.

59 BASIC BIOLOGICAL SCIENCES↗

ATCUN-like Copper Site in βB2-Crystallin Plays a Protective Role in Cataract-Associated Aggregation

Cataract is the leading cause of blindness worldwide, and it is caused by crystallin damage and aggregation. Senile cataractous lenses have relatively high levels of metals, while some metal ions can directly induce the aggregation of human γ-crystallins. Here, for this work, we evaluated the impact of divalent metal ions in the aggregation of human βB2-crystallin, one of the most abundant crystallins in the lens. Turbidity assays showed that Pb 2+ , Hg 2+ , Cu 2+ , and Zn 2+ ions induce the aggregation of βB2-crystallin. Metal-induced aggregation is partially reverted by a chelating agent, indicating the formation of metal-bridged species. Our study focused on the mechanism of copper-induced aggregation of βB2-crystallin, finding that it involves metal-bridging, disulfide-bridging, and loss of protein stability. Circular dichroism and electron paramagnetic resonance (EPR) revealed the presence of at least three Cu 2+ binding sites in βB2-crystallin, one of them with spectroscopic features typical for Cu 2+ bound to an amino-terminal copper and nickel (ATCUN) binding motif, which is found in Cu transport proteins. The ATCUN-like Cu binding site is located at the unstructured N-terminus of βB2-crystallin, and it could be modeled by a peptide with the first six residues in the protein sequence (NH 2 -ASDHQF-). Isothermal titration calorimetry indicates a nanomolar Cu 2+ binding affinity for the ATCUN-like site. An N-truncated form of βB2-crystallin is more susceptible to Cu-induced aggregation and is less thermally stable, indicating a protective role for the ATCUN-like site. EPR and X-ray absorption spectroscopy studies reveal the presence of a copper redox active site in βB2-crystallin that is associated with metal-induced aggregation and formation of disulfide-bridged oligomers. Our study demonstrates metal-induced aggregation of βB2-crystallin and the presence of putative copper binding sites in the protein. Whether the copper-transport ATCUN-like site in βB2-crystallin plays a functional/protective role or constitutes a vestige from its evolution as a lens structural protein remains to be elucidated.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Engineered Membrane Vesicle Production via oprF or oprI Deletion Has Distinct Phenotypic Effects in Pseudomonas putida -putative knockouts table

Table S1, putative gene knockout targets in P. putida KT2440 to enhance vesiculation; Table S2, protein sequence identity of OmpA from E. coli K12 to P. putida KT2440 genes; Table S3, strains utilized in this study and corresponding construction details; Table S4, oligonucleotides utilized in this study; Table S5, plasmids utilized in this study; Table S6, sequences for mNeonGreen, tags, and codon-optimized genes; Figure S1, particle count per gCDW for KT2440 and knockout strains corresponding to data presented in Figure 1B; Figure S2, OD600 measurements of extracted MVs from KT2440 and knockout strains; Figure S3, particle count per gCDW for WT, ΔPP_4669, and ΔPP_1502; Figure S4, particle count per gCDW for KT2440 and knockout strains corresponding to data presented in Figure 3C; Figure S5, sizes of MVs corresponding to particle counts in Figure S4; Figure S6, particle count per gCDW for KT2440 grown on 20 mM glucose alone or 20 mM glucose plus 12.5 mM p-coumarate and 12.5 mM ferulate; Figure S7 and Figure S8, principal component analysis of the cellular fractions; Figure S9, heatmap of outer membrane proteins with differential abundance; and Figure S10, mNeonGreen (mNG) fluorescence signal for the cellular fraction and the extracellular fraction

hypervesiculation↗

End-to-end learning of multiple sequence alignments with differentiable Smith–Waterman

Abstract Motivation Multiple sequence alignments (MSAs) of homologous sequences contain information on structural and functional constraints and their evolutionary histories. Despite their importance for many downstream tasks, such as structure prediction, MSA generation is often treated as a separate pre-processing step, without any guidance from the application it will be used for. Results Here, we implement a smooth and differentiable version of the Smith–Waterman pairwise alignment algorithm that enables jointly learning an MSA and a downstream machine learning system in an end-to-end fashion. To demonstrate its utility, we introduce SMURF (Smooth Markov Unaligned Random Field), a new method that jointly learns an alignment and the parameters of a Markov Random Field for unsupervised contact prediction. We find that SMURF learns MSAs that mildly improve contact prediction on a diverse set of protein and RNA families. As a proof of concept, we demonstrate that by connecting our differentiable alignment module to AlphaFold2 and maximizing predicted confidence, we can learn MSAs that improve structure predictions over the initial MSAs. Interestingly, the alignments that improve AlphaFold predictions are self-inconsistent and can be viewed as adversarial. This work highlights the potential of differentiable dynamic programming to improve neural network pipelines that rely on an alignment and the potential dangers of optimizing predictions of protein sequences with methods that are not fully understood. Availability and implementation Our code and examples are available at: https://github.com/spetti/SMURF. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

A Multiplexed Quantitative Analysis of Germline Single Amino Acid Variants by Targeted Proteomics in Nondepleted Human Plasma

Single amino acid variants (SAAVs) in protein sequences are often a direct result of single-nucleotide polymorphisms (SNPs). Certain germline SAAVs have shown biological relevance in different disease conditions but lack precise quantification in circulation, which could hinder functional investigations and progress in biomarker development. Here, we have developed a multiplexed liquid chromatography-selected reaction monitoring (LC-SRM) assay that monitors 5 wild-type and variant peptide pairs (Complement Factor B: CFB-R32Q/R32W, Clusterin: CLU-N317H, Fetuin B: FETUB-K360R, and Kininogen: KNG1-L212P) in nondepleted human plasma. The assay was optimized for imprecision, linearity, stability, and calibration assessments with CVs of under 20%. The wild-type and variant peptide pairs were characterized in a set of healthy individual plasma samples. These target identifications were also validated by SNP genotyping with more than 99% accuracy. For all protein targets, we observed significantly lower concentrations of WT species in the presence variant peptides. In CFB, the concentration of R32Q was significantly lower than its counterpart R32W variant and WT species. Furthermore, our results distinguished phenotypes of homozygosity and heterozygosity of the SAAV presence through direct concentration level characterization. These findings provide some insights into how SAAVs affect quantitative assessments of target peptides. The assay demonstrates a platform for proteogenomic analyses with potential applications in both research and clinical settings.

genetics↗

Phylogenomics and the first higher taxonomy of Placozoa, an ancient and enigmatic animal phylum

Placozoa is an ancient phylum of extraordinarily unusual animals: miniscule, ameboid creatures that lack most fundamental animal features. Despite high genetic diversity, only recently have the second and third species been named. While prior genomic studies suffer from incomplete placozoan taxon sampling, we more than double the count with protein sequences from seven key genomes and produce the first nuclear phylogenomic reconstruction of all major placozoan lineages. This leads us to the first complete Linnaean taxonomic classification of Placozoa, over a century after its discovery: This may be the only time in the 21st century when an entire higher taxonomy for a whole animal phylum is formalized. Our classification establishes 2 new classes, 4 new orders, 3 new families, 1 new genus, and 1 new species, namely classes Polyplacotomia and Uniplacotomia; orders Polyplacotomea, Trichoplacea, Cladhexea, and Hoilungea; families Polyplacotomidae, Cladtertiidae, and Hoilungidae; and genus Cladtertia with species Cladtertia collaboinventa, nov. Our likelihood and gene content tree topologies refine the relationships determined in previous studies. Adding morphological data into our phylogenomic matrices suggests sponges (Porifera) as the sister to other animals, indicating that modest data addition shifts this node away from comb jellies (Ctenophora). Furthermore, by adding the first genomic protein data of the exceptionally distinct and branching Polyplacotoma mediterranea , we solidify its position as sister to all other placozoans; a divergence we estimate to be over 400 million years old. Yet even this deep split sits on a long branch to other animals, suggesting a bottleneck event followed by diversification. Ancestral state reconstructions indicate large shifts in gene content within Placozoa, with Hoilungia hongkongensis and its closest relatives having the most unique genetics.

54 ENVIRONMENTAL SCIENCES↗

UniKP: a unified framework for the prediction of enzyme kinetic parameters

Prediction of enzyme kinetic parameters is essential for designing and optimizing enzymes for various biotechnological and industrial applications, but the limited performance of current prediction tools on diverse tasks hinders their practical applications. Here, we introduce UniKP, a unified framework based on pretrained language models for the prediction of enzyme kinetic parameters, including enzyme turnover number (k cat ), Michaelis constant (K m ), and catalytic efficiency (k cat / K m ), from protein sequences and substrate structures. A two-layer framework derived from UniKP (EF-UniKP) has also been proposed to allow robust k cat prediction in considering environmental factors, including pH and temperature. In addition, four representative re-weighting methods are systematically explored to successfully reduce the prediction error in high-value prediction tasks. We have demonstrated the application of UniKP and EF-UniKP in several enzyme discovery and directed evolution tasks, leading to the identification of new enzymes and enzyme mutants with higher activity. UniKP is a valuable tool for deciphering the mechanisms of enzyme kinetics and enables novel insights into enzyme engineering and their industrial applications.

59 BASIC BIOLOGICAL SCIENCES↗

UnigeneFinder: An Automated Pipeline for Gene Calling From Transcriptome Assemblies Without a Reference Genome

ABSTRACT For most species, transcriptome data are much more readily available than genome data. Without a reference genome, gene calling is cumbersome and inaccurate because of the high degree of redundancy in de novo transcriptome assemblies. To simplify and increase the accuracy of de novo transcriptome assembly in the absence of a reference genome, we developed UnigeneFinder. Combining several clustering methods, UnigeneFinder substantially reduces the redundancy typical of raw transcriptome assemblies. This pipeline offers an effective solution to the problem of inflated transcript numbers, achieving a closer representation of the actual underlying genome. UnigeneFinder performs comparably or better, compared with existing tools, on plant species with varying genome complexities. UnigeneFinder is the only available transcriptome redundancy solution that fully automates the generation of primary transcript, coding region, and protein sequences, analogous to those available for high‐quality reference genomes. These features, coupled with the pipeline’s cross‐platform implementation, focus on automation, and an accessible, user‐friendly interface, make UnigeneFinder a useful tool for many downstream sequence‐based analyses in nonmodel organisms lacking a reference genome, including differential gene expression analysis, accurate ortholog identification, functional enrichments, and evolutionary analyses. UnigeneFinder also runs efficiently both on high‐performance computing (HPC) systems and personal computers, further reducing barriers to use.

Xue, Bo [Plant Resilience Institute Michigan State↗

Flying Blind, or Just Flying Under the Radar? The Underappreciated Power of De Novo Methods of Mass Spectrometric Peptide Identification

Mass spectrometry-based proteomics is a popular and powerful method for precise and highly multiplexed protein identification. The most common method of analyzing untargeted proteomics data is called database searching, where the database is simply a collection of protein sequences from the target organism, derived from genome sequencing. Experimental peptide tandem mass spectra are compared to simplified models of theoretical spectra calculated from the translated genomic sequences. However, in several interesting application areas, such as forensics, archaeology, venomics and others, a genome sequence may not be available, or the correct genome sequence to use is not known. In these cases, de novo peptide identification can play an important role. De novo peptide identification infers peptide sequence directly from the tandem mass spectrum without reference to a sequence database, usually using graph-based or machine learning algorithms. In this review, we provide a basic overview of de novo peptide identification methods and applications, briefly covering de novo algorithms and tools, and focusing in more depth on recent applications from venomics, metaproteomics, forensics, and characterization of antibody drugs.

proteomics, mass spectrometry, forensics, de novo ↗

How hydrophobicity, side chains, and salt affect the dimensions of disordered proteins

Abstract Despite the generally accepted role of the hydrophobic effect as the driving force for folding, many intrinsically disordered proteins (IDPs), including those with hydrophobic content typical of foldable proteins, behave nearly as self‐avoiding random walks (SARWs) under physiological conditions. Here, we tested how temperature and ionic conditions influence the dimensions of the N‐terminal domain of pertactin (PNt), an IDP with an amino acid composition typical of folded proteins. While PNt contracts somewhat with temperature, it nevertheless remains expanded over 10–58°C, with a Flory exponent, ν , >0.50. Both low and high ionic strength also produce contraction in PNt, but this contraction is mitigated by reducing charge segregation. With 46% glycine and low hydrophobicity, the reduced form of snow flea anti‐freeze protein (red‐sfAFP) is unaffected by temperature and ionic strength and persists as a near‐SARW, ν ~ 0.54, arguing that the thermal contraction of PNt is due to stronger interactions between hydrophobic side chains. Additionally, red‐sfAFP is a proxy for the polypeptide backbone, which has been thought to collapse in water. Increasing the glycine segregation in red‐sfAFP had minimal effect on ν . Water remained a good solvent even with 21 consecutive glycine residues ( ν > 0.5), and red‐sfAFP variants lacked stable backbone hydrogen bonds according to hydrogen exchange. Similarly, changing glycine segregation has little impact on ν in other glycine‐rich proteins. These findings underscore the generality that many disordered states can be expanded and unstructured, and that the hydrophobic effect alone is insufficient to drive significant chain collapse for typical protein sequences.

59 BASIC BIOLOGICAL SCIENCES↗

Different chemical scaffolds bind to L-phe site in Mycobacterium tuberculosis Phe-tRNA synthetase

Tuberculosis (TB), caused by Mycobacterium tuberculosis (Mt), is one of the deadliest infectious diseases. The rise of multidrug-resistant strains represents a major public health threat, requiring new therapeutic options. Bacterial aminoacyl-tRNA synthetases (aaRS) have been shown to be highly promising drug targets, including for TB treatment. These enzymes play an essential role in translating the DNA gene code into protein sequence by attaching specific amino acid to their cognate tRNAs. They have multiple binding sites that can be targeted for inhibitor discovery: amino acid binding pocket, ATP binding pocket, tRNA binding site and an editing domain. Recently we reported several high-resolution structures of M. tuberculosis phenylalanyl-tRNA synthetase (MtPheRS) complexed with tRNA Phe and either L-Phe or a nonhydrolyzable phenylalanine adenylate analog. Here, in this study, using Nucleic Magnetic Resonance (NMR) and Surface Plasmon Resonance (SPR) we identified fragments that bind to MtPheRS and we determined crystal structures of their complexes with MtPheRS/tRNA Phe . All the binders interact with the L-Phe amino acid binding site. The analysis of interactions of the new compounds combined with adenylate analog structure provides insights for the rational design of antituberculosis drugs. The 3 ' arm of the tRNA Phe in all the structures was disordered with exception of one complex with D-735 compound. In this structure the 3' CCA end of the acceptor stem is observed in the editing domain of MtPheRS providing insights regarding the post-transfer editing activity of class II aaRS.

Gade, Priyanka [Univ. of Chicago, IL (United State↗

Chromosome-level genome assembly of Quercus variabilis provides insights into the molecular mechanism of cork thickness

Quercus variabilis is a deciduous woody species with high ecological and economic value and is a major source of cork in East Asia. Cork from thick softwood sheets have higher commercial value than those from thin sheets. It is extremely difficult to genetically improve Q. variabilis to produce high quality softwood due to the lack of genomic information. Here, we present a high-quality chromosomal genome assembly for Q. variabilis with length of 791,89 Mb and 54,606 predicted genes. Comparative analysis of protein sequences of Q. variabilis with 11 other species revealed that specific and expanded gene families were significantly enriched in the "fatty acid biosynthesis" pathway in Q. variabilis, which may contribute to the formation of its unique cork. Additionally, based on weighted correlation network analysis of time-course (i.e., five important developmental ages) gene expression data in thick-cork versus thin-cork genotypes of Q. variabilis, we identified one co-expression gene module associated with the thick-cork trait. Within this co-expression gene module, 10 hub genes were associated with suberin biosynthesis. Furthermore, we identified a total of 198 suberin biosynthesis-related new candidate genes that were up-regulated in trees with a thick cork layer relative to those with a thin cork layer. Also, we found that some genes related to cell expansion and cell division were highly expressed in trees with a thick cork layer. Collectively, our results revealed that two metabolic pathways (i.e., suberin biosynthesis, fatty acid biosynthesis), along with other genes involved in cell expansion, cell division, and transcriptional regulation, were associated with the thick-cork trait in Q. variabilis, providing insights into the molecular basis of cork development and knowledge for informing genetic improvement of cork thickness in Q. variabilis and closely related species.

59 BASIC BIOLOGICAL SCIENCES↗

An in silico method to assess antibody fragment polyreactivity

Antibodies are essential biological research tools and important therapeutic agents, but some exhibit non-specific binding to off-target proteins and other biomolecules. Such polyreactive antibodies compromise screening pipelines, lead to incorrect and irreproducible experimental results, and are generally intractable for clinical development. Here, we design a set of experiments using a diverse naïve synthetic camelid antibody fragment (nanobody) library to enable machine learning models to accurately assess polyreactivity from protein sequence (AUC > 0.8). Moreover, our models provide quantitative scoring metrics that predict the effect of amino acid substitutions on polyreactivity. We experimentally test our models’ performance on three independent nanobody scaffolds, where over 90% of predicted substitutions successfully reduced polyreactivity. Importantly, the models allow us to diminish the polyreactivity of an angiotensin II type I receptor antagonist nanobody, without compromising its functional properties. We provide a companion web-server that offers a straightforward means of predicting polyreactivity and polyreactivity-reducing mutations for any given nanobody sequence.

60 APPLIED LIFE SCIENCES↗

Detecting macroevolutionary genotype–phenotype associations using error-corrected rates of protein convergence

On macroevolutionary timescales, extensive mutations and phylogenetic uncertainty mask the signals of genotype–phenotype associations underlying convergent evolution. To overcome this problem, we extended the widely used framework of non-synonymous to synonymous substitution rate ratios and developed the novel metric ω C , which measures the error-corrected convergence rate of protein evolution. While ω C distinguishes natural selection from genetic noise and phylogenetic errors in simulation and real examples, its accuracy allows an exploratory genome-wide search of adaptive molecular convergence without phenotypic hypothesis or candidate genes. Using gene expression data, we explored over 20 million branch combinations in vertebrate genes and identified the joint convergence of expression patterns and protein sequences with amino acid substitutions in functionally important sites, providing hypotheses on undiscovered phenotypes. We further extended our method with a heuristic algorithm to detect highly repetitive convergence among computationally non-trivial higher-order phylogenetic combinations. Our approach allows bidirectional searches for genotype–phenotype associations, even in lineages that diverged for hundreds of millions of years.

59 BASIC BIOLOGICAL SCIENCES↗

Three mutations repurpose a plant karrikin receptor to a strigolactone receptor

Significance Parasitic plants like witchweed cause huge losses in crop yield in Africa. A key part to the success of witchweed is to start its life cycle upon sensing small molecules called strigolactones, which are exuded from roots of host plants into the soil. Witchweed sense host-derived strigolactones through receptors called HTLs. It is thought that the evolutionary origin of HTLs is a receptor called KAI2 in nonparasitic plants, which can respond to different small molecules such as karrikins. By making three changes in the protein sequence of KAI2, this hybrid receptor can now sense both strigolactones and karrikins. These results help in understanding how receptors can evolve to sense different signals and can lead to solutions for combating pesky witchweed.

Arellano-Saab, Amir↗

Machine learning predicts new anti-CRISPR proteins

The increasing use of CRISPR–Cas9 in medicine, agriculture, and synthetic biology has accelerated the drive to discover new CRISPR–Cas inhibitors as potential mechanisms of control for gene editing applications. Many anti-CRISPRs have been found that inhibit the CRISPR–Cas adaptive immune system. However, comparing all currently known anti-CRISPRs does not reveal a shared set of properties for facile bioinformatic identification of new anti-CRISPR families. Here, we describe AcRanker, a machine learning based method to aid direct identification of new potential anti-CRISPRs using only protein sequence information. Using a training set of known anti-CRISPRs, we built a model based on XGBoost ranking. We then applied AcRanker to predict candidate anti-CRISPRs from predicted prophage regions within self-targeting bacterial genomes and discovered two previously unknown anti-CRISPRs: AcrllA20 (ML1) and AcrIIA21 (ML8). We show that AcrIIA20 strongly inhibits Streptococcus iniae Cas9 (SinCas9) and weakly inhibits Streptococcus pyogenes Cas9 (SpyCas9). We also show that AcrIIA21 inhibits SpyCas9, Streptococcus aureus Cas9 (SauCas9) and SinCas9 with low potency. The addition of AcRanker to the anti-CRISPR discovery toolkit allows researchers to directly rank potential anti-CRISPR candidate genes for increased speed in testing and validation of new anti-CRISPRs. A web server implementation for AcRanker is available online at http://acranker.pythonanywhere.com/.

59 BASIC BIOLOGICAL SCIENCES↗