Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Zwartia hollandica gen. nov., sp. nov., Jezberella montanilacus gen. nov., sp. nov. and Sheuella amnicola gen. nov., comb. nov., representing the environmental GKS98 (betIII) cluster

We present two strains affiliated with the GKS98 cluster. This phylogenetically defined cluster is representing abundant, mainly uncultured freshwater bacteria, which were observed by many cultivation-independent studies on the diversity of bacteria in various freshwater lakes and streams. Bacteria affiliated with the GKS98 cluster were detected by cultivation-independent methods in freshwater systems located in Europe, Asia, Africa and the Americas. The two strains, LF4-65 T (=CCUG 56422 T =DSM 107630 T ) and MWH-P2sevCIIIb T (=CCUG 56420 T =DSM 107629 T ), are aerobic chemoorganotrophs, both with genome sizes of 3.2 Mbp and G+C values of 52.4 and 51.0 mol%, respectively. Phylogenomic analyses based on concatenated amino acid sequences of 120 proteins suggest an affiliation of the two strains with the family Alcaligenaceae and revealed Orrella amnicola and Orrella marina (= Algicoccus marinus ) as being the closest related, previously described species. However, the calculated phylogenomic trees clearly suggest that the current genus Orrella represents a polyphyletic taxon. Based on the branching order in the phylogenomic trees, as well as the revealed phylogenetic distances and chemotaxonomic traits, we propose to establish the new genus Zwartia gen. nov. and the new species Z. hollandica sp. nov. to harbour strain LF4-65 T and the new genus Jezberella gen. nov. and the new species J. montanilacus sp. nov. to harbour strain MWH-P2sevCIIIb T . Furthermore, we propose the reclassification of the species Orrella amnicola in the new genus Sheuella gen. nov. The new genera Zwartia, Jezberella and Sheuella together represent taxonomically the GKS98 cluster.

Microbiology↗

Protein identification from electron cryomicroscopy maps by automated model building and side-chain matching

Using single-particle electron cryo-microscopy (cryo-EM), it is possible to obtain multiple reconstructions showing the 3D structures of proteins imaged as a mixture. Here, it is shown that automatic map interpretation based on such reconstructions can be used to create atomic models of proteins as well as to match the proteins to the correct sequences and thereby to identify them. This procedure was tested using two proteins previously identified from a mixture at resolutions of 3.2 Å, as well as using 91 deposited maps with resolutions between 2 and 4.5 Å. The approach is found to be highly effective for maps obtained at resolutions of 3.5 Å and better, and to have some utility at resolutions as low as 4 Å.

59 BASIC BIOLOGICAL SCIENCES↗

OrthoPhylo

This software builds on PHAME developed at LANL to generate phylogenetic trees of bacterial whole genome sequences. Where PHAME uses whole genome alignments to generate informative sites to base tress on, PHAME-OuS annotates bacterial genes, identifies orthologous sequences, aligns related proteins, uses those alignments to inform transcript alignments, then builds trees with several methods. The first is a conventional gene concatenation and ML tree estemation method. The second attempts to reconcile gene tree with a unified species tree using quartets (ASTRAL). Both methods allow filtering of gene lists on number of species represented, length, and gappiness in order to tune noise-to-signal for tree estimation

Middlebrook, Earl↗

Rapid Computational Identification of Therapeutic Targets for Pathogens

Biological threats continue to persist and evolve as an important challenge to national security. There are multiple ways in which novel viral pathogens could emerge to pose a serious threat to human health. This project developed a pathogen target identification tool that can rapidly respond to a novel or emerging viral biological threat. A set of computational tools were developed that provide detailed information on the newly sequenced genes, their protein products and the drug target sites for the proteins that are best suited for biological countermeasure development. Three key innovations were developed in the project. 1) Development of a new extensive database of protein pocket structures with structure-based search algorithms to rapidly link novel protein targets with the complete collection of previously experimentally solved protein structures. 2) A novel clustering pipeline was introduced to group matching structures and associated small-molecule binding ligands into a consensus protein pocket with the associated small-molecule chemotypes predicted to fit in the pocket site. The matching experimentally solved structures were used to inform the value of different target sites. 3) Where there are viral protein targets with pockets structurally matched to similar human proteins, a biological knowledge graph, which links molecular interactions with human disease, was used to further assess the potential negative impact of a viral protein target with similarities to human proteins that could have important off target side effects. In total, the project produced a new resource for rapid and detailed assessment of promising targets for countermeasures, reflecting the ongoing wet lab, clinical, and computational data being collected. These capabilities will improve the ability to respond to a biological threat in multiple domains.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning guided selection of broad-spectrum epitope-specific functional antibodies for "Disease X"

Our project established and demonstrated a transfer learning framework that enables prediction of antibody–antigen interactions across related viruses. The approach focused on three major activities: 1. Conserved region and epitope identification – We compared viral protein structures and sequences to identify shared receptor-binding domains and neutralizing epitope regions across variants and related viruses. These conserved features formed the foundation for discovering broadly functional antibodies. 2. Machine learning model development – We built neural network–based models that integrate epitope features with antibody sequence information. Instead of relying solely on structural or physical properties, the models learned transferable patterns that describe antibody binding potential across different viral families. 3. Transfer learning and validation – Using SARS-CoV-2 and Ebola as source systems, we successfully transferred learned epitope features to predict antibody interactions for SARS CoV-1 and Marburg virus. Iterative cycles of dataset generation, retraining, and evaluation improved generalization and predictive power, ensuring the framework can adapt to new threats.

59 BASIC BIOLOGICAL SCIENCES↗

Ion Binding Properties of a Naturally Occurring Metalloantibody

LT1009 is a humanized version of murine LT1002 IgG1 that employs two bridging Ca2+ ions to bind its antigen, the biologically active lipid sphingosine-1-phosphate (S1P). We crystallized and determined the X-ray crystal structure of the LT1009 Fab fragment in 10 mM CaCl2 and found that it binds two Ca2+ in a manner similar to its antigen-bound state. Flame atomic absorption spectroscopy (FAAS) confirmed that murine LT1002 also binds Ca2+ in solution and inductively-coupled plasma-mass spectrometry (ICP-MS) revealed that, although Ca2+ is preferred, LT1002 can bind Mg2+ and, to much lesser extent, Ba2+. Isothermal titration calorimetry (ITC) indicated that LT1002 binds two Ca2+ ions endothermically with a measured dissociation constant (KD) of 171 μM. Protein and genome sequence analyses suggested that LT1002 is representative of a small class of confirmed and potential metalloantibodies and that Ca2+ binding is likely encoded for in germline variable chain genes. To test this hypothesis, we engineered, expressed, and purified a Fab fragment consisting of naïve murine germline-encoded light and heavy chain genes from which LT1002 is derived and observed that it binds Ca2+ in solution. We propose that LT1002 is representative of a class of naturally occurring metalloantibodies that are evolutionarily conserved across diverse mammalian genomes.

Farokhi, Elinaz↗

A comparative study of prebiotic and present day translational models

It is generally recognized that the understanding of the molecular basis of primitive translation is a fundamental step in developing a theory of the origin of life. However, even in modern molecular biology, the mechanism for the decoding of messenger RNA triplet codons into an amino acid sequence of a protein on the ribosome is understood incompletely. Most of the proposed models for prebiotic translation lack, not only experimental support, but also a careful theoretical scrutiny of their compatibility with well understood stereochemical and energetic principles of nucleic acid structure, molecular recognition principles, and the chemistry of peptide bond formation. Present studies are concerned with comparative structural modelling and mechanistic simulation of the decoding apparatus ranging from those proposed for prebiotic conditions to the ones involved in modern biology. Any primitive decoding machinery based on nucleic acids and proteins, and most likely the modern day system, has to satisfy certain geometrical constraints. The charged amino acyl and the peptidyl termini of successive adaptors have to be adjacent in space in order to satisfy the stereochemical requirements for amide bond formation. Simultaneously, the same adaptors have to recognize successive codons on the messenger. This translational complex has to be realized by components that obey nucleic acid conformational principles, stabilities, and specificities. This generalized condition greatly restricts the number of acceptable adaptor structures.

Rein, R.↗

A genomic timescale of prokaryote evolution: insights into the origin of methanogenesis, phototrophy, and the colonization of land

BACKGROUND: The timescale of prokaryote evolution has been difficult to reconstruct because of a limited fossil record and complexities associated with molecular clocks and deep divergences. However, the relatively large number of genome sequences currently available has provided a better opportunity to control for potential biases such as horizontal gene transfer and rate differences among lineages. We assembled a data set of sequences from 32 proteins (approximately 7600 amino acids) common to 72 species and estimated phylogenetic relationships and divergence times with a local clock method. RESULTS: Our phylogenetic results support most of the currently recognized higher-level groupings of prokaryotes. Of particular interest is a well-supported group of three major lineages of eubacteria (Actinobacteria, Deinococcus, and Cyanobacteria) that we call Terrabacteria and associate with an early colonization of land. Divergence time estimates for the major groups of eubacteria are between 2.5-3.2 billion years ago (Ga) while those for archaebacteria are mostly between 3.1-4.1 Ga. The time estimates suggest a Hadean origin of life (prior to 4.1 Ga), an early origin of methanogenesis (3.8-4.1 Ga), an origin of anaerobic methanotrophy after 3.1 Ga, an origin of phototrophy prior to 3.2 Ga, an early colonization of land 2.8-3.1 Ga, and an origin of aerobic methanotrophy 2.5-2.8 Ga. CONCLUSIONS: Our early time estimates for methanogenesis support the consideration of methane, in addition to carbon dioxide, as a greenhouse gas responsible for the early warming of the Earths' surface. Our divergence times for the origin of anaerobic methanotrophy are compatible with highly depleted carbon isotopic values found in rocks dated 2.8-2.6 Ga. An early origin of phototrophy is consistent with the earliest bacterial mats and structures identified as stromatolites, but a 2.6 Ga origin of cyanobacteria suggests that those Archean structures, if biologically produced, were made by anoxygenic photosynthesizers. The resistance to desiccation of Terrabacteria and their elaboration of photoprotective compounds suggests that the common ancestor of this group inhabited land. If true, then oxygenic photosynthesis may owe its origin to terrestrial adaptations.

Methane/metabolism↗

Sequence, structure prediction, and epitope analysis of the polymorphic membrane protein family in Chlamydia trachomatis

The polymorphic membrane proteins (Pmps) are a family of autotransporters that play an important role in infection, adhesion and immunity in Chlamydia trachomatis. Here we show that the characteristic GGA(I,L,V) and FxxN tetrapeptide repeats fit into a larger repeat sequence, which correspond to the coils of a large beta-helical domain in high quality structure predictions. Analysis of the protein using structure prediction algorithms provided novel insight to the chlamydial Pmp family of proteins. While the tetrapeptide motifs themselves are predicted to play a structural role in folding and close stacking of the beta-helical backbone of the passenger domain, we found many of the interesting features of Pmps are localized to the side loops jutting out from the beta helix including protease cleavage, host cell adhesion, and B-cell epitopes; while T-cell epitopes are predominantly found in the beta-helix itself. This analysis more accurately defines the Pmp family of Chlamydia and may better inform rational vaccine design and functional studies.

59 BASIC BIOLOGICAL SCIENCES↗

Structural Models and Sequence Alignment Results of the Desulfovibrio vulgaris Proteome

This dataset contains the structural models for the primary transcripts of the Desulfovibrio vulgaris proteome as well as sequence alignment results for a subset of the encoded proteins. For each protein, the five models inferred from AlphaFold 2 are provided. The largest pTM-scoring model for each protein was energy minimized; this minimized structure as well as its AlphaFold pickle output file are also provided. This set of structures represent an alternate source of models for the D. vulgaris proteome to those available in the AlphaFold Protein Structure Database (AFDB). This is a bit more complicated since the proteins reporting in the AFDB originate from an outdated form of the D. vulgaris sequence. The different versions of the D. vulgaris gene annotation are collected in the Chronology subdirectory; further consideration of these changes on the structural space of the proteome are currently underway. For proteins that have been annotated as hypothetical, sequence alignment results from the HHblits and SAdLSA alignment methods are provided. These methods are often more capable to resolve sequence homology than other methods. Therefore, the results from both HHblits and SAdLSA are provided to identify possible homologs for these challenging proteins. Numerous sequence databases are utilized for these alignments. References AlphaFold v2 Multimer: https://doi.org/10.1101/2021.10.04.463034. References HHblits: hhtps://doi.org/10.1186/s12859-019-3019-7. References SAdLSA: hhtps://doi.org/10.3389/fbinf.2021.689960.

59 BASIC BIOLOGICAL SCIENCES↗

cWINNOWER algorithm for finding fuzzy dna motifs

The cWINNOWER algorithm detects fuzzy motifs in DNA sequences rich in protein-binding signals. A signal is defined as any short nucleotide pattern having up to d mutations differing from a motif of length l. The algorithm finds such motifs if a clique consisting of a sufficiently large number of mutated copies of the motif (i.e., the signals) is present in the DNA sequence. The cWINNOWER algorithm substantially improves the sensitivity of the winnower method of Pevzner and Sze by imposing a consensus constraint, enabling it to detect much weaker signals. We studied the minimum detectable clique size qc as a function of sequence length N for random sequences. We found that qc increases linearly with N for a fast version of the algorithm based on counting three-member sub-cliques. Imposing consensus constraints reduces qc by a factor of three in this case, which makes the algorithm dramatically more sensitive. Our most sensitive algorithm, which counts four-member sub-cliques, needs a minimum of only 13 signals to detect motifs in a sequence of length N = 12,000 for (l, d) = (15, 4). Copyright Imperial College Press.

Evaluation Studies↗

nf-core/proteinfamilies: a scalable pipeline for the generation of protein families

The growth of metagenomics-derived amino acid sequence data has transformed our understanding of protein function, microbial diversity, and evolutionary relationships. However, the vast majority of these proteins remain functionally uncharacterized. Grouping the millions of such uncharacterized sequences with the few experimentally characterized ones allows the transfer of annotations, while the inspection of conserved residues with multiple sequence alignments can provide clues to function, even in the absence of existing functional information. To address the challenges associated with this data surge and the need to group sequences, we present a scalable, open-source, parametrizable Nextflow pipeline (nf-core/proteinfamilies) that generates nascent protein families or assigns new proteins to existing families. The computational benchmarks demonstrated that resource usage scales approximately linearly with input size, and the biological benchmarks showed that the generated protein families closely resemble manually curated families in widely used databases.

Nextflow↗

Structural and Biophysical Properties of a [4Fe-4S] Ferredoxin-Like Protein from Synechocystis sp. PCC 6803 with a Unique Two Domain Structure

Electron carrier proteins (ECPs), binding iron-sulfur clusters, are vital components within the intricate network of metabolic and photosynthetic reactions. They play a crucial role in the distribution of reducing equivalents. In Synechocystis sp. PCC 6803, the ECP network includes at least nine ferredoxins. Previous research, including global expression analyses and protein binding studies, has offered initial insights into the functional roles of individual ferredoxins within this network. This study primarily focuses on Ferredoxin 9 (slr2059). Through sequence analysis and computational modeling, Ferredoxin 9 emerges as a unique ECP with a distinctive two-domain architecture. It consists of a C-terminal iron-sulfur binding domain and an N-terminal domain with homology to Nil-domain proteins, connected by a structurally rigid 4-amino acid linker. Notably, in contrast to canonical [2Fe-2S] ferredoxins exemplified by PetF (ssl0020), which feature highly acidic surfaces facilitating electron transfer with photosystem I reaction centers, models of Ferredoxin 9 reveal a more neutral to basic protein surface. Using a combination of electron paramagnetic resonance spectroscopy and square-wave voltammetry on heterologously produced Ferredoxin 9, this study demonstrates that the protein coordinates 2x[4Fe-4S]2+/1+ redox-active and magnetically interacting clusters, with measured redox potentials of -420 +/- 9 mV and -516 +/- 10 mV vs SHE. A more in-depth analysis of Fdx9's unique structure and protein sequence suggests that this type of Nil-2[4Fe-4S] multi-domain ferredoxin is well conserved in cyanobacteria, bearing structural similarities to proteins involved in homocysteine synthesis in methanogens.

cyanobacteria↗

BMC Caller: a webtool to identify and analyze bacterial microcompartment types in sequence data

Bacterial microcompartments (BMCs) are protein-based organelles found across the bacterial tree of life. They consist of a shell, made of proteins that oligomerize into hexagonally and pentagonally shaped building blocks, that surrounds enzymes constituting a segment of a metabolic pathway. The proteins of the shell are unique to BMCs. They also provide selective permeability; this selectivity is dictated by the requirements of their cargo enzymes. We have recently surveyed the wealth of different BMC types and their occurrence in all available genome sequence data by analyzing and categorizing their components found in chromosomal loci using HMM (Hidden Markov Model) protein profiles. To make this a “do-it yourself” analysis for the public we have devised a webserver, BMC Caller (https://bmc-caller.prl.msu.edu), that compares user input sequences to our HMM profiles, creates a BMC locus visualization, and defines the functional type of BMC, if known. Shell proteins in the input sequence data are also classified according to our function-agnostic naming system and there are links to similar proteins in our database as well as an external link to a structure prediction website to easily generate structural models of the shell proteins, which facilitates understanding permeability properties of the shell. Additionally, the BMC Caller website contains a wealth of information on previously analyzed BMC loci with links to detailed data for each BMC protein and phylogenetic information on the BMC shell proteins. Our tools greatly facilitate BMC type identification to provide the user information about the associated organism’s metabolism and enable discovery of new BMC types by providing a reference database of all currently known examples.

59 BASIC BIOLOGICAL SCIENCES↗

Intramolecular interactions in aminoacyl nucleotides: Implications regarding the origin of genetic coding and protein synthesis

Cellular organisms store information as sequences of nucleotides in double stranded DNA. This information is useless unless it can be converted into the active molecular species, protein. This is done in contemporary creatures first by transcription of one strand to give a complementary strand of mRNA. The sequence of nucleotides is then translated into a specific sequence of amino acids in a protein. Translation is made possible by a genetic coding system in which a sequence of three nucleotides codes for a specific amino acid. The origin and evolution of any chemical system can be understood through elucidation of the properties of the chemical entities which make up the system. There is an underlying logic to the coding system revealed by a correlation of the hydrophobicities of amino acids and their anticodonic nucleotides (i.e., the complement of the codon). Its importance lies in the fact that every amino acid going into protein synthesis must first be activated. This is universally accomplished with ATP. Past studies have concentrated on the chemistry of the adenylates, but more recently we have found, through the use of NMR, that we can observe intramolecular interactions even at low concentrations, between amino acid side chains and nucleotide base rings in these adenylates. The use of this type of compound thus affords a novel way of elucidating the manner in which amino acids and nucleotides interact with each other. In aqueous solution, when a hydrophobic amino acid is attached to the most hydrophobic nucleotide, AMP, a hydrophobic interaction takes place between the amino acid side chain and the adenine ring. The studies to be reported concern these hydrophobic interactions.

Lacey, J. C., Jr.↗

Hidden Markov Model: a shortest unique representative approach to detect the protein toxins, virulence factors and antibiotic resistance genes

Objective: Currently, next generation sequencing (NGS) is widely used to decode potential novel or variant pathogens both in emergent outbreaks and in routine clinical practice. However, the efficient identification of novel or diverged pathogenomic compositions remains a big challenge. It is especially true for short DNA sequence fragments from NGS, since sequence similarity searching is vulnerable to false negatives or false positives, as is mismatching or matching with unrelated proteins. Therefore, this study aimed to establish a bioinformatics approach that can generate unique motif sequences for profiling searching, resulting in high specificity and sensitivity. Results: In this study, we introduced a Shortest Unique Representative Hidden Markov Model (HMM) approach to identify bacterial toxin, virulence factor (VF), and antimicrobial resistance (AR) in short sequence reads. We first construct unique representative domain sequences of toxin genes, VFs, and ARs to avoid potential false positives, and then to use HMM models to accurately identify potential toxin, VF, and AR fragments. The benchmark shows this approach can achieve relatively high specificity and sensitivity if the appropriate cutoff value is applied. Our approach can be used to recognize the protein sequences of known toxins and pathogens, identifies their common characteristics and then searches for similar sequences in other organisms.

59 BASIC BIOLOGICAL SCIENCES↗