Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “consensus sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Updated HIV-1 Consensus Sequences Change but Stay Within Similar Distance From Worldwide Samples

HIV consensus sequences are used in various bioinformatic, evolutionary, and vaccine related research. Since the previous HIV-1 subtype and CRF consensus sequences were constructed in 2002, the number of publicly available HIV-1 sequences have grown exponentially, especially from non-EU and US countries. Here, we reconstruct 90 new HIV-1 subtype and CRF consensus sequences from 3,470 high-quality, representative, full genome sequences in the LANL HIV database. While subtypes and CRFs are unevenly spread across the world, in total 89 countries were represented. For consensus sequences that were based on at least 20 genomes, we found that on average 2.3% (range 0.8–10%) of the consensus genome site states changed from 2002 to 2021, of which about half were nucleotide state differences and the rest insertions and deletions. Interestingly, the 2021 consensus sequences were shorter than in 2002, and compared to 4,674 HIV-1 worldwide genome sequences, the 2021 consensuses were somewhat closer to the worldwide genome sequences, i.e., showing on average fewer nucleotide state differences. Some subtypes/CRFs have had limited geographical spread, and thus sampling of subtypes/CRFs is uneven, at least in part, due to the epidemiological dynamics. Thus, taken as a whole, the 2021 consensus sequences likely are good representations of the typical subtype/CRF genome nucleotide states. The new consensus sequences are available at the LANL HIV database.

60 APPLIED LIFE SCIENCES↗

Monomer-scale design of functional protein polymers using consensus repeat sequences

Protein-based polymers possess chemically defined sequences that can encode diverse properties and functions into a new class of biopolymeric materials. However, sequence variation that emerges from evolution can obscure the sequence–function relationships of naturally derived polymers. One strategy to clarify these relationships is to identify common sequences between proteins with similar functions. These conserved sequences often emerge from repeat proteins, and “consensus repeat sequences” provide a convenient platform for systematic investigations of biopolymer sequence–property relationships. In this review, we highlight recent approaches to engineer tunable polymeric materials using monomer-scale design of consensus repeat proteins. Here, we explore established and emerging protein-based materials with mechanical resilience, thermodynamic phase behavior, chemical responsiveness, biomolecular transport, and hierarchical structure. Overall, recent advances in the monomer-scale design of repetitive protein polymers present exciting fundamental and translational opportunities for polymer scientists and engineers.

36 MATERIALS SCIENCE↗

Optimizing Cell-based Antimicrobials through Pooled Genomic Libraries

DNA synthesis and assembly technologies ushered in through synthetic biology have great promise for biomanufacturing, bioremediation, and the development of living therapeutics. Unfortunately, predicting sequence to function relationships, including for biosynthetic pathways expressed in a new host organism, is difficult and often requires many iterative cycles of design, construction, and testing. We are working to develop data-driven approaches to identify the genetic determinants of growth defects and productivity for the expression of a cell-based antimicrobial. We assayed the growth, pigment production, and antimicrobial activity of a collection of over 10,000 genetic mutants of the violacein biosynthetic pathway and sequenced the genetic variation of these mutants. Through this project, we have developed an innovative codebase to automate the determination of pigmentation and antimicrobial clearing diameter for tens of thousands of genetic mutants cultivated on agar dishes. Further, we have written DNA sequence analysis code to demultiplex & provide consensus sequences from high-throughput PacBio long-read circular consensus sequencing (CCS) datasets. From this foundation, we plan to map DNA sequence to function to predict an optimal genetic design to maximize antimicrobial activity while minimizing deleterious growth effects. The workflows and algorithms developed through this project can be broadly applied to other engineered functions in microbes, uncovering sequence to function relationships for complex phenotypes where function impacts fitness.

59 BASIC BIOLOGICAL SCIENCES↗

Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces

These data are from Bandopadhyay et al., "Soil microbial ecology and microbiome-metabolite linkages improve understanding of ecosystem states along terrestrial-aquatic interfaces". This study aims to understand the soil microbial ecology along terrestrial-aquatic interfaces of a freshwater and estuarine region and how it relates to organic matter. We analyzed soil microbial (16S rRNA gene) and organic matter (Fourier-transform ion cyclotron resonance mass spectrometry, FTICR-MS) composition from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie (freshwater) and Chesapeake Bay (estuarine) regions. This dataset includes 16S rRNA gene amplicon data (only processed file types included here) and organic matter composition from FTICR-MS data (raw and processed files included here) from upland (forested), transition (stressed forest), and wetland positions at three sites in each of the Lake Erie and Chesapeake Bay regions. These sites are part of the COMPASS-FME project (https://compass.pnnl.gov/FME/COMPASSFME). File formats and software needed to access files: 16S rRNA gene amplicon data: These files follow the format reported here https://ess-dive.gitbook.io/amplicon-sequencing-reporting-format#updates-in-v1.0.1. As per this format, there are four file types reported: 1. Taxon tables (also called sequence-by-sample or OTU (operational taxonomic unit)/ESV (exact sequence variant) tables) : available in a .txt file format and accessible using TextEdit or MS Excel. 2. Representative sequences (also called consensus sequences) : available in a .fasta format and accessible using TextEdit. 3. Sequencing metadata : available in a MS Excel workbook file format and CSV file format 4. Bioinformatic metadata : available in a MS Excel workbook file format and CSV file format FTICR-MS data: 1. Raw data converted to a processed file with intensities of the peaks in the given samples : available in a MS Excel CSV file format 2. Processed file used in analyses and visualizations (appended as icr_long_) : available in a MS Excel CSV file format 3. Metadata file for ICR features (appended as icr_meta) : available in a MS Excel CSV file format

54 ENVIRONMENTAL SCIENCES↗

In vitro selection of optimal DNA substrates for T4 RNA ligase

We have used in vitro selection techniques to characterize DNA sequences that are ligated efficiently by T4 RNA ligase. We find that the ensemble of selected sequences ligated about 10 times as efficiently as the random mixture of sequences used as the input for selection. Surprisingly, the majority of the selected sequences approximated a well-defined consensus sequence.

Harada, Kazuo↗

Pseudouridine synthase 7 is an opportunistic enzyme that binds and modifies substrates with diverse sequences and structures

Pseudouridine (Ψ) is a ubiquitous RNA modification incorporated by pseudouridine synthase (Pus) enzymes into hundreds of noncoding and protein-coding RNA substrates. Here, in this study, we determined the contributions of substrate structure and protein sequence to binding and catalysis by pseudouridine synthase 7 (Pus7), one of the principal messenger RNA (mRNA) modifying enzymes. Pus7 is distinct among the eukaryotic Pus proteins because it modifies a wider variety of substrates and shares limited homology with other Pus family members. We solved the crystal structure of Saccharomyces cerevisiae Pus7, detailing the architecture of the eukaryotic-specific insertions thought to be responsible for the expanded substrate scope of Pus7. Additionally, we identified an insertion domain in the protein that fine-tunes Pus7 activity both in vitro and in cells. These data demonstrate that Pus7 preferentially binds substrates possessing the previously identified UGUAR (R = purine) consensus sequence and that RNA secondary structure is not a strong requirement for Pus7-binding. In contrast, the rate constants and extent of Ψ incorporation are more influenced by RNA structure, with Pus7 modifying UGUAR sequences in less-structured contexts more efficiently both in vitro and in cells. Although less-structured substrates were preferred, Pus7 fully modified every transfer RNA, mRNA, and nonnatural RNA containing the consensus recognition sequence that we tested. Our findings suggest that Pus7 is a promiscuous enzyme and lead us to propose that factors beyond inherent enzyme properties (e.g., enzyme localization, RNA structure, and competition with other RNA-binding proteins) largely dictate Pus7 substrate selection.

59 BASIC BIOLOGICAL SCIENCES↗

An RNA motif that binds ATP

RNAs that contain specific high-affinity binding sites for small molecule ligands immobilized on a solid support are present at a frequency of roughly one in 10(10)-10(11) in pools of random sequence RNA molecules. Here we describe a new in vitro selection procedure designed to ensure the isolation of RNAs that bind the ligand of interest in solution as well as on a solid support. We have used this method to isolate a remarkably small RNA motif that binds ATP, a substrate in numerous biological reactions and the universal biological high-energy intermediate. The selected ATP-binding RNAs contain a consensus sequence, embedded in a common secondary structure. The binding properties of ATP analogues and modified RNAs show that the binding interaction is characterized by a large number of close contacts between the ATP and RNA, and by a change in the conformation of the RNA.

NASA Discipline Exobiology↗

Repetitive proteins that undergo large conformational changes evade structural prediction algorithms

Protein structure prediction algorithms, such as AlphaFold, have accelerated protein design and advanced the understanding of the relationship between amino acid sequence and protein structure. However, these algorithms are limited in their ability to predict the structures of conformationally dynamic, intrinsically disordered, and stimuli-responsive proteins. To evaluate sequence-to-structure predictions of such challenging proteins, we explored a class of conformationally dynamic, repeats-in-toxin (RTX) proteins. RTX proteins adopt intrinsically disordered conformations in the absence of calcium and undergo reversible folding into β-roll structures upon binding to calcium. RTX proteins are characterized by tandem repeats of the sequence GGXGXDXUX, in which X can be any amino acid and U is an aliphatic amino acid. We designed RTX sequence variants with global substitutions of nonconserved amino acids, tandem repeats of consensus sequences GGAGXDTLY, and tandem repeats of scrambled sequences GGAGXDTYL. AlphaFold2 and AlphaFold3 predicted that all of these RTX variants adopt β-roll structures, characteristic of wild-type RTX bound to calcium. However, modeling the predicted structures with molecular dynamics simulations and characterizing the protein variants with circular dichroism spectroscopy, small-angle x-ray scattering, and x-ray crystallography revealed that variants adopt diverse, sequence-dependent structures in the absence and presence of calcium. To better design proteins for applications in biotechnology and sustainability, it is critical to build predictive tools that consider intrinsically disordered protein states and validate these tools with multi-mode, multi-scale experimental data.

Chang, Marina P. [Stanford Univ., CA (United State↗

Sequence, molecular properties, and chromosomal mapping of mouse lumican

PURPOSE. Lumican is a major proteoglycan of vertebrate cornea. This study characterizes mouse lumican, its molecular form, cDNA sequence, and chromosomal localization. METHODS. Lumican sequence was determined from cDNA clones selected from a mouse corneal cDNA expression library using a bovine lumican cDNA probe. Tissue expression and size of lumican mRNA were determined using Northern hybridization. Glycosidase digestion followed by Western blot analysis provided characterization of molecular properties of purified mouse corneal lumican. Chromosomal mapping of the lumican gene (Lcn) used Southern hybridization of a panel of genomic DNAs from an interspecific murine backcross. RESULTS. Mouse lumican is a 338-amino acid protein with high-sequence identity to bovine and chicken lumican proteins. The N-terminus of the lumican protein contains consensus sequences for tyrosine sulfation. A 1.9-kb lumican mRNA is present in cornea and several other tissues. Antibody against bovine lumican reacted with recombinant mouse lumican expressed in Escherichia coli and also detected high molecular weight proteoglycans in extracts of mouse cornea. Keratanase digestion of corneal proteoglycans released lumican protein, demonstrating the presence of sulfated keratan sulfate chains on mouse corneal lumican in vivo. The lumican gene (Lcn) was mapped to the distal region of mouse chromosome 10. The Lcn map site is in the region of a previously identified developmental mutant, eye blebs, affecting corneal morphology. CONCLUSIONS. This study demonstrates sulfated keratan sulfate proteoglycan in mouse cornea and describes the tools (antibodies and cDNA) necessary to investigate the functional role of this important corneal molecule using naturally occurring and induced mutants of the murine lumican gene.

NASA Discipline Cell Biology↗

Archaebacterial rhodopsin sequences: Implications for evolution

It was proposed over 10 years ago that the archaebacteria represent a separate kingdom which diverged very early from the eubacteria and eukaryotes. It follows that investigations of archaebacterial characteristics might reveal features of early evolution. So far, two genes, one for bacteriorhodopsin and another for halorhodopsin, both from Halobacterium halobium, have been sequenced. We cloned and sequenced the gene coding for the polypeptide of another one of these rhodopsins, a halorhodopsin in Natronobacterium pharaonis. Peptide sequencing of cyanogen bromide fragments, and immuno-reactions of the protein and synthetic peptides derived from the C-terminal gene sequence, confirmed that the open reading frame was the structural gene for the pharaonis halorhodopsin polypeptide. The flanking DNA sequences of this gene, as well as those of other bacterial rhodopsins, were compared to previously proposed archaebacterial consensus sequences. In pairwise comparisons of the open reading frame with DNA sequences for bacterio-opsin and halo-opsin from Halobacterium halobium, silent divergences were calculated. These indicate very considerable evolutionary distance between each pair of genes, even in the dame organism. In spite of this, three protein sequences show extensive similarities, indicating strong selective pressures.

Lanyi, J. K.↗

Insights from a workplace SARS-CoV-2 specimen collection program, with genomes placed into global sequence phylogeny

In 2020, the Department of Energy established the National Virtual Biotechnology Laboratory (NVBL) to address key challenges associated with COVID-19. As part of that effort, Pacific Northwest National Laboratory (PNNL) established a capability to collect and analyze specimens from employees who self-reported symptoms consistent with the disease. During the spring and fall of 2021, 688 specimens were screened for SARS-CoV-2, with 64 (9.3%) testing positive using reverse-transcriptase quantitative PCR (RT-qPCR). Of these, 36 samples were released for research. All 36 positive samples released for research were sequenced and genotyped. Here, the relationship between patient age and viral load as measured by Ct values was measured and determined to be only weakly significant. Consensus sequences for each sample were placed into a global phylogeny and transmission dynamics were investigated, revealing that the closest relative for many samples was from outside of Washington state, indicating mixing of viral pools within geographic regions.

59 BASIC BIOLOGICAL SCIENCES↗

Structural dissection of sequence recognition and catalytic mechanism of human LINE-1 endonuclease

Abstract Long interspersed nuclear element-1 (L1) is an autonomous non-LTR retrotransposon comprising ∼20% of the human genome. L1 self-propagation causes genomic instability and is strongly associated with aging, cancer and other diseases. The endonuclease domain of L1’s ORFp2 protein (L1-EN) initiates de novo L1 integration by nicking the consensus sequence 5′-TTTTT/AA-3′. In contrast, related nucleases including structurally conserved apurinic/apyrimidinic endonuclease 1 (APE1) are non-sequence specific. To investigate mechanisms underlying sequence recognition and catalysis by L1-EN, we solved crystal structures of L1-EN complexed with DNA substrates. This showed that conformational properties of the preferred sequence drive L1-EN’s sequence-specificity and catalysis. Unlike APE1, L1-EN does not bend the DNA helix, but rather causes ‘compression’ near the cleavage site. This provides multiple advantages for L1-EN’s role in retrotransposition including facilitating use of the nicked poly-T DNA strand as a primer for reverse transcription. We also observed two alternative conformations of the scissile bond phosphate, which allowed us to model distinct conformations for a nucleophilic attack and a transition state that are likely applicable to the entire family of nucleases. This work adds to our mechanistic understanding of L1-EN and related nucleases and should facilitate development of L1-EN inhibitors as potential anticancer and antiaging therapeutics.

59 BASIC BIOLOGICAL SCIENCES↗

Multiple Mutations Associated with Emergent Variants Can Be Detected as Low-Frequency Mutations in Early SARS-CoV-2 Pandemic Clinical Samples

Genetic analysis of intra-host viral populations provides unique insight into pre-emergent mutations that may contribute to the genotype of future variants. Clinical samples positive for SARS-CoV-2 collected in California during the first months of the pandemic were sequenced to define the dynamics of mutation emergence as the virus became established in the state. Deep sequencing of 90 nasopharyngeal samples showed that many mutations associated with the establishment of SARS-CoV-2 globally were present at varying frequencies in a majority of the samples, even those collected as the virus was first detected in the US. A subset of mutations that emerged months later in consensus sequences were detected as subconsensus members of intra-host populations. Spike mutations P681H, H655Y, and V1104L were detected prior to emergence in variant genotypes, mutations were detected at multiple positions within the furin cleavage site, and pre-emergent mutations were identified in the nucleocapsid and the envelope genes. Because many of the samples had a very high depth of coverage, a bioinformatics pipeline, “Mappgene”, was established that uses both iVar and LoFreq variant calling to enable identification of very low-frequency variants. This enabled detection of a spike protein deletion present in many samples at low frequency and associated with a variant of concern.

60 APPLIED LIFE SCIENCES↗

Structure of the coding region and mRNA variants of the apyrase gene from pea (Pisum sativum)

Partial amino acid sequences of a 49 kDa apyrase (ATP diphosphohydrolase, EC 3.6.1.5) from the cytoskeletal fraction of etiolated pea stems were used to derive oligonucleotide DNA primers to generate a cDNA fragment of pea apyrase mRNA by RT-PCR and these primers were used to screen a pea stem cDNA library. Two almost identical cDNAs differing in just 6 nucleotides within the coding regions were found, and these cDNA sequences were used to clone genomic fragments by PCR. Two nearly identical gene fragments containing 8 exons and 7 introns were obtained. One of them (H-type) encoded the mRNA sequence described by Hsieh et al. (1996) (DDBJ/EMBL/GenBank Z32743), while the other (S-type) differed by the same 6 nucleotides as the mRNAs, suggesting that these genes may be alleles. The six nucleotide differences between these two alleles were found solely in the first exon, and these mutation sites had two types of consensus sequences. These mRNAs were found with varying lengths of 3' untranslated regions (3'-UTR). There are some similarities between the 3'-UTR of these mRNAs and those of actin and actin binding proteins in plants. The putative roles of the 3'-UTR and alternative polyadenylation sites are discussed in relation to their possible role in targeting the mRNAs to different subcellular compartments.

NASA Discipline Plant Biology↗

CovS inactivation reduces CovR promoter binding at diverse virulence factor encoding genes in group A Streptococcus

The control of virulence gene regulator (CovR), also called caspsule synthesis regulator (CsrR), is critical to how the major human pathogen group A Streptococcus fine-tunes virulence factor production. CovR phosphorylation (CovR~P) levels are determined by its cognate sensor kinase CovS, and functional abrogating mutations in CovS can occur in invasive GAS isolates leading to hypervirulence. Presently, the mechanism of CovR-DNA binding specificity is unclear, and the impact of CovS inactivation on global CovR binding has not been assessed. Thus, we performed CovR chromatin immunoprecipitation sequencing (ChIP-seq) analysis in the emm1 strain MGAS2221 and its CovS kinase deficient derivative strain 2221-CovS-E281A. We identified that CovR bound in the promoter regions of nearly all virulence factor encoding genes in the CovR regulon. Additionally, direct CovR binding was observed for numerous genes encoding proteins involved in amino acid metabolism, but we found limited direct CovR binding to genes encoding other transcriptional regulators. The consensus sequence AATRANAAAARVABTAAA was present in the promoters of genes directly regulated by CovR, and mutations of highly conserved positions within this motif relieved CovR repression of the hasA and MGAS2221_0187 promoters. Analysis of strain 2221-CovS-E281A revealed that binding of CovR at repressed, but not activated, promoters is highly dependent on CovR~P state. CovR repressed virulence factor encoding genes could be grouped dependent on how CovR~P dependent variation in DNA binding correlated with gene transcript levels. Taken together, the data show that CovR repression of virulence factor encoding genes is primarily direct in nature, involves binding to a newly-identified DNA binding motif, and is relieved by CovS inactivation. These data provide new mechanistic insights into one of the most important bacterial virulence regulators and allow for subsequent focused investigations into how CovR-DNA interaction at directly controlled promoters impacts GAS pathogenesis.

59 BASIC BIOLOGICAL SCIENCES↗

Benefits and Limits of Phasing Alleles for Network Inference of Allopolyploid Complexes

Abstract Accurately reconstructing the reticulate histories of polyploids remains a central challenge for understanding plant evolution. Although phylogenetic networks can provide insights into relationships among polyploid lineages, inferring networks may be hindered by the complexities of homology determination in polyploid taxa. We use simulations to show that phasing alleles from allopolyploid individuals can improve phylogenetic network inference under the multispecies coalescent by obtaining the true network with fewer loci compared with haplotype consensus sequences or sequences with heterozygous bases represented as ambiguity codes. Phased allelic data can also improve divergence time estimates for networks, which is helpful for evaluating allopolyploid speciation hypotheses and proposing mechanisms of speciation. To achieve these outcomes in empirical data, we present a novel pipeline that leverages a recently developed phasing algorithm to reliably phase alleles from polyploids. This pipeline is especially appropriate for target enrichment data, where the depth of coverage is typically high enough to phase entire loci. We provide an empirical example in the North American Dryopteris fern complex that demonstrates insights from phased data as well as the challenges of network inference. We establish that our pipeline (PATÉ: Phased Alleles from Target Enrichment data) is capable of recovering a high proportion of phased loci from both diploids and polyploids. These data may improve network estimates compared with using haplotype consensus assemblies by accurately inferring the direction of gene flow, but statistical nonidentifiability of phylogenetic networks poses a barrier to inferring the evolutionary history of reticulate complexes.

Evolutionary Biology↗

Binding profiles for 961 Drosophila and C. elegans transcription factors reveal tissue-specific regulatory relationships

A catalog of transcription factor (TF) binding sites in the genome is critical for deciphering regulatory relationships. Here, we present the culmination of the efforts of the modENCODE (model organism Encyclopedia of DNA Elements) and modERN (model organism Encyclopedia of Regulatory Networks) consortia to systematically assay TF binding events in vivo in two major model organisms,Drosophila melanogaster(fly) andCaenorhabditis elegans(worm). These data sets comprise 605 TFs identifying 3.6 M sites in the fly and 356 TFs identifying 0.9 M sites in the worm, and represent the majority of the regulatory space in each genome. We demonstrate that TFs associate with chromatin in clusters termed “metapeaks,” that larger metapeaks have characteristics of high-occupancy target (HOT) regions, and that the importance of consensus sequence motifs bound by TFs depends on metapeak size and complexity. Combining ChIP-seq data with single-cell RNA-seq data in a machine-learning model identifies TFs with a prominent role in promoting target gene expression in specific cell types, even differentiating between parent–daughter cells during embryogenesis. These data are a rich resource for the community that should fuel and guide future investigations into TF function. To facilitate data accessibility and utility, all strains expressing green fluorescent protein (GFP)-tagged TFs are available at the stock centers for each organism. The chromatin immunoprecipitation sequencing data are available through the ENCODE Data Coordinating Center, GEO, and through a direct interface that provides rapid access to processed data sets and summary analyses, as well as widgets to probe the cell-type-specific TF–target relationships.

Biochemistry & Molecular Biology↗