Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “protein sequence”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Water, Solute, and Ion Transport in De Novo-Designed Membrane Protein Channels

Biological organisms engineer peptide sequences to fold into membrane pore proteins capable of performing a wide variety of transport functions. Synthetic de novo-designed membrane pores can mimic this approach to achieve a potentially even larger set of functions. Here, in this work, we explore water, solute, and ion transport in three de novo designed β-barrel membrane channels in the 5–10 Å pore size range. We show that these proteins form passive membrane pores with high water transport efficiencies and size rejection characteristics consistent with the pore size encoded in the protein structure. Ion conductance and ion selectivity measurements also show trends consistent with the pore size, with the two larger pores showing weak cation selectivity. MD simulations of water and ion transport and solute size exclusion are consistent with the experimental trends and provide further insights into structure–function correlations in these membrane pores.

59 BASIC BIOLOGICAL SCIENCES↗

Protein remote homology detection and structural alignment using deep learning

Exploiting sequence–structure–function relationships in biotechnology requires improved methods for aligning proteins that have low sequence similarity to previously annotated proteins. We develop two deep learning methods to address this gap, TM-Vec and DeepBLAST. TM-Vec allows searching for structure–structure similarities in large sequence databases. It is trained to accurately predict TM-scores as a metric of structural similarity directly from sequence pairs without the need for intermediate computation or solution of structures. Once structurally similar proteins have been identified, DeepBLAST can structurally align proteins using only sequence information by identifying structurally homologous regions between proteins. It outperforms traditional sequence alignment methods and performs similarly to structure-based alignment methods. We show the merits of TM-Vec and DeepBLAST on a variety of datasets, including better identification of remotely homologous proteins compared with state-of-the-art sequence alignment and structure prediction methods.

59 BASIC BIOLOGICAL SCIENCES↗

Properties of protein unfolded states suggest broad selection for expanded conformational ensembles

Much attention is being paid to conformational biases in the ensembles of intrinsically disordered proteins. However, it is currently unknown whether or how conformational biases within the disordered ensembles of foldable proteins affect function in vivo. Recently, we demonstrated that water can be a good solvent for unfolded polypeptide chains, even those with a hydrophobic and charged sequence composition typical of folded proteins. These results run counter to the generally accepted model that protein folding begins with hydrophobicity-driven chain collapse. Here we investigate what other features, beyond amino acid composition, govern chain collapse. We found that local clustering of hydrophobic and/or charged residues leads to significant collapse of the unfolded ensemble of pertactin, a secreted autotransporter virulence protein from Bordetella pertussis , as measured by small angle X-ray scattering (SAXS). Sequence patterns that lead to collapse also correlate with increased intermolecular polypeptide chain association and aggregation. Crucially, sequence patterns that support an expanded conformational ensemble enhance pertactin secretion to the bacterial cell surface. Similar sequence pattern features are enriched across the large and diverse family of autotransporter virulence proteins, suggesting sequence patterns that favor an expanded conformational ensemble are under selection for efficient autotransporter protein secretion, a necessary prerequisite for virulence. More broadly, we found that sequence patterns that lead to more expanded conformational ensembles are enriched across water-soluble proteins in general, suggesting protein sequences are under selection to regulate collapse and minimize protein aggregation, in addition to their roles in stabilizing folded protein structures.

59 BASIC BIOLOGICAL SCIENCES↗

PTM‐Psi : A python package to facilitate the computational investigation of p ost‐ t ranslational m odification on p rotein s tructures and their i mpacts on dynamics and functions

Abstract Post‐translational modification (PTM) of a protein occurs after it has been synthesized from its genetic template, and involves chemical modifications of the protein's specific amino acid residues. Despite of the central role played by PTM in regulating molecular interactions, particularly those driven by reversible redox reactions, it remains challenging to interpret PTMs in terms of protein dynamics and function because there are numerous combinatorially enormous means for modifying amino acids in response to changes in the protein environment. In this study, we provide a workflow that allows users to interpret how perturbations caused by PTMs affect a protein's properties, dynamics, and interactions with its binding partners based on inferred or experimentally determined protein structure. This Python‐based workflow, called PTM‐Psi , integrates several established open‐source software packages, thereby enabling the user to infer protein structure from sequence, develop force fields for non‐standard amino acids using quantum mechanics, calculate free energy perturbations through molecular dynamics simulations, and score the bound complexes via docking algorithms. Using the S ‐nitrosylation of several cysteines on the GAP2 protein as an example, we demonstrated the utility of PTM‐Psi for interpreting sequence–structure–function relationships derived from thiol redox proteomics data. We demonstrate that the S ‐nitrosylated cysteine that is exposed to the solvent indirectly affects the catalytic reaction of another buried cysteine over a distance in GAP2 protein through the movement of the two ligands. Our workflow tracks the PTMs on residues that are responsive to changes in the redox environment and lays the foundation for the automation of molecular and systems biology modeling.

59 BASIC BIOLOGICAL SCIENCES↗

Modeling SARS-CoV-2 proteins in the CASP-commons experiment

Critical Assessment of Structure Prediction (CASP) is an organization aimed at advancing the state of the art in computing protein structure from sequence. In the spring of 2020, CASP launched a community project to compute the structures of the most structurally challenging proteins coded for in the SARS-CoV-2 genome. Forty-seven research groups submitted over 3000 three-dimensional models and 700 sets of accuracy estimates on 10 proteins. The resulting models were released to the public. CASP community members also worked together to provide estimates of local and global accuracy and identify structure-based domain boundaries for some proteins. Subsequently, two of these structures (ORF3a and ORF8) have been solved experimentally, allowing assessment of both model quality and the accuracy estimates. Models from the AlphaFold2 group were found to have good agreement with the experimental structures, with main chain GDT_TS accuracy scores ranging from 63 (a correct topology) to 87 (competitive with experiment).

59 BASIC BIOLOGICAL SCIENCES↗

Nop9 recognizes structured and single-stranded RNA elements of preribosomal RNA

Nop9 is an essential factor in the processing of preribosomal RNA. Its absence in yeast is lethal, and defects in the human ortholog are associated with breast cancer, autoimmunity, and learning/language impairment. PUF family RNA-binding proteins are best known for sequence-specific RNA recognition, and most contain eight α-helical repeats that bind to the RNA bases of single-stranded RNA. Nop9 is an unusual member of this family in that it contains eleven repeats and recognizes both RNA structure and sequence. Here we report a crystal structure of Saccharomyces cerevisiae Nop9 in complex with its target RNA within the 20S preribosomal RNA. This structure reveals that Nop9 brings together a carboxy-terminal module recognizing the 5' single-stranded region of the RNA and a bifunctional amino-terminal module recognizing the central double-stranded stem region. We further show that the 3' single-stranded region of the 20S target RNA adds sequence-independent binding energy to the RNA–Nop9 interaction. Both the amino- and carboxy-terminal modules retain the characteristic sequence-specific recognition of PUF proteins, but the amino-terminal module has also evolved a distinct interface, which allows Nop9 to recognize either single-stranded RNA sequences or RNAs with a combination of single-stranded and structured elements.

59 BASIC BIOLOGICAL SCIENCES↗

Identification and characterization of the WYL BrxR protein and its gene as separable regulatory elements of a BREX phage restriction system

Bacteriophage exclusion (‘BREX’) phage restriction systems are found in a wide range of bacteria. Various BREX systems encode unique combinations of proteins that usually include a site-specific methyltransferase; none appear to contain a nuclease. Here we describe the identification and characterization of a Type I BREX system from Acinetobacter and the effect of deleting each BREX ORF on growth, methylation, and restriction. We identified a previously uncharacterized gene in the BREX operon that is dispensable for methylation but involved in restriction. Biochemical and crystallographic analyses of this factor, which we term BrxR (‘BREX Regulator’), demonstrate that it forms a homodimer and specifically binds a DNA target site upstream of its transcription start site. Deletion of the BrxR gene causes cell toxicity, reduces restriction, and significantly increases the expression of BrxC. In contrast, the introduction of a premature stop codon into the BrxR gene, or a point mutation blocking its DNA binding ability, has little effect on restriction, implying that the BrxR coding sequence and BrxR protein play independent functional roles. We speculate that elements within the BrxR coding sequence are involved in cis regulation of anti-phage activity, while the BrxR protein itself plays an additional regulatory role, perhaps during horizontal transfer.

59 BASIC BIOLOGICAL SCIENCES↗

Bioinformatics Investigations of Universal Stress Proteins from Mercury-Methylating Desulfovibrionaceae

The presence of methylmercury in aquatic environments and marine food sources is of global concern. The chemical reaction for the addition of a methyl group to inorganic mercury occurs in diverse bacterial taxonomic groups including the Gram-negative, sulfate-reducing Desulfovibrionaceae family that inhabit extreme aquatic environments. The availability of whole-genome sequence datasets for members of the Desulfovibrionaceae presents opportunities to understand the microbial mechanisms that contribute to methylmercury production in extreme aquatic environments. We have applied bioinformatics resources and developed visual analytics resources to categorize a collection of 719 putative universal stress protein (USP) sequences predicted from 93 genomes of Desulfovibrionaceae. We have focused our bioinformatics investigations on protein sequence analytics by developing interactive visualizations to categorize Desulfovibrionaceae universal stress proteins by protein domain composition and functionally important amino acids. We identified 651 Desulfovibrionaceae universal stress protein sequences, of which 488 sequences had only one USP domain and 163 had two USP domains. The 488 single USP domain sequences were further categorized into 340 sequences with ATP-binding motif and 148 sequences without ATP-binding motif. The 163 double USP domain sequences were categorized into (1) both USP domains with ATP-binding motif (3 sequences); (2) both USP domains without ATP-binding motif (138 sequences); and (3) one USP domain with ATP-binding motif (21 sequences). We developed visual analytics resources to facilitate the investigation of these categories of datasets in the presence or absence of the mercury-methylating gene pair (hgcAB). Future research could utilize these functional categories to investigate the participation of universal stress proteins in the bacterial cellular uptake of inorganic mercury and methylmercury production, especially in anaerobic aquatic environments.

Isokpehi, Raphael D. (ORCID:0000000268770840)↗

Unraveling the functional dark matter through global metagenomics

Metagenomes encode an enormous diversity of proteins, reflecting a multiplicity of functions and activities1,2. Exploration of this vast sequence space has been limited to a comparative analysis against reference microbial genomes and protein families derived from those genomes. Here, to examine the scale of yet untapped functional diversity beyond what is currently possible through the lens of reference genomes, we develop a computational approach to generate reference-free protein families from the sequence space in metagenomes. We analyse 26,931 metagenomes and identify 1.17 billion protein sequences longer than 35 amino acids with no similarity to any sequences from 102,491 reference genomes or the Pfam database3. Using massively parallel graph-based clustering, we group these proteins into 106,198 novel sequence clusters with more than 100 members, doubling the number of protein families obtained from the reference genomes clustered using the same approach. We annotate these families on the basis of their taxonomic, habitat, geographical and gene neighbourhood distributions and, where sufficient sequence diversity is available, predict protein three-dimensional models, revealing novel structures. Overall, our results uncover an enormously diverse functional space, highlighting the importance of further exploring the microbial functional dark matter.

54 ENVIRONMENTAL SCIENCES↗

The His-tag as a decoy modulating preferred orientation in cryoEM

The His-tag is a widely used affinity tag that facilitates purification by means of affinity chromatography of recombinant proteins for functional and structural studies. We show here that His-tag presence affects how coproheme decarboxylase interacts with the air-water interface during grid preparation for cryoEM. Depending on His-tag presence or absence, we observe significant changes in patterns of preferred orientation. Our analysis of particle orientations suggests that His-tag presence can mask the hydrophobic and hydrophilic patches on a protein’s surface that mediate the interactions with the air-water interface, while the hydrophobic linker between a His-tag and the coding sequence of the protein may enhance other interactions with the air-water interface. Our observations suggest that tagging, including rational design of the linkers between an affinity tag and a protein of interest, offer a promising approach to modulating interactions with the air-water interface.

59 BASIC BIOLOGICAL SCIENCES↗

LevSeq: Rapid Generation of Sequence-Function Data for Directed Evolution and Machine Learning

Sequence-function data provides valuable information about the protein functional landscape but is rarely obtained during directed evolution campaigns. Here, we present Long-read every variant Sequencing (LevSeq), a pipeline that combines a dual barcoding strategy with nanopore sequencing to rapidly generate sequence-function data for entire protein-coding genes. LevSeq integrates into existing protein engineering workflows and comes with open-source software for data analysis and visualization. The pipeline facilitates data-driven protein engineering by consolidating sequence-function data to inform directed evolution and provide the requisite data for machine learning-guided protein engineering (MLPE). LevSeq enables quality control of mutagenesis libraries prior to screening, which reduces time and resource costs. Simulation studies demonstrate LevSeq’s ability to accurately detect variants under various experimental conditions. Lastly, we show LevSeq’s utility in engineering protoglobins for new-to-nature chemistry. Widespread adoption of LevSeq and sharing of the data will enhance our understanding of protein sequence-function landscapes and empower data-driven directed evolution.

59 BASIC BIOLOGICAL SCIENCES↗

Conserved unique peptide patterns (CUPP) online platform 2.0: implementation of +1000 JGI fungal genomes

Carbohydrate-processing enzymes, CAZymes, are classified into families based on sequence and three-dimensional fold. Because many CAZyme families contain members of diverse molecular function (different EC-numbers), sophisticated tools are required to further delineate these enzymes. Such delineation is provided by the peptide-based clustering method CUPP, Conserved Unique Peptide Patterns. CUPP operates synergistically with the CAZy family/subfamily categorizations to allow systematic exploration of CAZymes by defining small protein groups with shared sequence motifs. The updated CUPP library contains 21,930 of such motif groups including 3,842,628 proteins. The new implementation of the CUPP-webserver, https://cupp.info/, now includes all published fungal and algal genomes from the Joint Genome Institute (JGI), genome resources MycoCosm and PhycoCosm, dynamically subdivided into motif groups of CAZymes. This allows users to browse the JGI portals for specific predicted functions or specific protein families from genome sequences. Thus, a genome can be searched for proteins having specific characteristics. All JGI proteins have a hyperlink to a summary page which links to the predicted gene splicing including which regions have RNA support. The new CUPP implementation also includes an update of the annotation algorithm that uses only a fourth of the RAM while enabling multi-threading, providing an annotation speed below 1 ms/protein.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic and phenotypic comparison of two variants of multidrug-resistant Salmonella enterica serovar Heidelberg isolated during the 2015–2017 multi-state outbreak in cattle

Salmonella enterica subspecies enterica serovar Heidelberg (Salmonella Heidelberg) has caused several multistate foodborne outbreaks in the United States, largely associated with the consumption of poultry. However, a 2015–2017 multidrug-resistant (MDR) Salmonella Heidelberg outbreak was linked to contact with dairy beef calves. Traceback investigations revealed calves infected with outbreak strains of Salmonella Heidelberg exhibited symptoms of disease frequently followed by death from septicemia. To investigate virulence characteristics of Salmonella Heidelberg as a pathogen in bovine, two variants with distinct pulse-field gel electrophoresis (PFGE) patterns that differed in morbidity and mortality during the multistate outbreak were genotypically and phenotypically characterized and compared. Strain SX 245 with PFGE pattern JF6X01.0523 was identified as a dominant and highly pathogenic variant causing high morbidity and mortality in affected calves, whereas strain SX 244 with PFGE pattern JF6X01.0590 was classified as a low pathogenic variant causing less morbidity and mortality. Comparison of whole-genome sequences determined that SX 245 lacked ~200 genes present in SX 244, including genes associated with the IncI1 plasmid and phages; SX 244 lacked eight genes present in SX 245 including a second YdiV Anti-FlhC(2)FlhD(4) factor, a lysin motif domain containing protein, and a pentapeptide repeat protein. RNA-sequencing revealed fimbriae-related, flagella-related, and chemotaxis genes had increased expression in SX 245 compared to SX 244. Furthermore, SX 245 displayed higher invasion of human and bovine epithelial cells than SX 244. These data suggest that the presence and up-regulation of genes involved in type 1 fimbriae production, flagellar regulation and biogenesis, and chemotaxis may play a role in the increased pathogenicity and host range expansion of the Salmonella Heidelberg isolates involved in the bovine-related outbreak.

59 BASIC BIOLOGICAL SCIENCES↗

A Comment on “Deep Proteogenomics of a Photosynthetic Cyanobacterium”

Proteomic researchers strive to achieve complete annotation of protein-coding DNA sequences to provide a foundational context for their relevant biological data. A recent deep proteogenomic study using a photosynthetic cyanobacterium Synechocystis sp. PCC 6803 by Spät et al. proposed 64 refined open reading frames (ORFs). By searching LC-MS/MS data from affinity chromatography-isolated protein complexes, our laboratory identified that six of these high-abundance ORFs possess Nterminal initiation start sites that differ than those proposed in the alternative models. Our findings are supported by highly confident MS2 data, phylogenetic analysis, chemical labeling, and established data from two independent research groups. Based on these highquality experimental identifications, we subsequently propose a standardized strategy and set of criteria for future deep proteogenomic efforts to ensure accurate and stringent proteogenomic annotation.

cyanobacteria↗

Cadherin 11 Promotes Immunosuppression and Extracellular Matrix Deposition to Support Growth of Pancreatic Tumors and Resistance to Gemcitabine in Mice

Pancreatic ductal adenocarcinomas (PDACs) are characterized by fibrosis and an abundance of cancer-associated fibroblasts (CAFs). Here, we investigated strategies to disrupt interactions among CAFs, the immune system, and cancer cells, focusing on adhesion molecule CDH11, which has been associated with other fibrotic disorders and is expressed by activated fibroblasts. We compared levels of CDH11 messenger RNA in human pancreatitis and pancreatic cancer tissues and cells with normal pancreas, and measured levels of CDH11 protein in human and mouse pancreatic lesions and normal tissues. We crossed p48-Cre;LSL-Kras G12D/+ ;LSL-Trp53 R172H/+ (KPC) mice with CDH11-knockout mice and measured survival times of offspring. Pancreata were collected and analyzed by histology, immunohistochemistry, and (single-cell) RNA sequencing; RNA and proteins were identified by imaging mass cytometry. Some mice were given injections of PD1 antibody or gemcitabine and survival was monitored. Pancreatic cancer cells from KPC mice were subcutaneously injected into Cdh11 +/+ and Cdh11 –/– mice and tumor growth was monitored. Pancreatic cancer cells (mT3) from KPC mice (C57BL/6), were subcutaneously injected into Cdh11 +/+ (C57BL/6J) mice and mice were given injections of antibody against CDH11, gemcitabine, or small molecule inhibitor of CDH11 (SD133) and tumor growth was monitored. Levels of CDH11 messenger RNA and protein were significantly higher in CAFs than in pancreatic cancer epithelial cells, human or mouse pancreatic cancer cell lines, or immune cells. KPC/Cdh11 +/– and KPC/Cdh11 –/– mice survived significantly longer than KPC/Cdh11 +/+ mice. Markers of stromal activation entirely surrounded pancreatic intraepithelial neoplasias in KPC/Cdh11 +/+ mice and incompletely in KPC/Cdh11 +/– and KPC/Cdh11 –/– mice, whose lesions also contained fewer FOXP3 + cells in the tumor center. Compared with pancreatic tumors in KPC/Cdh11 +/+ mice, tumors of KPC/Cdh11 +/– mice had increased markers of antigen processing and presentation; more lymphocytes and associated cytokines; decreased extracellular matrix components; and reductions in markers and cytokines associated with immunosuppression. Administration of the PD1 antibody did not prolong survival of KPC mice with 0, 1, or 2 alleles of Cdh11. Gemcitabine extended survival of KPC/Cdh11 +/– and KPC/Cdh11 –/– mice only or reduced subcutaneous tumor growth in mT3 engrafted Cdh11 +/+ mice when given in combination with the CDH11 antibody. A small molecule inhibitor of CDH11 reduced growth of pre-established mT3 subcutaneous tumors only if T and B cells were present in mice. Knockout or inhibition of CDH11, which is expressed by CAFs in the pancreatic tumor stroma, reduces growth of pancreatic tumors, increases their response to gemcitabine, and significantly extends survival of mice. CDH11 promotes immunosuppression and extracellular matrix deposition, and might be developed as a therapeutic target for pancreatic cancer.

59 BASIC BIOLOGICAL SCIENCES↗

Defining upstream enhancing and inhibiting sequence patterns for plant peroxisome targeting signal type 1 using large–scale in silico and in vivo analyses

Peroxisomes are universal eukaryotic organelles essential to plants and animals. Most peroxisomal matrix proteins carry peroxisome targeting signal type 1 (PTS1), a C-terminal tripeptide. Studies from various kingdoms have revealed influences from sequence upstream of the tripeptide on peroxisome targeting, supporting the view that positive charges in the upstream region are the major enhancing elements. However, a systematic approach to better define the upstream elements influencing PTS1 targeting capability is needed. Here, we used protein sequences from 177 plant genomes to perform large-scale and in-depth analysis of the PTS1 domain, which includes the PTS1 tripeptide and upstream sequence elements. We identified and verified 12 low-frequency PTS1 tripeptides and revealed upstream enhancing and inhibiting sequence patterns for peroxisome targeting, which were subsequently validated in vivo. Follow-up analysis revealed that nonpolar and acidic residues have relatively strong enhancing and inhibiting effects, respectively, on peroxisome targeting. However, in contrast to the previous understanding, positive charges alone do not show the anticipated enhancing effect and that both the position and property of the residues within these patterns are important for peroxisome targeting. We further demonstrated that the three residues immediately upstream of the tripeptide are the core influencers, with a ‘basic-nonpolar-basic’ pattern serving as a strong and universal enhancing pattern for peroxisome targeting. These findings have significantly advanced our knowledge of the PTS1 domain in plants and likely other eukaryotic species as well. The principles and strategies employed in the present study may also be applied to deciphering auxiliary targeting signals for other organelles.

59 BASIC BIOLOGICAL SCIENCES↗

Programming Amphiphilic Peptoid Oligomers for Hierarchical Assembly and Inorganic Crystallization

Natural organisms make a wide variety of exquisitely complex, nano-, micro-, and macroscale structured materials in an energy-efficient and highly reproducible manner. During these processes, the information-carrying biomolecules (e.g., proteins, peptides, and carbohydrates) enable (1) hierarchical organization to assemble scaffold materials and execute high-level functions and (2) exquisite control over inorganic materials synthesis, generating biominerals whose properties are optimized for their functions. Inspired by nature, significant efforts have been devoted to developing functional materials that can rival those natural molecules by mimicking in vivo functions using engineered proteins, peptides, DNAs, sequence-defined synthetic molecules (e.g., peptoids), and other biomimetic polymers. Among them, peptoids, a new type of synthetic mimetics of peptides and proteins, have received particular attention because they combine the merits of both synthetic polymers (e.g., high chemical stability and efficient synthesis) and biomolecules (e.g., sequence programmability and biocompatibility). The lack of both chirality and hydrogen bonds in their backbone results in a highly designable peptoid-based system with reduced structural complexity and side chain-chemistry-dominated properties. Here in this Account, we present our recent efforts in this field by programming amphiphilic peptoid sequences for (1) the controlled self-assembly into different hierarchically structured nanomaterials with favorable properties and (2) manipulating inorganic (nano)crystal nucleation, growth, and assembly into superstructures. First, we designed a series of amphiphilic peptoids with controlled side chain chemistries that self-assembled into 1D highly stiff and dynamic nanotubes, 2D membrane-mimetic nanosheets, hexagonally patterned nanoribbons, and 3D nanoflowers. These crystalline nanostructures exhibited sequence-dependent properties and showed promise for different applications. The corresponding peptoid self-assembly pathways and mechanisms were also investigated by leveraging in situ atomic force microscopy studies and molecular dynamics simulations, which showed precise sequence dependency. Second, inspired by peptide- and protein-controlled formation of hierarchical inorganic nanostructures in nature, we developed peptoid-based biomimetic approaches for controlled synthesis of inorganic materials (e.g., noble metals and calcite), in which we took advantage of the substantial side chain chemistry of peptoids and investigated the relationship between the peptoid sequences and the morphology and growth kinetics of inorganic materials. For example, to overcome the challenges (e.g., complexity of protein- and peptide-folding, poor thermal and chemical stabilities) facing the area of protein- and peptide-controlled synthesis of inorganic materials, we recently reported the design of sequence-defined peptoids for controlled synthesis of highly branched plasmonic gold particles. Moreover, we developed a rule of thumb for designing peptoids that predictively enabled the morphological evolution from spherical to coral-shaped gold nanoparticles (NPs). With this Account, we hope to stimulate the research interest of chemists and materials scientists and promote the predictive synthesis of functional and robust materials through the design of sequence-defined synthetic molecules.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

jialiu232/MetaFunPrimer_paper_info

Genes belonging to the same functional group may include numerous and variable gene sequences, making characterizing and quantifying difficult. Therefore, high-throughput design tools are needed to simultaneously create primers for improved quantification of target genes. We developed MetaFunPrimer, a bioinformatic pipeline, to design primers for numerous genes of interest. This tool also enables gene target prioritization based on ranking the presence of genes in user-defined references, such as environment-specific metagenomes. Given inputs of protein and nucleotide sequences for gene targets of interest and an accompanying set of reference metagenomes or genomes, MetaFunPrimer generates primers for ranked genes of interest. To demonstrate the usage and benefits of MetaFunPrimer, a total of 78 primer pairs were designed to target observed ammonia monooxygenase subunit A (amoA) genes of ammonia-oxidizing bacteria (AOB) in 1,550 publicly available soil metagenomes. We demonstrate computationally that these amoA-AOB primers can cover 94% of the amoA-AOB genes observed in the 1,550 soil metagenomes compared with a 49% estimated coverage by previously published primers. Finally, we verified the utility of these primer sets in incubation experiments that used long-term nitrogen fertilized or unfertilized soils. High-throughput quantitative PCR (qPCR) results and statistical analyses showed significant differences in relative quantification patterns between the two soils, and subsequent absolute quantifications also confirmed that target genes enumerated by six selected primer pairs were significantly more abundant in the nitrogen-fertilized soils. This new tool gives microbial ecologists a new approach to assess functional gene abundance and related microbial community dynamics quickly and affordably.

Liu, Jia↗