Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “RNA sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

A minimal complex of KHNYN and zinc-finger antiviral protein binds and degrades single-stranded RNA

Detecting viral infection is a key role of the innate immune system. The genomes of some RNA viruses have a high CpG dinucleotide content relative to most vertebrate cell RNAs, making CpGs a molecular marker of infection. The human zinc-finger antiviral protein (ZAP) recognizes CpG, mediates clearance of the foreign CpG-rich RNA, and causes attenuation of CpG-rich RNA viruses. While ZAP binds RNA, it lacks enzymatic activity that might be responsible for RNA degradation and thus requires interacting cofactors for its function. One of these cofactors, KHNYN, has a predicted nuclease domain. Using biochemical approaches, we found that the KHNYN NYN domain is a single-stranded RNA ribonuclease that does not have sequence specificity and digests RNA with or without CpG dinucleotides equivalently in vitro. We show that unlike most KH domains, the KHNYN KH domain does not bind RNA. Indeed, a crystal structure of the KH region revealed a double-KH domain with a negatively charged surface that accounts for the lack of RNA binding. Rather, the KHNYN C-terminal domain (CTD) interacts with the ZAP RNA-binding domain (RBD) to provide target RNA specificity. We define a minimal complex composed of the ZAP RBD and the KHNYN NYN-CTD and use a fluorescence polarization assay to propose a model for how this complex interacts with a CpG dinucleotide-containing RNA. In the context of the cell, this module would represent the minimum ZAP and KHNYN domains required for CpG-recognition and ribonuclease activity essential for attenuation of viruses with clusters of CpG dinucleotides.

Yeoh, Zoe C. (ORCID:0000000226949068)↗

Developing Asparagaceae1726: An Asparagaceae‐specific probe set targeting 1726 loci for Hyb‐Seq and phylogenomics in the family

Abstract Premise Target sequence capture (Hyb‐Seq) is a cost‐effective sequencing strategy that employs RNA probes to enrich for specific genomic sequences. By targeting conserved low‐copy orthologs, Hyb‐Seq enables efficient phylogenomic investigations. Here, we present Asparagaceae1726—a Hyb‐Seq probe set targeting 1726 low‐copy nuclear genes for phylogenomics in the angiosperm family Asparagaceae—which will aid the often‐challenging delineation and resolution of evolutionary relationships within Asparagaceae. Methods Here we describe and validate the Asparagaceae1726 probe set (https://github.com/bentzpc/Asparagaceae1726) in six of the seven subfamilies of Asparagaceae. We perform phylogenomic analyses with these 1726 loci and evaluate how inclusion of paralogs and bycatch plastome sequences can enhance phylogenomic inference with target‐enriched data sets. Results We recovered at least 82% of target orthologs from all sampled taxa, and phylogenomic analyses resulted in strong support for all subfamilial relationships. Additionally, topology and branch support were congruent between analyses with and without inclusion of target paralogs, suggesting that paralogs had limited effect on phylogenomic inference. Discussion Asparagaceae1726 is effective across the family and enables the generation of robust data sets for phylogenomics of any Asparagaceae taxon. Asparagaceae1726 establishes a standardized set of loci for phylogenomic analysis in Asparagaceae, which we hope will be widely used for extensible and reproducible investigations of diversification in the family.

Plant Sciences↗

A potential role for RNA aminoacylation prior to its role in peptide synthesis

Coded ribosomal peptide synthesis could not have evolved unless its sequence and amino acid–specific aminoacylated tRNA substrates already existed. We therefore wondered whether aminoacylated RNAs might have served some primordial function prior to their role in protein synthesis. Here, we show that specific RNA sequences can be nonenzymatically aminoacylated and ligated to produce amino acid–bridged stem-loop RNAs. We used deep sequencing to identify RNAs that undergo highly efficient glycine aminoacylation followed by loop-closing ligation. The crystal structure of one such glycine-bridged RNA hairpin reveals a compact internally stabilized structure with the same eponymous T-loop architecture that is found in many noncoding RNAs, including the modern tRNA. We demonstrate that the T-loop-assisted amino acid bridging of RNA oligonucleotides enables the rapid template-free assembly of a chimeric version of an aminoacyl-RNA synthetase ribozyme. We suggest that the primordial assembly of amino acid–bridged chimeric ribozymes provides a direct and facile route for the covalent incorporation of amino acids into RNA. A greater functionality of covalently incorporated amino acids could contribute to enhanced ribozyme catalysis, providing a driving force for the evolution of sequence and amino acid–specific aminoacyl-RNA synthetase ribozymes in the RNA World. The synthesis of specifically aminoacylated RNAs, an unlikely prospect for nonenzymatic reactions but a likely one for ribozymes, could have set the stage for the subsequent evolution of coded protein synthesis.

Science & Technology - Other Topics↗

Populus VariantDB v3.2 facilitates CRISPR and functional genomics research

The success of CRISPR genome editing studies depends critically on the precision of guide RNA (gRNA) design. Sequence polymorphisms in outcrossing tree species pose design hazards that can render CRISPR genome editing ineffective. Despite recent advances in tree genome sequencing with haplotype resolution, sequence polymorphism information remains largely inaccessible to various functional genomics research efforts. The Populus VariantDB v3.2 addresses these challenges by providing a user-friendly search engine to query sequence polymorphisms of heterozygous genomes. The database accepts short sequences, such as gRNAs and primers, as input for searching against multiple poplar genomes, including hybrids, with customizable parameters. We provide examples to showcase the utilities of VariantDB in improving the precision of gRNA or primer design. The platform-agnostic nature of the probe search design makes Populus VariantDB v3.2 a versatile tool for the rapidly evolving CRISPR field and other sequence-sensitive functional genomics applications. The database schema is expandable and can accommodate additional tree genomes to broaden its user base.

59 BASIC BIOLOGICAL SCIENCES↗

Metastable multimeric G-quadruplex 2′FY-RNA aptamers that selectively bind pyoverdines

Two 2′FY-RNA aptamers with distinct sequences were selected for specific binding to pyoverdine-Pf5 (PVD-Pf5), increasing chromophore fluorescence upon binding. They also recognized the peptide portion of pyoverdines, as shown by their differential specificity for related variants. Computational analysis and experimental data (NMM binding, CD spectra) identified G-quadruplex structures that were thermally metastable but reformed in the presence of PVD-Pf5. Further structural studies mainly with one aptamer revealed imino proton peaks in 1D H-NMR and pressure stability up to 2 kbar. Electrophoretic evidence identified dimeric G-quadruplexes formed by the 2′FY-RNA aptamers and their RNA equivalents. While cations were necessary for PVD-Pf5 binding, they were not required for G-quadruplex formation. Given the established role of G-quadruplexes as protein interaction sites, multimeric G-quadruplexes offer a potential framework for structure-based regulatory mechanisms in cellular RNAs. In addition to previously characterized multimeric G-quadruplexes, these aptamers contribute novel sequences that expand the repertoire of known multimeric G-quadruplexes.

2′FY-RNA↗

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Zhou, Zhihan↗

Camelina circRNA landscape: Implications for gene regulation and fatty acid metabolism

Abstract Circular RNAs (circRNAs) are closed‐loop RNAs forming a covalent bond between their 3′ and 5′ ends, the back splice junction (BSJ), rendering them resistant to exonucleases and thus more stable compared to linear RNAs. Identification of circRNAs and distinction from their cognate linear RNA is only possible by sequencing the BSJ that is unique to the circRNA. CircRNAs are involved in the regulation of their cognate RNAs by increasing transcription rates, RNA stability, and alternative splicing. We have identified circRNAs from C. sativa that are associated with the regulation of germination, light response, and lipid metabolism. We sequenced light‐grown and etiolated seedlings after 5 or 7 days post‐germination and identified a total of 3447 circRNAs from 2763 genes. Most circRNAs originate from a single homeolog of the three subgenomes from allohexaploid camelina and correlate with higher ratios of alternative splicing of their cognate genes. A network analysis shows the interactions of select miRNA:circRNA:mRNAs for regulation of transcript stabilities where circRNA can act as a competing endogenous RNA. Several key lipid metabolism genes can generate circRNA, and we confirmed the presence of KASII circRNA as a true circRNA. CircRNA in camelina can be a novel target for breeding and engineering efforts.

Utley, Delecia [Department of Plant and Microbial ↗

Bacteria-mediated dsRNA delivery for mosquito-borne virus control

Mosquito-borne viruses represent an increasing global public health threat, exacerbated by urbanisation and climate change, thus making effective mosquito control essential. RNA interference (RNAi), a sequence-specific gene regulation mechanism, can be a flexible vector control tool. RNAi effectors, such as double-stranded RNA (dsRNA), can target mosquito genes or the viruses they carry, disrupting development or suppressing infection. However, current RNAi delivery methods are ineffective. Engineered bacterial symbionts offer a promising alternative for delivery, as they can produce dsRNA directly within mosquitoes. However, bacterial RNAi delivery in mosquitoes remains underexplored. We review emerging genetic tools, insights from RNAi and bacteria–mosquito interactions to outline priorities for realising bacterial RNAi as an efficient and sustainable vector control strategy.

Biological and medical sciences↗

Directing Nanoparticle Organization in Response to Diverse Chemical Inputs

Signaling cascades are crucial for transducing stimuli in biological systems, enabling multiple stimuli to regulate a downstream target with precisely controlled timing and amplifying signals through a series of intermediary reactions. Developing a robust signaling system with such capabilities would be pivotal for programming complex behaviors in synthetic DNA-based molecular devices. However, although “software” such as nucleic acid circuits could potentially be harnessed to relay signals to DNA-based nanostructure hardware, such explorations have been limited. Here, in this study, we develop a platform for transducing a variety of stimuli via messenger-mediated reactions to regulate the release and reloading of gold nanoparticles (AuNPs) in a 3D DNA framework. In the first step, an in vitro transcription circuit is engineered to sense and amplify chemical stimuli, including arbitrary DNA sequences and proteins, producing RNA. In the second step, the RNA releases the DNA-coated AuNPs from the DNA framework via a strand displacement reaction. AuNP reloading is controlled by a separate step driven by degradation of the RNA. Our platform holds promise for applications requiring dynamic multiagent control over DNA-based devices, offering a versatile tool for advanced molecular device engineering.

36 MATERIALS SCIENCE↗

Long-read sequencing transcriptome quantification with lr-kallisto

RNA abundance quantification has become routine and affordable thanks to high-throughput “short-read” technologies that provide accurate molecule counts at the gene level. Similarly accurate and affordable quantification of definitive full-length, transcript isoforms has remained a stubborn challenge, despite its obvious biological significance across a wide range of problems. “Long-read” sequencing platforms now produce data-types that can, in principle, drive routine definitive isoform quantification. However some particulars of contemporary long-read datatypes, together with isoform complexity and genetic variation, present bioinformatic challenges. We show here, using ONT data, that fast and accurate quantification of long-read data is possible and that it is improved by exome capture. To perform quantifications we developed lr-kallisto, which adapts the kallisto bulk and single-cell RNA-seq quantification methods for long-read technologies.

Loving, Rebekah K. (ORCID:0000000187250376)↗

Self-assembly and condensation of intermolecular poly(UG) RNA quadruplexes

Abstract Poly(UG) or ‘pUG’ dinucleotide repeats are highly abundant sequences in eukaryotic RNAs. In Caenorhabditis elegans, pUGs are added to RNA 3′ ends to direct gene silencing within Mutator foci, a germ granule condensate. Here, we show that pUG RNAs efficiently self-assemble into gel condensates through quadruplex (G4) interactions. Short pUG sequences form right-handed intermolecular G4s (pUG G4s), while longer pUGs form left-handed intramolecular G4s (pUG folds). We determined a 1.05 Å crystal structure of an intermolecular pUG G4, which reveals an eight stranded G4 dimer involving 48 nucleotides, 7 different G and U quartet conformations, 7 coordinated potassium ions, 8 sodium ions and a buried water molecule. A comparison of the intermolecular pUG G4 and intramolecular pUG fold structures provides insights into the molecular basis for G4 handedness and illustrates how a simple dinucleotide repeat sequence can form complex structures with diverse topologies.

Biochemistry & Molecular Biology↗

The direct and indirect drivers shaping RNA viral communities in grassland soils

ABSTRACT Recent studies have revealed diverse RNA viral communities in soils. Yet, how environmental factors influence soil RNA viruses remains largely unknown. Here, we recovered RNA viral communities from bulk metatranscriptomes sequenced from grassland soils managed for 5 years under multiple environmental conditions including water content, plant presence, cultivar type, and soil depth. More than half of the unique RNA viral contigs (64.6%) were assigned with putative hosts. About 74.7% of these classified RNA viral contigs are known as eukaryotic RNA viruses suggesting eukaryotic RNA viruses may outnumber prokaryotic RNA viruses by nearly three times in this grassland. Of the identified eukaryotic RNA viruses and the associated eukaryotic species, the most dominant taxa were Mitoviridae with an average relative abundance of 72.4%, and their natural hosts, Fungi with an average relative abundance of 56.6%. Network analysis and structural equation modeling support that soil water content, plant presence, and type of cultivar individually demonstrate a significant positive impact on eukaryotic RNA viral richness directly as well as indirectly on eukaryotic RNA viral abundance via influencing the co-existing eukaryotic members. A significant negative influence of soil depth on soil eukaryotic richness and abundance indirectly impacts soil eukaryotic RNA viral communities. These results provide new insights into the collective influence of multiple environmental and community factors that shape soil RNA viral communities and offer a structured perspective of how RNA virus diversity and ecology respond to environmental changes. IMPORTANCE Climate change has been reshaping the soil environment as well as the residing microbiome. This study provides field-relevant information on how environmental and community factors collectively shape soil RNA communities and contribute to ecological understanding of RNA viral survival under various environmental conditions and virus-host interactions in soil. This knowledge is critical for predicting the viral responses to climate change and the potential emergence of biothreats.

59 BASIC BIOLOGICAL SCIENCES↗

CABO-16S—a Combined Archaea, Bacteria, Organelle 16S rRNA database framework for amplicon analysis of prokaryotes and eukaryotes in environmental samples

Abstract Identification of both prokaryotic and eukaryotic microorganisms in environmental samples is currently challenged by the need for additional sequencing to obtain separate 16S and 18S ribosomal RNA (rRNA) amplicons or the constraints imposed by “universal” primers. Organellar 16S rRNA sequences are amplified and sequenced along with prokaryote 16S rRNA and provide an alternative method to identify eukaryotic microorganisms. CABO-16S combines bacterial and archaeal sequences from the SILVA database with 16S rRNA sequences of plastids and other organelles from the PR2 database to enable identification of all 16S rRNA sequences. Comparison of CABO-16S with SILVA 138.2 results in equivalent taxonomic classification of mock communities and increased classification of diverse environmental samples. In particular, identification of phototrophic eukaryotes in shallow seagrass environments, marine waters, and lake waters was increased. The CABO-16S framework allows users to add custom sequences for further classification of underrepresented clades and can be easily updated with future releases of reference databases. Addition of sequences obtained from Sanger sequencing of methane seep sediments and curated sequences of the polyphyletic SEEP-SRB1 clade resulted in differentiation of syntrophic and non-syntrophic SEEP-SRB1 in hydrothermal vent sediments. CABO-16S highlights the benefit of combining and amending existing training sets when studying microorganisms in diverse environments.

Eitel, Eryn M. (ORCID:0009000723919297)↗

Identification of shared viral sequences in peat moss metagenomes reveals elements of a possible Sphagnum core virome

Viruses are an understudied component of plant microbiomes. Identifying viruses that are shared between individual plants, or members of the “core virome”, could reveal stable viral populations with the potential to modulate the composition and function of the microbiome. Here, we examined the virome associated with Sphagnum mosses, a keystone species that has direct influence over the fate of peatland carbon stores. We analyzed bulk metagenomes and metatranscriptomes generated from Sphagnum field samples collected over a ten-month period to identify virus-like sequences shared among plants. Individual Sphagnum samples harbored distinct DNA and RNA viromes where only a small percentage (< 1%) of the total number of identified viral contigs were shared among all samples. Based on taxonomic classification, the shared viral contigs represent bacterial viruses, or phage (Caudoviricetes), as well as viruses of eukaryotes, namely nucleocytoplasmic large DNA viruses (Nucleocytoviricota) and RNA viruses (Riboviria). We linked the shared phage-like contigs to viral regions within sequenced genomes of bacterial taxa that are members of the Sphagnum core microbiome, suggesting that these contigs represent temperate phage or degraded prophage. The putative nucleocytoplasmic large DNA viruses and RNA viruses were phylogenetically diverse and showed sequence similarity to viruses associated with a broad range of hosts and environmental sources. The identification of shared viral contigs suggested that, despite the compositional heterogeneity between samples, Sphagnum mosses may harbor a core virome. Future work validating the presence of the core virome is warranted as it may aid in understanding how persistent viruses impact microbiome ecology and symbiont evolution within this climatically relevant keystone species.

Metagenomics↗

Structural insights into RNase H catalytic mechanism from room-temperature X-ray and neutron crystallography of apo- and RNA/DNA hybrid-bound enzyme

RNase H enzymes are sequence-nonspecific endonucleases that cleave RNA strands in RNA/DNA hybrid duplexes, an enzymatic process essential in DNA replication and repair in both prokaryotes and eukaryotes. Also, RNase H activity of the reverse transcriptase in human immunodeficiency viruses (HIV-1 and HIV-2) is indispensable for the viral replication cycle. RNase H enzymes play an central role in the development of gene therapies and are targets for novel antivirals. It is therefore of great importance to gain a detailed understanding of the RNase H catalytic mechanism to improve drug design. We utilized Bacillus halodurans RNase H1 (BhRNase H1) to shed light on its function and catalytic mechanism. Room-temperature neutron crystallography of the wild-type and inactive D132N mutant enzymes revealed that E109, belonging to the catalytic DEDD motif, can change its protonation state, allowing us to propose its role in the protonation of the leaving O3′ hydroxyl group of RNA. X-ray crystallography has demonstrated the ability of the RNA/DNA duplex to slide along the protein surface upon metal ion binding at site M A , transforming a product mimic into a Michaelis-like complex, which confirms an essential role of the M A metal ion in catalysis.

Enzyme mechanisms↗