Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “RNA sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Transcriptomics-based Machine Learning Analysis Predicts Space-Exposed Murine Livers

Limited sample sizes, high data dimensionality, and sensitivity to technical and biological variability of next generation sequencing (NGS), typically limits machine learning (ML) approaches in spaceflight studies that include radiation effects. However, pooling smaller studies while addressing intra- and inter-study variabilities allows for ML predictive modeling. Here, integration methods were applied to whole transcriptome shotgun sequencing (RNA-seq) data from six mouse liver GeneLab datasets (GLDS) (n ranging from 6 to 39 samples) from with a total of 81 spaceflight and ground-control samples to determine top features (i.e. genes) relevant to spaceflight including the effect of radiation exposure. RNASeq counts were normalized for each study, then merged and scaled across all datasets. Data dimensionality was reduced using a minimum redundancy maximum relevance (MRMR) methodology. Redundancy and relevance were computed using the Pearson correlation and F-statistic, respectively. The top 100 MRMR features were used to predict spaceflight vs. ground-control samples using Random Forest (RF), Support Vector Machine (SVM), and Linear Discriminant Analysis (LDA) classifiers with 5-fold cross validation (CV). Principal component analysis (PCA) on the complete feature set versus the MRMR features shows separation between spaceflight samples and ground controls (Figure 1A). The ML-based gene sets were compared against differential gene expression results obtained with DESeq2 from individual GLDS. Using all features or randomly sampled subsets at matching set sizes with MRMR, a maximum classifier accuracy of 69% was shown on the test set over 5 folds. For all classifiers, CV training using at least the top 30 MRMR genes show minimum 89% accuracy and 0.95 AUC value on the test set over 5 folds (Figure 1B). Baseline set analysis on differentially expressed genes (DEGs) identified using padj ≤ 0.05 show 295 DEGs that overlap at least two studies and 13 DEGs that overlap three studies (Figure 1C). Set analysis between the top 100 MRMR features and the DEGs showed 47 genes that overlap at least one study and 24 genes that overlap two studies. Over-representation analysis showed overlapping biological processes related to fatty acid and lipid metabolism which may indicate these processes in the response to spaceflight stressors. MRMR feature selection for the selected ML methods improve performance relative to a classifier built on all features or randomly sampled subsets. Permutation feature importance within the decorrelated MRMR features showed concordance in feature ranking between ML methods. A challenge of applying ML methods across heterogeneous NGS data is accounting for signal:noise. Here, signal validation across studies was shown by intersecting sets between top MRMR genes and DEGs from DESeq2 analysis. Non-intersecting sets introduce opportunity to explore genes relevant to differentiating space flight exposed groups and implementing ML methods across existing NGS datasets may overcome sample size limitations.

Machine Learning↗

Nanopores and nucleic acids: prospects for ultrarapid sequencing

DNA and RNA molecules can be detected as they are driven through a nanopore by an applied electric field at rates ranging from several hundred microseconds to a few milliseconds per molecule. The nanopore can rapidly discriminate between pyrimidine and purine segments along a single-stranded nucleic acid molecule. Nanopore detection and characterization of single molecules represents a new method for directly reading information encoded in linear polymers. If single-nucleotide resolution can be achieved, it is possible that nucleic acid sequences can be determined at rates exceeding a thousand bases per second.

Review↗

The rRNA evolution and procaryotic phylogeny

Studies of ribosomal RNA primary structure allow reconstruction of phylogenetic trees for prokaryotic organisms. Such studies reveal major dichotomy among the bacteria that separates them into eubacteria and archaebacteria. Both groupings are further segmented into several major divisions. The results obtained from 5S rRNA sequences are essentially the same as those obtained with the 16S rRNA data. In the case of Gram negative bacteria the ribosomal RNA sequencing results can also be directly compared with hybridization studies and cytochrome c sequencing studies. There is again excellent agreement among the several methods. It seems likely then that the overall picture of microbial phylogeny that is emerging from the RNA sequence studies is a good approximation of the true history of these organisms. The RNA data allow examination of the evolutionary process in a semi-quantitative way. The secondary structures of these RNAs are largely established. As a result it is possible to recognize examples of local structural evolution. Evolutionary pathways accounting for these events can be proposed and their probability can be assessed.

Fox, G. E.↗

Assembly of catalytic complexes from randomized oligonucleotides

The early evolution of life relied on catalytic RNAs (ribozymes) for central functions. To test whether early catalysts could have assembled from multiple short nucleic acid fragments in random sequence environments, we performed an in vitro selection from a short RNA library in the presence of 256 different DNA 20-nucleotide oligomers. High-throughput sequencing and biochemical analysis showed that most of the selected 1331 RNA sequences required at least one DNA for activity. Representatives for four of six RNA clusters that depended on DNA cofactors were active even when the 256 DNAs were replaced by completely random DNA 20-nucleotide oligomers. The formation of these catalytic complexes and the recruitment of oligonucleotide cofactors from completely random libraries demonstrate an important principle for the emergence of the earliest oligonucleotide catalysts.

Xu Han↗

Chance and necessity in the selection of nucleic acid catalysts

In Tom Stoppard's famous play [Rosencrantz and Guildenstern are Dead], the ill-fated heroes toss a coin 101 times. The first 100 times they do so the coin lands heads up. The chance of this happening is approximately 1 in 10(30), a sequence of events so rare that one might argue that it could only happen in such a delightful fiction. Similarly rare events, however, may underlie the origins of biological catalysis. What is the probability that an RNA, DNA, or protein molecule of a given random sequence will display a particular catalytic activity? The answer to this question determines whether a collection of such sequences, such as might result from prebiotic chemistry on the early earth, is extremely likely or unlikely to contain catalytically active molecules, and hence whether the origin of life itself is a virtually inevitable consequence of chemical laws or merely a bizarre fluke. The fact that a priori estimates of this probability, given by otherwise informed chemists and biologists, ranged from 10(-5) to 10(-50), inspired us to begin to address the question experimentally. As it turns out, the chance that a given random sequence RNA molecule will be able to catalyze an RNA polymerase-like phosphoryl transfer reaction is close to 1 in 10(13), rare enough, to be sure, but nevertheless in a range that is comfortably accessible by experiment. It is the purpose of this Account to describe the recent advances in combinatorial biochemistry that have made it possible for us to explore the abundance and diversity of catalysts existing in nucleic acid sequence space.

Review, Tutorial↗

Emergence of a replicating species from an in vitro RNA evolution reaction

The technique of self-sustained sequence replication allows isothermal amplification of DNA and RNA molecules in vitro. This method relies on the activities of a reverse transcriptase and a DNA-dependent RNA polymerase to amplify specific nucleic acid sequences. We have modified this protocol to allow selective amplification of RNAs that catalyze a particular chemical reaction. During an in vitro RNA evolution experiment employing this modified system, a unique class of "selfish" RNAs emerged and replicated to the exclusion of the intended RNAs. Members of this class of selfish molecules, termed RNA Z, amplify efficiently despite their inability to catalyze the target chemical reaction. Their amplification requires the action of both reverse transcriptase and RNA polymerase and involves the synthesis of both DNA and RNA replication intermediates. The proposed amplification mechanism for RNA Z involves the formation of a DNA hairpin that functions as a template for transcription by RNA polymerase. This arrangement links the two strands of the DNA, resulting in the production of RNA transcripts that contain an embedded RNA polymerase promoter sequence.

Non-NASA Center↗

A new version of the RDP (Ribosomal Database Project)

The Ribosomal Database Project (RDP-II), previously described by Maidak et al. [ Nucleic Acids Res. (1997), 25, 109-111], is now hosted by the Center for Microbial Ecology at Michigan State University. RDP-II is a curated database that offers ribosomal RNA (rRNA) nucleotide sequence data in aligned and unaligned forms, analysis services, and associated computer programs. During the past two years, data alignments have been updated and now include >9700 small subunit rRNA sequences. The recent development of an ObjectStore database will provide more rapid updating of data, better data accuracy and increased user access. RDP-II includes phylogenetically ordered alignments of rRNA sequences, derived phylogenetic trees, rRNA secondary structure diagrams, and various software programs for handling, analyzing and displaying alignments and trees. The data are available via anonymous ftp (ftp.cme.msu. edu) and WWW (http://www.cme.msu.edu/RDP). The WWW server provides ribosomal probe checking, approximate phylogenetic placement of user-submitted sequences, screening for possible chimeric rRNA sequences, automated alignment, and a suggested placement of an unknown sequence on an existing phylogenetic tree. Additional utilities also exist at RDP-II, including distance matrix, T-RFLP, and a Java-based viewer of the phylogenetic trees that can be used to create subtrees.

Non-NASA Center↗

Evolution of early life inferred from protein and ribonucleic acid sequences

The chemical structures of ferredoxin, 5S ribosomal RNA, and c-type cytochrome sequences have been employed to construct a phylogenetic tree which connects all major photosynthesizing organisms: the three types of bacteria, blue-green algae, and chloroplasts. Anaerobic and aerobic bacteria, eukaryotic cytoplasmic components and mitochondria are also included in the phylogenetic tree. Anaerobic nonphotosynthesizing bacteria similar to Clostridium were the earliest organisms, arising more than 3.2 billion years ago. Bacterial photosynthesis evolved nearly 3.0 billion years ago, while oxygen-evolving photosynthesis, originating in the blue-green algal line, came into being about 2.0 billion years ago. The phylogenetic tree supports the symbiotic theory of the origin of eukaryotes.

Dayhoff, M. O.↗

Analysis of in Vitro Evolution Reveals the Underlying Distribution of Catalytic Activity Among Random Sequences

The emergence of catalytic RNA is believed to have been a key event during the origin of life. Understanding how catalytic activity is distributed across random sequences is fundamental to estimating the probability that catalytic sequences would emerge. Here, we analyze the in vitro evolution of triphosphorylating ribozymes and translate their fitnesses into absolute estimates of catalytic activity for hundreds of ribozyme families. The analysis efficiently identified highly active ribozymes and estimated catalytic activity with good accuracy. The evolutionary dynamics follow Fisher’s Fundamental Theorem of Natural Selection and a corollary, permitting retrospective inference of the distribution of fitness and activity in the random sequence pool for the first time. The frequency distribution of rate constants appears to be log-normal, with a surprisingly steep dropoff at higher activity, consistent with a mechanism for the emergence of activity as the product of many independent contributions.

Abe Pressman↗

Evolution of heliobacteria: implications for photosynthetic reaction center complexes

The evolutionary position of the heliobacteria, a group of green photosynthetic bacteria with a photosynthetic apparatus functionally resembling Photosystem I of plants and cyanobacteria, has been investigated with respect to the evolutionary relationship to Gram-positive bacteria and cyanobacteria. On the basis of 16S rRNA sequence analysis, the heliobacteria appear to be most closely related to Gram-positive bacteria, but also an evolutionary link to cyanobacteria is evident. Interestingly, a 46-residue domain including the putative sixth membrane-spanning region of the heliobacterial reaction center protein show rather strong similarity (33% identity and 72% similarity) to a region including the sixth membrane-spanning region of the CP47 protein, a chlorophyll-binding core antenna polypeptide of Photosystem II. The N-terminal half of the heliobacterial reaction center polypeptide shows a moderate sequence similarity (22% identity over 232 residues) with the CP47 protein, which is significantly more than the similarity with the Photosystem I core polypeptides in this region. An evolutionary model for photosynthetic reaction center complexes is discussed, in which an ancestral homodimeric reaction center protein (possibly resembling the heliobacterial reaction center protein) with 11 membrane-spanning regions per polypeptide has diverged to give rise to the core of Photosystem I, Photosystem II, and of the photosynthetic apparatus in green, purple, and heliobacteria.

NASA Program Exobiology↗

Origins of the plant chloroplasts and mitochondria based on comparisons of 5S ribosomal RNAs

In this paper, we provide macromolecular comparisons utilizing the 5S ribosomal RNA structure to suggest extant bacteria that are the likely descendants of chloroplast and mitochondria endosymbionts. The genetic stability and near universality of the 5S ribosomal gene allows for a useful means to study ancient evolutionary changes by macromolecular comparisons. The value in current and future ribosomal RNA comparisons is in fine tuning the assignment of ancestors to the organelles and in establishing extant species likely to be descendants of bacteria involved in presumed multiple endosymbiotic events.

NASA Discipline Exobiology↗

The taxonomic status of "Halobacterium marismortui" from the Dead Sea: a comparison with Halobacterium vallismortis

A Halobacterium strain, isolated by Ginzburg et al. from the Dead Sea in the late 1960's, often referred to as "Halobacterium marismortui" or "Halobacterium of the Dead Sea" (deposited in the American Type Culture Collection as ATCC 43049) was compared with Halobacterium (Haloarcula) vallismortis ATCC 29715. The strains appeared to be very closely related, as shown by the near identity of their 5S and 16S ribosomal RNA's, and a large number of other common properties. Distinct differences exist, however, in cell morphology, and in their potency to utilize different sugars and other compounds.

NASA Discipline Exobiology↗

Abiotic Synthesis of Nucleic Acids: Hypochromicity and Future Research

The earliest forms of life would likely have a protocellular form, with a membrane encapsulating some form of linear charged polymer. These polymers could have enzymatic as well as genetic properties. We can simulate plausible prebiotic conditions in the laboratory to test hypotheses related to this concept. In earlier work we have shown that mononucleotides organized within a multilamellar lipid matrix can produce oligomers in the anhydrous phase of dehydration- rehydration cycles (Rajamani, 2008). If mononucleotides are in solution at millimolar concentrations, then oligomers resembling RNA are synthesized and exist in a steady state with their monomers DeGuzman, 2014). We have used conventional and novel techniques to demonstrate that secondary structures stabilized by hydrogen bonds may be present in the condensation products produced in dehydration- rehydration cycles that simulate hydrothermal fields that were present on the early Earth. Gel electrophoresis data corroborates the presence of up to 200-base pair length RNA fragments in products of Hydration-Dehydration experiments. Furthermore, hypochromicity measurements demonstrate a degree of hypochromicity found in single RNA strand of known sequence, as well as results that indicate this is true also for a sample of complementary strands of RNA. Analysis of ionic current signatures of known RNA hairpin molecule as measured using a nanopore detector indicate a significant variability in pattern, different from the signatures produced by DNA hairpin molecules. This informs how we may interpret nanopore data gathered from prebiotic simulations.

Glass, K.↗

Conserved gene clusters in bacterial genomes provide further support for the primacy of RNA

Five complete bacterial genome sequences have been released to the scientific community. These include four (eu)Bacteria, Haemophilus influenzae, Mycoplasma genitalium, M. pneumoniae, and Synechocystis PCC 6803, as well as one Archaeon, Methanococcus jannaschii. Features of organization shared by these genomes are likely to have arisen very early in the history of the bacteria and thus can be expected to provide further insight into the nature of early ancestors. Results of a genome comparison of these five organisms confirm earlier observations that gene order is remarkably unpreserved. There are, nevertheless, at least 16 clusters of two or more genes whose order remains the same among the four (eu)Bacteria and these are presumed to reflect conserved elements of coordinated gene expression that require gene proximity. Eight of these gene orders are essentially conserved in the Archaea as well. Many of these clusters are known to be regulated by RNA-level mechanisms in Escherichia coli, which supports the earlier suggestion that this type of regulation of gene expression may have arisen very early. We conclude that although the last common ancestor may have had a DNA genome, it likely was preceded by progenotes with an RNA genome.

Non-NASA Center↗

Continuous in vitro evolution of bacteriophage RNA polymerase promoters

Rapid in vitro evolution of bacteriophage T7, T3, and SP6 RNA polymerase promoters was achieved by a method that allows continuous enrichment of DNAs that contain functional promoter elements. This method exploits the ability of a special class of nucleic acid molecules to replicate continuously in the presence of both a reverse transcriptase and a DNA-dependent RNA polymerase. Replication involves the synthesis of both RNA and cDNA intermediates. The cDNA strand contains an embedded promoter sequence, which becomes converted to a functional double-stranded promoter element, leading to the production of RNA transcripts. Synthetic cDNAs, including those that contain randomized promoter sequences, can be used to initiate the amplification cycle. However, only those cDNAs that contain functional promoter sequences are able to produce RNA transcripts. Furthermore, each RNA transcript encodes the RNA polymerase promoter sequence that was responsible for initiation of its own transcription. Thus, the population of amplifying molecules quickly becomes enriched for those templates that encode functional promoters. Optimal promoter sequences for phage T7, T3, and SP6 RNA polymerase were identified after a 2-h amplification reaction, initiated in each case with a pool of synthetic cDNAs encoding greater than 10(10) promoter sequence variants.

Non-NASA Center↗

Atrogin-1, a muscle-specific F-box protein highly expressed during muscle atrophy

Muscle wasting is a debilitating consequence of fasting, inactivity, cancer, and other systemic diseases that results primarily from accelerated protein degradation by the ubiquitin-proteasome pathway. To identify key factors in this process, we have used cDNA microarrays to compare normal and atrophying muscles and found a unique gene fragment that is induced more than ninefold in muscles of fasted mice. We cloned this gene, which is expressed specifically in striated muscles. Because this mRNA also markedly increases in muscles atrophying because of diabetes, cancer, and renal failure, we named it atrogin-1. It contains a functional F-box domain that binds to Skp1 and thereby to Roc1 and Cul1, the other components of SCF-type Ub-protein ligases (E3s), as well as a nuclear localization sequence and PDZ-binding domain. On fasting, atrogin-1 mRNA levels increase specifically in skeletal muscle and before atrophy occurs. Atrogin-1 is one of the few examples of an F-box protein or Ub-protein ligase (E3) expressed in a tissue-specific manner and appears to be a critical component in the enhanced proteolysis leading to muscle atrophy in diverse diseases.

NASA Discipline Musculoskeletal↗

Archaeal phylogeny: reexamination of the phylogenetic position of Archaeoglobus fulgidus in light of certain composition-induced artifacts

A major and too little recognized source of artifact in phylogenetic analysis of molecular sequence data is compositional difference among sequences. The problem becomes particularly acute when alignments contain ribosomal RNAs from both mesophilic and thermophilic species. Among prokaryotes the latter are considerably higher in G + C content than the former, which often results in artificial clustering of thermophilic lineages and their being placed artificially deep in phylogenetic trees. In this communication we review archaeal phylogeny in the light of this consideration, focusing in particular on the phylogenetic position of the sulfate reducing species Archaeoglobus fulgidus, using both 16S rRNA and 23S rRNA sequences. The analysis shows clearly that the previously reported deep branching of the A. fulgidus lineage (very near the base of the euryarchaeal side of the archaeal tree) is incorrect, and that the lineage actually groups with a previously recognized unit that comprises the Methanomicrobiales and extreme halophiles.

NASA Discipline Exobiology↗