Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DNA sequencing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Fractals in biology and medicine

Our purpose is to describe some recent progress in applying fractal concepts to systems of relevance to biology and medicine. We review several biological systems characterized by fractal geometry, with a particular focus on the long-range power-law correlations found recently in DNA sequences containing noncoding material. Furthermore, we discuss the finding that the exponent alpha quantifying these long-range correlations ("fractal complexity") is smaller for coding than for noncoding sequences. We also discuss the application of fractal scaling analysis to the dynamics of heartbeat regulation, and report the recent finding that the normal heart is characterized by long-range "anticorrelations" which are absent in the diseased heart.

Review↗

Microbial monitoring of spacecraft and associated environments

Rapid microbial monitoring technologies are invaluable in assessing contamination of spacecraft and associated environments. Universal and widespread elements of microbial structure and chemistry are logical targets for assessing microbial burden. Several biomarkers such as ATP, LPS, and DNA (ribosomal or spore-specific), were targeted to quantify either total bioburden or specific types of microbial contamination. The findings of these assays were compared with conventional, culture-dependent methods. This review evaluates the applicability and efficacy of some of these methods in monitoring the microbial burden of spacecraft and associated environments. Samples were collected from the surfaces of spacecraft, from surfaces of assembly facilities, and from drinking water reservoirs aboard the International Space Station (ISS). Culture-dependent techniques found species of Bacillus to be dominant on these surfaces. In contrast, rapid, culture-independent techniques revealed the presence of many Gram-positive and Gram-negative microorganisms, as well as actinomycetes and fungi. These included both cultivable and noncultivable microbes, findings further confirmed by DNA-based microbial detection techniques. Although the ISS drinking water was devoid of cultivable microbes, molecular-based techniques retrieved DNA sequences of numerous opportunistic pathogens. Each of the methods tested in this study has its advantages, and by coupling two or more of these techniques even more reliable information as to microbial burden is rapidly obtained. Copyright 2004 Springer-Verlag.

Environmental Monitoring/methods↗

Novel Approach to Quantification of Telomere Length with Direct Nanopore Sequencing and PCR Amplification

The ends of human chromosomes contain telomeres, or tandem arrays of repeating DNA sequences capped by multiple associated proteins that protect chromosomal ends from degradation. Telomeres function to preserve genomic stability by preventing natural chromosomal ends from being recognized as broken DNA double-strand breaks and triggering inappropriate DNA damage responses. Mounting evidence shows telomere length is an inherited trait that decreases with cellular division and normal aging. In addition, telomere length also appears to be influenced by other factors such as cellular oxidative stress, radiation and mechanical unloading of tissues as in microgravity. To measure these potential effects of the space environment on telomere lengths and cellular aging and regenerative potential we developed a novel telomere measurement approach based on nanopore sequencing of PCR amplified bar-coded chromosome termini. Specifically, telomeres can be directly enriched using barcode sequences ligated to the end of a free end- repaired telomere using the WetLab-2 facility SmartCycler on ISS. Prior to the ligation and amplification protocol a proteinase K digestion of capping proteins followed by a single 95-degree C heat denaturation of the protease is included. After digestion and bar-code ligation, PCR amplification will initiate with the ligated barcoded sequence, suppressing amplification of intra-genomic fragments and resulting in long read barcoded telomere amplicons including the nanopore motor protein sequences. Purified PCR amplicons are then used for nanopore sequencing library generation by simple addition of motor proteins and sequencing library is loaded into the MinION nanopore DNA-sequencer. Amplicon sequence reads from the nanopore device can be base-called quickly on ISS due to barcoding ligation and subsequent PCR amplification enhancing the telomere sequence resolution. If successfully implemented on ISS this technique will provide a novel means of measuring regenerative ability of somatic stem cells in astronauts, and of determining whether spaceflight in microgravity alters their telomere lengths and causes premature cellular aging.

Ma, Kristin R.↗

Studying Functions of All Yeast Genes Simultaneously

A method of studying the functions of all the genes of a given species of microorganism simultaneously has been developed in experiments on Saccharomyces cerevisiae (commonly known as baker's or brewer's yeast). It is already known that many yeast genes perform functions similar to those of corresponding human genes; therefore, by facilitating understanding of yeast genes, the method may ultimately also contribute to the knowledge needed to treat some diseases in humans. Because of the complexity of the method and the highly specialized nature of the underlying knowledge, it is possible to give only a brief and sketchy summary here. The method involves the use of unique synthetic deoxyribonucleic acid (DNA) sequences that are denoted as DNA bar codes because of their utility as molecular labels. The method also involves the disruption of gene functions through deletion of genes. Saccharomyces cerevisiae is a particularly powerful experimental system in that multiple deletion strains easily can be pooled for parallel growth assays. Individual deletion strains recently have been created for 5,918 open reading frames, representing nearly all of the estimated 6,000 genetic loci of Saccharomyces cerevisiae. Tagging of each deletion strain with one or two unique 20-nucleotide sequences enables identification of genes affected by specific growth conditions, without prior knowledge of gene functions. Hybridization of bar-code DNA to oligonucleotide arrays can be used to measure the growth rate of each strain over several cell-division generations. The growth rate thus measured serves as an index of the fitness of the strain.

Stolc, Viktor↗

Mosaic organization of DNA nucleotides

Long-range power-law correlations have been reported recently for DNA sequences containing noncoding regions. We address the question of whether such correlations may be a trivial consequence of the known mosaic structure ("patchiness") of DNA. We analyze two classes of controls consisting of patchy nucleotide sequences generated by different algorithms--one without and one with long-range power-law correlations. Although both types of sequences are highly heterogenous, they are quantitatively distinguishable by an alternative fluctuation analysis method that differentiates local patchiness from long-range correlations. Application of this analysis to selected DNA sequences demonstrates that patchiness is not sufficient to account for long-range correlation properties.

NASA Discipline Number 14-10↗

Unexpected substrate specificity of T4 DNA ligase revealed by in vitro selection

We have used in vitro selection techniques to characterize DNA sequences that are ligated efficiently by T4 DNA ligase. We find that the ensemble of selected sequences ligates about 50 times as efficiently as the random mixture of sequences used as the input for selection. Surprisingly many of the selected sequences failed to produce a match at or close to the ligation junction. None of the 20 selected oligomers that we sequenced produced a match two bases upstream from the ligation junction.

Harada, Kazuo↗

Modeling the Activity of Single Genes

The central dogma of molecular biology states that information is stored in DNA, transcribed to messenger RNA (mRNA) and then translated into proteins. This picture is significantly augmentated when we consider the action of certain proteins in regulating transcription. These transcription factors provide a feedback pathway by which genes can regulate one another's expression as mRNA and then as protein. To review: DNA, RNA and proteins have different functions. DNA is the molecular storehouse of genetic information. When cells divide, the DNA is replicated, so that each daughter cell maintains the same genetic information as the mother cell. RNA acts as a go-between from DNA to proteins. Only a single copy of DNA is present, but multiple copies of the same piece of RNA may be present, allowing cells to make huge amounts of protein. In eukaryotes (organisms with a nucleus), DNA is found in the nucleus only. RNA is copied in the nucleus then translocates(moves) outside the nucleus, where it is transcribed into proteins. Along the way, the RNA may be spliced, i.e., may have pieces cut out. RNA then attaches to ribosomes and is translated to proteins. Proteins are the machinery of the cell other than DNA and RNA, all the complex molecules of the cell are proteins. Proteins are specialized machines, each of which fulfills its own task, which may be transporting oxygen, catalyzing reactions, or responding to extracellular signals, just to name a few. One of the more interesting functions a protein may have is binding directly or indirectly to DNA to perform transcriptional regulation, thus forming a closed feedback loop of gene regulation. The structure of DNA and the central dogma were understood in the 50s; in the early 80s it became possible to make arbitrary modifications to DNA and use cellular machinery to transcribe and translate the resulting genes; more recently, genomes (i.e., the complete DNA sequence) of many organisms have been sequenced. This large-scale sequencing began with simple organisms, viruses and bacteria, progressed to eukaryotes such as yeast, and more recently (1998) progressed to a multi-cellular animal, the nematode Caenorhabditis elegans. Sequencers have now moved on to the fruit fly Drosophila melanogaster, whose sequence is slated for completion by the end of 1999. The human genome project is expected to determine the complete sequence of all 3 billion bases of human DNA within the next five years. In the wake of genome-scale sequencing, further instrumentation is being developed to assay gene expression and function on a comparably large scale. Much of the work in computational biology focuses on computational tools used in sequencing, finding genes that are related to a particular gene, finding which parts of the DNA code for proteins and which do not, understanding what proteins will be formed from a given length of DNA, predicting how the proteins will fold from a one-dimensional structure into a three dimensional structure, and so on. Much less computational work has been done regarding the function of proteins. One reason for this is that different proteins function very differently, and so work on protein function is very specific to certain classes of proteins. There are, for example, proteins such enzymes that catalyze various intracellular reactions, receptors that respond to extracellular signals and ion channels that regulate the flow of charged particles into and out of the cell. In this chapter, we will consider a particular class of proteins called transcription factors(TFs), which are responsible for regulating when a certain gene is expressed in a certain cell, which cells it is express in, and how much is expressed. Understanding these processes will involve developing a deeper understanding of transcription, translation, and the cellular processes that control those processes. All of these elements fall under the aegis of gene regulation or more narrowly transcriptional regulation. Some of the key questions in gene regulation are: What genes are expressed in a certain cell at a certain time? How does gene expression differ from cell to cell in a multicellular organism? Which proteins act as transcription factors, i.e., are important in regulating gene expression? From questions like these, we hope to understand which genes are important for various macroscopic processes. Nearly all of the cells of a multicellular organism contain the same DNA. Yet this same genetic information yields a large number of different cell types. The fundamental difference between a neuron and a liver cell, for example, is which genes are expressed. Thus understanding gene regulation is an important step in understanding development. Furthermore, understanding the usual genes that are expressed in cells may give important clues about various diseases. Some diseases, such as sickle cell anemia and cystic fibrosis, are caused by defects in single, non-regulatory genes; others, such as certain cancers, are caused when the cellular control circuitry malfunctions - an understanding of these diseases will involve pathways of multiple interacting gene products. There are numerous challenges in the area of understanding and modeling gene regulation. First and foremost, biologists would like to develop a deeper understanding of the processes involved, including which genes and families of genes are important, how they interact, etc. From a computation point of view, there has been embarrassingly little work done. In this chapter there are many areas in which we can phrase meaningful, non-trivial computational questions, but questions that have not been addressed. Some of these are purely computational (what is a good algorithm for dealing with a model of type X) and others are more mathematical (given a system with certain characteristics, what sort of model can one use? How does one find biochemical parameters from system-level behavior using as few experiments as possible?). In addition to biological and algorithmic problems, there is also the ever-present issue of theoretical biology - what general principles can be derived from these systems, what can one do with models other than just simulate time-courses, what can be deduced about a class of systems without knowing all the details? The fundamental challenge to computationalists and theorists is to add value to the biology - to use models, modeling techniques and algorithms to understand the biology in new ways.

Mjolsness, Eric↗

Evidence of SV40 infections in hospitalized children

Simian virus 40 (SV40) is known to have contaminated poliovirus vaccines used between 1955 and 1963. Accumulating reports have described the presence of SV40 DNA in human tumors and normal tissues, although the significance of human infections by SV40 is unknown. We investigated whether unselected hospitalized children had evidence of SV40 infections and whether any clinical correlations were apparent. Serum samples were examined for SV40 neutralizing antibody using a specific plaque reduction test; of 337 samples tested, 20 (5.9%) had antibody to SV40. Seropositivity increased with age and was significantly associated with kidney transplants (6 of 15 [40%] positive, P < .001). Many of the antibody-positive patients had impaired immune systems. Molecular assays (polymerase chain reaction and DNA sequence analysis) on archival tissue specimens confirmed the presence of SV40 DNA in 4 of the antibody-positive patients. This study, using 2 independent assays, shows the presence of SV40 infections in children born after 1980. We conclude that SV40 causes natural infections in humans.

NASA Discipline Regulatory Physiology↗

Highly efficient and simple SSPER and rrPCR approaches for the accurate site-directed mutagenesis of large and small plasmids

Advances are needed in the site-directed mutagenesis of large plasmids for protein structure-function studies, as current methods are often inefficient, complicated and time-consuming. Here two new methods are reported that overcome these difficulties, namely the single primer extension reaction (SSPER) strategy that reaches 100% efficiency and the reduce recycle PCR (rrPCR) method that is advantageous in generating single and pairwise combinations of mutations. Both methods are distinguished from current technologies by the addition of a step that easily removes the oligonucleotide primer(s) after the first reaction, thus allowing for the addition of a second reaction in chronological sequence to generate and isolate the appropriate DNA product with the site-directed mutation(s). High efficiency of the methods is demonstrated by generating single and paired combinations of the 11 site-directed mutations targeted on 5 different plasmid DNA templates ranging from 10 to 12 kb and 57–60% GC-content at a rate of 50–100%. Overall, the methods are demonstrated to be (i) highly accurate, allowing for screening of plasmids by DNA sequencing, (ii) streamlined to generate the mutations within a single day, (iii) cost-effective in requiring only two primers and two enzymes (DpnI and a proofreading DNA polymerase), (iv) straightforward in primer design, (v) applicable for both large and small plasmids, and (vi) easily implemented by entry level researchers.

59 BASIC BIOLOGICAL SCIENCES↗

Biomolecular Analysis Capability for Cellular and Omics Research on the International Space Station

International Space Station (ISS) assembly complete ushered a new era focused on utilization of this state-of-the-art orbiting laboratory to advance science and technology research in a wide array of disciplines, with benefits to Earth and space exploration. ISS enabling capability for research in cellular and molecular biology includes equipment for in situ, on-orbit analysis of biomolecules. Applications of this growing capability range from biomedicine and biotechnology to the emerging field of Omics. For example, Biomolecule Sequencer is a space-based miniature DNA sequencer that provides nucleotide sequence data for entire samples, which may be used for purposes such as microorganism identification and astrobiology. It complements the use of WetLab-2 SmartCycler"TradeMark", which extracts RNA and provides real-time quantitative gene expression data analysis from biospecimens sampled or cultured onboard the ISS, for downlink to ground investigators, with applications ranging from clinical tissue evaluation to multigenerational assessment of organismal alterations. And the Genes in Space-1 investigation, aimed at examining epigenetic changes, employs polymerase chain reaction to detect immune system alterations. In addition, an increasing assortment of tools to visualize the subcellular distribution of tagged macromolecules is becoming available onboard the ISS. For instance, the NASA LMM (Light Microscopy Module) is a flexible light microscopy imaging facility that enables imaging of physical and biological microscopic phenomena in microgravity. Another light microscopy system modified for use in space to image life sciences payloads is initially used by the Heart Cells investigation ("Effects of Microgravity on Stem Cell-Derived Cardiomyocytes for Human Cardiovascular Disease Modeling and Drug Discovery"). Also, the JAXA Microscope system can perform remotely controllable light, phase-contrast, and fluorescent observations. And upcoming confocal microscopy capability will allow for optical sectioning of biological tissues to determine microanatomical localization of biomarkers. Furthermore, NASA's geneLAB effort addresses integration of genomic, epigenomic, transcriptomic, proteomic and metabolomic datasets, by applying an innovative open source science platform for multi-investigator high throughput utilization of the ISS. In sum, the expanding ISS capability for analysis of biomolecules is enabling innovative research in a broad spectrum of areas such as cellular and molecular biology, biotechnology, tissue engineering, biomedicine, and Omics, providing manifold benefits for humanity.

Guinart-Ramirez, Y.↗

Mitochondrial gene arrangement of the horseshoe crab Limulus polyphemus L.: conservation of major features among arthropod classes

Numerous complete mitochondrial DNA sequences have been determined for species within two arthropod groups, insects and crustaceans, but there are none for a third, the chelicerates. Most mitochondrial gene arrangements reported for crustaceans and insect species are identical or nearly identical to that of Drosophila yakuba. Sequences across 36 of the gene boundaries in the mitochondrial DNA (mtDNA) of a representative chelicerate. Limulus polyphemus L., also reveal an arrangement like that of Drosophila yakuba. Only the position of the tRNA(LEU)(UUR) gene differs; in Limulus it is between the genes for tRNA(LEU)(CUN) and ND1. This positioning is also found in onychophorans, mollusks, and annelids, but not in insects and crustaceans, and indicates that tRNA(LEU)(CUN)-tRNA(LEU)(UUR)-ND1 was the ancestral gene arrangement for these groups, as suggested earlier. There are no differences in the relative arrangements of protein-coding and ribosomal RNA genes between Limulus and Drosophila, and none have been observed within arthropods. The high degree of similarity of mitochondrial gene arrangements within arthropods is striking, since some taxa last shared a common ancestor before the Cambrian, and contrasts with the extensive mtDNA rearrangements occasionally observed within some other metazoan phyla (e.g., mollusks and nematodes).

Non-NASA Center↗

Sequence-specific dynamic DNA bending explains mitochondrial TFAM’s dual role in DNA packaging and transcription initiation

Abstract Mitochondrial transcription factor A (TFAM) employs DNA bending to package mitochondrial DNA (mtDNA) into nucleoids and recruit mitochondrial RNA polymerase (POLRMT) at specific promoter sites, light strand promoter (LSP) and heavy strand promoter (HSP). Herein, we characterize the conformational dynamics of TFAM on promoter and non-promoter sequences using single-molecule fluorescence resonance energy transfer (smFRET) and single-molecule protein-induced fluorescence enhancement (smPIFE) methods. The DNA-TFAM complexes dynamically transition between partially and fully bent DNA conformational states. The bending/unbending transition rates and bending stability are DNA sequence-dependent—LSP forms the most stable fully bent complex and the non-specific sequence the least, which correlates with the lifetimes and affinities of TFAM with these DNA sequences. By quantifying the dynamic nature of the DNA-TFAM complexes, our study provides insights into how TFAM acts as a multifunctional protein through the DNA bending states to achieve sequence specificity and fidelity in mitochondrial transcription while performing mtDNA packaging.

59 BASIC BIOLOGICAL SCIENCES↗

Genome dependent Cas9/gRNA search time underlies sequence dependent gRNA activity

Abstract CRISPR-Cas9 is a powerful DNA editing tool. A gRNA directs Cas9 to cleave any DNA sequence with a PAM. However, some gRNA sequences mediate cleavage at higher efficiencies than others. To understand this, numerous studies have screened large gRNA libraries and developed algorithms to predict gRNA sequence dependent activity. These algorithms do not predict other datasets as well as their training dataset and do not predict well between species. Here, to better understand these discrepancies, we retrospectively examine sequence features that impact gRNA activity in 44 published data sets. We find strong evidence that gRNA sequence dependent activity is largely influenced by the ability of the Cas9/gRNA complex to find the target site rather than activity at the target site and that this drives sequence dependent differences in gRNA activity between different species. This understanding will help guide future work to understand Cas9 activity as well as efforts to identify optimal gRNAs and improve Cas9 variants.

59 BASIC BIOLOGICAL SCIENCES↗

DL-TODA: A Deep Learning Tool for Omics Data Analysis

Metagenomics is a technique for genome-wide profiling of microbiomes; this technique generates billions of DNA sequences called reads. Given the multiplication of metagenomic projects, computational tools are necessary to enable the efficient and accurate classification of metagenomic reads without needing to construct a reference database. The program DL-TODA presented here aims to classify metagenomic reads using a deep learning model trained on over 3000 bacterial species. A convolutional neural network architecture originally designed for computer vision was applied for the modeling of species-specific features. Using synthetic testing data simulated with 2454 genomes from 639 species, DL-TODA was shown to classify nearly 75% of the reads with high confidence. The classification accuracy of DL-TODA was over 0.98 at taxonomic ranks above the genus level, making it comparable with Kraken2 and Centrifuge, two state-of-the-art taxonomic classification tools. DL-TODA also achieved an accuracy of 0.97 at the species level, which is higher than 0.93 by Kraken2 and 0.85 by Centrifuge on the same test set. Application of DL-TODA to the human oral and cropland soil metagenomes further demonstrated its use in analyzing microbiomes from diverse environments. Compared to Centrifuge and Kraken2, DL-TODA predicted distinct relative abundance rankings and is less biased toward a single taxon.

59 BASIC BIOLOGICAL SCIENCES↗

Scaling in nature: from DNA through heartbeats to weather

The purpose of this report is to describe some recent progress in applying scaling concepts to various systems in nature. We review several systems characterized by scaling laws such as DNA sequences, heartbeat rates and weather variations. We discuss the finding that the exponent alpha quantifying the scaling in DNA in smaller for coding than for noncoding sequences. We also discuss the application of fractal scaling analysis to the dynamics of heartbeat regulation, and report the recent finding that the scaling exponent alpha is smaller during sleep periods compared to wake periods. We also discuss the recent findings that suggest a universal scaling exponent characterizing the weather fluctuations.

Review↗

ULTRA-effective labeling of tandem repeats in genomic sequence

In the age of long read sequencing, genomics researchers now have access to accurate repetitive DNA sequence (including satellites) that, due to the limitations of short read-sequencing, could previously be observed only as unmappable fragments. Tools that annotate repetitive sequence are now more important than ever, so that we can better understand newly uncovered repetitive sequences, and also so that we can mitigate errors in bioinformatic software caused by those repetitive sequences. To that end, we introduce the 1.0 release of our tool for identifying and annotating locally repetitive sequence, ULTRA Locates Tandemly Repetitive Areas (ULTRA). ULTRA is fast enough to use as part of an efficient annotation pipeline, produces state-of-the-art reliable coverage of repetitive regions containing many mutations, and provides interpretable statistics and labels for repetitive regions.

59 BASIC BIOLOGICAL SCIENCES↗