Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Sequencing data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Evolution of the rhodospirillaceae and mitochondria - A view based on sequence data

New sequence data from several protein families and from 5S ribosomal RNA confirm and elaborate a previously proposed description of the phylogenetic connections between a variety of bacteria and the eukaryotes. Probably, the first organisms were nonphotosynthetic anaerobic prokaryotes, which were followed soon by photosynthetic anaerobes. From this photosynthetic stock, the aerobic line to Pseudomonadacae, Rhodospirillaceae, and blue-greens arose. The eukaryotes derived genetic material from the symbioses of at least three separate bacterial lines. Ancestors of Rhodopseudomonas globiformis gave rise to the eukaryote mitochondria, probably through at least three separate symbioses, one early on the flagellate line, one on the ciliate line, and one on the stem to the multicellular forms.

Dayhoff, M. O.

Elevating the Quality of Space Omics Sequencing Data: Innovations and Methodologies from NASA GeneLab Sample Processing Laboratory

NASA’s GeneLab, part of the NASA Open Science Data Repository, is a space-related database that hosts a diverse range of transcriptomics, proteomics, epigenomics and genomics data. The NASA GeneLab Sample Processing Laboratory (SPL) generates omics data from biological experiments conducted aboard the International Space Station, Space Shuttle and space related ground experiments, this omics data then hosted on the GeneLab repository. Samples generated such experiments pose numerous technical challenges such as small experimental sample size, variance in dissection times, limited tissue preservation methods, prolonged storage time, and more. GeneLab SPL team had developed specialized expertise in nucleic acid extraction, library preparation and sequencing of such biological samples via extensive training and years of experience. In order to ensure data accuracy and consistency across experiments, SPL has developed standardized protocols for each species and tissue type. These protocols in conjunction with quality control metrics and data standards are crucial in generating of high-quality data. SPL protocols and standards have been developed in collaboration with the scientific community and had been made publicly available on the GeneLab portal, guaranteeing comparability of datasets across spaceflight experiments. To ensure reliability of data generation, SPL leverages cutting-edge innovations in laboratory automation for sample processing. By leveraging these state-of-the-art platforms, SPL achieves high levels of data reproducibility while significantly minimizing sources of bias and variability, especially across experiments with large numbers of samples. Over the past few years, the space biology investigator community has accessed SPL-generated data from the Open Science Data Repository for a myriad of data re-analysis and re-use studies. We observe a trend that in-house SPL-generated data consistently outperforms outsourced sequencing data in terms of technical standards, quality control metrics, timeliness of data delivery, and sequencing and reagent efficiency. Superior data generation has and will continue to enable discoveries in disease, diagnostic tools, and the biological effects of long duration spaceflight.

GeneLab

Elevating the Quality of Space Omics Sequencing Data: Innovations and Methodologies from NASA GeneLab Sample Processing Laboratory

NASA’s GeneLab, part of the NASA Open Science Data Repository, is a space-related database that hosts a diverse range of transcriptomics, proteomics, epigenomics and genomics data. The NASA GeneLab Sample Processing Laboratory (SPL) generates omics data from biological experiments conducted aboard the International Space Station, Space Shuttle and space related ground experiments, this omics data then hosted on the GeneLab repository. Samples generated such experiments pose numerous technical challenges such as small experimental sample size, variance in dissection times, limited tissue preservation methods, prolonged storage time, and more. GeneLab SPL team had developed specialized expertise in nucleic acid extraction, library preparation and sequencing of such biological samples via extensive training and years of experience. In order to ensure data accuracy and consistency across experiments, SPL has developed standardized protocols for each species and tissue type. These protocols in conjunction with quality control metrics and data standards are crucial in generating of high-quality data. SPL protocols and standards have been developed in collaboration with the scientific community and had been made publicly available on the GeneLab portal, guaranteeing comparability of datasets across spaceflight experiments. To ensure reliability of data generation, SPL leverages cutting-edge innovations in laboratory automation for sample processing. By leveraging these state-of-the-art platforms, SPL achieves high levels of data reproducibility while significantly minimizing sources of bias and variability, especially across experiments with large numbers of samples. Over the past few years, the space biology investigator community has accessed SPL-generated data from the Open Science Data Repository for a myriad of data re-analysis and re-use studies. We observe a trend that in-house SPL-generated data consistently outperforms outsourced sequencing data in terms of technical standards, quality control metrics, timeliness of data delivery, and sequencing and reagent efficiency. Superior data generation has and will continue to enable discoveries in disease, diagnostic tools, and the biological effects of long duration spaceflight.

GeneLab

Iterative pass optimization of sequence data

The problem of determining the minimum-cost hypothetical ancestral sequences for a given cladogram is known to be NP-complete. This "tree alignment" problem has motivated the considerable effort placed in multiple sequence alignment procedures. Wheeler in 1996 proposed a heuristic method, direct optimization, to calculate cladogram costs without the intervention of multiple sequence alignment. This method, though more efficient in time and more effective in cladogram length than many alignment-based procedures, greedily optimizes nodes based on descendent information only. In their proposal of an exact multiple alignment solution, Sankoff et al. in 1976 described a heuristic procedure--the iterative improvement method--to create alignments at internal nodes by solving a series of median problems. The combination of a three-sequence direct optimization with iterative improvement and a branch-length-based cladogram cost procedure, provides an algorithm that frequently results in superior (i.e., lower) cladogram costs. This iterative pass optimization is both computation and memory intensive, but economies can be made to reduce this burden. An example in arthropod systematics is discussed. c2003 The Willi Hennig Society. Published by Elsevier Science (USA). All rights reserved.

NASA Discipline Evolutionary Biology

Space Flown Rodent Liver RNA Sequencing Data for Machine Learning in Space Biology Research

High-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. Data analysis has been accelerated in recent years by the adoption of artificial intelligence (AI) and machine learning (ML) techniques by biomedical researchers. In space biology research, RNAseq datasets from space-flown experimental samples are critical for characterizing the gene expression aberrations associated with exposure to spaceflight stressors. However, space biological experiments tend to be very low sample size, so identifying proper AI/ML algorithms for sequencing data analysis is an ongoing challenge since these algorithms typically require large sample size. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML”, focused on creating datasets meant for three main applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. These scientific benchmarks consist of an AI-ready dataset and a reference implementation on a specific scientific question. In this work, we focused on generating standardized datasets to allow the scientific community to benchmark AI/ML algorithms in the domain of space biology. We present here a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data as a collaboration between the NASA AI4LS (Artificial Intelligence for Life Sciences) working group. and NASA’s SMD. This dataset consists of space-flown and ground control mouse liver found in the NASA GeneLab omics database. However, to amplify the small sample number (n=112 samples) for ML purposes, we employ Gaussian noise and a generative adversarial network to extend this dataset to 6,000 synthetic samples, matching the original gene expression characteristics.

James Casaletto

Insights into the phylogenetic positions of photosynthetic bacteria obtained from 5S rRNA and 16S rRNA sequence data

Comparisons of complete 16S ribosomal ribonucleic acid (rRNA) sequences established that the secondary structure of these molecules is highly conserved. Earlier work with 5S rRNA secondary structure revealed that when structural conservation exists the alignment of sequences is straightforward. The constancy of structure implies minimal functional change. Under these conditions a uniform evolutionary rate can be expected so that conditions are favorable for phylogenetic tree construction.

Fox, G. E.

Homology and the optimization of DNA sequence data

Three methods of nucleotide character analysis are discussed. Their implications for molecular sequence homology and phylogenetic analysis are compared. The criterion of inter-data set congruence, both character based and topological, are applied to two data sets to elucidate and potentially discriminate among these parsimony-based ideas. c2001 The Willi Hennig Society.

Non-NASA Center

Discovering methylated DNA motifs in bacterial nanopore sequencing data with MIJAMP

Abstract Bacterial DNA methylation is involved in diverse cellular functions, including modulation of gene expression, DNA repair, and restriction–modification systems for defense against viruses and other foreign DNA. Restriction systems hinder efforts to engineer organisms to produce fuels and chemicals from waste and renewable feedstocks by degrading DNA during transformation. Methylome analysis allows identification of motifs within a bacterial chromosome that may be targeted by native restriction enzymes. Further expression of the corresponding methyltransferases in Escherichia coli allows plasmid DNA to be protected from restriction in the target organism, thereby drastically enhancing transformation efficiency. Nanopore sequencing can detect methylated bases, but software is needed to transform modified base coordinates into methylated motifs. Here, we develop MIJAMP (MIJAMP Is Just A MethylBED Parser), a software package that was developed to discover methylated motifs from the output of ONT’s Modkit or other data in the methylBED format. MIJAMP employs a human-driven refinement strategy that empirically validates all motifs against genome-wide methylation data, thus eliminating incorrect motifs. MIJAMP also reports methylation data on specific, user-defined motifs. Using MIJAMP, we determined the methylated motifs both in a control strain (wild-type E. coli) and in Synecococcus sp. strain PCC7002, laying the foundation for improved transformation in this organism. MIJAMP is available at https://code.ornl.gov/alexander-public/mijamp/. One Sentence Summary: Here we describe software written to discover DNA methylation motifs from nanopore sequencing data.

59 BASIC BIOLOGICAL SCIENCES

Multimodal framework for the joint analysis of single-cell RNA and T cell receptor sequencing data predicts T cell response to cancer immunotherapy

T cell states are prognostic in different cancer types. Recent technologies enable joint profiling of T cell RNA and T cell receptor (TCR) sequences at single-cell resolution. Here we present the TCR-RNA Integrating Model (TRIM), a multi-modal variational autoencoder framework that integrates RNA-TCR data and predicts T cell clonality and transcriptional states. TRIM learns a shared representation of the data conditioned on patient, tissue source, and treatment timepoint. We applied TRIM to three independent datasets that included T cells collected before and after checkpoint inhibitor treatment, sourced either from blood and tumor biopsies in patients with head and neck squamous cell carcinoma and colorectal cancer, or from tumor and adjacent tissue in a pan-cancer dataset. In all settings, TRIM accurately predicted intra-tumor T cell clonal expansion and transcriptional status based on T cells from blood or normal tissue before treatment, demonstrating its utility in modeling multimodal T cell data and predicting T cell response to treatment and disease progression.

60 APPLIED LIFE SCIENCES

Next-Generation Sequencing Data from a CUT&RUN Study of R. toruloides IFO0880 Cse4 and Orc1 Binding Sites

Rhodotorula toruloides has been increasingly explored as a host for bioproduction of lipids, fatty acid derivatives and terpenoids. Various genetic tools have been developed, but neither a centromere nor an autonomously replicating sequence (ARS), both necessary elements for stable episomal plasmid maintenance, has yet been reported. In this study, cleavage under targets and release using nuclease (CUT&RUN), a method used for genome-wide mapping of DNA–protein interactions, was used to identify R. toruloides IFO0880 genomic regions associated with the centromeric histone H3 protein Cse4, a marker of centromeric DNA. Fifteen putative centromeres ranging from 8 to 19 kb in length were identified and analyzed, and four were tested for, but did not show, ARS activity. These centromeric sequences contained below average GC content, corresponded to transcriptional cold spots, were primarily nonrepetitive and shared some vestigial transposon-related sequences but otherwise did not show significant sequence conservation. Future efforts to identify an ARS in this yeast can utilize these centromeric DNA sequences to improve the stability of episomal plasmids derived from putative ARS elements.

Genome Engineering

Deciphering Spaceflight Medical Risks Using High-Performance Computing and Next Generation Sequencing Data From Model Organisms

Somatic mutations are acquired point mutations and other forms of genetic alteration in the DNA of somatic cells in the body. Unlike germline mutations, which can be passed on from one individual to another, somatic mutations are not heritable. Somatic mutation (also called genetic sequence variation) has been recognized for decades as an important mechanism for initiating the development of cancer. Now, with the advent of next generation sequencing (NGS), the study of somatic mutation has become much more accessible, enabling genome scientists to characterize somatic mutations in a comprehensive fashion across the entire genome. From recent studies, it is now apparent that other disease processes besides cancer may be influenced by somatic mutation as well, including degenerative processes and inflammatory diseases. Cancer, degenerative diseases and inflammatory processes are all of concern in the setting of spaceflight. Study of somatic mutation analysis, therefore, is an important new tool for examining some of the earliest changes in the genome that lead to disease. This approach has tremendous potential for NASA, for analysis of both model organisms and humans.

Somatic mutation on ISS

Sequence data - Magnitude and implications of some ambiguities.

A stochastic model is applied to the divergence of the horse-pig lineage from a common ansestor in terms of the alpha and beta chains of hemoglobin and fibrinopeptides. The results are compared with those based on the minimum mutation distance model of Fitch (1972). Buckwheat and cauliflower cytochrome c sequences are analyzed to demonstrate their ambiguities. A comparative analysis of evolutionary rates for various proteins of horses and pigs shows that errors of considerable magnitude are introduced by Glx and Asx ambiguities into evolutionary conclusions drawn from sequences of incompletely analyzed proteins.

Holmquist, R.

Methods and apparatus for extraction and tracking of objects from multi-dimensional sequence data

An object tracking technique is provided which, given: (i) a potentially large data set; (ii) a set of dimensions along which the data has been ordered; and (iii) a set of functions for measuring the similarity between data elements, a set of objects are produced. Each of these objects is defined by a list of data elements. Each of the data elements on this list contains the probability that the data element is part of the object. The method produces these lists via an adaptive, knowledge-based search function which directs the search for high-probability data elements. This serves to reduce the number of data element combinations evaluated while preserving the most flexibility in defining the associations of data elements which comprise an object.

Hill, Matthew L.

Methods and apparatus for extraction and tracking of objects from multi-dimensional sequence data

An object tracking technique is provided which, given: (i) a potentially large data set; (ii) a set of dimensions along which the data has been ordered; and (iii) a set of functions for measuring the similarity between data elements, a set of objects are produced. Each of these objects is defined by a list of data elements. Each of the data elements on this list contains the probability that the data element is part of the object. The method produces these lists via an adaptive, knowledge-based search function which directs the search for high-probability data elements. This serves to reduce the number of data element combinations evaluated while preserving the most flexibility in defining the associations of data elements which comprise an object.

Hill, Matthew L.