Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “genomic methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Life in the fast lane for protein crystallization and X-ray crystallography

The common goal for structural genomic centers and consortiums is to decipher as quickly as possible the three-dimensional structures for a multitude of recombinant proteins derived from known genomic sequences. Since X-ray crystallography is the foremost method to acquire atomic resolution for macromolecules, the limiting step is obtaining protein crystals that can be useful of structure determination. High-throughput methods have been developed in recent years to clone, express, purify, crystallize and determine the three-dimensional structure of a protein gene product rapidly using automated devices, commercialized kits and consolidated protocols. However, the average number of protein structures obtained for most structural genomic groups has been very low compared to the total number of proteins purified. As more entire genomic sequences are obtained for different organisms from the three kingdoms of life, only the proteins that can be crystallized and whose structures can be obtained easily are studied. Consequently, an astonishing number of genomic proteins remain unexamined. In the era of high-throughput processes, traditional methods in molecular biology, protein chemistry and crystallization are eclipsed by automation and pipeline practices. The necessity for high-rate production of protein crystals and structures has prevented the usage of more intellectual strategies and creative approaches in experimental executions. Fundamental principles and personal experiences in protein chemistry and crystallization are minimally exploited only to obtain "low-hanging fruit" protein structures. We review the practical aspects of today's high-throughput manipulations and discuss the challenges in fast pace protein crystallization and tools for crystallography. Structural genomic pipelines can be improved with information gained from low-throughput tactics that may help us reach the higher-bearing fruits. Examples of recent developments in this area are reported from the efforts of the Southeast Collaboratory for Structural Genomics (SECSG).

Review↗

Tindallia californiensis sp. nov., a new anaerobic, haloalkaliphilic, spore-forming acetogen isolated from Mono Lake in California

A novel extremely haloalkaliphilic, strictly anaerobic, acetogenic bacterium strain APO was isolated from sediments of the athalassic, meromictic, alkaline Mono Lake in California. The Gram-positive, spore-forming, slightly curved rods with sizes 0.55- 0.7x1.7-3.0 microns were motile by a single laterally attached flagellum. Strain APO was mesophilic (range 10-48 C, optimum of 37 C); halophilic (NaCl range 1-20% (w/v) with optimum of 3-5% (w/v), and alkaliphilic (pH range 8.0-10.5, optimum 9.5). The novel isolate required sodium ions in the medium. Strain APO was an organotroph with a fermentative type of metabolism and used the substrates peptone, bacto-tryptone, casamino acid, yeast extract, L-serine, L-lysine, L-histidine, L-arginine, and pyruvate. The new isolate performed the Stickland reaction with the following amino acid pairs: proline + alanine, glycine + alanine, and tryptophan + valine. The main end product of growth was acetate. High activity of CO dehydrogenase and hydrogenase indicated the presence of a homoacetogenic, non-cycling acetyl-coA pathway. Strain APO was resistant to kanamycin but sensitive to chloramphenicol, tetracycline, and gentamycin. The G+C content of the genomic DNA was 44.4 mol% (by HPLC method). The sequence of the 16s rRNA gene of strain APO possessed 98.2% similarity with the sequence from Tindullia magadiensis Z-7934, but the DNA-DNA hybridization value between these organisms was only 55%. On the basis of these physiological and molecular properties, strain APO is proposed to be a novel species of the genus Tindallia with the name Tindallia californiensis sp. nov., (type strain APO = ATCC BAA-393 - DSM 14871).

Pikuta, E. V.↗

Genomic clocks and evolutionary timescales

For decades, molecular clocks have helped to illuminate the evolutionary timescale of life, but now genomic data pose a challenge for time estimation methods. It is unclear how to integrate data from many genes, each potentially evolving under a different model of substitution and at a different rate. Current methods can be grouped by the way the data are handled (genes considered separately or combined into a 'supergene') and the way gene-specific rate models are applied (global versus local clock). There are advantages and disadvantages to each of these approaches, and the optimal method has not yet emerged. Fortunately, time estimates inferred using many genes or proteins have greater precision and appear to be robust to different approaches.

Evolution, Molecular↗

Materials Genome Initiative Element

NASA is committed to developing new materials and manufacturing methods that can enable new missions with ever increasing mission demands. Typically, the development and certification of new materials and manufacturing methods in the aerospace industry has required more than 20 years of development time with a costly testing and certification program. To reduce the cost and time to mature these emerging technologies, NASA is developing computational materials tools to improve understanding of the material and guide the certification process.

Vickers, John↗

Reveal, A General Reverse Engineering Algorithm for Inference of Genetic Network Architectures

Given the immanent gene expression mapping covering whole genomes during development, health and disease, we seek computational methods to maximize functional inference from such large data sets. Is it possible, in principle, to completely infer a complex regulatory network architecture from input/output patterns of its variables? We investigated this possibility using binary models of genetic networks. Trajectories, or state transition tables of Boolean nets, resemble time series of gene expression. By systematically analyzing the mutual information between input states and output states, one is able to infer the sets of input elements controlling each element or gene in the network. This process is unequivocal and exact for complete state transition tables. We implemented this REVerse Engineering ALgorithm (REVEAL) in a C program, and found the problem to be tractable within the conditions tested so far. For n = 50 (elements) and k = 3 (inputs per element), the analysis of incomplete state transition tables (100 state transition pairs out of a possible 10(exp 15)) reliably produced the original rule and wiring sets. While this study is limited to synchronous Boolean networks, the algorithm is generalizable to include multi-state models, essentially allowing direct application to realistic biological data sets. The ability to adequately solve the inverse problem may enable in-depth analysis of complex dynamic systems in biology and other fields.

Liang, Shoudan↗

CRITICA: coding region identification tool invoking comparative analysis

Gene recognition is essential to understanding existing and future DNA sequence data. CRITICA (Coding Region Identification Tool Invoking Comparative Analysis) is a suite of programs for identifying likely protein-coding sequences in DNA by combining comparative analysis of DNA sequences with more common noncomparative methods. In the comparative component of the analysis, regions of DNA are aligned with related sequences from the DNA databases; if the translation of the aligned sequences has greater amino acid identity than expected for the observed percentage nucleotide identity, this is interpreted as evidence for coding. CRITICA also incorporates noncomparative information derived from the relative frequencies of hexanucleotides in coding frames versus other contexts (i.e., dicodon bias). The dicodon usage information is derived by iterative analysis of the data, such that CRITICA is not dependent on the existence or accuracy of coding sequence annotations in the databases. This independence makes the method particularly well suited for the analysis of novel genomes. CRITICA was tested by analyzing the available Salmonella typhimurium DNA sequences. Its predictions were compared with the DNA sequence annotations and with the predictions of GenMark. CRITICA proved to be more accurate than GenMark, and moreover, many of its predictions that would seem to be errors instead reflect problems in the sequence databases. The source code of CRITICA is freely available by anonymous FTP (rdp.life.uiuc.edu in/pub/critica) and on the World Wide Web (http:/(/)rdpwww.life.uiuc.edu).

Non-NASA Center↗

Ribosomal RNA: a key to phylogeny

As molecular phylogeny increasingly shapes our understanding of organismal relationships, no molecule has been applied to more questions than have ribosomal RNAs. We review this role of the rRNAs and some of the insights that have been gained from them. We also offer some of the practical considerations in extracting the phylogenetic information from the sequences. Finally, we stress the importance of comparing results from multiple molecules, both as a method for testing the overall reliability of the organismal phylogeny and as a method for more broadly exploring the history of the genome.

NASA Program Exobiology↗

Methods for determining the genetic affinity of microorganisms and viruses

Selecting which sub-sequences in a database of nucleic acid such as 16S rRNA are highly characteristic of particular groupings of bacteria, microorganisms, fungi, etc. on a substantially phylogenetic tree. Also applicable to viruses comprising viral genomic RNA or DNA. A catalogue of highly characteristic sequences identified by this method is assembled to establish the genetic identity of an unknown organism. The characteristic sequences are used to design nucleic acid hybridization probes that include the characteristic sequence or its complement, or are derived from one or more characteristic sequences. A plurality of these characteristic sequences is used in hybridization to determine the phylogenetic tree position of the organism(s) in a sample. Those target organisms represented in the original sequence database and sufficient characteristic sequences can identify to the species or subspecies level. Oligonucleotide arrays of many probes are especially preferred. A hybridization signal can comprise fluorescence, chemiluminescence, or isotopic labeling, etc.; or sequences in a sample can be detected by direct means, e.g. mass spectrometry. The method's characteristic sequences can also be used to design specific PCR primers. The method uniquely identifies the phylogenetic affinity of an unknown organism without requiring prior knowledge of what is present in the sample. Even if the organism has not been previously encountered, the method still provides useful information about which phylogenetic tree bifurcation nodes encompass the organism.

Fox, George E.↗

Genome-wide transcriptional analysis of flagellar regeneration in Chlamydomonas reinhardtii identifies orthologs of ciliary disease genes

The important role that cilia and flagella play in human disease creates an urgent need to identify genes involved in ciliary assembly and function. The strong and specific induction of flagellar-coding genes during flagellar regeneration in Chlamydomonas reinhardtii suggests that transcriptional profiling of such cells would reveal new flagella-related genes. We have conducted a genome-wide analysis of RNA transcript levels during flagellar regeneration in Chlamydomonas by using maskless photolithography method-produced DNA oligonucleotide microarrays with unique probe sequences for all exons of the 19,803 predicted genes. This analysis represents previously uncharacterized whole-genome transcriptional activity profiling study in this important model organism. Analysis of strongly induced genes reveals a large set of known flagellar components and also identifies a number of important disease-related proteins as being involved with cilia and flagella, including the zebrafish polycystic kidney genes Qilin, Reptin, and Pontin, as well as the testis-expressed tubby-like protein TULP2.

Polycystic Kidney Diseases/genetics↗

Decoding the effects of synonymous variants

Synonymous single nucleotide variants (sSNVs) are common in the human genome but are often overlooked. However, sSNVs can have significant biological impact and may lead to disease. Existing computational methods for evaluating the effect of sSNVs suffer from the lack of gold-standard training/evaluation data and exhibit over-reliance on sequence conservation signals. We developed synVep (synonymous Variant effect predictor), a machine learning-based method that overcomes both of these limitations. Our training data was a combination of variants reported by gnomAD (observed) and those unreported, but possible in the human genome (generated). We used positive-unlabeled learning to purify the generated variant set of any likely unobservable variants. We then trained two sequential extreme gradient boosting models to identify subsets of the remaining variants putatively enriched and depleted in effect. Our method attained 90% precision/recall on a previously unseen set of variants. Furthermore, although synVep does not explicitly use conservation, its scores correlated with evolutionary distances between orthologs in cross-species variation analysis. synVep was also able to differentiate pathogenic vs. benign variants, as well as splice-site disrupting variants (SDV) vs. non-SDVs. Thus, synVep provides an important improvement in annotation of sSNVs, allowing users to focus on variants that most likely harbor effects.

Zishuo Zeng↗

Nonlinear Homogenization of Finitely Deformed Viscoelastic-Viscoplastic Composites Using Mechanics of Structure Genome

The objective of this paper is to develop a micromechanics approach to homogenizing finitely deformed viscoelastic-viscoplastic composites using the mechanics of structure genome. The incremental constitutive relation for glassy polymers, formulated in the spatial configuration, is implemented in the present approach.This involves (1) pulling-back the constitutive model to the material configuration and (2)choosing the deformation gradient tensor and the first Piola–Kirchhoff stress tensor as the strain and the stress measures during homogenization, respectively. An Euler–Newton predictor–corrector method is developed for homogenization. Each step involves formulating a variational statement using the mechanics of structure genome, discretizing the statement in a finite-dimensional space, and solving the problem using an Euler/multilevel Newton method. The present approach is demonstrated by homogenizing fiber- and particle-reinforced composites undergoing uniaxial, biaxial, or shear deformation, at different stain rates.

Multi-scale modeling, High Strain Composites, Visc↗

Simultaneous measurement of multiple radiation-induced protein expression profiles using the Luminex(TM) system

Space flight results in the exposure of astronauts to a mixed field of radiation composed of energetic particles of varying energies, and biological indicators of space radiation exposure provides a better understanding of the associated long-term health risks. Current methods of biodosimetry have employed the use of cytogenetic analysis for biodosimetry, and more recently the advent of technological progression has led to advanced research in the use of genomic and proteomic expression profiling to simultaneously assess biomarkers of radiation exposure. We describe here the technical advantages of the Luminex(TM) 100 system relative to traditional methods and its potential as a tool to simultaneously profile multiple proteins induced by ionizing radiation. The development of such a bioassay would provide more relevant post-translational dynamics of stress response and will impart important implications in the advancement of space and other radiation contact monitoring. c2004 COSPAR. Published by Elsevier Ltd. All rights reserved.

NASA Center JSC↗

An Open-Science Approach to Address Individual Response to Simulated GCR In Genetically Diverse Populations of Mice and Humans

This project addresses the challenge of understanding and predicting individual radiation sensitivity by integrating genetics, demographics and biomarker characteristics across species (mice and humans). We hypothesize that ex vivo DNA repair response to GCR components is a central determinant of cancer risk from space radiation and can serve as a biomarker of radiation risk in combination with genetics. Automated image quantification of 53BP1+ radiation-induced foci (RIF) during the first 4-48 h post-irradiation was performed as a function of dose and LET in non-immortalized primary skin fibroblasts derived from 76 mice across 15 strains (5 inbred reference strains and 10 collaborative-cross strains) exposed to X rays (0.1, 1 and 4 Gy), 350 MeV/n 40Ar and 600 MeV/n 56Fe (1.1 and 3 particles/100sq. μm), as well as in peripheral blood mononuclear cells (PBMCs) from 768 healthy donors (matched ethnicity, 50/50 male/female, 18-70 years old) exposed to gamma rays (0.1 and 1 Gy), 350 MeV/n 28Si, 350 MeV/n 40Ar and 600 MeV/n 56Fe (1.1 and 3 particles/100sq. μm). A genome-wide association study (GWAS) was performed on the mouse strains between DNA damage responses to space radiation and single nucleotide polymorphisms (SNPs). We found SNPs, which were significantly associated to the RIF phenotype, mapped to genes and pathways that are functionally linked to health hazards for deep space exploration (e.g. carcinogenesis, nervous system damage and immune dysfunction). Some of these SNPs were located within protein coding regions, potentially interfering with protein functions and providing promising genetic targets for countermeasures. We also found correlations between both spontaneous and radiation-induced DNA damage and SNPs mapped to pathways associated with cellular metabolism. GWAS is undergoing for the human data. All data have been made available via the NASA Space Biology Open-Science database (genelab.nasa.gov) and we will discuss how various genomic and transcriptomic datasets can be accessed for modeling and integrated using machine learning methods for discovering new radiation biology.

Sylvain V Costes↗

Developing a Genetic Variant Calling Pipeline for Quantifying the Complex Mutagenic Load Accumulated in BioNutrients-1 Production Pack Samples

Microorganisms hold great promise for on demand production of labile nutrients and pharmaceuticals as well recycling and in situ resource utilization. The utilization of microorganisms for such tasks on space missions is hindered by the limited data on how microbes respond to spaceflight. For example, the genetic stability of microorganisms, and the genomic engineered traits added to deliver desired functions, over long-term storage in the spacecraft environment is poorly understood. The BioNutrients-1 (BN-1) mission conducted a 5-year study of desiccated storage in Low Earth Orbit (LEO) to evaluate the suitability of eight synthetic biology chassis organisms for long-duration space missions. We are employing high-depth, whole genome sequencing (WGS) to determine the mutagenic load that accumulated during long-term storage. Mutation analysis pipelines are well established for homogenous culture grown from a single colony, but the mutational landscape of the BN-1 samples present a unique analysis challenge, as every cell in the BN-1 samples had a unique genetic journey of DNA damage and repair. Consequently, sequence variants are expected at low allele frequency within samples. To address this genetic complexity, we apply two distinct computational approaches to identify mutations in pre-existing WGS data collected from populations of Chlamydomonas reinhardtii that were exposed to UV mutagenesis and growth in LEO. For reference genome free mutation detection, we utilized DiscoSNP++, which is a de Bruijn graph approach. For reference genome-based mutation detection we utilize GATK for Microbes, which is a Bayesian probabilistic approach. We will benchmark these approaches against the mutations originally identified using CRISP, a method optimized for pooled samples. Ultimately, quantifying the mutation load imposed by storage or growth on the ISS will help identify chassis organisms with both high levels of genome stability and viability, which are desirable traits for implementation of bioproduction in long-duration missions.

SNP↗

Fluorescent Approaches to High Throughput Crystallography

X-ray crystallography remains the primary method for determining the structure of macromolecules. The first requirement is to have crystals, and obtaining them is often the rate-limiting step. The numbers of crystallization trials that are set up for any one protein for structural genomics, and the rate at which they are being set up, now overwhelm the ability for strictly human analysis of the results. Automated analysis methods are now being implemented with varying degrees of success, but these typically cannot reliably extract intermediate results. By covalently modifying a subpopulation, less than or = 1%, of a macromolecule solution with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of a macromolecules purification. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination the crystals will show up as bright objects against a dark background. As crystalline packing is more dense than amorphous precipitate, the fluorescence intensity can be used as a guide in distinguishing different types of precipitated phases, even in the absence of obvious crystalline features, widening the available potential lead conditions in the absence of clear "bits." Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Also, brightly fluorescent crystals are readily found against less fluorescent precipitated phases, which under white light illumination may serve to obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries and by having the protein or protein structures all that show up. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons. This presentation will focus on the methodology for fluorescent labeling, the crystallization results, and the effects of the trace labeling on the crystal quality.

Minamitani, Elizabeth Forsythe↗

Fluorescent Approaches to High Throughput Crystallography

X-ray crystallography remains the primary method for determining the structure of macromolecules. The first requirement is to have crystals, and obtaining them is often the rate-limiting step. The numbers of crystallization trials that are set up for any one protein for structural genomics, and the rate at which they are being set up, now overwhelm the ability for strictly human analysis of the results. Automated analysis methods are now being implemented with varying degrees of success, but these typically can not reliably extract intermediate results. By covalently modifying a subpopulation, less than or = 1%, of a macromolecule solution with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination the crystals show up as bright objects against a dark background. As crystalline packing is more dense than amorphous precipitate, the fluorescence intensity can be used as a guide in distinguishing different types of precipitated phases, even in the absence of obvious crystalline features, widening the available potential lead conditions in the absence of clear "hits." Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Also, brightly fluorescent crystals are readily found against less fluorescent precipitated phases, which under white light illumination may serve to obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries and by having the protein or protein structures all that show up. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons. This presentation will focus on the methodology for fluorescent labeling, the crystallization results, and the effects of the trace labeling on the crystal quality.

Pusey, Marc L.↗

Fluorescent Approaches to High Throughput Crystallography

X-ray crystallography remains the primary method for determining the structure of macromolecules. The first requirement is to have crystals, and obtaining them is often the rate-limiting step. The numbers of crystallization trials that are set up for any one protein for structural genomics, and the rate at which they are being set up, now overwhelm the ability for strictly human analysis of the results. Automated analysis methods are now being implemented with varying degrees of success, but these typically cannot reliably extract intermediate results. By covalently modifying a subpopulation, 51%, of a macromolecule solution with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination the crystals show up as bright objects against a dark background. As crystalline packing is more dense than amorphous precipitate, the fluorescence intensity can be used as a guide in distinguishing different types of precipitated phases, even in the absence of obvious crystalline features, widening the available potential lead conditions in the absence of clear hits. Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Also, brightly fluorescent crystals are readily found against less fluorescent precipitated phases, which under white light illumination may serve to obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries and by having the protein or protein structures all that show up. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons. This presentation will focus on the methodology for fluorescent labeling, the crystallization results, and the effects of the trace labeling on the crystal quality.

Pusey, Marc L.↗

Fluorescent Approaches to High Throughput Crystallography

X-ray crystallography remains the primary method for determining the structure of macromolecules. The first requirement is to have crystals, and obtaining them is often the rate-limiting step. The numbers of crystallization trials that are set up for any one protein for structural genomics, and the rate at which they are being set up, now overwhelm the ability for strictly human analysis of the results. Automated analysis methods are now being implemented with varying degrees of success, but these typically cannot reliably extract intermediate results. By covalently modifying a subpopulation, less than or = 1 %, of a macromolecule solution with a fluorescent probe, the labeled material will add to a growing crystal as a microheterogeneous growth unit. Labeling procedures can be readily incorporated into the final stages of purification. The covalently attached probe will concentrate in the crystal relative to the solution, and under fluorescent illumination the crystals show up as bright objects against a dark background. As crystalline packing is more dense than amorphous precipitate, the fluorescence intensity can be used as a guide in distinguishing different types of precipitated phases, even in the absence of obvious crystalline features, widening the available potential lead conditions in the absence of clear "hits." Non-protein structures, such as salt crystals, will not incorporate the probe and will not show up under fluorescent illumination. Also, brightly fluorescent crystals are readily found against less fluorescent precipitated phases, which under white light illumination may serve to obscure the crystals. Automated image analysis to find crystals should be greatly facilitated, without having to first define crystallization drop boundaries and by having the protein or protein structures all that show up. The trace fluorescently labeled crystals will also emit with sufficient intensity to aid in the automation of crystal alignment using relatively low cost optics, further increasing throughput at synchrotrons. Preliminary experiments show that the presence of the fluorescent probe does not affect the nucleation process or the quality of the X-ray data obtained.

Pusey, Marc L.↗