Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Human Genome Project”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Modeling the Activity of Single Genes

The central dogma of molecular biology states that information is stored in DNA, transcribed to messenger RNA (mRNA) and then translated into proteins. The human genome project is expected to determine the complete sequence of all human genes, and the genomes of several other organisms are already completely sequenced.

gene↗

Modeling the Activity of Single Genes

The central dogma of molecular biology states that information is stored in DNA, transcribed to messenger RNA (mRNA) and then translated into proteins. This picture is significantly augmentated when we consider the action of certain proteins in regulating transcription. These transcription factors provide a feedback pathway by which genes can regulate one another's expression as mRNA and then as protein. To review: DNA, RNA and proteins have different functions. DNA is the molecular storehouse of genetic information. When cells divide, the DNA is replicated, so that each daughter cell maintains the same genetic information as the mother cell. RNA acts as a go-between from DNA to proteins. Only a single copy of DNA is present, but multiple copies of the same piece of RNA may be present, allowing cells to make huge amounts of protein. In eukaryotes (organisms with a nucleus), DNA is found in the nucleus only. RNA is copied in the nucleus then translocates(moves) outside the nucleus, where it is transcribed into proteins. Along the way, the RNA may be spliced, i.e., may have pieces cut out. RNA then attaches to ribosomes and is translated to proteins. Proteins are the machinery of the cell other than DNA and RNA, all the complex molecules of the cell are proteins. Proteins are specialized machines, each of which fulfills its own task, which may be transporting oxygen, catalyzing reactions, or responding to extracellular signals, just to name a few. One of the more interesting functions a protein may have is binding directly or indirectly to DNA to perform transcriptional regulation, thus forming a closed feedback loop of gene regulation. The structure of DNA and the central dogma were understood in the 50s; in the early 80s it became possible to make arbitrary modifications to DNA and use cellular machinery to transcribe and translate the resulting genes; more recently, genomes (i.e., the complete DNA sequence) of many organisms have been sequenced. This large-scale sequencing began with simple organisms, viruses and bacteria, progressed to eukaryotes such as yeast, and more recently (1998) progressed to a multi-cellular animal, the nematode Caenorhabditis elegans. Sequencers have now moved on to the fruit fly Drosophila melanogaster, whose sequence is slated for completion by the end of 1999. The human genome project is expected to determine the complete sequence of all 3 billion bases of human DNA within the next five years. In the wake of genome-scale sequencing, further instrumentation is being developed to assay gene expression and function on a comparably large scale. Much of the work in computational biology focuses on computational tools used in sequencing, finding genes that are related to a particular gene, finding which parts of the DNA code for proteins and which do not, understanding what proteins will be formed from a given length of DNA, predicting how the proteins will fold from a one-dimensional structure into a three dimensional structure, and so on. Much less computational work has been done regarding the function of proteins. One reason for this is that different proteins function very differently, and so work on protein function is very specific to certain classes of proteins. There are, for example, proteins such enzymes that catalyze various intracellular reactions, receptors that respond to extracellular signals and ion channels that regulate the flow of charged particles into and out of the cell. In this chapter, we will consider a particular class of proteins called transcription factors(TFs), which are responsible for regulating when a certain gene is expressed in a certain cell, which cells it is express in, and how much is expressed. Understanding these processes will involve developing a deeper understanding of transcription, translation, and the cellular processes that control those processes. All of these elements fall under the aegis of gene regulation or more narrowly transcriptional regulation. Some of the key questions in gene regulation are: What genes are expressed in a certain cell at a certain time? How does gene expression differ from cell to cell in a multicellular organism? Which proteins act as transcription factors, i.e., are important in regulating gene expression? From questions like these, we hope to understand which genes are important for various macroscopic processes. Nearly all of the cells of a multicellular organism contain the same DNA. Yet this same genetic information yields a large number of different cell types. The fundamental difference between a neuron and a liver cell, for example, is which genes are expressed. Thus understanding gene regulation is an important step in understanding development. Furthermore, understanding the usual genes that are expressed in cells may give important clues about various diseases. Some diseases, such as sickle cell anemia and cystic fibrosis, are caused by defects in single, non-regulatory genes; others, such as certain cancers, are caused when the cellular control circuitry malfunctions - an understanding of these diseases will involve pathways of multiple interacting gene products. There are numerous challenges in the area of understanding and modeling gene regulation. First and foremost, biologists would like to develop a deeper understanding of the processes involved, including which genes and families of genes are important, how they interact, etc. From a computation point of view, there has been embarrassingly little work done. In this chapter there are many areas in which we can phrase meaningful, non-trivial computational questions, but questions that have not been addressed. Some of these are purely computational (what is a good algorithm for dealing with a model of type X) and others are more mathematical (given a system with certain characteristics, what sort of model can one use? How does one find biochemical parameters from system-level behavior using as few experiments as possible?). In addition to biological and algorithmic problems, there is also the ever-present issue of theoretical biology - what general principles can be derived from these systems, what can one do with models other than just simulate time-courses, what can be deduced about a class of systems without knowing all the details? The fundamental challenge to computationalists and theorists is to add value to the biology - to use models, modeling techniques and algorithms to understand the biology in new ways.

Mjolsness, Eric↗

Mammalian Chromosome Analysis and Sorting by Flow Cytometry

The analysis of chromosomes by flow cytometry is termed flow cytogenetics, and it involves the analysis and sorting of single mitotic chromosomes in suspension. The study of flow karyograms provides insight into chromosome number and structure to provide information on chromosomal DNA content and can enable the detection of deletions, translocations, or any forms of aneuploidy. Beyond its clinical applications, flow cytogenetics greatly contributed to the Human Genome Project through the ability to sort pure populations of chromosomes for gene mapping, cloning, and the construction of DNA libraries. Maximizing the potential of these important applications of flow cytogenetics relies on precise instrument setup and optimal sample processing, both of which impact the accuracy and quality of the data that are generated. This article is a compilation of the existing protocols that describe the stepwise methodology of accumulating, isolating, and staining metaphase chromosomes to prepare single-chromosome suspensions for flow cytometric analysis and sorting. Although the chromosome preparation protocols have remained largely unchanged, cytometer technology has advanced dramatically since these protocols were originally developed. Advances in cytometry technologies offer new and exciting approaches for understanding and monitoring chromosomal aberrations, but the hallmark of these protocols remains their simplicity in methodologies and reagent requirements and the accuracy of data resolvable to every chromosome of the cell.

59 BASIC BIOLOGICAL SCIENCES↗

The Single Nucleotide Polymorphism Consortium

I want to discuss both the Single Nucleotide Polymorphism (SNP) Consortium and the Human Genome Project. I am afraid most of my presentation will be thin on law and possibly too high on rhetoric. Having been engaged in a personal and direct way with these issues as a trained scientist, I find it quite difficult to be always as objective as I ought to be.

Morgan, Michael↗

TPSAS-NF1676L-17800-DND

Currently, there are two national challenge problems that guide much of the research in durability and damage tolerance at NASA Langley. The first, Airframe Digital Twin, is a concept that combines as-built vehicle components, as-experienced loads and environments, and other vehicle-specific characteristics to enable ultrahigh fidelity modeling of aircraft and spacecraft throughout their service lives. The second, Materials Genome Initiative, is an analog to the Human Genome Project, and is intended to improve the rate at which materials scientists can discover, understand fundamental physics, and improve material systems. This presentation will highlight several research projects ongoing at NASA Langley that are in support of the above challenge problems. Two of those topics will be the subject of detailed discussion. First, investigations of microstructurally-small fatigue cracking (MSFC) in Al-2Cu and Al-4Cu, fabricated in-house, will be presented. Single- and oligo-crystals of Al-Cu specimens were loaded in uniaxial fatigue, while high-resolution in-situ measurements of deformation were made using image correlation (IC) in a scanning-electron microscope (SEM). The Al-Cu specimens were then replicated as crystal plasticity finite element models (CPFEM), where evolution of slip localization near grain boundaries was computed. Comparison among experiment and CPFEM is made. In addition, XRay diffraction measurements of the as-fabricated specimens were carried out, where direct measurements of the embedded copper precipitates were made, and their influence on growing MSFCs were directly observed. The second main topic will illustrate ongoing work in the area of so-called damage-sensing particles. In this work, shape-memory alloys are embedded in an aluminum alloy matrix. Upon the propagation of a fatigue crack, these particles undergo a strain-induced phase transformation which is detected using an acoustic sensor, providing real-time information on propagating cracks. Experiments and simulations regarding the development of this system will also be detailed.

Jacob Hochhalter↗

Human RNome Project draft human RNome sequence of GM12878, B-cell line, obtained by mass-spectrometry sequencing, long-read sequencing and short-read sequencing.

Here we report the first draft of the human RNome sequence, a reference map of RNA chemical modifications in a human B-cell line. RNA carries a diverse repertoire of chemical modifications that regulate gene expression, cellular function, and responses to physiological and pathological cues. Yet, unlike the genome, no reference map of RNA modifications is available for any human cell. To generate this resource, the Human RNome Project Consortium analyzed a shared RNA preparation from the well-characterized GM12878 B-cell line using short-read sequencing, long-read direct RNA sequencing, and mass spectrometry, generating more than 7.1 billion sequencing reads spanning approximately 1.2 trillion nucleotides. The resulting maps of the human RNome reveal that RNA modifications are organized according to function, transcript architecture, and cellular identity. Modifications concentrate at functional centers of ribosomal and transfer RNAs, follow the canonical topology of N6-methyladenosine in coding transcripts, and form coordinated hotspots in immune regulatory genes. This first reference human RNome provides a foundation for understanding how RNA chemistry shapes cellular identity, human disease, and the development of RNA-based therapeutics.

59 BASIC BIOLOGICAL SCIENCES↗

Exabiome: Advancing Microbial Science through Exascale Computing

The Exabiome project seeks to improve the understanding of microbiomes through the development of methods for accelerating metagenomic science using exascale computing. This article gives an overview of scientific impact of the three components of the project: metagenome assembly, protein family detection, and comparative analysis of metagenomes. Exabiome developed MetaHipMer, the only metagenome assembler capable of scaling to full exascale systems. MetaHipMer has enabled ground-breaking assemblies on the Frontier supercomputer, with many scientific benefits, such as the discovery of rare species and viral genomes. To investigate protein families, Exabiome developed two exascale tools, PASTIS and HipMCL. Together, these can utilize exascale resources to understand the functional diversity of billions of dark matter proteins and novel protein families. For comparative analysis, Exabiome developed kmerprof, a tool that can be used to compare huge metagenomes for many different scientific purposes, for example, grouping human microbiomes according to body location.

59 BASIC BIOLOGICAL SCIENCES↗

Author Correction: Expanded encyclopaedias of DNA elements in the human and mouse genomes

In the version of this article initially published, two members of the ENCODE Project Consortium were missing from the author list. Rizi Ai (Department of Chemistry and Biochemistry, University of California, San Diego, La Jolla, CA, USA) and Shantao Li (Program in Computational Biology and Bioinformatics, Yale University, New Haven, CT, USA) are now included in the author list. These errors have been corrected in the online version of the article.

59 BASIC BIOLOGICAL SCIENCES↗

The Analysis of the Patterns of Radiation-Induced DNA Damage Foci by a Stochastic Monte Carlo Model of DNA Double Strand Breaks Induction by Heavy Ions and Image Segmentation Software

To create a generalized mechanistic model of DNA damage in human cells that will generate analytical and image data corresponding to experimentally observed DNA damage foci and will help to improve the experimental foci yields by simulating spatial foci patterns and resolving problems with quantitative image analysis. Material and Methods: The analysis of patterns of RIFs (radiation-induced foci) produced by low- and high-LET (linear energy transfer) radiation was conducted by using a Monte Carlo model that combines the heavy ion track structure with characteristics of the human genome on the level of chromosomes. The foci patterns were also simulated in the maximum projection plane for flat nuclei. Some data analysis was done with the help of image segmentation software that identifies individual classes of RIFs and colocolized RIFs, which is of importance to some experimental assays that assign DNA damage a dual phosphorescent signal. Results: The model predicts the spatial and genomic distributions of DNA DSBs (double strand breaks) and associated RIFs in a human cell nucleus for a particular dose of either low- or high-LET radiation. We used the model to do analyses for different irradiation scenarios. In the beam-parallel-to-the-disk-of-a-flattened-nucleus scenario we found that the foci appeared to be merged due to their high density, while, in the perpendicular-beam scenario, the foci appeared as one bright spot per hit. The statistics and spatial distribution of regions of densely arranged foci, termed DNA foci chains, were predicted numerically using this model. Another analysis was done to evaluate the number of ion hits per nucleus, which were visible from streaks of closely located foci. In another analysis, our image segmentaiton software determined foci yields directly from images with single-class or colocolized foci. Conclusions: We showed that DSB clustering needs to be taken into account to determine the true DNA damage foci yield, which helps to determine the DSB yield. Using the model analysis, a researcher can refine the DSB yield per nucleus per particle. We showed that purely geometric artifacts, present in the experimental images, can be analytically resolved with the model, and that the quantization of track hits and DSB yields can be provided to the experimentalists who use enumeration of radiation-induced foci in immunofluorescence experiments using proteins that detect DNA damage. An automated image segmentaiton software can prove useful in a faster and more precise object counting for colocolized foci images.

Ponomarev, Artem↗

Overview of NASARTI (NASA Radiation Track Image) Program: Highlights of the Model Improvement and the New Results

This presentation summarizes several years of research done by the co-authors developing the NASARTI (NASA Radiation Track Image) program and supporting it with scientific data. The goal of the program is to support NASA mission to achieve a safe space travel for humans despite the perils of space radiation. The program focuses on selected topics in radiation biology that were deemed important throughout this period of time, both for the NASA human space flight program and to academic radiation research. Besides scientific support to develop strategies protecting humans against an exposure to deep space radiation during space missions, and understanding health effects from space radiation on astronauts, other important ramifications of the ionizing radiation were studied with the applicability to greater human needs: understanding the origins of cancer, the impact on human genome, and the application of computer technology to biological research addressing the health of general population. The models under NASARTI project include: the general properties of ionizing radiation, such as particular track structure, the effects of radiation on human DNA, visualization and the statistical properties of DSBs (DNA double-strand breaks), DNA damage and repair pathways models and cell phenotypes, chromosomal aberrations, microscopy data analysis and the application to human tissue damage and cancer models. The development of the GUI and the interactive website, as deliverables to NASA operations teams and tools for a broader research community, is discussed. Most recent findings in the area of chromosomal aberrations and the application of the stochastic track structure are also presented.

Ponomarev, Artem L.↗

Spatial top-down proteomics for the functional characterization of human kidney

Background: The Human Proteome Project has credibly detected nearly 93% of the roughly 20,000 proteins which are predicted by the human genome. However, the proteome is enigmatic, where alterations in amino acid sequences from polymorphisms and alternative splicing, errors in translation, and post-translational modifications result in a proteome depth estimated at several million unique proteoforms. Recently mass spectrometry has been demonstrated in several landmark efforts mapping the human proteoform landscape in bulk analyses. Herein, we developed an integrated workflow for characterizing proteoforms from human tissue in a spatially resolved manner by coupling laser capture microdissection, nanoliter-scale sample preparation, and mass spectrometry imaging. Results: Using healthy human kidney sections as the case study, we focused our analyses on the major functional tissue units including glomeruli, tubules, and medullary rays. After laser capture microdissection, these isolated functional tissue units were processed with microPOTS (microdroplet processing in one-pot for trace samples) for sensitive top-down proteomics measurement. This provided a quantitative database of 616 proteoforms that was further leveraged as a library for mass spectrometry imaging with near-cellular spatial resolution over the entire section. Notably, several mitochondrial proteoforms were found to be differentially abundant between glomeruli and convoluted tubules, and further spatial contextualization was provided by mass spectrometry imaging confirming unique differences identified by microPOTS, and further expanding the field-of-view for unique distributions such as enhanced abundance of a truncated form (1-74) of ubiquitin within cortical regions. Conclusions: We developed an integrated workflow to directly identify proteoforms and reveal their spatial distributions. Where of the 20 differentially abundant proteoforms identified as discriminate between tubules and glomeruli by microPOTS, the vast majority of tubular proteoforms were of mitochondrial origin (8 of 10) where discriminate proteoforms in glomeruli were primarily hemoglobin subunits (9 of 10). These trends were also identified within ion images demonstrating spatially resolved characterization of proteoforms that has the potential to reshape discovery-based proteomics because the proteoforms are the ultimate effector of cellular functions. Applications of this technology have the potential to unravel etiology and pathophysiology of disease states, informing on biologically active proteoforms, which remodel the proteomic landscape in chronic and acute disorders.

59 BASIC BIOLOGICAL SCIENCES↗

LinkFinder: An expert system that constructs phylogenic trees

An expert system has been developed using the C Language Integrated Production System (CLIPS) that automates the process of constructing DNA sequence based phylogenies (trees or lineages) that indicate evolutionary relationships. LinkFinder takes as input homologous DNA sequences from distinct individual organisms. It measures variations between the sequences, selects appropriate proportionality constants, and estimates the time that has passed since each pair of organisms diverged from a common ancestor. It then designs and outputs a phylogenic map summarizing these results. LinkFinder can find genetic relationships between different species, and between individuals of the same species, including humans. It was designed to take advantage of the vast amount of sequence data being produced by the Genome Project, and should be of value to evolution theorists who wish to utilize this data, but who have no formal training in molecular genetics. Evolutionary theory holds that distinct organisms carrying a common gene inherited that gene from a common ancestor. Homologous genes vary from individual to individual and species to species, and the amount of variation is now believed to be directly proportional to the time that has passed since divergence from a common ancestor. The proportionality constant must be determined experimentally; it varies considerably with the types of organisms and DNA molecules under study. Given an appropriate constant, and the variation between two DNA sequences, a simple linear equation gives the divergence time.

Inglehart, James↗

Tripal, a community update after 10 years of supporting open source, standards-based genetic, genomic and breeding databases

Abstract Online, open access databases for biological knowledge serve as central repositories for research communities to store, find and analyze integrated, multi-disciplinary datasets. With increasing volumes, complexity and the need to integrate genomic, transcriptomic, metabolomic, proteomic, phenomic and environmental data, community databases face tremendous challenges in ongoing maintenance, expansion and upgrades. A common infrastructure framework using community standards shared by many databases can reduce development burden, provide interoperability, ensure use of common standards and support long-term sustainability. Tripal is a mature, open source platform built to meet this need. With ongoing improvement since its first release in 2009, Tripal provides full functionality for searching, browsing, loading and curating numerous types of data and is a primary technology powering at least 31 publicly available databases spanning plants, animals and human data, primarily storing genomics, genetics and breeding data. Tripal software development is managed by a shared, inclusive governance structure including both project management and advisory teams. Here, we report on the most important and innovative aspects of Tripal after 11 years development, including integration of diverse types of biological data, successful collaborative projects across member databases, and support for implementing FAIR principles.

59 BASIC BIOLOGICAL SCIENCES↗

The Human Proteoform Project: Defining the human proteome

Proteins are the primary effectors of function in biology, and thus, complete knowledge of their structure and properties is fundamental to deciphering function in basic and translational research. The chemical diversity of proteins is expressed in their many proteoforms, which result from combinations of genetic polymorphisms, RNA splice variants, and posttranslational modifications. This knowledge is foundational for the biological complexes and networks that control biology yet remains largely unknown. We propose here an ambitious initiative to define the human proteome, that is, to generate a definitive reference set of the proteoforms produced from the genome. Several examples of the power and importance of proteoform-level knowledge in disease-based research are presented along with a call for improved technologies in a two-pronged strategy to the Human Proteoform Project.

59 BASIC BIOLOGICAL SCIENCES↗

NASA's GeneLab Phase II: Federated Search and Data Discovery

GeneLab is currently being developed by NASA to accelerate 'open science' biomedical research in support of the human exploration of space and the improvement of life on earth. Phase I of the four-phase GeneLab Data Systems (GLDS) project emphasized capabilities for submission, curation, search, and retrieval of genomics, transcriptomics and proteomics ('omics') data from biomedical research of space environments. The focus of development of the GLDS for Phase II has been federated data search for and retrieval of these kinds of data across other open-access systems, so that users are able to conduct biological meta-investigations using data from a variety of sources. Such meta-investigations are key to corroborating findings from many kinds of assays and translating them into systems biology knowledge and, eventually, therapeutics.

exobiology↗

NASAs GeneLab Phase II: Federated Search and Data Discovery

GeneLab is currently being developed by NASA to accelerate open science biomedical research in support of the human exploration of space and the improvement of life on earth. Phase I of the four-phase GeneLab Data Systems (GLDS) project emphasized capabilities for submission, curation, search, and retrieval of genomics, transcriptomics and proteomics (omics) data from biomedical research of space environments. The focus of development of the GLDS for Phase II has been federated data search for and retrieval of these kinds of data across other open-access systems, so that users are able to conduct biological meta-investigations using data from a variety of sources. Such meta-investigations are key to corroborating findings from many kinds of assays and translating them into systems biology knowledge and, eventually, therapeutics.

genome↗