Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Human Genome Project”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Mammalian Chromosome Analysis and Sorting by Flow Cytometry

The analysis of chromosomes by flow cytometry is termed flow cytogenetics, and it involves the analysis and sorting of single mitotic chromosomes in suspension. The study of flow karyograms provides insight into chromosome number and structure to provide information on chromosomal DNA content and can enable the detection of deletions, translocations, or any forms of aneuploidy. Beyond its clinical applications, flow cytogenetics greatly contributed to the Human Genome Project through the ability to sort pure populations of chromosomes for gene mapping, cloning, and the construction of DNA libraries. Maximizing the potential of these important applications of flow cytogenetics relies on precise instrument setup and optimal sample processing, both of which impact the accuracy and quality of the data that are generated. This article is a compilation of the existing protocols that describe the stepwise methodology of accumulating, isolating, and staining metaphase chromosomes to prepare single-chromosome suspensions for flow cytometric analysis and sorting. Although the chromosome preparation protocols have remained largely unchanged, cytometer technology has advanced dramatically since these protocols were originally developed. Advances in cytometry technologies offer new and exciting approaches for understanding and monitoring chromosomal aberrations, but the hallmark of these protocols remains their simplicity in methodologies and reagent requirements and the accuracy of data resolvable to every chromosome of the cell.

59 BASIC BIOLOGICAL SCIENCES↗

Human RNome Project draft human RNome sequence of GM12878, B-cell line, obtained by mass-spectrometry sequencing, long-read sequencing and short-read sequencing.

Here we report the first draft of the human RNome sequence, a reference map of RNA chemical modifications in a human B-cell line. RNA carries a diverse repertoire of chemical modifications that regulate gene expression, cellular function, and responses to physiological and pathological cues. Yet, unlike the genome, no reference map of RNA modifications is available for any human cell. To generate this resource, the Human RNome Project Consortium analyzed a shared RNA preparation from the well-characterized GM12878 B-cell line using short-read sequencing, long-read direct RNA sequencing, and mass spectrometry, generating more than 7.1 billion sequencing reads spanning approximately 1.2 trillion nucleotides. The resulting maps of the human RNome reveal that RNA modifications are organized according to function, transcript architecture, and cellular identity. Modifications concentrate at functional centers of ribosomal and transfer RNAs, follow the canonical topology of N6-methyladenosine in coding transcripts, and form coordinated hotspots in immune regulatory genes. This first reference human RNome provides a foundation for understanding how RNA chemistry shapes cellular identity, human disease, and the development of RNA-based therapeutics.

59 BASIC BIOLOGICAL SCIENCES↗

Exabiome: Advancing Microbial Science through Exascale Computing

The Exabiome project seeks to improve the understanding of microbiomes through the development of methods for accelerating metagenomic science using exascale computing. This article gives an overview of scientific impact of the three components of the project: metagenome assembly, protein family detection, and comparative analysis of metagenomes. Exabiome developed MetaHipMer, the only metagenome assembler capable of scaling to full exascale systems. MetaHipMer has enabled ground-breaking assemblies on the Frontier supercomputer, with many scientific benefits, such as the discovery of rare species and viral genomes. To investigate protein families, Exabiome developed two exascale tools, PASTIS and HipMCL. Together, these can utilize exascale resources to understand the functional diversity of billions of dark matter proteins and novel protein families. For comparative analysis, Exabiome developed kmerprof, a tool that can be used to compare huge metagenomes for many different scientific purposes, for example, grouping human microbiomes according to body location.

59 BASIC BIOLOGICAL SCIENCES↗

Author Correction: Expanded encyclopaedias of DNA elements in the human and mouse genomes

In the version of this article initially published, two members of the ENCODE Project Consortium were missing from the author list. Rizi Ai (Department of Chemistry and Biochemistry, University of California, San Diego, La Jolla, CA, USA) and Shantao Li (Program in Computational Biology and Bioinformatics, Yale University, New Haven, CT, USA) are now included in the author list. These errors have been corrected in the online version of the article.

59 BASIC BIOLOGICAL SCIENCES↗

Science and Technology Review (May 2021)

Lawrence Livermore has been on the forefront of cancer research for over 60 years. Early interest in cancer statistics stemmed from the nature of work, particularly how radiation affects humans. The Department of Energy funded research to investigate the effects of radiation on workers with long term exposure. This research quickly morphed into a wider breadth of cancer research topics, including the use of advanced computational models to investigate mutations in genes. Livermore is regarded as a leader in cancer research, from the Human Genome Center to its participation in the National Cancer Institute’s “Moonshot” project. The highly interdisciplinary Laboratory unites research in one more example: bringing together cancer biology, 3D printing, high-performance computing, big data, and materials science to address this pressing medical challenge.

36 MATERIALS SCIENCE↗

Spatial top-down proteomics for the functional characterization of human kidney

Background: The Human Proteome Project has credibly detected nearly 93% of the roughly 20,000 proteins which are predicted by the human genome. However, the proteome is enigmatic, where alterations in amino acid sequences from polymorphisms and alternative splicing, errors in translation, and post-translational modifications result in a proteome depth estimated at several million unique proteoforms. Recently mass spectrometry has been demonstrated in several landmark efforts mapping the human proteoform landscape in bulk analyses. Herein, we developed an integrated workflow for characterizing proteoforms from human tissue in a spatially resolved manner by coupling laser capture microdissection, nanoliter-scale sample preparation, and mass spectrometry imaging. Results: Using healthy human kidney sections as the case study, we focused our analyses on the major functional tissue units including glomeruli, tubules, and medullary rays. After laser capture microdissection, these isolated functional tissue units were processed with microPOTS (microdroplet processing in one-pot for trace samples) for sensitive top-down proteomics measurement. This provided a quantitative database of 616 proteoforms that was further leveraged as a library for mass spectrometry imaging with near-cellular spatial resolution over the entire section. Notably, several mitochondrial proteoforms were found to be differentially abundant between glomeruli and convoluted tubules, and further spatial contextualization was provided by mass spectrometry imaging confirming unique differences identified by microPOTS, and further expanding the field-of-view for unique distributions such as enhanced abundance of a truncated form (1-74) of ubiquitin within cortical regions. Conclusions: We developed an integrated workflow to directly identify proteoforms and reveal their spatial distributions. Where of the 20 differentially abundant proteoforms identified as discriminate between tubules and glomeruli by microPOTS, the vast majority of tubular proteoforms were of mitochondrial origin (8 of 10) where discriminate proteoforms in glomeruli were primarily hemoglobin subunits (9 of 10). These trends were also identified within ion images demonstrating spatially resolved characterization of proteoforms that has the potential to reshape discovery-based proteomics because the proteoforms are the ultimate effector of cellular functions. Applications of this technology have the potential to unravel etiology and pathophysiology of disease states, informing on biologically active proteoforms, which remodel the proteomic landscape in chronic and acute disorders.

59 BASIC BIOLOGICAL SCIENCES↗

Tripal, a community update after 10 years of supporting open source, standards-based genetic, genomic and breeding databases

Abstract Online, open access databases for biological knowledge serve as central repositories for research communities to store, find and analyze integrated, multi-disciplinary datasets. With increasing volumes, complexity and the need to integrate genomic, transcriptomic, metabolomic, proteomic, phenomic and environmental data, community databases face tremendous challenges in ongoing maintenance, expansion and upgrades. A common infrastructure framework using community standards shared by many databases can reduce development burden, provide interoperability, ensure use of common standards and support long-term sustainability. Tripal is a mature, open source platform built to meet this need. With ongoing improvement since its first release in 2009, Tripal provides full functionality for searching, browsing, loading and curating numerous types of data and is a primary technology powering at least 31 publicly available databases spanning plants, animals and human data, primarily storing genomics, genetics and breeding data. Tripal software development is managed by a shared, inclusive governance structure including both project management and advisory teams. Here, we report on the most important and innovative aspects of Tripal after 11 years development, including integration of diverse types of biological data, successful collaborative projects across member databases, and support for implementing FAIR principles.

59 BASIC BIOLOGICAL SCIENCES↗

The Human Proteoform Project: Defining the human proteome

Proteins are the primary effectors of function in biology, and thus, complete knowledge of their structure and properties is fundamental to deciphering function in basic and translational research. The chemical diversity of proteins is expressed in their many proteoforms, which result from combinations of genetic polymorphisms, RNA splice variants, and posttranslational modifications. This knowledge is foundational for the biological complexes and networks that control biology yet remains largely unknown. We propose here an ambitious initiative to define the human proteome, that is, to generate a definitive reference set of the proteoforms produced from the genome. Several examples of the power and importance of proteoform-level knowledge in disease-based research are presented along with a call for improved technologies in a two-pronged strategy to the Human Proteoform Project.

59 BASIC BIOLOGICAL SCIENCES↗

Investigating genomic prediction strategies for grain carotenoid traits in a tropical/subtropical maize panel

Abstract Vitamin A deficiency remains prevalent on a global scale, including in regions where maize constitutes a high percentage of human diets. One solution for alleviating this deficiency has been to increase grain concentrations of provitamin A carotenoids in maize (Zea mays ssp. mays L.)—an example of biofortification. The International Maize and Wheat Improvement Center (CIMMYT) developed a Carotenoid Association Mapping panel of 380 inbred lines adapted to tropical and subtropical environments that have varying grain concentrations of provitamin A and other health-beneficial carotenoids. Several major genes have been identified for these traits, 2 of which have particularly been leveraged in marker-assisted selection. This project assesses the predictive ability of several genomic prediction strategies for maize grain carotenoid traits within and between 4 environments in Mexico. Ridge Regression-Best Linear Unbiased Prediction, Elastic Net, and Reproducing Kernel Hilbert Spaces had high predictive abilities for all tested traits (β-carotene, β-cryptoxanthin, provitamin A, lutein, and zeaxanthin) and outperformed Least Absolute Shrinkage and Selection Operator. Furthermore, predictive abilities were higher when using genome-wide markers rather than only the markers proximal to 2 or 13 genes. These findings suggest that genomic prediction models using genome-wide markers (and assuming equal variance of marker effects) are worthwhile for these traits even though key genes have already been identified, especially if breeding for additional grain carotenoid traits alongside β-carotene. Predictive ability was maintained for all traits except lutein in between-environment prediction. The TASSEL (Trait Analysis by aSSociation, Evolution, and Linkage) Genomic Selection plugin performed as well as other more computationally intensive methods for within-environment prediction. The findings observed herein indicate the utility of genomic prediction methods for these traits and could inform their resource-efficient implementation in biofortification breeding programs.

59 BASIC BIOLOGICAL SCIENCES↗

Harmonizing model organism data in the Alliance of Genome Resources

The Alliance of Genome Resources (the Alliance) is a combined effort of 7 knowledgebase projects: Saccharomyces Genome Database, WormBase, FlyBase, Mouse Genome Database, the Zebrafish Information Network, Rat Genome Database, and the Gene Ontology Resource. The Alliance seeks to provide several benefits: better service to the various communities served by these projects; a harmonized view of data for all biomedical researchers, bioinformaticians, clinicians, and students; and a more sustainable infrastructure. The Alliance has harmonized cross-organism data to provide useful comparative views of gene function, gene expression, and human disease relevance. The basis of the comparative views is shared calls of orthology relationships and the use of common ontologies. The key types of data are alleles and variants, gene function based on gene ontology annotations, phenotypes, association to human disease, gene expression, protein–protein and genetic interactions, and participation in pathways. The information is presented on uniform gene pages that allow facile summarization of information about each gene in each of the 7 organisms covered (budding yeast, roundworm Caenorhabditis elegans, fruit fly, house mouse, zebrafish, brown rat, and human). The harmonized knowledge is freely available on the alliancegenome.org portal, as downloadable files, and by APIs. We expect other existing and emerging knowledge bases to join in the effort to provide the union of useful data and features that each knowledge base currently provides.

59 BASIC BIOLOGICAL SCIENCES↗

Building a FAIR data ecosystem for incorporating single-cell transcriptomics data into agricultural genome to phenome research

Introduction The agriculture genomics community has numerous data submission standards available, but the standards for describing and storing single-cell (SC, e.g., scRNA- seq) data are comparatively underdeveloped. Methods To bridge this gap, we leveraged recent advancements in human genomics infrastructure, such as the integration of the Human Cell Atlas Data Portal with Terra, a secure, scalable, open-source platform for biomedical researchers to access data, run analysis tools, and collaborate. In parallel, the Single Cell Expression Atlas at EMBL-EBI offers a comprehensive data ingestion portal for high-throughput sequencing datasets, including plants, protists, and animals (including humans). Developing data tools connecting these resources would offer significant advantages to the agricultural genomics community. The FAANG data portal at EMBL-EBI emphasizes delivering rich metadata and highly accurate and reliable annotation of farmed animals but is not computationally linked to either of these resources. Results Herein, we describe a pilot-scale project that determines whether the current FAANG metadata standards for livestock can be used to ingest scRNA-seq datasets into Terra in a manner consistent with HCA Data Portal standards. Importantly, rich scRNA-seq metadata can now be brokered through the FAANG data portal using a semi-automated process, thereby avoiding the need for substantial expert curation. We have further extended the functionality of this tool so that validated and ingested SC files within the HCA Data Portal are transferred to Terra for further analysis. In addition, we verified data ingestion into Terra, hosted on Azure, and demonstrated the use of a workflow to analyze the first ingested porcine scRNA-seq dataset. Additionally, we have also developed prototype tools to visualize the output of scRNA-seq analyses on genome browsers to compare gene expression patterns across tissues and cell populations. This JBrowse tool now features distinct tracks, showcasing PBMC scRNA-seq alongside two bulk RNA-seq experiments. Discussion We intend to further build upon these existing tools to construct a scientist-friendly data resource and analytical ecosystem based on Findable, Accessible, Interoperable, and Reusable (FAIR) SC principles to facilitate SC-level genomic analysis through data ingestion, storage, retrieval, re-use, visualization, and comparative annotation across agricultural species.

Genetics & Heredity↗

Expanding the genomic encyclopedia of Actinobacteria with 824 isolate reference genomes

The phylum Actinobacteria includes important human pathogens like Mycobacterium tuberculosis and Corynebacterium diphtheriae and renowned producers of secondary metabolites of commercial interest, yet only a small part of its diversity is represented by sequenced genomes. Here, we present 824 actinobacterial isolate genomes in the context of a phylum-wide analysis of 6,700 genomes including public isolates and metagenome-assembled genomes (MAGs). We estimate that only 30%–50% of projected actinobacterial phylogenetic diversity possesses genomic representation via isolates and MAGs. A comparison of gene functions reveals novel determinants of host-microbe interaction as well as environment-specific adaptations such as potential antimicrobial peptides. We identify plasmids and prophages across isolates and uncover extensive prophage diversity structured mainly by host taxonomy. Analysis of >80,000 biosynthetic gene clusters reveals that horizontal gene transfer and gene loss shape secondary metabolite repertoire across taxa. Our observations illustrate the essential role of and need for high-quality isolate genome sequences.

59 BASIC BIOLOGICAL SCIENCES↗

Microbial Tracking-2, a metagenomics analysis of bacteria and fungi onboard the International Space Station

The International Space Station (ISS) is a unique and complex built environment with the ISS surface microbiome originating from crew and cargo or from life support recirculation in an almost entirely closed system. The Microbial Tracking 1 (MT-1) project was the first ISS environmental surface study to report on the metagenome profiles without using whole-genome amplification. The study surveyed the microbial communities from eight surfaces over a 14-month period. The Microbial Tracking 2 (MT-2) project aimed to continue the work of MT-1, sampling an additional four flights from the same locations, over another 14 months. Eight surfaces across the ISS were sampled with sterile wipes and processed upon return to Earth. DNA extracted from the processed samples (and controls) were treated with propidium monoazide (PMA) to detect intact/viable cells or left untreated and to detect the total DNA population (free DNA/compromised cells/intact cells/viable cells). DNA extracted from PMA-treated and untreated samples were analyzed using shotgun metagenomics. Samples were cultured for bacteria and fungi to supplement the above results. Staphylococcus sp. and Malassezia sp. were the most represented bacterial and fungal species, respectively, on the ISS. Overall, the ISS surface microbiome was dominated by organisms associated with the human skin. Multi-dimensional scaling and differential abundance analysis showed significant temporal changes in the microbial population but no spatial differences. The ISS antimicrobial resistance gene profiles were however more stable over time, with no differences over the 5-year span of the MT-1 and MT-2 studies. Twenty-nine antimicrobial resistance genes were detected across all samples, with macrolide/lincosamide/streptogramin resistance being the most widespread. Metagenomic assembled genomes were reconstructed from the dataset, resulting in 82 MAGs. Functional assessment of the collective MAGs showed a propensity for amino acid utilization over carbohydrate metabolism. Co-occurrence analyses showed strong associations between bacterial and fungal genera. Culture analysis showed the microbial load to be on average 3.0 × 10 5 cfu/m 2 . Utilizing various metagenomics analyses and culture methods, we provided a comprehensive analysis of the ISS surface microbiome, showing microbial burden, bacterial and fungal species prevalence, changes in the microbiome, and resistome over time and space, as well as the functional capabilities and microbial interactions of this unique built microbiome. Data from this study may help to inform policies for future space missions to ensure an ISS surface microbiome that promotes astronaut health and spacecraft integrity.

59 BASIC BIOLOGICAL SCIENCES↗

DL-TODA: A Deep Learning Tool for Omics Data Analysis

Metagenomics is a technique for genome-wide profiling of microbiomes; this technique generates billions of DNA sequences called reads. Given the multiplication of metagenomic projects, computational tools are necessary to enable the efficient and accurate classification of metagenomic reads without needing to construct a reference database. The program DL-TODA presented here aims to classify metagenomic reads using a deep learning model trained on over 3000 bacterial species. A convolutional neural network architecture originally designed for computer vision was applied for the modeling of species-specific features. Using synthetic testing data simulated with 2454 genomes from 639 species, DL-TODA was shown to classify nearly 75% of the reads with high confidence. The classification accuracy of DL-TODA was over 0.98 at taxonomic ranks above the genus level, making it comparable with Kraken2 and Centrifuge, two state-of-the-art taxonomic classification tools. DL-TODA also achieved an accuracy of 0.97 at the species level, which is higher than 0.93 by Kraken2 and 0.85 by Centrifuge on the same test set. Application of DL-TODA to the human oral and cropland soil metagenomes further demonstrated its use in analyzing microbiomes from diverse environments. Compared to Centrifuge and Kraken2, DL-TODA predicted distinct relative abundance rankings and is less biased toward a single taxon.

59 BASIC BIOLOGICAL SCIENCES↗

A variant selection framework for genome graphs

Abstract Motivation Variation graph representations are projected to either replace or supplement conventional single genome references due to their ability to capture population genetic diversity and reduce reference bias. Vast catalogues of genetic variants for many species now exist, and it is natural to ask which among these are crucial to circumvent reference bias during read mapping. Results In this work, we propose a novel mathematical framework for variant selection, by casting it in terms of minimizing variation graph size subject to preserving paths of length α with at most δ differences. This framework leads to a rich set of problems based on the types of variants [e.g. single nucleotide polymorphisms (SNPs), indels or structural variants (SVs)], and whether the goal is to minimize the number of positions at which variants are listed or to minimize the total number of variants listed. We classify the computational complexity of these problems and provide efficient algorithms along with their software implementation when feasible. We empirically evaluate the magnitude of graph reduction achieved in human chromosome variation graphs using multiple α and δ parameter values corresponding to short and long-read resequencing characteristics. When our algorithm is run with parameter settings amenable to long-read mapping (α = 10 kbp, δ = 1000), 99.99% SNPs and 73% SVs can be safely excluded from human chromosome 1 variation graph. The graph size reduction can benefit downstream pan-genome analysis. Availability and implementation https://github.com/AT-CG/VF. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Developing a pipeline to expand the genetic code of diverse bacteria for microbial engineering

Microbial biotechnologies are key to addressing grand challenges to promote human health, reverse carbon emissions, recycle mixed plastic waste, remediate contaminated soils, and achieve sustainable economies. Synthetic biology has enabled design of diverse microbes and their proteins for useful purposes, but the narrowness of the natural genetic code limits functional diversity (e.g., biosynthesis) of engineered microbes. The natural genetic code defines the fundamental rules of translating genetic information into proteins comprised of 22 ‘canonical’ amino acids. However, using a technique called genetic code expansion (GCE), the chemical properties and therefore functions of proteins can be transformed by incorporation of one or more of ~200 chemically diverse ‘non-canonical’ amino acids. The effective application of genetic code expansion in diverse microbes has the potential to revolutionize biotechnology. However, despite over 50 years of research and its transformative potential, the application of genetic code expansion has been limited to a handful of bacterial species. In this project, we will perform three tasks to both overcome the barriers that prevent wide spread adoption of GCE as molecular tool and demonstrate its potential for biotechnological applications. Specifically, we will (1) develop a genetic engineering methodology that will enable use of GCE in a broad range of bacterial hosts, (2) use high-throughput functional genomics methods to identify physiological responses to both genetic code expansion and exposure to non-canonical amino acids in three different bacteria, and (3) demonstrate an application of GCE by selectively incorporate non-canonical amino acids into surface displayed peptides such as those used for biomining.

59 BASIC BIOLOGICAL SCIENCES↗

Data for EMSL Project 60929 from August 2023: PI Goemann MONet Request

Just as humans rely on a healthy gut microbiome for resilience to illness, plants rely on a healthy root microbiome for resilience to environmental abiotic stress (heat, drought). To achieve a healthy root microbiome, plants release carbon (C)-rich compounds as root exudates to stimulate microbial activity and increase local nutrient mineralization. However, the enhanced performance comes at a cost: up to 44% of a plant’s C can be lost to root exudates, diverting C from plant growth and respiration. Critical knowledge gaps include how the ‘C cost’ is managed and how root exudates alter the microbiome under different environmental conditions. In addition, historical climate conditions, particularly mean annual precipitation, is known to shape local soil microbiomes and alter their sensitivity to drought. Therefore, studies that better characterize the plant-microbe responses to environmental stress will aid in efforts to harness the microbiome to improve crop resilience. However, current knowledge gaps make it challenging to engineer beneficial plant-microbe interactions to improve plant productivity in agricultural systems and to predict how increased climate variability will alter terrestrial C fluxes and climate feedbacks. To fill this knowledge gap our research group at Montana State University – Bozeman is currently studying blue grama (Bouteloua gracilis), a prairie grass native across the Northern Great Plains, as a model for drought tolerance. Our goal is to investigate the above- and belowground responses of blue grama to drought and heat stress to improve our understanding of stress-induced carbon allocation and plant-microbe interactions. Most recently, we investigated the influence of climate history on the blue grama drought response. We collected soil from three blue grama-dominated sites (those proposed to sample here) across a 150 mm mean annual precipitation gradient in SW Montana, USA, to use as inoculum for a greenhouse drought experiment. Preliminary results indicate that soil climate history has a strong influence on the blue grama physiological response to drought as well as on the chemical composition of root exudates and rhizosphere microbiome composition. Metabarcoding data from this experiment is scheduled to be submitted to public databases within the next year. Having in-depth analyses of the soil biogeochemistry and metagenomic composition through the MONet project at each of the field sites associated with this experiment will allow us to link underlying ecological processes with observed patterns of plant growth and productivity at each site. In addition, we plan to utilize the MONet database for future meta-analyses to compare the genomic and biogeochemical signatures of our field sites to others across a wider precipitation gradient throughout the native range of blue grama. This will further provide critical insights into the mechanisms that drive ecosystem functioning and resilience to drought stress.

Peyton, Brent↗