Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

An expanded transcriptome atlas for Bacteroides thetaiotaomicron reveals a small RNA that modulates tetracycline sensitivity

Plasticity in gene expression allows bacteria to adapt to diverse environments. This is particularly relevant in the dynamic niche of the human intestinal tract; however, transcriptional networks remain largely unknown for gut-resident bacteria. Here we apply differential RNA sequencing (RNA-seq) and conventional RNA-seq to the model gut bacterium Bacteroides thetaiotaomicron to map transcriptional units and profile their expression levels across 15 in vivo-relevant growth conditions. We infer stress- and carbon source-specific transcriptional regulons and expand the annotation of small RNAs (sRNAs). Integrating this expression atlas with published transposon mutant fitness data, we predict conditionally important sRNAs. These include MasB, which downregulates tetracycline tolerance. Using MS2 affinity purification and RNA-seq, we identify a putative MasB target and assess its role in the context of the MasB-associated phenotype. These data—publicly available through the Theta-Base web browser (http://micromix.helmholtz-hiri.de/bacteroides/)—constitute a valuable resource for the microbiome community.

59 BASIC BIOLOGICAL SCIENCES↗

Microbial polyphenol metabolism is part of the thawing permafrost carbon cycle

Abstract With rising global temperatures, permafrost carbon stores are vulnerable to microbial degradation. The enzyme latch theory states that polyphenols should accumulate in saturated peatlands due to diminished phenol oxidase activity, inhibiting resident microbes and promoting carbon stabilization. Pairing microbiome and geochemical measurements along a permafrost thaw-induced saturation gradient in Stordalen Mire, a model Arctic peatland, we confirmed a negative relationship between phenol oxidase expression and saturation but failed to support other trends predicted by the enzyme latch. To inventory alternative polyphenol removal strategies, we built CAMPER, a gene annotation tool leveraging polyphenol enzyme knowledge gleaned across microbial ecosystems. Applying CAMPER to genome-resolved metatranscriptomes, we identified genes for diverse polyphenol-active enzymes expressed by various microbial lineages under a range of redox conditions. This shifts the paradigm that polyphenols stabilize carbon in saturated soils and highlights the need to consider both oxic and anoxic polyphenol metabolisms to understand carbon cycling in changing ecosystems.

54 ENVIRONMENTAL SCIENCES↗

A call for caution in the biological interpretation of viral auxiliary metabolic genes

Virus-encoded auxiliary metabolic genes (AMGs) are non-essential genes that increase viral fitness by maintaining or manipulating host metabolism during infection. AMGs are intriguing from an evolutionary perspective, as most viral genomes are highly compact and have limited coding capacity for accessory genes. Advances in viral (meta)genomics have expanded the detection of putative AMGs from viruses in diverse environments. However, this has also led to many instances of misannotation due to the limitations of annotation tools, resulting in misinterpretations about the roles of some viral genes. Here, we highlight studies that support claims about AMGs with more than just function predictions for guidance on best practices. We then propose the adoption of an expanded, inclusive view of all genes auxiliary to core viral functions with the term ‘auxiliary viral genes’ (AVGs), alongside an associated eco-evolutionary framework for considering the types of analyses that can better support claims made about AVGs.

Environmental microbiology↗

A genomic perspective on fungal diversity and evolution

Originating from aquatic unicellular ancestors, over the course of ~1 billion years, the fungi have evolved to occupy nearly all aerobic environments on the planet, diversified into millions of different ‘species’ and have developed complex multicellular structures. Their relatively small, simple genomes have facilitated massive-scale sequencing and allowed us to explore genome evolution across an ancient eukaryotic kingdom. With thousands of genomes from diverse lineages now available, this Review will discuss insights into fungal biology and evolution gleaned with genomics and other multi-omics approaches. Using published genomes available through GenBank and the Joint Genome Institute’s MycoCosm platform, we generated kingdom-wide phylogenies and used them to highlight how fungal genomes have changed over time. With this phylogeny as a guide, we also discuss major evolutionary transitions that occurred across the fungal kingdom. Although progress has been made, these efforts are hampered by biases in genome representation and limited characterization of gene functions. Here, in this study, we discuss these challenges and possible future directions to address them, including initiatives to characterize conserved genes of unknown function and scale up sequencing towards 10,000 annotated fungal genomes.

Mondo, Stephen J. [USDOE Joint Genome Institute (J↗

Modelling kidney disease using ontology: insights from the Kidney Precision Medicine Project

An important need exists to better understand and stratify kidney disease according to its underlying pathophysiology in order to develop more precise and effective therapeutic agents. National collaborative efforts such as the Kidney Precision Medicine Project are working towards this goal through the collection and integration of large, disparate clinical, biological and imaging data from patients with kidney disease. Ontologies are powerful tools that facilitate these efforts by enabling researchers to organize and make sense of different data elements and the relationships between them. Ontologies are critical to support the types of big data analysis necessary for kidney precision medicine, where heterogeneous clinical, imaging and biopsy data from diverse sources must be combined to define a patient’s phenotype. Here, the development of two new ontologies — the Kidney Tissue Atlas Ontology and the Ontology of Precision Medicine and Investigation — will support the creation of the Kidney Tissue Atlas, which aims to provide a comprehensive molecular, cellular and anatomical map of the kidney. These ontologies will improve the annotation of kidney-relevant data, and eventually lead to new definitions of kidney disease in support of precision medicine.

60 APPLIED LIFE SCIENCES↗

Genomic mechanisms of climate adaptation in polyploid bioenergy switchgrass

Long-term climate change and periodic environmental extremes threaten food and fuel security and global crop productivity. Although molecular and adaptive breeding strategies can buffer the effects of climatic stress and improve crop resilience, these approaches require sufficient knowledge of the genes that underlie productivity and adaptation—knowledge that has been limited to a small number of well-studied model systems. Here we present the assembly and annotation of the large and complex genome of the polyploid bioenergy crop switchgrass ( Panicum virgatum ). Analysis of biomass and survival among 732 resequenced genotypes, which were grown across 10 common gardens that span 1,800 km of latitude, jointly revealed extensive genomic evidence of climate adaptation. Climate–gene–biomass associations were abundant but varied considerably among deeply diverged gene pools. Furthermore, we found that gene flow accelerated climate adaptation during the postglacial colonization of northern habitats through introgression of alleles from a pre-adapted northern gene pool. The polyploid nature of switchgrass also enhanced adaptive potential through the fractionation of gene function, as there was an increased level of heritable genetic diversity on the nondominant subgenome. In addition to investigating patterns of climate adaptation, the genome resources and gene–trait associations developed here provide breeders with the necessary tools to increase switchgrass yield for the sustainable production of bioenergy.

09 BIOMASS FUELS↗

A mycobacterial ABC transporter mediates the uptake of hydrophilic compounds

Mycobacterium tuberculosis (Mtb) is an obligate human pathogen and the causative agent of tuberculosis. Although Mtb can synthesize vitamin B 12 (cobalamin) de novo, uptake of cobalamin has been linked to pathogenesis of tuberculosis. Mtb does not encode any characterized cobalamin transporter; however, the gene rv1819c was found to be essential for uptake of cobalamin. This result is difficult to reconcile with the original annotation of Rv1819c as a protein implicated in the transport of antimicrobial peptides such as bleomycin. In addition, uptake of cobalamin seems inconsistent with the amino acid sequence, which suggests that Rv1819c has a bacterial ATP-binding cassette (ABC)-exporter fold. In this paper, we present structures of Rv1819c, which reveal that the protein indeed contains the ABC-exporter fold, as well as a large water-filled cavity of about 7,700 Å, which enables the protein to transport the unrelated hydrophilic compounds bleomycin and cobalamin. On the basis of these structures, we propose that Rv1819c is a multi-solute transporter for hydrophilic molecules, analogous to the multidrug exporters of the ABC transporter family, which pump out structurally diverse hydrophobic compounds from cells.

59 BASIC BIOLOGICAL SCIENCES↗

An atlas of dynamic chromatin landscapes in mouse fetal development

The Encyclopedia of DNA Elements (ENCODE) project has established a genomic resource for mammalian development, profiling a diverse panel of mouse tissues at 8 developmental stages from 10.5 days after conception until birth, including transcriptomes, methylomes and chromatin states. Here we systematically examined the state and accessibility of chromatin in the developing mouse fetus. In total we performed 1,128 chromatin immunoprecipitation with sequencing (ChIP–seq) assays for histone modifications and 132 assay for transposase-accessible chromatin using sequencing (ATAC–seq) assays for chromatin accessibility across 72 distinct tissue-stages. We used integrative analysis to develop a unified set of chromatin state annotations, infer the identities of dynamic enhancers and key transcriptional regulators, and characterize the relationship between chromatin state and accessibility during developmental gene regulation. We also leveraged these data to link enhancers to putative target genes and demonstrate tissue-specific enrichments of sequence variants associated with disease in humans. The mouse ENCODE data sets provide a compendium of resources for biomedical researchers and achieve, to our knowledge, the most comprehensive view of chromatin dynamics during mammalian fetal development to date.

59 BASIC BIOLOGICAL SCIENCES↗

Perspectives on ENCODE

The Encyclopedia of DNA Elements (ENCODE) Project launched in 2003 with the long-term goal of developing a comprehensive map of functional elements in the human genome. These included genes, biochemical regions associated with gene regulation (for example, transcription factor binding sites, open chromatin, and histone marks) and transcript isoforms. The marks serve as sites for candidate cis-regulatory elements (cCREs) that may serve functional roles in regulating gene expression. The project has been extended to model organisms, particularly the mouse. Finally, in the third phase of ENCODE, nearly a million and more than 300,000 cCRE annotations have been generated for human and mouse, respectively, and these have provided a valuable resource for the scientific community.

59 BASIC BIOLOGICAL SCIENCES↗

Unraveling the functional dark matter through global metagenomics

Metagenomes encode an enormous diversity of proteins, reflecting a multiplicity of functions and activities1,2. Exploration of this vast sequence space has been limited to a comparative analysis against reference microbial genomes and protein families derived from those genomes. Here, to examine the scale of yet untapped functional diversity beyond what is currently possible through the lens of reference genomes, we develop a computational approach to generate reference-free protein families from the sequence space in metagenomes. We analyse 26,931 metagenomes and identify 1.17 billion protein sequences longer than 35 amino acids with no similarity to any sequences from 102,491 reference genomes or the Pfam database3. Using massively parallel graph-based clustering, we group these proteins into 106,198 novel sequence clusters with more than 100 members, doubling the number of protein families obtained from the reference genomes clustered using the same approach. We annotate these families on the basis of their taxonomic, habitat, geographical and gene neighbourhood distributions and, where sufficient sequence diversity is available, predict protein three-dimensional models, revealing novel structures. Overall, our results uncover an enormously diverse functional space, highlighting the importance of further exploring the microbial functional dark matter.

54 ENVIRONMENTAL SCIENCES↗

In vivo mapping of mutagenesis sensitivity of human enhancers

Distant-acting enhancers are central to human development1. However, our limited understanding of their functional sequence features prevents the interpretation of enhancer mutations in disease2. Here we determined the functional sensitivity to mutagenesis of human developmental enhancers in vivo. Focusing on seven enhancers that are active in the developing brain, heart, limb and face, we created over 1,700 transgenic mice for over 260 mutagenized enhancer alleles. Systematic mutation of 12-base-pair blocks collectively altered each sequence feature in each enhancer at least once. We show that 69% of all blocks are required for normal in vivo activity, with mutations more commonly resulting in loss (60%) than in gain (9%) of function. Using predictive modelling, we annotated critical nucleotides at the base-pair resolution. The vast majority of motifs predicted by these machine learning models (88%) coincided with changes in in vivo function, and the models showed considerable sensitivity, identifying 59% of all functional blocks. Taken together, our results reveal that human enhancers contain a high density of sequence features that are required for their normal in vivo function and provide a rich resource for further exploration of human enhancer logic.

Kosicki, Michael↗

MEMOTE for standardized genome-scale metabolic model testing

Reconstructing metabolic reaction networks enables the development of testable hypotheses of an organism’s metabolism under different conditions. State-of-the-art genome-scale metabolic models (GEMs) can include thousands of metabolites and reactions that are assigned to subcellular locations. Gene–protein–reaction (GPR) rules and annotations using database information can add meta-information to GEMs. GEMs with metadata can be built using standard reconstruction protocols, and guidelines have been put in place for tracking provenance and enabling interoperability, but a standardized means of quality control for GEMs is lacking. Here we report a community effort to develop a test suite named MEMOTE (for metabolic model tests) to assess GEM quality.

59 BASIC BIOLOGICAL SCIENCES↗

A unified catalog of 204,938 reference genomes from the human gut microbiome

Comprehensive, high-quality reference genomes are required for functional characterization and taxonomic assignment of the human gut microbiota. We present the Unified Human Gastrointestinal Genome (UHGG) collection, comprising 204,938 nonredundant genomes from 4,644 gut prokaryotes. These genomes encode >170 million protein sequences, which we collated in the Unified Human Gastrointestinal Protein (UHGP) catalog. The UHGP more than doubles the number of gut proteins in comparison to those present in the Integrated Gene Catalog. More than 70% of the UHGG species lack cultured representatives, and 40% of the UHGP lack functional annotations. Intraspecies genomic variation analyses revealed a large reservoir of accessory genes and single-nucleotide variants, many of which are specific to individual human populations. The UHGG and UHGP collections will enable studies linking genotypes to phenotypes in the human gut microbiome.

59 BASIC BIOLOGICAL SCIENCES↗

Auto-deconvolution and molecular networking of gas chromatography–mass spectrometry data

We engineered a machine learning approach, MSHub, to enable auto-deconvolution of gas chromatography–mass spectrometry (GC–MS) data. We then designed workflows to enable the community to store, process, share, annotate, compare and perform molecular networking of GC–MS data within the Global Natural Product Social (GNPS) Molecular Networking analysis platform. MSHub/GNPS performs auto-deconvolution of compound fragmentation patterns via unsupervised non-negative matrix factorization and quantifies the reproducibility of fragmentation patterns across samples.

47 OTHER INSTRUMENTATION↗

Cryo-EM model validation recommendations based on outcomes of the 2019 EMDataResource challenge

This paper describes outcomes of the 2019 Cryo-EM Model Challenge. The goals were to (1) assess the quality of models that can be produced from cryogenic electron microscopy (cryo-EM) maps using current modeling software, (2) evaluate reproducibility of modeling results from different software developers and users and (3) compare performance of current metrics used for model evaluation, particularly Fit-to-Map metrics, with focus on near-atomic resolution. Our findings demonstrate the relatively high accuracy and reproducibility of cryo-EM models derived by 13 participating teams from four benchmark maps, including three forming a resolution series (1.8 to 3.1 Å). The results permit specific recommendations to be made about validating near-atomic cryo-EM structures both in the context of individual experiments and structure data archives such as the Protein Data Bank. We recommend the adoption of multiple scoring parameters to provide full and objective annotation and assessment of the model, reflective of the observed cryo-EM map density.

59 BASIC BIOLOGICAL SCIENCES↗

Persistence and plasticity in bacterial gene regulation

Organisms orchestrate cellular functions through transcription factor (TF) interactions with their target genes, although these regulatory relationships are largely unknown in most species. Here we report a high-throughput approach for characterizing TF-target gene interactions across species and its application to 354 TFs across 48 bacteria, generating 17,000 genome-wide binding maps. This dataset revealed themes of ancient conservation and rapid evolution of regulatory modules. We observed rewiring, where the TF sensing and regulatory role is maintained while the arrangement and identity of target genes diverges, in some cases encoding entirely new functions. We further integrated phenotypic information to define new functional regulatory modules and pathways. So, we identified 242 new TF DNA binding motifs, including a 70% increase of known Escherichia coli motifs and the first annotation in Pseudomonas simiae, revealing deep conservation in bacterial promoter architecture. Our method provides a versatile tool for functional characterization of genetic pathways in prokaryotes and eukaryotes.

59 BASIC BIOLOGICAL SCIENCES↗

Cognitive analysis of metabolomics data for systems biology

Cognitive computing is revolutionizing the way big data are processed and integrated, with artificial intelligence (AI) natural language processing (NLP) platforms helping researchers to efficiently search and digest the vast scientific literature. Most available platforms have been developed for biomedical researchers, but new NLP tools are emerging for biologists in other fields and an important example is metabolomics. NLP provides literature-based contextualization of metabolic features that decreases the time and expert-level subject knowledge required during the prioritization, identification and interpretation steps in the metabolomics data analysis pipeline. Here, we describe and demonstrate four workflows that combine metabolomics data with NLP-based literature searches of scientific databases to aid in the analysis of metabolomics data and their biological interpretation. Additionally, the four procedures can be used in isolation or consecutively, depending on the research questions. The first, used for initial metabolite annotation and prioritization, creates a list of metabolites that would be interesting for follow-up. The second workflow finds literature evidence of the activity of metabolites and metabolic pathways in governing the biological condition on a systems biology level. The third is used to identify candidate biomarkers, and the fourth looks for metabolic conditions or drug-repurposing targets that the two diseases have in common. The protocol can take 1–4 h or more to complete, depending on the processing time of the various software used.

59 BASIC BIOLOGICAL SCIENCES↗

Small-wedge synchrotron and serial XFEL datasets for Cysteinyl leukotriene GPCRs

Structural studies of challenging targets such as G protein-coupled receptors (GPCRs) have accelerated during the last several years due to the development of new approaches, including small-wedge and serial crystallography. Here, we describe the deposition of seven datasets consisting of X-ray diffraction images acquired from lipidic cubic phase (LCP) grown microcrystals of two human GPCRs, Cysteinyl leukotriene receptors 1 and 2 (CysLT 1 R and CysLT 2 R), in complex with various antagonists. Five datasets were collected using small-wedge synchrotron crystallography (SWSX) at the European Synchrotron Radiation Facility with multiple crystals under cryo-conditions. Two datasets were collected using X-ray free electron laser (XFEL) serial femtosecond crystallography (SFX) at the Linac Coherent Light Source, with microcrystals delivered at room temperature into the beam within LCP matrix by a viscous media microextrusion injector. All seven datasets have been deposited in the open-access databases Zenodo and CXIDB. Here, we describe sample preparation and annotate crystallization conditions for each partial and full datasets. We also document full processing pipelines and provide wrapper scripts for SWSX and SFX data processing. A Correction to this paper has been published: https://doi.org/10.1038/s41597-020-00759-w

97 MATHEMATICS AND COMPUTING↗