Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Functional genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Resistance of virus to extinction on bottleneck passages: study of a decaying and fluctuating pattern of fitness loss

RNA viruses display high mutation rates and their populations replicate as dynamic and complex mutant distributions, termed viral quasispecies. Repeated genetic bottlenecks, which experimentally are carried out through serial plaque-to-plaque transfers of the virus, lead to fitness decrease (measured here as diminished capacity to produce infectious progeny). Here we report an analysis of fitness evolution of several low fitness foot-and-mouth disease virus clones subjected to 50 plaque-to-plaque transfers. Unexpectedly, fitness decrease, rather than being continuous and monotonic, displayed a fluctuating pattern, which was influenced by both the virus and the state of the host cell as shown by effects of recent cell passage history. The amplitude of the fluctuations increased as fitness decreased, resulting in a remarkable resistance of virus to extinction. Whereas the frequency distribution of fitness in control (independent) experiments follows a log-normal distribution, the probability of fitness values in the evolving bottlenecked populations fitted a Weibull distribution. We suggest that multiple functions of viral genomic RNA and its encoded proteins, subjected to high mutational pressure, interact with cellular components to produce this nontrivial, fluctuating pattern.

Serial Passage↗

Research from the NASA Twins Study and Omics in Support of Mars Missions

The NASA Twins Study, NASA's first foray into integrated omic studies in humans, illustrates how an integrated omics approach can be brought to bear on the challenges to human health and performance on a Mars mission. The NASA Twins Study involves US Astronaut Scott Kelly and his identical twin brother, Mark Kelly, a retired US Astronaut. No other opportunity to study a twin pair for a prolonged period with one subject in space and one on the ground is available for the foreseeable future. A team of 10 principal investigators are conducting the Twins Study, examining a very broad range of biological functions including the genome, epigenome, transcriptome, proteome, metabolome, gut microbiome, immunological response to vaccinations, indicators of atherosclerosis, physiological fluid shifts, and cognition. A novel aspect of the study is the integrated study of molecular, physiological, cognitive, and microbiological properties. Major sample and data collection from both subjects for this study began approximately six months before Scott Kelly's one year mission on the ISS, continue while Scott Kelly is in flight and will conclude approximately six months after his return to Earth. Mark Kelly will remain on Earth during this study, in a lifestyle unconstrained by this study, thereby providing a measure of normal variation in the properties being studied. An overview of initial results and the future plans will be described as well as the technological and ethical issues raised for spaceflight studies involving omics.

Kundrot, C.↗

Novel Approach to Quantification of Telomere Length with Direct Nanopore Sequencing and PCR Amplification

The ends of human chromosomes contain telomeres, or tandem arrays of repeating DNA sequences capped by multiple associated proteins that protect chromosomal ends from degradation. Telomeres function to preserve genomic stability by preventing natural chromosomal ends from being recognized as broken DNA double-strand breaks and triggering inappropriate DNA damage responses. Mounting evidence shows telomere length is an inherited trait that decreases with cellular division and normal aging. In addition, telomere length also appears to be influenced by other factors such as cellular oxidative stress, radiation and mechanical unloading of tissues as in microgravity. To measure these potential effects of the space environment on telomere lengths and cellular aging and regenerative potential we developed a novel telomere measurement approach based on nanopore sequencing of PCR amplified bar-coded chromosome termini. Specifically, telomeres can be directly enriched using barcode sequences ligated to the end of a free end- repaired telomere using the WetLab-2 facility SmartCycler on ISS. Prior to the ligation and amplification protocol a proteinase K digestion of capping proteins followed by a single 95-degree C heat denaturation of the protease is included. After digestion and bar-code ligation, PCR amplification will initiate with the ligated barcoded sequence, suppressing amplification of intra-genomic fragments and resulting in long read barcoded telomere amplicons including the nanopore motor protein sequences. Purified PCR amplicons are then used for nanopore sequencing library generation by simple addition of motor proteins and sequencing library is loaded into the MinION nanopore DNA-sequencer. Amplicon sequence reads from the nanopore device can be base-called quickly on ISS due to barcoding ligation and subsequent PCR amplification enhancing the telomere sequence resolution. If successfully implemented on ISS this technique will provide a novel means of measuring regenerative ability of somatic stem cells in astronauts, and of determining whether spaceflight in microgravity alters their telomere lengths and causes premature cellular aging.

Ma, Kristin R.↗

Integrating functional scoring and regulatory data to predict the effect of non-coding SNPs in a complex neurological disease

Abstract Most SNPs associated with complex diseases seem to lie in non-coding regions of the genome; however, their contribution to gene expression and disease phenotype remains poorly understood. Here, we established a workflow to provide assistance in prioritising the functional relevance of non-coding SNPs of candidate genes as susceptibility loci in polygenic neurological disorders. To illustrate the applicability of our workflow, we considered the multifactorial disorder migraine as a model to follow our step-by-step approach. We annotated the overlap of selected SNPs with regulatory elements and assessed their potential impact on gene expression based on publicly available prediction algorithms and functional genomics information. Some migraine risk loci have been hypothesised to reside in non-coding regions and to be implicated in the neurotransmission pathway. In this study, we used a set of 22 non-coding SNPs from neurotransmission and synaptic machinery-related genes previously suggested to be involved in migraine susceptibility based on our candidate gene association studies. After prioritising these SNPs, we focused on non-reported ones that demonstrated high regulatory potential: (1) VAMP2_rs1150 (3′ UTR) was predicted as a target of hsa-mir-5010-3p miRNA, possibly disrupting its own gene expression; (2) STX1A_rs6951030 (proximal enhancer) may affect the binding affinity of zinc-finger transcription factors (namely ZNF423) and disturb TBL2 gene expression; and (3) SNAP25_rs2327264 (distal enhancer) expected to be in a binding site of ONECUT2 transcription factor. This study demonstrated the applicability of our practical workflow to facilitate the prioritisation of potentially relevant non-coding SNPs and predict their functional impact in multifactorial neurological diseases.

Felício, Daniela↗

Multi-genome Phage Annotation Toolkit and Evaluator

Summary: To address the need for improved tools for annotation and comparative genomics of bacteriophage genomes, we developed multiPhATE2. As an extension of the multiPhATE code, multiPhATE2 includes comparative genomics codes for gene matching among sets of input bacteriophage genomes, and scales well to large input data sets due to incorporation of multiprocessing in the functional annotation and comparative genomics subsystems. Furthermore, additional search algorithms and databases have been added to the functional annotation subsystem. MultiPhATE2 was implemented in Python 3.7, and runs as a command-line code under Linux or MAC-OS.

Kimbrel, JeffreyA.↗

Cross-family and phage-specific gene requirements for Klebsiella infection revealed by scalable RB-TnSeq genetic screens.

Bacteriophages are being cataloged at an accelerating pace and are recognized as key players in nutrient and energy cycling across ecosystems. Yet the bacterial genetic determinants that govern phage-host specificity and infection success remain poorly understood, particularly in clinically and ecologically important genera such as Klebsiella where prior receptor characterization has been almost entirely limited to capsulated strains. Here we used a randomly barcoded, genome-wide, loss-of-function transposon mutant library (RB-TnSeq) of Klebsiella sp. M5al, a naturally acapsular, nitrogen-fixing rhizobacterium, to generate the first systematic, cross-family map of phage receptor gene dependencies in Klebsiella. Challenging the library against 25 double-stranded DNA phages spanning five families in 213 parallel assays, we identified 42 bacterial genes associated with phage infection, of which 15 had no prior association with phage infection in any bacterial system. Disruption of surface receptor biosynthesis genes conferred cross-resistance across multiple phage families, while intracellular gene disruptions had predominantly phage-specific effects. Clonal validation of eight genes confirmed LPS outer core biosynthesis genes as primary receptor determinants alongside additional host factors spanning outer membrane transport, cofactor biosynthesis, and two-component signaling. Comparative analysis across all 25 phages revealed that phage genus rather than family is the stronger predictor of host gene dependency profiles, a finding with direct implications for the functional annotation of uncharacterized phage isolates and rational phage cocktail design. Together, these findings provide a community resource for linking phage genomic diversity to functional host interaction space in this ecologically and clinically important genus.

Gittrich, Marissa R↗

A roadmap for the functional annotation of protein families: a community perspective

Over the last 25 years, biology has entered the genomic era and is becoming a science of ‘big data’. Most interpretations of genomic analyses rely on accurate functional annotations of the proteins encoded by more than 500 000 genomes sequenced to date. By different estimates, only half the predicted sequenced proteins carry an accurate functional annotation, and this percentage varies drastically between different organismal lineages. Such a large gap in knowledge hampers all aspects of biological enterprise and, thereby, is standing in the way of genomic biology reaching its full potential. A brainstorming meeting to address this issue funded by the National Science Foundation was held during 3–4 February 2022. Bringing together data scientists, biocurators, computational biologists and experimentalists within the same venue allowed for a comprehensive assessment of the current state of functional annotations of protein families. Further, major issues that were obstructing the field were identified and discussed, which ultimately allowed for the proposal of solutions on how to move forward.

59 BASIC BIOLOGICAL SCIENCES↗

Predictions of rhizosphere microbiome dynamics with a genome-informed and trait-based energy budget model

Abstract Soil microbiomes are highly diverse, and to improve their representation in biogeochemical models, microbial genome data can be leveraged to infer key functional traits. By integrating genome-inferred traits into a theory-based hierarchical framework, emergent behaviour arising from interactions of individual traits can be predicted. Here we combine theory-driven predictions of substrate uptake kinetics with a genome-informed trait-based dynamic energy budget model to predict emergent life-history traits and trade-offs in soil bacteria. When applied to a plant microbiome system, the model accurately predicted distinct substrate-acquisition strategies that aligned with observations, uncovering resource-dependent trade-offs between microbial growth rate and efficiency. For instance, inherently slower-growing microorganisms, favoured by organic acid exudation at later plant growth stages, exhibited enhanced carbon use efficiency (yield) without sacrificing growth rate (power). This insight has implications for retaining plant root-derived carbon in soils and highlights the power of data-driven, trait-based approaches for improving microbial representation in biogeochemical models.

59 BASIC BIOLOGICAL SCIENCES↗

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant↗

The Metaproteomics Initiative: a coordinated approach for propelling the functional characterization of microbiomes

Through connecting genomic and metabolic information, metaproteomics is an essential approach for understanding how microbiomes function in space and time. The international metaproteomics community is delighted to announce the launch of the Metaproteomics Initiative (www.metaproteomics.org), the goal of which is to promote dissemination of metaproteomics fundamentals, advancements, and applications through collaborative networking in microbiome research. The Initiative aims to be the central information hub and open meeting place where newcomers and experts interact to communicate, standardize, and accelerate experimental and bioinformatic methodologies in this field. We invite the entire microbiome community to join and discuss potential synergies at the interfaces with other disciplines, and to collectively promote innovative approaches to gain deeper insights into microbiome functions and dynamics.

59 BASIC BIOLOGICAL SCIENCES↗

Budding yeasts in the subphylum Saccharomycotina Genome sequencing and assembly

Eukaryotic life depends on the functional elements encoded by both the nuclear genome and organellar genomes, such as those contained within the mitochondria. The content, size, and structure of the mitochondrial genome varies across organisms with potentially large implications for phenotypic variance and resulting evolutionary trajectories. Among yeasts in the subphylum Saccharomycotina, extensive differences have been observed in various species relative to the model yeast Saccharomyces cerevisiae, but mitochondrial genome sampling across many groups has been scarce, even as hundreds of nuclear genomes have become available. By extracting mitochondrial reads from existing short-read genome sequence datasets, we have greatly expanded both the number of available genomes and the coverage across sparsely sampled clades. Comparison of 353 yeast mitochondrial genomes revealed that, while size and GC content were fairly consistent across species, those in the genera Metschnikowia and Saccharomyces trended larger, while several species in the order Saccharomycetales exhibited lower GC content. Extreme examples for both size and GC content were scattered throughout the subphylum. All mitochondrial genomes shared a core set of protein-coding genes for Complexes III, IV, and V, but they varied in the presence or absence of mitochondrially-encoded canonical Complex I genes. We traced the loss of Complex I genes to a major event in the ancestor of the orders Saccharomycetales and Saccharomycodales, but we also observed several independent losses in the orders Phaffomycetales, Pichiales, and Dipodascales. In contrast to prior hypotheses based on smaller-scale datasets, comparison of evolutionary rates in protein-coding genes showed no bias towards elevated rates among aerobically fermenting (Crabtree/Warburg-positive) yeasts. Mitochondrial introns were widely distributed, but highly enriched in some groups. The majority of mitochondrial introns were poorly conserved within groups, but several were shared within groups, between groups, and even across taxonomic orders, which is consistent with horizontal gene transfer, likely involving homing endonucleases acting as selfish elements. As the number of available fungal nuclear genomes continues to expand, the methods described here to retrieve mitochondrial genome sequences from these datasets will prove invaluable to ensuring that studies of fungal mitochondrial genomes keep pace with their nuclear counterparts.

diversity↗

Characterization of Mammalian In Vivo Enhancers Using Mouse Transgenesis and CRISPR Genome Editing [Book Chapter]

Embryonic morphogenesis is strictly dependent on tight spatiotemporal control of developmental gene expression, which is typically achieved through the concerted activity of multiple enhancers driving cell type-specific expression of a target gene. Mammalian genomes are organized in topologically associated domains, providing a preferred environment and framework for interactions between transcriptional enhancers and gene promoters. While epigenomic profiling and three-dimensional chromatin conformation capture have significantly increased the accuracy of identifying enhancers, assessment of subregional enhancer activities via transgenic reporter assays in mice remains the gold standard for assigning enhancer activity in vivo. Once this activity is defined, the ideal method to explore the functional necessity of a transcriptional enhancer and its contribution to target gene dosage and morphological or physiological processes is deletion of the enhancer sequence from the mouse genome. Here we present detailed protocols for efficient introduction of enhancer-reporter transgenes and CRISPR-mediated genomic deletions into the mouse genome, including a step-by-step guide for pronuclear microinjection of fertilized mouse eggs. We provide instructions for the assembly and genomic integration of enhancer-reporter cassettes that have been used for validation of thousands of putative enhancer sequences accessible through the VISTA enhancer browser, including a recently published method for robust site-directed transgenesis at the H11 safe-harbor locus. Together, these methods enable rapid and large-scale assessment of enhancer activities and sequence variants in mice, which is essential to understand mammalian genome function and genetic diseases.

cis-regulatory elements↗

Biases in genome reconstruction from metagenomic data

Background Advances in sequencing, assembly, and assortment of contigs into species-specific bins has enabled the reconstruction of genomes from metagenomic data (MAGs). Though a powerful technique, it is difficult to determine whether assembly and binning techniques are accurate when applied to environmental metagenomes due to a lack of complete reference genome sequences against which to check the resulting MAGs. Methods We compared MAGs derived from an enrichment culture containing ~20 organisms to complete genome sequences of 10 organisms isolated from the enrichment culture. Factors commonly considered in binning software—nucleotide composition and sequence repetitiveness—were calculated for both the correctly binned and not-binned regions. This direct comparison revealed biases in sequence characteristics and gene content in the not-binned regions. Additionally, the composition of three public data sets representing MAGs reconstructed from the Tara Oceans metagenomic data was compared to a set of representative genomes available through NCBI RefSeq to verify that the biases identified were observable in more complex data sets and using three contemporary binning software packages. Results Repeat sequences were frequently not binned in the genome reconstruction processes, as were sequence regions with variant nucleotide composition. Genes encoded on the not-binned regions were strongly biased towards ribosomal RNAs, transfer RNAs, mobile element functions and genes of unknown function. Our results support genome reconstruction as a robust process and suggest that reconstructions determined to be >90% complete are likely to effectively represent organismal function; however, population-level genotypic heterogeneity in natural populations, such as uneven distribution of plasmids, can lead to incorrect inferences.

54 ENVIRONMENTAL SCIENCES↗

Eco-evolutionary strategies for relieving carbon limitation under salt stress differ across microbial clades

With the continuous expansion of saline soils under climate change, understanding the eco-evolutionary tradeoff between the microbial mitigation of carbon limitation and the maintenance of functional traits in saline soils represents a significant knowledge gap in predicting future soil health and ecological function. Through shotgun metagenomic sequencing of coastal soils along a salinity gradient, we show contrasting eco-evolutionary directions of soil bacteria and archaea that manifest in changes to genome size and the functional potential of the soil microbiome. In salt environments with high carbon requirements, bacteria exhibit reduced genome sizes associated with a depletion of metabolic genes, while archaea display larger genomes and enrichment of salt-resistance, metabolic, and carbon-acquisition genes. This suggests that bacteria conserve energy through genome streamlining when facing salt stress, while archaea invest in carbon-acquisition pathways to broaden their resource usage. These findings suggest divergent directions in eco-evolutionary adaptations to soil saline stress amongst microbial clades and serve as a foundation for understanding the response of soil microbiomes to escalating climate change.

54 ENVIRONMENTAL SCIENCES↗

Finding the missing pieces: filling gaps that impede the translation of omics data into models

High-throughput omics technologies such as DNA sequencing have made the sequencing and computational assembly of microbial genomes recovered from the environment relatively routine. Computational inference of the protein products encoded by these genomes, and the associated biochemical functions, should enable the accurate prediction and modeling of microbial metabolism, organismal interactions, and ecosystem processes. However, a lack of scalable, probabilistic protein annotation tools limits the full potential of modeling for understanding the metabolism and biogeochemical cycles of microbial communities. Our approach to improve inference of protein annotations and metabolic models relied on learning from and emulating expert manual curation, leveraging software engineering and data science best practices to scale up the throughput and accuracy of annotations and metabolic model construction, building software to objectively evaluate different annotation strategies, and more closely linking the protein annotation and metabolic model inference process. Outcomes of this research include several improved or new computational tools, including DRAM (Distilled and Refined Annotation of Metabolism) for annotating microbial genomes with protein function and metabolic traits, CAMPER (Curated Annotations for Microbial Polyphenol Enzymes and Reactions) for annotating key polyphenol metabolisms, EC-Bench for comprehensive and unbiased benchmarking of annotation tools, and several apps available via the DOE Systems Biology Knowledgebase (KBase) for building genome-scale metabolic models. We demonstrate that these tools allow us to scalably annotate and understand thousands of genomes for microbial communities from a variety of systems and test cases, including rivers, thawing permafrost, and gut microbiomes. All of these computational tools are available as open-source software, with most broadly and easily accessible to the scientific community via KBase apps.

59 BASIC BIOLOGICAL SCIENCES↗

Genome-wide characterization of the soybean DOMAIN OF UNKNOWN FUNCTION 679 membrane protein gene family highlights their potential involvement in growth and stress response

The DMP (DUF679 membrane proteins) family is a plant-specific gene family that encodes membrane proteins. The DMP family genes are suggested to be involved in various programmed cell death processes and gamete fusion during double fertilization in Arabidopsis. However, their functional relevance in other crops remains unknown. This study identified 14 genes from the DMP family in soybean (Glycine max) and characterized their physiochemical properties, subcellular location, gene structure, and promoter regions using bioinformatics tools. Additionally, their tissue-specific and stress-responsive expressions were analyzed using publicly available transcriptome data. Phylogenetic analysis of 198 DMPs from monocots and dicots revealed six clades, with clade-I encoding senescence-related AtDMP1/2 orthologues and clade-II including pollen-specific AtDMP8/9 orthologues. The largest clade, clade-III, predominantly included monocot DMPs, while monocot- and dicot-specific DMPs were assembled in clade-IV and clade-VI, respectively. Evolutionary analysis suggests that soybean GmDMPs underwent purifying selection during evolution. Using 68 transcriptome datasets, expression profiling revealed expression in diverse tissues and distinct responses to abiotic and biotic stresses. The genes Glyma.09G237500 and Glyma.18G098300 showed pistil-abundant expression by qPCR, suggesting they could be potential targets for female organ-mediated haploid induction. Furthermore, cis-acting regulatory elements primarily related to stress-, hormone-, and light-induced pathways regulate GmDMPs, which is consistent with their divergent expression and suggests involvement in growth and stress responses. Overall, our study provides a comprehensive report on the soybean GmDMP family and a framework for further biological functional analysis of DMP genes in soybean or other crops.

59 BASIC BIOLOGICAL SCIENCES↗

A glycan receptor kinase facilitates intracellular accommodation of arbuscular mycorrhiza and symbiotic rhizobia in the legume Lotus japonicus

Receptors that distinguish the multitude of microbes surrounding plants in the environment enable dynamic responses to the biotic and abiotic conditions encountered. In this study, we identify and characterise a glycan receptor kinase, EPR3a, closely related to the exopolysaccharide receptor EPR3. Epr3a is up-regulated in roots colonised by arbuscular mycorrhizal (AM) fungi and is able to bind glucans with a branching pattern characteristic of surface-exposed fungal glucans. Expression studies with cellular resolution show localised activation of the Epr3a promoter in cortical root cells containing arbuscules. Fungal infection and intracellular arbuscule formation are reduced in epr3a mutants. In vitro , the EPR3a ectodomain binds cell wall glucans in affinity gel electrophoresis assays. In microscale thermophoresis (MST) assays, rhizobial exopolysaccharide binding is detected with affinities comparable to those observed for EPR3, and both EPR3a and EPR3 bind a well-defined β-1,3/β-1,6 decasaccharide derived from exopolysaccharides of endophytic and pathogenic fungi. Both EPR3a and EPR3 function in the intracellular accommodation of microbes. However, contrasting expression patterns and divergent ligand affinities result in distinct functions in AM colonisation and rhizobial infection in Lotus japonicus . The presence of Epr3a and Epr3 genes in both eudicot and monocot plant genomes suggest a conserved function of these receptor kinases in glycan perception.

59 BASIC BIOLOGICAL SCIENCES↗