Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gene prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Comparative genome analyses suggest a hemibiotrophic lifestyle and virulence differences for the beech bark disease fungal pathogens Neonectria faginata and Neonectria coccinea

Abstract Neonectria faginata and Neonectria coccinea are the causal agents of the insect-fungus disease complex known as beech bark disease (BBD), known to cause mortality in beech forest stands in North America and Europe. These fungal species have been the focus of extensive ecological and disease management studies, yet less progress has been made toward generating genomic resources for both micro- and macro-evolutionary studies. Here, we report a 42.1 and 42.7 mb highly contiguous genome assemblies of N. faginata and N. coccinea, respectively, obtained using Illumina technology. These species share similar gene number counts (12,941 and 12,991) and percentages of predicted genes with assigned functional categories (64 and 65%). Approximately 32% of the predicted proteomes of both species are homologous to proteins involved in pathogenicity, yet N. coccinea shows a higher number of predicted mitogen-activated protein kinase genes, virulence determinants possibly contributing to differences in disease severity between N. faginata and N. coccinea. A wide range of genes encoding for carbohydrate-active enzymes capable of degradation of complex plant polysaccharides and a small number of predicted secretory effector proteins, secondary metabolite biosynthesis clusters and cytochrome oxidase P450 genes were also found. This arsenal of enzymes and effectors correlates with, and reflects, the hemibiotrophic lifestyle of these two fungal pathogens. Phylogenomic analysis and timetree estimations indicated that the N. faginata and N. coccinea species divergence may have occurred at ∼4.1 million years ago. Differences were also observed in the annotated mitochondrial genomes as they were found to be 81.7 kb (N. faginata) and 43.2 kb (N. coccinea) in size. The mitochondrial DNA expansion observed in N. faginata is attributed to the invasion of introns into diverse intra- and intergenic locations. These first draft genomes of N. faginata and N. coccinea serve as valuable tools to increase our understanding of basic genetics, evolutionary mechanisms and molecular physiology of these two nectriaceous plant pathogenic species.

Salgado-Salazar, Catalina↗

Cancer genomics predicts disease relapse and therapeutic response to neoadjuvant chemotherapy of hormone sensitive breast cancers

Several studies provide insight into the landscape of breast cancer genomics with the genomic characterization of tumors offering exceptional opportunities in defining therapies tailored to the patient’s specific need. However, translating genomic data into personalized treatment regimens has been hampered partly due to uncertainties in deviating from guideline based clinical protocols. Here we report a genomic approach to predict favorable outcome to treatment responses thus enabling personalized medicine in the selection of specific treatment regimens. The genomic data were divided into a training set of N = 835 cases and a validation set consisting of 1315 hormone sensitive, 634 triple negative breast cancer (TNBC) and 1365 breast cancer patients with information on neoadjuvant chemotherapy responses. Patients were selected by the following criteria: estrogen receptor (ER) status, lymph node invasion, recurrence free survival. The k-means classification algorithm delineated clusters with low- and high- expression of genes related to recurrence of disease; a multivariate Cox’s proportional hazard model defined recurrence risk for disease. Classifier genes were validated by Immunohistochemistry (IHC) using tissue microarray sections containing both normal and cancerous tissues and by evaluating findings deposited in the human protein atlas repository. Based on the leave-on-out cross validation procedure of 4 independent data sets we identified 51-genes associated with disease relapse and selected 10, i.e. TOP2A, AURKA, CKS2, CCNB2, CDK1 SLC19A1, E2F8, E2F1, PRC1, KIF11 for in depth validation. Expression of the mechanistically linked disease regulated genes significantly correlated with recurrence free survival among ER-positive and triple negative breast cancer patients and was independent of age, tumor size, histological grade and node status. Importantly, the classifier genes predicted pathological complete responses to neoadjuvant chemotherapy (P < 0.001) with high expression of these genes being associated with an improved therapeutic response toward two different anthracycline-taxane regimens; thus, highlighting the prospective for precision medicine. Our study demonstrates the potential of classifier genes to predict risk for disease relapse and treatment response to chemotherapies. The classifier genes enable rational selection of patients who benefit best from a given chemotherapy thus providing the best possible care. The findings encourage independent clinical validation.

59 BASIC BIOLOGICAL SCIENCES↗

Draft genome of multiple resistance donor plant Sinapis alba: An insight into SSRs, annotations and phylogenetics

Sinapis alba is a wild member of the Brassicaceae family reported to possess genetic resistance against major biotic and abiotic stresses of oilseed brassicas. However, the resistance nature of S. alba was not exploited generously due to the unavailability of usable genome sequences in public databases. Therefore, the present study was conducted to assemble the first draft genome from raw whole genome shotgun sequences with annotation and develop simple sequence repeat markers for molecular genetics and marker-assisted breeding. Results The raw genome sequences had 96x coverage on the Illumina platform with 170 Gbp data. The developed assembly by SOAPdenovo2 has ~459 Mbp genome size covered in 403,423 contigs with an average size of 1138.04 bp. The assembly was BLASTX with Arabidopsis thaliana which showed 32.9% positive hits between both plants. The top hit species distribution analysis showed the highest similarity with A. thaliana. A total of 809,597 GO level annotations were recorded after BLASTX results, and 34,012 sequences were annotated with different enzyme codes grouped under seven classes. The gene prediction tool AUGUSTUS identified 113,107 probable genes with an average size of 684 bp. The biochemical pathway annotation assigned 16,119 potential genes to 152 KEGG maps and 1751 enzyme codes. The development of potential SSRs from the de-novo assembly yielded 70731 unique primer pairs. Out of 159 randomly selected SSR markers for validation, 149 successfully amplified in S. alba. However, 10 SSR markers did not amplify during the validation experiment. Conclusion The annotated genome assembly with a large number of SSRs was developed in the present study. To the best of our knowledge, this is the first report of S. alba genome assembly development, annotation, and SSRs mining to date. The data presented here will be a very important resource for future crop improvement programs, especially for resistant breeding.

59 BASIC BIOLOGICAL SCIENCES↗

Reekeekee- and roodoodooviruses, two different Microviridae clades constituted by the smallest DNA phages

Small circular single-stranded DNA viruses of the Microviridae family are both prevalent and diverse in all ecosystems. They usually harbor a genome between 4.3 and 6.3 kb, with a microvirus recently isolated from a marine Alphaproteobacteria being the smallest known genome of a DNA phage (4.248 kb). A subfamily, Amoyvirinae, has been proposed to classify this virus and other related small Alphaproteobacteria-infecting phages. Here, we report the discovery, in meta-omics data sets from various aquatic ecosystems, of sixteen complete microvirus genomes significantly smaller (2.991–3.692 kb) than known ones. Phylogenetic analysis reveals that these sixteen genomes represent two related, yet distinct and diverse, novel groups of microviruses—amoyviruses being their closest known relatives. We propose that these small microviruses are members of two tentatively named subfamilies Reekeekeevirinae and Roodoodoovirinae. As known microvirus genomes encode many overlapping and overprinted genes that are not identified by gene prediction software, we developed a new methodology to identify all genes based on protein conservation, amino acid composition, and selection pressure estimations. Surprisingly, only four to five genes could be identified per genome, with the number of overprinted genes lower than that in phiX174. These small genomes thus tend to have both a lower number of genes and a shorter length for each gene, leaving no place for variable gene regions that could harbor overprinted genes. Even more surprisingly, these two Microviridae groups had specific and different gene content, and major differences in their conserved protein sequences, highlighting that these two related groups of small genome microviruses use very different strategies to fulfill their lifecycle with such a small number of genes. The discovery of these genomes and the detailed prediction and annotation of their genome content expand our understanding of ssDNA phages in nature and are further evidence that these viruses have explored a wide range of possibilities during their long evolution.

59 BASIC BIOLOGICAL SCIENCES↗

The Crown Pearl: a draft genome assembly of the European freshwater pearl mussel Margaritifera margaritifera (Linnaeus, 1758)

Abstract Since historical times, the inherent human fascination with pearls turned the freshwater pearl mussel Margaritifera margaritifera (Linnaeus, 1758) into a highly valuable cultural and economic resource. Although pearl harvesting in M. margaritifera is nowadays residual, other human threats have aggravated the species conservation status, especially in Europe. This mussel presents a myriad of rare biological features, e.g. high longevity coupled with low senescence and Doubly Uniparental Inheritance of mitochondrial DNA, for which the underlying molecular mechanisms are poorly known. Here, the first draft genome assembly of M. margaritifera was produced using a combination of Illumina Paired-end and Mate-pair approaches. The genome assembly was 2.4 Gb long, possessing 105,185 scaffolds and a scaffold N50 length of 288,726 bp. The ab initio gene prediction allowed the identification of 35,119 protein-coding genes. This genome represents an essential resource for studying this species’ unique biological and evolutionary features and ultimately will help to develop new tools to promote its conservation.

Gomes-dos-Santos, André (ORCID:0000000199734861)↗

Genomic Analysis of Diverse Members of the Fungal Genus Monosporascus Reveals Novel Lineages, Unique Genome Content and a Potential Bacterial Associate

The genus Monosporascus represents an enigmatic group of fungi important in agriculture and widely distributed in natural arid ecosystems. Of the nine described species, two (M. cannonballus and M. eutypoides) are important pathogens on the roots of members of Cucurbitaceae in agricultural settings. The remaining seven species are capable of colonizing roots from a diverse host range without causing obvious disease symptoms. Recent molecular and culture studies have shown that members of the genus are nearly ubiquitous as root endophytes in arid environments of the Southwestern United States. Isolates have been obtained from apparently healthy roots of grasses, shrubs and herbaceous plants located in central New Mexico and other regions of the Southwest. Phylogenetic and genomic analyses reveal substantial diversity in these isolates. The New Mexico isolates include close relatives of M. cannonballus and M. ibericus, as well as isolates that represent previously unrecognized lineages. To explore evolutionary relationships within the genus and gain insights into potential ecological functions, we sequenced and assembled the genomes of three M. cannonballus isolates, one M. ibericus isolate, and six diverse New Mexico isolates. The assembled genomes were significantly larger than what is typical for the Sordariomycetes despite having predicted gene numbers similar to other members of the class. Differences in predicted genome content and organization were observed between endophytic and pathogenic lineages of Monosporascus. Several Monosporascus isolates appear to form associations with members of the bacterial genus Ralstonia (Burkholdariaceae).

59 BASIC BIOLOGICAL SCIENCES↗

A Sesquiterpene Synthase from the Endophytic Fungus Serendipita indica Catalyzes Formation of Viridiflorol

Interactions between plant-associated fungi and their hosts are characterized by a continuous crosstalk of chemical molecules. Specialized metabolites are often produced during these associations and play important roles in the symbiosis between the plant and the fungus, as well as in the establishment of additional interactions between the symbionts and other organisms present in the niche. Serendipita indica, a root endophytic fungus from the phylum Basidiomycota, is able to colonize a wide range of plant species, conferring many benefits to its hosts. The genome of S. indica possesses only few genes predicted to be involved in specialized metabolite biosynthesis, including a putative terpenoid synthase gene (SiTPS). In our experimental setup, SiTPS expression was upregulated when the fungus colonized tomato roots compared to its expression in fungal biomass growing on synthetic medium. Heterologous expression of SiTPS in Escherichia coli showed that the produced protein catalyzes the synthesis of a few sesquiterpenoids, with the alcohol viridiflorol being the main product. To investigate the role of SiTPS in the plant-endophyte interaction, an SiTPS-over-expressing mutant line was created and assessed for its ability to colonize tomato roots. Although overexpression of SiTPS did not lead to improved fungal colonization ability, an in vitro growth-inhibition assay showed that viridiflorol has antifungal properties. Addition of viridiflorol to the culture medium inhibited the germination of spores from a phytopathogenic fungus, indicating that SiTPS and its products could provide S. indica with a competitive advantage over other plant-associated fungi during root colonization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Differential ability of three bee species to move genes via pollen

Since the release of genetically engineered (GE) crops, there has been increased concern about the introduction of GE genes into non-GE fields of a crop and their spread to feral or wild cross-compatible relatives. More recently, attention has been given to the differential impact of distinct pollinators on gene flow, with the goal of developing isolation distances associated with specific managed pollinators. To examine the differential impact of bee species on gene movement, we quantified the relationship between the probability of getting a GE seed in a pod, and the order in which a flower was visited, or the cumulative distance traveled by a bee in a foraging bout. We refer to these relationships as ‘seed curves’ and compare these seeds curves among three bee species. The experiments used Medicago sativa L. plants carrying three copies of the glyphosate resistance (GR) allele as pollen donors (M. sativa is a tetraploid), such that each pollen grain carried the GR allele, and conventional plants as pollen recipients. Different foraging metrics, including the number of GR seeds produced over a foraging bout, were also quantified and contrasted among bee species. The lowest number of GR seeds set per foraging bout, and the GR seeds set at the shortest distances, were produced following leafcutting bee visits. In contrast, GR seeds were found at the longest distances following bumble bee visits. Values for honey bees were intermediate. The ranking of bee species based on seed curves correlated well with field-based gene flow estimates. Thus, differential seed curves of bee species, which describe patterns of seed production within foraging bouts, translated into distinct abilities of bee species to move genes at a landscape level. Bee behavior at a local scale (foraging bout) helps predict gene flow and the spread of GE genes at the landscape scale.

59 BASIC BIOLOGICAL SCIENCES↗

RatXcan: A framework for cross-species integration of genome-wide association and gene expression data

Genome-wide association studies (GWAS) have implicated specific alleles and genes as risk factors for numerous complex traits. However, translating GWAS results into biologically and therapeutically meaningful discoveries remains extremely challenging. Most GWAS results identify noncoding regions of the genome, suggesting that differences in gene regulation are the major driver of trait variability. To better integrate GWAS results with gene regulatory polymorphisms, we previously developed PrediXcan (also known as “transcriptome-wide association studies” orTWAS), which maps SNPs to predicted gene expression using GWAS data. In this study, we developed RatXcan, a framework that extends this methodology to outbred heterogeneous stock (HS) rats. RatXcan accounts for the close familial relationships among HS rats by modeling the relatedness with a random effect that encodes the genetic relatedness. RatXcan also corrects for polygenic-driven inflation because of the equivalence between a relatedness random effect and the infinitesimal polygenic model. To develop RatXcan, we trained transcript predictors for 8,934 genes using reference genotype and expression data from five rat brain regions. We found that the cis genetic architecture of gene expression in both rats and humans was sparse and similar across brain tissues. We tested the association between predicted expression in rats and two example traits (body length and BMI) using phenotype and genotype data from 5,401 densely genotyped HS rats and identified a significant enrichment between the genes associated with rat and human body length and BMI. Thus, RatXcan represents a valuable tool for identifying the relationship between gene expression and phenotypes across species and paves the way to explore shared biological mechanisms of complex traits.

Genetics & Heredity↗

Trichoderma harzianum transcriptome in response to the nematode Pratylenchus brachyurus

The root-lesion nematode Pratylenchus brachyurus causes extensive damage in several crops of economic importance. Fungi of Trichoderma genus have been highlighted as biopesticide agents in the control of several plant diseases. Although it is already widely used in agriculture, there are few studies, especially at the molecular level, that evaluate T. harzianum in the control of P. brachyurus . The aim of the present study was to investigate how interaction with the nematode P. brachyurus influences gene expression of T. harzianum , by using RNA-Seq analysis. Of the 13,932 predicted genes in the T. harzianum genome, 2,922 (21%) were differentially expressed in the presence of P. brachyurus , in relation to the absence of the nematode. Among the differentially expressed genes, we found genes encoding Carbohydrate Active EnZymes (CAZy), MEROPS peptidases, and proteins related to secondary metabolite synthesis. 118 pathways were identified as related to the biosynthesis of secondary metabolites. Additionally, the analysis identified 136 metabolic pathways related to these differentially expressed genes, among which we highlight: aminobenzoate degradation, xenobiotic metabolism by cytochrome P450, and sesquiterpenoid and triterpenoid biosynthesis. Our results contribute to a better understanding of the response of T. harzianum to the nematode P. bracyurus , the potential for biocontrol by this fungus.

59 BASIC BIOLOGICAL SCIENCES↗

Long-Term Effects of Very Low Dose Particle Radiation on Gene Expression in the Heart: Degenerative Disease Risks

Compared to low doses of gamma irradiation (γ-IR), high-charge-and-energy (HZE) particle IR may have different biological response thresholds in cardiac tissue at lower doses, and these effects may be IR type and dose dependent. Three- to four-month-old female CB6F1/Hsd mice were exposed once to one of four different doses of the following types of radiation: γ-IR 137Cs (40-160 cGy, 0.662 MeV), 14Si-IR (4-32 cGy, 260 MeV/n), or 22Ti-IR (3-26 cGy, 1 GeV/n). At 16 months post-exposure, animals were sacrificed and hearts were harvested and archived as part of the NASA Space Radiation Tissue Sharing Forum. These heart tissue samples were used in our study for RNA isolation and microarray hybridization. Functional annotation of twofold up/down differentially expressed genes (DEGs) and bioinformatics analyses revealed the following: (i) there were no clear lower IR thresholds for HZE- or γ-IR; (ii) there were 12 common DEGs across all 3 IR types; (iii) these 12 overlapping genes predicted various degrees of cardiovascular, pulmonary, and metabolic diseases, cancer, and aging; and (iv) these 12 genes revealed an exclusive non-linear DEG pattern in 14Si- and 22Ti-IR-exposed hearts, whereas two-thirds of γ-IR-exposed hearts revealed a linear pattern of DEGs. Thus, our study may provide experimental evidence of excess relative risk (ERR) quantification of low/very low doses of full-body space-type IR-associated degenerative disease development.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Genes encoding calmodulin-binding proteins in the Arabidopsis genome

Analysis of the recently completed Arabidopsis genome sequence indicates that approximately 31% of the predicted genes could not be assigned to functional categories, as they do not show any sequence similarity with proteins of known function from other organisms. Calmodulin (CaM), a ubiquitous and multifunctional Ca(2+) sensor, interacts with a wide variety of cellular proteins and modulates their activity/function in regulating diverse cellular processes. However, the primary amino acid sequence of the CaM-binding domain in different CaM-binding proteins (CBPs) is not conserved. One way to identify most of the CBPs in the Arabidopsis genome is by protein-protein interaction-based screening of expression libraries with CaM. Here, using a mixture of radiolabeled CaM isoforms from Arabidopsis, we screened several expression libraries prepared from flower meristem, seedlings, or tissues treated with hormones, an elicitor, or a pathogen. Sequence analysis of 77 positive clones that interact with CaM in a Ca(2+)-dependent manner revealed 20 CBPs, including 14 previously unknown CBPs. In addition, by searching the Arabidopsis genome sequence with the newly identified and known plant or animal CBPs, we identified a total of 27 CBPs. Among these, 16 CBPs are represented by families with 2-20 members in each family. Gene expression analysis revealed that CBPs and CBP paralogs are expressed differentially. Our data suggest that Arabidopsis has a large number of CBPs including several plant-specific ones. Although CaM is highly conserved between plants and animals, only a few CBPs are common to both plants and animals. Analysis of Arabidopsis CBPs revealed the presence of a variety of interesting domains. Our analyses identified several hypothetical proteins in the Arabidopsis genome as CaM targets, suggesting their involvement in Ca(2+)-mediated signaling networks.

NASA Discipline Plant Biology↗

MAPLE v.1.0

SAND2025-00659O MAPLE is a software tool that uses epigenomic data to predict gene expression. MAPLE uses a set of epigenomic modifications to determine the effect on gene expression in a specific subset of species. The algorithm can be trained on additional species and epigenomic modifications, enhancing its predictive capabilities. EAGLE employs a hybrid neural network architecture, featuring a convolutional front-end and a multi-head attention layer, to process pre-processed signal data as input. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.

Davis IV, Warren↗

Skeletal Micro-RNA Responses to Simulated Weightlessness

Astronauts lose bone structure during long-duration spaceflight. These changes are due, in part, to insufficient bone formation by the osteoblast cells. Little is known about the role that small (approximately 22 nucleotides), non-coding micro-RNAs (miRNAs) play in the osteoblast response to microgravity. We hypothesize that osteoblast-lineage cells alter their miRNA status during microgravity exposure, contributing to impaired bone formation during weightlessness. To simulate weightlessness, female mice (C57BL/6, Charles River, 10 weeks of age, n = 7) were hindlimb unloaded up to 12 days. Age-matched and normally ambulating mice served as controls (n=7). To assess the expression of miRNAs in skeletal tissue, the tibia was collected ex vivo and cleaned of soft-tissue and marrow. Total RNA was collected from tibial bone and relative abundance was measured for miRNAs of interest using quantitative real time PCR array looking at 372 unique and well-characterized mature miRNAs using the delta-delta Ct method. Transcripts of interest were normalized to an average of 6 reference RNAs. Preliminary results show that hindlimb unloading decreased the expression of 14 miRNAs to less than 0.5 times that of the control levels and increased the expression of 5 miRNAs relative to the control mice between 1.2-1.5-fold (p less than 0.05, respectively). Using the miRSystem we assessed overlapping target genes predicted to be regulated by multiple members of the 19 differentially expressed miRNAs as well as in silico predicted targets of our individual miRNAs. Our miRsystem results indicated that a number of our differentially expressed miRNAs were regulators of genes related to the Wnt-Beta Catenin pathway-a known regulator of bone health-and, interestingly, the estrogen-mediated cell-cycle regulation pathway, which may indicate that simulated weightlessness modulated systemic hormonal levels or hormonal transduction that additionally contributed to bone loss. We plan to follow up these findings by measuring gene expression of miRNA-regulated genes within these two pathways with the aim of furthering our understanding of the function of miRNAs in the skeletal response to spaceflight.

musculoskeletal system↗

Regulation of Bone Formation During Disuse by miRNA

Astronauts lose bone structure during long-duration spaceflight. These changes are due, in part, to insufficient bone formation by the osteoblast cells. Little is known about the role that small (approximately 22 nucleotide), non-coding micro-RNAs (miRNAs) play in the osteoblast response to microgravity. We hypothesize that osteoblast-lineage cells alter their miRNA status during microgravity exposure, contributing to impaired bone formation during weightlessness. To simulate weightlessness, female mice (C57BL/6, Charles River, 10 weeks of age, n = 6) were hindlimb unloaded for 12 days. Age-matched and normally ambulating mice served as controls (n=6). To assess the expression of miRNAs in skeletal tissue, the right and left tibia of the mice were collected ex vivo and cleaned of soft-tissue and marrow. Total RNA was collected from tibial bone and relative abundance was measured for miRNAs of interest using quantitative real time PCR array looking at 372 unique and well-characterized mature miRNAs using the delta-delta Ct method. Transcripts of interest were normalized to an average of 6 reference RNAs. Preliminary results show that hindlimb unloading decreased the expression of 14 miRNAs to less than 1.4-2.9X control levels and increased the expression of 5 miRNAs relative to the control mice greater than 1-2-1.5X (p less than 0.05, respectively). Using the miRSystem we assessed overlapping target genes predicted to be regulated by multiple members of the 19 differentially expressed miRNAs as well as in silico predicted targets of our individual miRNAs. Our miRSystem results indicated that a number of our differentially expressed miRNAs were regulators of genes related to the Wnt-Beta Catenin pathway-a known regulator of bone health-and, interestingly, the estrogen-mediated cell-cycle regulation pathway, which may indicate that simulated weightlessness induced systemic hormonal changes that contributed to bone loss. We plan to follow up these findings by measuring gene expression of miRNA-regulated genes within these two pathways with the aim of furthering our understanding of the function of miRNAs in the skeletal response to spaceflight.

bone↗

A cell type-aware framework for nominating non-coding variants in Mendelian regulatory disorders

Abstract Unsolved Mendelian cases often lack obvious pathogenic coding variants, suggesting potential non-coding etiologies. Here, we present a single cell multi-omic framework integrating embryonic mouse chromatin accessibility, histone modification, and gene expression assays to discover cranial motor neuron (cMN)cis-regulatory elements and subsequently nominate candidate non-coding variants in the congenital cranial dysinnervation disorders (CCDDs), a set of Mendelian disorders altering cMN development. We generate single cell epigenomic profiles for ~86,000 cMNs and related cell types, identifying ~250,000 accessible regulatory elements with cognate gene predictions for ~145,000 putative enhancers. We evaluate enhancer activity for 59 elements using an in vivo transgenic assay and validate 44 (75%), demonstrating that single cell accessibility can be a strong predictor of enhancer activity. Applying our cMN atlas to 899 whole genome sequences from 270 genetically unsolved CCDD pedigrees, we achieve significant reduction in our variant search space and nominate candidate variants predicted to regulate known CCDD disease genesMAFB, PHOX2A, CHN1, andEBF3– as well as candidates in recurrently mutated enhancers through peak- and gene-centric allelic aggregation. This work delivers non-coding variant discoveries of relevance to CCDDs and a generalizable framework for nominating non-coding variants of potentially high functional impact in other Mendelian disorders.

Science & Technology - Other Topics↗

The state of algal genome quality and diversity

The genomic era of biology has created unprecedented opportunity to study how life works, including understanding evolutionary principles, bioprospecting for novel antibiotics, and genetic manipulation of bioeconomically-relevant species. While genome sequencing and genomic analysis was previously restricted to large-scale projects and consortia, sequencing has become democratized. However, as genomics has become commonplace, the cataloging of sequenced organisms has become challenging, and standardized practices for sequencing, assembly, and annotation have not been adopted. This is equally true in research fields such as algal biology, despite the growing importance of algae in a bio-based economy. Here, in this study, we provide a comprehensive review of the state of eukaryotic algal genomics, explore the quality of algal genome assemblies, and identify the biases and gaps in the current species distribution to inform the development of future genome projects. Overall, we find a trend of declining quality of genomic resources, including a reduction in assembly quality, gene annotation quality, and genome completeness. Potential solutions to improve genome quality include widespread utilization of long read and scaffolding technologies, implementation of standards for assembly quality, evidence-based gene annotation and requisite publication, and support of continued development of gene prediction and genome assessment software.

59 BASIC BIOLOGICAL SCIENCES↗