Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Genetics & heredity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Binding profiles for 961 Drosophila and C. elegans transcription factors reveal tissue-specific regulatory relationships

A catalog of transcription factor (TF) binding sites in the genome is critical for deciphering regulatory relationships. Here, we present the culmination of the efforts of the modENCODE (model organism Encyclopedia of DNA Elements) and modERN (model organism Encyclopedia of Regulatory Networks) consortia to systematically assay TF binding events in vivo in two major model organisms,Drosophila melanogaster(fly) andCaenorhabditis elegans(worm). These data sets comprise 605 TFs identifying 3.6 M sites in the fly and 356 TFs identifying 0.9 M sites in the worm, and represent the majority of the regulatory space in each genome. We demonstrate that TFs associate with chromatin in clusters termed “metapeaks,” that larger metapeaks have characteristics of high-occupancy target (HOT) regions, and that the importance of consensus sequence motifs bound by TFs depends on metapeak size and complexity. Combining ChIP-seq data with single-cell RNA-seq data in a machine-learning model identifies TFs with a prominent role in promoting target gene expression in specific cell types, even differentiating between parent–daughter cells during embryogenesis. These data are a rich resource for the community that should fuel and guide future investigations into TF function. To facilitate data accessibility and utility, all strains expressing green fluorescent protein (GFP)-tagged TFs are available at the stock centers for each organism. The chromatin immunoprecipitation sequencing data are available through the ENCODE Data Coordinating Center, GEO, and through a direct interface that provides rapid access to processed data sets and summary analyses, as well as widgets to probe the cell-type-specific TF–target relationships.

Biochemistry & Molecular Biology↗

Streamlined spatial and environmental expression signatures characterize the minimalist duckweed Wolffia australiana

Single-cell genomics permits a new resolution in the examination of molecular and cellular dynamics, allowing global, parallel assessments of cell types and cellular behaviors through development and in response to environmental circumstances, such as interaction with water and the light–dark cycle of the Earth. Here, we leverage the smallest, and possibly most structurally reduced, plant, the semiaquaticWolffia australiana, to understand dynamics of cell expression in these contexts at the whole-plant level. We examined single-cell-resolution RNA-sequencing data and foundWolffiacells divide into four principal clusters representing the above- and below-water-situated parenchyma and epidermis. Although these tissues share transcriptomic similarity with model plants, they display distinct adaptations thatWolffiahas made for the aquatic environment. Within this broad classification, discrete subspecializations are evident, with select cells showing unique transcriptomic signatures associated with developmental maturation and specialized physiologies. Assessing this simplified biological system temporally at two key time-of-day (TOD) transitions, we identify additional TOD-responsive genes previously overlooked in whole-plant transcriptomic approaches and demonstrate that the core circadian clock machinery and its downstream responses can vary in cell-specific manners, even in this simplified system. Distinctions between cell types and their responses to submergence and/or TOD are driven by expression changes of unexpectedly few genes, characterizingWolffiaas a highly streamlined organism with the majority of genes dedicated to fundamental cellular processes.Wolffiaprovides a unique opportunity to apply reductionist biology to elucidate signaling functions at the organismal level, for which this work provides a powerful resource.

Biochemistry & Molecular Biology↗

Complete genomes of Asgard archaea reveal diverse integrated and mobile genetic elements

Asgard archaea are of great interest as the progenitors of Eukaryotes, but little is known about the mobile genetic elements (MGEs) that may shape their ongoing evolution. Here, we describe MGEs that replicate in Atabeyarchaeia, a wetland Asgard archaea lineage represented by two complete genomes. We used soil depth–resolved population metagenomic data sets to track 18 MGEs for which genome structures were defined and precise chromosome integration sites could be identified for confident host linkage. Additionally, we identified a complete 20.67 kbp circular plasmid and two family-level groups of viruses linked to Atabeyarchaeia, via CRISPR spacer targeting. Closely related 40 kbp viruses possess a hypervariable genomic region encoding combinations of specific genes for small cysteine-rich proteins structurally similar to restriction-homing endonucleases. One 10.9 kbp integrative conjugative element (ICE) integrates genomically into theAtabeyarchaeum deiterrae-1chromosome and has a 2.5 kbp circularizable element integrated within it. The 10.9 kbp ICE encodes an expressed Type IIG restriction-modification system with a sequence specificity matching an active methylation motif identified by Pacific Biosciences (PacBio) high-accuracy long-read (HiFi) metagenomic sequencing. Restriction-modification of Atabeyarchaeia differs from that of another coexisting Asgard archaea, Freyarchaeia, which has few identified MGEs but possesses diverse defense mechanisms, including DISARM and Hachiman, not found in Atabeyarchaeia. Overall, defense systems and methylation mechanisms of Asgard archaea likely modulate their interactions with MGEs, and integration/excision and copy number variation of MGEs in turn enable host genetic versatility.

Biochemistry & Molecular Biology↗

Quantile-Dependent Expressivity of Serum Uric Acid Concentrations

“Quantile-dependent expressivity” occurs when the effect size of a genetic variant depends upon whether the phenotype (e.g., serum uric acid) is high or low relative to its distribution. Analyses were performed to test whether serum uric acid heritability is quantile-specific and whether this could explain some reported gene-environment interactions. Methods. Serum uric acid concentrations were analyzed from 2151 sibships and 12,068 offspring-parent pairs from the Framingham Heart Study. Quantile-specific heritability from offspring-parent regression slopes (β OP , h 2 = 2β OP /(1 + r spouse )) and full-sib regression slopes (β FS ,h 2 = {(1 + 8r spouse β FS ) 0.5 - 1}/(2r spouse ) was robustly estimated by quantile regression with nonparametric significance assigned from 1000 bootstrap samples. Results. Quantile-specific h 2 (±SE) increased with increasing percentiles of the offspring’s sex- and age-adjusted uric acid distribution when estimated from β OP (P trend = 0.001): 0.34 ± 0.03 at the 10th, 0.36 ± 0.03 at the 25th, 0.41 ± 0.03 at the 50th, 0.46 ± 0.04 at the 75th, and 0.49 ± 0.05 at the 90th percentile and when estimated from β FS (P trend = 0.006). This is consistent with the larger genetic effect size of (1) the SLC2A9 rs11722228 polymorphism in gout patients vs. controls, (2) the ABCG2 rs2231142 polymorphism in men vs. women, (3) the SLC2A9 rs13113918 polymorphism in obese patients prior to bariatric surgery vs. two-year postsurgery following 29 kg weight loss, (4) the ABCG2 rs6855911 polymorphism in obese vs. nonobese women, and (5) the LRP2 rs2544390 polymorphism in heavier drinkers vs. abstainers. Quantile-dependent expressivity may also explain the larger genetic effect size of an SLC2A9/PKD2/ABCG2 haplotype for high vs. low intakes of alcohol, chicken, or processed meats. Conclusions. Heritability of serum uric acid concentrations is quantile-specific.

59 BASIC BIOLOGICAL SCIENCES↗

Role of diversity-generating retroelements for regulatory pathway tuning in cyanobacteria

Abstract Background Cyanobacteria maintain extensive repertoires of regulatory genes that are vital for adaptation to environmental stress. Some cyanobacterial genomes have been noted to encode diversity-generating retroelements (DGRs), which promote protein hypervariation through localized retrohoming and codon rewriting in target genes. Past research has shown DGRs to mainly diversify proteins involved in cell-cell attachment or viral-host attachment within viral, bacterial, and archaeal lineages. However, these elements may be critical in driving variation for proteins involved in other core cellular processes. Results Members of 31 cyanobacterial genera encode at least one DGR, and together, their retroelements form a monophyletic clade of closely-related reverse transcriptases. This class of retroelements diversifies target proteins with unique domain architectures: modular ligand-binding domains often paired with a second domain that is linked to signal response or regulation. Comparative analysis indicates recent intragenomic duplication of DGR targets as paralogs, but also apparent intergenomic exchange of DGR components. The prevalence of DGRs and the paralogs of their targets is disproportionately high among colonial and filamentous strains of cyanobacteria. Conclusion We find that colonial and filamentous cyanobacteria have recruited DGRs to optimize a ligand-binding module for apparent function in signal response or regulation. These represent a unique class of hypervariable proteins, which might offer cyanobacteria a form of plasticity to adapt to environmental stress. This analysis supports the hypothesis that DGR-driven mutation modulates signaling and regulatory networks in cyanobacteria, suggestive of a new framework for the utility of localized genetic hypervariation.

59 BASIC BIOLOGICAL SCIENCES↗

Genome-based analysis for the bioactive potential of Streptomyces yeochonensis CN732, an acidophilic filamentous soil actinobacterium

Background: Acidophilic members of the genus Streptomyces can be a good source for novel secondary metabolites and degradative enzymes of biopolymers. In this study, a genome-based approach on Streptomyces yeochonensis CN732, a representative neutrotolerant acidophilic streptomycete, was employed to examine the biosynthetic as well as enzymatic potential, and also presence of any genetic tools for adaptation in acidic environment. Results: A high quality draft genome (7.8Mb) of S. yeochonensis CN732 was obtained with a G+C content of 73.53% and 6549 protein coding genes. The in silico analysis predicted presence of multiple biosynthetic gene clusters (BGCs), which showed similarity with those for antimicrobial, anticancer or antiparasitic compounds. However, the low levels of similarity with known BGCs for most cases suggested novelty of the metabolites from those predicted gene clusters. The production of various novel metabolites was also confirmed from the combined high performance liquid chromatography-mass spectrometry analysis. Through comparative genome analysis with related Streptomyces species, genes specific to strain CN732 and also those specific to neutrotolerant acidophilic species could be identified, which showed that genes for metabolism in diverse environment were enriched among acidophilic species. In addition, the presence of strain specific genes for carbohydrate active enzymes (CAZyme) along with many other singletons indicated uniqueness of the genetic makeup of strain CN732. The presence of cysteine transpeptidases (sortases) among the BGCs was also observed from this study, which implies their putative roles in the biosynthesis of secondary metabolites. Conclusions: This study highlights the bioactive potential of strain CN732, an acidophilic streptomycete with regard to secondary metabolite production and biodegradation potential using genomics based approach. The comparative genome analysis revealed genes specific to CN732 and also those among acidophilic species, which could give some insights into the adaptation of microbial life in acidic environment.

59 BASIC BIOLOGICAL SCIENCES↗

De novo transcriptome in roots of switchgrass ( Panicum virgatum L. ) reveals gene expression dynamic and act network under alkaline salt stress

Background: Soil salinization is a major limiting factor for crop cultivation. Switchgrass is a perennial rhizomatous bunchgrass that is considered an ideal plant for marginal lands, including sites with saline soil. Here we investigated the physiological responses and transcriptome changes in the roots of Alamo (alkaline-tolerant genotype) and AM314/MS-155 (alkaline-sensitive genotype) under alkaline salt stress. Results: Alkaline salt stress significantly affected the membrane, osmotic adjustment and antioxidant systems in switchgrass roots, and the ASTTI values between Alamo and AM-314/MS-155 were divergent at different time points. A total of 108,319 unigenes were obtained after reassembly, including 73,636 unigenes in AM-314/MS-155 and 65,492 unigenes in Alamo. A total of 10,219 DEGs were identified, and the number of upregulated genes in Alamo was much greater than that in AM-314/MS-155 in both the early and late stages of alkaline salt stress. The DEGs in AM-314/MS-155 were mainly concentrated in the early stage, while Alamo showed greater advantages in the late stage. These DEGs were mainly enriched in plant-pathogen interactions, ubiquitin-mediated proteolysis and glycolysis/gluconeogenesis pathways. We characterized 1480 TF genes into 64 TF families, and the most abundant TF family was the C2H2 family, followed by the bZIP and bHLH families. A total of 1718 PKs were predicted, including CaMK, CDPK, MAPK and RLK. WGCNA revealed that the DEGs in the blue, brown, dark magenta and light steel blue 1 modules were associated with the physiological changes in roots of switchgrass under alkaline salt stress. The consistency between the qRT-PCR and RNA-Seq results confirmed the reliability of the RNA-seq sequencing data. A molecular regulatory network of the switchgrass response to alkaline salt stress was preliminarily constructed on the basis of transcriptional regulation and functional genes. Conclusions: Alkaline salt tolerance of switchgrass may be achieved by the regulation of ion homeostasis, transport proteins, detoxification, heat shock proteins, dehydration and sugar metabolism. These findings provide a comprehensive analysis of gene expression dynamic and act network induced by alkaline salt stress in two switchgrass genotypes and contribute to the understanding of the alkaline salt tolerance mechanism of switchgrass and the improvement of switchgrass germplasm.

59 BASIC BIOLOGICAL SCIENCES↗

HiFiAdapterFilt, a memory efficient read processing pipeline, prevents occurrence of adapter sequence in PacBio HiFi reads and their negative impacts on genome assembly

Abstract Background Pacific Biosciences HiFi read technology is currently the industry standard for high accuracy long-read sequencing that has been widely adopted by large sequencing and assembly initiatives for generation of de novo assemblies in non-model organisms. Though adapter contamination filtering is routine in traditional short-read analysis pipelines, it has not been widely adopted for HiFi workflows. Results Analysis of 55 publicly available HiFi datasets revealed that a read-sanitation step to remove sequence artifacts derived from PacBio library preparation from read pools is necessary as adapter sequences can be erroneously integrated into assemblies. Conclusions Here we describe the nature of adapter contaminated reads, their consequences in assembly, and present HiFiAdapterFilt, a simple and memory efficient solution for removing adapter contaminated reads prior to assembly.

59 BASIC BIOLOGICAL SCIENCES↗

Relationship and distribution of Salmonella enterica serovar I 4,[5],12:i:- strain sequences in the NCBI Pathogen Detection database

Background: Of the > 2600 Salmonella serovars, Salmonella enterica serovar I 4,[5],12:i:- (serovar I 4,[5],12:i:-) has emerged as one of the most common causes of human salmonellosis and the most frequent multidrug-resistant (MDR; resistance to ≥3 antimicrobial classes) nontyphoidal Salmonella serovar in the U.S. Serovar I 4,[5],12:i:- isolates have been described globally with resistance to ampicillin, streptomycin, sulfisoxazole, and tetracycline (R-type ASSuT) and an integrative and conjugative element with multi-metal tolerance named Salmonella Genomic Island 4 (SGI-4). Results: We analyzed 13,612 serovar I 4,[5],12:i:- strain sequences available in the NCBI Pathogen Detection database to determine global distribution, animal sources, presence of SGI-4, occurrence of R-type ASSuT, frequency of antimicrobial resistance (AMR), and potential transmission clusters. Genome sequences for serovar I 4,[5],12:i:- strains represented 30 countries from 5 continents (North America, Europe, Asia, Oceania, and South America), but sequences from the United States (59%) and the United Kingdom (28%) were dominant. The metal tolerance island SGI-4 and the R-type ASSuT were present in 71 and 55% of serovar I 4,[5],12:i:- strain sequences, respectively. Sixty-five percent of strain sequences were MDR which correlates to serovar I 4,[5],12:i:- being the most frequent MDR serovar. The distribution of serovar I 4,[5],12:i:- strain sequences in the NCBI Pathogen Detection database suggests that swine-associated strain sequences were the most frequent food-animal source and were significantly more likely to contain the metal tolerance island SGI-4 and genes for MDR compared to all other animal-associated isolate sequences. Conclusions: Our study illustrates how analysis of genomic sequences from the NCBI Pathogen Detection database can be utilized to identify the prevalence of genetic features such as antimicrobial resistance, metal tolerance, and virulence genes that may be responsible for the successful emergence of bacterial foodborne pathogens.

59 BASIC BIOLOGICAL SCIENCES↗

A transcriptome software comparison for the analyses of treatments expected to give subtle gene expression responses

Background: In this comparative study we evaluate the performance of four software tools: DNAstar-D (DESeq2), DNAstar-E (edgeR), CLC Genomics and Partek Flow for identification of differentially expressed genes (DEGs) using a transcriptome of E. coli. The RNA-seq data are from the effect of below-background radiation 5.5 nGy total dose (0.2nGy/hr) on E. coli grown shielded from natural radiation 655 m below ground in a pre-World War II steel vault. The gene expression response to three supplemented sources of radiation designed to mimic natural background, 1952 – 5720 nGy in total dose (71–208 nGy/hr), are compared to this “radiation-deprived” treatment. In addition, RNA-seq data of Caenorhabditis elegans nematode from similar radiation treatments was analyzed by three of the software packages. Results: In E. coli, the four software programs identified one of the supplementary sources of radiation (KCl) to evoke about 5 times more transcribed genes than the minus-radiation treatment (69–114 differentially expressed genes, DEGs), and so the rest of the analyses used this KCl vs “Minus” comparison. After imposing a 30-read minimum cutoff, one of the DNAStar options shared two of the three steps (mapping, normalization, and statistic) with Partek Flow (they both used median of ratios to normalize and the DESeq2 statistical package), and these two programs identified the highest number of DEGs in common with each other (53). In contrast, when the programs used different approaches in each of the three steps, between 31 and 40 DEGs were found in common. Regarding the extent of expression differences, three of the four programs gave high fold-change results (15–178 fold), but one (DNAstar’s DESeq2) resulted in more conservative fold-changes (1.5–3.5). In a parallel study comparing three qPCR commercial validation software programs, these programs also gave variable results as to which genes were significantly regulated. Similarly, the C. elegans analysis showed exaggerated fold-changes in CLC and DNAstar’s edgeR while DNAstar-D was more conservative. Conclusions: Regarding the extent of expression (fold-change), and considering the subtlety of the very low level radiation treatments, in E. coli three of the four programs gave what we consider exaggerated fold-change results (15 – 178 fold), but one (DNAstar’s DESeq2) gave more realistic fold-changes (1.5–3.5). When RT-qPCR validation comparisons to transcriptome results were carried out, they supported the more conservative DNAstar-D’s expression results. When another model organism’s (nematode) response to these radiation differences was similarly analyzed, DNAstar-D also resulted in the most conservative expression patterns. Therefore, we would propose DESeq2 (“DNAstar-D”) as an appropriate software tool for differential gene expression studies for treatments expected to give subtle transcriptome responses.

59 BASIC BIOLOGICAL SCIENCES↗

Human NCR3 gene variants rs2736191 and rs11575837 alter longitudinal risk for development of pediatric malaria episodes and severe malarial anemia

Background: Plasmodium falciparum malaria is a leading cause of pediatric morbidity and mortality in holoendemic transmission areas. Severe malarial anemia [SMA, hemoglobin (Hb) < 5.0 g/dL in children] is the most common clinical manifestation of severe malaria in such regions. Although innate immune response genes are known to influence the development of SMA, the role of natural killer (NK) cells in malaria pathogenesis remains largely undefined. As such, we examined the impact of genetic variation in the gene encoding a primary NK cell receptor, natural cytotoxicity-triggering receptor 3 (NCR3), on the occurrence of malaria and SMA episodes over time. Methods: Susceptibility to malaria, SMA, and all-cause mortality was determined in carriers of NCR3 genetic variants (i.e., rs2736191:C > G and rs11575837:C > T) and their haplotypes. The prospective observational study was conducted over a 36 mos. follow-up period in a cohort of children (n = 1,515, aged 1.9–40 mos.) residing in a holoendemic P. falciparum transmission region, Siaya, Kenya. Results: Poisson regression modeling, controlling for anemia-promoting covariates, revealed a significantly increased risk of malaria in carriers of the homozygous mutant allele genotype (TT) for rs11575837 after multiple test correction [Incidence rate ratio (IRR) = 1.540, 95% CI = 1.114–2.129, P = 0.009]. Increased risk of SMA was observed for rs2736191 in children who inherited the CG genotype (IRR = 1.269, 95% CI = 1.009–1.597, P = 0.041) and in the additive model (presence of 1 or 2 copies) (IRR = 1.198, 95% CI = 1.030–1.393, P = 0.019), but was not significant after multiple test correction. Modeling of the haplotypes revealed that the CC haplotype had a significant additive effect for protection against SMA (i.e., reduced risk for development of SMA) after multiple test correction (IRR = 0.823, 95% CI = 0.711–0.952, P = 0.009). Although increased susceptibility to SMA was present in carriers of the GC haplotype (IRR = 1.276, 95% CI = 1.030–1.581, P = 0.026) with an additive effect (IRR = 1.182, 95% CI = 1.018–1.372, P = 0.029), the results did not remain significant after multiple test correction. None of the NCR3 genotypes or haplotypes were associated with all-cause mortality. Conclusions: Variation in NCR3 alters susceptibility to malaria and SMA during the acquisition of naturally-acquired malarial immunity. These results highlight the importance of NK cells in the innate immune response to malaria.

60 APPLIED LIFE SCIENCES↗

Genetic variation and genetic complexity of nodule occupancy in soybean inoculated with USDA110 and USDA123 rhizobium strains

Abstract Background Symbiotic nitrogen fixation differs among Bradyrhizobium japonicum strains. Soybean inoculated with USDA123 has a lower yield than strains known to have high nitrogen fixation efficiency, such as USDA110. In the main soybean-producing area in the Midwest of the United States, USDA123 has a high nodule incidence in field-grown soybean and is competitive but inefficient in nitrogen fixation. In this study, a high-throughput system was developed to characterize nodule number among 1,321 Glycine max and 69 Glycine soja accessions single inoculated with USDA110 and USDA123. Results Seventy-three G. max accessions with significantly different nodule number of USDA110 and USDA123 were identified. After double inoculating 35 of the 73 accessions, it was observed that PI189939, PI317335, PI324187B, PI548461, PI562373, and PI628961 were occupied by USDA110 and double-strain nodules but not by USDA123 nodules alone. PI567624 was only occupied by USDA110 nodules, and PI507429 restricted all strains. Analysis showed that 35 loci were associated with nodule number in G. max when inoculated with strain USDA110 and 35 loci with USDA123. Twenty-three loci were identified in G. soja when inoculated with strain USDA110 and 34 with USDA123. Only four loci were common across two treatments, and each locus could only explain 0.8 to 1.5% of phenotypic variation. Conclusions High-throughput phenotyping systems to characterize nodule number and occupancy were developed, and soybean germplasm restricting rhizobium strain USDA123 but preferring USDA110 was identified. The larger number of minor effects and a small few common loci controlling the nodule number indicated trait genetic complexity and strain-dependent nodulation restriction. The information from the present study will add to the development of cultivars that limit USDA123, thereby increasing nitrogen fixation efficiency and productivity.

59 BASIC BIOLOGICAL SCIENCES↗

Evolutionary history of arbuscular mycorrhizal fungi and genomic signatures of obligate symbiosis

The colonization of land and the diversification of terrestrial plants is intimately linked to the evolutionary history of their symbiotic fungal partners. Extant representatives of these fungal lineages include mutualistic plant symbionts, the arbuscular mycorrhizal (AM) fungi in Glomeromycota and fine root endophytes in Endogonales (Mucoromycota), as well as fungi with saprotrophic, pathogenic and endophytic lifestyles. These fungal groups separate into three monophyletic lineages but their evolutionary relationships remain enigmatic confounding ancestral reconstructions. Their taxonomic ranks are currently fluid. In this study, we recognize these three monophyletic linages as phyla, and use a balanced taxon sampling and broad taxonomic representation for phylogenomic analysis that rejects a hard polytomy and resolves Glomeromycota as sister to a clade composed of Mucoromycota and Mortierellomycota. Low copy numbers of genes associated with plant cell wall degradation could not be assigned to the transition to a plant symbiotic lifestyle but appears to be an ancestral phylogenetic signal. Both plant symbiotic lineages, Glomeromycota and Endogonales, lack numerous thiamine metabolism genes but the lack of fatty acid synthesis genes is specific to AM fungi. Many genes previously thought to be missing specifically in Glomeromycota are either missing in all analyzed phyla, or in some cases, are actually present in some of the analyzed AM fungal lineages, e.g. the high affinity phosphorus transporter Pho89. Based on a broad taxon sampling of fungal genomes we present a well-supported phylogeny for AM fungi and their sister lineages. We show that among these lineages, two independent evolutionary transitions to mutualistic plant symbiosis happened in a genomic background profoundly different from that known from the emergence of ectomycorrhizal fungi in Dikarya. These results call for further reevaluation of genomic signatures associated with plant symbiosis.

59 BASIC BIOLOGICAL SCIENCES↗

Chromosome-level genome assemblies and genetic maps reveal heterochiasmy and macrosynteny in endangered Atlantic Acropora

Abstract Background Over their evolutionary history, corals have adapted to sea level rise and increasing ocean temperatures, however, it is unclear how quickly they may respond to rapid change. Genome structure and genetic diversity contained within may highlight their adaptive potential. Results We present chromosome-scale genome assemblies and linkage maps of the critically endangered Atlantic acroporids,Acropora palmataandA. cervicornis. Both assemblies and linkage maps were resolved into 14 chromosomes with their gene content and colinearity. Repeats and chromosome arrangements were largely preserved between the species. The family Acroporidae and the genusAcroporaexhibited many phylogenetically significant gene family expansions. Macrosynteny decreased with phylogenetic distance. Nevertheless, scleractinians shared six of the 21 cnidarian ancestral linkage groups as well as numerous fission and fusion events compared to other distantly related cnidarians. Genetic linkage maps were constructed from oneA. palmatafamily and 16A. cervicornisfamilies using a genotyping array. The consensus maps span 1,013.42 cM and 927.36 cM forA. palmataandA. cervicornis, respectively. Both species exhibited high genome-wide recombination rates (3.04 to 3.53 cM/Mb) and pronounced sex-based differences, known as heterochiasmy, with 2 to 2.5X higher recombination rates estimated in the female maps. Conclusions Together, the chromosome-scale assemblies and genetic maps we present here are the first detailed look at the genomic landscapes of the critically endangered Atlantic acroporids. These data sets revealed that adaptive capacity of Atlantic acroporids is not limited by their recombination rates. The sister species maintain macrosynteny with few genes with high sequence divergence that may act as reproductive barriers between them. In the AtlanticAcropora, hybridization between the two sister species yields an F1 hybrid with limited fertility despite the high levels of macrosynteny and gene colinearity of their genomes. Together, these resources now enable genome-wide association studies and discovery of quantitative trait loci, two tools that can aid in the conservation of these species.

Biotechnology & Applied Microbiology↗

A willow sex chromosome reveals convergent evolution of complex palindromic repeats

Background: Sex chromosomes have arisen independently in a wide variety of species, yet they share common characteristics, including the presence of suppressed recombination surrounding sex determination loci. Mammalian sex chromosomes contain multiple palindromic repeats across the non-recombining region that show sequence conservation through gene conversion and contain genes that are crucial for sexual reproduction. In plants, it is not clear if palindromic repeats play a role in maintaining sequence conservation in the absence of homologous recombination. Results: Here we present the first evidence of large palindromic structures in a plant sex chromosome, based on a highly contiguous assembly of the W chromosome of the dioecious shrub Salix purpurea . The W chromosome has an expanded number of genes due to transpositions from autosomes. It also contains two consecutive palindromes that span a region of 200 kb, with conspicuous 20-kb stretches of highly conserved sequences among the four arms that show evidence of gene conversion. Four genes in the palindrome are homologous to genes in the sex determination regions of the closely related genus Populus , which is located on a different chromosome. These genes show distinct, floral-biased expression patterns compared to paralogous copies on autosomes. Conclusion: The presence of palindromes in sex chromosomes of mammals and plants highlights the intrinsic importance of these features in adaptive evolution in the absence of recombination. Convergent evolution is driving both the independent establishment of sex chromosomes as well as their fine-scale sequence structure.

59 BASIC BIOLOGICAL SCIENCES↗

Structural variant analysis of a cancer reference cell line sample using multiple sequencing technologies

The cancer genome is commonly altered with thousands of structural rearrangements including insertions, deletions, translocation, inversions, duplications, and copy number variations. Thus, structural variant (SV) characterization plays a paramount role in cancer target identification, oncology diagnostics, and personalized medicine. As part of the SEQC2 Consortium effort, the present study established and evaluated a consensus SV call set using a breast cancer reference cell line and matched normal control derived from the same donor, which were used in our companion benchmarking studies as reference samples. We systematically investigated somatic SVs in the reference cancer cell line by comparing to a matched normal cell line using multiple NGS platforms including Illumina short-read, 10X Genomics linked reads, PacBio long reads, Oxford Nanopore long reads, and high-throughput chromosome conformation capture (Hi-C). We established a consensus SV call set of a total of 1788 SVs including 717 deletions, 230 duplications, 551 insertions, 133 inversions, 146 translocations, and 11 breakends for the reference cancer cell line. To independently evaluate and cross-validate the accuracy of our consensus SV call set, we used orthogonal methods including PCR-based validation, Affymetrix arrays, Bionano optical mapping, and identification of fusion genes detected from RNA-seq. We evaluated the strengths and weaknesses of each NGS technology for SV determination, and our findings provide an actionable guide to improve cancer genome SV detection sensitivity and accuracy. A high-confidence consensus SV call set was established for the reference cancer cell line. A large subset of the variants identified was validated by multiple orthogonal methods.

59 BASIC BIOLOGICAL SCIENCES↗

Niche-DE: niche-differential gene expression analysis in spatial transcriptomics data identifies context-dependent cell-cell interactions

Existing methods for analysis of spatial transcriptomic data focus on delineating the global gene expression variations of cell types across the tissue, rather than local gene expression changes driven by cell-cell interactions. We propose a new statistical procedure called niche-differential expression (niche-DE) analysis that identifies cell-type-specific niche-associated genes, which are differentially expressed within a specific cell type in the context of specific spatial niches. We further develop niche-LR, a method to reveal ligand-receptor signaling mechanisms that underlie niche-differential gene expression patterns. Niche-DE and niche-LR are applicable to low-resolution spot-based spatial transcriptomics data and data that is single-cell or subcellular in resolution.

59 BASIC BIOLOGICAL SCIENCES↗