Engineering PapersSearch

SEARCH · Engineering Papers

Results for “genomics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

UnigeneFinder: An Automated Pipeline for Gene Calling From Transcriptome Assemblies Without a Reference Genome

ABSTRACT For most species, transcriptome data are much more readily available than genome data. Without a reference genome, gene calling is cumbersome and inaccurate because of the high degree of redundancy in de novo transcriptome assemblies. To simplify and increase the accuracy of de novo transcriptome assembly in the absence of a reference genome, we developed UnigeneFinder. Combining several clustering methods, UnigeneFinder substantially reduces the redundancy typical of raw transcriptome assemblies. This pipeline offers an effective solution to the problem of inflated transcript numbers, achieving a closer representation of the actual underlying genome. UnigeneFinder performs comparably or better, compared with existing tools, on plant species with varying genome complexities. UnigeneFinder is the only available transcriptome redundancy solution that fully automates the generation of primary transcript, coding region, and protein sequences, analogous to those available for high‐quality reference genomes. These features, coupled with the pipeline’s cross‐platform implementation, focus on automation, and an accessible, user‐friendly interface, make UnigeneFinder a useful tool for many downstream sequence‐based analyses in nonmodel organisms lacking a reference genome, including differential gene expression analysis, accurate ortholog identification, functional enrichments, and evolutionary analyses. UnigeneFinder also runs efficiently both on high‐performance computing (HPC) systems and personal computers, further reducing barriers to use.

Xue, Bo [Plant Resilience Institute Michigan State

Gaia: An AI-enabled genomic context–aware platform for protein sequence annotation

Protein sequence similarity search is fundamental to biology research, but current methods are typically not able to consider crucial genomic context information indicative of protein function, especially in microbial systems. Here, we present Gaia (Genomic AI Annotator), a sequence annotation platform that enables rapid, context-aware protein sequence search across genomic datasets. Gaia leverages gLM2, a mixed-modality genomic language model trained on both amino acid sequences and their genomic neighborhoods to generate embeddings that integrate sequence-structure-context information. This approach allows for the identification of functionally and/or evolutionarily related genes that are found in conserved genomic contexts, which may be missed by traditional sequence- or structure-based search alone. Gaia enables real-time search of a curated database comprising more than 85 million protein clusters from 131,744 microbial genomes. We compare the homolog retrieval performance of Gaia search against other embedding and alignment-based approaches. We provide Gaia as a web-based, freely available tool.

Jha, Nishant

Genomic Language model for Annotation of Repetitive Elements (GLARE) v1.0

GLARE (Genomic Language model for Annotation of Repetitive Elements) is a tool that classifies transposable elements (TEs)—the mobile, repetitive DNA sequences that make up large fractions of eukaryotic genomes. GLARE fine-tunes the NTv3-650M genomic language model on a harmonized collection of curated TE sequences from the PanTEon and Repbase reference databases, assigning each input sequence to one of 11 orders and 32 superfamilies in a Wicker-compatible taxonomy. Features. From nucleotide FASTA input, GLARE outputs per-sequence predictions, class summaries, composition figures, and an annotated FASTA. It provides calibrated confidence scores with optional abstention and runs on CPU or GPU. Uses. GLARE serves as a classification component in genome-annotation pipelines, downstream of TE discovery, supporting genome annotation and comparative and evolutionary genomics. Advantages. GLARE is the first repeat-element classifier to leverage a pretrained genomic language model. Combined with multi-database training, this approach outperformed all nine classifiers in the PanTEon benchmark, generalized better to unseen taxonomic clades, and remained robust to sequence orientation—a common failure mode of existing tools.

Bruna, Tomas [Lawrence Berkeley National Laborator

An Innovative High Throughput Genome Releaser for Rapid and Efficient PCR Screening

High-throughput PCR screening is vital in synthetic biology and metabolic engineering as it allows researchers to rapidly analyze and detect numerous targeted genetic mutation in the genome. Current challenges for high-throughput PCR screening in synthetic biology include efficiently preparing genomic DNA, optimizing protocols for diverse sample types, managing contamination risks, and effectively analyzing the large volumes of data generated while ensuring consistent and accurate results. In this study, we present the development of a High Throughput Genome Releaser (HTGR), an innovative device addressing common challenges in screening PCR. This genome DNA releaser is designed based on a squash method for rapid, cost-effective, and efficient DNA release, optimized for subsequent PCR reactions. After experimenting with various synthetic materials, we selected a plastic that closely replicates the smooth surface and compression properties of microscope slides, ensuring reliable performance. We engineered a device featuring a 96-Well Plate and a shear applicator, operable both manually and automatically, and compatible with standard liquid-handling robot platform. This compatibility enhances ease of use in high-throughput PCR workflows. Additionally, we developed software to support its automatic functions. Our results demonstrated that the specially engineered 96-Well Plate and HTGR can effectively squash fungal spores , which release enough genome DNA for PCR screening. The genome releaser facilitates the preparation of PCR-amplifiable genomic DNA substrate from 96 samples within minutes, eliminates the need for extraction buffers, and is adaptable to a wide range of microorganisms and cells, which could significantly advance biomanufacturing processes.

Yuan, Guoliang [BATTELLE (PACIFIC NW LAB)]

The genomic footprints of wild Saccharum species trace domestication, diversification, and modern breeding of sugarcane

Sugarcane is a major crop of unclear origins due to its complex polyploid interspecific genome. We analyzed genome ancestries using whole-genome sequence data from 390 representative accessions based on repeated k-mers and chloroplast phylogeny. The results provided evidence that Saccharum officinarum was domesticated in the New Guinea region from the S. robustum wild species and revealed that its genome is a mosaic involving different S. robustum subgroups. We discovered a wild Saccharum contributor to most modern cultivars, likely originating from East Melanesia. We highlighted two early centers of sugarcane diversification associated with human transport, one in continental Asia through hybridization with different S. spontaneum subgroups and one in the Melanesian and Polynesian islands via hybridization with the discovered ancestor and Miscanthus. Finally, we revealed the genome ancestry of modern cultivars, highlighting untapped wild Saccharum diversity as a source of alleles for breeding programs.

Garsmeur, Olivier [CIRAD, Montpellier (France). Ag

Genomic Analysis of Aspergillus Section Terrei Reveals a High Potential in Secondary Metabolite Production and Plant Biomass Degradation

Aspergillus terreus has attracted interest due to its application in industrial biotechnology, particularly for the production of itaconic acid and bioactive secondary metabolites. As related species also seem to possess a prosperous secondary metabolism, they are of high interest for genome mining and exploitation. Here, we present draft genome sequences for six species from Aspergillus section Terrei and one species from Aspergillus section Nidulantes. Whole-genome phylogeny confirmed that section Terrei is monophyletic. Genome analyses identified between 70 and 108 key secondary metabolism genes in each of the genomes of section Terrei, the highest rate found in the genus Aspergillus so far. The respective enzymes fall into 167 distinct families with most of them corresponding to potentially unique compounds or compound families. Moreover, 53% of the families were only found in a single species, which supports the suitability of species from section Terrei for further genome mining. Intriguingly, this analysis, combined with heterologous gene expression and metabolite identification, suggested that species from section Terrei use a strategy for UV protection different to other species from the genus Aspergillus. Section Terrei contains a complete plant polysaccharide degrading potential and an even higher cellulolytic potential than other Aspergilli, possibly facilitating additional applications for these species in biotechnology.

60 APPLIED LIFE SCIENCES

The Ciona intestinalis genome: when the constraints are off

The recent genome sequencing of a non-vertebrate deuterostome, the ascidian tunicate Ciona intestinalis, makes a substantial contribution to the fields of evolutionary and developmental biology.1 Tunicates have some of the smallest bilaterian genomes, embryos with relatively few cells, fixed lineages and early determination of cell fates. Initial analyses of the C. intestinalis genome indicate that it has been evolving rapidly. Comparisons with other bilaterians show that C. intestinalis has lost a number of genes, and that many genes linked together in most other bilaterians have become uncoupled. In addition, a number of independent, lineage-specific gene duplications have been detected. These new results, although interesting in themselves, will take on a deeper significance once the genomes of additional invertebrate deuterostomes (e.g. echinoderms, hemichordates and amphioxus) have been sequenced. With such a broadened database, comparative genomics can begin to ask pointed questions about the relationship between the evolution of genomes and the evolution of body plans. Copyright 2003 Wiley Periodicals, Inc.

Review, Tutorial

An archaeal genomic signature

Comparisons of complete genome sequences allow the most objective and comprehensive descriptions possible of a lineage's evolution. This communication uses the completed genomes from four major euryarchaeal taxa to define a genomic signature for the Euryarchaeota and, by extension, the Archaea as a whole. The signature is defined in terms of the set of protein-encoding genes found in at least two diverse members of the euryarchaeal taxa that function uniquely within the Archaea; most signature proteins have no recognizable bacterial or eukaryal homologs. By this definition, 351 clusters of signature proteins have been identified. Functions of most proteins in this signature set are currently unknown. At least 70% of the clusters that contain proteins from all the euryarchaeal genomes also have crenarchaeal homologs. This conservative set, which appears refractory to horizontal gene transfer to the Bacteria or the Eukarya, would seem to reflect the significant innovations that were unique and fundamental to the archaeal "design fabric." Genomic protein signature analysis methods may be extended to characterize the evolution of any phylogenetically defined lineage. The complete set of protein clusters for the archaeal genomic signature is presented as supplementary material (see the PNAS web site, www.pnas.org).

Non-NASA Center

Comparative genomics of Aspergillus nidulans and section Nidulantes

Aspergillus nidulans is an important model organism for eukaryotic biology and the reference for the section Nidulantes in comparative studies. In this study, we de novo sequenced the genomes of 25 species of this section. Whole-genome phylogeny of 34 Aspergillus species and Penicillium chrysogenum clarifies the position of clades inside section Nidulantes. Comparative genomics reveals a high genetic diversity between species with 684 up to 2433 unique protein families. Furthermore, we categorized 2118 secondary metabolite gene clusters (SMGC) into 603 families across Aspergilli, with at least 40 % of the families shared between Nidulantes species. Genetic dereplication of SMGC and subsequent synteny analysis provides evidence for horizontal gene transfer of a SMGC. Proteins that have been investigated in A. nidulans as well as its SMGC families are generally present in the section Nidulantes, supporting its role as model organism. The set of genes encoding plant biomass-related CAZymes is highly conserved in section Nidulantes, while there is remarkable diversity of organization of MAT-loci both within and between the different clades. This study provides a deeper understanding of the genomic conservation and diversity of this section and supports the position of A. nidulans as a reference species for cell biology.

Theobald, Sebastian [Technical University of Denma

Reclassification of Botryococcus braunii chemical races into separate species based on a comparative genomics analysis

The colonial green microalga Botryococcus braunii is well known for producing liquid hydrocarbons that can be utilized as biofuel feedstocks. B. braunii is taxonomically classified as a single species made up of three chemical races, A, B, and L, that are mainly distinguished by the hydrocarbons produced. We previously reported a B race draft nuclear genome, and here we report the draft nuclear genomes for the A and L races. A comparative genomic study of the three B. braunii races and 14 other algal species within Chlorophyta revealed significant differences in the genomes of each race of B. braunii. Phylogenomically, there was a clear divergence of the three races with the A race diverging earlier than both the B and L races, and the B and L races diverging from a later common ancestor not shared by the A race. DNA repeat content analysis suggested the B race had more repeat content than the A or L races. Orthogroup analysis revealed the B. braunii races displayed more gene orthogroup diversity than three closely related Chlamydomonas species, with nearly 24-36% of all genes in each B. braunii race being specific to each race. This analysis suggests the three races are distinct species based on sufficient differences in their respective genomes. We propose reclassification of the three chemical races to the following species names: Botryococcus alkenealis (A race), Botryococcus braunii (B race), and Botryococcus lycopadienor (L race).

59 BASIC BIOLOGICAL SCIENCES

Genome-resolved analysis of Serratia marcescens strain SMTT infers niche specialization as a hydrocarbon-degrader

Abstract Bacteria that are chronically exposed to high levels of pollutants demonstrate genomic and corresponding metabolic diversity that complement their strategies for adaptation to hydrocarbon-rich environments. Whole genome sequencing was carried out to infer functional traits of Serratia marcescens strain SMTT recovered from soil contaminated with crude oil. The genome size (Mb) was 5,013,981 with a total gene count of 4,842. Comparative analyses with carefully selected S. marcescens strains, 2 of which are associated with contaminated soil, show conservation of central metabolic pathways in addition to intra-specific genetic diversity and metabolic flexibility. Genome comparisons also indicated an enrichment of genes associated with multidrug resistance and efflux pumps for SMTT. The SMTT genome contained genes that enable the catabolism of aromatic compounds via the protocatechuate para-degradation pathway, in addition to meta-cleavage of catechol (meta-cleavage pathway II); gene enrichment for aromatic compound degradation was markedly higher for SMTT compared to the other S. marcescens strains analysed. Our data presents a valuable genetic inventory for future studies on strains of S. marcescens and provides insights into those genomic features of SMTT with industrial potential.

Genetics & Heredity

Homologous recombination shapes the architecture and evolution of bacterial genomes

Homologous recombination is a key evolutionary force that varies considerably across bacterial species. However, how the landscape of homologous recombination varies across genes and within individual genomes has only been studied in a few species. Here, we used Approximate Bayesian Computation to estimate the recombination rate along the genomes of 145 bacterial species. Our results show that homologous recombination varies greatly along bacterial genomes and shapes many aspects of genome architecture and evolution. The genomic landscape of recombination presents several key signatures: rates are highest near the origin of replication in most species, patterns of recombination generally appear symmetrical in both replichores (i.e. replicational halves of circular chromosomes) and most species have genomic hotspots of recombination. Furthermore, many closely related species share conserved landscapes of recombination across orthologs indicating that recombination landscapes are conserved over significant evolutionary distances. We show evidence that recombination drives the evolution of GC-content through increasing the effectiveness of selection and not through biased gene conversion, thereby contributing to an ongoing debate. Finally, we demonstrate that the rate of recombination varies across gene function and that many hotspots of recombination are associated with adaptive and mobile regions often encoding genes involved in pathogenicity.

Torrance, Ellis L [University of North Carolina, G

Signatures of Selection for Resistance/Tolerance to Perkinsus olseni in Grooved Carpet Shell Clam ( Ruditapes decussatus ) Using a Population Genomics Approach

ABSTRACT The grooved carpet shell clam ( Ruditapes decussatus ) is a bivalve of high commercial value distributed throughout the European coast. Its production has suffered a decline caused by different factors, especially by the parasite Perkinsus olsenii . Improving production of R . decussatus requires genomic resources to ascertain the genetic factors underlying resistance/tolerance to P. olseni i . In this study, the first reference genome of R . decussatus was assembled through long‐ and short‐read sequencing (1677 contigs; 1.386 Mb) and further scaffolded at chromosome level with Hi‐C (19 superscaffolds; 95.4% of assembly). Repetitive elements were identified (32%) and masked for annotation of 38,276 coding‐ and 13,056 non‐coding genes. This genome was used as a reference to develop a 2bRAD‐Seq 13,438 SNP panel for a genomic screening on six shellfish beds distributed across the Atlantic Ocean and Mediterranean Sea. Beds were selected by perkinsosis prevalence and the infection level was individually evaluated in all the samples. Genetic diversity was significantly higher in the Mediterranean than in the Atlantic region. The main genetic breakage was detected between those regions (F ST = 0.224), being the Mediterranean more heterogeneous than the Atlantic. Several loci under divergent selection (394 outliers; 261 genomic windows) were detected across shellfish beds. Samples were also inspected to detect signals of selection for resistance/tolerance to P. olseni i by using infection‐level and population‐genomics approaches, and 90 common divergent outliers for resistance/tolerance to perkinsosis were identified and used for gene mining. Candidate genes and markers identified provide invaluable information for controlling perkinsosis and for improving production of the grooved carpet shell clam.

Sambade, Inés M. [Department of Zoology, Genetics

Structural and functional analyses of SARS-CoV-2 Nsp3 and its specific interactions with the 5’ UTR of the viral genome

ABSTRACT Non-structural protein 3 (Nsp3) is the largest open reading frame encoded in the SARS-CoV-2 genome, essential for the formation of double-membrane vesicles (DMV) wherein viral RNA replication occurs. We conducted an extensive structure-function analysis of Nsp3 and determined the crystal structures of the ubiquitin-like 1 (Ubl1), nucleic acid binding (NAB), β-coronavirus-specific marker (βSM) domains, and a sub-region of the Y domain of this protein. We show that the Ubl1, ADP-ribose phosphatase (ADRP), human SARS Unique (HSUD), NAB, and Y domains of Nsp3 bind the 5’ UTR of the viral genome and that the Ubl1 and Y domains possess affinity for recognition of this region, suggesting high specificity. The Ubl1-Nucleocapsid (N) protein complex binds the 5’ UTR with greater affinity than the individual proteins alone. Our results suggest that multiple domains of Nsp3, particularly Ubl1 and Y, shepherd the 5’ UTR of the viral genome during translocation through the DMV membrane, priming the Ubl1 domain to load the genome onto N protein. IMPORTANCE The largest protein encoded by the SARS-CoV-2 genome is Nsp3. In infected cells, this multi-domain protein forms a pore structure in the virus-induced double-membrane vesicles (DMV). We have incomplete data on Nsp3 molecular structure, and here, we describe crystal structures for multiple domains of Nsp3. It is thought that newly replicated viral RNA transits through the DMV pore; however, we possess incomplete data on which regions of Nsp3 actually interact with RNA. Here, we present data showing that five domains of Nsp3 interact with the 5’ UTR of the SARS-CoV-2 RNA, including the Y domain for which no function has ever been discovered. These data suggest that the pore structure plays an active role in recognizing the terminal end of the genome, transiting and loading the viral RNA onto the cytoplasmic nucleocapsid protein. These data help expand our knowledge of Nsp3 structure and function and the SARS-CoV-2 replication cycle.

Microbiology

Chromosome-level genome assemblies and genetic maps reveal heterochiasmy and macrosynteny in endangered Atlantic Acropora

Abstract Background Over their evolutionary history, corals have adapted to sea level rise and increasing ocean temperatures, however, it is unclear how quickly they may respond to rapid change. Genome structure and genetic diversity contained within may highlight their adaptive potential. Results We present chromosome-scale genome assemblies and linkage maps of the critically endangered Atlantic acroporids,Acropora palmataandA. cervicornis. Both assemblies and linkage maps were resolved into 14 chromosomes with their gene content and colinearity. Repeats and chromosome arrangements were largely preserved between the species. The family Acroporidae and the genusAcroporaexhibited many phylogenetically significant gene family expansions. Macrosynteny decreased with phylogenetic distance. Nevertheless, scleractinians shared six of the 21 cnidarian ancestral linkage groups as well as numerous fission and fusion events compared to other distantly related cnidarians. Genetic linkage maps were constructed from oneA. palmatafamily and 16A. cervicornisfamilies using a genotyping array. The consensus maps span 1,013.42 cM and 927.36 cM forA. palmataandA. cervicornis, respectively. Both species exhibited high genome-wide recombination rates (3.04 to 3.53 cM/Mb) and pronounced sex-based differences, known as heterochiasmy, with 2 to 2.5X higher recombination rates estimated in the female maps. Conclusions Together, the chromosome-scale assemblies and genetic maps we present here are the first detailed look at the genomic landscapes of the critically endangered Atlantic acroporids. These data sets revealed that adaptive capacity of Atlantic acroporids is not limited by their recombination rates. The sister species maintain macrosynteny with few genes with high sequence divergence that may act as reproductive barriers between them. In the AtlanticAcropora, hybridization between the two sister species yields an F1 hybrid with limited fertility despite the high levels of macrosynteny and gene colinearity of their genomes. Together, these resources now enable genome-wide association studies and discovery of quantitative trait loci, two tools that can aid in the conservation of these species.

Biotechnology & Applied Microbiology

Montane Conifer, Aspen, Meadow, and Sagebrush Metagenome Resolved Genomes and Traits in East River Watershed, Colorado, USA

Climate change is driving vegetation shifts in mountain watersheds, with unknown impacts on biogeochemical cycles. We hypothesize that these shifts will reshape soil microbiomes and associated biogeochemical processes. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed microbiome and microbial functional trait differences between soils under conifer, aspen, forby meadows, and sagebrush across the East River Watershed, CO, controlling for elevation and aspect.Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from soils 0-20cm in depth across three locations in the watershed—Headwaters, Upper Reaches, and Lower Reaches from August 3-11th 2016. Each location was further subdivided into two blocks, with one block on a west facing aspect, and two on the east aspect of the valley. Within blocks, two samples per vegetation type were taken (one at each depth). This resulted in 66 samples, which were sequenced at JGI and can be found under the Joint Genome Institute (JGI) Genomes Online Database (GOLD) sequencing project Gs0118068. Metagenomes were assembled through an inhouse pipeline (see methods), binned using four autobinners (concoct, maxbin2, metabat2, and vamb) and consolidated using dastool. The consolidated bins from all metagenomes were pooled, filtered by completeness (>75%) and contamination (<25%), and dereplicated at 95% ANI using drep. The dataset includes a zip file of 687 genomes (Vegtype_MAGS.zip), the accession numbers for the underlying metagenomes, a csv file with MAG quality metrics and taxonomy from Genome Taxonomy Database (GTDB) and National Center for Biotechnology Information (NCBI) taxonomic representative genome proteins (EastRiver_Vegtype_drep_genome_info.csv), and a file containing MAG quality metrics and taxonomy (gtdb_drep_bin_taxonomy.csv). The dataset additionally includes a sample metadata file (EastRiver_Vegtype_sample_metadata.csv), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a Google KML file for the sampled locations (sample_collection_sites.kml), a location metadata file (locations.csv), a file-level metadata file (flmd.csv), and a data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES

Diversity of Sordariales Fungi: Identification of Seven New Species of Naviculisporaceae Through Morphological Analyses and Genome Sequencing

Thanks to next-generation sequencing (NGS) technologies, the diversity of fungi can now be investigated through the analysis of their genome sequences. Naviculisporaceae is a family within the Sordariales, whose diversity is not well-known, with only one genome sequence published for this family. Here, we report on the isolation and cultivation of 20 new strains of Naviculisporaceae. Their genome sequences, as well as those of the five commercially available strains, were determined, thus providing complete genome sequences for 25 new Naviculisporaceae strains. Species delimitation was conducted using a combination of (1) ITS + LSU phylogenetic analysis of the new isolates along with other known species of the family, (2) comparisons between DNA barcode sequences of the new strains with those of the known species, and (3) average genome-wide nucleotide identity calculation. We built a phylogenomic tree and studied the organization of the mating-type locus. In vitro fruiting was obtained for 16 strains, enabling the definition of seven new species, namely Pseudorhypophila gallica, Pseudorhypophila guyanensis Rhypophila alpibus, Rhypophila brasiliensis, Rhypophila camarguensis, Rhypophila reunionensis and Rhypophila thailandica, as well as two new combinations, namely Pseudorhypophila latipes and Pseudorhypophila oryzae. Eight strains for which in vitro fruiting was not obtained may belong to additional new species. These results expand the known diversity of the Naviculisporaceae and greatly enlarge the genomic data available for the family.

Naviculisporaceae

Metagenome-assembled-genomes recovered from the Arctic drift expedition MOSAiC

The Multidisciplinary Observatory for Study of the Arctic Climate (MOSAiC) expedition consisted of a year-long drifting survey of the Central Arctic Ocean. The ecosystems component of MOSAiC included the sampling of molecular data, with metagenomes collected from a diverse range of environments. The generation of metagenome-assembled-genomes (MAGs) from metagenomes are a starting point for genome-resolved analyses. This dataset presents a catalogue of MAGs recovered from a set of 73 samples from MOSAiC, including 2407 prokaryotic and 56 eukaryotic MAGs, as well as annotations of a near complete eukaryotic MAG using the Joint Genome Institute (JGI) annotation pipeline. The metagenomic samples are from the surface ocean, chlorophyll maximum, mesopelagic and bathypelagic, within leads and under-ice ocean, as well as melt ponds, ice ridges, and first- and second-year sea ice. This set of MAGs can be used to benchmark microbial biodiversity in the Central Arctic Ocean, compare individual strains across space and time, and to study changes in Arctic microbial communities from the winter to summer, at a genomic level.

59 BASIC BIOLOGICAL SCIENCES