Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Genomic selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Investigating genomic prediction strategies for grain carotenoid traits in a tropical/subtropical maize panel

Abstract Vitamin A deficiency remains prevalent on a global scale, including in regions where maize constitutes a high percentage of human diets. One solution for alleviating this deficiency has been to increase grain concentrations of provitamin A carotenoids in maize (Zea mays ssp. mays L.)—an example of biofortification. The International Maize and Wheat Improvement Center (CIMMYT) developed a Carotenoid Association Mapping panel of 380 inbred lines adapted to tropical and subtropical environments that have varying grain concentrations of provitamin A and other health-beneficial carotenoids. Several major genes have been identified for these traits, 2 of which have particularly been leveraged in marker-assisted selection. This project assesses the predictive ability of several genomic prediction strategies for maize grain carotenoid traits within and between 4 environments in Mexico. Ridge Regression-Best Linear Unbiased Prediction, Elastic Net, and Reproducing Kernel Hilbert Spaces had high predictive abilities for all tested traits (β-carotene, β-cryptoxanthin, provitamin A, lutein, and zeaxanthin) and outperformed Least Absolute Shrinkage and Selection Operator. Furthermore, predictive abilities were higher when using genome-wide markers rather than only the markers proximal to 2 or 13 genes. These findings suggest that genomic prediction models using genome-wide markers (and assuming equal variance of marker effects) are worthwhile for these traits even though key genes have already been identified, especially if breeding for additional grain carotenoid traits alongside β-carotene. Predictive ability was maintained for all traits except lutein in between-environment prediction. The TASSEL (Trait Analysis by aSSociation, Evolution, and Linkage) Genomic Selection plugin performed as well as other more computationally intensive methods for within-environment prediction. The findings observed herein indicate the utility of genomic prediction methods for these traits and could inform their resource-efficient implementation in biofortification breeding programs.

59 BASIC BIOLOGICAL SCIENCES↗

Hyporheic zone, river, and groundwater metagenome resolved genomes and rpS3 genes in East River Watershed, Colorado USA Summer 2020, 2021

Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from water filter collected across 8 locations along the East River Watershed, CO, and 1 nearby groundwater well. The purpose was to look for connectivity and similarities across the network and to see the impact of the groundwater. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed community composition and strain similarities between the sites and we also compared it to previous metagenomic studies within the watershed looking at floodplain (Matheus Carnevali et al. 2021) and hillslope (Lavy et al. 2019) microbiomes. Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from filters across 8 locations during August 2020 and July 2021. This resulted in 32 samples. The groundwater sample was sequenced at UC Berkley's QB3. The other 31 samples were sequenced at University of Maryland. Metagenomes were assembled using four autobinners and the best bins were selected using dasTool. The genomes were dereplicated at 95% with dRep and the subset of winning genomes were manually curated based on visual inspection of taxonomic profile, GC content, coverage, and a set of 51 bacterial single copy genes (BSCG), and 38 archaeal signal copy genes (ASCG). The dataset includes a zip file of 311 genomes (HZ_River_SW_MAGS_Dereplicated_95.zip). The dataset additionally includes a zipped file of ribosomal protein small subunit 3 (rpS3) proteins from the hyporheic zone and river data (rpS3_Proteins_HZ_River.zip), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a location metadata file (locations.csv). This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

DNA↗

Nuclear and chloroplast genome engineering of a productive non-model alga Desmodesmus armatus: Insights into unusual and selective acquisition mechanisms for foreign DNA

Despite the tremendous potential of algae to contribute to a future bioeconomy, there are practical and theoretical limitations to how well naturally sourced species and strains can perform in an outdoor setting. The application of biotechnology to modulate and engineer algae metabolism or to increase performance, resilience, or produce novel compounds, offers opportunities to overcome some of the major commercialization barriers. There are numerous approaches reported in the literature having variable success on genetic engineering of algae with non-model algae often presenting unique challenges to genetic engineering. We report here on successful nuclear and chloroplast genomic integration of selection marker resistance in the non-model alga Desmodesmus armatus. Nuclear transformation was accomplished using both electroporation and Agrobacterium-mediated approaches. However, in all surviving transformants, DNA integration was accompanied by excision and/or rearrangement of the gene of interest and fluorescence reporter coding sequences. Similarly, chloroplast transformation was successfully accomplished using a biolistic DNA delivery method. For these transformants, we also observed off-target mutations in the chloroplast genome, not previously observed in other, more routinely used, algae species. Finally, we present insights into potential mechanisms for these observed truncations, rearrangements, and mutations in D. armatus.

59 BASIC BIOLOGICAL SCIENCES↗

Poplar: a phylogenomics pipeline

Motivation Generating phylogenomic trees from the genomic data is essential in understanding biological systems. Each step of this complex process has received extensive attention and has been significantly streamlined over the years. Given the public availability of data, obtaining genomes for a wide selection of species is straightforward. However, analyzing that data to generate a phylogenomic tree is a multistep process with legitimate scientific and technical challenges, often requiring a significant input from a domain-area scientist. Results We present Poplar, a new, streamlined computational pipeline, to address the computational logistical issues that arise when constructing the phylogenomic trees. It provides a framework that runs state-of-the-art software for essential steps in the phylogenomic pipeline, beginning from a genome with or without an annotation, and resulting in a species tree. Running Poplar requires no external databases. In the execution, it enables parallelism for execution for clusters and cloud computing. The trees generated by Poplar match closely with state-of-the-art published trees. The usage and performance of Poplar is far simpler and quicker than manually running a phylogenomic pipeline. Availability and implementation Freely available on GitHub at https://github.com/sandialabs/poplar. Implemented using Python and supported on Linux.

Koning, Elizabeth [Sandia National Laboratories (S↗

Combining GWAS and population genomic analyses to characterize coevolution in a legume‐rhizobia symbiosis

Abstract The mutualism between legumes and rhizobia is clearly the product of past coevolution. However, the nature of ongoing evolution between these partners is less clear. To characterize the nature of recent coevolution between legumes and rhizobia, we used population genomic analysis to characterize selection on functionally annotated symbiosis genes as well as on symbiosis gene candidates identified through a two‐species association analysis. For the association analysis, we inoculated each of 202 accessions of the legume host Medicago truncatula with a community of 88 Sinorhizobia (Ensifer) meliloti strains. Multistrain inoculation, which better reflects the ecological reality of rhizobial selection in nature than single‐strain inoculation, allows strains to compete for nodulation opportunities and host resources and for hosts to preferentially form nodules and provide resources to some strains. We found extensive host by symbiont, that is, genotype‐by‐genotype, effects on rhizobial fitness and some annotated rhizobial genes bear signatures of recent positive selection. However, neither genes responsible for this variation nor annotated host symbiosis genes are enriched for signatures of either positive or balancing selection. This result suggests that stabilizing selection dominates selection acting on symbiotic traits and that variation in these traits is under mutation‐selection balance. Consistent with the lack of positive selection acting on host genes, we found that among‐host variation in growth was similar whether plants were grown with rhizobia or N‐fertilizer, suggesting that the symbiosis may not be a major driver of variation in plant growth in multistrain contexts.

59 BASIC BIOLOGICAL SCIENCES↗

Molecular Architecture of Early Dissemination and Massive Second Wave of the SARS-CoV-2 Virus in a Major Metropolitan Area

We sequenced the genomes of 5,085 severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) strains causing two coronavirus disease 2019 (COVID-19) disease waves in metropolitan Houston, TX, an ethnically diverse region with 7 million residents. The genomes were from viruses recovered in the earliest recognized phase of the pandemic in Houston and from viruses recovered in an ongoing massive second wave of infections. The virus was originally introduced into Houston many times independently. Virtually all strains in the second wave have a Gly614 amino acid replacement in the spike protein, a polymorphism that has been linked to increased transmission and infectivity. Patients infected with the Gly614 variant strains had significantly higher virus loads in the nasopharynx on initial diagnosis. We found little evidence of a significant relationship between virus genotype and altered virulence, stressing the linkage between disease severity, underlying medical conditions, and host genetics. Some regions of the spike protein—the primary target of global vaccine efforts—are replete with amino acid replacements, perhaps indicating the action of selection. We exploited the genomic data to generate defined single amino acid replacements in the receptor binding domain of spike protein that, importantly, produced decreased recognition by the neutralizing monoclonal antibody CR3022. Our report represents the first analysis of the molecular architecture of SARS-CoV-2 in two infection waves in a major metropolitan region. The findings will help us to understand the origin, composition, and trajectory of future infection waves and the potential effect of the host immune response and therapeutic maneuvers on SARS-CoV-2 evolution.

59 BASIC BIOLOGICAL SCIENCES↗

Extreme dimensions — how big (or small) can tailed phages be?

Bacteriophages (phages), that is, viruses that infect bacteria, represent an extremely diverse, yet under-characterized, group of viruses. Although most known phages harbour genomes that are shorter than 200 kb packaged into capsids with a diameter under 100 Å, more and more ‘extremely large’ phages are being discovered. Early reports of phages with genomes larger than 200 kb defined them as ‘jumbo phages’1. Phylogenetic analysis of these jumbo phages revealed that they form distinct and cohesive groups with multiple independent origins. This suggested that jumbo phages are not merely mistakes or aberrations of smaller phages that got too big, but that large genome size is generally a stable trait. In fact, a larger genome can be advantageous; almost all jumbo phages encode their own specific transcription factors, at least some parts of the replication machinery, and many also have their own tRNA genes, which enables increased independence from their host. Here, large genomes also provide more flexibility in terms of transcription strategy; smaller phage genomes tend to have a strict modular genome structure to increase efficiency, whereas jumbo phage genomes may be much less organized in the absence of a strong selection pressure to maintain a compact genome.

59 BASIC BIOLOGICAL SCIENCES↗

Analysis of Salmonella enterica Isolated from a Mixed-Use Watershed in Georgia, USA: Antimicrobial Resistance, Serotype Diversity, and Genetic Relatedness to Human Isolates

As the cases of Salmonella enterica infections associated with contaminated water are increasing, this study was conducted to address the role of surface water as a reservoir of S. enterica serotypes. We sampled rivers and streams (n = 688) over a 3-year period (2015 to 2017) in a mixed-use watershed in Georgia, USA, and 70.2% of the total stream samples tested positive for Salmonella. A total of 1,190 isolates were recovered and characterized by serotyping, antimicrobial susceptibility testing, and pulsed-field gel electrophoresis (PFGE). A wide range of serotypes was identified, including those commonly associated with humans and animals, with S. enterica serotype Muenchen being predominant (22.7%) and each serotype exhibiting a high degree of strain diversity by PFGE. About half (46.1%) of the isolates had PFGE patterns indistinguishable from those of human clinical isolates in the CDC PulseNet database. A total of 52 isolates (4.4%) were resistant to antimicrobials, out of which 43 isolates were multidrug resistant (MDR; resistance to two or more classes of antimicrobials). These 52 resistant Salmonella isolates were screened for the presence of antimicrobial resistance genes, plasmid replicons, and class 1 integrons, out of which four representative MDR isolates were selected for whole-genome sequencing analysis. The results showed that 28 MDR isolates resistant to 10 antimicrobials had blacmy-2 on an A/C plasmid. Persistent contamination of surface water with a high diversity of Salmonella strains, some of which are drug resistant and genetically indistinguishable from human isolates, supports a role of environmental surface water as a reservoir for and transmission route of this pathogen.

59 BASIC BIOLOGICAL SCIENCES↗

Proteomic Tissue-Based Classifier for Early Prediction of Prostate Cancer Progression

Although ~40% of screen-detected prostate cancers (PCa) are indolent, advanced-stage PCa is a lethal disease with 5-year survival rates around 29%. Identification of biomarkers for early detection of aggressive disease is a key challenge. Starting with 52 candidate biomarkers, selected from existing PCa genomics datasets and known PCa driver genes, we used targeted mass spectrometry to quantify proteins that significantly differed in primary tumors from PCa patients treated with radical prostatectomy (RP) across three study outcomes: (i) metastasis ≥1-year post-RP, (ii) biochemical recurrence ≥1-year post-RP, and (iii) no progression after ≥10 years post-RP. Sixteen proteins that differed significantly in an initial set of 105 samples were evaluated in the entire cohort (n = 338). A five-protein classifier which combined FOLH1, KLK3, TGFB1, SPARC, and CAMKK2 with existing clinical and pathological standard of care variables demonstrated significant improvement in predicting distant metastasis, achieving an area under the receiver-operating characteristic curve of 0.92 (0.86, 0.99, p = 0.001) and a negative predictive value of 92% in the training/testing analysis. This classifier has the potential to stratify patients based on risk of aggressive, metastatic PCa that will require early intervention compared to low risk patients who could be managed through active surveillance.

60 APPLIED LIFE SCIENCES↗

Cancer genomics predicts disease relapse and therapeutic response to neoadjuvant chemotherapy of hormone sensitive breast cancers

Several studies provide insight into the landscape of breast cancer genomics with the genomic characterization of tumors offering exceptional opportunities in defining therapies tailored to the patient’s specific need. However, translating genomic data into personalized treatment regimens has been hampered partly due to uncertainties in deviating from guideline based clinical protocols. Here we report a genomic approach to predict favorable outcome to treatment responses thus enabling personalized medicine in the selection of specific treatment regimens. The genomic data were divided into a training set of N = 835 cases and a validation set consisting of 1315 hormone sensitive, 634 triple negative breast cancer (TNBC) and 1365 breast cancer patients with information on neoadjuvant chemotherapy responses. Patients were selected by the following criteria: estrogen receptor (ER) status, lymph node invasion, recurrence free survival. The k-means classification algorithm delineated clusters with low- and high- expression of genes related to recurrence of disease; a multivariate Cox’s proportional hazard model defined recurrence risk for disease. Classifier genes were validated by Immunohistochemistry (IHC) using tissue microarray sections containing both normal and cancerous tissues and by evaluating findings deposited in the human protein atlas repository. Based on the leave-on-out cross validation procedure of 4 independent data sets we identified 51-genes associated with disease relapse and selected 10, i.e. TOP2A, AURKA, CKS2, CCNB2, CDK1 SLC19A1, E2F8, E2F1, PRC1, KIF11 for in depth validation. Expression of the mechanistically linked disease regulated genes significantly correlated with recurrence free survival among ER-positive and triple negative breast cancer patients and was independent of age, tumor size, histological grade and node status. Importantly, the classifier genes predicted pathological complete responses to neoadjuvant chemotherapy (P < 0.001) with high expression of these genes being associated with an improved therapeutic response toward two different anthracycline-taxane regimens; thus, highlighting the prospective for precision medicine. Our study demonstrates the potential of classifier genes to predict risk for disease relapse and treatment response to chemotherapies. The classifier genes enable rational selection of patients who benefit best from a given chemotherapy thus providing the best possible care. The findings encourage independent clinical validation.

59 BASIC BIOLOGICAL SCIENCES↗

Sporophyte Stage Genes Exhibit Stronger Selection Than Gametophyte Stage Genes in Haplodiplontic Giant Kelp

Macrocystis pyrifera (giant kelp), a haplodiplontic brown macroalga that alternates between a macroscopic diploid (sporophyte) and a microscopic haploid (gametophyte) phase, provides an ideal system to investigate how ploidy background affects the evolutionary history of a gene. In M. pyrifera , the same genome is subjected to different selective pressures and environments as it alternates between haploid and diploid life stages. We assembled M. pyrifera gene models using available expression data and validated 8,292 genes models using the model alga Ectocarpus siliculosus . Differential expression analysis identified gene models expressed in either or both the haploid and diploid life stages while functional annotation identified processes enriched in each stage. Genes expressed preferentially or exclusively in the gametophyte stage were found to have higher nucleotide diversity (π = 2.3 × 10 –3 and 2.8 × 10 –3 , respectively) than those for sporophytes (π = 1.1 × 10 –3 and 1 × 10 –3 , respectively). While gametophyte-biased genes show faster sequence evolution, the sequence evolution exhibits less signatures of adaptations when compared to sporophyte-biased genes. Our findings contrast the standing masking hypothesis, which predicts higher standing genetic variation at the sporophyte stage, and support the strength of expression theory, which posits that genes expressed more strongly are expected to evolve slower. We argue that the sporophyte stage undergoes more stringent selection compared with the gametophyte stage, which carries a heavy genetic load associated with broadcast spawning. Furthermore, using whole-genome sequencing, we confirm the strong population structure in wild M. pyrifera populations previously established using microsatellite markers, and estimate population genetic parameters, such as pairwise genetic diversity and Tajima’s D , important for conservation and domestication of M. pyrifera .

Molano, Gary↗

Single nucleotide variants drive evolutionary phage-host arms race in anaerobic carbon dioxide-converting microbiome

Microbial bioconversions are shaped by environmental perturbations and the adaptation of resident microbiomes. Prokaryotes coexist with bacteriophages, yet their coevolutionary trajectories remain underexplored. Here, we investigate the effects of a cultivation vessel leak on an anaerobic consortium performing carbon dioxide reduction. Using time-series shotgun metagenomic sequencing, we reconstruct microbial and viral genomes to track community shifts. We further apply single-nucleotide variant profiling and CRISPR array analysis to monitor viral microdiversity and host defense mechanisms. After bioaugmentation restores bioconversion efficiency, the consortium undergoes pronounced restructuring, with new dominant taxa emerging from the rare biosphere. We identify patterns consistent with phage predation selectively removing certain species, while others exhibit resilience to infection. This shift aligns with a widespread viral outbreak and a transient increased frequency of single nucleotide variants in bacterial CRISPR–Cas defense genes. Expansion of CRISPR spacers further supports that CRISPR-mediated processes influence microbial resilience. Concurrently, phages infecting resilient hosts exhibited adaptive evolution, marked by high genetic heterogeneity. Selective pressure varies across their genomes, targeting infectivity genes and protospacer-adjacent motifs. These findings highlight a dynamic evolutionary arms race driven by the selection of beneficial genetic variants, providing a mechanistic framework for multi-omics investigations, and informing biotechnological applications, including phage-based microbiome manipulation.

Ghiotto, G↗

Engineering Citrobacter freundii using CRISPR/Cas9 system

The CRISPR/Cas9 (clustered regularly interspaced short palindromic repeats/CRISPR associated proteins) system is a useful tool to edit genomes quickly and efficiently. However, the use of CRISPR/Cas9 to edit bacterial genomes has been limited to select microbial chassis primarily used for bioproduction of high value products. Thus, expansion of CRISPR/Cas9 tools to other microbial organisms is needed. Here, our aim was to assess the suitability of CRISPR/Cas9 for genome editing of the Citrobacter freundii type strain ATCC 8090. We evaluated the commonly used two plasmid pCas/pTargetF system to enable gene deletions and insertions in C. freundii and determined editing efficiency. The CRISPR/Cas9 based method enabled high editing efficiency (~91%) for deletion of galactokinase (galk) and enabled deletion with various single guide RNA (sgRNA) sequences. To assess the ability of CRISPR/Cas9 tools to insert genes, we used the fluorescent reporter mNeonGreen, an endopeptidase (yebA), and a transcriptional regulator (xylS) and found successful insertion with high efficiency (81-100%) of each gene individually. These results strengthen and expand the use of CRISPR/Cas9 genome editing to C. freundii as an additional microbial chassis.

59 BASIC BIOLOGICAL SCIENCES↗

High-throughput genetic engineering of nonmodel and undomesticated bacteria via iterative site-specific genome integration

Efficient genome engineering is critical to understand and use microbial functions. Despite recent development of tools such as CRISPR-Cas gene editing, efficient integration of exogenous DNA with well-characterized functions remains limited to model bacteria. Here, we describe serine recombinase–assisted genome engineering, or SAGE, an easy-to-use, highly efficient, and extensible technology that enables selection marker–free, site-specific genome integration of up to 10 DNA constructs, often with efficiency on par with or superior to replicating plasmids. SAGE uses no replicating plasmids and thus lacks the host range limitations of other genome engineering technologies. We demonstrate the value of SAGE by characterizing genome integration efficiency in five bacteria that span multiple taxonomy groups and biotechnology applications and by identifying more than 95 heterologous promoters in each host with consistent transcription across environmental and genetic contexts. We anticipate that SAGE will rapidly expand the number of industrial and environmental bacteria compatible with high-throughput genetics and synthetic biology.

59 BASIC BIOLOGICAL SCIENCES↗

Genome collection processing for “Conserved upper thermal limits and small safety margins in soil copiotrophic bacteria”

We extracted the genomic DNA of 400 randomly selected isolates using a Quick-DNA Microprep Kit (Zymo Research D3020) according to the manufacturer’s protocol. We then submitted the extracted gDNA samples for short-read Illumina sequencing (200 Mbp) at SeqCoast Genomics (Portsmouth, NH, USA). After preprocessing the sequences using Trimmommatic (Bolger et al. 2014), we assembled the genomes using SPADES (Bankevich et al. 2012) and checked the quality of each assembly using QUAST (Gurevich et al. 2013). We processed the genome assemblies using a KBase (v1.4.0) pipeline (Allen et al. 2017; Arkin et al. 2018). Briefly, we used DRAM (v0.1.2) with default settings to annotate the genome assemblies. We then evaluated genome quality and possible contamination levels using CheckM (v1.0.18) (Parks et al. 2015) and retained genomes with completeness above 98% and contamination below 5% (n = 354), following the authors' guidelines. We then obtained taxonomic assignments for all remaining isolates using the Genome Taxonomy Database tool GTDB-Tk (v2.3.2, database version r214) (Chaumeil et al. 2019). We constructed a phylogenetic tree using the tool SpeciesTree (v2.2.0). We then trimmed the tree (using Trim SpeciesTree to GenomeSet- v1.4.0), retaining only tips within our collection with measured thermal performance.

59 BASIC BIOLOGICAL SCIENCES↗

Convergent reductive evolution and host adaptation in Mycoavidus bacterial endosymbionts of Mortierellaceae fungi

Intimate associations between fungi and intracellular bacterial endosymbionts are becoming increasingly well understood. Phylogenetic analyses demonstrate that bacterial endosymbionts of Mucoromycota fungi are related either to free-living Burkholderia or Mollicutes species. The so-called Burkholderia-related endosymbionts or BRE comprise Mycoavidus, Mycetohabitans and Candidatus Glomeribacter gigasporarum. These endosymbionts are marked by genome contraction thought to be associated with intracellular selection. However, the conclusions drawn thus far are based on a very small subset of endosymbiont genomes, and the mechanisms leading to genome streamlining are not well understood. The purpose of this study was to better understand how intracellular existence shapes Mycoavidus and BRE functionally at the genome level. To this end we generated and analyzed 14 novel draft genomes for Mycoavidus living within the hyphae of Mortierellomycotina fungi. We found that our novel Mycoavidus genomes were significantly reduced compared to free-living Burkholderiales relatives. Using a genome-scale phylogenetic approach including the novel and available existing genomes of Mycoavidus, we show that the genus is an assemblage composed of two independently derived lineages including three well supported clades of Mycoavidus. Using a comparative genomic approach, we shed light on the functional implications of genome reduction, documenting shared and unique gene loss patterns between the three Mycoavidus clades. We found that many endosymbiont isolates demonstrate patterns of vertical transmission and host-specificity, but others are present in phylogenetically disparate hosts. We discuss how reductive evolution and host specificity reflect convergent adaptation to the intrahyphal selective landscape, and commonalities of eukaryotic endosymbiont genome evolution.

59 BASIC BIOLOGICAL SCIENCES↗

Genome-wide toxicogenomic study of the lanthanides sheds light on the selective toxicity mechanisms associated with critical materials

Significance The growing use of lanthanides in various industries has increased the potential for human exposure to large concentrations of these heavy metals throughout the life cycle of new technologies, requiring more detailed investigations into their toxicological properties. The eukaryote model organism Saccharomyces cerevisiae is ideal to apply functional toxicogenomics tools at the system level and probe fundamental cellular functions disrupted by rare-earth metals. A comprehensive mechanistic assay was conducted to evaluate toxicity across the entire lanthanide series and provide information about which toxicological mechanisms may be shared by the different metals. We identified distinct characteristic behaviors between early and middle/late lanthanides, highlighting the powerful discrimination capabilities of natural systems and pointing to detailed mechanistic pathways associated with lanthanide exposure.

59 BASIC BIOLOGICAL SCIENCES↗

FeGenie: a comprehensive tool for the identification of iron genes and iron gene neighborhoods in genomes and metagenome assemblies

Iron is a micronutrient for nearly all life on Earth. It can be used as an electron donor and electron acceptor by iron-oxidizing and iron-reducing microorganisms and is used in a variety of biological processes, including photosynthesis and respiration. While it is the fourth most abundant metal in the Earth’s crust, iron is often limiting for growth in oxic environments because it is readily oxidized and precipitated. Much of our understanding of how microorganisms compete for and utilize iron is based on laboratory experiments. However, the advent of next-generation sequencing and surge in publicly available sequence data has made it possible to probe the structure and function of microbial communities in the environment. To bridge the gap between our understanding of iron acquisition, iron redox cycling, iron storage, and magnetosome formation in model microorganisms and the plethora of sequence data available from environmental studies, we have created a comprehensive database of hidden Markov models (HMMs) based on genes related to iron acquisition, storage, and reduction/oxidation in Bacteria and Archaea. Along with this database, we present FeGenie, a bioinformatics tool that accepts genome and metagenome assemblies as input and uses our comprehensive HMM database to annotate provided datasets with respect to iron-related genes and gene neighborhood. An important contribution of this tool is the efficient identification of genes involved in iron oxidation and dissimilatory iron reduction, which have been largely overlooked by standard annotation pipelines. We validated FeGenie against a selected set of 28 isolate genomes and showcase its utility in exploring iron genes present in 27 metagenomes, 4 isolate genomes from human oral biofilms, and 17 genomes from candidate organisms, including members of the candidate phyla radiation. We show that FeGenie accurately identifies iron genes in isolates. Furthermore, analysis of metagenomes using FeGenie demonstrates that the iron gene repertoire and abundance of each environment is correlated with iron richness. While this tool will not replace the reliability of culture-dependent analyses of microbial physiology, it provides reliable predictions derived from the most up-to-date genetic markers. FeGenie’s database will be maintained and continually updated as new genes are discovered.

59 BASIC BIOLOGICAL SCIENCES↗