Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Genomic selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Multi-Trait Regressor Stacking Increased Genomic Prediction Accuracy of Sorghum Grain Composition

Genomic prediction has enabled plant breeders to estimate breeding values of unobserved genotypes and environments. The use of genomic prediction will be extremely valuable for compositional traits for which phenotyping is labor-intensive and destructive for most accurate results. We studied the potential of Bayesian multi-output regressor stacking (BMORS) model in improving prediction performance over single trait single environment (STSE) models using a grain sorghum diversity panel (GSDP) and a biparental recombinant inbred lines (RILs) population. A total of five highly correlated grain composition traits—amylose, fat, gross energy, protein and starch, with genomic heritability ranging from 0.24 to 0.59 in the GSDP and 0.69 to 0.83 in the RILs were studied. Average prediction accuracies from the STSE model were within a range of 0.4 to 0.6 for all traits across both populations except amylose (0.25) in the GSDP. Prediction accuracy for BMORS increased by 41% and 32% on average over STSE in the GSDP and RILs, respectively. Prediction of whole environments by training with remaining environments in BMORS resulted in moderate to high prediction accuracy. Our results show regression stacking methods such as BMORS have potential to accurately predict unobserved individuals and environments, and implementation of such models can accelerate genetic gain.

54 ENVIRONMENTAL SCIENCES↗

Using intrahost single nucleotide variant data to predict SARS-CoV-2 detection cycle threshold values

Over the last four years, each successive wave of the COVID-19 pandemic has been caused by variants with mutations that improve the transmissibility of the virus. Despite this, we still lack tools for predicting clinically important features of the virus. In this study, we show that it is possible to predict the PCR cycle threshold (Ct) values from clinical detection assays using sequence data. Ct values often correspond with patient viral load and the epidemiological trajectory of the pandemic. Using a collection of 36,335 high quality genomes, we built models from SARS-CoV-2 intrahost single nucleotide variant (iSNV) data, computing XGBoost models from the frequencies of A, T, G, C, insertions, and deletions at each position relative to the Wuhan-Hu-1 reference genome. Our best model had an R 2 of 0.604 [0.593–0.616, 95% confidence interval] and a Root Mean Square Error (RMSE) of 5.247 [5.156–5.337], demonstrating modest predictive power. Overall, we show that the results are stable relative to an external holdout set of genomes selected from SRA and are robust to patient status and the detection instruments that were used. This study highlights the importance of developing modeling strategies that can be applied to publicly available genome sequence data for use in disease prevention and control.

COVID19↗

Introgressions of novel diseases resistance genes from Miscanthus into energycane

Sugarcane (Saccharum) is among the world’s leading bioenergy crops. However, modern sugarcane cultivars are derived from a relatively small set of founder genotypes, which has contributed to cultivar susceptibility to diseases. Miscanthus is a close relative of sugarcane that is genetically diverse and a potential source of genes for improving sugarcane. We found that Miscanthus is a source of resistance to four major diseases of sugarcane and we crossed these disease-resistance genes into a predominantly sugarcane genetic background (BC1 generation). Additionally, we found that Miscanthus could also confer genes for chilling-tolerant photosynthesis to sugarcane, which would be highly advantageous for production of energycane in subtropical environments, like the southern coastal plain of the US. Using advanced modeling techniques, we found that standard marker-assisted selection could be effective for breeding resistance conferred by a small number of genes each with large effect (i.e. vertical resistance), but to breed for many genes each of small effect (i.e. horizontal resistance), genomic selection would be the better strategy. Lastly, we learned that Miscanthus can be induced to flower sooner by giving the plants short days (long nights) but that the ideal day length for a given accession depends on its adaptation to its latitude of origin, with tropical accessions requiring shorter days to flower than accessions from high latitudes. This information will facilitate plant breeders’ ability to make crosses between Miscanthus and sugarcane, and enable greater use of Miscanthus as a genetic resource to improve sugarcane.

09 BIOMASS FUELS↗

Post-composing ontology terms for efficient phenotyping in plant breeding

Abstract Ontologies are widely used in databases to standardize data, improving data quality, integration, and ease of comparison. Within ontologies tailored to diverse use cases, post-composing user-defined terms reconciles the demands for standardization on the one hand and flexibility on the other. In many instances of Breedbase, a digital ecosystem for plant breeding designed for genomic selection, the goal is to capture phenotypic data using highly curated and rigorous crop ontologies, while adapting to the specific requirements of plant breeders to record data quickly and efficiently. For example, post-composing enables users to tailor ontology terms to suit specific and granular use cases such as repeated measurements on different plant parts and special sample preparation techniques. To achieve this, we have implemented a post-composing tool based on orthogonal ontologies providing users with the ability to introduce additional levels of phenotyping granularity tailored to unique experimental designs. Post-composed terms are designed to be reused by all breeding programs within a Breedbase instance but are not exported to the crop reference ontologies. Breedbase users can post-compose terms across various categories, such as plant anatomy, treatments, temporal events, and breeding cycles, and, as a result, generate highly specific terms for more accurate phenotyping.

Mathematical & Computational Biology↗

Physiological Responses of C4 Perennial Bioenergy Grasses to Climate Change: Causes, Consequences, and Constraints

C 4 perennial bioenergy grasses are an economically and ecologically important group whose responses to climate change will be important to the future bioeconomy. These grasses are highly productive and frequently possess large geographic ranges and broad environmental tolerances, which may contribute to the evolution of ecotypes that differ in physiological acclimation capacity and the evolution of distinct functional strategies. C 4 perennial bioenergy grasses are predicted to thrive under climate change—C 4 photosynthesis likely evolved to enhance photosynthetic efficiency under stressful conditions of low [CO 2 ], high temperature, and drought—although few studies have examined how these species will respond to combined stresses or to extremes of temperature and precipitation. Important targets for C 4 perennial bioenergy production in a changing world, such as sustainability and resilience, can benefit from combining knowledge of C 4 physiology with recent advances in crop improvement, especially genomic selection.

Plant Sciences↗

Accurate determination of genotypic variance of cell wall characteristics of a Populus trichocarpa pedigree using high-throughput pyrolysis-molecular beam mass spectrometry

Abstract Background Pyrolysis-molecular beam mass spectrometry (py-MBMS) analysis of a pedigree of Populus trichocarpa was performed to study the phenotypic plasticity and heritability of lignin content and lignin monomer composition. Instrumental and microspatial environmental variability were observed in the spectral features and corrected to reveal underlying genetic variance of biomass composition. Results Lignin-derived ions (including m/z 124, 154, 168, 194, 210 and others) were highly impacted by microspatial environmental variation which demonstrates phenotypic plasticity of lignin composition in Populus trichocarpa biomass. Broad-sense heritability of lignin composition after correcting for microspatial and instrumental variation was determined to be H 2 = 0.56 based on py-MBMS ions known to derive from lignin. Heritability of lignin monomeric syringyl/guaiacyl ratio ( S / G ) was H 2 = 0.81. Broad-sense heritability was also high (up to H 2 = 0.79) for ions derived from other components of the biomass including phenolics (e.g., salicylates) and C5 sugars (e.g., xylose). Lignin and phenolic ion abundances were primarily driven by maternal effects, and paternal effects were either similar or stronger for the most heritable carbohydrate-derived ions. Conclusions We have shown that many biopolymer-derived ions from py-MBMS show substantial phenotypic plasticity in response to microenvironmental variation in plantations. Nevertheless, broad-sense heritability for biomass composition can be quite high after correcting for spatial environmental variation. This work outlines the importance in accounting for instrumental and microspatial environmental variation in biomass composition data for applications in heritability measurements and genomic selection for breeding poplar for renewable fuels and materials.

09 BIOMASS FUELS↗

Selective Whole-Genome Amplification as a Tool to Enrich Specimens with Low Treponema pallidum Genomic DNA Copies for Whole-Genome Sequencing

Downstream next-generation sequencing (NGS) of the syphilis spirochete Treponema pallidum subspecies pallidum (T. pallidum) is hindered by low bacterial loads and the overwhelming presence of background metagenomic DNA in clinical specimens. In this study, we investigated selective whole-genome amplification (SWGA) utilizing multiple displacement amplification (MDA) in conjunction with custom oligonucleotides with an increased specificity for the T. pallidum genome and the capture and removal of 5'-C-phosphate-G-3' (CpG) methylated host DNA using the NEBNext Microbiome DNA enrichment kit followed by MDA with the REPLI-g single cell kit as enrichment methods to improve the yields of T. pallidum DNA in isolates and lesion specimens from syphilis patients. Sequencing was performed using the Illumina MiSeq v2 500 cycle or NovaSeq 6000 SP platform. These two enrichment methods led to 93 to 98% genome coverage at 5 reads/site in 5 clinical specimens from the United States and rabbit-propagated isolates, containing >14 T. pallidum genomic copies/μL of sample for SWGA and >129 genomic copies/μL for CpG methylation capture with MDA. Variant analysis using sequencing data derived from SWGA-enriched specimens showed that all 5 clinical strains had the A2058G mutation associated with azithromycin resistance. SWGA is a robust method that allows direct whole-genome sequencing (WGS) of specimens containing very low numbers of T. pallidum, which has been challenging until now.

59 BASIC BIOLOGICAL SCIENCES↗

Genome‐wide association and genomic prediction for yield and component traits of Miscanthus sacchariflorus

Abstract Accelerating biomass improvement is a major goal of Miscanthus breeding. The development and implementation of genomic‐enabled breeding tools, like marker‐assisted selection (MAS) and genomic selection, has the potential to improve the efficiency of Miscanthus breeding. The present study conducted genome‐wide association (GWA) and genomic prediction of biomass yield and 14 yield‐components traits in Miscanthus sacchariflorus . We evaluated a diversity panel with 590 accessions of M. sacchariflorus grown across 4 years in one subtropical and three temperate locations and genotyped with 268,109 single‐nucleotide polymorphisms (SNPs). The GWA study identified a total of 835 significant SNPs and 674 candidate genes across all traits and locations. Of the significant SNPs identified, 280 were localized in mapped quantitative trait loci intervals and proximal to SNPs identified for similar traits in previously reported Miscanthus studies, providing additional support for the importance of these genomic regions for biomass yield. Our study gave insights into the genetic basis for yield‐component traits in M. sacchariflorus that may facilitate marker‐assisted breeding for biomass yield. Genomic prediction accuracy for the yield‐related traits ranged from 0.15 to 0.52 across all locations and genetic groups. Prediction accuracies within the six genetic groupings of M. sacchariflorus were limited due to low sample sizes. Nevertheless, the Korea/NE China/Russia ( N = 237) genetic group had the highest prediction accuracy of all genetic groups (ranging 0.26–0.71), suggesting that with adequate sample sizes, there is strong potential for genomic selection within the genetic groupings of M. sacchariflorus . This study indicated that MAS and genomic prediction will likely be beneficial for conducting population‐improvement of M. sacchariflorus .

09 BIOMASS FUELS↗

A variant selection framework for genome graphs

Abstract Motivation Variation graph representations are projected to either replace or supplement conventional single genome references due to their ability to capture population genetic diversity and reduce reference bias. Vast catalogues of genetic variants for many species now exist, and it is natural to ask which among these are crucial to circumvent reference bias during read mapping. Results In this work, we propose a novel mathematical framework for variant selection, by casting it in terms of minimizing variation graph size subject to preserving paths of length α with at most δ differences. This framework leads to a rich set of problems based on the types of variants [e.g. single nucleotide polymorphisms (SNPs), indels or structural variants (SVs)], and whether the goal is to minimize the number of positions at which variants are listed or to minimize the total number of variants listed. We classify the computational complexity of these problems and provide efficient algorithms along with their software implementation when feasible. We empirically evaluate the magnitude of graph reduction achieved in human chromosome variation graphs using multiple α and δ parameter values corresponding to short and long-read resequencing characteristics. When our algorithm is run with parameter settings amenable to long-read mapping (α = 10 kbp, δ = 1000), 99.99% SNPs and 73% SVs can be safely excluded from human chromosome 1 variation graph. The graph size reduction can benefit downstream pan-genome analysis. Availability and implementation https://github.com/AT-CG/VF. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Genome engineering allows selective conversions of terephthalaldehyde to multiple valorized products in bacterial cells

Deconstruction of polyethylene terephthalate (PET) plastic waste generates opportunities for valorization to alternative products. We recently designed an enzymatic cascade that could produce terephthalaldehyde (TPAL) from terephthalic acid. Here, we showed that the addition of TPAL to growing cultures of Escherichia coli wild-type strain MG1655 and an engineered strain for reduced aromatic aldehyde reduction (RARE) strain resulted in substantial reduction. We then investigated if we could mitigate this reduction using multiplex automatable genome engineering (MAGE) to create an E. coli strain with 10 additional knockouts in RARE. Encouragingly, we found this newly engineered strain enabled a 2.5-fold higher retention of TPAL over RARE after 24 h. We applied this new strain for the production of para-xylylenediamine (pXYL) and observed a 6.8-fold increase in pXYL titer compared with RARE. Altogether, our study demonstrates the potential of TPAL as a versatile intermediate in microbial biosynthesis of chemicals that derived from waste PET.

59 BASIC BIOLOGICAL SCIENCES↗

Data for Genome-Wide Association and Genomic Prediction for Yield and Component Traits of Miscanthus sacchariflorus

Accelerating biomass improvement is a major goal of miscanthus breeding. The development and implementation of genomic-enabled breeding tools, like marker-assisted selection (MAS) and genomic selection, has the potential to improve the efficiency of miscanthus breeding. The present study conducted genome-wide association (GWA) and genomic prediction of biomass yield and 14 yield-components traits in Miscanthus sacchariflorus . We evaluated a diversity panel with 590 accessions of M. sacchariflorus grown across four years in one subtropical and three temperate locations and genotyped with 268,109 single-nucleotide polymorphisms (SNPs). The GWA study identified a total of 835 significant SNPs and 674 candidate genes across all traits and locations. Of the significant SNPs identified, 280 were localized in mapped quantitative trait loci intervals and proximal to SNPs identified for similar traits in previously reported miscanthus studies, providing additional support for the importance of these genomic regions for biomass yield. Our study gave insights into the genetic basis for yield-component traits in M. sacchariflorus that may facilitate marker-assisted breeding for biomass yield. Genomic prediction accuracy for the yield-related traits ranged from 0.15 to 0.52 across all locations and genetic groups. Prediction accuracies within the six genetic groupings of M. sacchariflorus were limited due to low sample sizes. Nevertheless, the Korea/NE China/Russia (N = 237) genetic group had the highest prediction accuracy of all genetic groups (ranging 0.26–0.71), suggesting that with adequate sample sizes, there is strong potential for genomic selection within the genetic groupings of M. sacchariflorus . This study indicated that MAS and genomic prediction will likely be beneficial for conducting population-improvement of M. sacchariflorus .

Biomass Analytics↗

Plastome evolution in annual Brachypodium species reveals widespread heteroplasmy and chloroplast capture, lineage-specific codon usage bias, and low positive selection

Comparative genomics and plastome phylogenomics have advanced significantly in recent years, highlighting the diversity, possible admixture, and non-neutral evolution of the predominantly considered non-recombinant chloroplast genomes in angiosperms. The grass genus Brachypodium serves as a powerful model for studying evolutionary processes in monocots. We analyzed 287 plastomes across the native circum-Mediterranean range of the three annual Brachypodium species ( B. distachyon, B. stacei, B. hybridum ), focusing on their structural variation, selection patterns and phylogenomic relationships. Our analyses confirmed the differentiation of the S and D plastomes, inherited respectively from the diploid progenitor species B. stacei and B. distachyon . We identified novel structural rearrangements and indels, and unique repeat motifs, along with widespread heteroplasmy, particularly in ancestral B. hybridum -D plastotypes. SNP diversity varied among plastotypes, reflecting population dynamics and evolutionary histories, with B. hybridum -D plastotypes showing the highest normalized diversity and B. hybridum -S the lowest. Positive selection was detected in 29 plastid genes by Tajima’s neutrality test, and in nine genes by site and branch-site evolutionary models, including matK, ndhF, rbcL, and rpoC2. Phylogenomic analyses revealed well-supported clades corresponding to the S and D plastome lineages, with frequent chloroplast capture events and long-distance dispersals shaping their evolutionary trajectories.

allopolyploidy↗

Comparative genomics of the Liberibacter genus reveals widespread diversity in genomic content and positive selection history

‘Candidatus Liberibacter’ is a group of bacterial species that are obligate intracellular plant pathogens and cause Huanglongbing disease of citrus trees and Zebra Chip in potatoes. Here, we examined the extent of intra- and interspecific genetic diversity across the genus using comparative genomics. Our approach examined a wide set of Liberibacter genome sequences including five pathogenic species and one species not known to cause disease. By performing comparative genomics analyses, we sought to understand the evolutionary history of this genus and to identify genes or genome regions that may affect pathogenicity. With a set of 52 genomes, we performed comparative genomics, measured genome rearrangement, and completed statistical tests of positive selection. We explored markers of genetic diversity across the genus, such as average nucleotide identity across the whole genome. These analyses revealed the highest intraspecific diversity amongst the ‘Ca. Liberibacter solanacearum’ species, which also has the largest plant host range. We identified sets of core and accessory genes across the genus and within each species and measured the ratio of nonsynonymous to synonymous mutations (dN/dS) across genes. We identified ten genes with evidence of a history of positive selection in the Liberibacter genus, including genes in the Tad complex, which have been previously implicated as being highly divergent in the ‘Ca. L. capsica’ species based on high values of dN.

59 BASIC BIOLOGICAL SCIENCES↗

Divergent selection and climate adaptation fuel genomic differentiation between sister species of Sphagnum (peat moss)

Abstract Background and Aims New plant species can evolve through the reinforcement of reproductive isolation via local adaptation along habitat gradients. Peat mosses (Sphagnaceae) are an emerging model system for the study of evolutionary genomics and have well-documented niche differentiation among species. Recent molecular studies have demonstrated that the globally distributed species Sphagnum magellanicum is a complex of morphologically cryptic lineages that are phylogenetically and ecologically distinct. Here, we describe the architecture of genomic differentiation between two sister species in this complex known from eastern North America: the northern S. diabolicum and the largely southern S. magniae. Methods We sampled plant populations from across a latitudinal gradient in eastern North America and performed whole genome and restriction-site associated DNA sequencing. These sequencing data were then analyzed computationally. Key Results Using sliding-window population genetic analyses we find that differentiation is concentrated within ‘islands’ of the genome spanning up to 400 kb that are characterized by elevated genetic divergence, suppressed recombination, reduced nucleotide diversity and increased rates of non-synonymous substitution. Sequence variants that are significantly associated with genetic structure and bioclimatic variables occur within genes that have functional enrichment for biological processes including abiotic stress response, photoperiodism and hormone-mediated signalling. Demographic modelling demonstrates that these two species diverged no more than 225 000 generations ago with secondary contact occurring where their ranges overlap. Conclusions We suggest that this heterogeneity of genomic differentiation is a result of linked selection and reflects the role of local adaptation to contrasting climatic zones in driving speciation. This research provides insight into the process of speciation in a group of ecologically important plants and strengthens our predictive understanding of how plant populations will respond as Earth’s climate rapidly changes.

58 GEOSCIENCES↗

Inferring demographic and selective histories from population genomic data using a 2-step approach in species with coding-sparse genomes: an application to human data

Abstract The demographic history of a population, and the distribution of fitness effects (DFE) of newly arising mutations in functional genomic regions, are fundamental factors dictating both genetic variation and evolutionary trajectories. Although both demographic and DFE inference has been performed extensively in humans, these approaches have generally either been limited to simple demographic models involving a single population, or, where a complex population history has been inferred, without accounting for the potentially confounding effects of selection at linked sites. Taking advantage of the coding-sparse nature of the genome, we propose a 2-step approach in which coalescent simulations are first used to infer a complex multi-population demographic model, utilizing large non-functional regions that are likely free from the effects of background selection. We then use forward-in-time simulations to perform DFE inference in functional regions, conditional on the complex demography inferred and utilizing expected background selection effects in the estimation procedure. Throughout, recombination and mutation rate maps were used to account for the underlying empirical rate heterogeneity across the human genome. Importantly, within this framework it is possible to utilize and fit multiple aspects of the data, and this inference scheme represents a generalized approach for such large-scale inference in species with coding-sparse genomes.

Soni, Vivak (ORCID:0000000294969562)↗

Case Study: Can you find Delftia?

Delftia is a genus of bacteria with a bunch of cool features! The best studied species, Delftia acidovorans can produce gold nanoparticles from gold ions in solution. This bacteria has been found living in biofolms with Cupriavidis metallidurans on gold nuggets. It is also found in soil, in sinks and in rhizospheres of different plants where it promotes their growth. Delftia acidovorans forms gold nuggets by producing a short nonribosomal peptide called delftibactin. The 16 genes responsible for the production of delftibactin are called the del cluster (delA-delP). These genes were originally discovered in Delftia acidovorans SPH-1, but our research shows that the del cluster appears to be present across the genus. We'll start with a number of raw reads, clean them up and do a little taxonomy to figure out what we might be able to assemble. Next, we'll assemble the reads using 4 different methods and pick out the best assembly. Then we'll take those reads and sort them by genome and assess the quality of those genomes. We'll select the high quality genomes to annotate and insert into phylogenetic trees to find relatives.

59 BASIC BIOLOGICAL SCIENCES↗

Major impacts of widespread structural variation on sorghum

Genetic diversity is critical to crop breeding and improvement, and dissection of the genomic variation underlying agronomic traits can both assist breeding and give insight into basic biological mechanisms. Although recent genome analyses in plants reveal many structural variants (SVs), most current studies of crop genetic variation are dominated by single-nucleotide polymorphisms (SNPs). The extent of the impact of SVs on global trait variation, as well as their utility in genome-wide selection, is not yet understood. In this study, we built an SV data set based on whole-genome resequencing of diverse sorghum lines (n = 363), validated the correlation of photoperiod sensitivity and variety type, and identified SV hotspots underlying the divergent evolution of cellulosic and sweet sorghum. In addition, we showed the complementary contribution of SVs for heritability of traits related to sorghum adaptation. Importantly, inclusion of SV polymorphisms in association studies revealed genotype–phenotype associations not observed with SNPs alone. Three-way genome-wide association studies (GWAS) based on whole-genome SNP, SV, and integrated SNP + SV data sets showed substantial associations between SVs and sorghum traits. The addition of SVs to GWAS substantially increased heritability estimates for some traits, indicating their important contribution to functional allelic variation at the genome level. Our discovery of the widespread impacts of SVs on heritable gene expression variation could render a plausible mechanism for their disproportionate impact on phenotypic variation. This study expands our knowledge of SVs and emphasizes the extensive impacts of SVs on sorghum.

59 BASIC BIOLOGICAL SCIENCES↗

The reference genome for the northeastern Pacific bull kelp, Nereocystis luetkeana

Bull kelp, Nereocystis luetkeana, is a northeastern Pacific kelp with broad distribution from Alaska to central California. Its population declines have caused severe concerns in northern California, the Salish Sea in Washington, and recently in some populations in Oregon. Despite bull kelp's accumulated ecological and physiological studies, an assembled and annotated genomic reference was still unavailable. Here, we report the complete and annotated genome of Nereocystis luetkeana, produced by the California Conservation Genomics Project (CCGP), which aims to reveal genomic diversity patterns across California by sequencing the complete genomes of approximately 150 carefully selected species. The genome was assembled into 1562 scaffolds with 449.82 Mb, 80x of coverage and 22 952 gene models. BUSCO assembly showed a completeness score of 72% for the stramenopiles gene set. The mitochondria and chloroplast genome sequences have 37 Kb and 131 Mb, respectively. The orthology analysis between 10 Phaeophycean genomes showed 1065 expanded and 286 unique orthogroups for this species. Pairwise comparisons showed 542 orthogroups present only in N. luetkeana and M. pyrifera, another large-body kelp. The enrichment analysis of these orthogroups showed important functions related to central metabolism and signaling due to ATPases enrichment in these two species. This genome assembly will provide an essential resource for the ecology, evolution, conservation, and breeding of bull kelp.

California Conservation Genomics Project—CCGP↗