Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “SnP”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Improving precision and accuracy of genetic mapping with genotyping‐by‐sequencing data in outcrossing species

Abstract Genotyping‐by‐sequencing (GBS) is a widely used strategy for obtaining large numbers of genetic markers in model and non‐model organisms. In crop plants, GBS‐derived marker datasets are frequently used to perform quantitative trait locus (QTL) mapping. In some plant species, however, high heterozygosity and complex genome structure mean that researchers must use care in handling GBS data to conduct QTL mapping most effectively. Such outbred crops include most of the perennial grass and tree species used for bioenergy. To identify strategies for increasing accuracy and precision of QTL mapping using GBS data in outbred crops, we conducted an empirical study of SNP‐calling and genetic map‐building pipeline parameters in a Miscanthus sinensis population, and a complementary simulation study to estimate the relationship between genome‐wide error rate, read depth, and marker number. The bioenergy grass Miscanthus is an obligate outcrossing species with a recent (diploidized) whole‐genome duplication. For the study of empirical M. sinensis data, we compared two SNP‐calling methods (one non‐reference‐based and one reference‐based), a series of depth filters (12×, 20×, 30×, and 40×) and two map‐construction methods (i.e., marker ordering: linkage‐only and order‐corrected based on a reference genome). We found that correcting the order of markers on a linkage map by using a high‐quality reference genome improved QTL precision (shorter confidence intervals). For typical GBS datasets of between 1000 and 5000 markers to build a genetic map for biparental populations, a depth filter set at 30× to 40× applied to outbred populations provided a genome‐wide genotype‐calling error rate of less than 1%, improved accuracy of QTL point estimates and minimized type I errors for identifying QTL. Based on these results, we recommend using a reference genome to correct the marker order of genetic maps and a robust genotype depth filter to improve QTL mapping for outbred crops.

59 BASIC BIOLOGICAL SCIENCES↗

Maast: genotyping thousands of microbial strains efficiently

Existing single nucleotide polymorphism (SNP) genotyping algorithms do not scale for species with thousands of sequenced strains, nor do they account for conspecific redundancy. Here we present a bioinformatics tool, Maast, which empowers population genetic meta-analysis of microbes at an unrivaled scale. Maast implements a novel algorithm to heuristically identify a minimal set of diverse conspecific genomes, then constructs a reliable SNP panel for each species, and enables rapid and accurate genotyping using a hybrid of whole-genome alignment and k-mer exact matching. We demonstrate Maast’s utility by genotyping thousands of Helicobacter pylori strains and tracking SARS-CoV-2 diversification.

59 BASIC BIOLOGICAL SCIENCES↗

Data for "Genetics of flooding tolerance in an F2 Miscanthus sacchariflorus ssp. lutarioriparius × M. sinensis population"

This dataset contains all data and supplementary materials from "Genetics of flooding tolerance in an F2 Miscanthus sacchariflorus ssp. lutarioriparius × M. sinensis population". 1. The dataset S1 table contains the raw phenotypic data collected during the experiment. 2. The dataset S2 table contains the LSmean values for the 24 traits studied. 3. The dataset S3 table contains the TASSEL GBSv2 map, marker information, and genotype data used for mapping. 4. The dataset S4 table contains information on candidate genes found in each of the QTL intervals. 5. The dataset S5 table contains the GO annotations and KEGG enrichment analyses for those candidate genes. 6. The dataset S6 table contains information on the sequences used to classify AP2 ERF transcription factors. 7. The dataset S7 table contains information on AP2 ERF orthologs between Miscanthus and rice based on synteny. 8. Supplementary file 1 contains the ANOVA results using the raw phenotypic data collected from protocol "A". 9. Supplementary file 2 contains the ANOVA results using the raw phenotypic data collected from protocol "B". 10. Supplementary file 3 contains notes on the comparison of SNP calling methods. 11. Supplementary file 4 is a script for analyzing candidate genes found in QTL intervals.

Miscanthus, flood, partial submergence, complete s↗

Enhanced Resistance Pines for Improved Renewable Biofuel and Chemical Production (Technical Report)

We completed phenotyping constitutive and inducible oleoresin flow across two seasons, constitutive resin canal number and density and wood terpene content in our ADEPT2 and CCLONES populations. We completed genetic association between 19 oleoresin phenotypes and a total of 523,192 SNP markers from ADEPT2 and 13,883 SNP markers in CCLONES using four mixed linear models. A total of 293 significant SNPs (FDR = 0.20) were identified. We used the MENTOR tool to mine mechanistic connections from a multiplex network constructed from poplar multi-omic data to construct a conceptual model for a subset of these significant SNPs. Our model contains 6 transcriptional regulators in addition to 3 monoterpene synthases. To generate more lines of evidence for these significant SNPs, we completed a time course RNAseq experiment after inducing vascular zone cells to differentiate into new resin canals with a methyl jasmonate treatment, a single nuclei RNAseq that identified differentiating resin canal epithelial cells and are completing analysis for a QTL study in a hybrid pine population. The time course identified 4634 significantly down and 1890 significantly up regulated transcripts after treatment with methyl jasmonate, an inducer of new resin canal formation in the vascular cambial meristem. To analyze this large set of differentially regulated genes, we created a predictive expression network and analyzed it with random walk restart using 6 seed genes coding for transcription factors regulating xylem differentiation in poplar. Of the top ranked 200 transcripts, 119 transcripts were significant differentially expressed supporting these transcripts as potential candidates regulating resin canal formation. Analysis of single nuclei sequencing of shoot tips that contain differentiating resin canals, identified 10 clusters. One cluster was highly enriched in transcripts coding for 9 of the enzymes in the MEP pathway 3 prenyl synthetases, and 3 monoterpene synthases strongly suggesting that this cluster represents resin canal epithelial cells. We are mining the additional transcripts to create a trajectory analysis. In summary, we have identified > 10 novel genes that are strongly supported candidates for further analysis in breeding lines and for genetic engineering over- and under- expressing lines to increase wood terpene content to improve resistance to insect and fungal pathogens while simultaneously increasing terpene supplies for renewable chemicals and biofuels.

59 BASIC BIOLOGICAL SCIENCES↗

Novel Mode of Molybdate Inhibition of Desulfovibrio vulgaris Hildenborough

Sulfate-reducing microorganisms (SRM) are found in multiple environments and play a major role in global carbon and sulfur cycling. Because of their growth capabilities and association with metal corrosion, controlling the growth of SRM has become of increased interest. One such mechanism of control has been the use of molybdate (MoO 4 2− ), which is thought to be a specific inhibitor of SRM. The way in which molybdate inhibits the growth of SRM has been enigmatic. It has been reported that molybdate is involved in a futile energy cycle with the sulfate-activating enzyme, sulfate adenylyl transferase (Sat), which results in loss of cellular ATP. However, we show here that a deletion of this enzyme in the model SRM, Desulfovibrio vulgaris Hildenborough, remained sensitive to molybdate. We performed several subcultures of the ∆ sat strain in the presence of increasing concentrations of molybdate and obtained a culture with increased resistance to the inhibitor (up to 3 mM). The culture was re-sequenced and three single nucleotide polymorphisms (SNPs) were identified that were not present in the parental strain. Two of the SNPs seemed unlikely candidates for molybdate resistance due to a lack of conservation of the mutated residues in homologous genes of closely related strains. The remaining SNP was located in DVU2210, a protein containing two domains: a YcaO-like domain and a tetratricopeptide-repeat domain. The SNP resulted in a change of a serine residue to arginine in the ATP-hydrolyzing motif of the YcaO-like domain. Deletion mutants of each of the three genes apparently enriched with SNPs in the presence of inhibitory molybdate and combinations of these genes were generated in the Δ sat and wild-type strains. Strains lacking both sat and DVU2210 became more resistant to molybdate. Deletions of the other two genes in which SNPs were observed did not result in increased resistance to molybdate. YcaO-like proteins are distributed across the bacterial and archaeal domains, though the function of these proteins is largely unknown. The role of this protein in D. vulgaris is unknown. Due to the distribution of YcaO-like proteins in prokaryotes, the veracity of molybdate as a specific SRM inhibitor should be reconsidered.

59 BASIC BIOLOGICAL SCIENCES↗

Dissemination of blaNDM-5 and mcr-8.1 in carbapenem-resistant Klebsiella pneumoniae and Klebsiella quasipneumoniae in an animal breeding area in Eastern China

Animal farms have become one of the most important reservoirs of carbapenem-resistant Klebsiella spp. (CRK) owing to the wide usage of veterinary antibiotics. “One Health”-studies observing animals, the environment, and humans are necessary to understand the dissemination of CRK in animal breeding areas. Based on the concept of “One-Health,” 263 samples of animal feces, wastewater, well water, and human feces from 60 livestock and poultry farms in Shandong province, China were screened for CRK. Five carbapenem-resistant Klebsiella pneumoniae (CRKP) and three carbapenem-resistant Klebsiella quasipneumoniae (CRKQ) strains were isolated from animal feces, human feces, and well water. The eight strains were characterized by antimicrobial susceptibility testing, plasmid conjugation assays, whole-genome sequencing, and bioinformatics analysis. All strains carried the carbapenemase-encoding gene bla NDM-5 , which was flanked by the same core genetic structure (IS 5 - bla NDM-5 - ble MBL - trpF - dsbD -IS 26 -IS Kox3 ) and was located on highly related conjugative IncX3 plasmids. The colistin resistance gene mcr-8.1 was carried by three CRKP and located on self-transmissible IncFII(K)/IncFIA(HI1) and IncFII(pKP91)/IncFIA(HI1) plasmids. The genetic context of mcr-8.1 consisted of IS 903 - orf - mcr-8.1-copR-baeS-dgkA - orf -IS 903 in three strains. Single nucleotide polymorphism (SNP) analysis confirmed the clonal spread of CRKP carrying- bla NDM-5 and mcr-8.1 between two human workers in the same chicken farm. Additionally, the SNP analysis showed clonal expansion of CRKP and CRKQ strains from well water in different farms, and the clonal CRKP was clonally related to isolates from animal farms and a wastewater treatment plant collected in other studies in the same province. These findings suggest that CRKP and CRKQ are capable of disseminating via horizontal gene transfer and clonal expansion and may pose a significant threat to public health unless preventative measures are taken.

Yang, Chengxia↗

Application of prophage sequence analysis to investigate a disease outbreak involving Salmonella Adjame, a rare serovar and implications for the population structure

Introduction Outbreak investigation of foodborne salmonellosis is hindered when the food source is contaminated by multiple strains of Salmonella , creating difficulties matching an incriminated organism recovered from patients with the specific strain in the suspect food. An outbreak of the rare Salmonella Adjame was caused by multiple strains of the organism as revealed by single-nucleotide polymorphism (SNP) variation. The use of highly discriminatory prophage analysis to characterize strains of Salmonella should enable a more precise strain characterization and aid the investigation of foodborne salmonellosis. Methods We have carried out genomic analysis of S. Adjame strains recovered during the course of a recent outbreak and compared them with other strains of the organism ( n = 38 strains), using SNPs to evaluate strain differences present in the core genome, and prophage sequence typing (PST) to evaluate the accessory genome. Phylogenetic analyses were performed using both total prophage content and conserved prophages. Results The PST analysis of the S. Adjame isolates showed a high degree of strain heterogeneity. We observed small clusters made up of 2-6 isolates ( n = 27) and singletons ( n = 11) in stark contrast with the three clusters observed by SNP analysis. In total, we detected 24 prophages of which only four were highly prevalent, namely: Entero_p88 (36/38 strains), Salmon_SEN34 (35/38 strains), Burkho_phiE255 (33/38 strains) and Edward_GF (28/38 strains). Despite the marked strain diversity seen with prophage analysis, the distribution of the four most common prophages matched the clustering observed using core genome. Discussion Mutations in the core and accessory genomes of S. Adjame have shed light on the evolutionary relationships among the Adjame strains and demonstrated a convergence of the variations observed in both fractions of the genome. We conclude that core and accessory genomes analyses should be adopted in foodborne bacteria outbreak investigations to provide a more accurate strain description and facilitate reliable matching of isolates from patients and incriminated food sources. The outcomes should translate to a better understanding of the microbial population structure and an 46 improved source attribution in foodborne illnesses.

Gao, Ruimin↗

Virulence and Genetic Diversity of Puccinia spp., Causal Agents of Rust on Switchgrass (Panicum virgatum L.) in the USA

Switchgrass (Panicum virgatum L.) is an important cellulosic biofuel grass native to North America. Rust, caused by Puccinia spp. is the most predominant disease of switchgrass and has the potential to impact biomass conversion. In this study, virulence patterns were determined on a set of 38 switchgrass genotypes for 14 single-spore rust isolates from 14 field samples collected in seven states. Single nucleotide polymorphism (SNP) variation was also assessed in 720 sequenced cloned amplicons representing 654 base pairs of the elongation factor 1-α gene from the field samples. Five major haplotypes were identified differing by 11 out of the 39 SNP positions identified. STRUCTURE, Principal Coordinate Analysis, and phylogenetic analyses divided the rust population into two genetic clusters. Virginia and Georgia had the highest and lowest rust genetic diversity, respectively. Only nine accessions showed a differential disease response between the 14 isolates, allowing the identification of eight races, differing by 1–3 virulence factors. Overall, the results suggested clonal reproduction of the pathogen and a North–South differentiation via local adaptation. However, similar haplotypes and races were also recovered from several states, suggesting migration events, and highlighting the need to further investigate the switchgrass rust population structure and evolution in the USA.

Bahri, Bochra A. (ORCID:0000000159055880)↗

Populus_trichocarpa_Breeding_Population_SNPs

These data are from the manuscript “Application of Genomic Prediction in a Populus trichocarpa Breeding Program”, by Brian J. Stanton, David Macaya-Sanz, Chanaka Roshan Abeyratne, David Kainer, Kathy Haiby, Austin Himes, Carlos Gantz, Gerald A. Tuskan, and Stephen P. DiFazio. The data are based on genome resequencing to approximately 10X depth on two collections of Populus trichocarpa trees from Oregon, Washington, California, and British Columbia. The first collection consists of 293 genets collected by Poplar Innovations LLC for a breeding program. The second collection consists of 961 trees collected for the purpose of genome-wide association studies. These genets were sequenced using short, paired-end Illumina sequence reads (Chhetri et al. 2019). Reads were aligned to the P. trichocarpa ′Stettler-14′ reference (Hofmeister et al. 2020), with minor modifications to correct mis-assemblies (Zhou et al. 2020), and variants were called as per methods described in (Abeyratne et al. 2023). Identified variants were filtered using GATK’s VariantFiltration tool (DePristo et al. 2011), with filter expression flag set to “AF < 0.01 || AF > 0.99 || QD < 10.0 || ExcessHet > 20.0 || FS > 10.0 || MQ < 58.0”. SNPs with severe departures from Hardy−Weinberg expectations (exact-test p< 0.01) were also removed using vcftools --hwe flag (Danecek et al. 2011), resulting in 15,627,211 bi-allelic SNPs. The data included here consist of 141,903 high quality bi-allelic genome-wide SNPs obtained by further filtering the original SNP dataset using vcftools with flags --maf 0.05, --max-maf 0.95, --max-missing 0.95, --min-meanDP 10.75, --max-meanDP 43.00, --thin 2000. Collectively, these filtering parameters removed SNPs with 1) a minor allele frequency ≤ 0.05; 2) proportion of missing data for individual loci exceeding 5%; 3) sequencing depth more than 2X mean-depth or less than 0.5X mean-depth; or 4) a distance of

09 BIOMASS FUELS↗

Genetic mapping of sugarcane aphid resistance in sorghum line SC112-14

Sugarcane aphid [Melanaphis sacchari (Zehntner)] is a destructive pest that has had an economic effect on sorghum in North America since 2013. The identification, development, and use of resistant sorghum germplasm is the most feasible strategy to control the pest. Nevertheless, the genetic control of sugarcane aphid (SCA) resistance is unknown for most sorghum resistant lines. To identify the genetic regions that confer SCA resistance in sorghum line SC112-14, 103 recombinant inbred lines (RILs) derived by its cross with the susceptible line PI 609251 were evaluated for their SCA resistance response in Georgia during two consecutive years. The resistance response was determined based on two ratings (2 wk apart) for aphid population size (APS) and aphid-induced plant damage (APD) each year. Segregation for SCA resistance was observed for the first APS and both APD ratings, and the broad-sense heritability estimate ranged from .71 to .76, respectively. A quantitative trait locus analysis using a high-density linkage map of 3,852 single nucleotide polymorphisms (SNPs) detected an 81-kb genomic region on chromosome 6 that explained 50–55% of the phenotypic variation. Comparative mapping analysis found that the resistance locus in SC112-14 is located 8- and 10-cM upstream of the Henong 16 (RMES1) and Tx2783 resistance loci, respectively, and encloses the SNP Sbv3.1_06_2316351 associated in Haitian resistant lines. Therefore, the line SC112-14 is an additional SCA resistance source that can be combined or strategically used with other resistance sources to assure a more robust host plant resistance to the SCA.

60 APPLIED LIFE SCIENCES↗

Genetic dissection of natural variation in oilseed traits of camelina by whole‐genome resequencing and QTL mapping

Abstract Camelina [ Camelina sativa (L.) Crantz] is an oilseed crop in the Brassicaceae family that is currently being developed as a source of bioenergy and healthy fatty acids. To facilitate modern breeding efforts through marker‐assisted selection and biotechnology, we evaluated genetic variation among a worldwide collection of 222 camelina accessions. We performed whole‐genome resequencing to obtain single nucleotide polymorphism (SNP) markers and to analyze genomic diversity. We also conducted phenotypic field evaluations in two consecutive seasons for variations in key agronomic traits related to oilseed production such as seed size, oil content (OC), fatty acid composition, and flowering time. We determined the population structure of the camelina accessions using 161,301 SNPs. Further, we identified quantitative trait loci (QTL) and candidate genes controlling the above field‐evaluated traits by genome‐wide association studies (GWAS) complemented with linkage mapping using a recombinant inbred line (RIL) population. Characterization of the natural variation at the genome and phenotypic levels provides valuable resources to camelina genetic studies and crop improvement. The QTL and candidate genes should assist in breeding of advanced camelina varieties that can be integrated into the cropping systems for the production of high yield of oils of desired fatty acid composition.

59 BASIC BIOLOGICAL SCIENCES↗

Linkage map construction using limited parental genotypic information

Abstract Genetic linkage maps based on single nucleotide polymorphisms (SNPs) represent an essential tool for a variety of genomic analyses. Today, next-generation sequencing (NGS) enables rapid genotyping of different mapping populations based on thousands of SNPs and the construction of highly saturated linkage maps. Nevertheless, missing data in the genotyping of the parental lines creates a bottleneck that determines the number of SNPs that can be used for the linkage map. As a proof of concept, a highly saturated genetic linkage map was constructed using the imputed genotypic data of a recombinant inbred line (RIL) population and the limited genotypic information of its parental lines. Two ABH genotype files were created from a pseudo-parental genotypic data set that includes all the SNPs present in the RIL population. In the first ABH file pseudo-parental 1 was considered parental A, while in the second pseudo-parental 1 was considered parental B. These two duplicate ABH genotype files were merged by chromosome and subjected to linkage map analysis. Since the ABH data were duplicated, two mirrored linkage groups were generated per chromosome. The correct linkage map was identified and selected based on the partial genotypic data of the parental lines. This strategy was effective for constructing a highly saturated linkage map of 33,421 SNPs based on the genotyping of 205 RILs and a limited number of 100 SNPs present in the parental lines. This strategy enables the use of all the NGS SNP data obtained from a low-coverage sequencing experiment in the mapping population.

59 BASIC BIOLOGICAL SCIENCES↗

Genetic architecture of leaf morphological and physiological traits in a Populus deltoides ‘Danhong’ × P. simonii ‘Tongliao1’ pedigree revealed by quantitative trait locus analysis

Understanding the genetic architecture of leaf morphological and physiological traits will help plant breeders develop high biomass poplar genotypes. Quantitative trait locus (QTL) studies combining next-generation sequencing techniques can advance our understanding of the genetic basis of complex traits. In this study, we measured 13 leaf morphological and physiological traits and identified quantitative trait loci (QTLs) in a Populus deltoides ‘Danhong’ × P. simonii ‘Tongliao1’ F1 population (500 progenies) using a high-density genetic map constructed by whole genome re-sequencing. This linkage map consisted of 5796 single nucleotide polymorphism (SNP) markers assigned to 19 linkage groups (LGs), spanning 2683.80 centimorgans (cM) of genetic length, with an average marker density of 0.46 cM. We identified 109 QTLs on 18 LGs for leaf morphological traits and 55 QTLs on 14 LGs for leaf physiological traits. One-hundred eight putative candidate genes were identified within the candidate genomic region. Co-expression network and gene ontology enrichment analyses suggested that these candidate genes were involved in the photosynthetic process. The differential expression patterns of the CYCLIN (Potri.015G112200) and RED CHLOROPHYLL REDUCTASE (Potri.007G043600) genes between two parents indicated their potential roles in leaf morphological and physiological traits. These findings decipher the genetic architecture of leaf morphological and physiological traits in the P. deltoides ‘Danhong’ × P. simonii ‘Tongliao1’ pedigree and provide candidate genes for future poplar genetic improvement.

59 BASIC BIOLOGICAL SCIENCES↗

A compact electromagnetic neutron nutator for precise neutron spin manipulation

A superconducting electromagnetic nutator (EMN) capable of generating a magnetic field vector along an arbitrary direction on a 2D plane has been designed. Its performance in precisely manipulating the neutron polarization vector has been tested at the HB2-D polarized development beamline at the High Flux Isotope Reactor. Unlike mechanical nutators that require physical handling or motor-driven actuation to rotate the magnetic field, the magnitude and orientation of the magnetic field produced by the EMN can be controlled electromagnetically. Further, the compact design (~15 mm depth, not including cryogenic housing) of this device ensures ease of coupling within existing superconducting neutron spin manipulation devices, such as magnetic Wollaston prisms (MWP), resonant radio frequency (RF) flippers, spherical neutron polarimetry (SNP) devices, etc.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Unraveling plant phenotype to genotype associations with daily hyperspectral traits in Populus trichocarpa

Hyperspectral remote sensing is a powerful, high-throughput phenotyping tool that quantifies physiologically and structurally relevant wavelengths across diverse genotypes and over varying temporal scales. In this study, we combined tower-based continuous hyperspectral sensing with genome-wide association studies to analyze 1423 wavebands (400-900 nm) and derivative vegetation indices across 505 genotypes and the genetic architecture of hyperspectral phenotypes over time in Populus trichocarpa Torr. & Gray grown under field conditions. Wavelengths related to chlorophyll and carotenoid absorption spectra exhibited the strongest genetic variation resulting in 98 significant SNP associations. Notably, we found substantial overlap in genetic association between the blue and red spectral regions, indicative of carotenoids and chlorophyll, respectively, and identified more than 10 candidate genes associated with chloroplast function, underpinning photosynthetic activity. Furthermore, fluctuations in associations for vegetative indices, such as the chlorophyll:carotenoid index (CCI), across the growing season reveal a temporally dynamic genetic architecture of physiological traits associated with fall senescence of this temperate tree species. Finally, we also observed correlations (spearman rho = 0.3, p < 1x10 −8 ) between individual wavebands or vegetative indices and growth rate, assessed as the relative change of tree height over the growing season. The growth rate prediction was substantially improved by a regularization multivariate model (spearman rho>0.5, p < 1x10 −16 ), reinforcing the value of hyperspectral measurements for predicting traits linked to tree productivity. These findings highlight the potential of high-throughput, rapid, hyperspectral genome wide association studies GWAS to uncover physiologically meaningful genetic variation and offer promising insights for future acceleration for plant breeding.

09 BIOMASS FUELS↗

A Multiplexed Quantitative Analysis of Germline Single Amino Acid Variants by Targeted Proteomics in Nondepleted Human Plasma

Single amino acid variants (SAAVs) in protein sequences are often a direct result of single-nucleotide polymorphisms (SNPs). Certain germline SAAVs have shown biological relevance in different disease conditions but lack precise quantification in circulation, which could hinder functional investigations and progress in biomarker development. Here, we have developed a multiplexed liquid chromatography-selected reaction monitoring (LC-SRM) assay that monitors 5 wild-type and variant peptide pairs (Complement Factor B: CFB-R32Q/R32W, Clusterin: CLU-N317H, Fetuin B: FETUB-K360R, and Kininogen: KNG1-L212P) in nondepleted human plasma. The assay was optimized for imprecision, linearity, stability, and calibration assessments with CVs of under 20%. The wild-type and variant peptide pairs were characterized in a set of healthy individual plasma samples. These target identifications were also validated by SNP genotyping with more than 99% accuracy. For all protein targets, we observed significantly lower concentrations of WT species in the presence variant peptides. In CFB, the concentration of R32Q was significantly lower than its counterpart R32W variant and WT species. Furthermore, our results distinguished phenotypes of homozygosity and heterozygosity of the SAAV presence through direct concentration level characterization. These findings provide some insights into how SAAVs affect quantitative assessments of target peptides. The assay demonstrates a platform for proteogenomic analyses with potential applications in both research and clinical settings.

genetics↗

Fast and accurate metagenotyping of the human gut microbiome with GT-Pro

Single nucleotide polymorphisms (SNPs) in metagenomics are used to quantify population structure, track strains and identify genetic determinants of microbial phenotypes. However, existing alignment-based approaches for metagenomic SNP detection require high-performance computing and enough read coverage to distinguish SNPs from sequencing errors. To address these issues, we developed the GenoTyper for Prokaryotes (GT-Pro), a suite of methods to catalog SNPs from genomes and use unique k-mers to rapidly genotype these SNPs from metagenomes. Compared to methods that use read alignment, GT-Pro is more accurate and two orders of magnitude faster. Here, using high-quality genomes, we constructed a catalog of 104 million SNPs in 909 human gut species and used unique k-mers targeting this catalog to characterize the global population structure of gut microbes from 7,459 samples. GT-Pro enables fast and memory-efficient metagenotyping of millions of SNPs on a personal computer.

59 BASIC BIOLOGICAL SCIENCES↗

A genomic data archive from the Network for Pancreatic Organ donors with Diabetes

The Network for Pancreatic Organ donors with Diabetes (nPOD) is the largest biorepository of human pancreata and associated immune organs from donors with type 1 diabetes (T1D), maturity-onset diabetes of the young (MODY), cystic fibrosis-related diabetes (CFRD), type 2 diabetes (T2D), gestational diabetes, islet autoantibody positivity (AAb+), and without diabetes. nPOD recovers, processes, analyzes, and distributes high-quality biospecimens, collected using optimized standard operating procedures, and associated de-identified data/metadata to researchers around the world. Herein describes the release of high-parameter genotyping data from this collection. 372 donors were genotyped using a custom precision medicine single nucleotide polymorphism (SNP) microarray. Data were technically validated using published algorithms to evaluate donor relatedness, ancestry, imputed HLA, and T1D genetic risk score. Additionally, 207 donors were assessed for rare known and novel coding region variants via whole exome sequencing (WES). These data are publicly-available to enable genotype-specific sample requests and the study of novel genotype:phenotype associations, aiding in the mission of nPOD to enhance understanding of diabetes pathogenesis to promote the development of novel therapies.

59 BASIC BIOLOGICAL SCIENCES↗