Engineering PapersSearch

DOE OSTI · 2477164

Machine Learning Framework for Conotoxin Class and Molecular Target Prediction

Abstract

Conotoxins are small and highly potent neurotoxic peptides derived from the venom of marine cone snails which have captured the interest of the scientific community due to their pharmacological potential. These toxins display significant sequence and structure diversity, which results in a wide range of specificities for several different ion channels and receptors. Despite the recognized importance of these compounds, our ability to determine their binding targets and toxicities remains a significant challenge. Predicting the target receptors of conotoxins, based solely on their amino acid sequence, remains a challenge due to the intricate relationships between structure, function, target specificity, and the significant conformational heterogeneity observed in conotoxins with the same primary sequence. We have previously demonstrated that the inclusion of post-translational modifications, collisional cross sections values, and other structural features, when added to the standard primary sequence features, improves the prediction accuracy of conotoxins against non-toxic and other toxic peptides across varied datasets and several different commonly used machine learning classifiers. Here, we present the effects of these features on conotoxin class and molecular target predictions, in particular, predicting conotoxins that bind to nicotinic acetylcholine receptors (nAChRs). We also demonstrate the use of the Synthetic Minority Oversampling Technique (SMOTE)-Tomek in balancing the datasets while simultaneously making the different classes more distinct by reducing the number of ambiguous samples which nearly overlap between the classes. In predicting the alpha, mu, and omega conotoxin classes, the SMOTE-Tomek PCA PLR model, using the combination of the SS and P feature sets establishes the best performance with an overall accuracy (OA) of 95.95%, with an average accuracy (AA) of 93.04%, and an f1 score of 0.959. Using this model, we obtained sensitivities of 98.98%, 89.66%, and 90.48% when predicting alpha, mu, and omega conotoxin classes, respectively. Similarly, in predicting conotoxins that bind to nAChRs, the SMOTE-Tomek PCA SVM model, which used the collisional cross sections (CCSs) and the P feature sets, demonstrated the highest performance with 91.3% OA, 91.32% AA, and an f1 score of 0.9131. The sensitivity when predicting conotoxins that bind to nAChRs is 91.46% with a 91.18% sensitivity when predicting conotoxins that do not bind to nAChRs.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Truong, Duc P., Monroe, Lyman K., Williams, Robert F., Nguyen, Hau B.. 2024-11-03. Machine Learning Framework for Conotoxin Class and Molecular Target Prediction. https://doi.org/10.3390/toxins16110475

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Soil metagenomics umbrella narrative

Implementing accessible, authentic research experiences in introductory courses is challenging, particularly at institutions serving diverse student populations. To address this gap, we developed and deployed a Course-based Undergraduate Research Experience (CURE) focused on plant-microbe interactions in General Biology II at Northeastern Illinois University (NEIU), a minority-serving institution with a diverse student body. Students grew sugar beets (Beta vulgaris), extracted DNA from the rhizoplane, and used the Department of Energy Systems Biology Knowledgebase (KBase) for bioinformatic analysis to compare microbial relative abundance in fertilized versus unfertilized soil. Over five semesters, the CURE engaged 103 students and leveraged the intuitive KBase platform to make complex sequencing data accessible. Pre/post-course survey data revealed significant increases in student self-assessed research skills, including the ability to explain results and determine the types of data to collect. Furthermore, students reported significant gains in confidence related to experimental design and hypothesis development, alongside a strong increase in familiarity with KBase. Informal faculty feedback indicated high student engagement and appreciation for the real-world connections (e.g. food systems, agriculture, and health). This scalable, low-cost model effectively integrates data science tools into the foundational curriculum, demonstrating a potent strategy for boosting research skills and broadening participation in authentic scientific inquiry among diverse undergraduate students.

59 BASIC BIOLOGICAL SCIENCES

Genome-resolved insights into microbial diversity and elemental cycling in Winogradsky columns

We retained 18 MAGs with ≥50% completion and <10% contamination (i.e., at least medium quality). Of these, 10 had >90% completion and <5% contamination; however, only one (Paceibacteria Bin.003_MG) can be described as high-quality, as the others lacked a full suite of 5S, 16S, and 23S rRNA genes. To maximize the diversity of our recovered MAGs, we also retained one MAG (Chromatiaceae Bin.008_AM) with >40% (but less than 50%) completion and <5% contamination, as well as one (Rhodopseudomonas Bin.015_MK) with >90% completion and <20% (but>10%) contamination. Interestingly, significant chimerism was not detected in this MAG (40) , suggesting that the elevated contamination (20%) may instead reflect two closely related strains collapsing into a single bin. Consistent with this, contig coverage was bimodal, with roughly 17% of the assembly at ~115x and the remaining 83% at ~282x, while GC content remained uniform across both groups (~64%), arguing against contamination from a taxonomically distinct source.

59 BASIC BIOLOGICAL SCIENCES