Engineering Papers⌕ Search

Engineering topics

Degnan, David J. (ORCID:0000000157377173)

Publications and source records attributed to Degnan, David J. (ORCID:0000000157377173).

Characterization of the biofilm landscape of Bacillus subtilis by spatial microproteomics

Bulk proteomics has been demonstrated to differentiate subpopulations within bacterial colonies, yet advanced analyses by mass spectrometry imaging (MSI) hold even greater promise for the future. This technology can enable high-throughput spatial phenotyping that can reshape biological discovery by providing visualization of components of various biomolecular mechanisms. With high mass resolving power and high spatial resolution analyses being routine, we can confidently enable intact protein imaging directly from samples with minimal preparation. Pairing those analyses with bulk experimental libraries can provide high confidence in annotations of post-translational modifications (PTMs) and truncations. Revealing PTM localization within the samples unlocks a direct window into unknown biology at the microscale. However, top-down proteomics (TDP) is not commonplace for microbial species, largely due to challenges in identifying detected peptides and proteins; considering the theoretical proteome of even the well-studied model bacterium Bacillus subtilis was only partially mapped recently. With little still known about the form and function of many of these proteins – let alone proteoforms, where PTMs and truncations of the same protein may possess unique physiological roles – there is a wealth of work to be done. Here we jointly apply TDP and MSI to describe the microscale spatial proteomic landscape within B. subtilis and further demonstrate the feasibility of detecting differentiated subpopulations through proteoforms across the biofilm landscape.

bacterial biofilms↗

Computationally efficient Bayesian estimation of graphical networks for omics data

Graphical networks are useful, widely-used modeling approaches to represent complex biological processes with biological measurements generated by platforms such as mass spectrometry. Bayesian analyses of graphical networks for omics data have several advantages over their frequentist counterparts, such as the inclusion of prior knowledge in the estimation of models. However, Bayesian approaches to date have only been feasible for data with a couple hundred biomolecules due to prohibitive computational time, but omics data often contains tens of thousands of biomolecules. Here, we present and illustrate a more computationally efficient approach named BPlane (Bayesian PseudoLikelihood-based Algorithm for Network Estimation) to extend Bayesian modeling capabilities for larger-sized datasets, such as most untargeted proteomics data. Via simulation, we demonstrate that BPlane produces substantial computational savings over a current state-of-the-art Bayesian algorithm while maintaining competitive edge detection accuracy. On a SARS-CoV2 proteomics data with 7000 proteins, the competing algorithm takes three times as long to complete the first iteration as BPlane takes to converge after over 100 iterations.

EM algorithm↗

OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data

Expression-based omics technologies (e.g. proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application SLOPE was designed to lower the barrier to ML modeling for omics data. SLOPE supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics datasets, such as proteomics, metabolomics, lipidomics, and transcriptomics. SLOPE offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user’s data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, SLOPE address critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, SLOPE is applied to data from a lignin exposure study to highlight the workflow for fitting both supervised and unsupervised models to data.

lipidomics↗