Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “proteome”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

PPI DataHub Project Data Package: S. elongatus PCC 7942 Limited Proteolysis and Thermal Proteome Profiling Structural Proteomics (JM-PB-DP3)

The purpose of this experiment was to investigate structural alterations in proteins involved in central carbon metabolism and photosynthetic electron transfer pathways in Synechococcus elongatus PCC 7942. Sample data was obtained from S. elongatus cell lysates using three complementary mass spectrometry (MS) techniques using limited proteolysis (LiP-MS), thermal proteome profiling (TPP-MS), and redox enrichment (Redox-MS) in evaluating alterations solvent accessibility and structural stability caused by light perturbation at the molecular level. Experimentally processed sample data for LiP and TPP proteomic datasets were derived from the same cell culture stock, prepared simultaneously in parallel, and acquired by mass spectrometry. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files, computed outputs, and supporting metadata materials. Experimental samples processed for LiP-MS label-free quantification (LFQ) or TPP-MS tandem mass tag (TMT) 10-plex were acquired using a Q-Exactive HF-X mass spectrometer and processed/compiled using either MSGF+ (v2024.03.26) or ​​​​PlexedPiper for proteome evaluation. Additional software supporting downstream proteomic analysis include FragPipe (v.4.0), MSFragger (v.22.1), and an adapted Microbial Isolate LiP Analysis Workflow (located at Zenodo). Processed proteomic data downloads include a sample naming key, normalized quantification results files, and processed protein annotated abundance files.

59 BASIC BIOLOGICAL SCIENCES↗

The promising role of proteomes and metabolomes in defining the single-cell landscapes of plants

The plant community has a strong track-record of RNA sequencing technology deployment, which combined with the recent advent of spatial platforms (e.g., 10x genomics), has resulted in an explosion of outstanding single cell and nuclei datasets that can be put in an in situ context within tissues (e.g., a cell atlas)1. In the genomics era, application of proteomics technologies in the plant sciences has always trailed behind that of RNA sequencing technologies, largely due to accessibility, ease-of-use and access to expertise along with depth of analysis benefits. On the other hand, the use of early analytical tools for characterizing small molecules (metabolites) from plant systems predates nucleic acid sequencing and proteomics analysis2, as the search for plant-based natural products has played a significant role in improving human health throughout history. However, the employment of proteomics and metabolomics assays for characterizing plant cell processes now remains significantly behind transcriptional approaches, even though both provide a direct functional readout of cell states and phenotypes.

Anderton, Christopher R. [BATTELLE (PACIFIC NW LAB↗

Predicted structural proteome of Sphagnum divinum and proteome-scale annotation

Sphagnum-dominated peatlands store a substantial amount of terrestrial carbon. The genus is undersampled and under-studied. No experimental crystal structure from any Sphagnum species exists in the Protein Data Bank and fewer than 200 Sphagnum-related genes have structural models available in the AlphaFold Protein Structure Database. Tools and resources are needed to help bridge these gaps, and to enable the analysis of other structural proteomes now made possible by accurate structure prediction. We present the predicted structural proteome (25,134 primary transcripts) of Sphagnum divinum computed using AlphaFold, structural alignment results of all high-confidence models against an annotated nonredundant crystallographic database of over 90,000 structures, a structure-based classification of putative Enzyme Commission (EC) numbers across this proteome, and the computational method to perform this proteome-scale structure-based annotation.

59 BASIC BIOLOGICAL SCIENCES↗

Proteomics Analysis of Human Contaminant Proteins

Complete characterization of unknowns via proteomics remains challenging. There exist regions of mass spectrometry-based proteomics data where empirical measurements are not attributed to peptides, and/or sequenced peptides from mass spectra are not attributed to any source. These uncharacterized regions are known as the “dark” proteome. Many proteomics tools rely on some a priori knowledge of sample composition; few tools allow for investigation of unknowns without relying on composition assumptions. Further, the potential low abundance of minor traces in these uncharacterized regions can make elucidation of the “dark” proteome challenging. Herein, we describe the development and evaluation of approaches to study the “dark” proteome and move towards an untargeted approach for more complete characterization, namely by studying minor human protein traces in non-human samples and combining that approach with non-human source organism identification without relying on assumptions. Human protein markers, in the form of genetically variant peptides, have been extensively examined in a variety of human matrices, including blood, plasma, and hair, but have yet to be investigated in non-human samples, such as cell cultures, as human contaminant traces. Genetically variant peptides are those that are found in proteins carrying single nucleotide polymorphisms. In this work, we aimed to (1) investigate the feasibility of detecting human contaminant genetically variant peptides (GVPs) in a diverse set of non-human organisms using public proteomics data and a computational pipeline, as well as to (2) develop a combined capability for untargeted source organism characterization and GVP detection. To our knowledge, this is the first report of applying these approaches towards a more complete proteomic characterization of unknowns. We successfully demonstrate the feasibility of broad human contaminant GVP detection in proteomics data, develop a better understanding of GVP detectability, characterize the sample-to-sample variability in GVP detection, and identify a core set of GVPs that can potentially be used as markers indicative of the human contaminant traces portion of the “dark” proteome. Further, we developed and evaluated a combined pipeline, MARLOWE-GVP, that enables both untargeted source organism characterization and GVP detection. We show high accuracy of correct source organism characterization and high degree of similarity of human contaminant GVP detection compared to the conventional approach. Success on both these efforts have allowed us to advance our understanding and characterization of the “dark” proteome.

59 BASIC BIOLOGICAL SCIENCES↗

Robust collection and processing for label-free single voxel proteomics

With advanced mass spectrometry (MS)-based proteomics, genome-scale proteome coverage can be achieved from bulk tissues. However, such bulk measurement lacks spatial resolution and obscures tissue heterogeneity, precluding proteome mapping of tissue microenvironment. Here we report an integrated $\underline{w}et$ $\underline{c}ollection$ of single microscale tissue voxels and $\underline{S}urfactant$$-assisted$ $\underline{O}ne$-$\underline{P}ot$ voxel processing method termed wcSOP for robust label-free single voxel proteomics. wcSOP capitalizes on buffer droplet-assisted wet collection of a single voxel dissected by LCM into the PCR tube cap and MS-compatible surfactant-assisted one-pot voxel processing in the collection cap. This convenient method allows reproducible label-free quantification of ~900 and ~4,600 proteins for single voxels at 20 µm × 20 µm × 10 µm (close to single cells) and 200 µm × 200 µm × 10 µm (~100 cells) from fresh frozen human spleen tissue, respectively. 100s-1000s of protein signatures were spatially resolved between spleen red and white pulp regions depending on the voxel size. Region-specific signaling pathways were enriched from single voxel proteomics data. To evaluate its broad applicability, we applied wcSOP-MS to two commonly accessible, OCT-embedded and FFPE, human archived tissues. It enabled to identify spatially resolved proteome changes and enriched pathways between diseased (breast cancer tumor or AD amyloid plaque) and adjacent normal regions. Antibody-based CODEX and IHC imaging validated label-free MS quantitation for single voxel analysis. The wcSOP-MS method paves the way for routine robust single voxel proteomics and spatial proteomics.

59 BASIC BIOLOGICAL SCIENCES↗

Identification of integrated proteomics and transcriptomics signature of alcohol-associated liver disease using machine learning

Distinguishing between alcohol-associated hepatitis (AH) and alcohol-associated cirrhosis (AC) remains a diagnostic challenge. In this study, we used machine learning with transcriptomics and proteomics data from liver tissue and peripheral mononuclear blood cells (PBMCs) to classify patients with alcohol-associated liver disease. The conditions in the study were AH, AC, and healthy controls. We processed 98 PBMC RNAseq samples, 55 PBMC proteomic samples, 48 liver RNAseq samples, and 53 liver proteomic samples. First, we built separate classification and feature selection pipelines for transcriptomics and proteomics data. The liver tissue models were validated in independent liver tissue datasets. Next, we built integrated gene and protein expression models that allowed us to identify combined gene-protein biomarker panels. For liver tissue, we attained 90% nested-cross validation accuracy in our dataset and 82% accuracy in the independent validation dataset using transcriptomic data. We attained 100% nested-cross validation accuracy in our dataset and 61% accuracy in the independent validation dataset using proteomic data. For PBMCs, we attained 83% and 89% accuracy with transcriptomic and proteomic data, respectively. The integration of the two data types resulted in improved classification accuracy for PBMCs, but not liver tissue. We also identified the following gene-protein matches within the gene-protein biomarker panels: CLEC4M-CLC4M, GSTA1-GSTA2 for liver tissue and SELENBP1-SBP1 for PBMCs. In this study, machine learning models had high classification accuracy for both transcriptomics and proteomics data, across liver tissue and PBMCs. The integration of transcriptomics and proteomics into a multi-omics model yielded improvement in classification accuracy for the PBMC data. The set of integrated gene-protein biomarkers for PBMCs show promise toward developing a liquid biopsy for alcohol-associated liver disease.

60 APPLIED LIFE SCIENCES↗

Robust surfactant-assisted one-pot sample preparation for label-free single-cell and nanoscale proteomics

With advanced mass spectrometry (MS)-based proteomics, genome-scale proteome coverage can be achieved from bulk cells. However, such bulk measurement obscures cell to cell heterogeneity, precluding proteome profiling of single cells and small numbers of cells of interest. To address this issue, in recent 5 years there are a surge of small sample preparation methods developed for robust effective collection and processing of single cells and small numbers of cells for in-depth MS-based proteome profiling. Based on their broad accessibility, they can be categorized into two types: specific device- and standard PCR tube- or multi-well plate-based methods. Herein we describe the detailed protocol of our recently developed, easily adoptable, Surfactant-assisted One-Pot (SOP) sample preparation coupled with MS method termed SOP-MS for label-free single-cell and nanoscale proteomics. SOP-MS capitalizes on the combination of a MS-compatible surfactant, DDM (n-Dodecyl-ß-D-maltoside), and standard low-bind PCR tube or multi-well plate for ‘all-in-one’ one-pot sample preparation without sample transfer. With its robust and convenient features, SOP-MS can be readily implemented in any MS laboratory for single-cell and nanoscale proteomics. With further improvements in MS detection sensitivity and sample throughput, we believe that SOP-MS could open an avenue for single-cell proteomics with broad applicability in the biological and biomedical research.

Single-cell proteomics, nanoscale proteomics, SOP-↗

Determining protein polarization proteome-wide using physical dissection of individual Stentor coeruleus cells

Cellular components are non-randomly arranged with respect to the shape and polarity of the whole cell. Patterning within cells can extend down to the level of individual proteins and mRNA. But how much of the proteome is actually localized with respect to cell polarity axes? Proteomics combined with cellular fractionation has shown that most proteins localize to one or more organelles but does not tell us how many proteins have a polarized localization with respect to the large-scale polarity axes of the intact cell. Genome-wide localization studies in yeast found that only a few percent of proteins have a localized position relative to the cell polarity axis defined by sites of polarized cell growth. Here, we describe an approach for analyzing protein distribution within a cell with a visibly obvious global patterning - the giant ciliate Stentor coeruleus. Ciliates, including Stentor, have highly polarized cell shapes with visible surface patterning. A Stentor cell is roughly 2mmlong, allowing a ‘‘proteomic dissection’’ in which microsurgery is used to separate cellular fragments along the anterior-posterior axis, followed by comparative proteomic analysis. In our analysis, 25% of the proteome, including signaling proteins, centrin/SFI proteins, and GAS2 orthologs, shows a polarized location along the cell’s anterior-posterior axis. We conclude that a large proportion of all proteins are polarized with respect to global cell polarity axes and that proteomic dissection provides a simple and effective approach for spatial proteomics.

59 BASIC BIOLOGICAL SCIENCES↗

Spatial top-down proteomics for the functional characterization of human kidney

Background: The Human Proteome Project has credibly detected nearly 93% of the roughly 20,000 proteins which are predicted by the human genome. However, the proteome is enigmatic, where alterations in amino acid sequences from polymorphisms and alternative splicing, errors in translation, and post-translational modifications result in a proteome depth estimated at several million unique proteoforms. Recently mass spectrometry has been demonstrated in several landmark efforts mapping the human proteoform landscape in bulk analyses. Herein, we developed an integrated workflow for characterizing proteoforms from human tissue in a spatially resolved manner by coupling laser capture microdissection, nanoliter-scale sample preparation, and mass spectrometry imaging. Results: Using healthy human kidney sections as the case study, we focused our analyses on the major functional tissue units including glomeruli, tubules, and medullary rays. After laser capture microdissection, these isolated functional tissue units were processed with microPOTS (microdroplet processing in one-pot for trace samples) for sensitive top-down proteomics measurement. This provided a quantitative database of 616 proteoforms that was further leveraged as a library for mass spectrometry imaging with near-cellular spatial resolution over the entire section. Notably, several mitochondrial proteoforms were found to be differentially abundant between glomeruli and convoluted tubules, and further spatial contextualization was provided by mass spectrometry imaging confirming unique differences identified by microPOTS, and further expanding the field-of-view for unique distributions such as enhanced abundance of a truncated form (1-74) of ubiquitin within cortical regions. Conclusions: We developed an integrated workflow to directly identify proteoforms and reveal their spatial distributions. Where of the 20 differentially abundant proteoforms identified as discriminate between tubules and glomeruli by microPOTS, the vast majority of tubular proteoforms were of mitochondrial origin (8 of 10) where discriminate proteoforms in glomeruli were primarily hemoglobin subunits (9 of 10). These trends were also identified within ion images demonstrating spatially resolved characterization of proteoforms that has the potential to reshape discovery-based proteomics because the proteoforms are the ultimate effector of cellular functions. Applications of this technology have the potential to unravel etiology and pathophysiology of disease states, informing on biologically active proteoforms, which remodel the proteomic landscape in chronic and acute disorders.

59 BASIC BIOLOGICAL SCIENCES↗

Proteomes reveal metabolic capabilities of Yarrowia lipolytica for biological upcycling of polyethylene into high-value chemicals

ABSTRACT Polyolefins derived from plastic wastes are recalcitrant for biological upcycling. However, chemical depolymerization of polyolefins can generate depolymerized plastic (DP) oil, comprising of a complex mixture of saturated, unsaturated, even, and odd hydrocarbons suitable for biological conversion. While DP oil contains a rich carbon and energy source, it is inhibitory to cells. Understanding and harnessing robust metabolic capabilities of microorganisms to upcycle the hydrocarbons in DP oil, both naturally and unnaturally occurring, into high-value chemicals are limited. Here, we discovered that an oleaginous yeast, Yarrowia lipolytica, undergoing short-term adaptation to DP oil robustly utilized a wide range of hydrocarbons for cell growth and production of citric acid and neutral lipids. When growing on hydrocarbons, Y. lipolytica partitioned into planktonic and oil-bound cells with each exhibiting distinct proteomes and amino acid distributions invested in establishing these proteomes. Significant proteome reallocation toward energy and lipid metabolism, belonging to 2 of the 23 Eukaryotic Orthologous Groups classes C and I, enabled robust growth of Y. lipolytica on hydrocarbons, with n-hexadecane as the preferential substrate. This investment was even higher for growth on DP oil where classes C and I were ranked one and two, respectively, and many associated proteins and pathways were expressed and upregulated including the hydrocarbon degradation pathway, Krebs cycle, glyoxylate shunt and, unexpectedly, propionate metabolism. However, a reduction in proteome allocation for protein biosynthesis, at the expense of the observed increase toward energy and lipid metabolisms, might have caused the inhibitory effect of DP oil on cell growth. IMPORTANCE Sustainable processes for biological upcycling of plastic wastes in a circular bioeconomy are needed to promote decarbonization and reduce environmental pollution due to increased plastic consumption, incineration, and landfill storage. Strain characterization and proteomic analysis revealed the robust metabolic capabilities of Yarrowia lipolytica to upcycle polyethylene into high-value chemicals. Significant proteome reallocation toward energy and lipid metabolisms was required for robust growth on hydrocarbons with n-hexadecane as the preferential substrate. However, an apparent over-investment in these same categories to utilize complex depolymerized plastic (DP) oil came at the expense of protein biosynthesis, limiting cell growth. Taken together, this study elucidates how Y. lipolytica activates its metabolism to utilize DP oil and establishes Y. lipolytica as a promising host for the upcycling of plastic wastes.

1-alkenes↗

Evaluating Linear Ion Trap for MS3-Based Multiplexed Single-Cell Proteomics

There is a growing demand to develop high-throughput and high-sensitivity mass spectrometry methods for single-cell proteomics. The commonly used isobaric labeling-based multiplexed single-cell proteomics approach suffers from distorted protein quantification due to co-isolated interfering ions during MS/MS fragmentation, also known as ratio compression. We reasoned that the use of MS3-based quantification could mitigate ratio compression and provide better quantification. However, previous studies indicated reduced proteome coverages in the MS3 method, likely due to long duty cycle time and ion losses during multilevel ion selection and fragmentation. Here, in this paper, we described an improved MS acquisition method for MS3-based single-cell proteomics by employing a linear ion trap to measure reporter ions. We demonstrated that linear ion trap can increase the proteome coverages for single-cell-level peptides with even higher gain obtained via the MS3 method. The optimized real-time search MS3 method was further applied to study the immune activation of single macrophages. Among a total of 126 single cells studied, over 1200 and 1000 proteins were quantifiable when at least 50 and 75% nonmissing data were required, respectively. Our evaluation also revealed several limitations of the low-resolution ion trap detector for multiplexed single-cell proteomics and suggested experimental solutions to minimize their impacts on single-cell analysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Coupling Microdroplet-Based Sample Preparation, Multiplexed Isobaric Labeling, and Nanoflow Peptide Fractionation for Deep Proteome Profiling of the Tissue Microenvironment

There is increasing interest in developing in-depth proteomic approaches for mapping tissue heterogeneity in a cell-type-specific manner to better understand and predict the function of complex biological systems such as human organs. Existing spatially resolved proteomics technologies cannot provide deep proteome coverage due to limited sensitivity and poor sample recovery. Herein, we seamlessly combined laser capture microdissection with a low-volume sample processing technology that includes a microfluidic device named microPOTS (microdroplet processing in one pot for trace samples), multiplexed isobaric labeling, and a nanoflow peptide fractionation approach. The integrated workflow allowed us to maximize proteome coverage of laser-isolated tissue samples containing nanogram levels of proteins. We demonstrated that the deep spatial proteomics platform can quantify more than 5000 unique proteins from a small-sized human pancreatic tissue pixel (∼60,000 μm2) and differentiate unique protein abundance patterns in pancreas. Furthermore, the use of the microPOTS chip eliminated the requirement for advanced microfabrication capabilities and specialized nanoliter liquid handling equipment, making it more accessible to proteomic laboratories.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A seed-like proteome in oil-rich tubers

There are numerous examples of plant organs or developmental stages that are desiccation-tolerant and can withstand extended periods of severe water loss. One prime example are seeds and pollen of many spermatophytes. However, in some plants, also vegetative organs can be desiccation-tolerant. One example are the tubers of yellow nutsedge ( Cyperus esculentus ), which also store large amounts of lipids similar to seeds. Interestingly, the closest known relative, purple nutsedge ( Cyperus rotundus ), generates tubers that do not accumulate oil and are not desiccation-tolerant. We generated nanoLC-MS/MS-based proteomes of yellow nutsedge in five replicates of four stages of tuber development and compared them to the proteomes of roots and leaves, yielding 2257 distinct protein groups. Our data reveal a striking upregulation of hallmark proteins of seeds in the tubers. A deeper comparison to the tuber proteome of the close relative purple nutsedge ( C. rotundus ) and a previously published proteome of Arabidopsis seeds and seedlings indicates that indeed a seed-like proteome was found in yellow but not purple nutsedge. This was further supported by an analysis of the proteome of a lipid droplet-enriched fraction of yellow nutsedge, which also displayed seed-like characteristics. One reason for the differences between the two nutsedge species might be the expression of certain transcription factors homologous to ABSCISIC ACID INSENSITIVE3, WRINKLED1, and LEAFY COTYLEDON1 that drive gene expression in Arabidopsis seed embryos.

59 BASIC BIOLOGICAL SCIENCES↗

Proteomic Assessment of Fluid Shifts and Association with Visual Impairment and Intracranial Pressure in Twin Astronauts

BACKGROUND: Astronauts participating in long duration space missions are at an increased risk of physiological disruptions. The development of visual impairment and intracranial pressure (VIIP) syndrome is one of the leading health concerns for crew members on long-duration space missions; microgravity-induced fluid shifts and chronic elevated cabin CO2 may be contributing factors. By studying physiological and molecular changes in one identical twin during his 1-year ISS mission and his ground-based co-twin, this work extends a current NASA-funded investigation to assess space flight induced "Fluid Shifts" in association with the development of VIIP. This twin study uniquely integrates physiological and -omic signatures to further our understanding of the molecular mechanisms underlying space flight-induced VIIP. We are: (i) conducting longitudinal proteomic assessments of plasma to identify fluid regulation-related molecular pathways altered by long-term space flight; and (ii) integrating physiological and proteomic data with genomic data to understand the genomic mechanism by which these proteomic signatures are regulated. PURPOSE: We are exploring proteomic signatures and genomic mechanisms underlying space flight-induced VIIP symptoms with the future goal of developing early biomarkers to detect and monitor the progression of VIIP. This study is first to employ a male monozygous twin pair to systematically determine the impact of fluid distribution in microgravity, integrating a comprehensive set of structural and functional measures with proteomic, metabolomic and genomic data. This project has a broader impact on Earth-based clinical areas, such as traumatic brain injury-induced elevations of intracranial pressure, hydrocephalus, and glaucoma. HYPOTHESIS: We predict that the space-flown twin will experience a space flight-induced alteration in proteins and peptides related to fluid balance, fluid control and brain injury as compared to his pre-flight protein/peptide signatures. Conversely, the trajectory of these protein signatures will remain relatively constant in his ground based co-twin. METHODS: We are using proteomic and standard immunoelectrophoresis techniques to delineate the change in protein signatures throughout the course of a long duration space flight in relation to the development of VIIP. We are also applying a novel cell-based metaboloic organ system assay ("Organs on a Plate") to address how these circulating biomarkers affect physiological processes at the cellular and organ level which could result in VIIP symptoms. These molecular data will be correlated with physiological measures (eg. extra and intracellular fluid volume, vascular filling/flow patterns, MRI, and Optic Coherence Tomography. DISCUSSION: Pre- and in-flight data collection is in progress for the space-flown twin, and similar data have been obtained from the ground-based twin. Biosamples will be batch processed when received from ISS after the conclusion of the 1-year mission. Omic and Physiological measures from the twin astronauts will be compared to similar data being collected on twin subjects who participated in simulated microgravity study. bed rest study.

Rana, Brinda K.↗

Thiol redox proteomics: Characterization of thiol-based post-translational modifications

Redox post-translational modifications on cysteine thiols (redox PTMs) have profound effects on protein structure and function, thus enabling regulation of various biological processes. Redox proteomics approaches aim to characterize the landscape of redox PTMs at the systems level. These approaches facilitate studies of condition-specific, dynamic processes implicating redox PTMs and have furthered our understanding of redox signaling and regulation. Mass spectrometry (MS) is a powerful tool for such analyses which has been demonstrated by significant advances in redox proteomics during the last decade. A group of well-established approaches involves the initial blocking of free thiols followed by selective reduction of oxidized PTMs and subsequent enrichment for downstream detection. Alternatively, novel chemoselective probe-based approaches have been developed for various redox PTMs. Direct detection of redox PTMs without any enrichment has also been demonstrated given the sensitivity of contemporary MS instruments. This review discusses the general principles behind different analytical strategies and covers recent advances in redox proteomics. Several applications of redox proteomics are also highlighted to illustrate how large-scale redox proteomics data can lead to novel biological insights.

59 BASIC BIOLOGICAL SCIENCES↗

MARLOWE: An Untargeted Proteomics, Statistical Approach to Taxonomic Classification for Forensics

General proteomics research for fundamental science typically addresses laboratory- or patient-derived samples of known origin and composition. However, in a few research areas, such as environmental proteomics, clinical identification of infectious organisms, archeology, art/cultural history, and forensics, attributing the origin of a protein-containing sample to the organisms that produced it is a central focus. A small number of groups have approached this problem and developed software tools for taxonomic characterization and/or identification using bottom-up proteomics. Most such tools identify peptides via database search, and many rely on organism-specific peptides as markers. Our group recently introduced MARLOWE, a software tool for taxonomic characterization of unknown samples based on de novo peptide identification and signal-erosion-resistant strong peptides, which are shared peptides distributed in a taxonomy-dependent manner. In the current work, we further characterize the utility of MARLOWE using publicly available proteomics data from forensically-relevant samples. MARLOWE characterizes samples based on their protein profile, and returns ranked organism lists of potential contributors and taxonomic scores based on shared strong peptides between organisms. Overall, the correct characterization rate ranges between 44 and 100%, depending on the sample type and data acquisition parameters (with lower numbers associated with lower-quality data sets). MARLOWE demonstrates successful characterization of true contributors and close relatives, and provides sufficient specificity to distinguish certain microbial species. MARLOWE demonstrates its ability to provide insight into potential taxonomic sources for a wide range of sample types without prior assumptions about sample contents. As a result, this approach can find utility in forensic science and also broadly in bioanalytical applications that utilize proteomics approaches for taxonomic characterization.

Bacteria↗