Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “proteome”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

MARLOWE: An Untargeted Proteomics, Statistical Approach to Taxonomic Classification for Forensics

General proteomics research for fundamental science typically addresses laboratory- or patient-derived samples of known origin and composition. However, in a few research areas, such as environmental proteomics, clinical identification of infectious organisms, archeology, art/cultural history, and forensics, attributing the origin of a protein-containing sample to the organisms that produced it is a central focus. A small number of groups have approached this problem and developed software tools for taxonomic characterization and/or identification using bottom-up proteomics. Most such tools identify peptides via database search, and many rely on organism-specific peptides as markers. Our group recently introduced MARLOWE, a software tool for taxonomic characterization of unknown samples based on de novo peptide identification and signal-erosion-resistant strong peptides, which are shared peptides distributed in a taxonomy-dependent manner. In the current work, we further characterize the utility of MARLOWE using publicly available proteomics data from forensically-relevant samples. MARLOWE characterizes samples based on their protein profile, and returns ranked organism lists of potential contributors and taxonomic scores based on shared strong peptides between organisms. Overall, the correct characterization rate ranges between 44 and 100%, depending on the sample type and data acquisition parameters (with lower numbers associated with lower-quality data sets). MARLOWE demonstrates successful characterization of true contributors and close relatives, and provides sufficient specificity to distinguish certain microbial species. MARLOWE demonstrates its ability to provide insight into potential taxonomic sources for a wide range of sample types without prior assumptions about sample contents. As a result, this approach can find utility in forensic science and also broadly in bioanalytical applications that utilize proteomics approaches for taxonomic characterization.

Bacteria↗

Data from a multi-year targeted proteomics study of a longitudinal birth cohort of type 1 diabetes

The deployment of liquid chromatography-mass spectrometry-based plasma proteomics experiments in a large cohort is sparse, leading to a lack of data available for benchmarking, method development or validation. Comprised of 6,426 plasma analyses, The Environmental Determinants of Diabetes in the Young (TEDDY) proteomics validation study constitutes one of the largest targeted proteomics experiments in the literature to date. The proteomics data from this study were generated over the course of 2.5 years from over 900 study subjects, each providing up to 29 longitudinal samples. The data also includes 916 quality control samples. The targeted mass spectrometry assay was comprised of 694 peptides mapping to 167 proteins and the panel was measured in each subject and QC sample. The targeted proteomic dataset presented here can be used as a resource for new computational method development, such as for batch correction, as well as for benchmarking and comparing the performance of different methods/tools.

60 APPLIED LIFE SCIENCES↗

Proteomic Analysis of Methanococcus voltae Grown in the Presence of Mineral and Nonmineral Sources of Iron and Sulfur

Iron sulfur (Fe-S) proteins are essential and ubiquitous across all domains of life, yet the mechanisms underpinning assimilation of iron (Fe) and sulfur (S) and biogenesis of Fe-S clusters are poorly understood. This is particularly true for anaerobic methanogenic archaea, which are known to employ more Fe-S proteins than other prokaryotes. Here, we utilized a deep proteomics analysis of Methanococcus voltae A3 cultured in the presence of either synthetic pyrite (FeS 2 ) or aqueous forms of ferrous iron and sulfide to elucidate physiological responses to growth on mineral or nonmineral sources of Fe and S. The liquid chromatography-mass spectrometry (LCMS) shotgun proteomics analysis included 77% of the predicted proteome. Through a comparative analysis of intra- and extracellular proteomes, candidate proteins associated with FeS 2 reductive dissolution, Fe and S acquisition, and the subsequent transport, trafficking, and storage of Fe and S were identified. The proteomic response shows a large and balanced change, suggesting that M. voltae makes physiological adjustments involving a range of biochemical processes based on the available nutrient source. Among the proteins differentially regulated were members of core methanogenesis, oxidoreductases, membrane proteins putatively involved in transport, Fe-S binding ferredoxin and radical S-adenosylmethionine proteins, ribosomal proteins, and intracellular proteins involved in Fe-S cluster assembly and storage. This work improves our understanding of ancient biogeochemical processes and can support efforts in biomining of minerals. Clusters of iron and sulfur are key components of the active sites of enzymes that facilitate microbial conversion of light or electrical energy into chemical bonds. The proteins responsible for transporting iron and sulfur into cells and assembling these elements into metal clusters are not well understood. Using a microorganism that has an unusually high demand for iron and sulfur, we conducted a global investigation of cellular proteins and how they change based on the mineral forms of iron and sulfur. Understanding this process will answer questions about life on early earth and has application in biomining and sustainable sources of energy.

59 BASIC BIOLOGICAL SCIENCES↗

Sorted-cell proteomics reveals an AT1-associated epithelial cornification phenotype and suggests endothelial redox imbalance in human bronchopulmonary dysplasia

Bronchopulmonary dysplasia (BPD) is a neonatal lung disease characterized by inflammation and scarring leading to long-term tissue damage. Previous whole tissue proteomics identified BPD-specific proteome changes and cell type shifts. Little is known about the proteome-level changes within specific cell populations in disease. Here, we sorted epithelial (EPI) and endothelial (ENDO) cell populations based on their differential surface markers from normal and BPD human lungs. Using a low-input compatible sample preparation method (MicroPOT), proteins were extracted and digested into peptides and subjected to liquid chromatography-tandem mass spectrometry (LC-MS/MS) proteome analysis. Of the 4,970 proteins detected, 293 were modulated in abundance or detection in the EPI population and 422 were modulated in ENDO cells. Modulation of proteins associated with actin-cytoskeletal function, such as SCEL, LMO7, and TBA1B was observed in the BPD EPIs. Using confocal imaging and analysis, we validated the presence of aberrant multilayer-like structures comprising SCEL and LMO7, known to be associated with epidermal cornification, in the human BPD lung. This is the first report of the accumulation of cornification-associated proteins in BPD. Their localization in the alveolar parenchyma, primarily associated with alveolar type 1 (AT1) cells, suggests a role in the BPD postinjury response. In the ENDOs, redox balance and mitochondrial function pathways were modulated. Alternative mRNA splicing and cell proliferative functions were elevated in both populations, suggesting potential dysregulation of cell progenitor fate. This study characterized the proteome of epithelial and endothelial cells from the BPD lung for the first time, identifying population-specific changes in BPD pathogenesis.

BPD↗

Plasma proteomic biomarkers of physical frailty in heart failure: a propensity score matched discovery-based pilot study

Background: Physical frailty is highly prevalent in heart failure (HF), but we lack an understanding of the underlying pathophysiology. Proteomics evaluation of plasma samples may elucidate potential mechanisms and biomarkers of physical frailty in HF. We aimed to identify plasma proteomic biomarkers that are differentially expressed between physically frail and non physically frail adults with HF. Methods: This was a secondary analysis of a subset of data and plasma samples from a study of frailty among patients with New York Heart Association (NYHA) Functional Classification I-IV HF. Physical frailty was measured using the Frailty Phenotype Criteria. Propensity score matching was used to match pairs of physically frail (n = 20) vs. non-physically frail (n = 20) patients on clinical characteristics. Plasma samples were processed using a sensitive liquid chromatography mass spectrometry platform, utilizing a multiplexed tandem mass tag-labeled quantitative proteomics approach. Differentially expressed proteins were quantified individually using paired t tests with associated log fold change of 0.3 and Fisher’s combined p values. Results: The sample (n = 40) was 62.8±16.9 years old, 58% female, and 55% NYHA Class III/IV. Proteomics analysis revealed 7 proteins differentially expressed using full differential criteria: matrix metalloproteinase-14 was downregulated in frailty, and copine-1, low affinity immunoglobulin gamma Fc region receptor III-A and III-B, probable non-functional immunoglobulin kappa variable 2D-24, glutathione S-transferase Mu 1, and argininosuccinate lyase were upregulated in frailty. Conclusions: Proteomic biomarkers related to the immune system, stress response, and detoxification were differentially expressed between physically frail and non-physically frail adults with HF.

Biomarkers↗

Human Host Cellular Response to HCoV-229E Infection Proteomics (ACS-JM-DP2)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5) nuclear extracts, immortalized human lung fibroblasts cells (MRC5) (MOI5) nuclear extracts, and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue and processed for proteome analysis. Processed datasets are openly accessible from the download button and contain secondary processed proteomic results files and supporting metadata materials. Experimental proteomics samples were prepared using Limited Proteolysis (LiP) methods for Label-free quantification (LFQ) and global proteomic evaluation. Sample data was acquired using a Q-Exactive HF-X mass spectrometer and was processed and compiled using MaxQuant software (v.1.6.17.0). Processed proteomic data downloads include a sample naming key, processed MaxQuant results/parameters, and protein annotated relative abundance files. See corresponding primary data accessions below and Viral Experiment LiP Analysis source code supporting data transparency and reuse. Experimental transcriptomics samples were collected in parallel and processed for RNA sequencing (RNA-Seq) as summarized under ACS-DP1 (https://data.pnnl.gov/group/nodes/dataset/34069).

59 BASIC BIOLOGICAL SCIENCES↗

Functional Meta-Analysis of the Proteomic Responses of Arabidopsis Seedlings to the Spaceflight Environment Reveals Multi-Dimensional Sources of Variability across Spaceflight Experiments

The human quest for sustainable habitation of extraterrestrial environments necessitates a robust understanding of life’s adaptability to the unique conditions of spaceflight. This study provides a comprehensive proteomic dissection of the Arabidopsis plant’s responses to the spaceflight environment through a meta-analysis of proteomics data from four separate spaceflight experiments conducted on the International Space Station (ISS) in different hardware configurations. Raw proteomics LC/MS spectra were analyzed for differential expression in MaxQuant and Perseus software. The analysis of dissimilarities among the datasets reveals the multidimensional nature of plant proteomic responses to spaceflight, impacted by variables such as spaceflight hardware, seedling age, lighting conditions, and proteomic quantification techniques. By contrasting datasets that varied in light exposure, we elucidated proteins involved in photomorphogenesis and skotomorphogenesis in plant spaceflight responses. Additionally, with data from an onboard 1 g control experiment, we isolated proteins that specifically respond to the microgravity environment and those that respond to other spaceflight conditions. This study identified proteins and associated metabolic pathways that are consistently impacted across the datasets. Notably, these shared proteins were associated with critical metabolic functions, including carbon metabolism, glycolysis, gluconeogenesis, and amino acid biosynthesis, underscoring their potential significance in Arabidopsis’ spaceflight adaptation mechanisms and informing strategies for successful space farming.

59 BASIC BIOLOGICAL SCIENCES↗

Aggregation Methods for Quantifying PTM and Structural Changes in Bottom-Up Proteomics

Bottom-up proteomic workflows rely on sequential preprocessing steps, commonly including peptide-to-protein aggregation (“roll-up”), to enhance data reliability and interpretability. While roll-up is effective for protein-centered analyses, it may be suboptimal for applications focused on post-translational modifications (PTMs) or protein structural changes, such as limited proteolysis–mass spectrometry (LiP-MS). Here, we investigate how different roll-up strategies influence site-level quantification in PTM differential analysis. Moreover, we introduce a novel site-centric roll-up approach tailored for LiP-MS, which quantifies proteolytic fragments rather than solely tryptic peptides. We benchmark these methods through simulation studies, comparing their sensitivity and specificity in detecting structural and PTM-driven changes. We found that the median and mean roll-up methods outperform the sum method in both PTM and LiP proteomics, and site-level quantification in LiP outperforms peptide-level quantification. Our findings offer the first systematic, data-driven guidance for selecting roll-up techniques in site-level proteomic analyses, with implications for both PTM-focused and structural proteomics studies.

aggregation↗

Applications of targeted proteomics in metabolic engineering: advances and opportunities

Optimization of metabolically engineered organisms requires good understanding of producing balanced level of pathway proteins. Targeted proteomics via selected-reaction monitoring (SRM) has been increasingly used in metabolic engineering studies to detect and quantify sets of proteins with high selectivity, multiplexity, and reproducibility. In combination with metabolomics and other omics tools, targeted proteomics has helped optimize the production of many bio-based chemicals in various metabolic engineering cell factories. In this review, we present recent applications of targeted proteomics in metabolic engineering studies and highlight several successful cases of targeted proteomics in boosting production of commodity and high value chemicals. Additionally, we also discuss challenges and limitations of current targeted proteomics and map opportunities for future research.

59 BASIC BIOLOGICAL SCIENCES↗

High-throughput Single-Cell Proteomics and Transcriptomics from the Same Cells with a Nanoliter-Scale Spin-Transfer Approach

Single-cell multiomic platforms provide a comprehensive snapshot of cellular states and cell types by offering critical insights into the spatiotemporal regulation of biomolecular networks at a systems level, thereby defining the basis of multicellularity. Here, we introduce nanoSPINS, an advanced platform that enables high-throughput profiling and integrative analysis of the transcriptome and proteome from the same single cells using RNA sequencing and isobaric labeling LC-MS-based proteomics, respectively. NanoSPINS can efficiently transfer mRNA-containing droplets across two microarrays via a centrifugation-based approach, while proteins are retained on the initial platform. Benchmarking of nanoSPINS on two cell lines demonstrates its ability to generate global proteomic and transcriptomic profiles that align well with previously established methodologies/platforms. The incorporation of isobaric TMTpro labeling into this single-cell multiomics platform significantly enhances the throughput of single-cell proteomic analyses. Through the high-throughput quantification of the proteome and transcriptome, nanoSPINS not only facilitates the identification of molecular features at both mRNA and protein level but also provides larger sample sizes for improved statistical power in clustering and differential abundance. Given the broad applicability of single-cell multiomics in biological research and clinical settings, we believe nanoSPINS represents a powerful platform for the characterization of heterogeneous cell populations.

multi 'omics↗

Lipid droplet-associated proteins in alcohol-associated fatty liver disease: A proteomic approach

The earliest manifestation of alcohol-associated liver disease (ALD) is steatosis characterized by deposition of fat in specialized organelles called lipid droplets (LDs). While alcohol administration causes a rise in LD numbers in the hepatocytes, little is known regarding their characteristics that allow their accumulation and size to increase. The aim of the present study is to gain insights into underlying pathophysiological mechanisms by investigating the ethanol-induced changes in hepatic LD proteome as a function of LD size. Adult male Wistar rats (180–200 g BW) were fed with ethanol liquid diet for 6 weeks. At sacrifice, large-, medium-, and small-sized hepatic LD subpopulations (LD1, LD2, and LD3, respectively) were isolated and subjected to morphological and proteomic analyses. Morphological analysis of LD1-LD3 fractions of ethanol-fed rats clearly demonstrated that LD1 contained larger LDs compared with LD2 and LD3 fractions. Our preliminary results from principal component analysis showed that the proteome of different-sized hepatic LD fractions was distinctly different. Proteomic data analysis identified over 2000 proteins in each LD fraction with significant alterations in protein abundance among the three LD fractions. Among the altered proteins, several were related to fat metabolism, including synthesis, incorporation of fatty acid, and lipolysis. Ingenuity pathway analysis revealed increased fatty acid synthesis, fatty acid incorporation, LD fusion, and reduced lipolysis in LD1 compared to LD3. Overall, the proteomic findings indicate that the increased level of protein that facilitates fusion of LDs combined with an increased association of negative regulators of lipolysis dictates the generation of large-sized LDs during the development of alcohol-associated hepatic steatosis. Several significantly altered proteins were identified in different-sized LDs isolated from livers of ethanol-fed rats. Ethanol-induced increases in specific proteins that hinder LD lipid metabolism led to the accumulation and persistence of large-sized LDs in the liver.

60 APPLIED LIFE SCIENCES↗

Proteomic and phosphoproteomic measurements enhance ability to predict ex vivo drug response in AML

Acute Myeloid Leukemia (AML) affects 20,000 patients in the US annually with a five-year survival rate of approximately 25%. One reason for the low survival rate is the high prevalence of clonal evolution that gives rise to heterogeneous sub-populations of leukemic cells with diverse mutation spectra, which eventually leads to disease relapse. This genetic heterogeneity drives the activation of complex signaling pathways that is reflected at the protein level. This diversity makes it difficult to treat AML with targeted therapy, requiring custom patient treatment protocols tailored to each individual’s leukemia. Toward this end, the Beat AML research program prospectively collected genomic and transcriptomic data from over 1000 AML patients and carried out ex vivo drug sensitivity assays to identify genomic signatures that could predict patient-specific drug responses. However, there are inherent weaknesses in using only genetic and transcriptomic measurements as surrogates of drug response, particularly the absence of direct information about phosphorylation-mediated signal transduction. As a member of the Clinical Proteomic Tumor Analysis Consortium, we have extended the molecular characterization of this cohort by collecting proteomic and phosphoproteomic measurements from a subset of these patient samples (38 in total) to evaluate the hypothesis that proteomic signatures can improve the ability to predict response to 26 drugs in AML ex vivo samples. In this work we describe our systematic, multi-omic approach to evaluate proteomic signatures of drug response and compare protein levels to other markers of drug response such as mutational patterns. We explore the nuances of this approach using two drugs that target key pathways activated in AML: quizartinib (FLT3) and trametinib (Ras/MEK), and show how patient-derived signatures can be interpreted biologically and validated in cell lines. In conclusion, this pilot study demonstrates strong promise for proteomics-based patient stratification to assess drug sensitivity in AML.

60 APPLIED LIFE SCIENCES↗

Mono-mix strategy enables comparative proteomics of a cross-kingdom microbial symbiosis

Cross-kingdom microbial symbioses, such as those between algae and bacteria, are key players in biogeochemical cycles. The molecular changes during initiation and establishment of symbiosis are of great interest, but quantitatively monitoring such changes can be challenging, particularly when the microorganisms differ greatly in size or are intimately associated. Here, we analyze output from label-free, data-dependent acquisition (DDA) LC-MS/MS proteomics experiments investigating the well-studied interaction between the alga Chlamydomonas reinhardtii and the heterotrophic bacterium Mesorhizobium japonicum. We found that detection of bacterial proteins decreased in coculture by 50% proteome-wide due to the abundance of algal proteins. As a result, standard differential expression analysis led to numerous false-positive reports of significantly downregulated proteins, where it was not possible to distinguish meaningful biological responses to symbiosis from artifacts of the reduced protein detection in coculture relative to monoculture. We show that data normalization alone does not eliminate the impact of altered detection on differential expression analysis of the cross-kingdom symbiosis. We assessed two additional strategies to overcome this methodological artifact inherent to DDA proteomics. In the first, we combined algal and bacterial monocultures at a relative abundance that mimicked the coculture, creating a “mono-mix” control to which the coculture could be compared. This approach enabled comparable detection of bacterial proteins in the coculture and the monoculture control. In the second strategy, we enhanced detection of lowly abundant bacterial proteins by using sample fractionation upstream of LC-MS/MS analysis. When these simple approaches were combined, they allowed for meaningful comparisons of nearly 10,000 algal proteins and over 4,000 bacterial proteins in response to symbiosis by DDA. They successfully recovered expected changes in the bacterial proteome in response to algal coculture, including upregulation of sugar-binding proteins and transporters. They also revealed novel proteomic responses to coculture that guide hypotheses about algal-bacterial interactions.

Dupuis, Sunnyjoy [University of California, Berkel↗

Spatially Resolved Top-Down Proteomics of Tissue Sections Based on a Microfluidic Nanodroplet Sample Preparation Platform

Conventional proteomics measures the averaged signal from mixed cell populations or bulk tissues, leading to the dilution of significant changes in subpopulations of cells that might serve as important biomarkers. Recent developments in bottom-up proteomics have enabled spatial mapping of cellular heterogeneity in tissue microenvironments. However, bottom-up proteomics cannot precisely infer the abundance changes of intact proteins, which are presented as proteoforms. Herein, we described a spatially resolved top-down proteomics (TDP) platform for proteoform identification and quantification directly on thin tissue sections. The spatial TDP platform consisted of a nanoPOTS (nanodroplet Processing in One pot for Trace Samples)-based sample preparation system and an LCM (laser capture microdissection)-based cell isolation system. We improved the nanoPOTS sample preparation by adding benzonase in the extraction buffer to enhance the coverage of nucleus proteins. Using ~200 cultured cells as model samples, the improved approach increased proteoform identifications from 493 to 700; newly identified proteoforms primarily corresponded to nuclear proteins. To demonstrate the spatial TDP platform in tissue samples, we analyzed LCM-isolated tissue voxels from rat brain cortex and hypothalamus regions. We quantified 426 proteoforms by combining identifications from TopPIC and TDPortal with the quantitation from ProMex. Several proteoforms corresponding to the same gene exhibited mixed abundance profiles between two tissue regions, suggesting potential PTM-specific spatial distributions. The spatial TDP workflow has prospects for biomarker discovery at proteoform level from small tissue sections.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluation of Differential Peptide Loading on Tandem Mass Tag-Based Proteomic and Phosphoproteomic Data Quality

Global and phosphoproteome profiling has demonstrated great utility for the analysis of clinical specimens. One major barrier to the broad clinical application of proteomic profiling is the large amount of biological material required, particularly for phosphoproteomics—currently on the order of 25 mg wet tissue weight, depending on tissue type. For hematopoietic cancers such as acute myeloid leukemia (AML), the sample requirement is in excess of 10 million (1E7) peripheral blood mononuclear cells (PBMCs). Throughout the course of a prospective study, this requirement will certainly exceed what is obtainable from many of the individual patients/timepoints. For this reason, we were interested in examining the impact of differential peptide loading across multiplex channels on proteomic data quality. Methods: To achieve this, we tested a range of channel loading amounts (20, 40, 100, 200, and 400 µg of tryptic peptides, or approximately the material obtainable from 5E5, 1E6, 2.5E6, 5E6, and 1E7 AML patient cells) to assess proteome coverage, quantification reproducibility and accuracy in experiments utilizing isobaric tandem mass tag (TMT) labeling. As expected, we found that fewer missing values are observed in TMT channels with higher peptide loading amounts compared to those with lower loading. Moreover, channels with lower loading amounts have greater quantitative variability than channels with higher loading amounts. Statistical analysis of the differences in means among the five loading groups showed that the 20 µg loading group was significantly different from the 400 µg loading group. However, no significant differences were detected among the 40, 100, 200 and 400 µg loading groups. Conclusions: These assessment data demonstrate the practical limits of loading differential quantities of peptides across channels in TMT multiplexes, and provide a basis for designing the optimal clinical proteomics study when specimen quantities are limited.

59 BASIC BIOLOGICAL SCIENCES↗

Spatial Proteomics towards cellular Resolution

Introduction: Spatial biology is an emerging interdisciplinary field facilitating biological discoveries through the use of spatial omics technologies. Recent advancements in spatial transcriptomics, spatial genomics (e.g. genetic mutations and epigenetic marks), multiplexed immunofluorescence, and spatial metabolomics/lipidomics have enabled high-resolution spatial profiling of gene expression, genetic variation, protein expression, and metabolites/lipids profiles in tissue. These developments contribute to a deeper understanding of the spatial organization within tissue microenvironments at the molecular level. Areas covered: This report provides an overview of the untargeted, bottom-up mass spectrometry (MS)-based spatial proteomics workflow. It highlights recent progress in tissue dissection, sample processing, bioinformatics, and liquid chromatography (LC)-MS technologies that are advancing spatial proteomics toward cellular resolution. Expert opinion: The field of untargeted MS-based spatial proteomics is rapidly evolving and holds great promise. To fully realize the potential of spatial proteomics, it is critical to advance data analysis and develop automated and intelligent tissue dissection at the cellular or subcellular level, along with high-throughput LC-MS analyses of thousands of samples. In conclusion, achieving these goals will necessitate significant advancements in tissue dissection technologies, LC-MS instrumentation, and computational tools.

59 BASIC BIOLOGICAL SCIENCES↗

A Comprehensive Urine Proteome Database Generated From Patients With Various Renal Conditions and Prostate Cancer

Urine proteins can serve as viable biomarkers for diagnosing and monitoring various diseases. A comprehensive urine proteome database, generated from a variety of urine samples with different disease conditions, can serve as a reference resource for facilitating discovery of potential urine protein biomarkers. Herein, we present a urine proteome database generated from multiple datasets using 2D LC-MS/MS proteome profiling of urine samples from healthy individuals (HI), renal transplant patients with acute rejection (AR) and stable graft (STA), patients with non-specific proteinuria (NS), and patients with prostate cancer (PC). A total of ~28,000 unique peptides spanning ~2,200 unique proteins were identified with a false discovery rate of <0.5% at the protein level. Over one third of the annotated proteins were plasma membrane proteins and another one third were extracellular proteins according to gene ontology analysis. Ingenuity Pathway Analysis of these proteins revealed 349 potential biomarkers in the literature-curated database. Forty-three percentage of all known cluster of differentiation (CD) proteins were identified in the various human urine samples. Interestingly, following comparisons with five recently published urine proteome profiling studies, which applied similar approaches, there are still ~400 proteins which are unique to this current study. These may represent potential disease-associated proteins. Among them, several proteins such as serpin B3, renin receptor, and periostin have been reported as pathological markers for renal failure and prostate cancer, respectively. Taken together, our data should provide valuable information for future discovery and validation studies of urine protein biomarkers for various diseases.

60 APPLIED LIFE SCIENCES↗

Reinspection of a Clinical Proteomics Tumor Analysis Consortium (CPTAC) Dataset with Cloud Computing Reveals Abundant Post-Translational Modifications and Protein Sequence Variants

The Clinical Proteomic Tumor Analysis Consortium (CPTAC) has provided some of the most in-depth analyses of the phenotypes of human tumors ever constructed. Today, the majority of proteomic data analysis is still performed using software housed on desktop computers which limits the number of sequence variants and post-translational modifications that can be considered. The original CPTAC studies limited the search for PTMs to only samples that were chemically enriched for those modified peptides. Similarly, the only sequence variants considered were those with strong evidence at the exon or transcript level. In this multi-institutional collaborative reanalysis, we utilized unbiased protein databases containing millions of human sequence variants in conjunction with hundreds of common post-translational modifications. Using these tools, we identified tens of thousands of high-confidence PTMs and sequence variants. We identified 4132 phosphorylated peptides in nonenriched samples, 93% of which were confirmed in the samples which were chemically enriched for phosphopeptides. In addition, our results also cover 90% of the high-confidence variants reported by the original proteogenomics study, without the need for sample specific next-generation sequencing. Finally, we report fivefold more somatic and germline variants that have an independent evidence at the peptide level, including mutations in ERRB2 and BCAS1. In this reanalysis of CPTAC proteomic data with cloud computing, we present an openly available and searchable web resource of the highest-coverage proteomic profiling of human tumors described to date.

60 APPLIED LIFE SCIENCES↗