Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “annotations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

deadtrees.earth — An open-access and interactive database for centimeter-scale aerial imagery to uncover global tree mortality dynamics

Excessive tree mortality is a global concern and remains poorly understood as it is a complex phenomenon. We lack global and temporally continuous coverage on tree mortality data. Ground-based observations on tree mortality, e.g., derived from national inventories, are very sparse, and may not be standardized or spatially explicit. Earth observation data, combined with supervised machine learning, offer a promising approach to map overstory tree mortality in a consistent manner over space and time. However, global-scale machine learning requires broad training data covering a wide range of environmental settings and forest types. Low altitude observation platforms (e.g., drones or airplanes) provide a cost-effective source of training data by capturing high-resolution orthophotos of overstory tree mortality events at centimeter-scale resolution. Here, we introduce deadtrees.earth, an open-access platform hosting more than two thousand centimeter-resolution orthophotos, covering more than 1,000,000 ha, of which more than 58,000 ha are manually annotated with live/dead tree classifications. This community-sourced and rigorously curated dataset can serve as a comprehensive reference dataset to uncover tree mortality patterns from local to global scales using space-based Earth observation data and machine learning models. This will provide the basis to attribute tree mortality patterns to environmental changes or project tree mortality dynamics to the future. The open nature of deadtrees.earth, together with its curation of high-quality, spatially representative, and ecologically diverse data will continuously increase our capacity to uncover and understand tree mortality dynamics.

Citizen science↗

An approach for broad molecular imaging of the root-soil interface via indirect matrix-assisted laser desorption/ionization mass spectrometry

Understanding of rhizospheric processes is limited by the need for imaging complex molecular transformations at relevant spatial scales within the root soil continuum. In this work, we demonstrate a method to enable this analysis by first extracting organic compounds from the rhizosphere onto a PVDF membrane while maintaining their 2D distribution. We then image the distribution of chemical compounds using matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS). This approach permitted us to visualize and identify compounds on the root surface and presumed root exudates in the rhizosphere. Within a 1.8 cm x 0.6 cm sampling area of a switchgrass rhizosphere, we could observe at least four chemically distinct zones. Using high performance Fourier transform ion cyclotron MS, we were able to accurately annotate numerous molecules co-localized to each of these zones.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Phosphate amendment drives bloom of RNA viruses after soil wet-up

Soil rewetting after a dry period results in a surge of activity and succession in both microbial and DNA virus communities. Less is known about the response of RNA viruses to soil rewetting—while they are highly diverse and widely distributed in soil, they remain understudied. We hypothesized that RNA viruses would show temporal succession following rewetting and that phosphate amendment would influence their trajectory, as viral proliferation may cause phosphorus limitation. Using 39 time-resolved metatranscriptomes and amplicon data, 2190 RNA viral populations were identified across five phyla, with 26 % of these predicted to infect bacteria, and 11 % fungi. Only 1.2 % of viral populations had annotated capsid genes, suggesting most persist via intracellular replication without a free virion phase. Phosphate amendment altered RNA viral community composition within the first week and amended vs. unamended communities remained distinguishable for up to three weeks. While the overall host community remained stable, certain bacterial populations showed reduced abundance in phosphate-amended soils, likely due to increased viral lysis, as RNA bacteriophages proliferated significantly. Notably, 60 % of the viruses with increased abundance under phosphate amendment belonged to basal Lenarviricota clades rather than well-known groups like Leviviricetes. We estimate RNA bacteriophage infections may affect 10 7 –10 9 bacteria per gram of soil, aligning with the total bacterial population (10 7 –10 10 g -1 soil), suggesting that RNA phages significantly influence bacterial communities post-wet-up, with phosphorus availability modulating this effect.

59 BASIC BIOLOGICAL SCIENCES↗

Root exudate lipids: Uncovering chemodiversity and carbon stability potential

Root-derived carbon has been shown to contribute more to soil carbon stocks than aboveground litter. Yet the molecular chemodiversity of root exudates remains poorly understood due to limited characterization and annotation. In this study, we characterized the molecular chemodiversity and production of metabolites and lipids in root exudates from field grown mature tall wheatgrass (Thinopyrum ponticum). We discovered a diversity of lipids, including substantial levels of triacylglycerols (∼19 μg/g fresh root per min), fatty acyls, sphingolipids, sterol lipids, and glycerophospholipids, some of which have not been previously documented in root exudates. By integrating tandem mass spectral library searching and deep learning-based chemical class assignment, our metabo-lipidomics approach significantly expanded the known molecular diversity of root exudates. Rates of lipid derived carbon production were approximately double that of polar metabolites (lipids: 81.52 ± 13.81 vs polar metabolites: 38.41 ± 5.93 μg C g −1 fresh root mass min −1 ) with an order of magnitude higher carbon to nitrogen ratios (lipids: 459 ± 90 vs polar metabolites: 14.40 ± 0.58). Exudate lipids displayed highly negative nominal oxidation state of carbon (−1.182 to −1.909), indicating that these compounds may be less favorable for microbial decomposition. Together our results suggest the potential of root exudate lipids to contribute to stable carbon pools in soil, supporting long-term carbon storage. This work advances understanding of plant-derived lipid inputs to soil and underscores the need for future studies on the functional roles of lipids in shaping root-microbe-soil interactions, microbial activity, soil structure, and nutrient availability – contributing to soil health.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Interpreting omics data with pathway enrichment analysis

Pathway enrichment analysis is indispensable for interpreting omics datasets and generating hypotheses. However, the foundations of enrichment analysis remain elusive to many biologists. Here, in this study, we discuss best practices in interpreting different types of omics data using pathway enrichment analysis and highlight the importance of considering intrinsic features of various types of omics data. We further explain major components that influence the outcomes of a pathway enrichment analysis, including defining background sets and choosing reference annotation databases. To improve reproducibility, we describe how to standardize reporting methodological details in publications. This article aims to serve as a primer for biologists to leverage the wealth of omics resources and motivate bioinformatics tool developers to enhance the power of pathway enrichment analysis.

60 APPLIED LIFE SCIENCES↗

Morphotype-resolved characterization of microalgal communities in a nutrient recovery process with ARTiMiS flow imaging microscopy

Microalgae-driven nutrient recovery represents a promising technology for phosphorus removal from wastewater while simultaneously generating biomass that can be valorized to offset treatment costs. As full-scale processes come online, system parameters including biomass composition must be carefully monitored to optimize performance and prevent culture crashes. In this study, flow imaging microscopy (FIM) was leveraged to characterize microalgal community composition in near real-time at a full-scale municipal wastewater treatment plant (WWTP) in Wisconsin, USA, and population and morphotype dynamics were examined to identify relationships between water chemistry, biomass composition, and system performance. Two FIM technologies, FlowCam and ARTiMiS, were evaluated as monitoring tools. ARTiMiS provided a more accurate estimate of total system biomass, and estimates derived from particle area as a proxy for biovolume yielded better approximations than particle counts. Deep learning classification models trained on annotated image libraries demonstrated equivalent performance between FlowCam and ARTiMiS, and convolutional neural network (CNN) classifiers proved significantly more accurate when compared to feature table-based dense neural network (DNN) models. Across a two-year study period, Scenedesmus spp. appeared most important for phosphorus removal, and were negatively impacted by elevated temperatures and increase in nitrite/nitrate concentrations. Chlorella and Monoraphidium also played an important role in phosphorus removal. For both Scenedesmus and Chlorella, smaller morphological types were more often associated with better system performance, whereas larger morphotypes likely associated with stress response(s) correlated with poor phosphorus recovery rates. Furthermore, these results demonstrate the potential of FIM as a critical technology for high-resolution characterization of industrial microalgal processes.

59 BASIC BIOLOGICAL SCIENCES↗

A HIF1A/miR-485–5p/SRPK1 axis modulates the aggressiveness of glioma cells upon hypoxia

The high aggressiveness of gliomas remains a huge challenge to clinical therapies, and the hypoxic microenvironment in the core region is a critical contributor to glioma aggressiveness. In this study, it was found that miR-485–5p was low expressed within glioma tissue samples and cells. GO enrichment annotation indicated that the predicted downstream targets miR-485–5p were enriched in hypoxia response and decreased oxygen level. In glioma cells, miR-485–5p overexpression suppressed cell viability, migratory ability, and invasive ability under both normoxic and hypoxic conditions. Through direct binding, miR-485–5p suppressed SRPK1 expression. Under hypoxia, SRPK1 overexpression enhanced hypoxia-induced glioma cell aggressiveness and significantly reversed the effects of miR-485–5p overexpression. Moreover, HIF1A could target the miR-485–5p promoter region to inhibit the transcription. HIF1A, miR-485–5p, and SRPK1 form a regulatory axis, which modulates glioma cell aggressiveness under hypoxia. In conclusion, we identify a HIF1A/miR-485–5p/SRPK1 axis that modulates the aggressiveness of glioma cells under hypoxia. The axis could potentially provide new research avenues in the treatment of gliomas considering the hypoxic environment in its core.

60 APPLIED LIFE SCIENCES↗

Disrupted NOS2 metabolism drives myoblast response to wasting-associated cytokines

Highlights: • Amino acid metabolism is impacted in myoblasts treated with cancer cell conditioned media. • Inflammatory cytokines induce nitric oxide synthase 2 in myoblasts. • Elevated nitric oxide synthase 2 activity impairs myoblast proliferation and differentiation. Skeletal muscle wasting drives negative clinical outcomes and is associated with a spectrum of pathologies including cancer. Cancer cachexia is a multi-factorial syndrome that encompasses skeletal muscle wasting and remains understudied, despite being a frequent and serious co-morbidity. Deviation from the homeostatic balance between breakdown and regeneration leads to muscle wasting disorders, such as cancer cachexia. Muscle stem cells (MuSCs) are the cellular compartment responsible for muscle regeneration, which makes MuSCs an intriguing target in the context of wasting muscle. Molecular studies investigating MuSCs and skeletal muscle wasting largely focus on transcriptional changes, but our group and others propose that metabolic changes are another layer of cellular regulation underlying MuSC dysfunction in cancer cachexia. In the present study, we combined gene expression and non-targeted metabolomic profiling of myoblasts exposed to wasting conditions (cancer cell conditioned media, CC-CM) to derive a more complete picture of the myoblast response to wasting factors. After mapping these features to annotated pathways, we found that more than half of the mapped pathways were amino acid-related, linking global amino acid metabolic disruption to conditioned media-induced myoblast defects. Notably, arginine metabolism was a highly enriched pathway in combined metabolomic and transcriptomic data. Arginine catabolism generates nitric oxide (NO), an important signaling molecule known to have negative effects on mature muscle. We hypothesize that tumor-derived disruptions in Nitric Oxide Synthase (NOS)2-regulated arginine catabolism impair differentiation of MuSCs. The work presented here further investigates the effect of NOS2 overactivity on myoblast proliferation and differentiation. We show that NOS2 inhibition is sufficient to rescue wasting phenotypes associated with inflammatory cytokines. Ultimately, this work provides new insights into MuSC biology and opens up potential therapeutic avenues for addressing disrupted MuSC dynamics in cancer cachexia.

60 APPLIED LIFE SCIENCES↗

Using low-coverage whole genome sequencing (genome skimming) to delineate three introgressed species of buffalofish ( Ictiobus )

Consumption of buffalofish has been sporadically associated with Haff disease-like illnesses involving sudden onset muscle pain and weakness due to skeletal muscle rhabdomyolysis, but determination of precisely which species are associated with these illnesses has been impeded by a lack of species-specific DNA-based markers. Here, three closely related species of buffalofish native to the Mississippi River Basin (Ictiobus bubalus, Ictiobus cyprinellus and Ictiobus niger) that have previously proven genetically indistinguishable using both mitochondrial and nuclear single-locus sequencing were reliably discriminated using low-coverage whole genome sequencing (‘genome skimming’). Using 44 specimens representing the three species collected from the mid/upper (Missouri) and lower (Louisiana) regions of the species’ native ranges, the SISRS (Site Identification from Short Read Sequences) bioinformatics pipeline was adapted to (1) identify over 620Mbp of putatively homologous nuclear sequence data and (2) isolate over 140,000 single-nucleotide polymorphisms (SNPs) that supported accurate species delimitation, all without the use of a reference genome or annotation data. These sites were used to classify Ictiobus spp. samples with genome-skim data, along with a larger set (n = 67) where ultraconserved elements (UCEs) were sequenced. Analyses of whole mitochondrial data revealed more limited signal. Nearly all samples matched their purported species based on morphologic identification, but two Missouri samples morphologically identified as I. niger grouped with samples of I. bubalus, albeit with significant enrichment of I. niger SNPs. To our knowledge this is the first report of a DNA-based tool to reliably discriminate these three morphologically distinct species.

59 BASIC BIOLOGICAL SCIENCES↗

RadioGalaxyNET: Dataset and novel computer vision algorithms for the detection of extended radio galaxies and infrared hosts

Abstract Creating radio galaxy catalogues from next-generation deep surveys requires automated identification of associated components of extended sources and their corresponding infrared hosts. In this paper, we introduce RadioGalaxyNET, a multimodal dataset, and a suite of novel computer vision algorithms designed to automate the detection and localization of multi-component extended radio galaxies and their corresponding infrared hosts. The dataset comprises 4 155 instances of galaxies in 2 800 images with both radio and infrared channels. Each instance provides information about the extended radio galaxy class, its corresponding bounding box encompassing all components, the pixel-level segmentation mask, and the keypoint position of its corresponding infrared host galaxy. RadioGalaxyNET is the first dataset to include images from the highly sensitive Australian Square Kilometre Array Pathfinder (ASKAP) radio telescope, corresponding infrared images, and instance-level annotations for galaxy detection. We benchmark several object detection algorithms on the dataset and propose a novel multimodal approach to simultaneously detect radio galaxies and the positions of infrared hosts.

Astronomy & Astrophysics↗

Constructing Self-Labeled Materials Imaging Datasets from Open Access Scientific Journals with EXSCLAIM!

Due to recent improvements in image resolution and acquisition speeds, materials microscopy is experiencing an explosion in imaging data. Yet, despite the volume of images generated, the overall accessibility landscape is highly fragmented, as researchers who do release images to the public, often only do so as snapshots of their larger private dataset in context of scientific journal publications. The effort to automatically consolidate images and descriptive information from web-based platforms has garnered broad attention from the computer vision, language technologies, and chemistry/materials informatics communities. However, these methods are problematic for scientific figures because over 30% of figures are compound in nature, and it is the individual images themselves, paired with relevant context, that are necessary for construction a proper labeled dataset. To this end, we outline in this paper the design of a software pipeline for the automatic EXtraction, Separation, and Caption-based natural Language Annotation of IMages from scientific figures (EXSCLAIM!). Successful consolidation of materials imaging across literature sources will enhance navigation and searchability of materials microscopy images for both novice and experienced researchers, as well as establish the framework necessary for users to search by images, text, or some combination of both.

36 MATERIALS SCIENCE↗

CFM-ID 4.0: More Accurate ESI MS/MS Spectral Prediction and Compound Identification

In the field of metabolomics, mass spectrometry (MS) is the method most commonly used for identifying and annotating metabolites. As this typically involves matching a given MS spectrum against an experimentally acquired reference spectral library, this approach is limited by the coverage and size of such libraries (which typically number in the thousands). These experimental libraries can be greatly extended by predicting the MS spectra of known chemical structures (which number in the millions) to create computational reference spectral libraries. To facilitate the generation of predicted spectral reference libraries we developed CFM-ID, a computer program that can accurately predict ESI-MS/MS spectrum for a given compound structure. CFM-ID is one of the best-performing methods for compound-to-mass-spectrum prediction, and also one of the top tools for in silico mass-spectrum-to-compound identification. This work improves CFM-ID’s ability to predict ESI-MS/MS spectra from compounds by: (1) learning parameters from features based on the molecular topology, (2) adding a new approach to ring cleavage that models such cleavage as a sequence of simple chemical bond dissociations and (3) expanding its hand-written rule-based predictor to cover more chemical classes, including acylcarnitines, acylcholines, flavonols, flavones, flavanones, and flavonoid glycosides. We demonstrate that this new version of CFM-ID (version 4.0) is significantly more accurate than previous CFM-ID versions, in terms of both EI-MS/MS spectral prediction and compound identification. CFM-ID 4.0 is available at http://cfmid4.wishartlab.com/ as a webservice and docker images can be downloaded at https://hub.docker.com/r/wishartlab/cfmid

Wang, Fei↗

Experimental Assessment of Mammalian Lipidome Complexity Using Multimodal 21 T FTICR Mass Spectrometry Imaging

Herein, we assess the complementarity and complexity of data that can be detected within mammalian lipidome mass spectrometry imaging (MSI) via matrix-assisted laser desorption ionization (MALDI) and nanospray desorption electrospray ionization (nano-DESI). We do so by employing 21 T Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) with absorption mode FT processing in both cases, allowing unmatched mass resolving power per unit time (≥613k at m/z 760, 1.536 s transients). While our results demonstrated that molecular coverage and dynamic range capabilities were greater in MALDI analysis, nano-DESI provided superior mass error, and all annotations for both modes had sub-ppm error. Taken together, these experiments highlight the coverage of 1676 lipids and serve as a functional guide for expected lipidome complexity within nano-DESI-MSI and MALDI-MSI. To further assess the lipidome complexity, mass splits (i.e., the difference in mass between neighboring peaks) within single pixels were collated across all pixels from each respective MSI experiment. The spatial localization of these mass splits was powerful in informing whether the observed mass splits were biological or artificial (e.g., matrix related). Mass splits down to 2.4 mDa were observed (i.e., sodium adduct ambiguity) in each experiment, and both modalities highlighted comparable degrees of lipidome complexity. Further, we highlight the persistence of certain mass splits (e.g., 8.9 mDa; double bond ambiguity) independent of ionization biases. Here, we also evaluate the need for ultrahigh mass resolving power for mass splits ≤4.6 mDa (potassium adduct ambiguity) at m/z > 1000, which may only be resolved by advanced FTICR-MS instrumentation.

21T-FTICR-MS↗

Increasing the Scale of the Mass Spectrometry Query Language Compendium with Explainable AI

A significant bottleneck in metabolomics data interpretation is the effective use of domain knowledge to assign structural information based on fragmentation patterns. The mass spectrometry query language (MassQL) aims to make this process accessible and applicable across multiple analysis platforms. While advanced computational methods are capable of predicting compound structures from fragmentation data, AI/ML approaches often rely on complex, opaque criteria that are difficult to interpret or modify. As a result, their predictive patterns cannot be readily translated into human-readable rules, such as those used in MassQL. Here, in this study, we introduce ChemEcho, a machine learning embedding method that converts tandem mass spectrometry data into sparse feature vectors containing peak and neutral mass subformulae to enhance explainable AI/ML-based methods. An advantage of this approach is that decision trees trained using these feature vectors can be directly translated to MassQL. Using a battery of decision trees trained using ChemEcho embeddings to predict molecular attributes, we generated over 1500 MassQL queries for 765 molecular features and evaluated their precision and recall. From these queries, the 50 highest-performing queries were integrated into the MassQL compendium. This set of generated MassQL queries included environmentally and biologically relevant classes such as PFAS and molecules containing phosphate or sulfate substructures. To illustrate the impact these queries would have on a typical metabolomics experiment, these MassQL queries were applied to a public metabolomics data set─resulting in a marked increase in the structural information derived from tandem mass spectra. Access and reuse of these queries is expected to enhance structural annotation in untargeted experiments, leading to more specific claims and advancing many applications in metabolomics.

Harwood, Thomas V. [USDOE Joint Genome Institute (↗

High-Resolution Tandem Mass Spectrometry-Based Analysis of Model Lignin–Iron Complexes: Novel Pipeline and Complex Structures

Understanding the chemical nature of soil organic carbon (SOC) with great potential to bind iron (Fe) minerals is critical for predicting the stability of SOC. Organic ligands of Fe are among the top candidates for SOCs able to strongly sorb on Fe minerals, but most of them are still molecularly uncharacterized. To shed insights into the chemical nature of organic ligands in soil and their fate, this study developed a protocol for identifying organic ligands using ultrahigh-performance liquid chromatography-high-resolution tandem mass spectrometry (UHPLC-HRMS/MS) and metabolomic tools. The protocol was used for investigating the Fe complexes formed by model compounds of lignin-derived organic ligands, namely, caffeic acid (CA), p-coumaric acid (CMA), vanillin (VNL), and cinnamic acid (CNA). Isotopologue analysis of 54/56 Fe was used to screen out the potential UHPLC-HRMS (m/z) features for complexes formed between organic ligands and Fe, with multiple features captured for CA, CMA, VNL, and CNA when 35/37 Cl isotopologue analysis was used as supplementary evidence for the complexes with Cl. MS/MS spectra, fragment analysis, and structure prediction with SIRIUS were used to annotate the structures of mono/bidentate mono/biligand complexes. The analysis determined the structures of monodentate and bidentate complexes of FeL x Cl y (L: organic ligand, x = 1–4, y = 0–3) formed by model compounds. The protocol developed in this study can be used to identify unknown organic ligands occurring in complex environmental samples and shed light on the molecular-level processes governing the stability of the SOC.

54 ENVIRONMENTAL SCIENCES↗

Assessing Metal Ion Assignment Accuracy in Protein Data Bank Models via Elemental Spectroscopy

Accurate representation of metal ions in macromolecular structures is critical for chemical interpretation, computational modeling, and machine-learning methods that rely on Protein Data Bank (PDB) entries. However, the elemental identity of metals modeled in crystallographic structures is often inferred indirectly and rarely validated experimentally. Here, we combine Particle Induced X-ray Emission (PIXE) and X-ray Fluorescence Spectroscopy (XRFS) to determine the elemental composition of protein samples used to generate 70 deposited metalloprotein crystal structures. By analyzing the original protein material employed for crystallization, but before the addition of crystallization buffer solutions, we assess whether the modeled metal ions in deposited structures are consistent with experimentally detectable elemental content. We find that in a majority of cases, the metals modeled in the corresponding PDB entries are inconsistent with the metals present in the protein samples before crystallization, or that additional metals are present but not represented in the structural models. Spectroscopic results were integrated with automated crystallographic validation metrics, including real-space Z-difference (RSZD) analysis and systematic rerefinement, to evaluate atomic-number mismatch at metal sites. PIXE and XRFS show strong agreement for dominant elemental signals and provide complementary, scalable approaches for identifying suspect metal assignments. This work does not address physiological or functional metalation but instead highlights a widespread data integrity issue in deposited macromolecular structures, PDB-wide. These results establish an experimentally corroborated link between elemental identity and crystallographic validation metrics, enabling the large-scale detection of chemically inconsistent annotations in structural databases used for computational modeling and machine learning.

Crystallization↗

PSpecteR: A User-Friendly and Interactive Application for Visualizing Top-Down and Bottom-Up Proteomics Data in R

Visual examination of mass spectrometry data is necessary to assess data quality and to facilitate data exploration. Graphics provide the means to evaluate spectral properties, test alternative peptide/protein sequence matches, prepare annotated spectra for publication, and fine-tune parameters during wet lab procedures. Visual inspection of MS data is hindered by proprietary proteomics visualization software designed for particular workflows and academic software that lack visualization tools. We built PSpecteR, an open-source and interactive R Shiny web application to address these issues, with support for several steps of proteomics data processing, including: reading various mass spectrometry files, running open-source database search tools, labelling spectra with fragmentation patterns, testing post-translational modifications, plotting where identified fragments map to reference sequences, and visualizing algorithmic output and metadata. All figures, tables, and spectra are exportable within one easy-to-use graphical user interface. Our current software provides a flexible and modern R framework to support fast implementation of additional features. The open source code is readily available (https://github.com/EMSL-Computing/PSpecteR), and a PSpecteR Docker container (https://hub.docker.com/r/emslcomputing) is available for easy local installation.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

IsoForma: An R Package for Quantifying and Visualizing Positional Isomers in Top-Down LC-MS/MS Data

Proteoforms, the different forms of a protein with sequence variations including post-translational modifications (PTMs), execute vital functions in biological systems such as cell signaling and epigenetic regulation. Precisely defining the stoichiometry of PTMs has been challenging because, in the widely used bottom-up proteomics methods, the detection occurs at the peptide level and thus the link between peptides and their specific modification site is lost, resulting in proteoform ambiguity. Advances in top-down mass spectrometry (MS) technology have permitted the direct characterization of intact proteoforms and their exact number of modification sites, allowing for the relative quantification of positional isomers (PI). Proteins with positional isomers refers to proteoforms with identical total mass and set of modifications but varying PTM site combinations. The relative abundance of PI can be estimated by matching proteoform-specific fragment ions to top-down tandem MS (MS2) data to localize and quantify modifications. However, current approaches heavily rely on manual annotation. Here, we present IsoForma, an open-source R package for relative quantification of PI within a single tool. We benchmarked IsoForma’s performance against two existing workflows and highlight the similarity of the results and improvements in speed. Overall, IsoForma provides a streamlined process, reduces the time of conducting isoform-based analyses, and offers an essential framework for developing customized proteoform analysis workflows. Finally, the software is open source and available at https://github.com/EMSL-Computing/isoforma-lib.

59 BASIC BIOLOGICAL SCIENCES↗