Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Correlation coefficient”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Assessing free tropospheric quasi-equilibrium for different GCM resolutions using a cloud-resolving model simulation of tropical convection

Abstract This study examines the free-tropospheric quasi-equilibrium at different global climate model (GCM) resolutions using the simulation of tropical convection by a cloud-resolving model during the Tropical Western Pacific International Cloud Experiment. The simulated dynamic and thermodynamic fields within the model domain are averaged over subdomains of different sizes equivalent to different GCM resolutions. These coarse-grained fields are then used to compute CAPE and its change with time, and their relationships with simulated convection. Results show that CAPE change with time is controlled predominantly by variations of thermodynamic properties in the planetary boundary layer for all subdomain sizes ranging from 64 to 4 km. Lag correlation analysis shows that CAPE generation by the free-tropospheric dynamical advection (dCAPE ls ) leads convective precipitation but is in phase with convective mass flux at 600 mb and 500 mb vertical velocity for all subdomain sizes. However, the correlation coefficients and regression slopes decrease as the subdomain size decreases for subdomain sizes smaller than 16 km. This is probably due to increased randomness of convection and more scale-dependence of the relationships when the subdomain size reaches the grey zone. By examining the sensitivity of the relationships of convection with dCAPE ls to temporal scales in different subdomain size, it shows that the quasi-equilibrium between dCAPE ls and convection holds well for timescales of 30 min or longer at all subdomain sizes. These results suggest that the free tropospheric quasi-equilibrium assumption may still be useable even for GCM resolutions in the grey zone.

54 ENVIRONMENTAL SCIENCES↗

The substantial role of May soil temperature over Central Asia for summer surface air temperature variation and prediction over Northeastern China

The slowly varying soil temperature can exert local and nonlocal influences on regional climate system, and may thus provide a critical source of subseasonal-to-seasonal climate prediction. In this study, we identify that soil temperature in May over the key region of Central Asia (42 °N–50 °N, 62 °E–80 °E, KRCA) from Noah, Mosaic, CLM and ERA-interim datasets is closely linked to variations of the surface air temperature, daily maximum temperature and hot days over Northeastern China in summer (June-July-August), with correlation coefficients of regional average detrended time series ranging from 0.42 to 0.54, and all significant at the 99% confidence level for the period of 1979-2018. The possible physical mechanism behind the substantial downstream impacts of soil temperature over Central Asia are explored via diagnostical analysis combined with regional climate model experiments. Warmer soil temperature in May over the KRCA tends to cause positive anomalies of geopotential height in summer over Northeastern China through the Rossby wave propagation, and associated stronger subsidence warming, less cloud cover, more solar radiation reaching the surface, higher planetary boundary layer, and stronger thermal advection at 850 hPa, which provide favorable conditions for warmer surface air temperature particularly in the daytime as well as more hot days. Here this study further reveals that soil temperature over Central Asia in May makes important contribution to prediction of summer surface air temperature, daily maximum temperature and hot days over Northeastern China in terms of regional average time series and spatial patterns. Our findings highlight the previously-unknown substantial role of antecedent soil temperature condition over Central Asia for summer surface air temperature variation and prediction over Northeastern China.

54 ENVIRONMENTAL SCIENCES↗

High precision control and deep learning-based corn stand counting algorithms for agricultural robot

This paper presents high precision control and deep learning-based corn stand counting algorithms for a low-cost, ultra-compact 3D printed and autonomous field robot for agricultural operations. Currently, plant traits, such as emergence rate, biomass, vigor, and stand counting, are measured manually. This is highly labor-intensive and prone to errors. The robot, termed TerraSentia, is designed to automate the measurement of plant traits for efficient phenotyping as an alternative to manual measurements. In this paper, we formulate a Nonlinear Moving Horizon Estimator that identifies key terrain parameters using onboard robot sensors and a learning-based Nonlinear Model Predictive Control that ensures high precision path tracking in the presence of unknown wheel-terrain interaction. Moreover, we develop a machine vision algorithm designed to enable an ultra-compact ground robot to count corn stands by driving through the fields autonomously. The algorithm leverages a deep network to detect corn plants in images, and a visual tracking model to re-identify detected objects at different time steps. We collected data from 53 corn plots in various fields for corn plants around 14 days after emergence (stage V3 - V4). The robot predictions have agreed well with the ground truth with C robot =1.02×C human -0.86 and a correlation coefficient R=0.96. The mean relative error given by the algorithm is -3.78%, and the standard deviation is 6.76%. These results indicate a first and significant step towards autonomous robot-based real-time phenotyping using low-cost, ultra-compact ground robots for corn and potentially other crops.

97 MATHEMATICS AND COMPUTING↗

High fidelity simulations of contaminant dispersion in an urban environment with comparison to magnetic resonance imaging measurements

The dispersion of a contaminant in an urban environment has the potential to impact a large population of people. In this work, a complex urban canopy flow based on the Oklahoma City downtown business district circa 2003 is studied using Magnetic Resonance Imaging (MRI) and high-fidelity Large Eddy Simulations (LES). MRI is a novel experimental technique that can provide high-resolution measurements in four dimensions (three spatial and temporal) for lab scale models. The experiments and simulations use the same geometry and boundary conditions providing a one-to-one comparison of the two methods. Results are presented on the time-averaged velocity and concentration fields, the temporal dynamics of the concentration plumes for a transient release, and a novel Cloud Identification Algorithm that can separate plumes produced by periodic contaminant releases used for ensemble averaging over many releases. The MRI and LES datasets both include millions of measurement voxels and the comparisons highlight the complex 3D nature of the flow including strong vertical velocities in spanwise street canyons and flow acceleration in streamwise street canyons. The concentration fields are qualitatively similar albeit the LES shows larger dispersion. A quantitative analysis with performance measures compares the datasets pointwise and demonstrates that the two 3D datasets are similar with respect to many measures including a fractional bias of 0.02 (ideal=0.0), correlation coefficient of 0.87 (ideal = 1.0), and the fraction points within a factor of 2 is 0.98 (ideal = 1.0). Plume analysis compares the arrival and residence time of contaminant and is found to vary significantly with location within the urban environment with arrival times between 0 and 1.25 and differences within the contaminant cloud less than 10% at most locations.

54 ENVIRONMENTAL SCIENCES↗

A Performance Comparison of Low-Cost Near-Infrared (NIR) Spectrometers to a Conventional Laboratory Spectrometer for Rapid Biomass Compositional Analysis

The performance of a conventional laboratory near-infrared (NIR) spectrometer and two NIR spectrometer prototypes (a Texas Instruments NIRSCAN Nano evaluation model (EVM) and an InnoSpectra NIR-M-R2 spectrometer) are compared by collecting reflectance spectra of 270 well-characterized herbaceous biomass samples, building calibration models using the partial least squares (PLS-2) algorithm to predict five constituents of the samples from the reflectance spectra, and comparing the resulting model statistics. The prediction models developed using spectra from the Foss XDS spectrometer were slightly better than the prediction models developed using spectra from either the TI NIRSCAN Nano EVM and the InnoSpectra NIR-M-R2 as measured by the root mean square error (RMSECV) and the correlation coefficient (R 2 _cv) for “leave-one-out” cross-validation (CV). The models built from the two prototype units were not statistically significantly different from each other (p = 0.05). The Foss spectrometer has a larger wavelength range (400–2500 nm) compared with the two prototypes (900–1700 nm). When the spectra from the Foss XDS spectrometer were truncated so their wavelength range matched the wavelength range of the two prototype units, the resulting model was not statistically significantly different from the models from either prototype.

47 OTHER INSTRUMENTATION↗

Data Mining and Visualization of High-Dimensional ICME Data for Additive Manufacturing

Integrated computational materials engineering (ICME) methods combining CALPHAD with process-based simulations can produce rich, high-dimensional data for alloy and process design. In ICME methods for metallurgical applications, the visualization and interpretation of such high-dimensional data has previously been through heat maps represented in 2 or 3 dimensions. While such an approach is ideal when one variable is varied at a time, in the case of high-dimensional data with multiple variables varied simultaneously, as is the case in additive manufacturing, interpreting the trends through two- or three-dimensional heat maps becomes challenging. Here, we propose a strategy of mixed visual data mining and quantitative analysis for high-dimensional metallurgical and process data using high-throughput thermodynamic calculations. Two case studies show the application of the proposed approach. The first case study investigated the effects of feedstock chemistry on the δ ferrite formation in 316L stainless steel powders used for binder jet additive manufacturing. The second case study linked Scheil–Gulliver calculations to a process model for dissimilar joining of aluminum alloys 5356 and 6111 during laser hot-wire additive manufacturing. Both cases contained thousands of calculated data points, showcasing the utility of visual data analysis through parallel coordinate plotting, Pearson correlation coefficient matrices, and scatter matrices compared to traditional process maps. These visualization techniques can be extended to many additive manufacturing problems to capture process–structure–property relationships for additively manufactured components.

36 MATERIALS SCIENCE↗

Calcium fluoride as a dominating matrix for quantitative analysis by laser ablation-inductively coupled plasma-mass spectrometry (LA-ICP-MS): A feasibility study

Here, calcium fluoride formed by the reaction between ammonium bifluoride and calcium chloride was investigated as a dominating matrix for quantitative analysis by laser ablation inductively coupled plasma mass spectrometry (LA-ICP-MS). Transformation from a solid sample to the calcium fluoride-based matrix permitted quantitative analysis based on calibration standards made from elemental standards. A low abundance stable calcium isotope, i.e. 44 Ca + , was monitored as the internal standard for quantitative analysis by LA-ICP-MS. Correlation coefficient factors for multiple elements were obtained with values over 0.999. The results for multiple elements in a certified reference material of soil (NIST SRM 2710a) agreed with the certified values in the range of expanded uncertainty, indicating the present method was valid for quantitation of elements in solid samples.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The effect of price-based demand response on carbon emissions in European electricity markets: The importance of adequate carbon prices

Price-based demand response (PBDR) has recently been attributed great economic but also environmental potential. However, the determination of its short-term effects on carbon emissions requires the knowledge of marginal emission factors (MEFs), which compared to grid mix emission factors (XEFs), are cumbersome to calculate due to the complex characteristics of national electricity markets. This study, therefore, proposes two merit order-based methods to approximate hourly MEFs and applies them to readily available datasets from 20 European countries for the years 2017–2019. Based on the calculated electricity prices, standardized daily load shifts were simulated which indicated that carbon emissions increased for 8 of the 20 countries and by 2.1% on average. Thus, under specific circumstances, PBDR leads to carbon emissions increases, mainly due to the economic advantage fuel sources such as lignite and coal have in the merit order. MEF-based load shifts reduced the mean resulting carbon emissions by 35%, albeit with 56% lower monetary cost savings compared to price-based load shifts. Finally, by repeating the load shift simulations for different carbon price levels, the impact of the carbon price on the resulting carbon emissions was analyzed. The Spearman correlation coefficient between carbon intensity and marginal cost along the German merit order substantially increased with increasing carbon price. The coefficients were -0.13 for the 2019 carbon price of 24.9 €/t, 0 for 42.6 €/t, and 0.4 for 100.0 €/t. Therefore, with adequate carbon prices, PBDR can be an effective tool for both economical and environmental improvement.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Evaluation of CMIP6 models in simulating the statistics of extreme precipitation over Eastern Africa

We report the Eastern Africa region experiences frequent extreme precipitation events that can cause destruction of property and environment, and loss of lives. Thus, there is a need to understand how these events may change in the future and how well the global climate models that are used to make projections can simulate precipitation extremes in this region before they can be used in downscaling or flood and drought impact assessment studies. In this work, we evaluated the ability of sixteen Coupled Model Intercomparison Project Phase 6 (CMIP6) models to simulate present-day precipitation extremes over the Eastern Africa region during the two rainy seasons (March–May and September–November). We used nine extreme precipitation indices (including seven (one) indices of wet (dry) extremes) defined by the Expert Team on Climate Change Detection and Indices. The CMIP6 models were evaluated against two gridded observation datasets: Global Precipitation Climatology Project One-Degree Daily Dataset and Tropical Rainfall Measuring Mission Multi-satellite Precipitation Analysis 3B42. Three model performance metrics (percentage bias, normalized root-mean-square error, and pattern correlation coefficient) were employed to further assess the strengths and weakness of the models. Our results show that the multi-model ensemble mean generally provides a better representation of observed precipitation and related extremes compared to individual models when considering all metrics and seasons. Several consistent biases are evident across CMIP6 models, which tend to overestimate the total-wet day precipitation and consecutive wet days, and underestimate very wet days and maximum 5-day precipitation in both seasons. Furthermore, no single model consistently performs best, model performance varies with the season and index under consideration and is generally independent of horizontal resolution.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Scaling up carboxylic acid production from cheese whey and brewery wastewater via methane-arrested anaerobic digestion

In a circular economy, organic waste streams are valuable resources for sustainably producing chemicals and fuels. This study investigates a new methane-arrested anaerobic digestion (MAAD) process that converts high-strength cheese whey and brewery wastewater into carboxylic acids. The process was developed and optimized under various bench-scale semi-continuous (fed-batch) operating conditions (e.g., retention time, organic loading rate, pH, and feed/harvest frequency). The MAAD responses to various control-failure scenarios were also systematically investigated. The highest total acid productivity was 26g/(L liq ·d) with a substrate conversion of 0.79g COD digested /g COD fed at a hydraulic retention time (HRT) of approximately 2 d in a 14-L digester. The most stable conditions for digester operation (HRT 3 d at pH 6.0 and 40°C) were selected for process scale-up to 100gal (~380 L). Semi-continuous, pilot-scale MAAD successfully produced a total acid concentration of 40.6+/-1.1 g/L with 8.1% acetic acid, 45.1% butyric acid, and 44.5% lactic acid. The links between wastewater characteristics, operation mode, digester scale, and microbial community structure were statistically analyzed. The results show genera Sporolactobacillus and Clostridium positively correlate with butyric acid production (Pearson correlation coefficient>0.5). Moreover, four kinetic models were developed and fit to batch MAAD experimental datasets (R 2 >95%) and were successfully applied to predict the total acid production in both bench-and pilot-scale semi-continuous MAADs. In conclusion, this study shows MAAD has the potential for industrial-scale applications and is a robust platform to valorize low- or negative-value waste streams into high-value bioproducts.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Deep learning approaches to semantic segmentation of fatigue cracking within cyclically loaded nickel superalloy

Improvements to synchrotron-based micro-computed tomography scanning capabilities have gifted researchers the ability to characterize 4D material thermomechanical responses more thoroughly than ever before. These advancements, however, have brought about new challenges in analyzing the resulting deluge of data. We report on a nickel-based superalloy specimen imaged 26 times in-situ during cyclic loading at Argonne National Laboratory Advanced Photon Source beamline 1ID, in order to monitor crack growth within the microstructure. Therefore, several deep learning approaches which utilize convolutional neural networks are implemented to segment crack features from reconstructed tomography scans. U-Net architecture implementations are found to be especially effective, achieving IoU = 0.995 +/- 0.004 and Matthews correlation coefficient scores of Φ = 0.826 +/- 0.085. These advancements broaden possibilities for scientists seeking to automate segmentation analyses of similar large datasets.

36 MATERIALS SCIENCE↗

From observation to replication: machine-learning-driven quantification and replication of fine-scale fish kinematics and behavior

Long-term quantification of fish behavior is essential for aquatic ecology, wildlife telemetry, and biomechanical device development. However, the observation duration required to obtain reliable behavioral and kinematic metrics remains unclear, and few tools exist to physically reproduce natural swimming motion for controlled experimentation. We address these challenges by developing a generalizable framework that models behavioral reliability (Spearman–Brown reliability index) as a function of observation duration and derives metric-specific monitoring thresholds. Using juvenile white sturgeon (Acipenser transmontanus) as a case study, we demonstrate that the minimum duration needed for reliable estimates varies substantially across kinematic features: to exceed a reliability of 0.8, total distance traveled requires 12 days, average curvature (mm?¹) 15 days, tail-beat frequency (Hz) 8 days, and average speed (body length/s) 17 days. We further bridge digital analysis and physical testing by developing a hardware-in-the-loop simulator that reconstructs machine-learning-derived swimming kinematics with high fidelity (correlation coefficient 0.98–0.99, RMSE 1.22–1.27 mm over a 5-minute segment). This platform enables realistic, repeatable motion stimuli for evaluating aquatic sensing technologies and bio-integrated devices under controlled conditions. Together, these contributions provide a scalable approach for designing long-term behavioral studies and a data-driven connection between ecological observation and robotic experimentation.

Hwang, SungJoo↗

Interfacial thermal conductance between multi-layer graphene sheets and solid/liquid octadecane: A molecular dynamics study

Mixtures of paraffin and carbon nanofillers have promising potential for thermal storage, as paraffin (the matrix) possesses high latent heat and the nanofillers compensate for the low thermal conductivity (TC) of paraffin. Understanding thermal transport in these materials is essential for practical applications, as weak thermal transport hinders fast charge/discharge of thermal energy. Here, we use non-equilibrium molecular dynamics (NEMD) simulations to study the interfacial thermal conductance (ITC) between graphene sheets and octadecane (C18H38) matrix under the limiting conditions of the sheets being parallel or perpendicular to the direction of the imposed heat flux. The findings show that the systems containing thin graphene layers exhibit higher values of ITC. This study captures the asymptotic saturation of thermal conductance for the liquid phase of the perpendicular structure. Besides, given the greater number of structured layers of paraffin upon phase change, the ITC for the solid paraffin-graphene system is higher than the conductance of the liquid paraffin-graphene interface. We use the Pearson correlation coefficients of the vibrational power spectrum (VPS) of interfacing materials to explain the orders of magnitude variations of the observed ITC.

25 ENERGY STORAGE↗

Aerosol emissions from water-lean solvents for post-combustion CO 2 capture

Advanced water-lean solvents (WLS) for post-combustion CO 2 capture have been gaining interest due to their ability to reduce the parasitic penalty from energy needed for solvent regeneration. Commercial implementation of these novel CO 2 capture technologies hinges on successful control of amine emissions. RTI conducted a parametric study of fundamental and operational variables influence on overall amine aerosol and vapor emissions from our water-lean solvent eCO 2 Sol™ using our 6-kW equivalent bench-scale gas absorption system. The parametric testing used a simulated flue gas with 15 % CO 2 , 2.3–4.2 % H 2 O, and 0–6 ppm sulfite (SO 3 ) to examine the impact of the presence of aerosols to the capture performance and amine emissions from the system. The SO 3 reacts with water in the flue gas to create H 2 SO 4 , which forms liquid aerosol droplets and provide nucleation sites for growth of aerosols. Scanning Mobility Particle Sizer and Aerodynamic Particle Sizer instruments monitored the aerosol particle size distribution. Parametric testing results suggested that the presence of the aerosols in the flue gas could increase the overall amine emissions by 10X compared to the baseline emissions from WLS’s vapor pressure. Principal component analysis (PCA) and projection to latent squares (PLS) developed models to predict the aerosol-based amine emissions from process data. The predictive PLS model had a correlation coefficient (Q 2 ) of 0.92 and could predict the aerosol-based emissions from the NAS process with ±15 % accuracy (average absolute deviation, AAD). The PLS regression model also identified key variables affecting aerosol-based emissions from WLS.

42 ENGINEERING↗

Exploring drought-responsive crucial genes in Sorghum

Drought severely affects global food production. Sorghum is a typical drought-resistant model crop. Based on RNA-seq data for Sorghum with multiple time points and the gray correlation coefficient, this paper firstly selects candidate genes via mean variance test and constructs weighted gene differential co-expression networks (WGDCNs); then, based on guilt-by-rewiring principle, the WGDCNs and the hidden Markov random field model, drought-responsive crucial genes are identified for five developmental stages respectively. Enrichment and sequence alignment analysis reveal that the screened genes may play critical functional roles in drought responsiveness. A multilayer differential co-expression network for the screened genes reveals that Sorghum is very sensitive to pre-flowering drought. Furthermore, a crucial gene regulatory module is established, which regulates drought responsiveness via plant hormone signal transduction, MAPK cascades, and transcriptional regulations. The proposed method can well excavate crucial genes through RNA-seq data, which have implications in breeding of new varieties with improved drought tolerance.

60 APPLIED LIFE SCIENCES↗

Detection and attribution of long-term and fine-scale changes in spring phenology over urban areas: A case study in New York State

Spring phenology plays an essential role in climate change, terrestrial ecosystem, and public health. Field-based monitoring and understanding of changes in spring phenology for long periods and in large regions are challenging due to the limited in-site observations. Space-based remotely sensed observations offer great potentials for monitoring decadal spring phenology changes from regional to global scales. However, the coarse-scale remotely sensed observations are insufficient to capture fine-scale spring phenology dynamics, especially in urban areas, and this makes it challenging for understanding the combined effects of climate change and urbanization on spring phenology. We derived the start of phenology season (SOS) in New York State using 30 m Landsat observations from 1990 to 2015 to understand the impact of the environment and urbanization on SOS. The results show that SOS for different years reveals heterogeneous spatial distribution. Most regions of New York State have been experiencing significant spring phenology changes in form of earlier onset of vegetation greening, ranging from 0.2 to 0.6 day/year during 1990 to 2015, and this trend varies slightly with latitudes and urbanization levels. Further, spatial correlation analysis shows that the increase in temperature and urbanization could both promote the advancement of SOS. However, the effect of urbanization (partial correlation coefficient (R) ranges from −0.289 to −0.542) on SOS is greater than the effect of temperature (R ranges from 0.006 to −0.192). The study generates a high spatio-temporal resolution spring phenology dataset for ecological, environmental and public health studies, especially in urban areas, and reveals the importance of better accounting for the urbanization effects when quantifying the SOS dynamics in phenology models.

Landsat↗

Rapid Coal-Ash Characterization using Geophysical Methods & Machine Learning

Coal combustion products (CCP) are challenging to delineate in heterogeneous field settings. Conventional methods (test pits, coring, and laboratory analyses) are labor-intensive, slow, invasive, and provide sparse spatial coverage. This study evaluates whether rapid non-invasive geophysical screening methods—induced polarization (IP), magnetic susceptibility, and nuclear magnetic resonance (NMR) —combined with surface colorimetry (RGB_24), can discriminate CCP-soil mixtures and provide reliable estimates of CCP content. Laboratory measurements were collected on five CCP-soil mixtures (series) and modeled using (i) a linear baseline, (ii) a calibrated non-linear (power-mean) model, and (iii) a machine-learning (ML) Random Forest approach, with validation via leave-one-series-out and site-specific tests. Across the five series, individual signals—particularly IP and magnetic susceptibility—were strongly predictive of ash content but were consistently outperformed by combined models. The pooled calibrated non-linear and ML models captured the observed non-linearity and achieved high accuracy and precision, improving on linear fits. Colorimetry showed the weakest direct relationship with ash content for the tested samples but improved performance when included in multi-signal models. At pre-selected 3.5% decision threshold, calibrated and ML approaches yielded near-perfect classification (Matthews correlation coefficient ˜ 1), suggesting strong practical operability for field screening. Additionally, field-analog tests highlighted the role of endmembers—accuracy declined without access to end-member measurements but was largely recovered by collecting a minimal labeled pair for local recalibration. With end members, accuracy remained high. Globally trained models performed well on three operational unknowns; however, series-specific refits provided the most accurate predictions. Overall, these results highlight the potential of combining rapid geophysics and minimal local calibration for improved coal-ash delineation.

Peshtani, Klaudio↗

Deep-Learning-Derived Evaluation Metrics Enable Effective Benchmarking of Computational Tools for Phosphopeptide Identification

Tandem mass spectrometry (MS/MS)-based phosphoproteomics is a powerful technology for global phosphorylation analysis. However, applying four computational pipelines to a typical mass spectrometry (MS)-based phosphoproteomic dataset from a human cancer study, we observed a large discrepancy among the reported phosphopeptide identification and phosphosite localization results, underscoring a critical need for benchmarking. While efforts have been made to compare performance of computational pipelines using data from synthetic phosphopeptides, evaluations involving real application data have been largely limited to comparing the numbers of phosphopeptide identifications due to the lack of appropriate evaluation metrics. We investigated three deep learning-derived features as potential evaluation metrics: phosphosite probability, Delta RT and spectral similarity. Predicted phosphosite probability is computed by MusiteDeep, which provides high accuracy as previously reported; Delta RT is defined as the absolute retention time (RT) difference between RTs observed and predicted by AutoRT; and spectral similarity is defined as the Pearson’s correlation coefficient between spectra observed and predicted by pDeep2. Using a synthetic peptide dataset, we found that both Delta RT and spectral similarity provided excellent discrimination between correct and incorrect peptide-spectrum matches (PSMs) both when incorrect PSMs involved wrong peptide sequences and even when incorrect PSMs were caused by only incorrect phosphosite localization. Based on these results, we used all the three deep learning-derived features as evaluation metrics to compare different computational pipelines on diverse set of phosphoproteomic datasets and showed their utility in benchmarking performance of the pipelines. The benchmark metrics demonstrated in this study will enable users to select computational pipelines and parameters for routine analysis of phosphoproteomics data and will offer guidance for developers to improve computational methods.

59 BASIC BIOLOGICAL SCIENCES↗