Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “correlation coefficient”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

From observation to replication: machine-learning-driven quantification and replication of fine-scale fish kinematics and behavior

Long-term quantification of fish behavior is essential for aquatic ecology, wildlife telemetry, and biomechanical device development. However, the observation duration required to obtain reliable behavioral and kinematic metrics remains unclear, and few tools exist to physically reproduce natural swimming motion for controlled experimentation. We address these challenges by developing a generalizable framework that models behavioral reliability (Spearman–Brown reliability index) as a function of observation duration and derives metric-specific monitoring thresholds. Using juvenile white sturgeon (Acipenser transmontanus) as a case study, we demonstrate that the minimum duration needed for reliable estimates varies substantially across kinematic features: to exceed a reliability of 0.8, total distance traveled requires 12 days, average curvature (mm?¹) 15 days, tail-beat frequency (Hz) 8 days, and average speed (body length/s) 17 days. We further bridge digital analysis and physical testing by developing a hardware-in-the-loop simulator that reconstructs machine-learning-derived swimming kinematics with high fidelity (correlation coefficient 0.98–0.99, RMSE 1.22–1.27 mm over a 5-minute segment). This platform enables realistic, repeatable motion stimuli for evaluating aquatic sensing technologies and bio-integrated devices under controlled conditions. Together, these contributions provide a scalable approach for designing long-term behavioral studies and a data-driven connection between ecological observation and robotic experimentation.

Hwang, SungJoo↗

Interfacial thermal conductance between multi-layer graphene sheets and solid/liquid octadecane: A molecular dynamics study

Mixtures of paraffin and carbon nanofillers have promising potential for thermal storage, as paraffin (the matrix) possesses high latent heat and the nanofillers compensate for the low thermal conductivity (TC) of paraffin. Understanding thermal transport in these materials is essential for practical applications, as weak thermal transport hinders fast charge/discharge of thermal energy. Here, we use non-equilibrium molecular dynamics (NEMD) simulations to study the interfacial thermal conductance (ITC) between graphene sheets and octadecane (C18H38) matrix under the limiting conditions of the sheets being parallel or perpendicular to the direction of the imposed heat flux. The findings show that the systems containing thin graphene layers exhibit higher values of ITC. This study captures the asymptotic saturation of thermal conductance for the liquid phase of the perpendicular structure. Besides, given the greater number of structured layers of paraffin upon phase change, the ITC for the solid paraffin-graphene system is higher than the conductance of the liquid paraffin-graphene interface. We use the Pearson correlation coefficients of the vibrational power spectrum (VPS) of interfacing materials to explain the orders of magnitude variations of the observed ITC.

25 ENERGY STORAGE↗

Aerosol emissions from water-lean solvents for post-combustion CO 2 capture

Advanced water-lean solvents (WLS) for post-combustion CO 2 capture have been gaining interest due to their ability to reduce the parasitic penalty from energy needed for solvent regeneration. Commercial implementation of these novel CO 2 capture technologies hinges on successful control of amine emissions. RTI conducted a parametric study of fundamental and operational variables influence on overall amine aerosol and vapor emissions from our water-lean solvent eCO 2 Sol™ using our 6-kW equivalent bench-scale gas absorption system. The parametric testing used a simulated flue gas with 15 % CO 2 , 2.3–4.2 % H 2 O, and 0–6 ppm sulfite (SO 3 ) to examine the impact of the presence of aerosols to the capture performance and amine emissions from the system. The SO 3 reacts with water in the flue gas to create H 2 SO 4 , which forms liquid aerosol droplets and provide nucleation sites for growth of aerosols. Scanning Mobility Particle Sizer and Aerodynamic Particle Sizer instruments monitored the aerosol particle size distribution. Parametric testing results suggested that the presence of the aerosols in the flue gas could increase the overall amine emissions by 10X compared to the baseline emissions from WLS’s vapor pressure. Principal component analysis (PCA) and projection to latent squares (PLS) developed models to predict the aerosol-based amine emissions from process data. The predictive PLS model had a correlation coefficient (Q 2 ) of 0.92 and could predict the aerosol-based emissions from the NAS process with ±15 % accuracy (average absolute deviation, AAD). The PLS regression model also identified key variables affecting aerosol-based emissions from WLS.

42 ENGINEERING↗

Exploring drought-responsive crucial genes in Sorghum

Drought severely affects global food production. Sorghum is a typical drought-resistant model crop. Based on RNA-seq data for Sorghum with multiple time points and the gray correlation coefficient, this paper firstly selects candidate genes via mean variance test and constructs weighted gene differential co-expression networks (WGDCNs); then, based on guilt-by-rewiring principle, the WGDCNs and the hidden Markov random field model, drought-responsive crucial genes are identified for five developmental stages respectively. Enrichment and sequence alignment analysis reveal that the screened genes may play critical functional roles in drought responsiveness. A multilayer differential co-expression network for the screened genes reveals that Sorghum is very sensitive to pre-flowering drought. Furthermore, a crucial gene regulatory module is established, which regulates drought responsiveness via plant hormone signal transduction, MAPK cascades, and transcriptional regulations. The proposed method can well excavate crucial genes through RNA-seq data, which have implications in breeding of new varieties with improved drought tolerance.

60 APPLIED LIFE SCIENCES↗

Detection and attribution of long-term and fine-scale changes in spring phenology over urban areas: A case study in New York State

Spring phenology plays an essential role in climate change, terrestrial ecosystem, and public health. Field-based monitoring and understanding of changes in spring phenology for long periods and in large regions are challenging due to the limited in-site observations. Space-based remotely sensed observations offer great potentials for monitoring decadal spring phenology changes from regional to global scales. However, the coarse-scale remotely sensed observations are insufficient to capture fine-scale spring phenology dynamics, especially in urban areas, and this makes it challenging for understanding the combined effects of climate change and urbanization on spring phenology. We derived the start of phenology season (SOS) in New York State using 30 m Landsat observations from 1990 to 2015 to understand the impact of the environment and urbanization on SOS. The results show that SOS for different years reveals heterogeneous spatial distribution. Most regions of New York State have been experiencing significant spring phenology changes in form of earlier onset of vegetation greening, ranging from 0.2 to 0.6 day/year during 1990 to 2015, and this trend varies slightly with latitudes and urbanization levels. Further, spatial correlation analysis shows that the increase in temperature and urbanization could both promote the advancement of SOS. However, the effect of urbanization (partial correlation coefficient (R) ranges from −0.289 to −0.542) on SOS is greater than the effect of temperature (R ranges from 0.006 to −0.192). The study generates a high spatio-temporal resolution spring phenology dataset for ecological, environmental and public health studies, especially in urban areas, and reveals the importance of better accounting for the urbanization effects when quantifying the SOS dynamics in phenology models.

Landsat↗

Rapid Coal-Ash Characterization using Geophysical Methods & Machine Learning

Coal combustion products (CCP) are challenging to delineate in heterogeneous field settings. Conventional methods (test pits, coring, and laboratory analyses) are labor-intensive, slow, invasive, and provide sparse spatial coverage. This study evaluates whether rapid non-invasive geophysical screening methods—induced polarization (IP), magnetic susceptibility, and nuclear magnetic resonance (NMR) —combined with surface colorimetry (RGB_24), can discriminate CCP-soil mixtures and provide reliable estimates of CCP content. Laboratory measurements were collected on five CCP-soil mixtures (series) and modeled using (i) a linear baseline, (ii) a calibrated non-linear (power-mean) model, and (iii) a machine-learning (ML) Random Forest approach, with validation via leave-one-series-out and site-specific tests. Across the five series, individual signals—particularly IP and magnetic susceptibility—were strongly predictive of ash content but were consistently outperformed by combined models. The pooled calibrated non-linear and ML models captured the observed non-linearity and achieved high accuracy and precision, improving on linear fits. Colorimetry showed the weakest direct relationship with ash content for the tested samples but improved performance when included in multi-signal models. At pre-selected 3.5% decision threshold, calibrated and ML approaches yielded near-perfect classification (Matthews correlation coefficient ˜ 1), suggesting strong practical operability for field screening. Additionally, field-analog tests highlighted the role of endmembers—accuracy declined without access to end-member measurements but was largely recovered by collecting a minimal labeled pair for local recalibration. With end members, accuracy remained high. Globally trained models performed well on three operational unknowns; however, series-specific refits provided the most accurate predictions. Overall, these results highlight the potential of combining rapid geophysics and minimal local calibration for improved coal-ash delineation.

Peshtani, Klaudio↗

Deep-Learning-Derived Evaluation Metrics Enable Effective Benchmarking of Computational Tools for Phosphopeptide Identification

Tandem mass spectrometry (MS/MS)-based phosphoproteomics is a powerful technology for global phosphorylation analysis. However, applying four computational pipelines to a typical mass spectrometry (MS)-based phosphoproteomic dataset from a human cancer study, we observed a large discrepancy among the reported phosphopeptide identification and phosphosite localization results, underscoring a critical need for benchmarking. While efforts have been made to compare performance of computational pipelines using data from synthetic phosphopeptides, evaluations involving real application data have been largely limited to comparing the numbers of phosphopeptide identifications due to the lack of appropriate evaluation metrics. We investigated three deep learning-derived features as potential evaluation metrics: phosphosite probability, Delta RT and spectral similarity. Predicted phosphosite probability is computed by MusiteDeep, which provides high accuracy as previously reported; Delta RT is defined as the absolute retention time (RT) difference between RTs observed and predicted by AutoRT; and spectral similarity is defined as the Pearson’s correlation coefficient between spectra observed and predicted by pDeep2. Using a synthetic peptide dataset, we found that both Delta RT and spectral similarity provided excellent discrimination between correct and incorrect peptide-spectrum matches (PSMs) both when incorrect PSMs involved wrong peptide sequences and even when incorrect PSMs were caused by only incorrect phosphosite localization. Based on these results, we used all the three deep learning-derived features as evaluation metrics to compare different computational pipelines on diverse set of phosphoproteomic datasets and showed their utility in benchmarking performance of the pipelines. The benchmark metrics demonstrated in this study will enable users to select computational pipelines and parameters for routine analysis of phosphoproteomics data and will offer guidance for developers to improve computational methods.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning for reactor power monitoring with limited labeled data

Real-time reactor power monitoring is critical for a variety of nuclear applications, spanning safety, security, operations, and maintenance. While machine learning methods have shown promise in monitoring reactor power levels, there is limited research on their efficacy in label-starved environments. The goal of this work is to assess the feasibility of classifying nuclear reactor power level using multisource data in scenarios with limited labels. Data were collected using low-resolution multisensors at four nuclear reactor facilities: two large research reactors and two TRIGA reactors. Within each pair, one reactor dataset served as the source and the other as the target in a transfer learning paradigm. Twenty-three supervised models were trained on labeled sequences of magnetic field and acceleration data from each of the target sites. Self-learning and transfer learning methods were applied to the top performing models to assess their classification performance with increasing amounts of labeled data. While reactor power level classification was achieved with a Matthews Correlation Coefficient of up to 0.739 ± 0.003 and 0.622 ± 0.009 with only 400 sequences per power state for the large research reactor and TRIGA target sites, respectively, self-learning and transfer learning leveraging source site data did not improve target classification performance. These findings suggest that alternative methods, such as higher sensitivity sensors, digital twins, or the use of physics-informed models, are required to enable high-performance classification in machine learning approaches to reactor monitoring with a dearth of target ground truth.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Characterization of dynamics and decay in the St. Benedict Paul trap

The St. Benedict experiment includes a linear Paul trap designed to measure the beta-neutrino angular correlation coefficient 𝑎 𝛽ν of mixed superallowed 𝛽-decay transitions between mirror nuclei via coincidence detection of 𝛽-particles and recoiling ions. The emitted 𝛽 particle and daughter ion are detected with plastic scintillators and micro-channel plate detectors, respectively, allowing for accurate measurements of their time-of-flights. From the shape of the coincidence time-of-flight distribution, a value of 𝑎 𝛽ν can be determined. This manuscript presents detailed simulations of the St. Benedict Paul trap, with a focus on ion cloud dynamics and recoiling daughter trajectories.

Linear Paul trap↗

Isoscalar and isovector giant resonances in 44 Ca, 54 Fe, 64,68 Zn and 56,58,60,68 Ni

We have studied the uncharacteristic behavior of the measured values of the isoscalar and isovector centroid energies, E CEN , of nuclear giant resonances of multipolarity from L = 0 to L = 3 in 44 Ca, 54 Fe, 64,68 Zn and 56,58,60,68 Ni. For this purpose, we carried out calculations of E CEN within the spherical Hartree-Fock (HF)-based random phase approximation (RPA) theory with 33 distinct Skyrme-like effective nucleon-nucleon interactions. We have also determined the Pearson linear correlation coefficients between centroid energies, obtained from the HF-RPA, and the various properties of nuclear matter (NM) of each interaction and determined the sensitivity of E CEN to NM properties. We compared the theoretical values of E CEN obtained from the HF-RPA calculations with experimental data and discuss the results, pointing out significant disagreements between theoretical and experimental values. We note in particular, that we obtain good agreement for the theoretical E CEN of the isovector giant dipole resonance and the available experimental data.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Pseudo-viscous modeling of transport in dense granular flows for thermal energy storage applications

Dense, granular flows were examined to effectively capture and model bulk viscous properties in thin packed beds. A modified Couette cell with particle image velocimetry was used to experimentally determine pseudo-viscosity properties of four particulate media with varying morphologies: (1) iron oxide-coated SiO 2 particles, (2) CARBOBEAD CP30-60 particles, (3) CARBOBEAD CP40-100 particles, and (4) Al 2 O 3 beads. The pseudo-viscosity functions were fitted using a power law to correlate the measured shear stress as a function of measured shear rate. The pseudo-viscous functions were used as inputs to computation fluid dynamics models for a single-phase viscous fluid to predict granular flow profiles. Steady-state free surface velocity profiles at angular velocities <7 rad/s predicted by the model were in good agreement with the experimental particle image velocimetry measurements, resulting in Pearson correlation coefficients of 0.97 for iron-oxide coated SiO 2 particles and 0.95 for CP30-60 particles. As a result, this alternative approach to measuring pseudo-viscous properties under shearing and modeling bulk transport behavior of granular flow using computation fluid dynamics model offered significant reduction in computational load compared to discrete element methods.

14 SOLAR ENERGY↗

Citizen science coupled with machine learning to quantify green-blue infrastructure cooling potential in Maricopa County, Arizona

Here, this study investigates the spatiotemporal cooling performance of green and blue infrastructure (GBI) in the Dobson Ranch urban neighborhood in Phoenix, Arizona. We leveraged citizen science near-surface (2 m) air temperature (Tair) measurements to train a highly accurate Tair predicting LightGBM machine learning model (R 2 : 0.986, MAE: 0.251 °C, RMSE: 0.585 °C). On June 16, 2024, the park area exhibited approximately 1 °C cooling effect (relative to the neighborhood mean) during both day and night. In contrast, the nearby artificial lake exhibited a stronger cooling effect of 2.4 °C during the day but a slight warming of 0.3 °C at night. At 00:00, locations 50 m downwind of the park were 0.3 °C warmer than the park, while locations 50 m upwind were 0.8 °C warmer. At 11:00, we observed that the downwind area is 0.8 °C cooler and the upwind area is 0.6 °C warmer—at the same 50 m distances relative to the park. We also observed 1 °C cooler and warmer effects respectively at the same 50 m downwind and upwind locations at 19:00 on June 17, 2024. Our data-driven analysis highlights potential limitations of car-traverse measurements, showing that failure to account for temporal variations during the traverse can lead to overestimation of Tair at night and underestimation during the day. Our analysis also showed only a weak correlation (coefficient: 0.48) between Landsat-derived land surface temperature (LST) and model predicted Tair at the time of the local Landsat overpass (∼11.00). This highlights the potential error of relying solely on LST for human thermal exposure analysis—particularly within the heterogenous built-environment.

54 ENVIRONMENTAL SCIENCES↗

Performance evaluation of CMIP6 models on the Arctic-Siberian Plain teleconnection affecting the East Asian heat waves

The frequency and intensity of summer heat waves in East Asia have increased sharply in recent decades, significantly impacting public health and the economy. The Arctic-Siberian Plain (ASP) teleconnection pattern has been identified as a key driver, with ASP warming amplifying atmospheric circulation patterns conducive to extreme temperatures. This study evaluates the ability of Coupled Model Inter-comparison Project phase 6 models to simulate the ASP pattern across interannual variability (IAV) and intra-seasonal variability (ISV) timescales using the Common Basis Function method. The multi-model mean shows statistically significant pattern correlations with ERA5 reanalysis, with correlation coefficients of 0.90 and 0.99 for IAV and ISV, respectively. While the ASP pattern is generally well captured, models exhibit substantial inter-model diversity in the intensity and position of anticyclonic anomalies over the ASP and East Asia. Models with ASP pattern variability similar to reanalysis better reproduce extreme East Asian temperatures, whereas those over- or underestimating ASP variability exhibit lower skill. These performance differences are related to differences in simulating key variables associated with the development of the ASP pattern. Our findings highlight the role of the ASP pattern in modulating extreme heat events, as models with improved ASP simulations align more closely with observed temperature extremes. Refining ASP representations in models could enhance seasonal heat wave predictions, improving climate adaptation strategies.

Arctic-Siberian Plain (ASP)↗

Differential localization patterns of pyruvate kinase isoforms in murine naïve, formative, and primed pluripotent states

Highlights: • PKM1/2 protein abundance is greater in formative mEpiLCs compared to naïve mESCs or primed mEpiSCs. The ratio of PKM1/2 is maintained. • PKM1/2 are localized in both nuclear and cytoplasmic regions relative to GAPDH and OCT4 across the pluripotent continuum. • PKM1 localization is strongly correlated to OCT4 localization, and moderately to GAPDH in formative mEpiLCs. • PKM1/2 localization is moderately correlated to OCT4 and GAPDH localization in naïve mESCs. • PKM1/2 localization are moderately correlated to GAPDH localization in primed mEpiSCs. Mouse embryonic stem cells (mESCs) and mouse epiblast stem cells (mEpiSCs) represent opposite ends of the pluripotency continuum, referred to as naïve and primed pluripotent states, respectively. These divergent pluripotent states differ in several ways, including growth factor requirements, transcription factor expression, DNA methylation patterns, and metabolic profiles. Naïve cells employ both glycolysis and oxidative phosphorylation (OXPHOS), whereas primed cells preferentially utilize aerobic glycolysis, a trait shared with cancer cells referred to as the Warburg Effect. Until recently, metabolism has been regarded as a by-product of cell fate, however, evidence now supports metabolism as being a driver of stem cell state and fate decisions. Pyruvate kinase muscle isoforms (PKM1 and PKM2) are important for generating and maintaining pluripotent stem cells (PSCs) and mediating the Warburg Effect. Both isoforms catalyze the final, rate limiting step of glycolysis, generating adenosine triphosphate and pyruvate, however, the precise role(s) of PKM1/2 in naïve and primed pluripotency is not well understood. The primary objective of this study was to characterize the cellular expression and localization patterns of PKM1 and PKM2 in mESCs, chemically transitioned epiblast-like cells (mEpiLCs) representing formative pluripotency, and mEpiSCs using immunoblotting and confocal microscopy. The results indicate that PKM1 and PKM2 are not only localized to the cytoplasm, but also accumulate in differential subnuclear regions of mESC, mEpiLCs, and mEpiSCs as determined by a quantitative confocal microscopy employing orthogonal projections and airyscan processing. Importantly, we discovered that the subnuclear localization of PKM1/2 changes during the transition from mESCs, mEpiLCs, and mEpiSCs. Finally, we have comprehensively validated the appropriateness and power of the Pearson's correlation coefficient and Manders's overlap coefficient for assessing nuclear and cytoplasmic protein colocalization in PSCs by immunofluorescence confocal microscopy. We propose that nuclear PKM1/2 may assist with distinct pluripotency state maintenance and lineage priming by non-canonical mechanisms. These results advance our understanding of the overall mechanisms controlling naïve, formative, and primed pluripotency.

60 APPLIED LIFE SCIENCES↗

Reynolds-number scaling of wall-pressure–velocity correlations in wall-bounded turbulence

Wall-pressure fluctuations are a practically robust input for real-time control systems aimed at modifying wall-bounded turbulence. The scaling behaviour of the wall-pressure–velocity coupling requires investigation to properly design a controller with such input data so that it can actuate upon the desired turbulent structures. A comprehensive database from direct numerical simulations (DNS) of turbulent channel flow is used for this purpose, spanning a Reynolds-number range$Re_\tau \approx 550\unicode{x2013}5200$. Spectral analysis reveals that the streamwise velocity is most strongly coupled to the linear term of the wall pressure, at a Reynolds-number invariant distance-from-the-wall scaling of$\lambda _x/y \approx 14$(and$\lambda _x/y \approx 8$for the wall-normal velocity). When extending the analysis to both homogeneous directions in$x$and$y$, the peak coherence is centred at$\lambda _x/\lambda _z \approx 2$and$\lambda _x/\lambda _z \approx 1$for$p_w$and$u$, and$p_w$and$v$, respectively. A stronger coherence is retrieved when the quadratic term of the wall pressure is concerned, but there is only little evidence for a wall-attached-eddy type of scaling. An experimental dataset comprising simultaneous measurements of wall pressure and velocity complements the DNS-based findings at one value of$Re_\tau \approx 2$k, with ample evidence that the DNS-inferred correlations can be replicated with experimental pressure data subject to significant levels of (acoustic) facility noise. It is furthermore shown that velocity-state estimations can be achieved with good accuracy by including both the linear and quadratic terms of the wall pressure. An accuracy of up to 72 % in the binary state of the streamwise velocity fluctuations in the logarithmic region is achieved; this corresponds to a correlation coefficient of$\approx$0.6. This thus demonstrates that wall-pressure sensing for velocity-state estimation – e.g. for use in real-time control of wall-bounded turbulence – has merit in terms of its realization at a range of Reynolds numbers.

Mechanics↗

Automated Coupling of Nanodroplet Sample Preparation with Liquid Chromatography–Mass Spectrometry for High-Throughput Single-Cell Proteomics

Single-cell proteomics can provide critical biological insight into the cellular heterogeneity that is masked by bulk-scale analysis. Here, we have developed a nanoPOTS (nanodroplet processing in one pot for trace samples) platform and demonstrated its broad applicability for single-cell proteomics. However, because of nanoliter-scale sample volumes, the nanoPOTS platform is not compatible with automated LC-MS systems, which significantly limits sample throughput and robustness. To address this challenge, we have developed a nanoPOTS autosampler allowing fully automated sample injection from nanowells to LC-MS systems. We also developed a sample drying, extraction, and loading workflow to enable reproducible and reliable sample injection. The sequential analysis of 20 samples containing 10 ng tryptic peptides demonstrated high reproducibility with correlation coefficients of >0.995 between any two samples. The nanoPOTS autosampler can provide analysis throughput of 9.6, 16, and 24 single cells per day using 120, 60, and 30 min LC gradients, respectively. As a demonstration for single-cell proteomics, the autosampler was first applied to profiling protein expression in single MCF10A cells using a label-free approach. At a throughput of 24 single cells per day, an average of 256 proteins was identified from each cell and the number was increased to 731 when the Match Between Runs algorithm of MaxQuant was used. Using a multiplexed isobaric labeling approach (TMT-11plex), ~77 single cells could be analyzed per day. We analyzed 152 cells from three acute myeloid leukemia cell lines, resulting in a total of 2558 identified proteins with 1465 proteins quantifiable (70% valid values) across the 152 cells. These data showed quantitative single-cell proteomics can cluster cells to distinct groups and reveal functionally distinct differences.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Examination of How Well Long-Range-Corrected Density Functionals Satisfy the Ionization Energy Theorem

For this work, we calculated the vertical ionization energies (VIE) of 99 species in two ways to examine the accuracy of several long-range-corrected (LC) hybrid meta functionals in comparison with a gradient approximation (GA), global hybrids, and doubly hybrids. In the category of LC functionals, we examined both those with meta ingredients (i.e., that depend on the kinetic energy density) and those without them. The LC-hybrid meta functionals examined are M11, revM11, M11plus, and ωB97M-V. The reference data used to assess accuracy consist of 95 molecules and 4 atoms in the GW100 set. The two methods studied are the ΔSCF method (involving the difference of neutral and cation self-consistent field (SCF) energies) and the ionization energy theorem (involving the orbital energy of the highest occupied molecular orbital, HOMO). We calculated linear correlation coefficients (r 2 ) and mean absolute deviations (MADs) between each approach and the reference VIE value from the CCSD(T)/def2-TZVPP level of theory. We compared the new LC-hybrid meta calculations to calculations with the 10 functionals in a previous VIE study by Brémond et al. and to the calculations with LC-BLYP (LC-Becke, Lee–Yang–Parr), CAM-B3LYP (Coulomb-attenuating-method Becke-3-parameter Lee–Yang–Parr), LC-ωHPBE, and ωB97X-D. The results show that Minnesota LC-hybrid meta functionals have the smallest mean absolute deviation of ionization energy theorem VIEs with the reference data; the LC-ωHPBE functional also does quite well in this test. This is very encouraging and indicates that LC-hybrid meta functionals would be the best starting points for the tuning strategy that has been shown to be a very good procedure for improving time-dependent density functional calculations, and it also helps explain the good success of LC-hybrid meta functionals for molecular excitation energies.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Alchemical Free Energy Estimators and Molecular Dynamics Engines: Accuracy, Precision, and Reproducibility

The binding free energy between a ligand and its target protein is an essential quantity to know at all stages of the drug discovery pipeline. Assessing this value computationally can offer insight into where efforts should be focused in the pursuit of effective therapeutics to treat a myriad of diseases. In this work, we examine the computation of alchemical relative binding free energies with an eye for assessing reproducibility across popular molecular dynamics packages and free energy estimators. The focus of this work is on 54 ligand transformations from a diverse set of protein targets: MCL1, PTP1B, TYK2, CDK2, and thrombin. These targets are studied with three popular molecular dynamics packages: OpenMM, NAMD2, and NAMD3 alpha. Trajectories collected with these packages are used to compare relative binding free energies calculated with thermodynamic integration and free energy perturbation methods. The resulting binding free energies show good agreement between molecular dynamics packages with an average mean unsigned error between them of 0.50 kcal/mol. The correlation between packages is very good, with the lowest Spearman’s, Pearson’s and Kendall’s tau correlation coefficients being 0.92, 0.91, and 0.76, respectively. Agreement between thermodynamic integration and free energy perturbation is shown to be very good when using ensemble averaging.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗