Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Correlation coefficient”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Validation of the IRI-2016 model with Indian NavIC data for future navigation applications

The position accuracy of Navigation with Indian Constellation (NavIC) system is affected by several sources of errors. Among them, the Ionospheric Time Delay (ITD) error is the most predominant one which depends upon the total electron content (TEC) present in the ionosphere. The ITD variations are more intense over low latitude regions due to the equatorial anomaly effects. Hence, modelling of ITD error is necessary. The International Reference Ionosphere (IRI)-2016 model is one of the standard global ionospheric models to estimate the Vertical TEC (VTEC). This paper discusses about the VTEC deviations due to the IRI-2016 model over low latitude Hyderabad station, Indian region using NavIC, Global Positioning System and Global Navigation Satellite System signals at corresponding Ionospheric Pierce Point latitude and longitudes for all months and various Kp indices during the low solar activity year 2017. In this work, cross correlation coefficient, the metric norm (L2N), Symmetric Kullbacke Leibler Distance metrics are used to evaluate the performance of IRI-2016 model TEC with NavIC, GPS and GLONASS data. From the results, it is found that TEC predicted by the IRI-2016 model produced smaller estimation errors with NavIC data over Indian region. The obtained results will be helpful for future updates of IRI model.

42 ENGINEERING↗

Climate-Related Trends of Within-Storm Intensities Using Dimensionless Temporal-Storm Distributions

Huff curves are probabilistic time distributions of rainfall expressed as dimensionless cumulative percentages of storm depth and duration. Previous studies have documented development factors, spatial robustness, and the utility of Huff curves in practical applications. However, the effects of trending rainfall on Huff curve intensity patterns have not yet been studied. As such, the goal of this paper is to fill this gap by studying Huff curve patterns in a watershed with demonstrated increasing trends of temperature and precipitation, with the intention that it can be generalized to other areas in the US and the world. To achieve this goal, the high temporal resolution precipitation data collected from a high spatial density, 72-year precipitation-gauge network on the 4.25-km 2 North Appalachian Experimental Watershed in east-central Ohio were used. Seasonal storm pattern trends from 1939 to 2010 were investigated using dimensionless depth (with the frequency of 50%, $d_{50}$) and the curve variability $(V = d_{80} - d_{20})$ at three dimensionless within-storms time periods (three verticals). The Spearman rank correlation procedure (correlation coefficient, ρ and significance probability, p) was used to statistically determine trends over time using 8 periods of 4-season sets of Huff curves over the 72 years. Two averages of ρ and p were computed: (1) by averaging the individual ρ and p obtained from 10 gauges (AvgI); and (2) by grouped averaging of all individual gauge values of $d_{50}$ and V and then computing ρ and p (AvgG). The test results of individual gauges showed that 4 cases for $d_{50}$ and 23 cases for V were significant for all seasons and verticals (total of 120 cases for each variable). The test results of AvgI for $d_{50}$ and V and AvgG for $d_{50}$ showed no significant trends in all seasons and verticals. Only the AvgG for V led to a significant trend for V in spring and fall at different times within storm patterns. The data do not provide sufficient evidence at the p=0.05 significance level to reject the null hypothesis of unchanging position of the dimensionless depth of the 50% Huff curves for individual or averages for all seasons and verticals. Also, there is insufficient evidence to reject the null hypothesis of unchanging variability, V, using the average ρ of individual gauges (AvgI for V) for all seasons and verticals. AvgG results showed a significant trend in V; however, this analysis may be affected by the nonindependence of storm data. The results suggest that it is likely that there is little if any effect of trending climate over approximately 70 years on Huff curve patterns. The results of this study add to the robustness characteristics of Huff curves and to their potential use in hydrological practice as design storms, as the foundation of stochastic storm generation, and for storm disaggregation, and they deserve further investigation. These different forms of inputs to watershed models have the potential to improve runoff estimation. The results suggest that, if verified in other studies, they have applicability to provide useful stationary precipitation patterns across the US and other areas of the world in areas of nonstationary climate. Also, individual rain gauge data may not be representative of trends even over small areas, and seasonal differences were noticeable as found in other studies. Recommendations are provided.

54 ENVIRONMENTAL SCIENCES↗

Reproducibility of protein x-ray diffuse scattering and potential utility for modeling atomic displacement parameters

Protein structure and dynamics can be probed using x-ray crystallography. Whereas the Bragg peaks are only sensitive to the average unit-cell electron density, the signal between the Bragg peaks—diffuse scattering—is sensitive to spatial correlations in electron-density variations. Although diffuse scattering contains valuable information about protein dynamics, the diffuse signal is more difficult to isolate from the background compared to the Bragg signal, and the reproducibility of diffuse signal is not yet well understood. We present a systematic study of the reproducibility of diffuse scattering from isocyanide hydratase in three different protein forms. Both replicate diffuse datasets and datasets obtained from different mutants were similar in pairwise comparisons (Pearson correlation coefficient ≥0.8). The data were processed in a manner inspired by previously published methods using custom software with modular design, enabling us to perform an analysis of various data processing choices to determine how to obtain the highest quality data as assessed using unbiased measures of symmetry and reproducibility. The diffuse data were then used to characterize atomic mobility using a liquid-like motions (LLM) model. This characterization was able to discriminate between distinct anisotropic atomic displacement parameter (ADP) models arising from different anisotropic scaling choices that agreed comparably with the Bragg data. Our results emphasize the importance of data reproducibility as a model-free measure of diffuse data quality, illustrate the ability of LLM analysis of diffuse scattering to select among alternative ADP models, and offer insights into the design of successful diffuse scattering experiments.

59 BASIC BIOLOGICAL SCIENCES↗

Novel angular velocity estimation technique for plasma filaments

Magnetic field aligned filaments such as blobs and edge localized mode filaments carry significant amounts of heat and particles to the plasma facing components and they decrease their lifetime. The dynamics of these filaments determine at least a part of the heat and particle loads. These dynamics can be characterized by their translation and rotation. In this paper, we present an analysis method novel for fusion plasmas, which can estimate the angular velocity of the filaments on frame-by-frame time resolution. After pre-processing, the frames are two-dimensional (2D) Fourier-transformed, then the resulting 2D Fourier magnitude spectra are transformed to log-polar coordinates, and finally the 2D cross-correlation coefficient function (CCCF) is calculated between the consecutive frames. The displacement of the CCCF’s peak along the angular coordinate estimates the angle of rotation of the most intense structure in the frame. Further, the proposed angular velocity estimation method is tested and validated for its accuracy and robustness by applying it to rotating Gaussian-structures. The method is also applied to gas-puff imaging measurements of filaments in National Spherical Torus Experiment plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Determining best practices for using genetic algorithms in molecular discovery

Genetic algorithms (GAs) are a powerful tool to search large chemical spaces for inverse molecular design. However, GAs have multiple hyperparameters that have not been thoroughly investigated for chemical space searches. In this tutorial, we examine the general effects of a number of hyperparameters, such as population size, elitism rate, selection method, mutation rate, and convergence criteria, on key GA performance metrics. Here, we show that using a self-termination method with a minimum Spearman’s rank correlation coefficient of 0.8 between generations maintained for 50 consecutive generations along with a population size of 32, a 50% elitism rate, three-way tournament selection, and a 40% mutation rate provides the best balance of finding the overall champion, maintaining good coverage of elite targets, and improving relative speedup for general use in molecular design GAs.

36 MATERIALS SCIENCE↗

Quantifying uncertainty in analysis of shockless dynamic compression experiments on platinum. I. Inverse Lagrangian analysis

Absolute measurements of solid-material compressibility by magnetically driven shockless dynamic compression experiments to multi-megabar pressures have the potential to greatly improve the accuracy and precision of pressure calibration standards for use in diamond anvil cell experiments. Here, to this end, we apply characteristics-based inverse Lagrangian analysis (ILA) to 11 sets of ramp-compression data on pure platinum (Pt) metal and then reduce the resulting weighted-mean stress–strain curve to the principal isentrope and room-temperature isotherm using simple models for yield stress and Grüneisen parameter. We introduce several improvements to methods for ILA and quasi-isentrope reduction, the latter including calculation of corrections in wave speed instead of stress and pressure to render results largely independent of initial yield stress while enforcing thermodynamic consistency near zero pressure. More importantly, we quantify in detail the propagation of experimental uncertainty through ILA and model uncertainty through quasi-isentrope reduction, considering all potential sources of error except the electrode and window material models used in ILA. Compared to previous approaches, we find larger uncertainty in longitudinal stress. Monte Carlo analysis demonstrates that uncertainty in the yield-stress model constitutes by far the largest contribution to uncertainty in quasi-isentrope reduction corrections. We present a new room-temperature isotherm for Pt up to 444 GPa, with 1-sigma uncertainty at that pressure of just under ±1.2%; the latter is about a factor of three smaller than uncertainty previously reported for multi-megabar ramp-compression experiments on Pt. The result is well represented by a Vinet-form compression curve with (isothermal) bulk modulus K 0 = 270.3 ± 3.8 GPa, pressure derivative K$^{'}_{0}$= 5.66 ± 0.10, and correlation coefficient $R_{K_{0},K^{'}_{0}}$= –0.843.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A method for examining ensemble averaging forms during the transition to turbulence in HED systems for application to RANS models

This paper discusses a strategy to initialize a two-dimensional (2D) Reynolds-averaged Navier–Stokes model [LANL's Besnard–Harlow–Rauenzahn (BHR) model] in order to describe an unsteady transitional Richtmyer–Meshkov (RM)-induced flow observed in on-going high-energy-density ensemble experiments performed on the OMEGA-EP facility. The experiments consist of a nominal single-mode perturbation (initial amplitude a 0 ≈ 10 and wavelength $λ$ = 100μm) with target-to-target variations in the surface roughness subjected to the RM instability with delayed Rayleigh–Taylor in a heavy-to-light configuration. Our strategy leverages high-resolution three-dimensional (3D) implicit large eddy simulations (ILES) simulations to initialize BHR-relevant parameters and subsequently validate the 2D BHR results against the 3D ILES simulations. A suite of five 3D ILES simulations corresponding to five experimental target profiles is undertaken to generate an ensemble dataset. Using ensemble averages from the 3D simulations to initialize the turbulent kinetic energy in the BHR model ( K 0 ) demonstrates the ability of the model to predict the time evolution of the interface as well as the density-specific-volume covariance, b . To quantify the sensitivity of the BHR results to the choice of K 0 and the initial turbulent length scale, S 0 , we execute a parameter sweep spanning four orders of magnitude for both S 0 and K 0 , generating a parameter space consisting of 26 simulations. The Pearson's correlation coefficient is used as a measure of discrepancy between the 2D BHR and 3D ILES simulations and reveals that the ranges 8≲S 0 ≲20 μm and 10 9 ≲K 0 ≲10 10 cm 2 /s 2 produce predictions that agree best with the 3D ILES results.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Computational prediction of the effect of amino acid changes on the binding affinity between SARS-CoV-2 spike RBD and human ACE2

The association of the receptor binding domain (RBD) of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) spike protein with human angiotensin-converting enzyme 2 (hACE2) represents the first required step for cellular entry. SARS-CoV-2 has continued to evolve with the emergence of several novel variants, and amino acid changes in the RBD have been implicated with increased fitness and potential for immune evasion. Reliably predicting the effect of amino acid changes on the ability of the RBD to interact more strongly with the hACE2 can help assess the implications for public health and the potential for spillover and adaptation into other animals. Here, we introduce a two-step framework that first relies on 48 independent 4-ns molecular dynamics (MD) trajectories of RBD-hACE2 variants to collect binding energy terms decomposed into Coulombic, covalent, van der Waals, lipophilic, generalized Born solvation, hydrogen bonding, π-π packing, and self-contact correction terms. The second step implements a neural network to classify and quantitatively predict binding affinity changes using the decomposed energy terms as descriptors. The computational base achieves a validation accuracy of 82.8% for classifying single–amino acid substitution variants of the RBD as worsening or improving binding affinity for hACE2 and a correlation coefficient of 0.73 between predicted and experimentally calculated changes in binding affinities. Both metrics are calculated using a fivefold cross-validation test. Our method thus sets up a framework for screening binding affinity changes caused by unknown single– and multiple–amino acid changes offering a valuable tool to predict host adaptation of SARS-CoV-2 variants toward tighter hACE2 binding.

60 APPLIED LIFE SCIENCES↗

Availability of Critical Benchmark Experiments for the Pebble Tanker Transportation Model for Nuclear Criticality Safety Validation of TRISO Pebbles

This study addresses the need for comprehensive investigations into TRi-structural ISOtropic (TRISO) fuel pebble transportation validation. In this work, an exploratory model, the pebble tanker(PT), was developed with the aim of facilitating the validation of nuclear criticality safety calculations in the context of industrial-scale transportation of TRISO fuel. The PT model was designed to investigate the availability and applicability of critical benchmark experiments crucial for assessing the transportation of these pebbles. This work incorporated sensitivity/uncertainty (S/U) similarity studies to quantify the applicability of critical benchmark experiments and to address nuclear data uncertainties in the context of TRISO transportation. Two container models were investigated: one for the Hermes-type pebble and one for the Pebble Bed Modular Reactor (PBMR)–type pebble. The models were simplified, considering fuel, containment, and either water or air, to enable a focus on the underlying physics of applications involving TRISO fuel pebbles using the PT model. A crucial aspect under consideration was the capacity of the transport package to hold pebbles while ensuring subcriticality in the flooded state. An approach in the criticality validation process involves assessing the similarity between systems through an integral index parameter evaluation. This involves calculating a correlation coefficient (referred to as c k ) based on shared nuclear data–induced uncertainty between a benchmark experiment and the application of the PT model. To facilitate this analysis, the SCALE tools, particularly the CSAS6-Shift, TSUNAMI-3D-Shift, and TSUNAMI-IP sequences, were employed for comprehensive studies in neutronics and S/U analysis. Our findings showed that there are sufficient critical experimental benchmarks to perform this validation of the PT model in the most reactive state, i.e. when the tanker is flooded. This paper provides valuable insights into validating a transport package for Generation IV TRISO fuel pebbles.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Assessment of Critical Experiment Benchmark Applicability to a Large-Capacity HALEU Transportation Package Concept

This work presents an assessment of the applicability of existing benchmark critical experiments to the criticality safety code validation for a large-capacity high-assay low-enriched uranium (HALEU) transportation package concept. Numerous next-generation nuclear reactor designs require HALEU fuel, which is characterized by an enrichment between 5 and 20 wt% 235 U. The U.S. Department of Energy (DOE) has proposed to recover and downblend highly enriched uranium from DOE-owned used nuclear fuel to accelerate the demonstration of commercially viable microreactor technologies. One element of the infrastructure needed to demonstrate HALEU-fueled reactors is the ability to safely transport enriched product to be used for fuel fabrication. There is uncertainty as to whether existing critical benchmark experiment data are sufficient to support criticality safety code validation for HALEU transportation applications. The anticipated chemical form of the HALEU in the proposed transportation concept is UO 2 with 20 wt% 235 U/U. The concept uses a combination of an existing transportation packaging design and a novel basket design, including borated aluminum flux traps. The basket provides space for 18 reusable, stainless steel canisters that contain the HALEU. In 10 CFR 71, normal conditions of transport (NCTs) and hypothetical accident conditions (HACs) are defined for fissile material transportation packages. NCT and HAC KENO-VI models of the transportation package were developed using the Standardized Computer Analyses for Licensing Evaluation (SCALE) 6.2.3 computer code package, and optimum moderation conditions were determined using the SCALE SAMPLER sequence. The SCALE Tools for Sensitivity and Uncertainty Analysis Methodology Implementation (TSUNAMI) sequences were then used to compare the neutronic characteristics of 1584 International Criticality Safety Benchmark Evaluation Project benchmark critical experiments with the NCT and HAC HALEU transportation models. The TSUNAMI integral correlation coefficient c k was the criterion used to rank neutronic similarity. Thirty-four experiments were identified as similar (c k ≥ 0.9) to the NCT model, and 55 experiments were identified as similar to the HAC model. Hundreds of experiments were also identified as at least marginally similar (c k ≥ 0.8) to both models. The results indicate that additional critical experiments are unlikely to be needed to support HALEU transportation criticality safety analyses for package concepts similar to the concept package analyzed.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

High mass and halo resolution from fast low resolution simulations

Generating mocks for future sky surveys requires large volumes and high resolutions, which is computationally expensive even for fast simulations. Here we try to develop numerical schemes to calibrate various halo and matter statistics in fast low resolution simulations compared to high resolution N-body and hydrodynamic simulations. For the halos, we improve the initial condition resolution and develop a halo finder "relaxed-FoF", where we allow different linking lengths for different halo mass and velocity dispersions. We show that our relaxed-FoF halo finder improves the common statistics, such as halo bias, halo mass function, halo auto power spectrum, cross correlation coefficient with the reference halo catalog, and halo-matter cross power spectrum. We also calibrate small-scale velocities of small halos to improve the power spectrum in redshift space. For the matter statistics, we incorporate the potential gradient descent (PGD) method into fast simulations to improve the matter distribution at nonlinear scales. By building a lightcone output, we show that the PGD method significantly improves the weak lensing convergence tomographic power spectrum. With these improvements FastPM is comparable to the high resolution full N-body simulation of the same mass resolution, with two orders of magnitude fewer time steps. These techniques can be used to improve the halo and matter statistics of FastPM simulations for mock catalogs of future surveys such as DESI and LSST.

79 ASTRONOMY AND ASTROPHYSICS↗

Multi-tracer intensity mapping: cross-correlations, line noise & decorrelation

Line intensity mapping (LIM) is a rapidly emerging technique for constraining cosmology and galaxy formation using multi-frequency, low angular resolution maps. Many LIM applications crucially rely on cross-correlations of two line intensity maps, or of intensity maps with galaxy surveys or galaxy/CMB lensing. We present a consistent halo model to predict all these cross-correlations and enable joint analyses, in 3D redshift-space and for 2D projected maps. We extend the conditional luminosity function formalism to the multi-line case, to consistently account for correlated scatter between multiple galaxy line luminosities. This allows us to model the scale-dependent decorrelation between two line intensity maps, a key input for foreground rejection and for approaches that estimate auto-spectra from cross-spectra. This also enables LIM cross-correlations to reveal astrophysical properties of the interstellar medium inacessible with LIM auto-spectra. We expose the different sources of luminosity scatter or "line noise" in LIM, and clarify their effects on the 1-halo and galaxy shot noise terms. In particular, we show that the effective number density of halos can in some cases exceed that of galaxies, counterintuitively. Using observational and simulation input, we implement this halo model for the Hα, [Oiii], Lyman-α, CO and [Cii] lines. Here, we encourage observers and simulators to measure galaxy luminosity correlation coefficients for pairs of lines whenever possible. Our code is publicly available at https://github.com/EmmanuelSchaan/HaloGen/tree/LIM. In a companion paper, we use this halo model formalism and code to highlight the degeneracies between cosmology and astrophysics in LIM, and to compare the LIM observables to galaxy detection for a number of surveys.

79 ASTRONOMY AND ASTROPHYSICS↗

Full calibration of the tomographic redshift distribution from the HSC PDR3 Shape Catalog with DESI

The calibration of tomographic redshift distributionsis essential for cosmological analysis of weak lensing data.In this work, we calibrate all four tomographic bins of the Hyper Suprime Camera (HSC) weak lensing catalog with the Dark Energy Spectroscopic Instrument (DESI) Data Release 1 and 2 using the clustering redshifts technique. We include z > 1.2 redshift sources such as emission line galaxies (ELG) and quasars (QSO) sources in our calibration, which were not available in the previous HSC calibration (Rau et al. (2022), Mon. Not. Roy. Astron. Soc. 524 (2023) 5109), allowing a complete calibration of all the redshift bins. We find the first tomographic bin exhibits a small shift towards low redshifts. The second bin is in good agreement with the photometric calibration, while third and fourth bin exhibit a shift towards higher redshifts. However, these shifts are considerably smaller than the shifts obtained in the HSC Year 3 cosmic shear analyses. We evaluate the impact of galaxy bias and magnification effects from all the samples on the measurements, finding them to be small, and we propose corrections to reduce them further. Specifically, we relax the assumption of linear bias and only assume no redshift evolution of the cross-correlation coefficient, allowing us to leverage smaller clustering scales. We model the redshift distributions with splines and compare our results to previous analyses as well as to other parameterizations found in literature. For the two high-redshift tomographic bins, we find the shifts to higher redshifts with respect to the measurements performed in Rau+2022 to be Δz$_{3}$ =-0.039$^{+0.020}$$_{-0.021}$ and Δz$_{4}$ = -0.048$^{+0.012}$$_{-0.012}$.

Choppin de Janvry, J. [LBL, Berkeley; UC, Berkeley↗

Energetic particle-induced geodesic acoustic modes on DIII-D

Various properties of the energetic particle-induced geodesic acoustic mode (EGAM) are explored in this large database analysis of DIII-D experimental data. EGAMs are n = 0 modes with m = 0 electrostatic potential fluctuations (where n/m = toroidal/poloidal mode number), m = 1 density fluctuations, and m = 2 magnetic fluctuations. The fundamental frequency (~20–40 kHz) of the mode is typically below that of the traditional geodesic acoustic mode frequency. EGAMs are most easily destabilized by beams in the counter plasma current (counter-I p ) direction as compared to co-Ip and off-axis beams. During counter beam injection, the mode frequency is found to have the strongest linear dependence (correlation coefficient r = –0.71) with the safety factor (q). Here, the stability of the mode in the space of q and poloidal beta (β p ) shows a clear boundary for the mode stability. The stability of the mode depends more strongly on damping rate than on fast-ion drive for a given injection geometry.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Performance assessment of PHITS simulations for the inverse-kinematic p( 7 Li,n) 7 Be reaction based on fast-neutron measurements with a diamond detector

The inverse-kinematics p( 7 Li,n) 7 Be reaction produces forward-focused neutron emission, offering enhanced usable flux and reduced shielding requirements. Reliable simulation of such neutron fields is essential for the development of compact accelerator-based neutron sources. In this study, a PHITS-based simulation framework for the reaction was experimentally assessed using fast-neutron measurements. Forward-directed neutrons were measured with a diamond neutron detector and quantitatively compared with simulations with newly prepared IK-Frag cross-section file based on the proton-induced reaction data in ENDF/B-VIII.1, TENDL-2023, and JENDL-4.0/HE. Measurements and simulations were performed for incident 7 Li 3+ energies ranging from 15.0 to 25.0 MeV using a 50 μm-thick polypropylene target. For all conditions, the PHITS-based simulation framework reproduced the deposited energy spectra at the correct order of magnitude. The comparison of deposited energy spectra in the diamond detector showed high correlation coefficients across all investigated energies, indicating reasonable agreement in spectral shape between simulations and measurements. This work represents an initial step toward establishing a benchmark for PHITS simulations of the inverse kinematic reaction between an incident lithium-ion and a proton target.

43 PARTICLE ACCELERATORS↗

Circumventing data imbalance in magnetic ground state data for magnetic moment predictions

Abstract Magnetic materials play a crucial role in the transition to more sustainable forms of energy and electric vehicles. There is an anticipated shortage in magnetic materials in the future, and as a result there is an urgent need to discover and design new magnetic materials. Computational magnetic material design using density functional theory is daunting because of the challenge in identifying magnetic ground states from a combinatorially large set of possibilities. Machine learning offers a path forward by enabling efficient surrogate models that can more readily enumerate these states, but there is a dearth of training data available, and what is available tends to be imbalanced with too much non-magnetic data. In this work we show that the discrete and previously tackled data imbalance that exists at the level of the magnetic ordering leads to an imbalanced continuous distribution with many zeros when the data is unraveled at the atomic magnetic moment level, which subsequently leads to models with low accuracy for magnetic properties. We mitigate this by using a two-part model framework. Our scheme is able to classify atoms into magnetic and non-magnetic with an F1 score and Matthew’s correlation coefficient (MCC) of ~91% and then to provide an implicit embedding representation that maps directly onto the magnitude of the magnetic moment with a mean absolute error of 0.1 μ B . Beyond screening for new magnetic materials, we demonstrate an additional practical use case of our scheme: the provision of good initial guesses for magnetic moments in first-principles electronic relaxations. Such initialization is shown to lead to faster convergence to configurations that lie closer to the ground state.

Computer Science↗

PRODeepSyn: predicting anticancer synergistic drug combinations by embedding cell lines with protein–protein interaction network

Abstract Although drug combinations in cancer treatment appear to be a promising therapeutic strategy with respect to monotherapy, it is arduous to discover new synergistic drug combinations due to the combinatorial explosion. Deep learning technology holds immense promise for better prediction of in vitro synergistic drug combinations for certain cell lines. In methods applying such technology, omics data are widely adopted to construct cell line features. However, biological network data are rarely considered yet, which is worthy of in-depth study. In this study, we propose a novel deep learning method, termed PRODeepSyn, for predicting anticancer synergistic drug combinations. By leveraging the Graph Convolutional Network, PRODeepSyn integrates the protein–protein interaction (PPI) network with omics data to construct low-dimensional dense embeddings for cell lines. PRODeepSyn then builds a deep neural network with the Batch Normalization mechanism to predict synergy scores using the cell line embeddings and drug features. PRODeepSyn achieves the lowest root mean square error of 15.08 and the highest Pearson correlation coefficient of 0.75, outperforming two deep learning methods and four machine learning methods. On the classification task, PRODeepSyn achieves an area under the receiver operator characteristics curve of 0.90, an area under the precision–recall curve of 0.63 and a Cohen’s Kappa of 0.53. In the ablation study, we find that using the multi-omics data and the integrated PPI network’s information both can improve the prediction results. Additionally, the case study demonstrates the consistency between PRODeepSyn and previous studies.

Wang, Xiaowen↗

The impact of curation errors in the PDBBind Database on machine learning predictions of protein–protein binding affinity

The PDBBind database has been widely utilized for the computational prediction of protein–protein binding affinities. While the accuracy of the PDBBind-curated equilibrium dissociation constants (K D ) has been reported for the protein–ligand subset of the PDBBind database, the curation accuracy has not been reported for the protein–protein subset. Here, we present a detailed manual analysis for the subset of PDBBind records with PubMed Central Open Access primary publications and find that ~19% of these records had K D values that were not supported by their primary publications. The impact of these putative curation errors on the machine learning-based prediction of K D from experimental protein–protein 3D structures was evaluated and correcting the curation errors improved the Pearson correlation coefficient between measured and random forest-predicted log 10 (K D ) values by ~8 percentage points. This finding underscores the importance of dataset accuracy for computational modelling and highlights the need for more stringent curation processes when extracting information from the scientific literature.

59 BASIC BIOLOGICAL SCIENCES↗