Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Correlation coefficient”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Day-to-day reliability of basal heart rate and short-term and ultra short-term heart rate variability assessment by the Equivital eq02+ LifeMonitor in US Army soldiers

Introduction The present study determined the (1) day-to-day reliability of basal heart rate (HR) and HR variability (HRV) measured by the Equivital eq02+ LifeMonitor and (2) agreement of ultra short-term HRV compared with short-term HRV. Methods Twenty-three active-duty US Army Soldiers (5 females, 18 males) completed two experimental visits separated by >48 hours with restrictions consistent with basal monitoring (eg, exercise, dietary), with measurements after supine rest at minutes 20–21 (ultra short-term) and minutes 20–25 (short-term). HRV was assessed as the SD of R–R intervals (SDNN) and the square root of the mean squared differences between consecutive R–R intervals (RMSSD). Results The day-to-day reliability (intraclass correlation coefficient (ICC)) using linear-mixed model approach was good for HR (0.849, 95% CI: 0.689 to 0.933) and RMSSD (ICC: 0.823, 95% CI: 0.623 to 0.920). SDNN had moderate day-to-day reliability with greater variation (ICC: 0.689, 95% CI: 0.428 to 0.858). The reliability of RMSSD was slightly improved when considering the effect of respiration (ICC: 0.821, 95% CI: 0.672 to 0.944). There was no bias for HR measured for 1 min versus 5 min (p=0.511). For 1 min measurements versus 5 min, there was a very modest mean bias of −4 ms for SDNN and −1 ms for RMSSD (p≤0.023). Conclusion When preceded by a 20 min stabilisation period using restrictions consistent with basal monitoring and measuring respiration, military personnel can rely on the eq02+ for basal HR and RMSSD monitoring but should be more cautious using SDNN. These data also support using ultra short-term measurements when following these procedures.

General & Internal Medicine↗

Predicting September Arctic Sea Ice: A Multimodel Seasonal Skill Comparison

This study quantifies the state of the art in the rapidly growing field of seasonal Arctic sea ice prediction. A novel multimodel dataset of retrospective seasonal predictions of September Arctic sea ice is created and analyzed, consisting of community contributions from 17 statistical models and 17 dynamical models. Prediction skill is compared over the period 2001–20 for predictions of pan-Arctic sea ice extent (SIE), regional SIE, and local sea ice concentration (SIC) initialized on 1 June, 1 July, 1 August, and 1 September. This diverse set of statistical and dynamical models can individually predict linearly detrended pan-Arctic SIE anomalies with skill, and a multimodel median prediction has correlation coefficients of 0.79, 0.86, 0.92, and 0.99 at these respective initialization times. Regional SIE predictions have similar skill to pan-Arctic predictions in the Alaskan and Siberian regions, whereas regional skill is lower in the Canadian, Atlantic, and central Arctic sectors. The skill of dynamical and statistical models is generally comparable for pan-Arctic SIE, whereas dynamical models outperform their statistical counterparts for regional and local predictions. The prediction systems are found to provide the most value added relative to basic reference forecasts in the extreme SIE years of 1996, 2007, and 2012. SIE prediction errors do not show clear trends over time, suggesting that there has been minimal change in inherent sea ice predictability over the satellite era. Overall, this study demonstrates that there are bright prospects for skillful operational predictions of September sea ice at least 3 months in advance.

54 ENVIRONMENTAL SCIENCES↗

Hydrometeor Shape Variability in Snowfall as Retrieved from Polarimetric Radar Measurements

A polarimetric radar–based method for retrieving atmospheric ice particle shapes is applied to snowfall measurements by a scanning Ka-band radar deployed at Oliktok Point, Alaska (70.495°N, 149.883°W). The mean aspect ratio, which is defined by the hydrometeor minor-to-major dimension ratio for a spheroidal particle model, is retrieved as a particle shape parameter. The radar variables used for aspect ratio profile retrievals include reflectivity, differential reflectivity, and the copolar correlation coefficient. The retrievals indicate that hydrometeors with mean aspect ratios below 0.2–0.3 are usually present in regions with air temperatures warmer than approximately from –17° to –15°C, corresponding to a regime that has been shown to be favorable for growth of pristine ice crystals of planar habits. Radar reflectivities corresponding to the lowest mean aspect ratios are generally between –10 and 10 dBZ. For colder temperatures, mean aspect ratios are typically in a range between 0.3 and 0.8. There is a tendency for hydrometeor aspect ratios to increase as particles transition from altitudes in the temperature range from –17° to –15°C toward the ground. Furthermore, this increase is believed to result from aggregation and riming processes that cause particles to become more spherical and is associated with areas demonstrating differential reflectivity decreases with increasing reflectivity. Aspect ratio retrievals at the lowest altitudes are consistent with in situ measurements obtained using a surface-based multiangle snowflake camera. Pronounced gradients in particle aspect ratio profiles are observed at altitudes at which there is a change in the dominant hydrometeor species, as inferred by spectral measurements from a vertically pointing Doppler radar.

54 ENVIRONMENTAL SCIENCES↗

Projecting Future Energy Production from Operating Wind Farms in North America. Part I: Dynamical Downscaling

Abstract New simulations at 12-km grid spacing with the Weather and Research Forecasting (WRF) Model nested in the MPI Earth System Model (ESM) are used to quantify possible changes in wind power generation potential as a result of global warming. Annual capacity factors (CF; measures of electrical power production) computed by applying a power curve to hourly wind speeds at wind turbine hub height from this simulation are also used to illustrate the pitfalls in seeking to infer changes in wind power generation directly from low-spatial-resolution and time-averaged ESM output. WRF-derived CF are evaluated using observed daily CF from operating wind farms. The spatial correlation coefficient between modeled and observed mean CF is 0.65, and the root-mean-square error is 5.4 percentage points. Output from the MPI-WRF Model chain also captures some of the seasonal variability and the probability distribution of daily CF at operating wind farms. Projections of mean annual CF (CF A ) indicate no change to 2050 in the southern Great Plains and Northeast. Interannual variability of CF A increases in the Midwest, and CF A declines by up to 2 percentage points in the northern Great Plains. The probability of wind droughts (extended periods with anomalously low production) and wind bonus periods (high production) remains unchanged over most of the eastern United States. The probability of wind bonus periods exhibits some evidence of higher values over the Midwest in the 2040s, whereas the converse is true over the northern Great Plains. Significance Statement Wind energy is playing an increasingly important role in low-carbon-emission electricity generation. It is a “weather dependent” renewable energy source, and thus changes in the global atmosphere may cause changes in regional wind power production (PP) potential. We use PP data from operating wind farms to demonstrate that regional simulations exhibit skill in capturing actual power production. Projections to the middle of this century indicate that over most of North America east of the Rocky Mountains annual expected PP is largely unchanged, as is the probability of extended periods of anomalously high or low production. Any small declines in annual PP are of much smaller magnitude than changes due to technological innovation over the last two decades.

Meteorology & Atmospheric Sciences↗

A Process Model for ITCZ Narrowing under Warming Highlights Clear-Sky Water Vapor Feedbacks and Gross Moist Stability Changes in AMIP Models

Tropical areas with mean upward motion—and as such the zonal-mean intertropical convergence zone (ITCZ)—are projected to contract under global warming. To understand this process, a simple model based on dry static energy and moisture equations is introduced for zonally symmetric overturning driven by sea surface temperature (SST). Processes governing ascent area fraction and zonal mean precipitation are examined for insight into Atmospheric Model Intercomparison Project (AMIP) simulations. Bulk parameters governing radiative feedbacks and moist static energy transport in the simple model are estimated from the AMIP ensemble. Uniform warming in the simple model produces ascent area contraction and precipitation intensification—similar to observations and climate models. Contributing effects include stronger water vapor radiative feedbacks, weaker cloud-radiative feedbacks, stronger convection-circulation feedbacks, and greater poleward moisture export. The simple model identifies parameters consequential for the inter-AMIP-model spread; an ensemble generated by perturbing parameters governing shortwave water vapor feedbacks and gross moist stability changes under warming tracks inter-AMIP-model variations with a correlation coefficient ~0.46. Here, the simple model also predicts the multimodel mean changes in tropical ascent area and precipitation with reasonable accuracy. Furthermore, the simple model reproduces relationships among ascent area precipitation, ascent strength, and ascent area fraction observed in AMIP models. A substantial portion of the inter-AMIP-model spread is traced to the spread in how moist static energy and vertical velocity profiles change under warming, which in turn impact the gross moist stability in deep convective regions—highlighting the need for observational constraints on these quantities.

54 ENVIRONMENTAL SCIENCES↗

PERSIANN Dynamic Infrared–Rain Rate (PDIR-Now): A Near-Real-Time, Quasi-Global Satellite Precipitation Dataset

This study presents the Precipitation Estimation from Remotely Sensed Information Using Artificial Neural Networks–Dynamic Infrared Rain Rate (PDIR-Now) near-real-time precipitation dataset. This dataset provides hourly, quasi-global, infrared-based precipitation estimates at 0.04° × 0.04° spatial resolution with a short latency (15–60 min). It is intended to supersede the PERSIANN–Cloud Classification System (PERSIANN-CCS) dataset previously produced as the near-real-time product of the PERSIANN family. We first provide a brief description of the algorithm’s fundamentals and the input data used for deriving precipitation estimates. Second, we provide an extensive evaluation of the PDIR-Now dataset over annual, monthly, daily, and subdaily scales. Last, the article presents information on the dissemination of the dataset through the Center for Hydrometeorology and Remote Sensing (CHRS) web-based interfaces. The evaluation, conducted over the period 2017–18, demonstrates the utility of PDIR-Now and its improvement over PERSIANN-CCS at all temporal scales. Specifically, PDIR-Now improves the estimation of rain/no-rain days as demonstrated by a critical success index (CSI) of 0.53 compared to 0.47 of PERSIANN-CCS. In addition, PDIR-Now improves the estimation of seasonal and diurnal cycles of precipitation as well as regional precipitation patterns erroneously estimated by PERSIANN-CCS. Finally, an evaluation is carried out to examine the performance of PDIR-Now in capturing two extreme events, Hurricane Harvey and a cluster of summer thunderstorms that occurred over the Netherlands, where it is shown that PDIR-Now adequately represents spatial precipitation patterns as well as subdaily precipitation rates with a correlation coefficient (CORR) of 0.64 for Hurricane Harvey and 0.76 for the Netherlands thunderstorms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evaluation of coupled wind-wave model simulations of offshore winds in the Mid-Atlantic Bight using lidar-equipped buoys.

From 2014 to 2017, two Department of Energy buoys equipped with Doppler lidar were deployed off the U.S. East Coast to provide long term measurements of hub-height wind speed in the marine environment. In this study, we performed simulations of selected cases from the deployment using a 5-km configuration of the Weather Research and Forecasting (WRF) model, to see if simulated hub height speeds could produce closer agreement with the observations than existing reanalysis products. For each case we performed two additional simulations: one in which marine surface roughness height was one-way coupled to forecast wave parameters from a standalone WaveWatch III (WW3) simulation, and another in which WRF and WW3 were two-way coupled using the Coupled-Ocean-Atmosphere-Wave-Sediment-Transport (COAWST) framework. It was found that all the 5-km WRF simulations improved 90-m wind speed statistics for the tropical cyclone case of 08 May 2015 and the cold frontal case of 25 Mar 2016, but not the nor-easter of 18 Jan 2016. The impact of wave coupling on buoy-level (4 m) wind speed was modest and case dependent, but when present, the impact was typically seen at 90 m as well, being as large as 10% in stable conditions. One-way wave coupling consistently reduced wind speeds, improving biases for 25 Mar 2016 but worsening them for 08 May 2015. Two-way wave coupling mitigated these negative biases, improved wave field representation and statistics, and mostly improved 4-m wind field correlation coefficients, at least at the VA buoy, largely due to greater self-consistency between wind and wave fields.

54 ENVIRONMENTAL SCIENCES↗

Development and Evaluation of Ensemble Learning-based Environmental Methane Detection and Intensity Prediction Models

The environmental impacts of global warming driven by methane (CH 4 ) emissions have catalyzed significant research initiatives in developing novel technologies that enable proactive and rapid detection of CH 4 . Several data-driven machine learning (ML) models were tested to determine how well they identified fugitive CH 4 and its related intensity in the affected areas. Various meteorological characteristics, including wind speed, temperature, pressure, relative humidity, water vapor, and heat flux, were included in the simulation. We used the ensemble learning method to determine the best-performing weighted ensemble ML models built upon several weaker lower-layer ML models to (i) detect the presence of CH 4 as a classification problem and (ii) predict the intensity of CH 4 as a regression problem. The classification model performance for CH 4 detection was evaluated using accuracy, F1 score, Matthew’s Correlation Coefficient (MCC), and the area under the receiver operating characteristic curve (AUC ROC), with the top-performing model being 97.2%, 0.972, 0.945 and 0.995, respectively. The R 2 score was used to evaluate the regression model performance for CH 4 intensity prediction, with the R 2 score of the best-performing model being 0.858. The ML models developed in this study for fugitive CH 4 detection and intensity prediction can be used with fixed environmental sensors deployed on the ground or with sensors mounted on unmanned aerial vehicles (UAVs) for mobile detection.

Majumder, Reek↗

Modeling the Metabolic Costs of Heavy Military Backpacking

Existing predictive equations underestimate the metabolic costs of heavy military load carriage. Metabolic costs are specific to each type of military equipment, and backpack loads often impose the most sustained burden on the dismounted warfighter. This study aimed to develop and validate an equation for estimating metabolic rates during heavy backpacking for the US Army Load Carriage Decision Aid (LCDA), an integrated software mission planning tool. Thirty healthy, active military-age adults (3 women, 27 men; age, 25 ± 7 yr; height, 1.74 ± 0.07 m; body mass, 77 ± 15 kg) walked for 6–21 min while carrying backpacks loaded up to 66% body mass at speeds between 0.45 and 1.97 m·s -1 . A new predictive model, the LCDA backpacking equation, was developed on metabolic rate data calculated from indirect calorimetry. Model estimation performance was evaluated internally by k-fold cross-validation and externally against seven historical reference data sets. We tested if the 90% confidence interval of the mean paired difference was within equivalence limits equal to 10% of the measured metabolic rate. Estimation accuracy and level of agreement were also evaluated by the bias and concordance correlation coefficient (CCC), respectively. Estimates from the LCDA backpacking equation were statistically equivalent ( P < 0.01) to metabolic rates measured in the current study (bias, -0.01 ± 0.62 W·kg -1 ; CCC, 0.965) and from the seven independent data sets (bias, -0.08 ± 0.59 W·kg -1 ; CCC, 0.926). The newly derived LCDA backpacking equation provides close estimates of steady-state metabolic energy expenditure during heavy load carriage. These advances enable further optimization of thermal-work strain monitoring, sports nutrition, and hydration strategies.

59 BASIC BIOLOGICAL SCIENCES↗

Metabolic Costs of Walking with Weighted Vests

ABSTRACT Introduction The US Army Load Carriage Decision Aid (LCDA) metabolic model is used by militaries across the globe and is intended to predict physiological responses, specifically metabolic costs, in a wide range of dismounted warfighter operations. However, the LCDA has yet to be adapted for vest-borne load carriage, which is commonplace in tactical populations, and differs in energetic costs to backpacking and other forms of load carriage. Purpose The purpose of this study is to develop and validate a metabolic model term that accurately estimates the effect of weighted vest loads on standing and walking metabolic rate for military mission-planning and general applications. Methods Twenty healthy, physically active military-age adults (4 women, 16 men; age, 26 ± 8 yr old; height, 1.74 ± 0.09 m; body mass, 81 ± 16 kg) walked for 6 to 21 min with four levels of weighted vest loading (0 to 66% body mass) at up to 11 treadmill speeds (0.45 to 1.97 m·s −1 ). Using indirect calorimetry measurements, we derived a new model term for estimating metabolic rate when carrying vest-borne loads. Model estimates were evaluated internally byk-fold cross-validation and externally against 12 reference datasets (264 total participants). We tested if the 90% confidence interval of the mean paired difference was within equivalence limits equal to 10% of the measured walking metabolic rate. Estimation accuracy, precision, and level of agreement were also evaluated by the bias, standard deviation of paired differences, and concordance correlation coefficient (CCC), respectively. Results Metabolic rate estimates using the new weighted vest term were statistically equivalent (P< 0.01) to measured values in the current study (bias, −0.01 ± 0.54 W·kg −1 ; CCC, 0.973) as well as from the 12 reference datasets (bias, −0.16 ± 0.59 W·kg −1 ; CCC, 0.963). Conclusions The updated LCDA metabolic model calculates accurate predictions of metabolic rate when carrying heavy backpack and vest-borne loads. Tactical populations and recreational athletes that train with weighted vests can confidently use the simplified LCDA metabolic calculator provided as Supplemental Digital Content to estimate metabolic rates for work/rest guidance, training periodization, and nutritional interventions.

Sport Sciences↗

A Generalizable Evaluated Approach, Applying Advanced Geospatial Statistical Methods, to Identify High Lead Exposure Locations at Census Tract Scale: Michigan Case Study

BACKGROUND: Despite great progress in reducing environmental lead (Pb) levels, many children in the United States are still being exposed. OBJECTIVE: Our aim was to develop a generalizable approach for systematically identifying, verifying, and analyzing locations with high prevalence of children’s elevated blood Pb levels (EBLLs) and to assess available Pb models/indices as surrogates, using a Michigan case study. METHODS: We obtained ~1:9 million BLL test results of children <6 years of age in Michigan from 2006–2016; we then evaluated them for data representativeness by comparing two percentage EBLL (%EBLL) rates (number of children tested with EBLL divided by both number of children tested and total population). We analyzed %EBLLs across census tracts over three time periods and between two EBLL reference values (≥5 vs. ≥10 μg/dL) to evaluate consistency. Locations with high %EBLLs were identified by a top 20 percentile method and a Getis-Ord Gi* geospatial cluster “hotspot” analysis. For the locations identified, we analyzed convergences with three available Pb exposure models/indices based on old housing and sociodemographics. RESULTS: Analyses of 2014–2016 %EBLL data identified 11 Michigan locations via cluster analysis and 80 additional locations via the top 20 percentile method and their associated census tracts. Data representativeness and consistency were supported by a 0.93 correlation coefficient between the two EBLL rates over 11 y, and a Kappa score of ~0:8 of %EBLL hotspots across the time periods (2014–2016) and reference values. Many EBLL hotspot locations converge with current Pb exposure models/indices; others diverge, suggesting additional Pb sources for targeted interventions. DISCUSSION: This analysis confirmed known Pb hotspot locations and revealed new ones at a finer geographic resolution than previously available, using advanced geospatial statistical methods and mapping/visualization. It also assessed the utility of surrogates in the absence of blood Pb data. This approach could be applied to other states to inform Pb mitigation and prevention efforts. https://doi.org/10.1289/EHP9705

54 ENVIRONMENTAL SCIENCES↗

Performing k eff Validation of As-Loaded Criticality Safety Calculations Using UNF-ST&DARDS: Applicable Experiment Selection

The general method for performing validation of as loaded criticality safety calculations using UNF ST&DARDS is presented in a paper by Clarity, which includes a description of the UNF-ST&DARDS system. Proof-of-principle analyses were performed in the summer of 2019 for MPC-32 dual purpose canisters (DPCs) containing pressurized water reactor (PWR) fuel assemblies. Summaries of these results are presented in this and a companion paper for this conference. The current paper describes the TSUNAMI-IP calculations performed to select applicable experiments for validation of 11 MPC-32 DPCs. The companion paper discusses the TSUNAMI-3D calculations used to generate sensitivity data to support the experiment selections discussed here. Experiment selection is based on the sensitivity/uncertainty (S/U) methods used to validate criticality safety calculations of as-loaded DPCs containing pressurized water reactor (PWR) spent nuclear fuel (SNF). This process has been demonstrated and is summarized in this paper. The approach is similar to that used in NUREG/CR-7109, which provides an approach for validation of PWR burnup credit (BUC), including major and minor actinides and major fission products. The premise of S/U-based validation is that applicable experiments—those having a similar bias to a given application system—will have similar sensitivities for each isotope and reaction in the two systems. It is assumed that cross sections with larger uncertainties are more likely to contain data errors which contribute to the bias. The integral index c k thus propagates the system sensitivities with the nuclear covariance data to calculate a correlation coefficient representing the similarity of the two systems. In this work, a c k value of 0.8 or higher is interpreted as identifying an experiment with sufficient similarity for use in validation. This paper presents a brief summary of the characteristics of the 11 MPC-32 DPCs used in the proof of-principle analysis for as-loaded criticality safety calculation validation and an overview of the critical experiment suite with which each of these DPC models was compared. A summary and discussion of c k results is also presented, followed by conclusions and a discussion of future work to be performed for validation of UNF ST&DARDS as-loaded criticality safety calculations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Hierarchical Data Format for Nuclear Data Sensitivities

The SCALE code system includes capabilities for sensitivity and uncertainty (S/U) analysis as part of its TSUNAMI code suite. The sensitivity of a quantity of interest (for example, an application’s $k_{eff}$) to nuclear data is stored as a profile in a text-based file, which is known as a sensitivity data file (SDF). The sensitivity profile can be used to calculate uncertainties, correlation coefficients, and similarity indices. One of the goals of the present work was to seek general performance improvements in the TSUNAMI code suite, starting with the TSUNAMI-IP code for calculating similarity indices. Through profiling, it was found that reading the text-based sensitivity files was a performance bottleneck in the TSUNAMI-IP code. In a typical TSUNAMI-IP calculation, an application might be compared to thousands of benchmarks, thus requiring the reading of thousands of SDFs. Reading of binary-based data is generally faster than reading text-based data. Hierarchical Data Format 5 (HDF5) is a binary-based format that also benefits from being portable, and it can be inspected with nonproprietary tools. This paper describes an HDF5-based file format that has been introduced for SDFs.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Impact of Thermal Scattering Law on Similarity Assessment in Light-Water or Polyethylene-Moderated Systems

A collaborative effort between Pacific Northwest National Laboratory (PNNL) and Oak Ridge National Laboratory (ORNL) is underway to provide a technical basis and methodology for the criticality safety community to use the sum of fractions (SoF) method for generating limits for mixtures of “selected actinide nuclides” included in the ANSI/ANS 8.15 standard. The PNNL scope in this project is to define a range of mixtures of 233 U, 235 U, and 239 Pu moderated with either light water or polyethylene and to examine the critical masses for these mixtures. The ORNL scope is primarily to provide validation support for these studies. More complete discussion of the project and its validation aspects will be presented at the upcoming International Conference on Nuclear Criticality Safety (ICNC) this Fall in Sendai, Japan. Clear differences in benchmark similarity to application systems as assessed by the integral parameter ck are noted in the validation studies performed as part of this project as a function of moderator. The c k value is a correlation coefficient that represents that amount of shared uncertainty in k eff due to cross sections between two systems. Individual nuclide-reaction contributions between the two systems can be simply summed to arrive at the total c k value. Specifically, the c k values for light-water–moderated solution experiments are higher for a water-moderated application than for a polyethylene-moderated application. This result is neither totally unexpected nor surprising, but the magnitude of the difference was difficult to anticipate. The TSUNAMI sequence, in the SCALE 6.2.4 code package developed by ORNL, was used to generate eigenvalues and reactivity effects with perturbation-theory based approach through sensitivity coefficients for all nuclides in the system with all reactions and energy groups. The TSUNAMI-Indices and Parameters (IP) sequence then uses the sensitivity data generated through TSUNAMI to generate relational parameters (i.e., c k ) to determine the degree of similarity between systems. One detail of the SCALE material and data implementation must be discussed at this point. Several thermal scattering laws (TSLs) are available for 1 H. SCALE uses a different nuclide ID number for each TSL; essentially, each version of 1 H is treated as a unique nuclide. For example, 1 H bound in water ( 1 H-H2O) is assigned the nuclide ID 1001, whereas 1 H bound in polyethylene (h-poly) is assigned the nuclide ID 9001001. The same cross section data are used for all reactions in 1 H, regardless of TSL, except for scattering below the TSL cutoff energy. TSUNAMI-IP treats different nuclide IDs as different nuclides; thus, no uncertainty is shared between 1 H-H2O and h-poly, despite much of the same data, including covariance data, being used for both nuclides. This presents a question: how much of the difference in assessed similarity between water- and polyethylene-moderated systems is due to the moderators, and how much is caused by the treatment of 1 H-H2O and h poly with cross section and covariance data. The extended edits generated by TSUANMI-IP allow for an investigation of this issue specifically, as well as a demonstration of the general techniques available within TSUNAMI to understand the results of the similarity assessment. This paper presents and analyzes the similarity assessment of both water- and polyethylene-moderated systems for a single benchmark: PST-002-001.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Assessment of software methods for estimating protein-protein relative binding affinities

A growing number of computational tools have been developed to accurately and rapidly predict the impact of amino acid mutations on protein-protein relative binding affinities. Such tools have many applications, for example, designing new drugs and studying evolutionary mechanisms. In the search for accuracy, many of these methods employ expensive yet rigorous molecular dynamics simulations. By contrast, non-rigorous methods use less exhaustive statistical mechanics, allowing for more efficient calculations. However, it is unclear if such methods retain enough accuracy to replace rigorous methods in binding affinity calculations. This trade-off between accuracy and computational expense makes it difficult to determine the best method for a particular system or study. Here, eight non-rigorous computational methods were assessed using eight antibody-antigen and eight non-antibody-antigen complexes for their ability to accurately predict relative binding affinities (ΔΔG) for 654 single mutations. In addition to assessing accuracy, we analyzed the CPU cost and performance for each method using a variety of physico-chemical structural features. This allowed us to posit scenarios in which each method may be best utilized. Most methods performed worse when applied to antibody-antigen complexes compared to non-antibody-antigen complexes. Rosetta-based JayZ and EasyE methods classified mutations as destabilizing (ΔΔG < -0.5 kcal/mol) with high (83–98%) accuracy and a relatively low computational cost for non-antibody-antigen complexes. Some of the most accurate results for antibody-antigen systems came from combining molecular dynamics with FoldX with a correlation coefficient (r) of 0.46, but this was also the most computationally expensive method. Overall, our results suggest these methods can be used to quickly and accurately predict stabilizing versus destabilizing mutations but are less accurate at predicting actual binding affinities. This study highlights the need for continued development of reliable, accessible, and reproducible methods for predicting binding affinities in antibody-antigen proteins and provides a recipe for using current methods.

59 BASIC BIOLOGICAL SCIENCES↗

Coarse Woody Debris Decomposition Assessment Tool: Model validation and application

Coarse woody debris (CWD) is a significant component of the forest biomass pool; hence a model is warranted to predict CWD decomposition and its role in forest carbon (C) and nutrient cycling under varying management and climatic conditions. A process-based model, CWDDAT (Coarse Woody Debris Decomposition Assessment Tool) was calibrated and validated using data from the FACE (Free Air Carbon Dioxide Enrichment) Wood Decomposition Experiment utilizing pine ( Pinus taeda ), aspen ( Populous tremuloides ) and birch ( Betula papyrifera ) on nine Experimental Forests (EF) covering a range of climate, hydrology, and soil conditions across the continental USA. The model predictions were evaluated against measured FACE log mass loss over 6 years. Four widely applied metrics of model performance demonstrated that the CWDDAT model can accurately predict CWD decomposition. The R 2 (squared Pearson’s correlation coefficient) between the simulation and measurement was 0.80 for the model calibration and 0.82 for the model validation ( P <0.01). The predicted mean mass loss from all logs was 5.4% lower than the measured mass loss and 1.4% lower than the calculated loss. The model was also used to assess the decomposition of mixed pine-hardwood CWD produced by Hurricane Hugo in 1989 on the Santee Experimental Forest in South Carolina, USA. The simulation reflected rapid CWD decomposition of the forest in this subtropical setting. The predicted dissolved organic carbon (DOC) derived from the CWD decomposition and incorporated into the mineral soil averaged 1.01 g C m -2 y -1 over the 30 years. The main agents for CWD mass loss were fungi (72.0%) and termites (24.5%), the remainder was attributed to a mix of other wood decomposers. These findings demonstrate the applicability of CWDDAT for large-scale assessments of CWD dynamics, and fine-scale considerations regarding the fate of CWD carbon.

54 ENVIRONMENTAL SCIENCES↗

Graph-based machine learning improves just-in-time defect prediction

The increasing complexity of today’s software requires the contribution of thousands of developers. This complex collaboration structure makes developers more likely to introduce defect-prone changes that lead to software faults. Determining when these defect-prone changes are introduced has proven challenging, and using traditional machine learning (ML) methods to make these determinations seems to have reached a plateau. In this work, we build contribution graphs consisting of developers and source files to capture the nuanced complexity of changes required to build software. By leveraging these contribution graphs, our research shows the potential of using graph-based ML to improve Just-In-Time (JIT) defect prediction. We hypothesize that features extracted from the contribution graphs may be better predictors of defect-prone changes than intrinsic features derived from software characteristics. We corroborate our hypothesis using graph-based ML for classifying edges that represent defect-prone changes. This new framing of the JIT defect prediction problem leads to remarkably better results. We test our approach on 14 open-source projects and show that our best model can predict whether or not a code change will lead to a defect with an F1 score as high as 77.55% and a Matthews correlation coefficient (MCC) as high as 53.16%. This represents a 152% higher F1 score and a 3% higher MCC over the state-of-the-art JIT defect prediction. We describe limitations, open challenges, and how this method can be used for operational JIT defect prediction.

59 BASIC BIOLOGICAL SCIENCES↗

Utah FORGE: Interferometric Synthetic Aperture Radar Data from 2023 and 2024

The dataset comprises Interferometric Synthetic Aperture Radar (InSAR) data from the TerraSAR-X and TanDEM-X satellite missions, covering the Utah FORGE site. This data includes interferometric pairs created using GMT-SAR processing software, chosen for their short orbital separations between May 1, 2023, and June 30, 2024. Included are various data and metadata, including Digital Elevation Models, unit vectors, and correlation coefficients. The dataset is packaged in several compressed tar files and formatted in NetCDF. To utilize this dataset, users will need software capable of handling NetCDF files and tools for decompressing tar files.

15 GEOTHERMAL ENERGY↗