Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “predictive”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Toward Predictive RANS and SRS Computations of Turbulent External Flows of Practical Interest

In this work, we investigate the main challenges to prediction of turbulent external flows of practical interest with Reynolds-Averaged Navier–Stokes equations (RANS) and Scale-Resolving Simulation (SRS) models. This represents a crucial step toward further developing and establishing these formulations so they can be confidently utilized in engineering problems without reference data. The study initiates by identifying the major challenges to prediction. A literature review is performed to illustrate their effects in RANS and SRS computations. Afterward, we evaluate the impact of the challenges to prediction by analyzing representative statistically steady and unsteady flows with prominent RANS and SRS methods. These include multiple turbulent viscosity and second-moment RANS closures, and hybrid and bridging SRS models. The results demonstrate the potential of the selected SRS models to predict engineering flows. Yet, they also show the importance of considering the challenges to prediction during the setup and conduction of numerical experiments. These can suppress the advantages of using SRS formulations. The data also indicate that only SRS models can confidently predict statistically unsteady flows. In contrast, the results demonstrate that mean-flow quantities of statistically steady flows can be efficiently calculated with RANS closures, especially second-moment closures. Among the selected SRS methods, bridging models reveal better suited for prediction due to their ability to prevent commutation errors and enable the robust evaluation of numerical and modeling errors. This last property allows the use of a new validation technique that does not require reference data.

42 ENGINEERING↗

Accuracy of predictions made by machine learned models for biocrude yields obtained from hydrothermal liquefaction of organic wastes

Hydrothermal liquefaction (HTL) has potential for converting abundant wet organic wastes into renewable fuels. Because HTL consists of a complex reaction network, deterministic, physics-based prediction of its biocrude yield is prohibitively difficult. Data-driven methods provide an alternative to the physics-based approach; however, rigorous testing must be performed to ensure the accuracy of predictions made by data-driven methods. To this end, a data set was assembled consisting of 570 data points appearing in the open literature. The data set was divided into training, validation, and test sub-sets and used for evaluating different machine learning regression approaches to predict biocrude yield. Among the tested algorithms, Random Forest and eXtreme Gradient Boosting (XGBoost) predicted biocrude yields in a test set that had not been used for training with the greatest accuracy, with root mean square errors (RMSE) of 8.34 and 8.57, respectively. Further refinement of the Random Forest model reduced its RMSE to 8.07. In comparison, predictions of a series of literature models resulted in RMSE ranging from 9.16 in the most accurate case to 27.6 in the least accurate; most literature models yielded RMSE values > 10. Using biocrude yield predictions from the most accurate Random Forest model and a probabilistic economic analysis found that the model accuracy is sufficient to prioritize allocation of resources based on projected minimum fuel selling price. In our report the models and analysis represent a major advance in the ability to use readily available data to predict biocrude yields on new feedstocks that have not previously been studied.

42 ENGINEERING↗

Evaluating performance of different generative adversarial networks for large-scale building power demand prediction

We report as an unsupervised-learning data-driven model, Generative Adversarial Networks (GANs) have recently attracted a lot of attention for various applications. There is potential to apply GANs for large-scale building power demand prediction, which is needed for power grid operation. However, there are many GAN variations and it is unclear which GAN is suitable for this application. To answer this question, this paper identifies five promising GANs (Original GAN, cGAN, SGAN, InfoGAN, and ACGAN) and evaluates their performance for predicting building power demand at a large scale. Physics-based building energy models are developed to generate training and reference data. A new evaluation indicator that combines accuracy and reproducibility is proposed to evaluate the performance of different GANs in predicting building power demand. The results show that SGAN and InfoGAN are not suitable because they cannot control the number of generated building samples for different building types. The prediction performance among the Original GAN, cGAN, and ACGAN can vary depending on training sample sizes and number of building types. If the training sample size is sufficiently large, Original GAN and cGAN can predict building power demand more accurately than ACGAN with the same number of samples. If training samples are limited, Original GAN provides better accuracy than cGAN and ACGAN. When the number of building types increase, the prediction accuracy increases for cGAN, decreases for ACGAN, and remains the same for Original GAN. As a result, cGAN and Original GAN are recommended for large-scale building power demand prediction.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Performance and power modeling and prediction using MuMMI and 10 machine learning methods

Energy-efficient scientific applications require insight into how high performance computing system features impact the applications' power and performance. This insight can result from the development of performance and power models. Here, in this article, we use the modeling and prediction tool MuMMI (Multiple Metrics Modeling Infrastructure) and 10 machine learning methods to model and predict performance and power consumption and compare their prediction error rates. We use an algorithm-based fault-tolerant linear algebra code and a multilevel checkpointing fault-tolerant heat distribution code to conduct our modeling and prediction study on the Cray XC40 Theta and IBM BG/Q Mira at Argonne National Laboratory and the Intel Haswell cluster Shepard at Sandia National Laboratories. Our experimental results show that the prediction error rates in performance and power using MuMMI are less than 10% for most cases. By utilizing the models for runtime, node power, CPU power, and memory power, we identify the most significant performance counters for potential application optimizations, and we predict theoretical outcomes of the optimizations. Based on two collected datasets, we analyze and compare the prediction accuracy in performance and power consumption using MuMMI and 10 machine learning methods.

97 MATHEMATICS AND COMPUTING↗

Early season prediction of within-field crop yield variability by assimilating CubeSat data into a crop model

Accurate early season predictions of crop yield at the within-field scale can be used to address a range of crop production, management, and precision agricultural challenges. While the remote sensing of within-field insights has been a research goal for many years, it is only recently that observations with the required spatio-temporal resolutions, together with efficient assimilation methods to integrate these into modeling frameworks, have become available to advance yield prediction efforts. Here we explore a yield prediction approach that combines daily high-resolution CubeSat imagery with the APSIM crop model. The approach employs APSIM to train a linear regression that relates simulated yield to simulated leaf area index (LAI). That relationship is then used to identify the optimal regression date at which the LAI provides the best prediction of yield: in this case, approximately 14 weeks prior to harvest. Instead of applying the regression on satellite imagery that is coincident, or closest to, the regression date, our method implements a particle filter that integrates CubeSat-based LAI into APSIM to provide end-of-season high-resolution (3 m) yield maps weeks before the optimal regression date. The approach is demonstrated on a rainfed maize field located in Nebraska, USA, where suitable collections of both imagery and in-situ data were available for assessment. The procedure does not require in-field data to calibrate the regression model, with results showing that even with a single assimilation step, it is possible to provide yield estimates with good accuracy up to 21 days before the optimal regression date. Yield spatial variability was reproduced reasonably well, with a strong correlation to independently collected measurements (R 2 = 0.73 and rRMSE = 12%). When the field averaged yield was compared, our approach reduced yield prediction error from 1 Mg/ha (control case based on a calibrated APSIM model), to 0.5 Mg/ha (using satellite imagery alone), and then to 0.2 Mg/ha (results with assimilation up to three weeks prior to the optimal regression date). Such a capacity to provide spatially explicit yield predictions early in the season has considerable potential to enhance digital agricultural goals and improve end-of-season yield predictions.

54 ENVIRONMENTAL SCIENCES↗

Reservoir Computing as a Tool for Climate Predictability Studies

Reduced-order dynamical models play a central role in developing our understanding of predictability of climate irrespective of whether we are dealing with the actual climate system or surrogate climate models. In this context, the linear inverse modeling (LIM) approach, by capturing a few essential interactions between dynamical components of the full system, has proven valuable in providing insights into predictability of the full system. We demonstrate that reservoir computing (RC), a form of learning suitable for systems with chaotic dynamics, provides an alternative nonlinear approach that improves on the predictive skill of the LIM approach. We do this in the example setting of predicting sea surface temperature in the North Atlantic in the preindustrial control simulation of a popular earth system model, the Community Earth System Model so that we can compare the performance of the new RC-based approach with the traditional LIM approach both when learning data are plentiful and when such data are more limited. The improved predictive skill of the RC approach over a wide range of conditions—larger number of retained EOF coefficients, extending well into the limited data regime, etc.—suggests that this machine-learning technique may have a use in climate predictability studies. While the possibility of developing a climate emulator—the ability to continue the evolution of the system on the attractor long after failing to be able to track the reference trajectory—is demonstrated in the Lorenz-63 system, it is suggested that further development of the RC approach may permit such uses of the new approach in more realistic predictability studies.

54 ENVIRONMENTAL SCIENCES↗

Prediction of inter-chain distance maps of protein complexes with 2D attention-based deep neural networks

Residue-residue distance information is useful for predicting tertiary structures of protein monomers or quaternary structures of protein complexes. Many deep learning methods have been developed to predict intra-chain residue-residue distances of monomers accurately, but few methods can accurately predict inter-chain residue-residue distances of complexes. We develop a deep learning method CDPred (i.e., Complex Distance Prediction) based on the 2D attention-powered residual network to address the gap. Tested on two homodimer datasets, CDPred achieves the precision of 60.94% and 42.93% for top L/5 inter-chain contact predictions (L: length of the monomer in homodimer), respectively, substantially higher than DeepHomo’s 37.40% and 23.08% and GLINTER’s 48.09% and 36.74%. Tested on the two heterodimer datasets, the top Ls/5 inter-chain contact prediction precision (Ls: length of the shorter monomer in heterodimer) of CDPred is 47.59% and 22.87% respectively, surpassing GLINTER’s 23.24% and 13.49%. Moreover, the prediction of CDPred is complementary with that of AlphaFold2-multimer.

59 BASIC BIOLOGICAL SCIENCES↗

Machine learning approaches to predict gestational age in normal and complicated pregnancies via urinary metabolomics analysis

The elucidation of dynamic metabolomic changes during gestation is particularly important for the development of methods to evaluate pregnancy status or achieve earlier detection of pregnancy-related complications. Some studies have constructed models to evaluate pregnancy status and predict gestational age using omics data from blood biospecimens; however, less invasive methods are desired. Here we propose a model to predict gestational age, using urinary metabolite information. In our prospective cohort study, we collected 2741 urine samples from 187 healthy pregnant women, 23 patients with hypertensive disorders of pregnancy, and 14 patients with spontaneous preterm birth. Using gas chromatography-tandem mass spectrometry, we identified 184 urinary metabolites that showed dynamic systematic changes in healthy pregnant women according to gestational age. A model to predict gestational age during normal pregnancy progression was constructed; the correlation coefficient between actual and predicted weeks of gestation was 0.86. The predicted gestational ages of cases with hypertensive disorders of pregnancy exhibited significant progression, compared with actual gestational ages. This is the first study to predict gestational age in normal and complicated pregnancies by using urinary metabolite information. Minimally invasive urinary metabolomics might facilitate changes in the prediction of gestational age in various clinical settings.

60 APPLIED LIFE SCIENCES↗

Machine learning prediction of incidence of Alzheimer’s disease using large-scale administrative health data

Nationwide population-based cohort provides a new opportunity to build an automated risk prediction model based on individuals’ history of health and healthcare beyond existing risk prediction models. We tested the possibility of machine learning models to predict future incidence of Alzheimer’s disease (AD) using large-scale administrative health data. From the Korean National Health Insurance Service database between 2002 and 2010, we obtained de-identified health data in elders above 65 years (N = 40,736) containing 4,894 unique clinical features including ICD-10 codes, medication codes, laboratory values, history of personal and family illness and socio-demographics. To define incident AD we considered two operational definitions: “definite AD” with diagnostic codes and dementia medication (n = 614) and “probable AD” with only diagnosis (n = 2026). We trained and validated random forest, support vector machine and logistic regression to predict incident AD in 1, 2, 3, and 4 subsequent years. For predicting future incidence of AD in balanced samples (bootstrapping), the machine learning models showed reasonable performance in 1-year prediction with AUC of 0.775 and 0.759, based on “definite AD” and “probable AD” outcomes, respectively; in 2-year, 0.730 and 0.693; in 3-year, 0.677 and 0.644; in 4-year, 0.725 and 0.683. The results were similar when the entire (unbalanced) samples were used. Important clinical features selected in logistic regression included hemoglobin level, age and urine protein level. This study may shed a light on the utility of the data-driven machine learning model based on large-scale administrative health data in AD risk prediction, which may enable better selection of individuals at risk for AD in clinical trials or early detection in clinical settings.

97 MATHEMATICS AND COMPUTING↗

Neural Network-Based Electric Vehicle Range Prediction for Smart Charging Optimization

Range prediction is a standard feature in most modern road vehicles, allowing drivers to make informed decisions about when to refuel. Most vehicles make range predictions through data- or model-driven means, monitoring the average fuel consumption rate or using a tuned vehicle model to predict fuel consumption. The uncertainty of future driving conditions makes the range prediction problem challenging, particularly for less pervasive battery electric vehicles (BEV). Most contemporary machine learning-based methods attempt to forecast the battery SOC discharge profile to predict vehicle range. In this work, we propose a novel approach using two recurrent neural networks (RNNs) to predict the remaining range of BEVs and the minimum charge required to safely complete a trip. Each RNN has two outputs that can be used for statistical analysis to account for uncertainties; the first loss function leads to mean and variance estimation (MVE), while the second results in bounded interval estimation (BIE). These outputs of the proposed RNNs are then used to predict the probability of a vehicle completing a given trip without charging, or if charging is needed, the remaining range and minimum charging required to finish the trip with high probability. Training data was generated using a low-order physics model to estimate vehicle energy consumption from historical drive cycle data collected from medium-duty last-mile delivery vehicles. Here, the proposed method demonstrated high accuracy in the presence of day-to-day route variability, with the root-mean-square error (RMSE) below 6% for both RNN models.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Integrating multi-modal remote sensing, deep learning, and attention mechanisms for yield prediction in plant breeding experiments

In both plant breeding and crop management, interpretability plays a crucial role in instilling trust in AI-driven approaches and enabling the provision of actionable insights. The primary objective of this research is to explore and evaluate the potential contributions of deep learning network architectures that employ stacked LSTM for end-of-season maize grain yield prediction. A secondary aim is to expand the capabilities of these networks by adapting them to better accommodate and leverage the multi-modality properties of remote sensing data. In this study, a multi-modal deep learning architecture that assimilates inputs from heterogeneous data streams, including high-resolution hyperspectral imagery, LiDAR point clouds, and environmental data, is proposed to forecast maize crop yields. The architecture includes attention mechanisms that assign varying levels of importance to different modalities and temporal features that, reflect the dynamics of plant growth and environmental interactions. The interpretability of the attention weights is investigated in multi-modal networks that seek to both improve predictions and attribute crop yield outcomes to genetic and environmental variables. This approach also contributes to increased interpretability of the model's predictions. The temporal attention weight distributions highlighted relevant factors and critical growth stages that contribute to the predictions. The results of this study affirm that the attention weights are consistent with recognized biological growth stages, thereby substantiating the network's capability to learn biologically interpretable features. Accuracies of the model's predictions of yield ranged from 0.82-0.93 R 2 ref in this genetics-focused study, further highlighting the potential of attention-based models. Further, this research facilitates understanding of how multi-modality remote sensing aligns with the physiological stages of maize. The proposed architecture shows promise in improving predictions and offering interpretable insights into the factors affecting maize crop yields, while demonstrating the impact of data collection by different modalities through the growing season. By identifying relevant factors and critical growth stages, the model's attention weights provide valuable information that can be used in both plant breeding and crop management. The consistency of attention weights with biological growth stages reinforces the potential of deep learning networks in agricultural applications, particularly in leveraging remote sensing data for yield prediction. To the best of our knowledge, this is the first study that investigates the use of hyperspectral and LiDAR UAV time series data for explaining/interpreting plant growth stages within deep learning networks and forecasting plot-level maize grain yield using late fusion modalities with attention mechanisms.

59 BASIC BIOLOGICAL SCIENCES↗

Using Ultrasound Image Augmentation and Ensemble Predictions to Prevent Machine-Learning Model Overfitting

Deep learning predictive models have the potential to simplify and automate medical imaging diagnostics by lowering the skill threshold for image interpretation. However, this requires predictive models that are generalized to handle subject variability as seen clinically. Here, we highlight methods to improve test accuracy of an image classifier model for shrapnel identification using tissue phantom image sets. Using a previously developed image classifier neural network—termed ShrapML—blind test accuracy was less than 70% and was variable depending on the training/test data setup, as determined by a leave one subject out (LOSO) holdout methodology. Introduction of affine transformations for image augmentation or MixUp methodologies to generate additional training sets improved model performance and overall accuracy improved to 75%. Further improvements were made by aggregating predictions across five LOSO holdouts. This was done by bagging confidences or predictions from all LOSOs or the top-3 LOSO confidence models for each image prediction. Top-3 LOSO confidence bagging performed best, with test accuracy improved to greater than 85% accuracy for two different blind tissue phantoms. This was confirmed by gradient-weighted class activation mapping to highlight that the image classifier was tracking shrapnel in the image sets. Overall, data augmentation and ensemble prediction approaches were suitable for creating more generalized predictive models for ultrasound image analysis, a critical step for real-time diagnostic deployment.

60 APPLIED LIFE SCIENCES↗

Using Machine-Learning Methods and Expert Prediction Probabilities to Forecast Solar Flares

It has long been known that studying connection between solar flares and properties of magnetic field in active regions is very important for understanding the flare physics and developing space weather forecasts. The Helioseismic and Magnetic Imager onboard the Solar Dynamics Observatory (SDO/HMI) obtains tremendous amounts of magnetic field data products. However the operational NOAA Space Weather Prediction Center (SWPC) forecasts of solar flares still represent prediction probabilities issued by the experts. In this research we investigate the possibilities to enhance the daily operational flare forecasts performed at the SWPC by developing a synergy of the expert predictions and physics-based criteria, and by employing machine-learning methods. Among the physics-based criteria we consider the descriptors of the Polarity Inversion Line (PIL) and Space weather HMI Active Region Patches (SHARP), and derive from them daily characteristics of the entire Sun. We also consider the daily descriptors of the GOES Soft X-Ray (SXR) 1-8 Angstroms flux such as the flare history of the previous days and averaged X-Ray flux. We estimate the effectiveness in separation of flaring and non-flaring cases for each characteristic, as well as for the expert prediction probabilities, and find that some PIL, SHARP and SXR descriptors are as effective as the expert prediction probabilities and should be considered to issue the flare forecast. Finally, we train and test several Machine-Learning classification algorithms (Support Vector Classifiers with various kernel functions, k-Nearest Neighbor Classifier, Random Forest Classifier, and Neural Networks) using the most effective descriptors and expert prediction probabilities, and compare the obtained predictions with the current SWPC forecasts.

Machine-Learning↗

Trajectory Prediction Accuracy and Error Sources for Regional Jet Descents: Results of a 2010 Flight Trial at Denver International Airport using a Global 5000 Test Aircraft - Part I

The Efficient Descent Advisor (EDA) controller automation tool generates trajectory-based speed, path, and altitude-profile advisories to facilitate efficient, continuous descents into congested terminal airspace. While prior field trials have assessed the trajectory-prediction accuracy for large jet (i.e., Boeing and Airbus) types, smaller (i.e., regional and business) jet types present unique challenges involving different descent procedures and Flight Management System (FMS) capabilities. A small-jet field trial was conducted at Denver in the fall of 2010 with the objective of measuring trajectory prediction accuracy and quantifying the primary sources of error. This paper uses data collected onboard a Bombardier Global 5000 test aircraft to quantify the size and sources of trajectory prediction error. Error sources were quantified for the 44 runs by incrementally replacing predicted data with data collected onboard the aircraft and measuring the effect on time error. Results for en-route descents, from prior to top of descent to the meter fix 60-120 nmi downstream, indicate that the aircraft arrived an average 15 seconds earlier than predicted, with a standard deviation of 10 seconds. Target Mach and CAS deceleration were found to be the two largest error sources. If CAS deceleration error was reduced using a typical, more predictable level flight deceleration then the arrival time prediction error in 2010 would be on par with a 2009 flight trial of Airbus and Boeing revenue flights. Four of the error sources, tracker jumps, CAS deceleration, target Mach, and path distance, lend themselves to significant reductions with modest to no changes to ATC automation andor procedures. Wind error and its impact on arrival time error was significantly reduced in 2010 compared to a 1994 flight test using NASAs Boeing 737 test aircraft.

Trajectory Prediction Error↗

Tau Positron Emission Tomography for Predicting Dementia in Individuals With Mild Cognitive Impairment

An accurate prognosis is especially pertinent in mild cognitive impairment (MCI), when individuals experience considerable uncertainty about future progression. To evaluate the prognostic value of tau positron emission tomography (PET) to predict clinical progression from MCI to dementia. This was a multicenter cohort study with external validation and a mean (SD) follow-up of 2.0 (1.1) years. Data were collected from centers in South Korea, Sweden, the US, and Switzerland from June 2014 to January 2024. Participant data were retrospectively collected and inclusion criteria were a baseline clinical diagnosis of MCI; longitudinal clinical follow-up; a Mini-Mental State Examination (MMSE) score greater than 22; and available tau PET, amyloid-β (Aβ) PET, and magnetic resonance imaging (MRI) scan less than 1 year from diagnosis. A total of 448 eligible individuals with MCI were included (331 in the discovery cohort and 117 in the validation cohort). None of these participants were excluded over the course of the study. Exposures included Tau PET, Aβ PET, and MRI. Positive results on tau PET (temporal meta–region of interest), Aβ PET (global; expressed in the standardized metric Centiloids), and MRI (Alzheimer disease [AD] signature region) was assessed using quantitative thresholds and visual reads. Clinical progression from MCI to all-cause dementia (regardless of suspected etiology) or to AD dementia (AD as suspected etiology) served as the primary outcomes. The primary analyses were receiver operating characteristics. In the discovery cohort, the mean (SD) age was 70.9 (8.5) years, 191 (58%) were male, the mean (SD) MMSE score was 27.1 (1.9), and 110 individuals with MCI (33%) converted to dementia (71 to AD dementia). Only the model with tau PET predicted all-cause dementia (area under the receiver operating characteristic curve [AUC], 0.75; 95% CI, 0.70-0.80) better than a base model including age, sex, education, and MMSE score (AUC, 0.71; 95% CI, 0.65-0.77; P = .02), while the models assessing the other neuroimaging markers did not improve prediction. In the validation cohort, tau PET replicated in predicting all-cause dementia. Compared to the base model (AUC, 0.75; 95% CI, 0.69-0.82), prediction of AD dementia in the discovery cohort was significantly improved by including tau PET (AUC, 0.84; 95% CI, 0.79-0.89; P < .001), tau PET visual read (AUC, 0.83; 95% CI, 0.78-0.88; P = .001), and Aβ PET Centiloids (AUC, 0.83; 95% CI, 0.78-0.88; P = .03). In the validation cohort, only the tau PET and the tau PET visual reads replicated in predicting AD dementia. In this study, tau-PET showed the best performance as a stand-alone marker to predict progression to dementia among individuals with MCI. This suggests that, for prognostic purposes in MCI, a tau PET scan may be the best currently available neuroimaging marker.

59 BASIC BIOLOGICAL SCIENCES↗

Utility of anthesis–silking interval information to predict grain yield under water and nitrogen limited conditions

Delayed silking relative to pollen shed, measured as the anthesis–silking interval (ASI, the period between pollen shed and silking), is a good indicator of response to abiotic stresses in maize ( Zea mays L.). This research was conducted to investigate how ASI is affected by nitrogen (N) and water availability and to assess the utility of ASI to indirectly predict grain yield (GY) under contrasting water and N treatments. Two experiments were conducted in Hancock, WI, in 2018 and 2019. One experiment (Diverse hybrids) included 302 hybrids resulting from the cross of diverse inbred lines by a single tester evaluated at four different treatment levels resulting from combining nonlimited and low N with nonlimited and low water treatments. The second experiment (NSS FAC) included a set of 408 hybrids derived from the cross of biparental doubled-haploid lines from 13 factorial populations and evaluated under nonlimited and low N treatments. Anthesis and silk time in growing degree days, and GY (Mg ha -1 ) were measured. Genomic prediction was assessed using a genomic best linear unbiased prediction model, and predictive ability was calculated as the correlation between genomic predictions and adjusted means in the different treatments. Predictive ability ranged from .15 to .49 for NSS FAC and from .06 to .51 for Diverse hybrids across traits and treatments. The ASI was a good indicator of stress and showed higher heritability than GY in the limited treatments for both experiments; however, it did not improve yield predictability.

59 BASIC BIOLOGICAL SCIENCES↗

Functional Relevance of CASP16 Nucleic Acid Predictions as Evaluated by Structure Providers

ABSTRACT Accurate biomolecular structure prediction enables the prediction of mutational effects, the speculation of function based on predicted structural homology, the analysis of ligand binding modes, experimental model building, and many other applications. Such algorithms to predict essential functional and structural features remain out of reach for biomolecular complexes containing nucleic acids. Here, we report a quantitative and qualitative evaluation of nucleic acid structures for the CASP16 blind prediction challenge by 12 of the experimental groups who provided nucleic acid targets. Blind predictions accurately model secondary structure and some aspects of tertiary structure, including reasonable global folds for some complex RNAs; however, predictions often lack accuracy in the regions of highest functional importance. All models have inaccuracies in non‐canonical regions where, for example, the nucleic‐acid backbone bends, deviating from an A‐form helix geometry, or a base forms a non‐standard hydrogen bond (not a Watson‐Crick base pair). These bends and non‐canonical interactions are integral to forming functionally important regions such as RNA enzymatic active sites. Additionally, the modeling of conserved and functional interfaces between nucleic acids and ligands, proteins, or other nucleic acids remains poor. For some targets, the experimental structures may not represent the only structure the biomolecular complex occupies in solution or in its functional life cycle, posing a future challenge for the community.

Biochemistry & Molecular Biology↗

Empirical relationships between environmental factors and soil organic carbon produce comparable prediction accuracy as the Machine Learning

Accurate representation of environmental controllers of soil organic carbon (SOC) stocks in Earth System Model (ESM) land models could reduce uncertainties in future carbon-climate feedback projections. Using empirical relationships between environmental factors and SOC stocks to evaluate land models can help modelers understand prediction biases beyond what can be achieved with the observed SOC stocks alone. In this study, we used 31 observed environmental factors, field SOC observations (n = 6,213) from the continental US, and two Machine Learning approaches [Random Forest (RF) and Generalized Additive Modeling (GAM)] to (1) select important environmental predictors of SOC stocks, (2) derive empirical relationships between environmental factors and SOC stocks, and (3) use the derived relationships to predict SOC stocks and compare the prediction accuracy of simpler model developed with the machine learning predictions. Out of the 31 environmental factors we investigated, 12 were identified as important predictors of SOC stocks by the RF approach. In contrast, the GAM approach identified six (of those 12) environmental factors as important controllers of SOC stocks: potential evapotranspiration, normalized difference vegetation index, soil drainage condition, precipitation, elevation, and net primary productivity. The GAM approach showed minimal SOC predictive importance of the remaining six environmental factors identified by the RF approach. Our derived empirical relations produced comparable prediction accuracy as the GAM and RF approach using only a subset of environmental factors. The empirical relationships we derived using the GAM approach can serve as important benchmarks to evaluate environmental control representations of SOC stocks in ESMs, which could reduce uncertainty in predicting future carbon-climate feedbacks.

54 ENVIRONMENTAL SCIENCES↗