Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “predictive”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Prediction of Geomagnetic Activity and Key Parameters in High-latitude Ionosphere

Prediction of geomagnetic activity and related events in the Earth's magnetosphere and ionosphere are important tasks of US Space Weather Program. Prediction reliability is dependent on the prediction method, and elements included in the prediction scheme. Two of the main elements of such prediction scheme are: an appropriate geomagnetic activity index, and an appropriate coupling function (the combination of solar wind parameters providing the best correlation between upstream solar wind data and geomagnetic activity). We have developed a new index of geomagnetic activity, the Polar Magnetic (PM) index and an improved version of solar wind coupling function. PM index is similar to the existing polar cap PC index but it shows much better correlation with upstream solar wind/IMF data and other events in the magnetosphere and ionosphere. We investigate the correlation of PM index with upstream solar wind/IMF data for 10 years (1995-2004) that include both low and high solar activity. We also have introduced a new prediction function for the predicting of cross-polar-cap voltage and Joule heating based on using both PM index and upstream solar wind/IMF data. As we show such prediction function significantly increase the reliability of prediction of these important parameters. The correlation coefficients between the actual and predicted values of these parameters are approx. 0.9 and higher.

Khazanov, George V.↗

Uncertainty and Sensitivity Analyses of a Two-Parameter Impedance Prediction Model

This paper presents comparisons of predicted impedance uncertainty limits derived from Monte-Carlo-type simulations with a Two-Parameter (TP) impedance prediction model and measured impedance uncertainty limits based on multiple tests acquired in NASA Langley test rigs. These predicted and measured impedance uncertainty limits are used to evaluate the effects of simultaneous randomization of each input parameter for the impedance prediction and measurement processes. A sensitivity analysis is then used to further evaluate the TP prediction model by varying its input parameters on an individual basis. The variation imposed on the input parameters is based on measurements conducted with multiple tests in the NASA Langley normal incidence and grazing incidence impedance tubes; thus, the input parameters are assigned uncertainties commensurate with those of the measured data. These same measured data are used with the NASA Langley impedance measurement (eduction) processes to determine the corresponding measured impedance uncertainty limits, such that the predicted and measured impedance uncertainty limits (95% confidence intervals) can be compared. The measured reactance 95% confidence intervals encompass the corresponding predicted reactance confidence intervals over the frequency range of interest. The same is true for the confidence intervals of the measured and predicted resistance at near-resonance frequencies, but the predicted resistance confidence intervals are lower than the measured resistance confidence intervals (no overlap) at frequencies away from resonance. A sensitivity analysis indicates the discharge coefficient uncertainty is the major contributor to uncertainty in the predicted impedances for the perforate-over-honeycomb liner used in this study. This insight regarding the relative importance of each input parameter will be used to guide the design of experiments with test rigs currently being brought on-line at NASA Langley.

Jones, M. G.↗

Predictive Information: Status or Alert Information?

Previous research investigating the efficacy of predictive information for detecting and diagnosing aircraft system failures found that subjects like to have predictive information concerning when a parameter would reach an alert range. This research focused on where the predictive information should be located, whether the information should be more closely associated with the parameter information or with the alert information. Each subject saw 3 forms of predictive information: (1) none, (2) a predictive alert message, and (3) predictive information on the status display. Generally, subjects performed better and preferred to have predictive information available although the difference between status and alert predictive information was minimal. Overall, for detection and recalling what happened, status predictive information is best; however for diagnosis, alert predictive information holds a slight edge.

Trujillo, Anna C.↗

A Recursive Multi-step Machine Learning Approach for Airport Configuration Prediction

Airport configuration selection is a complex decision-making process that involves several operational and human factors. In this paper we propose a novel recursive multi-step machine learning (ML) approach to predict airport configuration. The multi-step approach guarantees stability of the predicted configuration by taking as input the configuration predicted at the previous time step. The features of the proposed model include weather data, future arrival and departure counts and current configuration. Due to the importance of arrival and departure counts in predicting the airport configuration, arrival counts are calculated using landing time predictions selected from physics-based landing time predictions available in FAA System Wide Information Management data feeds for each flight. The selection rules were developed and refined to select the most accurate time for different phases of flight. The proposed model predicts the airport configurations up to 6 hours ahead. In this paper we show the predictive performance of the proposed model for six major US airports, including Charlotte Douglas International Airport (CLT), Dallas/Fort Worth International Airport (DFW), John F. Kennedy International Airport (JFK), Newark Liberty International Airport (EWR), LaGuardia Airport (LGA) and Dallas Love Field Airport (DAL). We trained and evaluated models on 2019 and 2020 data in order to study the effect of the pandemic and how changes in traffic patterns affected the performance of the proposed model. Results are compared with a baseline assuming no airport configuration changes. In our results for DFW, we obtained a prediction accuracy of 89.3% for 3 hours ahead prediction, and 82.8% for 6 hours ahead when applied on 2019 data.

machine learning↗

A Recursive Multi-step Machine Learning Approach for Airport Configuration Prediction

Airport configuration selection is a complex decision-making process that involves several operational and human factors. In this paper we propose a novel recursive multi-step machine learning (ML) approach to predict airport configuration. The multi-step approach guarantees stability of the predicted configuration by taking as input the configuration predicted at the previous time step. The features of the proposed model include weather data, future arrival and departure counts and current configuration. Due to the importance of arrival and departure counts in predicting the airport configuration, arrival counts are calculated using landing time predictions selected from physics-based landing time predictions available in FAA System Wide Information Management data feeds for each flight. The selection rules were developed and refined to select the most accurate time for different phases of flight. The proposed model predicts the airport configurations up to 6 hours ahead. In this paper we show the predictive performance of the proposed model for six major US airports, including Charlotte Douglas International Airport (CLT), Dallas/Fort Worth International Airport (DFW), John F. Kennedy International Airport (JFK), Newark Liberty International Airport (EWR), LaGuardia Airport (LGA) and Dallas Love Field Airport (DAL). We trained and evaluated models on 2019 and 2020 data in order to study the effect of the pandemic and how changes in traffic patterns affected the performance of the proposed model. Results are compared with a baseline assuming no airport configuration changes. In our results for DFW, we obtained a prediction accuracy of 89.3% for 3 hours ahead prediction, and 82.8% for 6 hours ahead when applied on 2019 data.

machine learning↗

Assimilation of SMAP Observations Over Land Improves the Simulation and Prediction of Tropical Cyclone Idai

This work is focused on the role of soil moisture in the prediction of tropical cyclones (TCs) approaching land and after landfall. Soil moisture conditions can impact the circulation and structure of an existing tropical cyclone (TC) when part or all of the circulation is over land. For example, dry land surface conditions may lead to faster dissipation of a TC over land (often associated with changes in precipitation structure), whereas very wet conditions may help sustain or in rare cases re-intensify a TC. Moreover, the presence of strong soil moisture gradients may affect the symmetry and development of the TC circulation leading to changes in its over-land track. While the link between soil moisture conditions and TC evolution in proximity to land is relatively well understood in theory, applications of these findings in the context of numerical weather prediction (NWP) have been limited. Here we present a case study that explores the potential of improving TC predictions through an improved soil moisture initialization in an NWP framework. Specifically, we examine the impact of assimilating observations from the NASA Soil Moisture Active Passive (SMAP) mission into the NASA Goddard Earth Observing System (GEOS) global weather model on the prediction of South-West Indian Ocean TC Idai (2019). SMAP provides accurate L-band (1.4 GHz) brightness temperatures (Tb) observations that are sensitive to soil moisture globally and at high revisit times of 2-3 days. It has previously been shown that the assimilation of SMAP Tb observations significantly improves modeled land surface states. Thus, it is expected that SMAP can be used to constrain land surface initial conditions and potentially benefit TC forecasts. Here we present two sets of retrospective forecasts of TC Idai that are compared in an Observing System Experiment framework at ¼ degree resolution: (i) forecasts initialized from an analysis that is comparable to the GEOS operational analysis (without SMAP Tb assimilation) and (ii) forecasts initialized from an analysis that additionally assimilates SMAP brightness temperature observations over land using a weakly-coupled land analysis. We find that the assimilation of SMAP meaningfully improves the representation of TC Idai’s structure as well as the prediction of its intensity and track. The analyzed TC size, as measured by the wind speed radius, is improved by up to 18% in the analysis with SMAP assimilation relative to the control run. The forecast intensity error, measured against the observed intensity, is reduced by up to 23%. At the 1/4-degree resolution used here, GEOS unavoidably under-estimates TC intensity and over-estimates TC size. The SMAP assimilation therefore corrects the model in the right direction, leading to a storm that is more energetic and more compact. Furthermore, we find that the along-track forecast error is reduced by up to 34%, indicating a more accurate propagation speed, which is consistent with the fact that TC speed over land is strongly affected by surface processes. The impact of SMAP assimilation on the forecast cross-track error is neutral. Across the TC forecast skill metrics used here, the improvements from SMAP DA are largest at lead times of 36 to 72 hours, suggesting that the predictability of forecasts at shorter lead times may be dominated by short-term convective processes, while the land and its longer memory gains in importance as a source of predictability on a 2-3 day timescale. We further investigated the underlying mechanisms leading to the skill improvements from SMAP data assimilation by isolating the land areas that directly influence TC Idai using a back trajectory analysis. We find that the assimilation of SMAP leads to wetter soil moisture conditions that cause an increased latent heat flux, which ultimately results in TC analyzed representation that has higher column-integrated total moisture content and total energy compared to the analysis in the control run without SMAP assimilation. Overall, the results highlight that the assimilation of SMAP observations into a global numerical weather prediction model can lead to pronounced improvements of TC predictions. This is a crucial step towards a better mitigation of the socio-economic impact of landfalling TCs and thus safeguarding human lives. Finally, our study presents an event-based approach that assesses the impact of land data assimilation for a particular weather event rather than by globally averaging differences in skill. We argue that global skill assessments – while necessary – can mute the impact of land data assimilation, because the land’s influence on the atmosphere is constrained to certain locations and certain times. Instead, the event-based approach better highlights the true potential of land data assimilation in the context of NWP, especially for extreme events when accurate predictions are critical.

Jana Kolassa↗

Assimilation of SMAP Observations Over Land Improves the Simulation and Prediction of Tropical Cyclone Idai

This work is focused on the role of soil moisture in the prediction of tropical cyclones (TCs) approaching land and after landfall. Soil moisture conditions can impact the circulation and structure of an existing tropical cyclone (TC) when part or all of the circulation is over land. For example, dry land surface conditions may lead to faster dissipation of a TC over land (often associated with changes in precipitation structure), whereas very wet conditions may help sustain or in rare cases re-intensify a TC. Moreover, the presence of strong soil moisture gradients may affect the symmetry and development of the TC circulation leading to changes in its over-land track. While the link between soil moisture conditions and TC evolution in proximity to land is relatively well understood in theory, applications of these findings in the context of numerical weather prediction (NWP) have been limited. Here we present a case study that explores the potential of improving TC predictions through an improved soil moisture initialization in an NWP framework. Specifically, we examine the impact of assimilating observations from the NASA Soil Moisture Active Passive (SMAP) mission into the NASA Goddard Earth Observing System (GEOS) global weather model on the prediction of South-West Indian Ocean TC Idai (2019). SMAP provides accurate L-band (1.4 GHz) brightness temperatures (Tb) observations that are sensitive to soil moisture globally and at high revisit times of 2-3 days. It has previously been shown that the assimilation of SMAP Tb observations significantly improves modeled land surface states. Thus, it is expected that SMAP can be used to constrain land surface initial conditions and potentially benefit TC forecasts. Here we present two sets of retrospective forecasts of TC Idai that are compared in an Observing System Experiment framework at ¼ degree resolution: (i) forecasts initialized from an analysis that is comparable to the GEOS operational analysis (without SMAP Tb assimilation) and (ii) forecasts initialized from an analysis that additionally assimilates SMAP brightness temperature observations over land using a weakly-coupled land analysis. We find that the assimilation of SMAP meaningfully improves the representation of TC Idai’s structure as well as the prediction of its intensity and track. The analyzed TC size, as measured by the wind speed radius, is improved by up to 18% in the analysis with SMAP assimilation relative to the control run. The forecast intensity error, measured against the observed intensity, is reduced by up to 23%. At the 1/4-degree resolution used here, GEOS unavoidably under-estimates TC intensity and over-estimates TC size. The SMAP assimilation therefore corrects the model in the right direction, leading to a storm that is more energetic and more compact. Furthermore, we find that the along-track forecast error is reduced by up to 34%, indicating a more accurate propagation speed, which is consistent with the fact that TC speed over land is strongly affected by surface processes. The impact of SMAP assimilation on the forecast cross-track error is neutral. Across the TC forecast skill metrics used here, the improvements from SMAP DA are largest at lead times of 36 to 72 hours, suggesting that the predictability of forecasts at shorter lead times may be dominated by short-term convective processes, while the land and its longer memory gains in importance as a source of predictability on a 2-3 day timescale. We further investigated the underlying mechanisms leading to the skill improvements from SMAP data assimilation by isolating the land areas that directly influence TC Idai using a back trajectory analysis. We find that the assimilation of SMAP leads to wetter soil moisture conditions that cause an increased latent heat flux, which ultimately results in TC analyzed representation that has higher column-integrated total moisture content and total energy compared to the analysis in the control run without SMAP assimilation. Overall, the results highlight that the assimilation of SMAP observations into a global numerical weather prediction model can lead to pronounced improvements of TC predictions. This is a crucial step towards a better mitigation of the socio-economic impact of landfalling TCs and thus safeguarding human lives. Finally, our study presents an event-based approach that assesses the impact of land data assimilation for a particular weather event rather than by globally averaging differences in skill. We argue that global skill assessments – while necessary – can mute the impact of land data assimilation, because the land’s influence on the atmosphere is constrained to certain locations and certain times. Instead, the event-based approach better highlights the true potential of land data assimilation in the context of NWP, especially for extreme events when accurate predictions are critical.

Jana Kolassa↗

Comparing Magnetopause Predictions From Two MHD Models During A Geomagnetic Storm and A Quiet Period

Magnetopause location is an important prediction of numerical simulations of the magnetosphere, yet the models can err, either under-predicting or over-predicting the motion of the boundary. This study compares results from two of the most widely used magnetohydrodynamic (MHD) models, the Lyon–Fedder–Mobarry (LFM) model and the Space Weather Modeling Framework (SWMF), to data from the GOES 13 and 15 satellites during the geomagnetic storm on 22 June 2015, and to THEMIS A, D, and E during a quiet period on 31 January 2013. The models not only reproduce the magnetopause crossings of the spacecraft during the storm, but they also predict spurious magnetopause motion after the crossings seen in the GOES data. We investigate the possible causes of the over-predictions during the storm and find the following. First, using different ionospheric conductance models does not significantly alter predictions of the magnetopause location. Second, coupling the Rice Convection Model (RCM) to the MHD codes improves the SWMF magnetopause predictions more than it does for the LFM predictions. Third, the SWMF produces a stronger ring current than LFM, both with and without the RCM and regardless of the LFM spatial resolution. During the non-storm event, LFM predicts the THEMIS magnetopause crossings due to the southward interplanetary magnetic field better than the SWMF. Additionally, increasing the LFM spatial grid resolution improves the THEMIS predictions, while increasing the SWMF grid resolutions does not.

magnetohydrodynamics↗

DNCON2_Inter: predicting interchain contacts for homodimeric and homomultimeric protein complexes using multiple sequence alignments of monomers and deep learning

Deep learning methods that achieved great success in predicting intrachain residue-residue contacts have been applied to predict interchain contacts between proteins. However, these methods require multiple sequence alignments (MSAs) of a pair of interacting proteins (dimers) as input, which are often difficult to obtain because there are not many known protein complexes available to generate MSAs of sufficient depth for a pair of proteins. In recognizing that multiple sequence alignments of a monomer that forms homomultimers contain the co-evolutionary signals of both intrachain and interchain residue pairs in contact, we applied DNCON2 (a deep learning-based protein intrachain residue-residue contact predictor) to predict both intrachain and interchain contacts for homomultimers using multiple sequence alignment (MSA) and other co-evolutionary features of a single monomer followed by discrimination of interchain and intrachain contacts according to the tertiary structure of the monomer. We name this tool DNCON2_Inter. Allowing true-positive predictions within two residue shifts, the best average precision was obtained for the Top-L/10 predictions of 22.9% for homodimers and 17.0% for higher-order homomultimers. In some instances, especially where interchain contact densities are high, DNCON2_Inter predicted interchain contacts with 100% precision. We also developed Con_Complex, a complex structure reconstruction tool that uses predicted contacts to produce the structure of the complex. Using Con_Complex, we show that the predicted contacts can be used to accurately construct the structure of some complexes. Our experiment demonstrates that monomeric multiple sequence alignments can be used with deep learning to predict interchain contacts of homomeric proteins.

59 BASIC BIOLOGICAL SCIENCES↗

Enhancing alphafold-multimer-based protein complex structure prediction with MULTICOM in CASP15

To enhance the AlphaFold-Multimer-based protein complex structure prediction, we developed a quaternary structure prediction system (MULTICOM) to improve the input fed to AlphaFold-Multimer and evaluate and refine its outputs. MULTICOM samples diverse multiple sequence alignments (MSAs) and templates for AlphaFold-Multimer to generate structural predictions by using both traditional sequence alignments and Foldseek-based structure alignments, ranks structural predictions through multiple complementary metrics, and refines the structural predictions via a Foldseek structure alignment-based refinement method. The MULTICOM system with different implementations was blindly tested in the assembly structure prediction in the 15th Critical Assessment of Techniques for Protein Structure Prediction (CASP15) in 2022 as both server and human predictors. MULTICOM_qa ranked 3 rd among 26 CASP15 server predictors and MULTICOM_human ranked 7 th among 87 CASP15 server and human predictors. The average TM-score of the first predictions submitted by MULTICOM_qa for CASP15 assembly targets is ~0.76, 5.3% higher than ~0.72 of the standard AlphaFold-Multimer. The average TM-score of the best of top 5 predictions submitted by MULTICOM_qa is ~0.80, about 8% higher than ~0.74 of the standard AlphaFold-Multimer. Moreover, the Foldseek Structure Alignment-based Multimer structure Generation (FSAMG) method outperforms the widely used sequence alignment-based multimer structure generation.

59 BASIC BIOLOGICAL SCIENCES↗

A Genome-Based Model to Predict the Virulence of Pseudomonas aeruginosa Isolates

ABSTRACT: Variation in the genome of Pseudomonas aeruginosa , an important pathogen, can have dramatic impacts on the bacterium’s ability to cause disease. We therefore asked whether it was possible to predict the virulence of P. aeruginosa isolates based on their genomic content. We applied a machine learning approach to a genetically and phenotypically diverse collection of 115 clinical P. aeruginosa isolates using genomic information and corresponding virulence phenotypes in a mouse model of bacteremia. We defined the accessory genome of these isolates through the presence or absence of accessory genomic elements (AGEs), sequences present in some strains but not others. Machine learning models trained using AGEs were predictive of virulence, with a mean nested cross-validation accuracy of 75% using the random forest algorithm. However, individual AGEs did not have a large influence on the algorithm’s performance, suggesting instead that virulence predictions are derived from a diffuse genomic signature. These results were validated with an independent test set of 25 P. aeruginosa isolates whose virulence was predicted with 72% accuracy. Machine learning models trained using core genome single-nucleotide variants and whole-genome k-mers also predicted virulence. Our findings are a proof of concept for the use of bacterial genomes to predict pathogenicity in P. aeruginosa and highlight the potential of this approach for predicting patient outcomes. IMPORTANCE Pseudomonas aeruginosa is a clinically important Gram-negative opportunistic pathogen. P. aeruginosa shows a large degree of genomic heterogeneity both through variation in sequences found throughout the species (core genome) and through the presence or absence of sequences in different isolates (accessory genome). P. aeruginosa isolates also differ markedly in their ability to cause disease. In this study, we used machine learning to predict the virulence level of P. aeruginosa isolates in a mouse bacteremia model based on genomic content. We show that both the accessory and core genomes are predictive of virulence. This study provides a machine learning framework to investigate relationships between bacterial genomes and complex phenotypes such as virulence.

59 BASIC BIOLOGICAL SCIENCES↗

Wave Prediction using X-Band Radar (Final Summary Report)

Model Predictive Control (MPC) is the best controls framework available to wave energy today. It allows for the maximization of energy yield from a WEC system while respecting various constraints and losses in the system. This type of constrained optimization is a critical capability for the techno-economic optimization of WEC devices and is difficult to do effectively with causal controls frameworks. MPC requires a future prediction of wave excitation forces, requiring wave prediction on a future time horizon of up to 30 seconds. The future horizon required is highly dependent on the WEC and PTO type available. To our knowledge no-one has been able to implement MPC on a WEC device at sea, due to the fact that phase-resolved wave-prediction is not a capability that has been sufficiently developed to date. The key objectives of this project are focused on developing “industry-ready” wave prediction technology building blocks that can be applied to a wide range of different wave energy conversion topologies. Our collaborative work with device developers has demonstrated that the level of improvement attainable for any particular device is heavily dependent on the device and PTO topology chosen. The range of annual average power capture improvements is on the order of 10% to over 100%. However, the important aspect is that we can get within about 10% of theoretical upper limits even when considering errors in the wave prediction and more importantly, we can optimize performance while respecting various constraints such as motion amplitude or peak structural loads. The combination of improved performance while limiting structural loads has a net effect of reducing the levelized cost of energy from these emerging technologies and allows us to design control laws that meet an economic optimum. Efforts under a previous project on controls has focused on developing a set of wave-prediction algorithms that work by leveraging a network of wave sensing measurement buoys that are equipped with real-time telemetry. The present work is focusing primarily on using X-band radar to measure the wave-field around the WEC to predict the future excitation forces on the structure and combining the measurements of these different sources to create improvements in the wave-prediction accuracy.

16 TIDAL AND WAVE POWER↗

Frost prediction using machine learning and deep neural network models

This study describes accurate, computationally efficient models that can be implemented for practical use in predicting frost events for point-scale agricultural applications. Frost damage in agriculture is a costly burden to farmers and global food security alike. Timely prediction of frost events is important to reduce the cost of agricultural frost damage and traditional numerical weather forecasts are often inaccurate at the field-scale in complex terrain. In this paper, we developed machine learning (ML) algorithms for the prediction of such frost events near Alcalde, NM at the point-scale. ML algorithms investigated include deep neural network, convolution neural networks, and random forest models at lead-times of 6–48 h. Our results show promising accuracy (6-h prediction RMSE = 1.53–1.72°C) for use in frost and minimum temperature prediction applications. Seasonal differences in model predictions resulted in a slight negative bias during Spring and Summer months and a positive bias in Fall and Winter months. Additionally, we tested the model transferability by continuing training and testing using data from sensors at a nearby farm. We calculated the feature importance of the random forest models and were able to determine which parameters provided the models with the most useful information for predictions. We determined that soil temperature is a key parameter in longer term predictions (>24 h), while other temperature related parameters provide the majority of information for shorter term predictions. The model error compared favorable to previous ML based frost studies and outperformed the physically based High Resolution Rapid Refresh forecasting system making our ML-models attractive for deployment toward real-time monitoring of frost events and damage at commercial farming operations.

97 MATHEMATICS AND COMPUTING↗

Automated Cloud Based Long Short-Term Memory Neural Network Based SWE Prediction

Snow derived water is a critical component of the US water supply. Measurements of the Snow Water Equivalent (SWE) and associated predictions of peak SWE and snowmelt onset are essential inputs for water management efforts. This paper aims to develop an integrated framework for real-time data ingestion, estimation, prediction and visualization of SWE based on daily snow datasets. In particular, we develop a data-driven approach for estimating and predicting SWE dynamics using the Long Short-Term Memory neural network (LSTM) method. Our approach uses historical datasets (precipitation, air temperature, SWE, and snow thickness) collected at NRCS Snow Telemetry (SNOTEL) stations to train the LSTM network and current year data to predict SWE behavior. The performance of our prediction was compared for different prediction dates and prediction training datasets. Our results suggest that the proposed LSTM network can be an efficient tool for forecasting the SWE timeseries, as well as Peak SWE and snowmelt timing. Results showed that the window size impacts the model performance (where the Nash Sutcliffe efficiency (NSE) ranged from 0.96 to 0.85 and the Rooted Mean Square Error (RMSE) ranged from 0.038 to 0.07) with an optimum number that should be calibrated for different stations and climate conditions. In addition, by implementing the LSTM prediction capability in a cloud based site-monitoring platform, we automate model-data integration. By making the data accessible through a graphical web interface and an underlying API which exposes both training and prediction capabilities. The associated results can be made easily accessible to a broad range of stakeholders.

54 ENVIRONMENTAL SCIENCES↗

Enhanced Co-Expression Extrapolation (COXEN) Gene Selection Method for Building Anti-Cancer Drug Response Prediction Models

The co-expression extrapolation (COXEN) method has been successfully used in multiple studies to select genes for predicting the response of tumor cells to a specific drug treatment. Here, we enhance the COXEN method to select genes that are predictive of the efficacies of multiple drugs for building general drug response prediction models that are not specific to a particular drug. The enhanced COXEN method first ranks the genes according to their prediction power for each individual drug and then takes a union of top predictive genes of all the drugs, among which the algorithm further selects genes whose co-expression patterns are well preserved between cancer cases for building prediction models. We apply the proposed method on benchmark in vitro drug screening datasets and compare the performance of prediction models built based on the genes selected by the enhanced COXEN method to that of models built on genes selected by the original COXEN method and randomly picked genes. Models built with the enhanced COXEN method always present a statistically significantly improved prediction performance (adjusted p-value ≤ 0.05). Our results demonstrate the enhanced COXEN method can dramatically increase the power of gene expression data for predicting drug response.

60 APPLIED LIFE SCIENCES↗

Multi-Temporal Predictive Modelling of Sorghum Biomass Using UAV-Based Hyperspectral and LiDAR Data

High-throughput phenotyping using high spatial, spectral, and temporal resolution remote sensing (RS) data has become a critical part of the plant breeding chain focused on reducing the time and cost of the selection process for the “best” genotypes with respect to the trait(s) of interest. In this paper, the potential of accurate and reliable sorghum biomass prediction using visible and near infrared (VNIR) and short-wave infrared (SWIR) hyperspectral data as well as light detection and ranging (LiDAR) data acquired by sensors mounted on UAV platforms is investigated. Predictive models are developed using classical regression-based machine learning methods for nine experiments conducted during the 2017 and 2018 growing seasons at the Agronomy Center for Research and Education (ACRE) at Purdue University, Indiana, USA. The impact of the regression method, data source, timing of RS and field-based biomass reference data acquisition, and the number of samples on the prediction results are investigated. R2 values for end-of-season biomass ranged from 0.64 to 0.89 for different experiments when features from all the data sources were included. Geometry-based features derived from the LiDAR point cloud to characterize plant structure and chemistry-based features extracted from hyperspectral data provided the most accurate predictions. Evaluation of the impact of the time of data acquisition during the growing season on the prediction results indicated that although the most accurate and reliable predictions of final biomass were achieved using remotely sensed data from mid-season to end-of-season, predictions in mid-season provided adequate results to differentiate between promising varieties for selection. The analysis of variance (ANOVA) of the accuracies of the predictive models showed that both the data source and regression method are important factors for a reliable prediction; however, the data source was more important with 69% significance, versus 28% significance for the regression method.

09 BIOMASS FUELS↗

Machine Learning Framework for Conotoxin Class and Molecular Target Prediction

Conotoxins are small and highly potent neurotoxic peptides derived from the venom of marine cone snails which have captured the interest of the scientific community due to their pharmacological potential. These toxins display significant sequence and structure diversity, which results in a wide range of specificities for several different ion channels and receptors. Despite the recognized importance of these compounds, our ability to determine their binding targets and toxicities remains a significant challenge. Predicting the target receptors of conotoxins, based solely on their amino acid sequence, remains a challenge due to the intricate relationships between structure, function, target specificity, and the significant conformational heterogeneity observed in conotoxins with the same primary sequence. We have previously demonstrated that the inclusion of post-translational modifications, collisional cross sections values, and other structural features, when added to the standard primary sequence features, improves the prediction accuracy of conotoxins against non-toxic and other toxic peptides across varied datasets and several different commonly used machine learning classifiers. Here, we present the effects of these features on conotoxin class and molecular target predictions, in particular, predicting conotoxins that bind to nicotinic acetylcholine receptors (nAChRs). We also demonstrate the use of the Synthetic Minority Oversampling Technique (SMOTE)-Tomek in balancing the datasets while simultaneously making the different classes more distinct by reducing the number of ambiguous samples which nearly overlap between the classes. In predicting the alpha, mu, and omega conotoxin classes, the SMOTE-Tomek PCA PLR model, using the combination of the SS and P feature sets establishes the best performance with an overall accuracy (OA) of 95.95%, with an average accuracy (AA) of 93.04%, and an f1 score of 0.959. Using this model, we obtained sensitivities of 98.98%, 89.66%, and 90.48% when predicting alpha, mu, and omega conotoxin classes, respectively. Similarly, in predicting conotoxins that bind to nAChRs, the SMOTE-Tomek PCA SVM model, which used the collisional cross sections (CCSs) and the P feature sets, demonstrated the highest performance with 91.3% OA, 91.32% AA, and an f1 score of 0.9131. The sensitivity when predicting conotoxins that bind to nAChRs is 91.46% with a 91.18% sensitivity when predicting conotoxins that do not bind to nAChRs.

59 BASIC BIOLOGICAL SCIENCES↗

Rolling Bearing Life Prediction, Theory, and Application

A tutorial is presented outlining the evolution, theory, and application of rolling-element bearing life prediction from that of A. Palmgren, 1924; W. Weibull, 1939; G. Lundberg and A. Palmgren, 1947 and 1952; E. Ioannides and T. Harris, 1985; and E. Zaretsky, 1987. Comparisons are made between these life models. The Ioannides-Harris model without a fatigue limit is identical to the Lundberg-Palmgren model. The Weibull model is similar to that of Zaretsky if the exponents are chosen to be identical. Both the load-life and Hertz stress-life relations of Weibull, Lundberg and Palmgren, and Ioannides and Harris reflect a strong dependence on the Weibull slope. The Zaretsky model decouples the dependence of the critical shear stress-life relation from the Weibull slope. This results in a nominal variation of the Hertz stress-life exponent. For 9th- and 8th-power Hertz stress-life exponents for ball and roller bearings, respectively, the Lundberg-Palmgren model best predicts life. However, for 12th- and 10th-power relations reflected by modern bearing steels, the Zaretsky model based on the Weibull equation is superior. Under the range of stresses examined, the use of a fatigue limit would suggest that (for most operating conditions under which a rolling-element bearing will operate) the bearing will not fail from classical rolling-element fatigue. Realistically, this is not the case. The use of a fatigue limit will significantly overpredict life over a range of normal operating Hertz stresses. (The use of ISO 281:2007 with a fatigue limit in these calculations would result in a bearing life approaching infinity.) Since the predicted lives of rolling-element bearings are high, the problem can become one of undersizing a bearing for a particular application. Rules had been developed to distinguish and compare predicted lives with those actually obtained. Based upon field and test results of 51 ball and roller bearing sets, 98 percent of these bearing sets had acceptable life results using the Lundberg- Palmgren equations with life adjustment factors to predict bearing life. That is, they had lives equal to or greater than that predicted. The Lundberg-Palmgren model was used to predict the life of a commercial turboprop gearbox. The life prediction was compared with the field lives of 64 gearboxes. From these results, the roller bearing lives exhibited a load-life exponent of 5.2, which correlated with the Zaretsky model. The use of the ANSI/ABMA and ISO standards load-life exponent of 10/3 to predict roller bearing life is not reflective of modern roller bearings and will underpredict bearing lives.

Life prediction↗