Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “predictive”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

DRA/NASA/ONERA Collaboration on Icing Research: Prediction of Airfoil Ice Accretion - Part 2

This report presents results from a joint study by DRA, NASA, and ONERA for the purpose of comparing, improving, and validating the aircraft icing computer codes developed by each agency. These codes are of three kinds: (1) water droplet trajectory prediction, (2) ice accretion modeling, and (3) transient electrothermal deicer analysis. In this joint study, the agencies compared their code predictions with each other and with experimental results. These comparison exercises were published in three technical reports, each with joint authorship. DRA published and had first authorship of Part 1 - Droplet Trajectory Calculations, NASA of Part 2 - Ice Accretion Prediction, and ONERA of Part 3 - Electrothermal Deicer Analysis. The results cover work done during the period from August 1986 to late 1991. As a result, all of the information in this report is dated. Where necessary, current information is provided to show the direction of current research. In this present report on ice accretion, each agency predicted ice shapes on two dimensional airfoils under icing conditions for which experimental ice shapes were available. In general, all three codes did a reasonable job of predicting the measured ice shapes. For any given experimental condition, one of the three codes predicted the general ice features (i.e., shape, impingement limits, mass of ice) somewhat better than did the other two. However, no single code consistently did better than the other two over the full range of conditions examined, which included rime, mixed, and glaze ice conditions. In several of the cases, DRA showed that the user's knowledge of icing can significantly improve the accuracy of the code prediction. Rime ice predictions were reasonably accurate and consistent among the codes, because droplets freeze on impact and the freezing model is simple. Glaze ice predictions were less accurate and less consistent among the codes, because the freezing model is more complex and is critically dependent upon unsubstantiated heat transfer and surface roughness models. Thus, heat transfer prediction methods used in the codes became the subject for a separate study in this report to compare predicted heat transfer coefficients with a limited experimental database of heat transfer coefficients for cylinders with simulated glaze and rime ice shapes. The codes did a good job of predicting heat transfer coefficients near the stagnation region of the ice shapes. But in the region of the ice horns, all three codes predicted heat transfer coefficients considerably higher than the measured values. An important conclusion of this study is that further research is needed to understand the finer detail of of the glaze ice accretion process and to develop improved glaze ice accretion models.

Wright, William B.↗

Assessment of Current Jet Noise Prediction Capabilities

An assessment was made of the capability of jet noise prediction codes over a broad range of jet flows, with the objective of quantifying current capabilities and identifying areas requiring future research investment. Three separate codes in NASA s possession, representative of two classes of jet noise prediction codes, were evaluated, one empirical and two statistical. The empirical code is the Stone Jet Noise Module (ST2JET) contained within the ANOPP aircraft noise prediction code. It is well documented, and represents the state of the art in semi-empirical acoustic prediction codes where virtual sources are attributed to various aspects of noise generation in each jet. These sources, in combination, predict the spectral directivity of a jet plume. A total of 258 jet noise cases were examined on the ST2JET code, each run requiring only fractions of a second to complete. Two statistical jet noise prediction codes were also evaluated, JeNo v1, and Jet3D. Fewer cases were run for the statistical prediction methods because they require substantially more resources, typically a Reynolds-Averaged Navier-Stokes solution of the jet, volume integration of the source statistical models over the entire plume, and a numerical solution of the governing propagation equation within the jet. In the evaluation process, substantial justification of experimental datasets used in the evaluations was made. In the end, none of the current codes can predict jet noise within experimental uncertainty. The empirical code came within 2dB on a 1/3 octave spectral basis for a wide range of flows. The statistical code Jet3D was within experimental uncertainty at broadside angles for hot supersonic jets, but errors in peak frequency and amplitude put it out of experimental uncertainty at cooler, lower speed conditions. Jet3D did not predict changes in directivity in the downstream angles. The statistical code JeNo,v1 was within experimental uncertainty predicting noise from cold subsonic jets at all angles, but did not predict changes with heating of the jet and did not account for directivity changes at supersonic conditions. Shortcomings addressed here give direction for future work relevant to the statistical-based prediction methods. A full report will be released as a chapter in a NASA publication assessing the state of the art in aircraft noise prediction.

Hunter, Craid A.↗

A multiscale recurrent neural network model for predicting energy production from geothermal reservoirs

Optimization of energy production from geothermal reservoirs requires reliable prediction of energy production performance under alternative operation and development scenarios. Traditionally, reservoir simulation models are used for the evaluation and screening of alternative production and development plans. However, simulation models require extensive data collection and modeling efforts and are time-consuming to build, run, and update. Data-driven predictive models, on the other hand, can serve as efficient prediction tools that can be used for decision support and management of daily operations and surveillance activities. Data-driven models become particularly attractive when a reservoir simulation model for a field does not exist and/or is difficult to build. Machine learning (ML)-based data-driven models that have recently become popular in several fields exploit statistical patterns and relations in training data to generate predictions. As such, they tend to perform better in interpolation problems (that is, prediction within the training data range) than when they are used to extrapolate beyond the training data. Production data from geothermal reservoirs tend to exhibit short-term variabilities as well as long-term trends, such as monotonically declining production temperatures. Capturing both short-term features and long-term trends with ML-based models is not trivial. We evaluate the use of recurrent neural networks (RNN) for the prediction of energy production from geothermal reservoirs. RNN is a class of ML architectures that are used to represent and predict sequential/dynamic data. Thus, it can be challenging to apply RNN to problems where long-term trends must be captured and extrapolation beyond the training data range is needed. We introduce the multiscale RNN architecture to extend the application of RNN to detect and predict both short-term variabilities and long-term trends in geothermal data. The developed architecture consists of a long-term component to only capture low-frequency data patterns, and a short-term component to detect features with higher frequency and more nonlinearity. The final prediction is obtained by combining the long-term and short-term predictions. Both synthetic and field data are used to evaluate the presented multiscale RNN model. The prediction performance of the multiscale RNN is compared against those obtained from the regular RNN and the autoregressive (AR) model. The results suggest that the multiscale architecture improves the long-term prediction performance of the regular RNN and enhances its robustness against noise.

15 GEOTHERMAL ENERGY↗

Assessment of Arctic and Antarctic Sea Ice Predictability in CMIP5 Decadal Hindcasts

This paper examines the ability of coupled global climate models to predict decadal variability of Arctic and Antarctic sea ice. We analyze decadal hindcasts/predictions of 11 Coupled Model Intercomparison Project Phase 5 (CMIP5) models. Decadal hindcasts exhibit a large multimodel spread in the simulated sea ice extent, with some models deviating significantly from the observations as the predicted ice extent quickly drifts away from the initial constraint. The anomaly correlation analysis between the decadal hindcast and observed sea ice suggests that in the Arctic, for most models, the areas showing significant predictive skill become broader associated with increasing lead times. This area expansion is largely because nearly all the models are capable of predicting the observed decreasing Arctic sea ice cover. Sea ice extent in the North Pacific has better predictive skill than that in the North Atlantic (particularly at a lead time of 3-7 years), but there is a reemerging predictive skill in the North Atlantic at a lead time of 6-8 years. In contrast to the Arctic, Antarctic sea ice decadal hindcasts do not show broad predictive skill at any timescales, and there is no obvious improvement linking the areal extent of significant predictive skill to lead time increase. This might be because nearly all the models predict a retreating Antarctic sea ice cover, opposite to the observations. For the Arctic, the predictive skill of the multi-model ensemble mean outperforms most models and the persistence prediction at longer timescales, which is not the case for the Antarctic. Overall, for the Arctic, initialized decadal hindcasts show improved predictive skill compared to uninitialized simulations, although this improvement is not present in the Antarctic.

sea ice↗

EPIsembleVis: A geo-visual analysis and comparison of the prediction ensembles of multiple COVID-19 models

In this work, we present EPIsembleVis, a web-based comparative visual analysis tool for evaluating the consistency of multiple COVID-19 prediction models. Our approach analyzes a collection of COVID-19 predictions from different epidemiological models as an ensemble and utilizes two metrics to quantify model performance. These metrics include (a) prediction uncertainty (represented as the dispersion of predictions in each ensemble) and (b) prediction error (calculated by comparing individual model predictions with the recorded data). Through an interactive visual interface, our approach provides a data-driven workflow for (a) selecting and constructing the COVID-19 model prediction ensemble based on the spatiotemporal overlap of available predictions of multiple epidemiological models, (b) quantifying the model performance using both the uncertainty of each model prediction ensemble, and the error of each ensemble member that represents individual model predictions, and (c) visualizing the spatiotemporal variability in the projection performance of individual models using a suite of novel ensemble visualization techniques, such as the data availability map, a spatiotemporal textured-tile calendar, multivariate rose chart, and time-series leaflet glyph. We demonstrate the capability of our ensemble visual interface through a case study that investigates the performance of weekly COVID-19 predictions, which are provided through the COVID-19 Forecast Hub UMass-Amherst Influenza Forecasting Center of Excellence [47] for the United States and United States Territories. The EPIsembleVis tool is implemented using open-source web technologies and adaptive system design, rendering it interoperable with Elasticsearch and Kibana for automatically ingesting COVID-19 predictions from online repositories, and it is generalizable for analyzing worldwide projections from more epidemiological models.

60 APPLIED LIFE SCIENCES↗

Denoising Autoencoder for Reconstructing Sensor Observation Data and Predicting Evapotranspiration: Noisy and Missing Values Repair and Uncertainty Quantification

Abstract Machine learning (ML) methods applied in scientific research often deal with interrelated features in high‐dimensional data. Reducing data noise and redundancy is needed to increase prediction accuracy and efficiency especially when dealing with data from field sensors. We explored an unsupervised learning method, the denoising autoencoder (DAE), to extract the underlying data structure from noisy raw data in the context of predicting hydrologic quantities from multiple field sensors. These sensors have intrinsic instrumental noise and occasional malfunctions that cause missing values. Our DAE neural network reconstructed meteorological sensor data containing noise and missing values to predict evapotranspiration in a mountainous watershed. The DAE reconstructed the sensor variables with a mean coefficient of determination value of 0.77 across 15 dimensions representing individual sensors. It reduced variance and bias uncertainties compared to a classical autoencoder model. The reconstruction quality varied across dimensions depending on their cross‐correlation and alignment with the underlying data structure. Uncertainties arising from the model structure were overall higher than those resulting from data corruption. We attached the DAE structure to a downstream ET‐prediction neural network in three formats and achieved reasonably accurate ET predictions . The use of the DAE notably reduced variance uncertainty in ET prediction. However, excessive variance reduction may be accompanied by an increase in bias due to the intrinsic bias‐variance tradeoff. Our method of evaluating and reducing uncertainties in aggregated data from different sources can be used to improve predictive models, process understanding, and uncertainty quantification for better water resource management. Plain Language Summary We present a machine learning method, namely the denoising autoencoder, which reduces the effects of data noise and missing values typically present in scientific data sets collected through sensor measurements. This method selects the most relevant information from noisy raw data collected by the instruments and fills in missing values. To demonstrate the effectiveness of our method, we applied it to predict evapotranspiration, a hydrologic variable that represents the water moved from the land surface to the atmosphere through a combination of evaporation and plant water use (transpiration). We also used a random sampling technique (the Monte Carlo method) to compare the uncertainty in the predictions when using the raw and noisy data versus the reconstructed data. The denoising process produced more accurate predictions of evapotranspiration with less uncertainty. Improved predictions of evapotranspiration can lead to a better understanding and accounting of water budgets. This ML approach is broadly suitable for a wide variety of applications that involve noisy sensor data with missing values. Key Points We used a denoising autoencoder (DAE) neural network to reduce noise in meteorological and soil sensor observations by on average We used Monte Carlo sampling to estimate the bias and variance of all model outputs, including uncertainty sources from data and the model We attached the DAE component to a downstream neural network to predict ET with the variance reduced by , compared to that without the DAE

denoising autoencoder↗

Subseasonal Forecasting and MJO Teleconnections in Machine Learning Weather Prediction Models

Abstract In recent years, machine‐learning (ML) models trained on reanalysis data have rivaled physics‐based forecast models in terms of performance skill for global weather forecasting. With increased rollout stability, the question of how these models perform for subseasonal to seasonal (S2S, week 3–8) forecasting has emerged. In this study we run a large set of subseasonal hindcasts over 2004–2023 to evaluate two ML weather forecast models at the S2S time scale, SFNO‐HENS (Nvidia, fully ML) and NeuralGCM (Google Research, hybrid). Corresponding hindcasts from the European Centre for Medium‐Range Weather Forecasts (ECMWF) are used as a baseline for comparison to a physics‐based model. Because our focus is on predicting moisture transport over the Western United States between October and March, we evaluate the models' prediction skill for the Madden‐Julian Oscillation (MJO) and its associated teleconnections in the North Pacific. We find that both ML models are competitive with the ECWMF model, with comparable skill in predicting the North Pacific large‐scale circulation and the MJO at week 3 and beyond. Even though overall the mid‐latitude subseasonal prediction skill remains low, the ML models exhibit interesting behavior such as a realistic propagation of the MJO across the Maritime Continent and realistic teleconnections. A SFNO‐HENS sensitivity experiment with altered initial conditions in the tropics demonstrates the stability of the model, and it illustrates the capability of ML models to represent important physical processes of the atmosphere at the S2S time scale. Plain Language Summary Predicting weather patterns and precipitation a few weeks in advance (subseasonal time scale) is of great interest for stakeholders such as water managers in the Southwest United States (US), where arid conditions prevail. Subseasonal forecasts from traditional weather forecast models exhibit low skill in the region, limiting their applicability. Here we examine whether the recent breakthrough in weather forecasting made with machine learning/artificial intelligence models can translate to improved subseasonal forecasts. Recently‐developed machine learning models exhibit comparable skill to a state‐of‐the‐art physics‐based model for predicting weather patterns in the North Pacific/North America region, and associated moisture transport. The same applies to their skill in predicting the tropical pattern, the Madden‐Julian Oscillation, and its important remote perturbations over the midlatitude East Pacific and Southwest US. Additionally, a perturbation experiment carried out with one of the machine learning models illustrates their ability to not only predict the evolution of atmospheric fields, but also to learn and represent physical processes such as tropics‐extratropics Rossby wave propagation. Key Points Two machine learning weather forecast models exhibit state‐of‐the‐art prediction skill at the subseasonal time scale in the Pacific sector The models equal ECWMF in terms of Madden‐Julian oscillation (MJO) prediction skill, and they accurately predict the MJO propagation and associated teleconnections The two machine‐learning models represent key physical processes for subseasonal prediction, despite being trained for weather forecasting

Peings, Yannick↗

Benchmarking core turbulence and transport predictions for an inductive compact tokamak reactor plasma

Motivated by the need for accurate, timely, and efficient calculations of plasma transport, predictions of plasma turbulence properties made using different TGLF saturation rules are benchmarked against corresponding predictions from linear and nonlinear gyrokinetic CGYRO simulations. This benchmarking is carried out using parameters taken from an inductive burning plasma scenario in a hypothetical compact high-field (R maj = 4 m, B T = 8 T) tokamak, lying in a much different regime of parameter space than either the TGLF calibration regime or current-day experiments. The core turbulent transport in this scenario is predicted to be dominated by ion temperature gradient (ITG) turbulence. In general, the ITG critical gradients predicted by various TGLF saturation rules are quite close to the CGYRO predictions. Both codes predict similar linear ITG growth rates and frequency spectra, as well as their scaling with R/L T i = −Rd ln(T i )/dr. However, TGLF systematically predicts unstable trapped-electron modes (TEMs) above k y ρ s ≃ 0.5 not seen by CGYRO for the same parameters, due to TGLF predicting a lower threshold in R/L T e than CGYRO for TEM onset. It is shown that for this scenario, nonlinear CGYRO simulations predict stiffer ITG turbulence than the TGLF SAT0 and SAT1 saturation rules, with energy fluxes close in magnitude and scaling with R/L T i to what is predicted by the SAT2 saturation rule. Self-consistent core profiles calculated using nonlinear CGYRO flux predictions and the PORTALS transport solver are shown to agree fairly well with corresponding predictions made using the TGLF SAT2 model, including a similar level of density peaking.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Mechanical loading prediction through accelerometry data during walking and running

ABSTRACT Currently, there is no way to assess mechanical loading variables such as peak ground reaction forces (pGRF) and peak loading rate (pLR) in clinical settings. The purpose of this study was to develop accelerometry‐based equations to predict both pGRF and pLR during walking and running. One hundred and thirty one subjects (79 females; 76.9 ± 19.6 kg) walked and ran at different speeds (2–14 km·h −1 ) on a force plate–instrumented treadmill while wearing accelerometers at their ankle, lower back and hip. Regression equations were developed to predict pGRF and pLR from accelerometry data. Leave‐one‐out cross‐validation was used to calculate prediction accuracy and Bland–Altman plots. Our pGRF prediction equation was compared with a reference equation previously published. Body mass and peak acceleration were included for pGRF prediction and body mass and peak acceleration rate for pLR prediction. All pGRF equation coefficients of determination were above 0.96, and a good agreement between actual and predicted pGRF was observed, with a mean absolute percent error (MAPE) below 7.3%. Accuracy indices from our equations were better than previously developed equations. All pLR prediction equations presented a lower accuracy compared to those developed to predict pGRF. Walking and running pGRF can be predicted with high accuracy by accelerometry‐based equations, representing an easy way to determine mechanical loading in free‐living conditions. The pLR prediction equations yielded a somewhat lower prediction accuracy compared with the pGRF equations.

Veras, Lucas↗

Predicting Volume of Distribution in Humans: Performance of In Silico Methods for a Large Set of Structurally Diverse Clinical Compounds

Volume of distribution at steady state (V D,ss ) is one of the key pharmacokinetic parameters estimated during the drug discovery process. Despite considerable efforts to predict V D,ss , accuracy and choice of prediction methods remain a challenge, with evaluations constrained to a small set (<150) of compounds. To address these issues, a series of in silico methods for predicting human V D,ss directly from structure were evaluated using a large set of clinical compounds. Machine learning (ML) models were built to predict V D,ss directly and to predict input parameters required for mechanistic and empirical V D,ss predictions. In addition, log D, fraction unbound in plasma (fup), and blood-to-plasma partition ratio (BPR) were measured on 254 compounds to estimate the impact of measured data on predictive performance of mechanistic models. Furthermore, the impact of novel methodologies such as measuring partition (Kp) in adipocytes and myocytes (n = 189) on V D,ss predictions was also investigated. In predicting V D,ss directly from chemical structures, both mechanistic and empirical scaling using a combination of predicted rat and dog V D,ss demonstrated comparable performance (62%–71% within 3-fold). The direct ML model outperformed other in silico methods (75% within 3-fold, r 2 = 0.5, AAFE = 2.2) when built from a larger data set. Scaling to human from predicted V D,ss of either rat or dog yielded poor results (<47% within 3-fold). Measured fup and BPR improved performance of mechanistic V D,ss predictions significantly (81% within 3-fold, r 2 = 0.6, AAFE = 2.0). Adipocyte intracellular Kp showed good correlation to the V D,ss but was limited in estimating the compounds with low V D,ss .

59 BASIC BIOLOGICAL SCIENCES↗

Validation of finite element and boundary element methods for predicting structural vibration and radiated noise

Analytical and experimental validation of methods to predict structural vibration and radiated noise are presented. A rectangular box excited by a mechanical shaker was used as a vibrating structure. Combined finite element method (FEM) and boundary element method (BEM) models of the apparatus were used to predict the noise radiated from the box. The FEM was used to predict the vibration, and the surface vibration was used as input to the BEM to predict the sound intensity and sound power. Vibration predicted by the FEM model was validated by experimental modal analysis. Noise predicted by the BEM was validated by sound intensity measurements. Three types of results are presented for the total radiated sound power: (1) sound power predicted by the BEM modeling using vibration data measured on the surface of the box; (2) sound power predicted by the FEM/BEM model; and (3) sound power measured by a sound intensity scan. The sound power predicted from the BEM model using measured vibration data yields an excellent prediction of radiated noise. The sound power predicted by the combined FEM/BEM model also gives a good prediction of radiated noise except for a shift of the natural frequencies that are due to limitations in the FEM model.

Seybert, A. F.↗

Accuracy Assessment of Two Gps Fidelity Prediction Services in Urban Terrain

Low altitude flight in urban areas is susceptible to degraded GNSS-based navigation system performance due to terrain interference with radio signals from orbital positioning satellites. Predictive navigation performance fidelity tools are needed a) in preflight planning to assist in the creation of safe flight paths and b) in-flight to provide contingency management agents with the navigation risk of proximal flight corridors. Two navigation fidelity prediction services are validated by comparison with over 6000 readings from GNSS sensors collected along a five-mile path through urban areas of Corpus Christi, Texas, on three dates in 2022. Predictions are based on satellite line of sight through 3D terrain data collected in 2018. Each service predicts a set of navigation fidelity metrics over a user-specified time period. One metric estimated by both is the number of visible satellites. A direct comparison of the number of predicted visible satellites with the number sensed by the receiver is used to validate the prediction services. Results show an exact match in the number of predicted satellites for 60% of the measurements, and a match within +/- 4 satellites for 95% of the measurements. As expected, agreement improves away from vertical blocking terrain. Most cases of mismatch are due a lower predicted count than measured (false negatives), and can be accounted for by receiver pickup of stray signals caused by multipath propagation. About 10% of mismatches are false positives and are mostly accounted for by foliage effects. The two services predict visibility of the same set of satellites 80% of the time, differ by two or less satellites 95% of the time, and can compute predictions for one hour of observations in one minute or less. Validation is analyzed statistically and in detailed case studies of selected observation times. The prediction services validated in this study run fast enough for preflight safety planning. The more stringent challenge of inflight navigation fidelity prediction for contingency management requires both a speedup of the current level of modeling and equally fast stray signal modeling.

Andrew Moore↗

Accuracy Assessment of Two GPS Fidelity Prediction Services in Urban Terrain

Low altitude flight in urban areas is susceptible to degraded GNSS-based navigation system performance due to terrain interference with radio signals from orbital positioning satellites. Predictive navigation performance fidelity tools are needed a) in preflight planning to assist in the creation of safe flight paths and b) in-flight to provide contingency management agents with the navigation risk of proximal flight corridors. Two navigation fidelity prediction services are validated by comparison with over 6000 readings from GNSS sensors collected along a five-mile path through urban areas of Corpus Christi, Texas, on three dates in 2022. Predictions are based on satellite line of sight through 3D terrain data collected in 2018. Each service predicts a set of navigation fidelity metrics over a user-specified time period. One metric estimated by both is the number of visible satellites. A direct comparison of the number of predicted visible satellites with the number sensed by the receiver is used to validate the prediction services. Results show an exact match in the number of predicted satellites for 60% of the measurements, and a match within +/- 4 satellites for 95% of the measurements. As expected, agreement improves away from vertical blocking terrain. Most cases of mismatch are due a lower predicted count than measured (false negatives), and can be accounted for by receiver pickup of stray signals caused by multipath propagation. About 10% of mismatches are false positives and are mostly accounted for by foliage effects. The two services predict visibility of the same set of satellites 80% of the time, differ by two or less satellites 95% of the time, and can compute predictions for one hour of observations in one minute or less. Validation is analyzed statistically and in detailed case studies of selected observation times. The prediction services validated in this study run fast enough for preflight safety planning. The more stringent challenge of inflight navigation fidelity prediction for contingency management requires both a speedup of the current level of modeling and equally fast stray signal modeling.

Andrew J. Moore↗

Celebrating 10 Years of the Sub-Seasonal to Seasonal Prediction Project and Looking to the Future

The conference clearly demonstrated the increasing interest and growth of the scientific community working on the development and application of sub-seasonal to seasonal prediction since the start of the World Weather Research Programme (WWRP)/World Climate Research Programme (WCRP) sub-seasonal to seasonal (S2S) prediction project in 2013. The conference, which was held at the University of Reading (United Kingdom), was organized into three main themes as briefly summarized below, with eleven invited talks, 74 oral contributed talks, and 101 posters. The conference also included a two-hour breakout session, wherein eight groups discussed the current state and prospect for S2S prediction, and an early career researcher event. A summary of these discussions and recommendations is presented below. The conference web page (https://research.reading.ac.uk/s2s-summit2023/) is archived at the University of Reading. Introductory comments by representatives of the World Meteorological Organization (WMO) WWRP and WCRP emphasized the importance of the weather–climate linkage, targeted by S2S forecasts (from 2 weeks to a season ahead), addressing the challenges of creating “end-to-end” forecasts that encompass the entire climate-services chain from the prediction science and forecast, to the development and issuing of forecast products tailored to informing user-decisions. They also emphasized the efficacy of multi-model ensemble efforts and databases to foster collaborations internationally and between operational centres and academia. Although the WWRP/WCRP S2S project comes to an end in 2023, S2S prediction will remain an important focus for WWRP and WCRP. In WWRP, a new project called SAGE (Sub-seasonal to seasonal predictions for Agriculture and Environment) will start in 2024. Another important legacy of the S2S project will be the maintenance of the S2S database (Vitart et al. 2017) and the establishment of a WMO Lead Center for sub-seasonal prediction multi-model ensemble (LC-SSPMME) which will provide real-time multi-model S2S climate information. In two keynote presentations, Prof. Brian Hoskins (University of Reading) and Dr. Gilbert Brunet (Australian Bureau of Meteorology) discussed the potential of S2S predictability and the ongoing journey for understanding and improving these predictions. This conference was a sequel to the International Conference on Sub-seasonal to Seasonal Prediction (Robertson et al., 2014) which took place in College Park (Maryland, USA) in February 2014 to celebrate the start of the WWRP/WCRP S2S project, and to WCRP and WWRP conferences in Boulder, USA, in 2018 (Merryfield et al., 2020). A significant development compared to the previous S2S conferences was the large number of presentations on research to operation (R2O) and S2S applications and on the use of artificial intelligence and machine learning (AI/ML) methods for S2S prediction. Some of these methods provide empirical S2S forecasts which are competitive with state-of-the-art dynamical models. Other presentations demonstrated that AI/ML can provide alternative calibration of dynamical model outputs to traditional methods. Several talks and posters highlighted the increasing use of AI/ML, including deep learning, in S2S forecast post-processing and using AI to identify higher flow-dependent skill. Finally, some presentations demonstrated the value of AI/ML methods for a better understanding of S2S sources of predictability and attribution of extreme events.

S. J. Woolnough↗

CyProduct: A software tool for accurately predicting the byproducts of human cytochrome P450 metabolism

In silico metabolism prediction is a cheminformatic task of autonomously predicting the set of metabolic byproducts produced from a specified molecule and a set of enzymes or reactions. Here we describe a novel machine-learned in silico cytochrome P450 (CYP450) metabolism prediction suite, called CyProduct, that accurately predicts metabolic byproducts for a specified molecule and a human CYP450 isoform. It includes three modules: (1) CypReact, a tool that predicts if the query compound reacts with a given CYP450 enzyme; (2) CypBoM, a tool that accurately predicts the “bond site” of the reaction (i.e., which specific bonds within the query molecule react with the CYP isoform); and (3) MetaboGen, a tool that generates the metabolic byproducts based on CypBoM’s bond-site prediction. CyProduct predicts metabolic biotransformation products for each of the nine most important human CYP450 enzymes. CypBoM uses an important new concept called “Bond of Metabolism” (BoM), which extends the traditional “Site of Metabolism" (SoM) by specifying the information about the set of chemical bonds that is modified or formed in a metabolic reaction (rather than the specific atom). We created a BoM database for 3487 CYP450-mediated Phase I reactions, then used this to train the CypBoM Predictor to predict the reactive bond locations on substrate molecules. CypBoM Predictor’s cross-validated Jaccard score for reactive bond prediction ranged from 0.380 to 0.452 over the nine CYP450 enzymes. Over variants of a test set of 72 known CYP450 substrates and 30 non-reactants, CyProduct outperformed the other packages -- including ADMET Predictor, BioTransformer and GLORY -- by an average of 200% (wrt Jaccard score) in terms of predicting metabolites. The CyProduct suite and the datasets are freely available at https://bitbucket.org/wishartlab/cyproduct/src/master/.

Machine learning, Cytochrome P450, Metabolism pred↗

Ensemble transfer learning for the prediction of anti-cancer drug response

Abstract Transfer learning, which transfers patterns learned on a source dataset to a related target dataset for constructing prediction models, has been shown effective in many applications. In this paper, we investigate whether transfer learning can be used to improve the performance of anti-cancer drug response prediction models. Previous transfer learning studies for drug response prediction focused on building models to predict the response of tumor cells to a specific drug treatment. We target the more challenging task of building general prediction models that can make predictions for both new tumor cells and new drugs. Uniquely, we investigate the power of transfer learning for three drug response prediction applications including drug repurposing, precision oncology, and new drug development, through different data partition schemes in cross-validation. We extend the classic transfer learning framework through ensemble and demonstrate its general utility with three representative prediction algorithms including a gradient boosting model and two deep neural networks. The ensemble transfer learning framework is tested on benchmark in vitro drug screening datasets. The results demonstrate that our framework broadly improves the prediction performance in all three drug response prediction applications with all three prediction algorithms.

60 APPLIED LIFE SCIENCES↗

Tandem Predictions for HPC Jobs

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

HPC↗

Tandem Predictions for HPC Jobs: Preprint

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

97 MATHEMATICS AND COMPUTING↗