Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “predictive”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Subseasonal Forecasting and MJO Teleconnections in Machine Learning Weather Prediction Models

Abstract In recent years, machine‐learning (ML) models trained on reanalysis data have rivaled physics‐based forecast models in terms of performance skill for global weather forecasting. With increased rollout stability, the question of how these models perform for subseasonal to seasonal (S2S, week 3–8) forecasting has emerged. In this study we run a large set of subseasonal hindcasts over 2004–2023 to evaluate two ML weather forecast models at the S2S time scale, SFNO‐HENS (Nvidia, fully ML) and NeuralGCM (Google Research, hybrid). Corresponding hindcasts from the European Centre for Medium‐Range Weather Forecasts (ECMWF) are used as a baseline for comparison to a physics‐based model. Because our focus is on predicting moisture transport over the Western United States between October and March, we evaluate the models' prediction skill for the Madden‐Julian Oscillation (MJO) and its associated teleconnections in the North Pacific. We find that both ML models are competitive with the ECWMF model, with comparable skill in predicting the North Pacific large‐scale circulation and the MJO at week 3 and beyond. Even though overall the mid‐latitude subseasonal prediction skill remains low, the ML models exhibit interesting behavior such as a realistic propagation of the MJO across the Maritime Continent and realistic teleconnections. A SFNO‐HENS sensitivity experiment with altered initial conditions in the tropics demonstrates the stability of the model, and it illustrates the capability of ML models to represent important physical processes of the atmosphere at the S2S time scale. Plain Language Summary Predicting weather patterns and precipitation a few weeks in advance (subseasonal time scale) is of great interest for stakeholders such as water managers in the Southwest United States (US), where arid conditions prevail. Subseasonal forecasts from traditional weather forecast models exhibit low skill in the region, limiting their applicability. Here we examine whether the recent breakthrough in weather forecasting made with machine learning/artificial intelligence models can translate to improved subseasonal forecasts. Recently‐developed machine learning models exhibit comparable skill to a state‐of‐the‐art physics‐based model for predicting weather patterns in the North Pacific/North America region, and associated moisture transport. The same applies to their skill in predicting the tropical pattern, the Madden‐Julian Oscillation, and its important remote perturbations over the midlatitude East Pacific and Southwest US. Additionally, a perturbation experiment carried out with one of the machine learning models illustrates their ability to not only predict the evolution of atmospheric fields, but also to learn and represent physical processes such as tropics‐extratropics Rossby wave propagation. Key Points Two machine learning weather forecast models exhibit state‐of‐the‐art prediction skill at the subseasonal time scale in the Pacific sector The models equal ECWMF in terms of Madden‐Julian oscillation (MJO) prediction skill, and they accurately predict the MJO propagation and associated teleconnections The two machine‐learning models represent key physical processes for subseasonal prediction, despite being trained for weather forecasting

Peings, Yannick↗

Benchmarking core turbulence and transport predictions for an inductive compact tokamak reactor plasma

Motivated by the need for accurate, timely, and efficient calculations of plasma transport, predictions of plasma turbulence properties made using different TGLF saturation rules are benchmarked against corresponding predictions from linear and nonlinear gyrokinetic CGYRO simulations. This benchmarking is carried out using parameters taken from an inductive burning plasma scenario in a hypothetical compact high-field (R maj = 4 m, B T = 8 T) tokamak, lying in a much different regime of parameter space than either the TGLF calibration regime or current-day experiments. The core turbulent transport in this scenario is predicted to be dominated by ion temperature gradient (ITG) turbulence. In general, the ITG critical gradients predicted by various TGLF saturation rules are quite close to the CGYRO predictions. Both codes predict similar linear ITG growth rates and frequency spectra, as well as their scaling with R/L T i = −Rd ln(T i )/dr. However, TGLF systematically predicts unstable trapped-electron modes (TEMs) above k y ρ s ≃ 0.5 not seen by CGYRO for the same parameters, due to TGLF predicting a lower threshold in R/L T e than CGYRO for TEM onset. It is shown that for this scenario, nonlinear CGYRO simulations predict stiffer ITG turbulence than the TGLF SAT0 and SAT1 saturation rules, with energy fluxes close in magnitude and scaling with R/L T i to what is predicted by the SAT2 saturation rule. Self-consistent core profiles calculated using nonlinear CGYRO flux predictions and the PORTALS transport solver are shown to agree fairly well with corresponding predictions made using the TGLF SAT2 model, including a similar level of density peaking.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Mechanical loading prediction through accelerometry data during walking and running

ABSTRACT Currently, there is no way to assess mechanical loading variables such as peak ground reaction forces (pGRF) and peak loading rate (pLR) in clinical settings. The purpose of this study was to develop accelerometry‐based equations to predict both pGRF and pLR during walking and running. One hundred and thirty one subjects (79 females; 76.9 ± 19.6 kg) walked and ran at different speeds (2–14 km·h −1 ) on a force plate–instrumented treadmill while wearing accelerometers at their ankle, lower back and hip. Regression equations were developed to predict pGRF and pLR from accelerometry data. Leave‐one‐out cross‐validation was used to calculate prediction accuracy and Bland–Altman plots. Our pGRF prediction equation was compared with a reference equation previously published. Body mass and peak acceleration were included for pGRF prediction and body mass and peak acceleration rate for pLR prediction. All pGRF equation coefficients of determination were above 0.96, and a good agreement between actual and predicted pGRF was observed, with a mean absolute percent error (MAPE) below 7.3%. Accuracy indices from our equations were better than previously developed equations. All pLR prediction equations presented a lower accuracy compared to those developed to predict pGRF. Walking and running pGRF can be predicted with high accuracy by accelerometry‐based equations, representing an easy way to determine mechanical loading in free‐living conditions. The pLR prediction equations yielded a somewhat lower prediction accuracy compared with the pGRF equations.

Veras, Lucas↗

Predicting Volume of Distribution in Humans: Performance of In Silico Methods for a Large Set of Structurally Diverse Clinical Compounds

Volume of distribution at steady state (V D,ss ) is one of the key pharmacokinetic parameters estimated during the drug discovery process. Despite considerable efforts to predict V D,ss , accuracy and choice of prediction methods remain a challenge, with evaluations constrained to a small set (<150) of compounds. To address these issues, a series of in silico methods for predicting human V D,ss directly from structure were evaluated using a large set of clinical compounds. Machine learning (ML) models were built to predict V D,ss directly and to predict input parameters required for mechanistic and empirical V D,ss predictions. In addition, log D, fraction unbound in plasma (fup), and blood-to-plasma partition ratio (BPR) were measured on 254 compounds to estimate the impact of measured data on predictive performance of mechanistic models. Furthermore, the impact of novel methodologies such as measuring partition (Kp) in adipocytes and myocytes (n = 189) on V D,ss predictions was also investigated. In predicting V D,ss directly from chemical structures, both mechanistic and empirical scaling using a combination of predicted rat and dog V D,ss demonstrated comparable performance (62%–71% within 3-fold). The direct ML model outperformed other in silico methods (75% within 3-fold, r 2 = 0.5, AAFE = 2.2) when built from a larger data set. Scaling to human from predicted V D,ss of either rat or dog yielded poor results (<47% within 3-fold). Measured fup and BPR improved performance of mechanistic V D,ss predictions significantly (81% within 3-fold, r 2 = 0.6, AAFE = 2.0). Adipocyte intracellular Kp showed good correlation to the V D,ss but was limited in estimating the compounds with low V D,ss .

59 BASIC BIOLOGICAL SCIENCES↗

CyProduct: A software tool for accurately predicting the byproducts of human cytochrome P450 metabolism

In silico metabolism prediction is a cheminformatic task of autonomously predicting the set of metabolic byproducts produced from a specified molecule and a set of enzymes or reactions. Here we describe a novel machine-learned in silico cytochrome P450 (CYP450) metabolism prediction suite, called CyProduct, that accurately predicts metabolic byproducts for a specified molecule and a human CYP450 isoform. It includes three modules: (1) CypReact, a tool that predicts if the query compound reacts with a given CYP450 enzyme; (2) CypBoM, a tool that accurately predicts the “bond site” of the reaction (i.e., which specific bonds within the query molecule react with the CYP isoform); and (3) MetaboGen, a tool that generates the metabolic byproducts based on CypBoM’s bond-site prediction. CyProduct predicts metabolic biotransformation products for each of the nine most important human CYP450 enzymes. CypBoM uses an important new concept called “Bond of Metabolism” (BoM), which extends the traditional “Site of Metabolism" (SoM) by specifying the information about the set of chemical bonds that is modified or formed in a metabolic reaction (rather than the specific atom). We created a BoM database for 3487 CYP450-mediated Phase I reactions, then used this to train the CypBoM Predictor to predict the reactive bond locations on substrate molecules. CypBoM Predictor’s cross-validated Jaccard score for reactive bond prediction ranged from 0.380 to 0.452 over the nine CYP450 enzymes. Over variants of a test set of 72 known CYP450 substrates and 30 non-reactants, CyProduct outperformed the other packages -- including ADMET Predictor, BioTransformer and GLORY -- by an average of 200% (wrt Jaccard score) in terms of predicting metabolites. The CyProduct suite and the datasets are freely available at https://bitbucket.org/wishartlab/cyproduct/src/master/.

Machine learning, Cytochrome P450, Metabolism pred↗

Ensemble transfer learning for the prediction of anti-cancer drug response

Abstract Transfer learning, which transfers patterns learned on a source dataset to a related target dataset for constructing prediction models, has been shown effective in many applications. In this paper, we investigate whether transfer learning can be used to improve the performance of anti-cancer drug response prediction models. Previous transfer learning studies for drug response prediction focused on building models to predict the response of tumor cells to a specific drug treatment. We target the more challenging task of building general prediction models that can make predictions for both new tumor cells and new drugs. Uniquely, we investigate the power of transfer learning for three drug response prediction applications including drug repurposing, precision oncology, and new drug development, through different data partition schemes in cross-validation. We extend the classic transfer learning framework through ensemble and demonstrate its general utility with three representative prediction algorithms including a gradient boosting model and two deep neural networks. The ensemble transfer learning framework is tested on benchmark in vitro drug screening datasets. The results demonstrate that our framework broadly improves the prediction performance in all three drug response prediction applications with all three prediction algorithms.

60 APPLIED LIFE SCIENCES↗

Tandem Predictions for HPC Jobs

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

HPC↗

Tandem Predictions for HPC Jobs: Preprint

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

97 MATHEMATICS AND COMPUTING↗

Sensitivity analysis for characterizing the impact of HNGD model on the prediction of hydrogen redistribution in Zircaloy cladding using BISON code

Hydrogen in zirconium cladding is able to precipitate into zirconium hydrides which impacts cladding integrity. The Hydride Nucleation-Growth-Dissolution (HNGD) model in the BISON code accounts for the precipitation and dissolution kinetics of hydride in Zircaloy material. This paper presents global sensitivity analyses of the HNGD model aiming to enhance our understanding of the hydride precipitation phenomena by quantifying the variance that key parameters have on the prediction of hydrogen behavior under various environmental conditions. Here, model predictions are compared to experimental data obtained under two different conditions: 1) with uniformly precharged specimens subjected to a linear thermal gradient, and 2) specimens precharged with a cathodically applied hydride rim at one end of the sample and subjected to an asymmetric thermal gradient. The Sobol sensitivity analysis identifies the key parameters in the HNGD model for both types of specimens. For linear temperature cases, the heat of transport dominates the accuracy of predictions when no precipitation occurs at the cold end, while Terminal Solid Solubility for Dissolution (TSSD) is the most important parameter when precipitation occurs. A large variation in the predicted hydrogen concentration profiles is found in the range of high TSSD due to the occurrence of precipitation. For asymmetric temperature cases, the solubility coefficient gives the largest impact on the predicted hydrogen distribution, as it determines the amount of solute hydrogen dissolved from the initially applied hydride rim. A large discrepancy in hydrogen distribution between simulations and experiments exists with the asymmetric specimens because BISON simulations fail to predict the precipitation of hydride at the cooler end. Comparative studies using former and updated models verifies the significant impact of the hydride growth mechanism on predicted hydrogen concentration profiles. In particular, when hydride initially exists, changes in TSSD generate a large variation in the predicted amount of precipitation by hydride growth, giving large uncertainty in predicting the hydrogen distribution over the sample length. The outputs characterize the significant impact of the hydride growth mechanism in the HNGD model on predicting hydrogen behavior, and improve the understanding of the precipitation of hydride in Zircaloy cladding within a range of expected environmental conditions. The analyses indicate work is still needed to improve the hydride solvus models in the BISON code to accurately predict experimentally observed hydride concentrations and distributions.

36 MATERIALS SCIENCE↗

Long-Term Vehicle Speed Prediction via Historical Traffic Data Analysis for Improved Energy Efficiency of Connected Electric Vehicles

Connected and automated vehicles (CAVs) are expected to provide enhanced safety, mobility, and energy efficiency. While abundant evidence has been accumulated showing substantial energy saving potentials of CAVs through eco-driving, traffic condition prediction has remained to be the main challenge in capitalizing the gains. The coupled power and thermal subsystems of CAVs necessitate the use of different speed preview windows for effective and integrated power and thermal management. Real-time vehicle-to-infrastructure (V2I) communications can provide an accurate speed prediction over a short prediction horizon (e.g., 30 s to 60 s), but not for a long range (e.g., over 180 s). Therefore, advanced approaches are required to develop detailed speed prediction for robust optimization-based energy management of CAVs. This paper presents an integrated speed prediction framework based on historical traffic data classification and real-time V2I communications for efficient energy management of electrified CAVs. The proposed framework provides multi-range speed predictions with different fidelity over short and long horizons. The proposed multi-range speed prediction is integrated with an economic model predictive control (MPC) strategy for the battery thermal management (BTM) of connected and automated electric vehicles (EVs). The simulation results over real-world urban driving cycles confirm the enhanced prediction performance of the proposed data classification strategy over a long prediction horizon. Despite the uncertainty in long-range CAVs’ speed predictions, the vehicle-level simulation results show that 14% and 19% energy savings can be accumulated sequentially through eco-driving and BTM optimization (eco-cooling), respectively, when compared with normal driving (i.e., human driver) and conventional BTM strategy.

Engineering↗

Impact of volcanic eruptions on CMIP6 decadal predictions: a multi-model analysis

Abstract. In recent decades, three major volcanic eruptions of different intensity have occurred (Mount Agung in 1963, El Chichón in 1982 and Mount Pinatubo in 1991), with reported climate impacts on seasonal to decadal timescales that could have been potentially predicted with accurate and timely estimates of the associated stratospheric aerosol loads. The Decadal Climate Prediction Project component C (DCPP-C) includes a protocol to investigate the impact of volcanic aerosols on the climate experienced during the years that followed those eruptions through the use of decadal predictions. The interest of conducting this exercise with climate predictions is that, thanks to the initialisation, they start from the observed climate conditions at the time of the eruptions, which helps to disentangle the climatic changes due to the initial conditions and internal variability from the volcanic forcing. The protocol consists of repeating the retrospective predictions that are initialised just before the last three major volcanic eruptions but without the inclusion of their volcanic forcing, which are then compared with the baseline predictions to disentangle the simulated volcanic effects upon climate. We present the results from six Coupled Model Intercomparison Project Phase 6 (CMIP6) decadal prediction systems. These systems show strong agreement in predicting the well-known post-volcanic radiative effects following the three eruptions, which induce a long-lasting cooling in the ocean. Furthermore, the multi-model multi-eruption composite is consistent with previous work reporting an acceleration of the Northern Hemisphere polar vortex and the development of El Niño conditions the first year after the eruption, followed by a strengthening of the Atlantic Meridional Overturning Circulation the subsequent years. Our analysis reveals that all these dynamical responses are both model- and eruption-dependent. A novel aspect of this study is that we also assess whether the volcanic forcing improves the realism of the predictions. Comparing the predicted surface temperature anomalies in the two sets of hindcasts (with and without volcanic forcing) with observations we show that, overall, including the volcanic forcing results in better predictions. The volcanic forcing is found to be particularly relevant for reproducing the observed sea surface temperature (SST) variability in the North Atlantic Ocean following the 1991 eruption of Pinatubo.

Bilbao, Roberto (ORCID:0000000307294980)↗

Controlling prediction functional blocks used by a branch predictor in a processor

An electronic device includes a processor, a branch predictor in the processor, and a predictor controller in the processor. The branch predictor includes multiple prediction functional blocks, each prediction functional block configured for generating predictions for control transfer instructions (CTIs) in program code based on respective prediction information, the branch predictor configured to select, from among predictions generated by the prediction functional blocks for each CTI, a selected prediction to be used for that CTI. The predictor controller keeps a record of prediction functional blocks from which the branch predictor previously selected predictions for CTIs. The predictor controller uses information from the record for controlling which prediction functional blocks are used by the branch predictor for generating predictions for CTIs.

97 MATHEMATICS AND COMPUTING↗

Unveiling the Potential of MeshGraphNets for Predicting Subsurface Evolution in Carbon Storage Projects

This is the conference paper accompanying an oral presentation “Unveiling the Potential of MeshGraphNets for Predicting Subsurface Evolution in Carbon Storage Projects” at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24 , 2024. Carbon capture and storage (CCS) technology is critical for mitigating climate change but requires effective subsurface reservoir management to ensure safe containment of injected CO2. Accurate predictions of reservoir pressure and saturation are essential for assessing long-term CCS performance. Traditional numerical simulations, while effective, are computationally intensive, time-consuming, and constrained by data discretization. Previous work has shown the effectiveness of MeshGraphNets (MGN), a graph-based machine learning framework, as an innovative alternative for predicting reservoir behavior. MGN leverages graph neural networks (GNNs) and mesh representations to model complex geological formations, offering superior adaptability across different discretizations and reservoir configurations. Classic MGN implementations utilize an autoregressive technique to predict future behavior based on current predictions, but this technique is hampered by error accumulation over time. To enhance the model accuracy in time-series predictions, this study implemented a multi-step rollout strategy that integrates autoregressive predictions during training to stabilize prediction of saturation over time. Using the Illinois Basin – Decatur Project (IBDP) dataset, comprising 100 simulations of CO2 injection, pressure, and saturation changes, the framework demonstrated its ability to learn spatial dependencies and temporal dynamics. With inputs including permeabilities, porosities, and injection rates, MGN accurately predicted CO2 plume evolution over time, even with limited training data. Moreover, the addition of a multi-step rollout procedure during training improved the ability of MGN to predict stably over time by ~15%. This research positions MGN, enhanced with multi-step rollout capabilities, as a robust and efficient tool for CCS applications. It advances the field by enabling precise, computationally efficient predictions of reservoir behavior, providing a foundation for the broader adoption of machine learning frameworks in CCS and other geoscience domains.

Holcomb, Paul↗

AlphaFold predictions are valuable hypotheses and accelerate but do not replace experimental structure determination

Abstract Artificial intelligence-based protein structure prediction methods such as AlphaFold have revolutionized structural biology. The accuracies of these predictions vary, however, and they do not take into account ligands, covalent modifications or other environmental factors. Here, we evaluate how well AlphaFold predictions can be expected to describe the structure of a protein by comparing predictions directly with experimental crystallographic maps. In many cases, AlphaFold predictions matched experimental maps remarkably closely. In other cases, even very high-confidence predictions differed from experimental maps on a global scale through distortion and domain orientation, and on a local scale in backbone and side-chain conformation. We suggest considering AlphaFold predictions as exceptionally useful hypotheses. We further suggest that it is important to consider the confidence in prediction when interpreting AlphaFold predictions and to carry out experimental structure determination to verify structural details, particularly those that involve interactions not included in the prediction.

59 BASIC BIOLOGICAL SCIENCES↗

Investigating genomic prediction strategies for grain carotenoid traits in a tropical/subtropical maize panel

Abstract Vitamin A deficiency remains prevalent on a global scale, including in regions where maize constitutes a high percentage of human diets. One solution for alleviating this deficiency has been to increase grain concentrations of provitamin A carotenoids in maize (Zea mays ssp. mays L.)—an example of biofortification. The International Maize and Wheat Improvement Center (CIMMYT) developed a Carotenoid Association Mapping panel of 380 inbred lines adapted to tropical and subtropical environments that have varying grain concentrations of provitamin A and other health-beneficial carotenoids. Several major genes have been identified for these traits, 2 of which have particularly been leveraged in marker-assisted selection. This project assesses the predictive ability of several genomic prediction strategies for maize grain carotenoid traits within and between 4 environments in Mexico. Ridge Regression-Best Linear Unbiased Prediction, Elastic Net, and Reproducing Kernel Hilbert Spaces had high predictive abilities for all tested traits (β-carotene, β-cryptoxanthin, provitamin A, lutein, and zeaxanthin) and outperformed Least Absolute Shrinkage and Selection Operator. Furthermore, predictive abilities were higher when using genome-wide markers rather than only the markers proximal to 2 or 13 genes. These findings suggest that genomic prediction models using genome-wide markers (and assuming equal variance of marker effects) are worthwhile for these traits even though key genes have already been identified, especially if breeding for additional grain carotenoid traits alongside β-carotene. Predictive ability was maintained for all traits except lutein in between-environment prediction. The TASSEL (Trait Analysis by aSSociation, Evolution, and Linkage) Genomic Selection plugin performed as well as other more computationally intensive methods for within-environment prediction. The findings observed herein indicate the utility of genomic prediction methods for these traits and could inform their resource-efficient implementation in biofortification breeding programs.

59 BASIC BIOLOGICAL SCIENCES↗

A Multi-Scale Computational Platform for Predictive Modeling of Corrosion in Al-Steel Joints (Final Report)

The research team proposed to develop innovative multi-scale models to predict corrosion and the resulting mechanical performances in aluminum-steel joints. The methods of joining considered are resistance spot welding, self-piercing riveting, and rivet-welding, all suitable for mass production applications. The multi-scale models integrate high throughput first-principle calculations based on density functional theory (DFT), high throughput calculation of phase diagrams (CALPHAD) modeling, and finite element method (FEM) simulations. These models are to be validated through laboratory experiments. Furthermore, the models are available as open source so as to enable scientists and engineers in the community to adapt and contribute to the development and application. The approaches rely on the research team’s extensive experience on the prediction of properties of individual phases at finite temperatures and variable compositions through DFT calculations, and our broad expertise on dissimilar material joining and their corrosion. The proposed computational framework enables high throughput computations for improved predictions of corrosion and the associated mechanical performance in dissimilar material joints, resulting in significant reduction in computational time needed by the current state-of-the-art methods. With the participation of researchers from three universities, an auto manufacturer, two manufacturing technology/equipment suppliers, and a software developer/vendor, the interdisciplinary research team applies the technical development on both phase-based modeling and laboratory experiments into the automobile body joining processes for validation and technology demonstration. The global cost of corrosion was estimated at about 3.4% of the global GDP in 2013. By using available corrosion control practices, it is estimated a saving between 15-35% of the cost of corrosion. In the U.S., more than $276 billion is spent repairing corrosion damage. Prediction of the corrosion and its impact on performance of the dissimilar material joints is critical for reducing the massive number of the current corrosion-based recalls for automobiles. Thus, the project goal is to develop models to enable predictive maintenance and end-of-life planning of multi-metal joints with risk of corrosion under different conditions such as exposure to high temperatures in summer and salt solutions in winter, quantified through its pH. An academia-industry consortium led by the University of Michigan and including Pennsylvania State University, University of Illinois Urbana-Champaign, University of Georgia, General Motors Company, Livermore Software Technology Corporation, and Optimal Process Technologies, LLC. created multi-scale models for prediction of corrosion in aluminum-steel joint structures such of them used in vehicle subassemblies – chassis and transmission systems. Starting from the first principle calculations, the team developed mathematical and data-driven models to predict the metallic components, which are formed during joining of two metals, for example aluminum and steel - a lightweight multilateral system which is currently used in more than 60% car bodies. These models were used for simulating chemical reactions that are happening when the joining metallic components are exposed to high temperatures and different pH values. The team was able to predict how the corrosion installs on the metallic components and how they lead to a sudden failure of components in cars. Newly developed machine learning algorithms combining Science, Technology, Engineering and Math disciplines, advanced finite element simulation and experimental validations have been integrated in a platform for prediction of the corrosion evolution and prediction the failure of joints under mechanical loadings and fatigue. Moreover, based on machine learning and inverse analysis, the team proposed solutions for designing new metallic alloys less susceptible to corrosion when joining multi-material assembles. An average of 4% error compared with experiments was achieved for the most common joints that are used in vehicle subassemblies.

36 MATERIALS SCIENCE↗

Predicting Intensive Care Unit Length of Stay and Mortality Using Patient Vital Signs: Machine Learning Model Development and Validation

Background: Patient monitoring is vital in all stages of care. In particular, intensive care unit (ICU) patient monitoring has the potential to reduce complications and morbidity, and to increase the quality of care by enabling hospitals to deliver higher-quality, cost-effective patient care, and improve the quality of medical services in the ICU. Objective: We here report the development and validation of ICU length of stay and mortality prediction models. The models will be used in an intelligent ICU patient monitoring module of an Intelligent Remote Patient Monitoring (IRPM) framework that monitors the health status of patients, and generates timely alerts, maneuver guidance, or reports when adverse medical conditions are predicted. Methods: We utilized the publicly available Medical Information Mart for Intensive Care (MIMIC) database to extract ICU stay data for adult patients to build two prediction models: one for mortality prediction and another for ICU length of stay. For the mortality model, we applied six commonly used machine learning (ML) binary classification algorithms for predicting the discharge status (survived or not). For the length of stay model, we applied the same six ML algorithms for binary classification using the median patient population ICU stay of 2.64 days. For the regression-based classification, we used two ML algorithms for predicting the number of days. We built two variations of each prediction model: one using 12 baseline demographic and vital sign features, and the other based on our proposed quantiles approach, in which we use 21 extra features engineered from the baseline vital sign features, including their modified means, standard deviations, and quantile percentages. Results: We could perform predictive modeling with minimal features while maintaining reasonable performance using the quantiles approach. The best accuracy achieved in the mortality model was approximately 89% using the random forest algorithm. The highest accuracy achieved in the length of stay model, based on the population median ICU stay (2.64 days), was approximately 65% using the random forest algorithm. Conclusions: The novelty in our approach is that we built models to predict ICU length of stay and mortality with reasonable accuracy based on a combination of ML and the quantiles approach that utilizes only vital signs available from the patient’s profile without the need to use any external features. This approach is based on feature engineering of the vital signs by including their modified means, standard deviations, and quantile percentages of the original features, which provided a richer dataset to achieve better predictive power in our models.

59 BASIC BIOLOGICAL SCIENCES↗

Uncertainty quantification of machine learning models to improve streamflow prediction under changing climate and environmental conditions

Machine learning (ML) models, and Long Short-Term Memory (LSTM) networks in particular, have demonstrated remarkable performance in streamflow prediction and are increasingly being used by the hydrological research community. However, most of these applications do not include uncertainty quantification (UQ). ML models are data driven and can suffer from large extrapolation errors when applied to changing climate/environmental conditions. UQ is required to quantify the influence of data noises on model predictions and avoid overconfident projections in extrapolation. In this work, we integrate a novel UQ method, called PI3NN, with LSTM networks for streamflow prediction. PI3NN calculates Prediction Intervals by training 3 Neural Networks. It can precisely quantify the predictive uncertainty caused by the data noise and identify out-of-distribution (OOD) data in a non-stationary condition to avoid overconfident predictions. We apply the PI3NN-LSTM method in the snow-dominant East River Watershed in the western US and in the rain-driven Walker Branch Watershed in the southeastern US. Results indicate that for the prediction data which have similar features as the training data, PI3NN precisely quantifies the predictive uncertainty with the desired confidence level; and for the OOD data where the LSTM network fails to make accurate predictions, PI3NN produces a reasonably large uncertainty indicating that the results are not trustworthy and should avoid overconfidence. PI3NN is computationally efficient, robust in performance, and generalizable to various network structures and data with no distributional assumptions. It can be broadly applied in ML-based hydrological simulations for credible prediction.

54 ENVIRONMENTAL SCIENCES↗