Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “predictive”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

CyProduct: A software tool for accurately predicting the byproducts of human cytochrome P450 metabolism

In silico metabolism prediction is a cheminformatic task of autonomously predicting the set of metabolic byproducts produced from a specified molecule and a set of enzymes or reactions. Here we describe a novel machine-learned in silico cytochrome P450 (CYP450) metabolism prediction suite, called CyProduct, that accurately predicts metabolic byproducts for a specified molecule and a human CYP450 isoform. It includes three modules: (1) CypReact, a tool that predicts if the query compound reacts with a given CYP450 enzyme; (2) CypBoM, a tool that accurately predicts the “bond site” of the reaction (i.e., which specific bonds within the query molecule react with the CYP isoform); and (3) MetaboGen, a tool that generates the metabolic byproducts based on CypBoM’s bond-site prediction. CyProduct predicts metabolic biotransformation products for each of the nine most important human CYP450 enzymes. CypBoM uses an important new concept called “Bond of Metabolism” (BoM), which extends the traditional “Site of Metabolism" (SoM) by specifying the information about the set of chemical bonds that is modified or formed in a metabolic reaction (rather than the specific atom). We created a BoM database for 3487 CYP450-mediated Phase I reactions, then used this to train the CypBoM Predictor to predict the reactive bond locations on substrate molecules. CypBoM Predictor’s cross-validated Jaccard score for reactive bond prediction ranged from 0.380 to 0.452 over the nine CYP450 enzymes. Over variants of a test set of 72 known CYP450 substrates and 30 non-reactants, CyProduct outperformed the other packages -- including ADMET Predictor, BioTransformer and GLORY -- by an average of 200% (wrt Jaccard score) in terms of predicting metabolites. The CyProduct suite and the datasets are freely available at https://bitbucket.org/wishartlab/cyproduct/src/master/.

Machine learning, Cytochrome P450, Metabolism pred↗

Ensemble transfer learning for the prediction of anti-cancer drug response

Abstract Transfer learning, which transfers patterns learned on a source dataset to a related target dataset for constructing prediction models, has been shown effective in many applications. In this paper, we investigate whether transfer learning can be used to improve the performance of anti-cancer drug response prediction models. Previous transfer learning studies for drug response prediction focused on building models to predict the response of tumor cells to a specific drug treatment. We target the more challenging task of building general prediction models that can make predictions for both new tumor cells and new drugs. Uniquely, we investigate the power of transfer learning for three drug response prediction applications including drug repurposing, precision oncology, and new drug development, through different data partition schemes in cross-validation. We extend the classic transfer learning framework through ensemble and demonstrate its general utility with three representative prediction algorithms including a gradient boosting model and two deep neural networks. The ensemble transfer learning framework is tested on benchmark in vitro drug screening datasets. The results demonstrate that our framework broadly improves the prediction performance in all three drug response prediction applications with all three prediction algorithms.

60 APPLIED LIFE SCIENCES↗

Tandem Predictions for HPC Jobs

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

HPC↗

Tandem Predictions for HPC Jobs: Preprint

At the core of the predictive analytics applied to High Performance Computing (HPC), the most prominent tasks are the prediction of job runtimes and the prediction of job queue times, both of which have the potential for informing HPC users during their every-day decision making. Accurate runtime predictions can help users better choose so-called wallclock times at job submission, decreasing the odds of their jobs waiting in queues longer than necessary. The accurate and timely queue time predictions offered for the available partitions can inform the favorable selection of partitions for running jobs. This potential is well understood as we see in the abundance of research studies that propose solutions for these tasks, including the work published in the last several years. These tasks are seemingly receptive to the Machine Learning (ML) solutions, considering that there is no shortage of training data where HPC centers over time run millions and millions of jobs. However, we study the existing research literature, as well as look for examples in the toolchains supported on the exemplar HPC facilities, and, surprisingly, do not find any practical solutions that are ready to be adopted. We interpret this as a manifestation of the shortage of UX/UI efforts that support HPC analytics and also as a sign that the research has not come to the consensus on solving these tasks. In this study, we aim to shed new light on the long-running task of job queue time prediction by exploring the utility of runtime predictions in improving prediction accuracy and, actually, predicting these two metrics together, in tandem. In other words, we show how runtime predictions become valuable input in the queue time modeling. We challenge the existing approaches to feature engineering for the queue time prediction and describe promising results we obtained for a large dataset of HPC jobs from a supercomputer at the National Renewable Energy Laboratory.

97 MATHEMATICS AND COMPUTING↗

Sensitivity analysis for characterizing the impact of HNGD model on the prediction of hydrogen redistribution in Zircaloy cladding using BISON code

Hydrogen in zirconium cladding is able to precipitate into zirconium hydrides which impacts cladding integrity. The Hydride Nucleation-Growth-Dissolution (HNGD) model in the BISON code accounts for the precipitation and dissolution kinetics of hydride in Zircaloy material. This paper presents global sensitivity analyses of the HNGD model aiming to enhance our understanding of the hydride precipitation phenomena by quantifying the variance that key parameters have on the prediction of hydrogen behavior under various environmental conditions. Here, model predictions are compared to experimental data obtained under two different conditions: 1) with uniformly precharged specimens subjected to a linear thermal gradient, and 2) specimens precharged with a cathodically applied hydride rim at one end of the sample and subjected to an asymmetric thermal gradient. The Sobol sensitivity analysis identifies the key parameters in the HNGD model for both types of specimens. For linear temperature cases, the heat of transport dominates the accuracy of predictions when no precipitation occurs at the cold end, while Terminal Solid Solubility for Dissolution (TSSD) is the most important parameter when precipitation occurs. A large variation in the predicted hydrogen concentration profiles is found in the range of high TSSD due to the occurrence of precipitation. For asymmetric temperature cases, the solubility coefficient gives the largest impact on the predicted hydrogen distribution, as it determines the amount of solute hydrogen dissolved from the initially applied hydride rim. A large discrepancy in hydrogen distribution between simulations and experiments exists with the asymmetric specimens because BISON simulations fail to predict the precipitation of hydride at the cooler end. Comparative studies using former and updated models verifies the significant impact of the hydride growth mechanism on predicted hydrogen concentration profiles. In particular, when hydride initially exists, changes in TSSD generate a large variation in the predicted amount of precipitation by hydride growth, giving large uncertainty in predicting the hydrogen distribution over the sample length. The outputs characterize the significant impact of the hydride growth mechanism in the HNGD model on predicting hydrogen behavior, and improve the understanding of the precipitation of hydride in Zircaloy cladding within a range of expected environmental conditions. The analyses indicate work is still needed to improve the hydride solvus models in the BISON code to accurately predict experimentally observed hydride concentrations and distributions.

36 MATERIALS SCIENCE↗

Long-Term Vehicle Speed Prediction via Historical Traffic Data Analysis for Improved Energy Efficiency of Connected Electric Vehicles

Connected and automated vehicles (CAVs) are expected to provide enhanced safety, mobility, and energy efficiency. While abundant evidence has been accumulated showing substantial energy saving potentials of CAVs through eco-driving, traffic condition prediction has remained to be the main challenge in capitalizing the gains. The coupled power and thermal subsystems of CAVs necessitate the use of different speed preview windows for effective and integrated power and thermal management. Real-time vehicle-to-infrastructure (V2I) communications can provide an accurate speed prediction over a short prediction horizon (e.g., 30 s to 60 s), but not for a long range (e.g., over 180 s). Therefore, advanced approaches are required to develop detailed speed prediction for robust optimization-based energy management of CAVs. This paper presents an integrated speed prediction framework based on historical traffic data classification and real-time V2I communications for efficient energy management of electrified CAVs. The proposed framework provides multi-range speed predictions with different fidelity over short and long horizons. The proposed multi-range speed prediction is integrated with an economic model predictive control (MPC) strategy for the battery thermal management (BTM) of connected and automated electric vehicles (EVs). The simulation results over real-world urban driving cycles confirm the enhanced prediction performance of the proposed data classification strategy over a long prediction horizon. Despite the uncertainty in long-range CAVs’ speed predictions, the vehicle-level simulation results show that 14% and 19% energy savings can be accumulated sequentially through eco-driving and BTM optimization (eco-cooling), respectively, when compared with normal driving (i.e., human driver) and conventional BTM strategy.

Engineering↗

Impact of volcanic eruptions on CMIP6 decadal predictions: a multi-model analysis

Abstract. In recent decades, three major volcanic eruptions of different intensity have occurred (Mount Agung in 1963, El Chichón in 1982 and Mount Pinatubo in 1991), with reported climate impacts on seasonal to decadal timescales that could have been potentially predicted with accurate and timely estimates of the associated stratospheric aerosol loads. The Decadal Climate Prediction Project component C (DCPP-C) includes a protocol to investigate the impact of volcanic aerosols on the climate experienced during the years that followed those eruptions through the use of decadal predictions. The interest of conducting this exercise with climate predictions is that, thanks to the initialisation, they start from the observed climate conditions at the time of the eruptions, which helps to disentangle the climatic changes due to the initial conditions and internal variability from the volcanic forcing. The protocol consists of repeating the retrospective predictions that are initialised just before the last three major volcanic eruptions but without the inclusion of their volcanic forcing, which are then compared with the baseline predictions to disentangle the simulated volcanic effects upon climate. We present the results from six Coupled Model Intercomparison Project Phase 6 (CMIP6) decadal prediction systems. These systems show strong agreement in predicting the well-known post-volcanic radiative effects following the three eruptions, which induce a long-lasting cooling in the ocean. Furthermore, the multi-model multi-eruption composite is consistent with previous work reporting an acceleration of the Northern Hemisphere polar vortex and the development of El Niño conditions the first year after the eruption, followed by a strengthening of the Atlantic Meridional Overturning Circulation the subsequent years. Our analysis reveals that all these dynamical responses are both model- and eruption-dependent. A novel aspect of this study is that we also assess whether the volcanic forcing improves the realism of the predictions. Comparing the predicted surface temperature anomalies in the two sets of hindcasts (with and without volcanic forcing) with observations we show that, overall, including the volcanic forcing results in better predictions. The volcanic forcing is found to be particularly relevant for reproducing the observed sea surface temperature (SST) variability in the North Atlantic Ocean following the 1991 eruption of Pinatubo.

Bilbao, Roberto (ORCID:0000000307294980)↗

Artificial neural network implementation of a near-ideal error prediction controller

A theory has been developed at the University of Virginia which explains the effects of including an ideal predictor in the forward loop of a linear error-sampled system. It has been shown that the presence of this ideal predictor tends to stabilize the class of systems considered. A prediction controller is merely a system which anticipates a signal or part of a signal before it actually occurs. It is understood that an exact prediction controller is physically unrealizable. However, in systems where the input tends to be repetitive or limited, (i.e., not random) near ideal prediction is possible. In order for the controller to act as a stability compensator, the predictor must be designed in a way that allows it to learn the expected error response of the system. In this way, an unstable system will become stable by including the predicted error in the system transfer function. Previous and current prediction controller include pattern recognition developments and fast-time simulation which are applicable to the analysis of linear sampled data type systems. The use of pattern recognition techniques, along with a template matching scheme, has been proposed as one realizable type of near-ideal prediction. Since many, if not most, systems are repeatedly subjected to similar inputs, it was proposed that an adaptive mechanism be used to 'learn' the correct predicted error response. Once the system has learned the response of all the expected inputs, it is necessary only to recognize the type of input with a template matching mechanism and then to use the correct predicted error to drive the system. Suggested here is an alternate approach to the realization of a near-ideal error prediction controller, one designed using Neural Networks. Neural Networks are good at recognizing patterns such as system responses, and the back-propagation architecture makes use of a template matching scheme. In using this type of error prediction, it is assumed that the system error responses be known for a particular input and modeled plant. These responses are used in the error prediction controller. An analysis was done on the general dynamic behavior that results from including a digital error predictor in a control loop and these were compared to those including the near-ideal Neural Network error predictor. This analysis was done for a second and third order system.

Mcvey, Eugene S.↗

Prediction of Transitional Flows in the Low Pressure Turbine

Current turbulence models tend to give too early and too short a length of flow transition to turbulence, and hence fail to predict flow separation induced by the adverse pressure gradients and streamline flow curvatures. Our discussion will focus on the development and validation of transition models. The baseline data for model comparisons are the T3 series, which include a range of free-stream turbulence intensity and cover zero-pressure gradient to aft-loaded turbine pressure gradient flows. The method will be based on the conditioned N-S equations and a transport equation for the intermittency factor. First, several of the most popular 2-equation models in predicting flow transition are examined: k-e [Launder-Sharina], k-w [Wilcox], Lien-Leschiziner and SST [Menter] models. All models fail to predict the onset and the length of transition, even for the simplest flat plate with zero-pressure gradient(T3A). Although the predicted onset position of transition can be varied by providing different inlet turbulent energy dissipation rates, the appropriate inlet conditions for turbulence quantities should be adjusted to match the decay of the free-stream turbulence. Arguably, one may adjust the low-Reynolds-number part of the model to predict transition. This approach has so far not been very successful. However, we have found that the low-Reynolds-number model of Launder and Sharma [1974], which is an improved version of Jones and Launder [1972] gave the best overall performance. The Launder and Sharma model was designed to capture flow re-laminarization (a reverse of flow transition), but tends to give rise to a too early and too fast transition in comparison with the physical transition. The three test cases were for flows with zero pressure gradient but with different free-stream turbulent intensities. The same can be said about the model when considering flows subject to pressure gradient(T3C1). To capture the effects of transition using existing turbulence models, one approach is to make use of the concept of the intermittency to predict the flow transition. It was originally based on the intermittency distribution of Narasimha [1957], and then gradually evolved into a transport equation for the intermittency factor. Gostelow and associates [1994,1995] have made some improvements to Narasimha's method in an attempt to account for both favorable and adverse pressure gradients. Their approach is based on a linear, explicit combination of laminar and turbulent solutions. This approach fails to predict the overshoot of the skin friction on a flat plate near the end of transition zone, even though the length of transition is well predicted. The major flaw of Gostelow's approach is that it assumes the non-turbulent part being the laminar solution and the turbulent part being the turbulent solution and they do not interact across the transitional region. The technique in condition averaging the flow equations in intermittent flows was first introduced by Libby [1975] and Dopazo [1977] and further refined by Dick and associates [1988, 1996]. This approach employs two sets of transport equations for the non-turbulent part and the other for the turbulent part. The advantage of this approach is that it allows the interaction of non-turbulent and turbulent velocities through the introduction of additional source terms in the continuity and momentum equations for the non-turbulent and turbulent velocities. However, the strong coupling of the two sets of equations has caused some numerical difficulties, which requires special attention. The prediction of the skin friction can be improved by this approach via the implicit coupling of non-turbulent and turbulent velocity flelds. Another improvement of the interrmittency model can be further made by allowing the intermittency to vary in the cross-stream direction. This is one step prior to testing any proposal for the transport equation for the intermittency factor. Instead of solving the transport equation for the intermittency factor, the distribution for the intermittency factor is prescribed by Klebanoff's empirical formula [1955]. The skin friction is very well predicted by this new modification, including the overshoot of the profile near the end of the transition zone. The outcome of this study is very encouraging since it indicates that the proper description of the intermittency distribution is the key to the success of the model prediction. This study will be used to guide us on the modelling of the intermittency transport equation.

Huang, George↗

All Recent Mars Landers Have Landed Downrange - Are Mars Atmosphere Models Mis-Predicting Density?

All recent Mars landers (Mars Pathfinder, the two Mars Exploration Rovers Spirit and Opportunity, and the Mars Phoenix Lander) have landed further downrange than their pre-entry predictions. Mars Pathfinder landed 27 km downrange of its prediction [1], Spirit and Opportunity landed 13.4 km and 14.9 km, respectively, downrange from their predictions [2], and Phoenix landed 21 km downrange from its prediction [3]. Reconstruction of their entries revealed a lower density profile than the best a priori atmospheric model predictions. Do these results suggest that there is a systemic issue in present Mars atmosphere models that predict a higher density than observed on landing day? Spirit Landing: The landing location for Spirit was 13.4 km downrange of the prediction as shown in Fig. 1. The navigation errors upon Mars arrival were very small [2]. As such, the entry interface conditions were not responsible for this downrange landing. Consequently, experiencing a lower density during the entry was the underlying cause. The reconstructed density profile that Spirit experienced is shown in Fig. 2, which is plotted as a fraction of the pre-entry baseline prediction that was used for all the entry, descent, and landing (EDL) design analyses. The reconstructed density is observed to be less dense throughout the descent reaching a maximum reduction of 15% at 21 km. This lower density corresponded to approximately a 1- low profile relative to the dispersions predicted. Nearly all the deceleration during the entry occurs within 10- 50 km. As such, prediction of density within this altitude band is most critical for entry flight dynamics analyses and design (e.g., aerodynamic and aerothermodynamic predictions, landing location, etc.).

Desai, Prasun N.↗

Predictions of Solar Cycle 24: How are We Doing?

Predictions of solar activity are an essential part of our Space Weather forecast capability. Users are requiring usable predictions of an upcoming solar cycle to be delivered several years before solar minimum. A set of predictions of the amplitude of Solar Cycle 24 accumulated in 2008 ranged from zero to unprecedented levels of solar activity. The predictions formed an almost normal distribution, centered on the average amplitude of all preceding solar cycles. The average of the current compilation of 105 predictions of the annual-average sunspot number is 106 +/- 31, slightly lower than earlier compilations but still with a wide distribution. Solar Cycle 24 is on track to have a below-average amplitude, peaking at an annual sunspot number of about 80. Our need for solar activity predictions and our desire for those predictions to be made ever earlier in the preceding solar cycle will be discussed. Solar Cycle 24 has been a below-average sunspot cycle. There were peaks in the daily and monthly averaged sunspot number in the Northern Hemisphere in 2011 and in the Southern Hemisphere in 2014. With the rapid increase in solar data and capability of numerical models of the solar convection zone we are developing the ability to forecast the level of the next sunspot cycle. But predictions based only on the statistics of the sunspot number are not adequate for predicting the next solar maximum. I will describe how we did in predicting the amplitude of Solar Cycle 24 and describe how solar polar field predictions could be made more accurate in the future.

Pesnell, William D.↗

Navigation Prediction Performance During OSIRIS-REx Proximity Operations at (101955) Bennu

The OSIRIS-REx (Origins, Spectral Interpretation, Resource Identification and Security–Regolith Explorer) Orbit Determination team performed covariance analyses prior to the commencement of proximity operations (ProxOps) at (101955) Bennu to determine the expected predicted trajectory performance in order to meet trajectory knowledge requirements throughout each phase of the mission. One of the primary requirements placed on the predicted trajectory performance was based on the performance during orbital phases leading up to the maneuver to initiate the Touch-and-Go (TAG) trajectory descent. Throughout ProxOps the nominal force models being used to predict the spacecraft trajectory were updated in an effort to improve the prediction performance. The most significant models that contributed to prediction performance were of solar radiation pressure, thermal reradiation of the spacecraft, predicted attitude errors, and desaturation maneuvers. Efforts were made throughout all of ProxOps to monitor, trend, predict, and update spacecraft modeling to improve the prediction performance. These efforts were vital to reduce the spacecraft knowledge errors necessary to achieve a TAG target smaller than pre-launch analysis allowed due to the rough terrain of Bennu. Increased precision in predicted trajectory errors allowed for refined uncertainties to be used for future phase planning throughout the mission. The navigation team successfully predicted the spacecraft trajectory throughout all of ProxOps achieving predicted trajectories errors less than originally analyzed.

Jason M. Leonard↗

Controlling prediction functional blocks used by a branch predictor in a processor

An electronic device includes a processor, a branch predictor in the processor, and a predictor controller in the processor. The branch predictor includes multiple prediction functional blocks, each prediction functional block configured for generating predictions for control transfer instructions (CTIs) in program code based on respective prediction information, the branch predictor configured to select, from among predictions generated by the prediction functional blocks for each CTI, a selected prediction to be used for that CTI. The predictor controller keeps a record of prediction functional blocks from which the branch predictor previously selected predictions for CTIs. The predictor controller uses information from the record for controlling which prediction functional blocks are used by the branch predictor for generating predictions for CTIs.

97 MATHEMATICS AND COMPUTING↗

Unveiling the Potential of MeshGraphNets for Predicting Subsurface Evolution in Carbon Storage Projects

This is the conference paper accompanying an oral presentation “Unveiling the Potential of MeshGraphNets for Predicting Subsurface Evolution in Carbon Storage Projects” at the 17th International Conference on Greenhouse Gas Control Technologies GHGT-17 held in Calgary, Canada, October 20-24 , 2024. Carbon capture and storage (CCS) technology is critical for mitigating climate change but requires effective subsurface reservoir management to ensure safe containment of injected CO2. Accurate predictions of reservoir pressure and saturation are essential for assessing long-term CCS performance. Traditional numerical simulations, while effective, are computationally intensive, time-consuming, and constrained by data discretization. Previous work has shown the effectiveness of MeshGraphNets (MGN), a graph-based machine learning framework, as an innovative alternative for predicting reservoir behavior. MGN leverages graph neural networks (GNNs) and mesh representations to model complex geological formations, offering superior adaptability across different discretizations and reservoir configurations. Classic MGN implementations utilize an autoregressive technique to predict future behavior based on current predictions, but this technique is hampered by error accumulation over time. To enhance the model accuracy in time-series predictions, this study implemented a multi-step rollout strategy that integrates autoregressive predictions during training to stabilize prediction of saturation over time. Using the Illinois Basin – Decatur Project (IBDP) dataset, comprising 100 simulations of CO2 injection, pressure, and saturation changes, the framework demonstrated its ability to learn spatial dependencies and temporal dynamics. With inputs including permeabilities, porosities, and injection rates, MGN accurately predicted CO2 plume evolution over time, even with limited training data. Moreover, the addition of a multi-step rollout procedure during training improved the ability of MGN to predict stably over time by ~15%. This research positions MGN, enhanced with multi-step rollout capabilities, as a robust and efficient tool for CCS applications. It advances the field by enabling precise, computationally efficient predictions of reservoir behavior, providing a foundation for the broader adoption of machine learning frameworks in CCS and other geoscience domains.

Holcomb, Paul↗

AlphaFold predictions are valuable hypotheses and accelerate but do not replace experimental structure determination

Abstract Artificial intelligence-based protein structure prediction methods such as AlphaFold have revolutionized structural biology. The accuracies of these predictions vary, however, and they do not take into account ligands, covalent modifications or other environmental factors. Here, we evaluate how well AlphaFold predictions can be expected to describe the structure of a protein by comparing predictions directly with experimental crystallographic maps. In many cases, AlphaFold predictions matched experimental maps remarkably closely. In other cases, even very high-confidence predictions differed from experimental maps on a global scale through distortion and domain orientation, and on a local scale in backbone and side-chain conformation. We suggest considering AlphaFold predictions as exceptionally useful hypotheses. We further suggest that it is important to consider the confidence in prediction when interpreting AlphaFold predictions and to carry out experimental structure determination to verify structural details, particularly those that involve interactions not included in the prediction.

59 BASIC BIOLOGICAL SCIENCES↗

Investigating genomic prediction strategies for grain carotenoid traits in a tropical/subtropical maize panel

Abstract Vitamin A deficiency remains prevalent on a global scale, including in regions where maize constitutes a high percentage of human diets. One solution for alleviating this deficiency has been to increase grain concentrations of provitamin A carotenoids in maize (Zea mays ssp. mays L.)—an example of biofortification. The International Maize and Wheat Improvement Center (CIMMYT) developed a Carotenoid Association Mapping panel of 380 inbred lines adapted to tropical and subtropical environments that have varying grain concentrations of provitamin A and other health-beneficial carotenoids. Several major genes have been identified for these traits, 2 of which have particularly been leveraged in marker-assisted selection. This project assesses the predictive ability of several genomic prediction strategies for maize grain carotenoid traits within and between 4 environments in Mexico. Ridge Regression-Best Linear Unbiased Prediction, Elastic Net, and Reproducing Kernel Hilbert Spaces had high predictive abilities for all tested traits (β-carotene, β-cryptoxanthin, provitamin A, lutein, and zeaxanthin) and outperformed Least Absolute Shrinkage and Selection Operator. Furthermore, predictive abilities were higher when using genome-wide markers rather than only the markers proximal to 2 or 13 genes. These findings suggest that genomic prediction models using genome-wide markers (and assuming equal variance of marker effects) are worthwhile for these traits even though key genes have already been identified, especially if breeding for additional grain carotenoid traits alongside β-carotene. Predictive ability was maintained for all traits except lutein in between-environment prediction. The TASSEL (Trait Analysis by aSSociation, Evolution, and Linkage) Genomic Selection plugin performed as well as other more computationally intensive methods for within-environment prediction. The findings observed herein indicate the utility of genomic prediction methods for these traits and could inform their resource-efficient implementation in biofortification breeding programs.

59 BASIC BIOLOGICAL SCIENCES↗

A Multi-Scale Computational Platform for Predictive Modeling of Corrosion in Al-Steel Joints (Final Report)

The research team proposed to develop innovative multi-scale models to predict corrosion and the resulting mechanical performances in aluminum-steel joints. The methods of joining considered are resistance spot welding, self-piercing riveting, and rivet-welding, all suitable for mass production applications. The multi-scale models integrate high throughput first-principle calculations based on density functional theory (DFT), high throughput calculation of phase diagrams (CALPHAD) modeling, and finite element method (FEM) simulations. These models are to be validated through laboratory experiments. Furthermore, the models are available as open source so as to enable scientists and engineers in the community to adapt and contribute to the development and application. The approaches rely on the research team’s extensive experience on the prediction of properties of individual phases at finite temperatures and variable compositions through DFT calculations, and our broad expertise on dissimilar material joining and their corrosion. The proposed computational framework enables high throughput computations for improved predictions of corrosion and the associated mechanical performance in dissimilar material joints, resulting in significant reduction in computational time needed by the current state-of-the-art methods. With the participation of researchers from three universities, an auto manufacturer, two manufacturing technology/equipment suppliers, and a software developer/vendor, the interdisciplinary research team applies the technical development on both phase-based modeling and laboratory experiments into the automobile body joining processes for validation and technology demonstration. The global cost of corrosion was estimated at about 3.4% of the global GDP in 2013. By using available corrosion control practices, it is estimated a saving between 15-35% of the cost of corrosion. In the U.S., more than $276 billion is spent repairing corrosion damage. Prediction of the corrosion and its impact on performance of the dissimilar material joints is critical for reducing the massive number of the current corrosion-based recalls for automobiles. Thus, the project goal is to develop models to enable predictive maintenance and end-of-life planning of multi-metal joints with risk of corrosion under different conditions such as exposure to high temperatures in summer and salt solutions in winter, quantified through its pH. An academia-industry consortium led by the University of Michigan and including Pennsylvania State University, University of Illinois Urbana-Champaign, University of Georgia, General Motors Company, Livermore Software Technology Corporation, and Optimal Process Technologies, LLC. created multi-scale models for prediction of corrosion in aluminum-steel joint structures such of them used in vehicle subassemblies – chassis and transmission systems. Starting from the first principle calculations, the team developed mathematical and data-driven models to predict the metallic components, which are formed during joining of two metals, for example aluminum and steel - a lightweight multilateral system which is currently used in more than 60% car bodies. These models were used for simulating chemical reactions that are happening when the joining metallic components are exposed to high temperatures and different pH values. The team was able to predict how the corrosion installs on the metallic components and how they lead to a sudden failure of components in cars. Newly developed machine learning algorithms combining Science, Technology, Engineering and Math disciplines, advanced finite element simulation and experimental validations have been integrated in a platform for prediction of the corrosion evolution and prediction the failure of joints under mechanical loadings and fatigue. Moreover, based on machine learning and inverse analysis, the team proposed solutions for designing new metallic alloys less susceptible to corrosion when joining multi-material assembles. An average of 4% error compared with experiments was achieved for the most common joints that are used in vehicle subassemblies.

36 MATERIALS SCIENCE↗

Predicting Intensive Care Unit Length of Stay and Mortality Using Patient Vital Signs: Machine Learning Model Development and Validation

Background: Patient monitoring is vital in all stages of care. In particular, intensive care unit (ICU) patient monitoring has the potential to reduce complications and morbidity, and to increase the quality of care by enabling hospitals to deliver higher-quality, cost-effective patient care, and improve the quality of medical services in the ICU. Objective: We here report the development and validation of ICU length of stay and mortality prediction models. The models will be used in an intelligent ICU patient monitoring module of an Intelligent Remote Patient Monitoring (IRPM) framework that monitors the health status of patients, and generates timely alerts, maneuver guidance, or reports when adverse medical conditions are predicted. Methods: We utilized the publicly available Medical Information Mart for Intensive Care (MIMIC) database to extract ICU stay data for adult patients to build two prediction models: one for mortality prediction and another for ICU length of stay. For the mortality model, we applied six commonly used machine learning (ML) binary classification algorithms for predicting the discharge status (survived or not). For the length of stay model, we applied the same six ML algorithms for binary classification using the median patient population ICU stay of 2.64 days. For the regression-based classification, we used two ML algorithms for predicting the number of days. We built two variations of each prediction model: one using 12 baseline demographic and vital sign features, and the other based on our proposed quantiles approach, in which we use 21 extra features engineered from the baseline vital sign features, including their modified means, standard deviations, and quantile percentages. Results: We could perform predictive modeling with minimal features while maintaining reasonable performance using the quantiles approach. The best accuracy achieved in the mortality model was approximately 89% using the random forest algorithm. The highest accuracy achieved in the length of stay model, based on the population median ICU stay (2.64 days), was approximately 65% using the random forest algorithm. Conclusions: The novelty in our approach is that we built models to predict ICU length of stay and mortality with reasonable accuracy based on a combination of ML and the quantiles approach that utilizes only vital signs available from the patient’s profile without the need to use any external features. This approach is based on feature engineering of the vital signs by including their modified means, standard deviations, and quantile percentages of the original features, which provided a richer dataset to achieve better predictive power in our models.

59 BASIC BIOLOGICAL SCIENCES↗