Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “predictive”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Fouling modeling and prediction approach for heat exchangers using deep learning

In this article, we develop a generalized and scalable statistical model for accurate prediction of fouling resistance using commonly measured parameters of industrial heat exchangers. This prediction model is based on deep learning where a scalable algorithmic architecture learns non-linear functional relationships between a set of target and predictor variables from large number of training samples. Here, the efficacy of this modeling approach is demonstrated for predicting fouling in an analytically modeled cross-flow heat exchanger, designed for waste heat recovery from flue-gas using room temperature water. The performance results of the trained models demonstrate that the mean absolute prediction errors are under 10 –4 KW –1 for flue-gas side, water side and overall fouling resistances. The coefficients of determination (R 2 ), which characterize the goodness of fit between the predictions and observed data, are over 99%. Even under varying levels of measurement noise in the inputs, we demonstrate that predictions over an ensemble of multiple neural networks achieves better accuracy and robustness to noise. We find that the proposed deep-learning fouling prediction framework learns to follow heat exchanger flow and heat transfer physics, which we confirm using locally interpretable model agnostic explanations around randomly selected operating points. Overall, we provide a robust algorithmic framework for fouling prediction that can be generalized and scaled to various types of industrial heat exchangers.

42 ENGINEERING↗

Data-driven prediction of geometry- and toolpath sequence-dependent intra-layer process conditions variations in laser powder bed fusion

Geometrical features and toolpath sequence are two important factors that cause process condition variations, such as variations in the meltpool temperature or meltpool size, that might lead to undesired material properties in the laser powder bed fusion (LPBF) process. Due to the high dynamics and complex physics of the LPBF process, it is difficult to predict variations in process conditions with simulations alone. Advances in measurement technology and computational technologies open up new possibilities for smart manufacturing. In this paper, a data-driven method to predict intra-layer variations in the processing conditions that source from the toolpath sequence and part geometry is presented. The approach is demonstrated using two-color on-axis pyrometer measurements. Three demonstration cases are presented in which it is demonstrated (1) how the trained predictive model can be used as a filter to ease the interpretation of process variations and discover patterns related to toolpath and part geometry, and (2) how to generate predictions that can be used for feedforward control, i.e., for adjusting laser power or scanning speed along the toolpath using a meltpool temperature prediction model generated based on on-axis measurements. Results show that the developed prediction model is able to meaningfully predict process variations resulted from toolpath sequence and geometry. Predictions are aligned with the results from the related work of others and for the case of 180° laser path turnarounds in our high-speed X-ray imaging experiments. In conclusion, the potential issues related to the current maturity status of the process and measuring equipment that could in practice affect the performance of the proposed solutions are also discussed.

42 ENGINEERING↗

An investigation on machine learning predictive accuracy improvement and uncertainty reduction using VAE-based data augmentation

The confluence of ultrafast computers with large memory, rapid progress in Machine Learning (ML) algorithms, and the availability of large datasets place multiple engineering fields at the threshold of dramatic progress. However, a unique challenge in nuclear engineering is data scarcity because experimentation on nuclear systems is usually more expensive and time-consuming than most other disciplines. One potential way to resolve the data scarcity issue is deep generative learning, which uses certain ML models to learn the underlying distribution of existing data and generate synthetic samples that resemble the real data. In this way, one can significantly expand the dataset to train more accurate predictive ML models. In this study, our objective is to evaluate the effectiveness of data augmentation using variational autoencoder (VAE)-based deep generative models. We investigated whether the data augmentation leads to improved accuracy in the predictions of a deep neural network (DNN) model trained using the augmented data. Additionally, the DNN prediction uncertainties are quantified using Bayesian Neural Networks (BNN) and conformal prediction (CP) to assess the impact on predictive uncertainty reduction. To test the proposed methodology, we used TRACE simulations of steady-state void fraction data based on the NUPEC Boiling Water Reactor Full-size Fine-mesh Bundle Test (BFBT) benchmark. Here, we found that augmenting the training dataset using VAEs has improved the DNN model’s predictive accuracy, improved the prediction confidence intervals, and reduced the prediction uncertainties.

Bayesian neural network↗

Comparative Analysis of TCR and TCR-pMHC Complex Structure Prediction Tools

The rapid development of computational approaches for predicting the structures of T cell receptors (TCRs) and TCR-peptide-major histocompatibility (TCR-pMHC) complexes, accelerated by AI breakthroughs such as AlphaFold, has made it feasible to calculate these structures with increasing accuracy. Although these tools show great potential, their relative accuracy and limitations remain unclear due to the lack of standardized benchmarks. Here, we systematically evaluate seven tools for predicting isolated TCR structures together with six tools for predicting TCR-pMHC complex structures. The methods include homology-based approaches, general prediction tools using AlphaFold, TCR-specific tools derived from AlphaFold2, and the newly developed tFold-TCR model. The evaluation uses a post-training data set comprising 40 αβ TCRs and 27 TCR-pMHC complexes (21 Class I and 6 Class II). Model accuracy is assessed at global, local, and interface levels using a variety of metrics. We find that each tool offers distinct advantages in various aspects of its predictions. AlphaFold2, AlphaFold3, and tFold-TCR excel in overall accuracy of TCR structure prediction, and TCRmodel2 and AlphaFold2 perform well in overall accuracy of TCR-pMHC structure prediction. However, TCR-specific tools derived from AlphaFold2 show lower accuracy in the framework region than both homology-based methods and general-purpose tools such as AlphaFold, and challenges remain for all in modeling CDR3 loops, docking orientations, TCR-peptide interfaces, and Class II MHC-peptide interfaces. Furthermore, these findings will guide researchers in selecting appropriate tools, emphasize the importance of using multiple evaluation metrics to assess model performance, and offer suggestions for improving TCR and TCR-pMHC structure prediction tools.

Chemical structure↗

CMIP6 skill at predicting interannual to multi-decadal summer monsoon precipitation variability

Monsoons affect the economy, agriculture, and human health of two thirds of the world’s population. Therefore, predicting variations in monsoon precipitation is societally important. We explore the ability of climate models from the sixth phase of the Climate Model Intercomparison Project to predict summer monsoon precipitation variability by using hindcasts from the Decadal Climate Prediction Project (Component A). The multi-model ensemble-mean shows significant skill at predicting summer monsoon precipitation from one year to 6–9 years ahead. However, this skill is dependent on the model, monsoon domain, and lead-time. In general, the skill of the multi-model ensemble-mean prediction is low in year 1 but increases for longer-lead times and is largely consistent with externally forced changes. The best captured region is northern Africa for the 2–5 and 6–9 year forecast lead times. In contrast, there is no significant skill using the ensemble-mean over East and South Asia and, furthermore, there is significant spread in skill among models for these domains. By sub-sampling the ensemble we show that the difference in skill between models is tied to the simulation of the externally forced response over East and South Asia, with models with a more skilful forced response capable of better predictions. A further contribution is from skilful prediction of Pacific Ocean temperatures for the South Asian summer monsoon at longer lead-times. Therefore, these results indicate that predictions of the East and South Asian monsoons could be significantly improved.

54 ENVIRONMENTAL SCIENCES↗

Putting AlphaFold models to work with phenix.process_predicted_model and ISOLDE

AlphaFold has recently become an important tool in providing models for experimental structure determination by X-ray crystallography and cryo-EM. Large parts of the predicted models typically approach the accuracy of experimentally determined structures, although there are frequently local errors and errors in the relative orientations of domains. Importantly, residues in the model of a protein predicted by AlphaFold are tagged with a predicted local distance difference test score, informing users about which regions of the structure are predicted with less confidence. AlphaFold also produces a predicted aligned error matrix indicating its confidence in the relative positions of each pair of residues in the predicted model. The phenix.process_predicted_model tool downweights or removes low-confidence residues and can break a model into confidently predicted domains in preparation for molecular replacement or cryo-EM docking. These confidence metrics are further used in ISOLDE to weight torsion and atom–atom distance restraints, allowing the complete AlphaFold model to be interactively rearranged to match the docked fragments and reducing the need for the rebuilding of connecting regions.

59 BASIC BIOLOGICAL SCIENCES↗

Biomass yield improvement in switchgrass through genomic prediction of flowering time

The seasonal timing of the transition from vegetative to reproductive growth has a major impact on biomass accumulation in switchgrass. Late-flowering switchgrass cultivars produce greater biomass, a critical trait for sustainable bioenergy production. Genomic prediction (GP) may allow rapid selection of late-flowering individuals with reduced time and expense for field evaluations. To evaluate GP, two flowering time traits (heading date and anthesis date) were collected on 1532 genotypes from four breeding populations: Midwest, Gulf, Atlantic, and Hybrid. These were sequenced using genotype-by-sequencing (530,792 single-nucleotide polymorphisms). Predictive ability of single-trait and multi-trait models were evaluated by cross-validation, by prediction of a progeny trial (n = 122), and through prediction of yield performance in a parallel experiment (n = 52). Predictive ability was not improved by sharing information among breeding groups. Overall, multi-trait models provided an advantage during cross-validation, but a smaller advantage during progeny prediction. Within populations, GP resulted in lower per-cycle progress than previously reported field evaluations (3.1 vs. 5.0 day –1 cycle –1 ). However, GP cycles are potentially much faster than field evaluations. When directly predicting biomass yield, the Hybrid training population had a predictive ability of 0.54–0.63. This reinforces the strong linkage between biomass yields in swards and flowering time. Furthermore, these results highlight the value of GP for rapid yield improvement in switchgrass, particularly in a breeding program designed to share information between biomass yield trials and low-cost flowering time evaluations.

09 BIOMASS FUELS↗

Multi-Trait Regressor Stacking Increased Genomic Prediction Accuracy of Sorghum Grain Composition

Genomic prediction has enabled plant breeders to estimate breeding values of unobserved genotypes and environments. The use of genomic prediction will be extremely valuable for compositional traits for which phenotyping is labor-intensive and destructive for most accurate results. We studied the potential of Bayesian multi-output regressor stacking (BMORS) model in improving prediction performance over single trait single environment (STSE) models using a grain sorghum diversity panel (GSDP) and a biparental recombinant inbred lines (RILs) population. A total of five highly correlated grain composition traits—amylose, fat, gross energy, protein and starch, with genomic heritability ranging from 0.24 to 0.59 in the GSDP and 0.69 to 0.83 in the RILs were studied. Average prediction accuracies from the STSE model were within a range of 0.4 to 0.6 for all traits across both populations except amylose (0.25) in the GSDP. Prediction accuracy for BMORS increased by 41% and 32% on average over STSE in the GSDP and RILs, respectively. Prediction of whole environments by training with remaining environments in BMORS resulted in moderate to high prediction accuracy. Our results show regression stacking methods such as BMORS have potential to accurately predict unobserved individuals and environments, and implementation of such models can accelerate genetic gain.

54 ENVIRONMENTAL SCIENCES↗

Direct Prediction of Phonon Density of States With Euclidean Neural Networks

Abstract Machine learning has demonstrated great power in materials design, discovery, and property prediction. However, despite the success of machine learning in predicting discrete properties, challenges remain for continuous property prediction. The challenge is aggravated in crystalline solids due to crystallographic symmetry considerations and data scarcity. Here, the direct prediction of phonon density‐of‐states (DOS) is demonstrated using only atomic species and positions as input. Euclidean neural networks are applied, which by construction are equivariant to 3D rotations, translations, and inversion and thereby capture full crystal symmetry, and achieve high‐quality prediction using a small training set of examples with over 64 atom types. The predictive model reproduces key features of experimental data and even generalizes to materials with unseen elements, and is naturally suited to efficiently predict alloy systems without additional computational cost. The potential of the network is demonstrated by predicting a broad number of high phononic specific heat capacity materials. The work indicates an efficient approach to explore materials' phonon structure, and can further enable rapid screening for high‐performance thermal storage materials and phonon‐mediated superconductors.

97 MATHEMATICS AND COMPUTING↗

Predicting concrete compressive strength using hybrid ensembling of surrogate machine learning models

This study aims to implement a hybrid ensemble surrogate machine learning technique in predicting the compressive strength (CS) of concrete, an important parameter used for durability design and service life prediction of concrete structures in civil engineering projects. For this purpose, an experimental database consisting of 1030 records has been compiled from the machine learning repository of the University of California, Irvine. The database was used to train and validate four conventional machine learning (CML) models, namely Artificial Neural Network (ANN), Linear and Non-Linear Multivariate Adaptive Regression Splines (MARS-L and MARS-C), Gaussian Process Regression (GPR), and Minimax Probability Machine Regression (MPMR). Subsequently, the predicted outputs of CML models were combined and trained using ANN to construct the Hybrid Ensemble Model (HENSM). It is observed that the proposed HENSM produces higher predictive accuracy compared to the CML models used in the present study. The predictive performance of all models for CS prediction was compared using the testing dataset and it is found that the HENSM model attained the highest predictive accuracy in both phases. Based on the experimental results, the newly constructed HENSM model is very potential to be a new alternative in handling the overfitting issues of CML models and hence, can be used to predict the concrete CS, including the design of less polluting and more sustainable concrete constructions.

36 MATERIALS SCIENCE↗

Predicting the proximity to macroscopic failure using local strain populations from dynamic in situ X-ray tomography triaxial compression experiments on rocks

Predicting the proximity of large-scale dynamic failure is a critical concern in the engineering and geophysical sciences. Here we use evolving contractive, dilatational, and shear strain deformation preceding failure in dynamic X-ray tomography experiments to examine which strain components best predict the proximity to failure. We develop machine learning models to predict the proximity to failure using time series of three-dimensional local incremental strain tensor fields acquired in rock deformation experiments under stress conditions of the upper crust. Three-dimensional scans acquired in situ throughout triaxial compression experiments provide a distribution of density contrasts from which we estimate the three-dimensional incremental strain that accumulates between each scan acquisition. Training machine learning models on multiple experiments of six rock types provides suites of feature importance that indicate the predictive power of each feature. Comparing the average importance of groups of features that include information about each strain component quantifies the ability of the contractive, dilatational and shear strain to predict the proximity of macroscopic failure. A total of 24 models of four machine learning algorithms with six rock types indicate that 1) the dilatational strain provides the best predictive power of the strain components, and 2) the intermediate values (25th-75th percentile) of the strain population provide the best predictive power of the statistics of the strain populations. In addition, the success of the predictions of models trained on one rock type and tested on other rock types quantifies the similarities and differences of the precursory strain accumulation process in the six rock types. These similarities suggest the potential existence of a unified theory of brittle rock deformation for a range of rock types.

58 GEOSCIENCES↗

PySIDT: Subgraph Isomorphic Decision Trees for Molecular Property Prediction

Accurate molecular property prediction is important across all fields of chemistry. Deep neural networks (DNNs) have become increasingly popular due to their ability to train automatically, avoiding the incredibly tedious process of constructing and extending traditional property estimation schemes. However, DNNs require large amounts of training data, are challenging to interpret, require large amounts of memory to load even during inference, and have severe difficulties incorporating qualitative chemical knowledge, which are often desired for molecular property prediction tasks. Here, in this study, we present PySIDT (https://github.com/zadorlab/PySIDT), a software for training and running inference on Subgraph Isomorphic Decision Trees (SIDTs). SIDTs are graph-based decision trees made of nodes associated with molecular substructures. Inference is done by descending target molecular structures down the decision tree to nodes with matching subgraph isomorphic substructures and making predictions based on the final (most specific) nodes matched. SIDTs scale down well to dataset sizes much smaller than is feasible for DNNs. As trees of molecular substructures, SIDTs are inherently readable and easy to visualize, making them easy to analyze. They are also straightforward to extend and retrain, facilitate uncertainty estimation, and enable easy integration of expert knowledge. We demonstrate the SIDT approach discussing its application to a diverse range of molecular prediction tasks: rate coefficient estimation, diffusion coefficient estimation, thermochemistry estimation, transition state bond stretch prediction, p K a prediction, stability of molecular structures, stability of surface structures, and prediction of surface lateral interaction energetics. Additionally, we demonstrate the power of the SIDT algorithms in two direct learning curve vanilla comparisons with the popular DNN-based software Chemprop and the popular gradient boosted trees-based software XGBoost on enthalpy of formation and rate coefficient prediction tasks. In particular, in the enthalpy of formation case, vanilla PySIDT is able to outperform vanilla Chemprop and XGBoost across the full range of training/validation set sizes out to 11,560 data points.

Johnson, Matthew Sean [Sandia National Laboratorie↗

A Bayesian Deep Learning Approach to Near-Term Climate Prediction

Since model bias and associated initialization shock are serious shortcomings that reduce prediction skills in state-of-the-art decadal climate prediction efforts, we pursue a complementary machine-learning-based approach to climate prediction. The example problem setting we consider consists of predicting natural variability of the North Atlantic sea surface temperature on the interannual timescale in the pre-industrial control simulation of the Community Earth System Model. While previous works have considered the use of recurrent networks such as convolutional LSTMs and reservoir computing networks in this and other similar problem settings, we currently focus on the use of feedforward convolutional networks. In particular, we find that a feedforward convolutional network with a Densenet architecture is able to outperform a convolutional LSTM in terms of predictive skill. Next, we go on to consider a probabilistic formulation of the same network based on Stein variational gradient descent and find that in addition to providing useful measures of predictive uncertainty, the probabilistic (Bayesian) version improves on its deterministic counterpart in terms of predictive skill. Finally, we characterize the reliability of the ensemble of machine learning models obtained in the probabilistic setting by using analysis tools developed in the context of ensemble numerical weather prediction.

54 ENVIRONMENTAL SCIENCES↗

Machine-learning informed prediction of high-entropy solid solution formation: Beyond the Hume-Rothery rules

The empirical rules for the prediction of solid solution formation proposed so far in the literature usually have very compromised predictability. Some rules with seemingly good predictability were, however, tested using small data sets. Based on an unprecedented large dataset containing 1252 multicomponent alloys, machine-learning methods showed that the formation of solid solutions can be very accurately predicted (93%). The machine-learning results help identify the most important features, such as molar volume, bulk modulus, and melting temperature. As such a new thermodynamics-based rule was developed to predict solid–solution alloys. The new rule is nonetheless slightly less accurate (73%) but has roots in the physical nature of the problem. The new rule is employed to predict solid solutions existing in the three blocks, each of which consists of 9 elements. The predictions encompass face-centered cubic (FCC), body-centered cubic (BCC), and hexagonal closest packed (HCP) structures in a high throughput manner. The validity of the prediction is further confirmed by CALculations of PHAse Diagram (CALPHAD) calculations with high consistency (94%). Since the new thermodynamics-based rule employs only elemental properties, applicability in screening for solid solution high-entropy alloys is straightforward and efficient.

36 MATERIALS SCIENCE↗

Soil microbiome predictability increases with spatial and taxonomic scale

Soil microorganisms shape ecosystem function, yet it remains an open question whether we can predict the composition of the soil microbiome in places before observing it. Furthermore, it is unclear whether the predictability of microbial life exhibits taxonomic- and spatial-scale dependence, as it does for macrobiological communities. Here, we leverage multiple large-scale soil microbiome surveys to develop predictive models of bacterial and fungal community composition in soil, then test these models against independent soil microbial community surveys from across the continental United States. We find remark- able scale dependence in community predictability. The predictability of bacterial and fungal communities increases with the spatial scale of observation, and fungal predictability increases with taxonomic scale. These patterns suggest that there is an increasing importance of deterministic versus stochastic processes with scale, consistent with findings in plant and animal communities, suggesting a general scaling relationship across biology. Biogeochemical functional groups and high-level taxonomic groups of microorganisms were equally predictable, indicating that traits and taxonomy are both powerful lenses for understanding soil communities. Here, by focusing on out-of-sample prediction, these findings suggest an emerging generality in our understanding of the soil microbiome, and that this understanding is fundamentally scale dependent

Biogeography↗

Incorrect computation of Madden-Julian oscillation prediction skill

The Madden–Julian oscillation (MJO) is a major tropical weather system and one of the largest sources of predictability for subseasonal-to-seasonal weather forecasts. Skillful prediction of the MJO has been a highly active area of research due to its large socio-economic impacts. Silini et al., herein S21, developed a machine learning model to predict the MJO, which they claimed to have an MJO prediction skill of 26–27 days over all seasons and 45 days for December–February (DJF) winter. If true, this would make the skill of their model competitive with that of the state-of-the-art dynamical MJO prediction systems at 20–35 days. However, here we show that the MJO prediction was calculated incorrectly in S21, which spuriously increased the performance of their model. Correctly computed skill of their model was substantially lower than that reported in S21; the skill for all seasons drops to 11–12 days and the skill for forecasts initialized during DJF drops to 15 days. Our findings clarify that the S21 machine learning model is not competitive with state-of-the-art numerical weather prediction models in predicting the MJO.

54 ENVIRONMENTAL SCIENCES↗

Machine learned synthesizability predictions aided by density functional theory

Abstract A grand challenge of materials science is predicting synthesis pathways for novel compounds. Data-driven approaches have made significant progress in predicting a compound’s synthesizability; however, some recent attempts ignore phase stability information. Here, we combine thermodynamic stability calculated using density functional theory with composition-based features to train a machine learning model that predicts a material’s synthesizability. Our model predicts the synthesizability of ternary 1:1:1 compositions in the half-Heusler structure, achieving a cross-validated precision of 0.82 and recall of 0.82. Our model shows improvement in predicting non-half-Heuslers compared to a previous study’s model, and identifies 121 synthesizable candidates out of 4141 unreported ternary compositions. More notably, 39 stable compositions are predicted unsynthesizable while 62 unstable compositions are predicted synthesizable; these findings otherwise cannot be made using density functional theory stability alone. This study presents a new approach for accurately predicting synthesizability, and identifies new half-Heuslers for experimental synthesis.

Lee, Andrew (ORCID:0000000153014295)↗

Advancing energy storage through solubility prediction: leveraging the potential of deep learning

Solubility prediction plays a crucial role in energy storage applications, such as redox flow batteries, because it directly affects the efficiency and reliability. Researchers have developed various methods that utilize quantum calculations and descriptors to predict the aqueous solubilities of organic molecules. Notably, machine learning models based on descriptors have shown promise for solubility prediction. As deep learning tools, graph neural networks (GNNs) have emerged to capture complex structure–property relationships for material property prediction. Specifically, MolGAT, a type of GNN model, was designed to incorporate n-dimensional edge attributes, enabling the modeling of intricacies in molecular graphs and enhancing the prediction capabilities. In a previous study, MolGAT successfully screened 23 467 promising redox-active molecules from a database of over 500 000 compounds, based on redox potential predictions. This study focused on applying the MolGAT model to predict the aqueous solubility (log S) of a broad range of organic compounds, including those previously screened for redox activity. The model was trained on a diverse sample of 8494 organic molecules from AqSolDB and benchmarked against literature data, demonstrating superior accuracy compared with other state of the art graph-based and descriptor-based models. Subsequently, the trained MolGAT model was employed to screen redox-active organic compounds identified in the first phase of high-throughput virtual screening, targeting favorable solubility in energy storage applications. The second round of screening, which considered solubility, yielded 12 332 promising redox-active and soluble organic molecules suitable for use in aqueous redox flow batteries. Thus, the two-phase high-throughput virtual screening approach utilizing MolGAT, specifically trained for redox potential and solubility, is an effective strategy for selecting suitable intrinsically soluble redox-active molecules from extensive databases, potentially advancing energy storage through reliable material development. This indicates that the model is reliable for predicting the solubility of various molecules and provides valuable insights for energy storage, pharmaceutical, environmental, and chemical applications.

25 ENERGY STORAGE↗