Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Human limits in machine learning: prediction of potato yield and disease using soil microbiome data

Abstract Background The preservation of soil health is a critical challenge in the 21st century due to its significant impact on agriculture, human health, and biodiversity. We provide one of the first comprehensive investigations into the predictive potential of machine learning models for understanding the connections between soil and biological phenotypes. We investigate an integrative framework performing accurate machine learning-based prediction of plant performance from biological, chemical, and physical properties of the soil via two models: random forest and Bayesian neural network. Results Prediction improves when we add environmental features, such as soil properties and microbial density, along with microbiome data. Different preprocessing strategies show that human decisions significantly impact predictive performance. We show that the naive total sum scaling normalization that is commonly used in microbiome research is one of the optimal strategies to maximize predictive power. Also, we find that accurately defined labels are more important than normalization, taxonomic level, or model characteristics. ML performance is limited when humans can’t classify samples accurately. Lastly, we provide domain scientists via a full model selection decision tree to identify the human choices that optimize model prediction power. Conclusions Our study highlights the importance of incorporating diverse environmental features and careful data preprocessing in enhancing the predictive power of machine learning models for soil and biological phenotype connections. This approach can significantly contribute to advancing agricultural practices and soil health management.

Aghdam, Rosa↗

Leveraging structure-informed machine learning for fast steric zipper propensity prediction across whole proteomes

Predicting the amyloid fold and the propensity of peptide segments to adopt amyloid-like structures remain a challenge. However, recent progress has facilitated structure-based prediction of steric zipper propensity and the use of machine learning to accelerate the calculation of predictive models across many scientific areas. Leveraging these advances, we have developed a new approach for rapid proteome-wide assessment of zipper profiles that is informed by four million steric zipper predictions collected over ten years. This collection is used to build a machine learning model capable of rapidly predicting steric zipper propensity, and allowing for the assessment of zippers at both the protein and proteome level. Our predictions show enrichment for zipper forming segments in proteins involved in cell wall reorganization in yeast, highlighting a potential category of interest for experimental characterization. Overall, our predictive model allows for the exploration of amyloid formation across the tree of life and provides a tool for assessment of both novel and designed sequences for zipper density.

Biochemistry & Molecular Biology↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

Probabilistic Predictions for Fastener Failure in the Sandia Mechanics Challenge Using the Discrete-Direct Uncertainty Quantification Approach

This paper documents the blind and post-blind analysis predictions for the 2023 Sandia Mechanics Challenge (SMC), which involved predicting the behavior of a threaded fastener joint structure subjected to shock loading. Utilizing repeat sets of fastener calibration data from various experimental configurations including tension, double shear, and joint tension, we developed a library of calibrated models which were propagated through the application model using the Discrete-Direct (DD) uncertainty quantification (UQ) approach. Although the initial blind predictions did not incorporate spare-sample processing to quantify fastener failure probabilities, the analyses yielded reasonable conclusions aligned with experimental results. In the post-blind analysis phase, we focused on enhancing the fidelity of the aluminum constitutive model and innovating the DD approach to obtain probabilistic predictions for fastener failure, particularly when quantities of interest (QoIs) approach their bounds. The improved aluminum model captures the behavior of the cantilever under shock loading more accurately, predicting both partial and complete cracks, although it tends to underpredict failure propagation. The enhanced DD approach facilitates probabilistic predictions that reflect the interdependent failure mechanisms of the fasteners and the cantilever, revealing that while certain fasteners are more likely to fail, the failure does not necessarily follow a progressive pattern. Overall, the post-blind analyses significantly improved the predictive capabilities of the model, providing valuable insights into the SMC application and establishing a robust foundation for informed engineering decisions. The methodology demonstrates a cost-effective and extensible approach suitable for a wide range of applications, highlighting the importance of uncertainty quantification to provide context for engineering decision making.

42 ENGINEERING↗

Prediction of Silicon Content in a Blast Furnace via Machine Learning: A Comprehensive Processing and Modeling Pipeline

Silicon content plays an important role in determining the operational efficiency of blast furnaces (BFs) and their downstream processes in integrated steelmaking; however, existing sampling methods and first-principles models are somewhat limited in their capability and flexibility. Current data-based prediction models primarily rely on a limited set of manually selected furnace parameters. Additionally, different BFs present a diverse set of operating parameters and state variables that are known to directly influence the hot metal’s silicon content, such as fuel injection, blast temperature, and raw material charge composition, among other process variables that have their own impacts. The expansiveness of the parameter set adds complexity to parameter selection and processing. This highlights the need for a comprehensive methodology to integrate and select from all relevant parameters for accurate silicon content prediction. Providing accurate silicon content predictions would enable operators to adjust furnace conditions dynamically, improving safety and reducing economic risk. To address these issues, a two-stage approach is proposed. First, a generalized data processing scheme is proposed to accommodate diverse furnace parameters. Second, a robust modeling pipeline is used to establish a machine learning (ML) model capable of predicting hot metal silicon content with reasonable accuracy. The method employed herein predicted the average Si content of the upcoming furnace cast with an accuracy of 91% among 200 target predictions for a specific furnace provisioned by the XGBoost model. This prediction is achieved using only the past shift’s operating conditions, which should be available in real time. This performance provides a strong baseline for the modeling approach with potential for further improvement through provision of real-time features.

Chemistry↗

New framework for benchmarking decadal predictions leveraging the PCMDI Metric Package with interactive visualization

Reliable climate predictions across multiple timescales are increasingly critical as climate-related risks continue to rise. With the growing number and diversity of climate prediction systems, systematic intercomparison has become essential. Here, we present a comprehensive evaluation framework based on the PCMDI Metric Package to assess the performance of multiple decadal climate prediction systems. Unlike uninitialized simulations, initialized predictions exhibit bias and predictive skill that evolve with forecast lead time. To address this, we introduce (1) model-by-lead-time portrait plots, which efficiently summarize metrics of global temperature, precipitation, and Arctic/Antarctic sea-ice extent, and (2) an HTML-based interactive visualization platform that provides detailed regional and seasonal diagnostics of model bias, skill scores, and ensemble spread for each model and lead time. Comparisons with uninitialized simulations further quantify the relative impacts of initialization and external forcing on prediction skill. The proposed framework provides a scalable and transparent approach for multi-model climate prediction assessments and can be readily extended to a wide range of operational and research forecasting systems.

54 ENVIRONMENTAL SCIENCES↗

Foaming prediction in pure liquids from dimensionless numbers inspired by the theory of fluid behavior for drops

Foaming prediction is critical for selecting materials and designing processes in industries such as bioprocessing and gas processing. Existing models lack the generality needed for a wide range of materials and overlook the foaming behavior in pure liquids. Here, this work presents a novel method for predicting foaming in pure liquids based on their density, surface tension, and viscosity, using Reynolds ( Re ) and Ohnesorge ( Oh ) numbers. A foaming prediction map, leveraging the theory of fluid drop behavior, was developed by plotting these numbers. This map delineates distinct non-foaming and foaming regions, functioning as a binary classifier for foaming predictions. The map was fitted and validated through shake test experiments on 46 liquids, demonstrating reliable predictions, except for a specific region characterized by small Oh and large Re numbers. This region corresponded to relatively low foam stability and high turbulence, making foaming predictions challenging for liquids in this category.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A data-driven framework for predicting machining stability: employing simulated data, operational modal analysis, and enhanced transfer learning

Chatter, a self-excited vibration phenomenon, presents a significant challenge in machining operations, particularly in high-speed milling, where it can degrade tool life, reduce material removal efficiency, and compromise workpiece quality. Addressing this challenge requires a reliable predictive model that can accommodate the complex dynamics of various machining scenarios. This study introduces a novel, data-driven approach to predicting machining stability, leveraging over 140,000 simulated datasets and employing advanced techniques such as operational modal analysis (OMA), enhanced transfer learning (TL), and receptance coupling substructure analysis (RCSA). By integrating these methodologies, the framework effectively classifies and predicts chatter across diverse operational modes, achieving robust and accurate outcomes. Our model utilizes a Random Forest (RF) classifier trained with the comprehensive dataset, which demonstrates substantial improvements in both predictive accuracy and robustness. Specifically, the RF model achieved an accuracy rate of 85%, an area under the curve (AUC) of 0.90, and an F1 score of 0.88, underscoring its capability to adapt to varying machining configurations. These results highlight the framework’s potential to enhance operational efficiency and machining quality by providing reliable chatter predictions across a broad range of machining parameters. In conclusion, this research thus offers a significant advancement in predictive maintenance for machining processes, enabling more stable and efficient manufacturing operations.

42 ENGINEERING↗

An accelerated framework for predicting creep rupture lifetimes in engineering alloys

Confidently predicting high-temperature deformation, including creep and creep rupture, is paramount for the design and commercialization of candidate materials for advanced nuclear energy systems. To accelerate creep quantification, we introduce a framework that enables rapid, cost-effective, and reliable prediction of creep rupture lifetimes, minimizing reliance on time-intensive bulk creep testing. Unlike conventional creep analysis, which requires extensive time and resources, our method leverages a maximum of four short-term bulk creep tests as training data for prediction. This framework combines high-throughput nanoindentation up to 700 °C with these targeted bulk tests to inform our creep rupture model in order to predict rupture lifetimes. The strong agreement between our predictions and conventional experimental data demonstrates the effectiveness of our approach for accelerated creep analysis and lifetime prediction of structural components in high-temperature applications. Our multi-pronged approach motivates further integration of computational tools and advanced instrumentation to establish a universal framework for understanding high-temperature material responses.

36 MATERIALS SCIENCE↗

EC-Bench: A Benchmark for Enzyme Commission Number Prediction

Enzymes are proteins that catalyze specific biochemical reactions in cells. Enzyme Commission (EC) numbers are used to annotate enzymes in a four-level hierarchy that classifies enzymes based on the specific chemical reactions they catalyze. Accurate EC number prediction is essential for understanding enzyme functions. Despite the availability of numerous methods for predicting EC numbers from protein sequences, there is no unified framework for evaluating and studying such methods systematically. This gap limits the ability of the community to identify the most effective approaches for enzyme annotation. We introduce EC-Bench, a benchmark for EC number prediction, consisting of 1) an initial representative set of existing methods (including homology-based, deep learning, contrastive learning, and language model methods), 2) existing and novel accuracy and efficiency performance metrics, and 3) selected datasets to allow for comprehensive comparative study. EC-Bench is open-source and provides a framework for researchers to not only compare among existing methods objectively under uniform conditions, but also to introduce and effectively evaluate performance of new methods in a comparative framework. To demonstrate the utility of EC-Bench, we perform extensive experimentation to compare the existing EC number prediction methods and establish their advantages and disadvantages in a variety of prediction tasks, namely “exact EC number prediction”, “EC number completion” and (partial or additional) “EC number recommendation”. We find wide variation in the performance of different methods, but also subtle but potentially useful differences in the performance of different methods across tasks and for different parts of the EC hierarchy.

59 BASIC BIOLOGICAL SCIENCES↗

Genomic prediction of regional-scale performance in switchgrass ( Panicum virgatum ) by accounting for genotype-by-environment variation and yield surrogate traits

Switchgrass is a potential crop for bioenergy or carbon capture schemes, but further yield improvements through selective breeding are needed to encourage commercialization. To identify promising switchgrass germplasm for future breeding efforts, we conducted multisite and multitrait genomic prediction with a diversity panel of 630 genotypes from 4 switchgrass subpopulations (Gulf, Midwest, Coastal, and Texas), which were measured for spaced plant biomass yield across 10 sites. Our study focused on the use of genomic prediction to share information among traits and environments. Specifically, we evaluated the predictive ability of cross-validation (CV) schemes using only genetic data and the training set (cross-validation 1: CV1), a subset of the sites (cross-validation 2: CV2), and/or with 2 yield surrogates (flowering time and fall plant height). We found that genotype-by-environment interactions were largely due to the north–south distribution of sites. The genetic correlations between the yield surrogates and the biomass yield were generally positive (mean height r = 0.85; mean flowering time r = 0.45) and did not vary due to subpopulation or growing region (North, Middle, or South). Genomic prediction models had CV predictive abilities of –0.02 for individuals using only genetic data (CV1), but 0.55, 0.69, 0.76, 0.81, and 0.84 for individuals with biomass performance data from 1, 2, 3, 4, and 5 sites included in the training data (CV2), respectively. To simulate a resource-limited breeding program, we determined the predictive ability of models provided with the following: 1 site observation of flowering time (0.39); 1 site observation of flowering time and fall height (0.51); 1 site observation of fall height (0.52); 1 site observation of biomass (0.55); and 5 site observations of biomass yield (0.84). The ability to share information at a regional scale is very encouraging, but further research is required to accurately translate spaced plant biomass to commercial-scale sward biomass performance.

09 BIOMASS FUELS↗

Optimizing genomic prediction for complex traits via investigating multiple factors in switchgrass

Genomic prediction has accelerated breeding processes and provided mechanistic insights into the genetic bases of complex traits. To further optimize genomic prediction, we assess the impact of genome assemblies, genotyping approaches, variant types, allelic complexities, polyploidy levels, and population structures on the prediction of 20 complex traits in switchgrass (Panicum virgatum L.), a perennial biofuel feedstock. Surprisingly, short read-based genome assembly performs comparably to or even better than long read-based assembly. Due to higher gene coverage, exome capture and multi-allelic variants outperform genotyping-by-sequencing and bi-allelic variants, respectively. Tetraploid models show higher prediction accuracy than octoploid models for most traits, likely due to the greater genetic distances among tetraploids. Depending on the trait in question, different types of variants need to be integrated for optimal predictions. Furthermore, our study provides insights into the factors influencing genomic prediction outcomes, guiding best practices for future studies and for improving agronomic traits in switchgrass and other species through selective breeding.

60 APPLIED LIFE SCIENCES↗

Neural-Network-Enhanced COTSIM: Advancing Predictive Capabilities for Fast DIII-D Simulations

Sustaining fusion reactions in tokamaks requires heating plasma to thermonuclear temperatures while maintaining confinement and stability. Neutral beam injection (NBI) provides heating, current drive, torque, and fueling, while electron cyclotron (EC) waves are widely used for heating and current drive; together, these actuators shape the plasma current, temperature, and density profiles. The control-oriented tokamak simulator (COTSIM), a predictive, control-oriented code, has been enhanced with neural-network surrogates for transport and sources. Turbulent transport is predicted by MMMnet—a neural-network version of the updated multimode model (MMM 9.0.10)—with significantly reduced computation time relative to MMM; neoclassical transport follows the Chang–Hinton model. NUBEAMnet, a surrogate of the Monte Carlo NUBEAM module, predicts beam-driven heating, current, and torque. EC heating and current drive use a control-oriented, empirically scaled source model; plasma resistivity follows the Spitzer formulation; bootstrap current uses the Sauter model. Equilibrium is computed using both prescribed and fixed-boundary solvers (FBSs), and the pedestal structure is modeled with an empirical pedestal model. For a representative DIII-D discharge, COTSIM predicts electron and ion temperature and safety-factor profiles in close agreement with TRANSP predictive and interpretive simulations while extending predictions through the pedestal region to the plasma edge (versus 80% of the minor radius in TRANSP). Furthermore, the equivalent COTSIM simulation runs in under 3 min compared to about 2 h for TRANSP, enabling rapid scenario planning, optimization of tokamak operation, and between-pulse control design.

Control-oriented tokamak simulator (COTSIM)↗

Quantifying Uncertainty in HPC Job Queue Time Predictions

High Performance Computing (HPC) has developed at an unprecedented pace in recent decades. This growth has demanded corresponding development in the area of HPC Operational Data Analytics (ODA), which encompasses a wide range of data analysis techniques, ML/AI efforts, tools, and visualizations. Published studies in ODA offer a variety of practical ways to inform HPC users, administrators, procurement managers, and other stakeholders. Uncertainty analysis, however, is rare in the related published literature. For instance, we identify only 1 out of 14 existing studies focused on job queue time prediction that investigates the uncertainty aspect of their proposed predictions. We recognize the utmost importance uncertainty quantification can have in such predictive analytics solutions, with consequences in how users interpret information they receive, and attempt to bridge this gap. With the goal of improving access to such insights, we develop a process for determining upper and lower bounds of the predicted queue times of a regression model at a specified confidence level. Our current research is focused on the uncertainty in predicting job queue times, yet our approach may be employed in predicting other metrics.

HPC↗

Prediction of non-intuitive metabolic targets with bayesian metabolic control analysis to improve 3-hydroxypropionic acid production in Aspergillus niger

Development of efficient bioconversion processes is limited by the ability to predictably improve metabolic flux. Here we deployed Bayesian Metabolic Control Analysis as a platform to integrate multi-omics data with metabolic modeling and evaluated its ability to predict genetic interventions that improve metabolic flux. Global Metabolomics and proteomics data was collected from 17 Aspergillus niger strains engineered to produce the platform biochemical 3-hydroxypropionic acid from which seven actional genetic interventions were predicted from significant flux control coefficients. Of the suggested genetic interventions, two were present within the intuitively designed strains used for training (malonic semialdehyde dehydrogenase and pyruvate carboxylase) while five predicted targets were present within non-intuitive areas of the metabolic network including 5-formyltetrahydrofolate deformylase and four mitochondrial enzymes, alcohol dehydrogenase, succinyl-CoA ligase, aspartate aminotransferase, and malate dehydrogenase. Six of the targets were validated in the highest performing 3-HP strain used for multi-omics data generation which contained a prior disruption of the highest scoring target malonic semialdehyde dehydrogenase. Predicted directional perturbation of five of the six tested targets significantly improved titer and rate of 3-HP production and two significantly improved yield. The greatest improvements were observed following disruption of the non-intuitive target succinyl-CoA ligase which increased titer by 39% and yield by 29% (to 20.4 g/L 3-HP and 0.31 g 3-HP/g glucose) over the strains used for training. This study demonstrates the utility of Bayesian Metabolic Control Analysis and highlights the ability to predict meaningful genetic targets in unexpected areas of metabolism to improve engineered strains for bioconversion.

3-hydroxypropionic acid↗

Deposition Height Prediction in Directed Energy Deposition

Using 316L stainless steel as a model material, reduced-order models are developed to predict capture efficiency, deposition height, and site-specific hardness in directed energy deposition. Capture efficiency is predicted over a 15 to 55 pct range using a dimensionless number derived from processing conditions and thermophysical properties. Deposition height is predicted over a 0.3 to 1.3 mm range without in situ sensing or prior training data, using two models based on the same mass and energy-balance principles. Predictions are compared with machine learning approaches. A quantitative relationship links deposition height, primary dendrite arm spacing (PDAS), and hardness: heights of 0.3 to 1.1 mm correspond to PDAS values of 2.7 to 5.1 µm and Vickers hardness (HV) of 160 to 219. Thinner layers cool more rapidly, producing finer microstructures and higher hardness. Samples fabricated with in situ variations in deposition height exhibited up to 55 HV differences between thick and thin regions, demonstrating that local control of deposition height enables predictive, site-specific hardness within a single build. These results establish deposition height prediction as a pathway for a priori process design and property control in directed energy deposition for 316L stainless steel.

Kunkel, William [Univ. of Wisconsin, Madison, WI (↗

Thermodynamics and its prediction and CALPHAD modeling: Review, state of the art, and perspectives

Thermodynamics is a science concerning the state of a system, whether it is stable, metastable, or unstable, when interacting with its surroundings. The combined law of thermodynamics derived by Gibbs about 150 years ago laid the foundation of thermodynamics. In Gibbs combined law, the entropy production due to internal processes was not included, and the 2nd law was thus practically removed from the Gibbs combined law, so it is only applicable to systems under equilibrium, thus commonly termed as equilibrium or Gibbs thermodynamics. Gibbs further derived the classical statistical thermodynamics in terms of the probability of configurations in a system in the later 1800's and early 1900's. With the quantum mechanics (QM) developed in 1920's, the QM-based statistical thermodynamics was established and connected to classical statistical thermodynamics at the classical limit as shown by Landau in the 1940's. In 1960's the development of density functional theory (DFT) by Kohn and co-workers enabled the QM prediction of properties of the ground state of a system. On the other hand, the entropy production due to internal processes in non-equilibrium systems was studied separately by Onsager in 1930's and Prigogine and co-workers in the 1950's. In 1960's to 1970's the digitization of thermodynamics was developed by Kaufman in the framework of the CALculation of PHAse Diagrams (CALPHAD) modeling of individual phases with internal degrees of freedom. CALPHAD modeling of thermodynamics and atomic transport properties has enabled computational design of complex materials in the last 50 years. Our recently termed zentropy theory integrates DFT and statistical mechanics through the replacement of the internal energy of each individual configuration by its DFT-predicted free energy. The zentropy theory is capable of accurately predicting the free energy of individual phases, transition temperatures and properties of magnetic and ferroelectric materials with free energies of individual configurations solely from DFT-based calculations and without fitting parameters, and is being tested for other phenomena including superconductivity, quantum criticality, and black holes. Those predictions include the singularity at critical points with divergence of physical properties, negative thermal expansion, and the strongly correlated physics. Furthermore, those individual configurations may thus be considered as the genomic building blocks of individual phases in the spirit of the materials genome®. This has the potential to shift the paradigm of CALPHAD modeling from being heavily dependent on experimental inputs to becoming fully predictive with inputs solely from DFT-based calculations and machine learning models built on those calculations and existing experimental data through newly developed and future open-source tools. Furthermore, through the combined law of thermodynamics including the internal entropy production, it is shown that the kinetic coefficient matrix of independent internal processes is diagonal with respect to the conjugate potentials in the combined law, and the cross phenomena that the phenomenological Onsager flux and reciprocal relationships are due to the dependence of the conjugate potential of a molar quantity on nonconjugate molar quantities and other potentials, which can be predicted by the zentropy theory and CALPHAD modeling.

42 ENGINEERING↗

Transferable predictions of energetic and structural properties for refractory solid solution alloys across chemical compositions

We present a data-efficient approach to train graph neural networks (GNNs) on density functional theory (DFT) data for accurate and transferable predictions of energetic and structural properties of refractory solid solution alloys in the niobium-tantalum-vanadium (Nb-Ta-V) chemical space. We start by training the GNN model only on DFT data that describes refractory binary alloys niobium-tantalum (Nb-Ta), niobium-vanadium (Nb-V), and tantalum-vanadium (Ta-V) to predict formation enthalpy and root mean squared displacement. Once trained, the GNN predictions are tested on DFT data describing refractory ternary alloys Nb-Ta-V. While, unsurprisingly, direct transferability from binary to ternary is not sufficiently accurate, augmenting the training with only 1% of the available ternary data (uniformly distributed across the entire range of chemical compositions) improves significantly the quality of the GNN predictions. For comparison, we assess the transferability in the opposite direction by training GNN models on ternary Nb-Ta-V data and making predictions on binaries Nb-Ta, Nb-V, and Ta-V, which exhibits notably higher predictive errors. The proposed methodology, which favors transferability from lower-component to higher-component alloys, offers an efficient path towards avoiding the curse of dimensionality incurred when collecting DFT data for discovery and design of multi-component disordered alloys.

Density functional theory calculations↗