Physiochemical machine learning models predict operational lifetimes of CH 3 NH 3 PbI 3 perovskite solar cells
First machine learning predictions of perovskite solar cell service lifetimes.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
First machine learning predictions of perovskite solar cell service lifetimes.
Our objective is to apply machine learning (ML) algorithms for the prediction of molecular catalysis descriptors from geometric properties derived from experimental crystallographic databases. Catalysis is often considered a “low-data” discipline that is poorly suited for ML methods. An exception is the extensive structural information that is available for molecular catalysts through the Cambridge Structural Database (CSD), which contains atomically precise molecular structures from X-ray diffraction analysis for >600K metal complexes. As a proof-of-principle, we targeted the prediction of hydricity, a thermodynamic property that provides understanding and control of catalytic hydride transfer. We built a training set composed of ~100 molecular complexes with a known hydricity and structural information from the CSD. This data set was converted into a machine-readable format using the smooth overlap of atomic positions (SOAP) representation and further labeled with simple electronic descriptors for the metal centers. Multiple different neural networks were trained on this data set, and the accuracy of the hydricity predictions ranged from < 2 kcal/mol to 20 kcal/mol. The accuracy of each model was highly sensitive to which compounds were in the train versus test set, underscoring the challenges associated with small and chemically diverse data sets. Finally, to further augment the data set, we attempted to experimentally measure several new hydricity values, however these experiments were unsuccessful due to undesired chemical reactivity of the selected complexes.
Figure data for the paper 'Machine learning enhanced predictions of ICRF heating: overcoming numerical limitations via data curation' accepted in Physics of Plasmas, DOI TBD
Machine learning (ML) has been gaining interest in the metabolic engineering community as a means to automate prediction tasks. In this work, we introduce and study the task of using ML to recommend high-fitness triplet mutants as candidates for wet-lab experiments. We first utilize individual fitness and digenic fitness scores as features and train machine learning models that produce a ranked list, from high to low fitness scores, for triplet gene mutants of S. cerevisiae. Then, we incorporate prior metabolic knowledge from an existing gene ontology, by designing a novel graph representation and deducing features that can capture gene similarity and gene interactions. Lastly, experimental results show that our proposed gene ontology enriched model, termed TriGORank, improves both performance and explainability.
This paper presents a physics-informed machine learning (ML) framework to construct reduced-order models (ROMs) for reactive-transport quantities of interest (QoIs) based on high-fidelity numerical simu-lations. QoIs include species decay, product yield, and degree of mixing. The ROMs for QoIs are applied to quantify and understand how the chemical species evolve over time. First, high-resolution datasets for constructing ROMs are generated by solving anisotropic reaction-di?usion equations using a non-negative finite element formulation for di?erent input parameters. The reactive-mixing model input parameters are: time-scale associated with flipping of velocity, spatial-scale controlling small/large vortex structures of velocity, perturbation parameter of the vortex-based velocity, anisotropic dispersion strength/contrast, and molecular diffusion. Second, random forests, F-test, and mutual information criterion are used to evaluate the importance of model inputs/features with respect to QoIs. We observed that anisotropic dispersion strength/contrast is the most important feature and time-scale associated with flipping of velocity is the least important feature. Third, Support Vector Machines (SVM) and Support Vector Regression (SVR) are used to construct ROMs based on the model inputs. The constructed SVR-ROMs are then used to predict scaling of QoIs. We also present estimates and inequalities on the QoIs, which inform that the species decay, mix, and produce in an exponential fashion. These inequalities also inform that a radial basis function is the most suitable kernel for the SVM/SVR models for QoIs. It is observed that R2-score for SVR-ROMs on unseen data is greater than 0.9, implying that the SVR-ROMs are able to predict the reaction-diffusion system state reasonably well. Finally, in terms of the computational cost, the proposed SVM-ROMs are O(107) times faster than running a high-fidelity finite element simulation for evaluating QoIs. This makes the proposed ML-based ROMs attractive for reactive-transport sensing and real-time monitoring applications as they are significantly faster yet reasonably accurate.
Operation and maintenance (O&M) costs for wind turbines pose a risk to competitiveness and asset owners. With machine-learning technologies and digitalization rapidly maturing, the wind industry is actively investigating these new technologies to optimize O&M practices and reduce costs. This paper reviews recent work on machine-learning approaches to generator bearing failure prediction and presents a relevant real-world case study through a collaboration between the National Renewable Energy Laboratory and Envision Digital Corporation. In the case study, we evaluate the performance of representative machine-learning algorithms for predicting wind turbine generator bearing failures. Operational supervisory control and data acquisition data from one wind power plant was used to train and test the machine-learning models. The investigated data channels are chosen based on whether physically they reflect the failed generator bearing conditions and the component historical usage, including both environmental and operational conditions. Benefits and drawbacks of different methods are identified.
Machine learning has proven to be an invaluable tool for characterizing the stability of planets in simplified planetary systems. In this work, we investigate the performance of a machine learning classifier on tightlypacked systems containing a rich diversity of planets, from Earths to Jupiters. Using information derived from short numerical simulations about a planet’s early orbital evolution and its relationship with the most massive planets in the system, we train a random forest classifier to predict instability with a > 88 percent accuracy. Our classifier relies on relative planet masses and the standard deviation of eccentricity for much of its predictive power. Most misclassified planets lie along a multi-dimensional boundary between stable and unstable planets, indicating that their early orbital evolution is ambiguous. The major reason for misclassification in this work is timescale: because our classifier uses information from only the first 137 years of simulation data, it is blind to late time interactions that cause or prevent instability. Machine learning methods like those utilized in this work provide powerful tools to complement numerical simulations across a wide range of planetary architectures.
A major obstacle for machine learning (ML) in chemical science is the lack of physically informed feature representations that provide both accurate prediction and easy interpretability of the ML model. In this work, we describe adsorption systems using novel two-dimensional energy histogram (2D-EH) features, which are obtained from the probe-adsorbent energies and energy gradients at grid points located throughout the adsorbent. The 2D-EH features encode both energetic and structural information of the material and lead to highly accurate ML models (coefficient of determination R2 ~ 0.94–0.99) for predicting single-component adsorption capacity in metal–organic frameworks (MOFs). Here, we consider the adsorption of spherical molecules (Kr and Xe), linear alkanes with a wide range of aspect ratios (ethane, propane, n-butane, and n-hexane), and a branched alkane (2,2-dimethylbutane) over a wide range of temperatures and pressures. The interpretable 2D-EH features enable the ML model to learn the basic physics of adsorption in pores from the training data. We show that these MOF-data-trained ML models are transferrable to different families of amorphous nanoporous materials. We also identify several adsorption systems where capillary condensation occurs, and ML predictions are more challenging. Nevertheless, our 2D-EH features still outperform structural features including those derived from persistent homology. The novel 2D-EH features may help accelerate the discovery and design of advanced nanoporous materials using ML for gas storage and separation in the future.
Abstract Invasive plant pathogenic fungi have a global impact, with devastating economic and environmental effects on crops and forests. Biosurveillance, a critical component of threat mitigation, requires risk prediction based on fungal lifestyles and traits. Recent studies have revealed distinct genomic patterns associated with specific groups of plant pathogenic fungi. We sought to establish whether these phytopathogenic genomic patterns hold across diverse taxonomic and ecological groups from the Ascomycota and Basidiomycota, and furthermore, if those patterns can be used in a predictive capacity for biosurveillance. Using a supervised machine learning approach that integrates phylogenetic and genomic data, we analyzed 387 fungal genomes to test a proof-of-concept for the use of genomic signatures in predicting fungal phytopathogenic lifestyles and traits during biosurveillance activities. Our machine learning feature sets were derived from genome annotation data of carbohydrate-active enzymes (CAZymes), peptidases, secondary metabolite clusters (SMCs), transporters, and transcription factors. We found that machine learning could successfully predict fungal lifestyles and traits across taxonomic groups, with the best predictive performance coming from feature sets comprising CAZyme, peptidase, and SMC data. While phylogeny was an important component in most predictions, the inclusion of genomic data improved prediction performance for every lifestyle and trait tested. Plant pathogenicity was one of the best-predicted traits, showing the promise of predictive genomics for biosurveillance applications. Furthermore, our machine learning approach revealed expansions in the number of genes from specific CAZyme and peptidase families in the genomes of plant pathogens compared to non-phytopathogenic genomes (saprotrophs, endo- and ectomycorrhizal fungi). Such genomic feature profiles give insight into the evolution of fungal phytopathogenicity and could be useful to predict the risks of unknown fungi in future biosurveillance activities.
Machine learning can be used to predict fault properties such as shear stress, friction, and time to failure using continuous records of fault zone acoustic emissions. The files are extracted features and labels from lab data (experiment p4679). The features are extracted with a non-overlapping window from the original acoustic data. The first column is the time of the window. The second and third columns are the mean and the variance of the acoustic data in this window, respectively. The 4th-11th column is the the power spectrum density ranging from low to high frequency. And the last column is the corresponding label (shear stress level). The name of the file means which driving velocity the sequence is generated from. Data were generated from laboratory friction experiments conducted with a biaxial shear apparatus. Experiments were conducted in the double direct shear configuration in which two fault zones are sheared between three rigid forcing blocks. Our samples consisted of two 5-mm-thick layers of simulated fault gouge with a nominal contact area of 10 by 10 cm^2. Gouge material consisted of soda-lime glass beads with initial particle size between 105 and 149 micrometers. Prior to shearing, we impose a constant fault normal stress of 2 MPa using a servo-controlled load-feedback mechanism and allow the sample to compact. Once the sample has reached a constant layer thickness, the central block is driven down at constant rate of 10 micrometers per second. In tandem, we collect an AE signal continuously at 4 MHz from a piezoceramic sensor embedded in a steel forcing block about 22 mm from the gouge layer The data from this experiment can be used with the deep learning algorithm to train it for future fault property prediction.
The empirical rules for the prediction of solid solution formation proposed so far in the literature usually have very compromised predictability. Some rules with seemingly good predictability were, however, tested using small data sets. Based on an unprecedented large dataset containing 1252 multicomponent alloys, machine-learning methods showed that the formation of solid solutions can be very accurately predicted (93%). The machine-learning results help identify the most important features, such as molar volume, bulk modulus, and melting temperature. As such a new thermodynamics-based rule was developed to predict solid–solution alloys. The new rule is nonetheless slightly less accurate (73%) but has roots in the physical nature of the problem. The new rule is employed to predict solid solutions existing in the three blocks, each of which consists of 9 elements. The predictions encompass face-centered cubic (FCC), body-centered cubic (BCC), and hexagonal closest packed (HCP) structures in a high throughput manner. The validity of the prediction is further confirmed by CALculations of PHAse Diagram (CALPHAD) calculations with high consistency (94%). Since the new thermodynamics-based rule employs only elemental properties, applicability in screening for solid solution high-entropy alloys is straightforward and efficient.
Abstract In this study, we evaluated the performance of machine learning (ML) models (XGBoost) in predicting low‐cloud fraction (LCF), compared to two generations of the community atmospheric model (CAM5 and CAM6) and ERA5 reanalysis data, each having a different cloud scheme. ML models show a substantial enhancement in predicting LCF regarding root mean squared errors and correlation coefficients. The good performance is consistent across the full spectrums of atmospheric stability and large‐scale vertical velocity. Employing an explainable ML approach, we revealed the importance of including the amount of available moisture in ML models for representing spatiotemporal variations in LCF in the midlatitudes. Also, ML models demonstrated marked improvement in capturing the LCF variations during the stratocumulus‐to‐cumulus transition (SCT). This study suggests ML models' great potential to address the longstanding issues of “too few” low clouds and “too rapid” SCT in global climate models.
Abstract High-temperature alloy design requires a concurrent consideration of multiple mechanisms at different length scales. We propose a workflow that couples highly relevant physics into machine learning (ML) to predict properties of complex high-temperature alloys with an example of the 9–12 wt% Cr steels yield strength. We have incorporated synthetic alloy features that capture microstructure and phase transformations into the dataset. Identified high impact features that affect yield strength of 9Cr from correlation analysis agree well with the generally accepted strengthening mechanism. As a part of the verification process, the consistency of sub-datasets has been extensively evaluated with respect to temperature and then refined for the boundary conditions of trained ML models. The predicted yield strength of 9Cr steels using the ML models is in excellent agreement with experiments. The current approach introduces physically meaningful constraints in interrogating the trained ML models to predict properties of hypothetical alloys when applied to data-driven materials.
This study describes accurate, computationally efficient models that can be implemented for practical use in predicting frost events for point-scale agricultural applications. Frost damage in agriculture is a costly burden to farmers and global food security alike. Timely prediction of frost events is important to reduce the cost of agricultural frost damage and traditional numerical weather forecasts are often inaccurate at the field-scale in complex terrain. In this paper, we developed machine learning (ML) algorithms for the prediction of such frost events near Alcalde, NM at the point-scale. ML algorithms investigated include deep neural network, convolution neural networks, and random forest models at lead-times of 6–48 h. Our results show promising accuracy (6-h prediction RMSE = 1.53–1.72°C) for use in frost and minimum temperature prediction applications. Seasonal differences in model predictions resulted in a slight negative bias during Spring and Summer months and a positive bias in Fall and Winter months. Additionally, we tested the model transferability by continuing training and testing using data from sensors at a nearby farm. We calculated the feature importance of the random forest models and were able to determine which parameters provided the models with the most useful information for predictions. We determined that soil temperature is a key parameter in longer term predictions (>24 h), while other temperature related parameters provide the majority of information for shorter term predictions. The model error compared favorable to previous ML based frost studies and outperformed the physically based High Resolution Rapid Refresh forecasting system making our ML-models attractive for deployment toward real-time monitoring of frost events and damage at commercial farming operations.
This study explores the use of machine learning (ML) as a computational tool to accelerate the design of multi-principal element alloys (MPEAs) with improved tensile elongation. An ML model was trained using available experimental data from the literature along with theoretically derived features to predict yield strength (YS) and ductility. A subset of ML-predicted compositions—CoNiVFe, CoNiVTi, CoNiVTiFe, and CoCrNiVTi—was synthesized and evaluated through tensile testing. The ML model underpredicted YS by approximately 20–30 % and overpredicted ductility by 60–70 % for Ti-containing alloys. Microstructural analysis revealed that Ti segregation at interdendritic regions contributed to early fracture, leading to discrepancies in ductility predictions. Ti segregation at these regions likely drives the increased YS due to segregation strengthening. In contrast, the CoNiVFe alloy showed good agreement with both experimental YS and elongation, with prediction errors of ∼10.2 % and ∼20.7 %, respectively. Microstructural characterization revealed minimal segregation in this alloy, suggesting that the ML model can reliably predict the properties of alloys with little to no segregation. These findings highlight the capability of ML in predicting YS with good accuracy but underscore its limitations in capturing defect-driven failure mechanisms such as segregation-induced embrittlement.
Atmospheric corrosion of metallic parts is a widespread materials degradation phenomena that is challenging to predict given its dependence on many factors (e.g. environmental, physiochemical, and part geometry). For materials with long expected service lives, accurately predicting the degree to which corrosion will degrade part performance is especially difficult due to the stochastic nature of corrosion damage spread across years or decades of service. The Finite Element Method (FEM) is a computational technique capable of providing accurate estimates of corrosion rate by numerically solving complex differential Eqs. characterizing this phenomena. Nevertheless, given the iterative nature of FEM and the computational expense required to solve these complex equations, FEM is ill-equipped for an efficient exploration of the design space to identify factors that accelerate or deter corrosion, despite its accuracy. In this work, a machine learning based surrogate model capable of providing accurate predictions of corrosion with significant computational savings is introduced. Specifically, this work leverages AdaBoosted Decision trees to provide an accurate estimate of corrosion current per width given different values of temperature, water layer thickness, molarity of the solution, and the length of the cathode for a galvanic couple of aluminum and stainless steel.
Abstract A grand challenge of materials science is predicting synthesis pathways for novel compounds. Data-driven approaches have made significant progress in predicting a compound’s synthesizability; however, some recent attempts ignore phase stability information. Here, we combine thermodynamic stability calculated using density functional theory with composition-based features to train a machine learning model that predicts a material’s synthesizability. Our model predicts the synthesizability of ternary 1:1:1 compositions in the half-Heusler structure, achieving a cross-validated precision of 0.82 and recall of 0.82. Our model shows improvement in predicting non-half-Heuslers compared to a previous study’s model, and identifies 121 synthesizable candidates out of 4141 unreported ternary compositions. More notably, 39 stable compositions are predicted unsynthesizable while 62 unstable compositions are predicted synthesizable; these findings otherwise cannot be made using density functional theory stability alone. This study presents a new approach for accurately predicting synthesizability, and identifies new half-Heuslers for experimental synthesis.
Thermally anisotropic building envelope (TABE) is a novel active building envelope that can save energy use to maintain thermal comfort in buildings by redirecting heat and coolness from building envelopes to thermal loops. Finite element models (FEMs) can be used to compute the heat fluxes through TABEs, but the high computational cost of finite element simulations has prevented parametric studies and design optimizations. This paper proposes a domain knowledge–informed, finite element–based machine learning framework to reduce the computation cost for the energy management of buildings installed with TABE that uses a ground thermal loop. First, the training heat flux data set was generated by FEM simulations with different thermal loop schedules. Then, both shallow learning models (i.e., multivariate linear regression and eXtreme Gradient Boost, or XGBoost) and a deep learning model (i.e., deep neural network, or DNN) were trained to predict the heat fluxes. Domain knowledge was used for data preprocessing and feature selection. Finally, the suitability of the selected machine learning model was tested under different thermal loop schedules. Herein, the case study results showed that: (1) XGBoost can be as accurate as DNN (coefficient of determination equal to 0.81) with much less training time; (2) the annual energy cost savings for different thermal loop schedules obtained by the XGBoost-predicted and FEM-calculated heat fluxes are consistent, having a difference of only 4%; and (3) XGBoost can reduce the computation time for the annual energy analysis of the case study building with a given thermal loop schedule from around 12 h by using FEM to less than 1 min.