Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “functional principal component analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Global teleconnections influencing large-scale drought in the United States using SVDI

Understanding recent large-scale drought patterns and the mechanisms producing extreme drought events is vital for future drought forecasts and understanding future drought risks. Increasingly, vapor pressure deficit (VPD) has been used as an important measure of evaporative demand and proxy for drought detection. In this study, VPD is used to calculate the new Standardized VPD Drought Index (SVDI) with NASA North American Land Data Assimilation System (NLDAS) data. Previous studies have shown that SVDI accurately identifies the timing and magnitude short-term droughts in the United States (U.S). In the present study, SVDI is now used to identify large-scale drought patterns between 1980 and 2021 and drought variability driven by selected global teleconnections originating in the Pacific and Atlantic Oceans. Spatial drought characteristics were extracted from SVDI using empirical orthogonal function (EOF) analysis. Then a k-means clustering algorithm was applied to both EOF principal components and primary teleconnections, including the El Nino-Southern Oscillation (ENSO) and Pacific Decadal Oscillation (PDO) to identify drought events driven by the Pacific Ocean. Results show that the SVDI is useful in evaluating large-scale drought variability in the U.S. related to global teleconnections, and that mechanisms influencing summer drought patterns in the Western and Southwestern U.S. are driven by a tropical-extratropical interactions originating in the equatorial Pacific Ocean related to ENSO dynamics with interdecadal variability modulated by PDO. The large-scale droughts in the Central and Southern U.S., like those in 2011 and 2012, on the other hand, are driven by the North Pacific Ocean warm pool during a strong negative PDO, which subsequently influenced variability in the Bermuda-Azores High in the Atlantic Ocean. In summer 2011, the Bermuda-Azores High weakened, reducing the onshore winds and moisture transport along the eastern Gulf of Mexico and contributing to ongoing drought in the region. The Northern Pacific and Atlantic Ocean sea surface temperatures (SSTs) have increased between 1980 and 2021. In conclusion, as SSTs continue to rise in the Northern Pacific Ocean, one consequence of the coupled North Pacific warm pool and atmospheric dynamics, is to increase summer drought variability over a large region in the southern and midwestern U.S. under global warming.

54 ENVIRONMENTAL SCIENCES↗

Structural features of xylan dictate reactivity and functionalization potential for bio-based materials

Plant-based materials have the potential to replace some petroleum-based products, offering compostability and biodegradability as critical advantages. Xylan-rich biomass sources are gaining recognition due to their abundance and underutilization in current industrial applications. Research of potential xylan applications has been complicated by the complex and heterogeneous structure that varies for different xylan feedstocks. Acylation is a broadly used reaction in functionalization of polysaccharides at an industrial scale. However, the efficiency of this reaction varies with the xylan source. To optimize xylan valorization, a systematic understanding of structure–reactivity relationships is essential. This study explores, characterizes, and compares various xylan feedstocks in the acylation process. Xylan feedstocks were analyzed for their chemical composition, degree of polymerization, branching, solubility, and presence of impurities. These features were correlated with xylan glycotypes’ reactivity toward functionalization with succinic anhydride in an optimized DMSO/KOH condition, achieving carboxyl contents of up to 1.46. We used principal component analysis and hierarchical clustering to identify key structural features of xylan that promote its reactivity. Our findings reveal that xylans with higher xylose content and lower degrees of branching exhibit enhanced reactivity, achieving higher carboxyl content and yields. Structural analyses confirmed successful modification, and light scattering analyses showed dramatic changes in the solution properties. Succinylation improves the solubility and film-forming properties of native xylans. This study shows key structure–reactivity relationships in xylan succinylation, establishing that low branching, high xylose content, and reduced lignin impurity enhance chemical functionalization. The results offer a framework for selecting optimal biomass feedstocks and support future efforts in genetic and synthetic biology to design plants with tunable xylan architectures. These findings advance the hemicellulose valorization for applications in coatings and packaging.

Acylation↗

Unraveling Fundamental Activity–Stability Relationships in Rutile Oxides

The oxygen evolution reaction (OER) is a key anodic half-cell reaction that accompanies several critical electrochemical reduction reactions of interest to a variety of applications. Despite steady advances in understanding and qualitatively predicting OER activity and selectivity trends, a comprehensive description or prediction of material aqueous (in)stability and degradation mechanisms remains elusive, even though these processes critically influence device lifetime and economic feasibility. In this work, we investigate the interplay, or lack thereof, between OER activity and material aqueous stability across rutile oxides, with a particular focus on iridium oxide (IrO 2 ). By applying a Born–Haber cycle, we calculate the thermodynamic driving force for metal dissolution as a function of the applied bias and electrolyte conditions. We apply interpretable machine learning techniques, including principal component analysis and symbolic regression, to analyze trends across rutile oxides and find that key thermodynamic descriptors for OER activity and surface stability are only very weakly correlated. Instead, the local atomic environment─especially electronic structure signatures for interactions between the active site and its neighbors─plays a more important role in predicting material stability. Leveraging these insights, we investigate the impact of doping IrO 2 with a range of transition metals and show that the stability of Ir active sites can be tuned largely independently of its predicted OER activity. These insights lay the foundation for material design to improve stability with respect to corrosion, with the ultimate aim to enhance long-term stability without sacrificing catalytic performance in the OER.

evolution reactions↗

Operando pair distribution function analysis of nanocrystalline functional materials: the case of TiO 2 -bronze nanocrystals in Li-ion battery electrodes

Structural modelling of operando pair distribution function (PDF) data of complex functional materials can be highly challenging. To aid the understanding of complex operando PDF data, this article demonstrates a toolbox for PDF analysis. The tools include denoising using principal component analysis together with the structureMining , similarityMapping and nmfMapping apps available through the online service `PDF in the cloud' ( PDFitc , https://pdfitc.org/). The toolbox is used for both ex situ and operando PDF data for 3 nm TiO 2 -bronze nanocrystals, which function as the active electrode material in a Li-ion battery. The tools enable structural modelling of the ex situ and operando PDF data, revealing two pristine TiO 2 phases (bronze and anatase) and two lithiated Li x TiO 2 phases (lithiated versions of bronze and anatase), and the phase evolution during galvanostatic cycling is characterized.

Chemistry↗

Reinforcement learning for real-time process control in high-temperature superconductor manufacturing

With high efficiency and low energy loss, high-temperature superconductors (HTS) have demonstrated their profound applications in various fields, such as medical imaging, transportation, accelerators, microwave devices, and power systems. The high-field applications of HTS tapes have raised the demand for producing cost-effective tapes with long lengths in superconductor manufacturing. However, achieving the uniform and enhanced performance of a long HTS tape is challenging due to the unstable growth conditions in the manufacturing process. Although it is confirmed that the process parameters during the advanced metal organic chemical vapor deposition (A-MOCVD) process influence the uniformity of the produced HTS tapes, the high-dimensional process parameter signals and their complicated interactions make it difficult to develop an effective control policy. In this paper, we propose a local measure for the uniformity of HTS tapes to provide instant feedback for our control policy. Then, we model the manufacturing of HTS tapes as a Markov decision process (MDP) with continuous state and action spaces to assess the instant reward in real time in our feedback control model. As our MDP involves continuous and high-dimensional state and action spaces, a neural fitted Q-iteration (NFQ) algorithm is adopted to solve the MDP with artificial neural network (ANN) function approximation. The collinearity of process parameters can restrict our capability of adjusting the process parameters, which is addressed by the principal component analysis (PCA) in our method. The control policy adjusts the PCA of process parameters using the NFQ algorithm. In conclusion, based on our case studies on real A-MOCVD dataset, the obtained control policy increases the average uniformity of tapes by 5.6% and performs especially well on sample HTS tapes with a low uniformity.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Decomposition of Irganox 1010 in plastic bonded explosives

Abstract Degradation pathways of Irganox 1010 in aged plastic bonded explosive (PBX) 9501 were investigated using ultrahigh performance liquid chromatography coupled to quadrupole time of flight mass spectrometry (UHPLC‐QTOF). Using a targeted approach, a total of 44 Irganox 1010 decomposition products were discovered. These decomposition products were formed through hydrolysis, scission, and/or oxidation of Irganox 1010. The hydrolytic decomposition of Irganox is a straightforward process resulting in the cleavage of the ester group(s) while oxidation and scission are more complicated and can happen at multiple locations on the Irganox 1010 molecule. Moreover, due to the symmetric nature of Irganox 1010, multiple decomposition reactions can occur. Indeed some decomposition products exhibited hydrolysis, oxidation, and scission. In order to probe any trends in the aged PBX 9501 samples, principal component analysis (PCA) was implemented. The greatest chemical differences between the aged PBX samples was hydrolysis of the ester functional groups on Irganox 1010. Despite the negative connotations of hydrolysis, the Irganox 1010 decomposition products are still able to function as a radical scavenger in PBX 9501 as intended.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Evaluation of CMIP6 GCMs Over the CONUS for Downscaling Studies

Despite the necessity of Global Climate Models (GCMs) sub-selection in downscaling studies, an objective approach for their selection is currently lacking. Building on the previously established concepts in GCMs evaluation frameworks, we develop a weighted averaging technique to remove the redundancy in the evaluation criteria and rank 37 GCMs from the sixth phase of the Coupled Models Intercomparison Project over the contiguous United States. GCMs are rated based on their average performance across 66 evaluation measures in the historical period (1981–2014) after each metric is weighted between zero and one, depending on its uniqueness. The robustness of the outcome is tested by repeating the process with the empirical orthogonal function analysis in which each GCM is ranked based on its sum of distances from the reference in the principal component space. The two methodologies work in contrasting ways to remove the metrics redundancy but eventually develop similar GCMs rankings. A disparity in GCMs' behavior related to their sensitivity to the size of the evaluation suite is observed, highlighting the need for comprehensive multi-variable GCMs evaluation at varying timescales for determining their skillfulness over a region. The sub-selection goal is to use a representative set of skillful models over the region of interest without substantial overlap in their future climate responses and modeling errors in representing historical climate. Additional analyses of GCMs' independence and spread in their future projections provide the necessary information to objectively select GCMs while keeping all aspects of necessity in view.

54 ENVIRONMENTAL SCIENCES↗

Accelerating the Structure Exploration of Diverse Bi–Pt Nanoclusters via Physics‐Informed Machine Learning Potential and Particle Swarm Optimization

Bimetallic Bi–Pt nanoclusters exhibit diverse structural motifs, including core-shell, Janus, and mixed alloy configurations, due to the unique bonding characteristics between Bi and Pt atoms. Using density functional theory refinements from ChIMES physically machine-learned potential and CALYPSO particle swarm optimization global searches, 34 Bi20-Pt20 nanoclusters are systematically classified. The results reveal that Bi atoms predominantly occupy surface sites, driven by charge transfer effects. Cohesive energy trends alone prove insufficient for structure differentiation, necessitating a data-driven approach employing principal component analysis and K-means clustering. Furthermore, vibrational, electronic, and infrared spectral analyses provide additional insights into structure-property relationships. The findings offer an original framework for the automated classification and analysis of bimetallic nanoclusters, enhancing the understanding of their stability and functional properties.

bimetallic nanoparticles↗

Mallat Scattering Transformation based surrogate for Magnetohydrodynamics

Abstract A Machine and Deep Learning (MLDL) methodology is developed and applied to give a high fidelity, fast surrogate for 2D resistive MagnetoHydroDynamic (MHD) simulations of Magnetic Liner Inertial Fusion (MagLIF) implosions. The resistive MHD code is used to generate an ensemble of implosions with different liner aspect ratios, initial gas preheat temperatures (that is, different adiabats), and different liner perturbations. The liner density and magnetic field as functions of x , y , and z were generated. The Mallat Scattering Transformation (MST) is taken of the logarithm of both fields and a Principal Components Analysis (PCA) is done on the logarithm of the MST of both fields. The fields are projected onto the PCA vectors and a small number of these PCA vector components are kept. Singular Value Decompositions of the cross correlation of the input parameters to the output logarithm of the MST of the fields, and of the cross correlation of the SVD vector components to the PCA vector components are done. This allows the identification of the PCA vectors vis-a-vis the input parameters. Finally, a Multi Layer Perceptron (MLP) neural network with ReLU activation and a simple three layer encoder/decoder architecture is trained on this dataset to predict the PCA vector components of the fields as a function of time. Details of the implosion, stagnation, and the disassembly are well captured. Examination of the PCA vectors and a permutation importance analysis of the MLP show definitive evidence of an inverse turbulent cascade into a dipole emergent behavior. The orientation of the dipole is set by the initial liner perturbation. The analysis is repeated with a version of the MST which includes phase, called Wavelet Phase Harmonics (WPH). While WPH do not give the physical insight of the MST, they can and are inverted to give field configurations as a function of time, including field-to-field correlations.

97 MATHEMATICS AND COMPUTING↗

Characterizing different motility-induced regimes in active matter with machine learning and noise

Here we examine motility-induced phase separation (MIPS) in two-dimensional run-and-tumble disk systems using both machine learning and noise fluctuation analysis. Our measures suggest that within the MIPS state there are several distinct regimes as a function of density and run time, so that systems with MIPS transitions exhibit an active fluid, an active crystal, and a critical regime. The different regimes can be detected by combining an order parameter extracted from principal component analysis with a cluster stability measurement. The principal component-derived order parameter is maximized in the critical regime, remains low in the active fluid, and has an intermediate value in the active crystal regime. We demonstrate that machine learning can better capture dynamical properties of the MIPS regimes compared to more standard structural measures such as the maximum cluster size. The different regimes can also be characterized via changes in the noise power of the fluctuations in the average speed. In the critical regime, the noise power passes through a maximum and has a broad spectrum with a 1/f 1.6 signature, similar to the noise observed near depinning transitions or for solids undergoing plastic deformation.

97 MATHEMATICS AND COMPUTING↗

MODE: A Web Application for Interactive Visualization and Exploration of Omics Data

Studies generating transcriptomics, proteomics, lipidomics, and metabolomics (colloquially referred to as “omics”) data allow researchers to find biomarkers or molecular targets, or understand complex biological structures and functions by identifying changes in biomolecule abundance and expression between experimental conditions. Omics data is multi-dimensional and oftentimes summarization techniques such as principal component analysis (PCA) are used to identify high-level patterns in data. Though useful, these summaries don’t allow exploration of detailed patterns in omics data that may have biological relevance. The use of interactive HTML displays with plots allows researchers to interact with omics data at a detailed level, but building these displays requires significant coding expertise. To overcome this barrier, the software MODE was built to empower users to build their own interactive HTML displays to support scientific discovery. These displays are easily shareable, do not depend on a specific operating system, and allow users to effortlessly sort and filter plots by categorical or numerical variables. MODE allows users to build and share these displays with several options for plot design and meta selection. In conclusion, the MODE web application and its capabilities are presented and then demonstrated on lipidomics data from a leaf wounding study.

lipidomics↗

Novel Cell-Type-Specific Drought-Responsive Proteins in Root Tips of Field-Grown Perennial Switchgrass

The root-tip region of plants, including the root cap, forms the most basal terminal of the root and exhibits a high degree of cellular complexity in terms of morphology, cytological function, and interaction with environmental cues in the soil. Cells in this region follow a developmental trajectory, transitioning from stem cells to meristematic cells, and ultimately to fully differentiated cell types. However, our understanding of root-tip cell-type specific proteomic responses to abiotic stresses, such as drought, particularly under field conditions, remains limited. This study aimed to identify spatially resolved, cell type-specific proteomes in switchgrass (Panicum virgatum) root tips under drought stress. Root tips were collected from seven-year-old, field-grown switchgrass ‘Alamo’ plants excavated under both well-watered and long-term drought conditions. Cell type-specific proteins were identified using laser capture microdissection (LCM) coupled with nanoPOTS (Nanodroplet Processing in One Pot for Trace Samples) and nano-LC-MS proteomics analysis. Five distinct cell types were targeted: (1) cells in the quiescent center and stem cell niche (QuC), (2) protodermal epidermal cells (PEC) in the meristematic zone, (3) epidermal cells in the transition and elongation zones above the root cap (Epi), (4) peripheral root cap cells (PRC), forming 2–3 layers below the PEC and 1–2 layers above the root border cells, and (5) columella root cap cells (Col) comprising of the columella initials and a single underlying layer of cells undergoing active growth. Principal component analysis (PCA) revealed clear separation among the five targeted cell types, confirming distinct proteomic profiles. Proteins predominantly enriched in each cell type were linked to distinct cellular functions, with QuC cells showing involvement in chromosomal behavior, DNA replication, and mitosis—key processes for stem cell niche regulation. Drought stress resulted in alterations of proteostasis, as evidenced by significant decreases in ribosomal proteins and increases in protein synthesis inhibitors. Moreover, drought stress induced unique cell-type–specific proteins involved in phytohormone biosynthesis and signaling pathways, including auxin, cytokinin, and jasmonic acid. In particular, QuC cells were more highly enriched in proteins associated with DNA repair and mitotic processes. Metabolic pathways related to amino acids, carbohydrates, and lipids were differentially affected in a cell-type–dependent manner, whereas general stress-responsive proteins exhibited consistent changes across all five cell types. Overall, this study provides unique spatially resolved, cell-type-specific proteomic profiles in root tips, representing a significant advancement in our understanding of the cellular mechanisms underlying plant responses to drought stress in natural field conditions.

perennial grass↗

Training and projecting: A reduced basis method emulator for many-body physics

Here, we present the reduced basis method as a tool for developing emulators for equations with tun able parameters within the context of the nuclear many-body problem. The method uses a basis expansion informed by a set of solutions for a few values of the model parameters and then projects the equations over a well-chosen low-dimensional subspace. We connect some of the results in the eigenvector continuation literature to the formalism of reduced basis methods and show how these methods can be applied to a broad set of problems. As we illustrate, the possible success of the formalism on such problems can be diagnosed beforehand by a principal component analysis. We apply the reduced basis method to the one-dimensional Gross-Pitaevskii equation with a harmonic trap ping potential and to nuclear density functional theory for 48 Ca, achieving speed-ups of more than x150 and x250, respectively, when compared to traditional solvers. The outstanding performance of the approach, together with its straightforward implementation, show promise for its application to the emulation of computationally demanding calculations, including uncertainty quantification.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Dimensional Reduction for Sampled Priors and Application to Photometric Redshift Distributions

A typical Bayesian inference on the values of some parameters of interest q from some data D involves running a Markov Chain (MC) to sample from the posterior $p$($q$,$n$|$D$) $\propto$ $\mathcal{L}$($D$|$q$,$n$)$p$(q)$p$($n$), where n are some nuisance parameters with a separable prior. In some cases, the nuisance parameters are high-dimensional, and their prior p(n) is itself defined only by a set of samples that have been drawn from some other MC. The MC for the posterior will typically require evaluation of p(n) at arbitrary values of n, i.e., one needs to provide a density estimator over the full n space from the provided samples. But the high dimensionality of n hinders both the density estimation and the efficiency of the MC for the posterior. We describe a solution to this problem: a linear compression of the n space into a much lower-dimensional space u, which projects away directions in n space that cannot appreciably alter $\mathcal{L}$. The algorithm for doing so is a slight modification to principal components analysis, and is less restrictive on p(n) than other proposed solutions to this issue. We demonstrate this “mode projection” technique using the analysis of 2-point correlation functions of weak lensing fields and galaxy density in the Dark Energy Survey, where n is a binned representation of the redshift distribution n(z) of the galaxies.

79 ASTRONOMY AND ASTROPHYSICS↗

Exploring the dependence of gas cooling and heating functions on the incident radiation field with machine learning

ABSTRACT Gas cooling and heating functions play a crucial role in galaxy formation. But, it is computationally expensive to exactly compute these functions in the presence of an incident radiation field. These computations can be greatly sped up by using interpolation tables of pre-computed values, at the expense of making significant and sometimes even unjustified approximations. Here, we explore the capacity of machine learning to approximate cooling and heating functions with a generalized radiation field. Specifically, we use the machine learning algorithm XGBoost to predict cooling and heating functions calculated with the photoionization code cloudy at fixed metallicity, using different combinations of photoionization rates as features. We perform a constrained quadratic fit in metallicity to enable a fair comparison with traditional interpolation methods at arbitrary metallicity. We consider the relative importance of various photoionization rates through both a principal component analysis (PCA) and calculation of SHapley Additive exPlanation (shap) values for our XGBoost models. We use feature importance information to select different subsets of rates to use in model training. Our XGBoost models outperform a traditional interpolation approach at each fixed metallicity, regardless of feature selection. At arbitrary metallicity, we are able to reduce the frequency of the largest cooling and heating function errors compared to an interpolation table. We find that the primary bottleneck to increasing accuracy lies in accurately capturing the metallicity dependence. This study demonstrates the potential of machine learning methods such as XGBoost to capture the non-linear behaviour of cooling and heating functions.

79 ASTRONOMY AND ASTROPHYSICS↗

The relationship between below average cognitive ability at age 5 years and the child’s experience of school at age 9

Background At age 5, while only embarking on their educational journey, substantial differences in children’s cognitive ability will already exist. The aim of this study was to examine the causal association between below average cognitive ability at age 5 years and child-reported experience of school and self-concept, and teacher-reported class engagement and emotional-behavioural function at age 9 years. Methods This longitudinal cohort study used data from 7,392 children in the Growing Up in Ireland Infant Cohort, who had completed the Picture Similarities and Naming Vocabulary subtests of the British Abilities Scales at age 5. Principal components analysis was used to produce a composite general cognitive ability score for each child. Children with a general cognitive ability score more than 1 standard deviation (SD) below the mean at age 5 were categorised as ‘Below Average Cognitive Ability’ (BACA), and those scoring above this as ‘Typical Cognitive Development’ (TCD). The outcomes of interest, measured at age 9, were child-reported experience of school, child’s self-concept, teacher-reported class engagement, and teacher-reported emotional behavioural function. Binary and multinomial logistic regression models were used to examine the association between BACA and these outcomes. Results Compared to those with TCD, those with BACA had significantly higher odds of never liking school [Adjusted odds ratio (AOR) 1.82, 95% CI 1.37–2.43, p < 0.001], of being picked on (AOR 1.27, 95% CI 1.09–1.48) and of picking on others (AOR 1.53, 95% CI 1.27–1.84). They had significantly higher odds of experiencing low self-concept (AOR 1.20, 95% CI 1.02–1.42) and emotional-behavioural difficulties (AOR 1.34, 95% CI 1.10–1.63, p = 0.003). Compared to those with TCD, children with BACA had significantly higher odds of hardly ever or never being interested, motivated and excited to learn (AOR 2.29, 95% CI 1.70–3.10). Conclusion Children with BACA at school-entry had significantly higher odds of reporting a negative school experience and low self-concept at age 9. They had significantly higher odds of having teacher-reported poor class engagement and problematic emotional-behavioural function at age 9. The findings of this study suggest BACA has a causal role in these adverse outcomes. Early childhood policy and intervention design should be cognisant of the important role of cognitive ability in school and childhood outcomes.

Bowe, Andrea K.↗

Approach to using 3D laser-induced breakdown spectroscopy (LIBS) data to explore the interaction of FLiNaK and FLiBe molten salts with nuclear-grade graphite

Nuclear graphite has historically been a key component of many nuclear reactor designs and has emerged as key to numerous advanced nuclear reactor design concepts. Molten salt reactors (MSRs) are one broad group of advanced reactor designs currently being pursued by industry for commercialization. Several MSR designs under consideration use graphitic materials that directly interface with a molten salt, whether it is a fuel salt, coolant salt, or both. Therefore, the interaction of graphite materials with molten salts must be understood. To gain this required understanding, a range of data is needed including porosity, strength, and composition as a function of different salt exposure parameters. In this study, a laser-induced breakdown spectroscopy (LIBS) measurement and data analysis methodology was developed to obtain spatially resolved elemental composition information for graphite samples exposed to a molten fluoride salt. Traditional univariate emission line analysis of atomic, ionic, and molecular optical emission signals was coupled via correlation analysis with spectral decomposition of the data using principal component analysis. Elemental depth profiling and elemental mapping were also performed to visualize salt–graphite interactions. LIBS was demonstrated to be useful for measuring key analytes such as fluorine and hydrogen, which are troublesome for other analysis techniques. Evidence for complex behavior was found, thereby demonstrating the usefulness of the developed approach for future systematic studies.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Maximizing machine learning interatomic potential transferability for the discovery of the novel stellated octadecagon Bi18-Pt24 cage structure

Achieving true transferability remains the central challenge for Machine Learning Interatomic Potentials (ML-IAPs) in modeling complex bimetallic nanoclusters across their vast potential energy surfaces. We systematically investigate data selection strategies to optimize the Chebyshev Interaction Model for Efficient Simulation (ChIMES) potential for the Bi-Pt nanoclusters by comparing three innovative sampling methods: Principal Component Analysis (PCA)/k-means (structural diversity), t-distributedStochasticNeighborEmbedding (t-SNE)/k-means (force-space diversity), and hierarchical clustering. Quantitatively, the PCA/k-means strategy proved most effective for global accuracy, yielding the lowest force errors and achieving energy root mean square errors (RMSE) values competitive with Density Functional Theory (DFT), demonstrating excellent accuracy (19.16meV/atom). Structural validation on 34 unique DFT-optimized isomers further confirmed the potential’s high fidelity, with the best model PCA/k-means reproducing structures with an average root mean square deviation (RMSD) of 0.10 Å. However, the t-SNE methods, by maximizing diversity in the force space, demonstrated superior extrapolative power, leading to the more precise prediction of a novel stellated octadecagon Bi18⁢Pt24 cage structure, demonstrating the potential for exploring previously unseen morphologies. Our results establish a clear methodology for strategic data sampling that successfully maximizes ML-IAP transferability, providing an accurate and computationally efficient tool that accelerates the theoretical discovery of complex bimetallic architectures.

Vangheluwe, Raphaël [Université Paris-Saclay, CNRS↗