Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Bayesian model selection”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Multi-objective Bayesian optimization of ferroelectric materials with interfacial control for memory and energy storage applications

Optimization of materials’ performance for specific applications often requires balancing multiple aspects of materials’ functionality. Even for the cases where a generative physical model of material behavior is known and reliable, this often requires search over multidimensional function space to identify low-dimensional manifold corresponding to the required Pareto front. In this work, we introduce the multi-objective Bayesian optimization (MOBO) workflow for the ferroelectric/antiferroelectric performance optimization for memory and energy storage applications based on the numerical solution of the Ginzburg–Landau equation with electrochemical or semiconducting boundary conditions. MOBO is a low computational cost optimization tool for expensive multi-objective functions, where we update posterior surrogate Gaussian process models from prior evaluations and then select future evaluations from maximizing an acquisition function. Using the parameters for a prototype bulk antiferroelectric (PbZrO 3 ), we first develop a physics-driven decision tree of target functions from the loop structures. We further develop a physics-driven MOBO architecture to explore multidimensional parameter space and build Pareto-frontiers by maximizing two target functions jointly—energy storage and loss. This approach allows for rapid initial materials and device parameter selection for a given application and can be further expanded toward the active experiment setting. The associated notebooks provide both the tutorial on MOBO and allow us to reproduce the reported analyses and apply them to other systems (https://github.com/arpanbiswas52/MOBO_AFI_Supplements).

36 MATERIALS SCIENCE↗

The Dark Energy Survey supernova program: cosmological biases from supernova photometric classification

ABSTRACT Cosmological analyses of samples of photometrically identified type Ia supernovae (SNe Ia) depend on understanding the effects of ‘contamination’ from core-collapse and peculiar SN Ia events. We employ a rigorous analysis using the photometric classifier SuperNNova on state-of-the-art simulations of SN samples to determine cosmological biases due to such ‘non-Ia’ contamination in the Dark Energy Survey (DES) 5-yr SN sample. Depending on the non-Ia SN models used in the SuperNNova training and testing samples, contamination ranges from 0.8 to 3.5 per cent, with a classification efficiency of 97.7–99.5 per cent. Using the Bayesian Estimation Applied to Multiple Species (BEAMS) framework and its extension BBC (‘BEAMS with Bias Correction’), we produce a redshift-binned Hubble diagram marginalized over contamination and corrected for selection effects, and use it to constrain the dark energy equation-of-state, w. Assuming a flat universe with Gaussian ΩM prior of 0.311 ± 0.010, we show that biases on w are <0.008 when using SuperNNova, with systematic uncertainties associated with contamination around 10 per cent of the statistical uncertainty on w for the DES-SN sample. An alternative approach of discarding contaminants using outlier rejection techniques (e.g. Chauvenet’s criterion) in place of SuperNNova leads to biases on w that are larger but still modest (0.015–0.03). Finally, we measure biases due to contamination on w0 and wa (assuming a flat universe), and find these to be <0.009 in w0 and <0.108 in wa, 5 to 10 times smaller than the statistical uncertainties for the DES-SN sample.

79 ASTRONOMY AND ASTROPHYSICS↗

Bayesian tensorized neural networks with automatic rank selection

Tensor decomposition is an effective approach to compress over-parameterized neural networks and to enable their deployment on resource-constrained hardware platforms. However, directly applying tensor compression in the training process is a challenging task due to the difficulty of choosing a proper tensor rank. In order to address this challenge, this paper proposes a low-rank Bayesian tensorized neural network. Our Bayesian method performs automatic model compression via an adaptive tensor rank determination. We also present approaches for posterior density calculation and maximum a posteriori (MAP) estimation for the end-to-end training of our tensorized neural network. Here, we provide experimental validation on a two-layer fully connected neural network, a 6-layer CNN and a 110-layer residual neural network where our work produces 7.4x to 137x more compact neural networks directly from the training while achieving high prediction accuracy.

Low-rank tensor↗

Optimizing long-term monitoring of radiation air-dose rates after the Fukushima Daiichi Nuclear Power Plant

Radiation air dose rates near the Fukushima Daiichi Nuclear Power Plant (FDNPP) have been steadily decreasing over the past eight years since the release of radioactive elements in March 2011. Currently, the radiation monitoring program is expected to transition to long-term monitoring after most of the remediation activities are completed. The main long-term monitoring objectives are to (1) confirm the continuing reduction of contaminant and hazard levels, (2) provide assurance for the public, (3) accumulate the basic datasets for scientific knowledge and future preparation, and (4) detect changes or anomalies in contaminant mobility (if they occur), or any unexpected processes or events. In this work, we have developed a methodology for optimizing the monitoring locations of radiation air dose-rate monitoring. Our approach consists of three steps in order to determine monitoring locations in a systematic manner: (1) prioritizing the critical locations, such as schools or regulatory requirement locations, (2) diversifying locations that cover the key environmental controls that are known to influence contaminant mobility and distributions, and (3) capturing the heterogeneity of radiation air-dose rates across the domain. Therefore, for the second step, we use a Gaussian mixture model to identify the representative locations among multiple environmental variables, such as elevation and land-cover types. For the third step, we use a Gaussian process model to capture and estimate the heterogeneity of air-dose rates across the domain. Employing an integrated dose-rate map derived from Bayesian geostatistical methods as a reference map, we distribute the monitoring locations in such a way as to capture the heterogeneity of the reference map. Our results have shown that this approach allows us to select monitoring locations in a systematic manner such that the heterogeneity of air dose rates is captured by the minimal number of monitoring locations.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Towards a New Supply Chain Cybersecurity Risk Analysis Technique

Supply chain cyber-attacks, such as the SolarWinds Orion attack, are occurring with greater frequency. These attacks compromise a digital device before it is sent to customers, bypassing traditional security controls to remain persistent and undetected in operational environments. While supply chain attacks are prevalent, methods for analyzing the risk of these attacks are currently unavailable. This paper proposes new supply chain cyber-attack difficulty and risk metrics to evaluate the relative risk of an attack throughout the supply chain lifecycle. Difficulty metrics for each stakeholder in a digital device’s supply chain (e.g., hardware manufacturing, firmware development, software development, storage, and distribution entities) are calculated using scores from cybersecurity maturity questionnaires in a Bayesian Network leaky Noisy-MAX model. These difficulty metrics are then used to calculate an overall supply chain cyber-attack risk. Vulnerability and recoverability metrics are also proposed to evaluate the relative stakeholder influence in the attack risk. These proposed relative risk metrics enable continuous supply chain monitoring, provide decision-makers with information necessary for improved supplier selection, and help drive improvements in the cybersecurity posture of the stakeholders in their supply chain.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Machine Learning Guided Synthesis of Flash Graphene

Advances in nanoscience have enabled the synthesis of nanomaterials, such as graphene, from low-value or waste materials through flash Joule heating. Though this capability is promising, the complex and entangled variables that govern nanocrystal formation in the Joule heating process remain poorly understood. In this work, machine learning (ML) models are constructed to explore the factors that drive the transformation of amorphous carbon into graphene nanocrystals during flash Joule heating. An XGBoost regression model of crystallinity achieves an r 2 score of 0.8051 ± 0.054. Feature importance assays and decision trees extracted from these models reveal key considerations in the selection of starting materials and the role of stochastic current fluctuations in flash Joule heating synthesis. Furthermore, partial dependence analyses demonstrate the importance of charge and current density as predictors of crystallinity, implying a progression from reaction-limited to diffusion-limited kinetics as flash Joule heating parameters change. Finally, a practical application of the ML models is shown by using Bayesian meta-learning algorithms to automatically improve bulk crystallinity over many Joule heating reactions. Furthermore, these results illustrate the power of ML as a tool to analyze complex nanomanufacturing processes and enable the synthesis of 2D crystals with desirable properties by flash Joule heating.

01 COAL, LIGNITE, AND PEAT↗

Uncertainty quantification for deep learning in particle accelerator applications

With the advent of increased computational resources and improved algorithms, machine learning-based models are being increasingly applied to complex problems in particle accelerators. However, such data-driven models may provide overly confident predictions with unknown errors and uncertainties. For reliable deployment of machine learning models in high-regret and safety-critical systems such as particle accelerators, estimates of prediction uncertainty are needed along with accurate point predictions. In this investigation, we evaluate Bayesian neural networks (BNN) as an approach that can provide accurate predictions along with reliably quantified uncertainties for particle accelerator problems, and compare their performance with bootstrapped ensembles of neural networks. We select three accelerator setups for this evaluation: a storage ring, a photoinjector, and a linac. The problems span different data volumes and dimensionalities (e.g., scalar predictions as well as image outputs). It is found that BNN provide accurate predictions of the mean along with reliable estimates of predictive uncertainty across the test cases. In this vein, BNN may offer an attractive alternative to deterministic deep learning tools to generate accurate predictions with quantified uncertainties in particle accelerator applications.

43 PARTICLE ACCELERATORS↗

How to estimate soil organic carbon stocks of agricultural fields? perspectives using ex-ante evaluation

Estimating soil organic carbon (SOC) stocks of agricultural fields has a range of important applications from development of sustainable management practices to monitoring carbon stocks. There are many estimation strategies with the potential for more reliable estimates of SOC stock and more efficient use of soil sampling and analysis resources, especially by leveraging readily available auxiliary information such as remote sensing. However, concrete guidance for strategy selection is lacking. This study narrows this gap with a comparison of strategies for estimating deep SOC stock (0–60 cm) in a prototypical field. Using high density SOC stock measurements and simulation, we built on past studies by 1) ex-ante evaluating a large number of strategy options, 2) using a Bayesian approach to quantify the uncertainty of the comparison, and 3) considering multiple Bayesian models to assess sensitivity to this modeling choice. We found that, using readily available auxiliary information, both balanced and stratified sampling offer substantial improvements over simple random sampling. The auxiliary information most important for this improvement is a Sentinel-2 SOC index = blue / (green × red), followed by the topographic wetness index. We found that these results are robust to the choice of mapping method, but that there is uncertainty in the magnitude of improvement. Here, we recommend future studies implement this Bayesian approach for simulated ex-ante evaluation of SOC stock estimation strategies across more fields to investigate the generalizability of these findings.

54 ENVIRONMENTAL SCIENCES↗

Bayesian High-Rank Hankel Matrix Completion for Nonlinear Synchrophasor Data Recovery

Phasor measurement units (PMUs) provide high temporal-resolution synchrophasor measurements for power system monitoring and control. The frequent data quality issues, such as missing and bad data, prevent the incorporation of synchrophasor data in real-time operations. Most existing data-driven data recovery methods assume the power system dynamics can be approximated by a linear dynamical system, and the recovery performance degrades significantly when the power system is experiencing nonlinear dynamics during significant events. Here, this paper proposes a data-driven Bayesian nonlinear synchrophasor data recovery method (Ba-NSDR) that can recover a consecutive time period of simultaneous data losses or errors across all channels, even when the underlying system is highly nonlinear. The idea is to lift the Hankel matrix of the spatial-temporal synchrophasor data to a higher dimension such that the lifted Hankel matrix is low-rank in that space and can be processed with the kernel trick. Our proposed Bayesian method then infers the probabilistic distributions of synchrophasor from the partial observations. Some distinctive features of Ba-NSDR include an uncertainty index to measure the accuracy of the recovery result and the robustness to parameter selections. Our method is verified on both synthetic and recorded event datasets.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Footprints of the QCD Crossover on Cosmological Gravitational Waves at Pulsar Timing Arrays

Pulsar timing arrays (PTAs) have reported evidence for a stochastic gravitational wave (GW) background at nanohertz frequencies, possibly originating in the early Universe. We show that the spectral shape of the low-frequency (causality) tail of GW signals sourced at temperatures around T ≳ 1 GeV is distinctively affected by confinement of strong interactions (QCD), due to the corresponding sharp decrease in the number of relativistic species, and significantly deviates from ∼ f 3 commonly adopted in the literature. Bayesian analyses in the NANOGrav 15 years and the previous international PTA datasets reveal a significant improvement in the fit with respect to cubic power-law spectra, previously employed for the causality tail. While no conclusion on the nature of the signal can be drawn at the moment, our results show that the inclusion of standard model effects on cosmological GWs can have a decisive impact on model selection. Published by the American Physical Society 2024

Franciolini, Gabriele (ORCID:0000000268929145)↗

Measurement of $\nu_\mu$ CC Interactions With Two-Proton Final State in MINERvA

This dissertation presents a measurement of charged–current (CC) muon–neutrino interactions with exactly two protons and no pions in the final state (CC~$2p\,0\pi$), using data collected by the MINERvA detector in the NuMI medium–energy beam at Fermilab. Such two–proton topologies are a sensitive probe of nuclear dynamics in the few–GeV regime, including multi–nucleon correlations (npnh, notably $2p2h$) and intranuclear final–state interactions (FSI) such as pion absorption and nucleon rescattering. A precise experimental characterization of these processes is essential both for neutrino–interaction theory and for reducing systematic uncertainties in oscillation experiments that rely on accurate modeling of neutrino–nucleus interactions. Events are selected by requiring a $\nu_\mu$ CC interaction with a reconstructed $\mu^-$ and two proton tracks originating from a common vertex in MINERvA’s finely segmented scintillator tracker, with no reconstructed mesons. Muon charge and momentum are constrained by matching to the MINOS Near Detector, while proton identification exploits energy–loss profiles and stopping–proton features. Backgrounds from pion–producing channels that enter the signal region through FSI or reconstruction effects are constrained with data–driven sidebands (Michel–electron and isolated–cluster “blob” samples) and tuned via a simultaneous fit across signal and sideband regions. To correct detector resolution and acceptance effects, the analysis employs iterative Bayesian unfolding with extensive validation: statistical pseudo–experiments, and robustness checks against generator systematic “universes” and additional strong shape warps. Single–differential cross sections are reported for three observables tailored to the two–proton final state: the opening–angle cosine $\cos\!\left(\theta_{pp}\right)$, the leading–proton momentum, and the subleading–proton momentum. Systematic uncertainties include contributions from neutrino flux, interaction modeling (e.g., npnh and resonance parameters, pion FSI), and detector response (calibration, reconstruction efficiencies). The resulting distributions provide targeted constraints on the interplay of multi–nucleon dynamics and FSI that shape CC~$2p\,0\pi$ final states on hydrocarbon. Comparisons to modern GENIE–based simulations highlight kinematic regions where model components require refinement. These measurements thus inform generator tuning and improve the reliability of neutrino–energy reconstruction strategies for current and future long–baseline oscillation programs.

Syrotenko, Vladyslav S. [Tufts U.]↗

The eROSITA Final Equatorial-Depth Survey (eFEDS): The AGN catalog and its X-ray spectral properties

The eROSITA Final Equatorial Depth Survey (eFEDS), observed with eROSITA ahead of its planned 4-yr all-sky survey, is the largest contiguous-field X-ray survey at present. It yielded a large sample of X-ray sources with very rich multiband photometric and spectroscopic coverage. We present here the eFEDS active galactic nuclei (AGN) catalog and the eROSITA X-ray spectral properties of the eFEDS sources. Using a Bayesian method, we performed a systematic X-ray spectral analysis for all the eFEDS sources. We adopted multiple spectral models, including single-component power-law or hot-plasma models and double-component models of a power law plus soft excess. We investigated the capacity of eROSITA X-ray spectra for constraining AGN spectral shapes through a detailed analysis of the posterior parameter probability distribution functions. Hierarchical Bayesian modeling was used to recover the spectral parameter distribution of the sample. The source fluxes and luminosities were measured from the posterior of the spectral fitting. The eFEDS AGN catalog (22 079 sources) comprises ~80% of the eFEDS point sources. Despite a large number of faint sources, our spectral fitting provides reasonable measurements of spectral shapes and intrinsic luminosities for a majority of the sources. Because of sample selection bias, this AGN catalog is dominated by X-ray unobscured sources, with an obscured (logN H > 21.5) fraction of 8%; the power-law emission of the hot corona is also relatively soft, with a typical slope of 2.0. For type-I AGN, the X-ray emission is well correlated with the UV emission with the usual anticorrelation between the X-ray to UV spectral slope α OX and the UV luminosity. The X-ray spectral properties measured with various models are presented for all the eFEDS sources.

79 ASTRONOMY AND ASTROPHYSICS↗

BatteryPro: A Python Toolkit for Battery Data Analysis and Machine Learning Predictions

Analyzing battery test data for research & development can be time-consuming since battery tests often run on the order of months to years, generating large volumes of data. BatteryPro is a comprehensive Python package and software designed to facilitate advanced analysis and performance predictions for battery test data. Developed for battery researchers, it supports data types from widely used battery testing instruments, including MACCOR and Biologic cycling systems. The software provides a variety of tools for extracting and plotting key battery parameters such as time, voltage, capacity, current, and pressure. In addition to its extensive data analysis capabilities, BatteryPro features a dedicated machine learning module that employs a Bayesian Gaussian Mixture Model (GMM) to predict battery performance and degradation. Users can generate synthetic capacity fade data, calculate fade metrics, and leverage predictive models to forecast long-term battery behavior. The software's graphical user interface (GUI) enhances usability, allowing researchers to upload, merge, and analyze multiple data files with full customizability. The GUI also supports machine learning predictions, enabling users to fit models and make predictions based on selected data and parameters. BatteryPro is built using QtDesigner, scikit-learn, matplotlib, and pandas, ensuring a high level of customization, flexibility, and accuracy in battery data analysis. This tool aims to empower researchers with the ability to perform detailed battery analysis and make informed predictions, ultimately advancing the field of battery research.

25 - ENERGY STORAGE↗

Model form and sensitivity analysis of CALPHAD-based nucleation models in b-stabilized Ti alloys

Accurate prediction of α-phase nucleation and growth in β-stabilized titanium alloys is crucial for designing heat treatments to optimize mechanical properties in additively manufactured lightweight components. Ideally, predictions of nucleation and growth would incorporate both top-down observations of past experimental heat treatments and bottom-up modeling of phase transformations; however, the appropriate method of combining these information sources is not self-evident. Combining top-down and bottom-up information requires a unified form of model that can connect between spatiotemporal scales, as well as sets of fitting parameters that can be identified by each data source. The selection of which parameters to fit to which data source can be made based on expert opinion, or by performing a sensitivity analysis. In solid-solid nucleation, direct observation of the nucleation and growth process is challenging. Most data on the heat treatment-controlled phase transformations are not in-situ. To predict the process and outcome of the nucleation, growth and coarsening of precipitates, theoretical models of the nucleation pathway are used to bridge the gap. Many sources of uncertainty affect the modeling of this nucleation process. It can be influenced by small variations in the thermomechanical processing history, chemical composition, and initial microstructure. If molecular dynamics (MD) simulations are used to determine thermodynamic quantities and inform CALPHAD modeling, additional uncertainty can be introduced and accounted for using Bayesian methods. Top-down uncertainties require additional steps to quantify. The influence of nucleation model form on the sensitivity of predictions to input parameters and physical conditions is the focus of this study. Classical nucleation theory (CNT) allows modeling to formulate the nucleation as homogeneous or, more commonly, heterogeneous. Non-classical nucleation models are also increasingly explored as a means of reconciling top-down and bottom-up data. In this study, the sensitivity of the intragranular nucleation of α in a β-annealed, slow-cooled aging (BASCA) heat treatment of β-stabilized Ti5553 alloy is explored using CNT and both heterogeneous and homogeneous assumptions. The Kampmann-Wagner Numerical model of precipitate nucleation and growth is employed. Using open-source tools (pyCalphad and thermodynamic modeling of TiMo as a surrogate system, a sensitivity analysis is performed to measure variations in key parameters, including chemical driving force, interfacial energy, and diffusivity, as they relate to predictions of precipitate number density. The inclusion of top-down and bottom-up data in selection of nucleation model form is discussed.

Rodriguez Negron, A. M.↗

The dark matter halo masses of elliptical galaxies as a function of observationally robust quantities

Context. The assembly history of the stellar component of a massive elliptical galaxy is closely related to that of its dark matter halo. Measuring how the properties of galaxies correlate with their halo mass can therefore help to understand their evolution. Aims. We investigate how the dark matter halo mass of elliptical galaxies varies as a function of their properties, using weak gravitational lensing observations. To minimise the chances of biases, we focus on the following galaxy properties that can be determined robustly: the surface brightness profile and the colour. Methods. We selected 2409 central massive elliptical galaxies (log M*/M ⊙ ≳ 11.4) from the Sloan Digital Sky Survey spectroscopic sample. We first measured their surface brightness profile and colours by fitting Sérsic models to photometric data from the Kilo-Degree Survey (KiDS). We fitted their halo mass distribution as a function of redshift, rest-frame r-band luminosity, half-light radius, and rest-frame u - g colour, using KiDS weak lensing measurements and a Bayesian hierarchical approach. For the sake of robustness with respect to assumptions on the large-radii behaviour of the surface brightness, we repeated the analysis replacing the total luminosity and half-light radius with the luminosity within a 10 kpc aperture, L r, 10 , and the light-weighted surface brightness slope, Γ 10 . Results. We did not detect any correlation between the halo mass and either the half-light radius or colour at fixed redshift and luminosity. Using the robust surface brightness parameterisation, we found that the halo mass correlates weakly with L r,10 and anti-correlates with Γ 10 . At fixed redshift, L r, 10 and Γ 10 , the difference in the average halo mass between galaxies at the 84th percentile and 16th percentile of the colour distribution is 0.00 ± 0.11 dex. Conclusion. Our results indicate that the average star formation efficiency of massive elliptical galaxies has little dependence on their final size or colour. This suggests that the origin of the diversity in the size and colour distribution of these objects lies with properties other than the halo mass.

79 ASTRONOMY AND ASTROPHYSICS↗

Optimal sizing of battery energy storage systems for peak shaving and demand response using a degradation-aware Bayesian Optimization-Mixed-Integer Linear Programming framework

The increasing integration of renewable energy and rising electricity demand highlight the importance of battery energy storage systems for peak shaving and demand response. Unlike prior approaches that overlook operational impacts on degradation, this study proposes a Bayesian Optimization–Mixed Integer Linear Programming framework for optimal battery energy storage system sizing. In this framework, Mixed Integer Linear Programming determines short-term scheduling while a calibrated electrochemical model iteratively evaluates degradation. The central hypothesis is that the framework can efficiently identify optimal sizes that yield realistic and economically robust outcomes. The method is tested across three scenarios: peak shaving, peak shaving with energy-reduction demand response, and peak shaving with power-reduction demand response. Results show that the framework converge to the optimum within 20 iterations out of 150 possible sizes. Under baseline conditions, the framework consistently selects the smallest feasible system, minimizing unnecessary degradation costs from oversized storage. Sensitivity analyses reveal that larger systems are favored as demand rates or incentives increase. Comparisons of demand response programs indicate that power-reduction demand response offers greater economic benefits than energy-reduction demand response, although demand savings from peak shaving remain the dominant contributor to overall performance. This study demonstrates that the proposed framework balances computational tractability with degradation fidelity, identifies critical economic thresholds for investment, and offers a practical, flexible tool to guide industrial stakeholders in cost-effective battery energy storage system deployment.

Batteries↗

Noise-aware optimization in nominally identical manufacturing and measuring systems for high-throughput parallel workflows

Device-to-device variability in experimental noise critically impacts reproducibility, especially in automated, high-throughput systems like additive manufacturing farms. While manageable in small labs, such variability can escalate into serious risks at larger scales, such as architectural 3D printing, where noise may cause structural or economic failures. This contribution presents a noise-aware decision-making algorithm that quantifies and models device-specific noise profiles to manage variability adaptively. It uses distributional analysis and pairwise divergence metrics with clustering to choose between single-device and robust multi-device Bayesian optimization strategies. Unlike conventional methods that assume homogeneous devices or enforce generic robustness, the proposed framework explicitly determines whether shared optimization across devices is appropriate based on the degree of inter-device noise heterogeneity. This enables improved performance, reproducibility, and efficiency. An experimental case study involving three nominally identical 3D printers (same brand, model, and close serial numbers) demonstrates reduced redundancy, lower resource usage, and improved reliability, along with improved convergence stability and solution quality through the selection of the appropriate optimization strategy based on the degree of inter-device noise heterogeneity. Overall, this framework establishes a general approach for precision- and resource-aware optimization in scalable, automated experimental platforms, demonstrated here on a representative multi-device 3D printing case study.

Schenk, Christina↗

Hybrid Data‐Driven Discovery of High‐Performance Silver Selenide‐Based Thermoelectric Composites

Optimizing material compositions often enhances thermoelectric performances. However, the large selection of possible base elements and dopants results in a vast composition design space that is too large to systematically search using solely domain knowledge. To address this challenge, a hybrid data-driven strategy that integrates Bayesian optimization (BO) and Gaussian process regression (GPR) is proposed to optimize the composition of five elements (Ag, Se, S, Cu, and Te) in AgSe-based thermoelectric materials. Data is collected from the literature to provide prior knowledge for the initial GPR model, which is updated by actively collected experimental data during the iteration between BO and experiments. Within seven iterations, the optimized AgSe-based materials prepared using a simple high-throughput ink mixing and blade coating method deliver a high power factor of 2100 µW m −1 K −2 , which is a 75% improvement from the baseline composite (nominal composition of Ag 2 Se 1 ). In conclusion, the success of this study provides opportunities to generalize the demonstrated active machine learning technique to accelerate the development and optimization of a wide range of material systems with reduced experimental trials.

36 MATERIALS SCIENCE↗