Engineering Papers⌕ Search

DOE OSTI · 2281536

Chapter 6: Surrogate Model Guided Optimization Algorithms and Their Potential Use in Autonomous Experimentation

Abstract

This chapter reviews the basics of derivative-free optimization methods based on surrogate models and outlines how these methods can straightforwardly be applied to autonomously steering experimentation. It summarizes general solution approaches that use surrogate models and active learning. Surrogate modeling is often combined with active learning strategies, where in each iteration of the optimization algorithm, the surrogate model is used to identify which new inputs should be evaluated next and given the new input-output pair, the surrogate model is updated. Regardless of feasibility, the surrogate models for the constraints are updated in each iteration of the optimization algorithm, while the surrogate model for the objective function is only updated when a feasible point has been found. Similarly to the case of computationally cheap constraints, the surrogate models of the constraints should be incorporated into the definition of the auxiliary optimization problem that is solved to select new sample points.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mueller, Juliane. 2023-12-13. Chapter 6: Surrogate Model Guided Optimization Algorithms and Their Potential Use in Autonomous Experimentation. https://doi.org/10.1201/9781003359593

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

Data‐Driven Engineering of Thermostable Collagen‐Mimetic Peptoid Triple Helices

Collagen-mimetic peptides (CMPs) are engineered molecules designed to replicate the triple-helical structure of natural collagen. A repeating x–y-Gly sequence is the defining motif of CMPs and is critical to their triple-helical structure and stability. Substitutions to the residues occupying the x and y positions present a means to modulate the CMP structure and properties. Peptoid residues—N-substituted glycine derivatives—present an attractive potential substitution due to their thermal stability, proteolytic resistance, biocompatibility, and diverse palette of non-natural side chains, but also tend to introduce a high degree of backbone flexibility that can diminish the stability of the triple helix. In this work, we report a computational active learning cycle comprising molecular dynamics simulation, Gaussian process regression, and Bayesian optimization to computationally identify a number of promising peptoid substitutions predicted to stabilize the desired quaternary structure through side chain interactions and produce stable peptoid-based collagen-like triple helices. To experimentally test the computational predictions, a top candidate identified by the screen was synthesized and imaged using scanning electron microscopy to resolve fibril-like bundles consistent with collagen-like triple helices. This work predicts a number of CMP peptoid substitutions capable of forming stable triple-helical structures, presents a generalizable design strategy for engineering desired peptoid structures, and opens new avenues for the design of peptoid-based biomimetic materials.

active learning↗

Factorization Machine‐Based Active Learning for Functional Materials Design with Optimal Initial Data

The optimization of functional materials is important to enhance their properties, but their complex geometries pose great challenges to optimization. Data-driven algorithms efficiently navigate such complex design spaces by learning relationships between material structures and performance metrics to discover high-performance functional materials. Surrogate-based active learning, continually improving its surrogate model by iteratively including high-quality data points, has emerged as a cost-effective data-driven approach. Furthermore, it can be coupled with quantum computing to enhance optimization processes, especially when paired with a special form of surrogate model (i.e., quadratic unconstrained binary optimization), formulated by factorization machine (FM). However, current practices often overlook the variability in design space sizes when determining the initial data size for optimization. In this work, we investigate the optimal initial data sizes required for efficient convergence across various design space sizes. By employing averaged piecewise linear regression, we identify initiation points where convergence begins, highlighting the crucial role of employing adequate initial data in achieving efficient optimization. These results contribute to the efficient optimization of functional materials by ensuring faster convergence and reducing computational costs in FM-based active learning.

active learning↗

D–MOPH–25: diverse MOF–molecule pairs for Henry’s constants prediction

Computational methods like grand-canonical Monte Carlo simulations and machine learning (ML) have accelerated metal–organic frameworks (MOF) exploration but are typically limited to a narrow range of adsorbates due to data availability and force field constraints. In this study, we introduce a dataset of diverse MOF–molecule pairs for Henry’s constant prediction, D–MOPH–25, which systematically explores a diverse chemical space by combining 113 molecular adsorbates with over 5000 MOF structures through an active learning process. D–MOPH–25 constitutes the most diverse adsorbate dataset used in any ML study of molecular adsorption in MOFs to date. Our workflow builds a benchmark for predicting Henry’s constants at 300 K, leveraging conformal prediction for uncertainty quantification. Assessment through Shannon entropy and uniform manifold approximation and projection confirms the comprehensiveness of D–MOPH–25 while highlighting the importance of robust classification to filter out unphysical data points in regression tasks. Although future enhancements in model architecture and sampling criteria could improve predictive performance, our dataset already spans the target space using only 2.31% of total possibilities. This comprehensive dataset facilitates assessment of model generalizability across adsorbate species and can establish a foundation for high-throughput MOF screening and ML-driven separation processes.

active learning↗