Engineering PapersSearch

SEARCH · Engineering Papers

Results for “active learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Automated Knowledge Discovery From Simulators

A computational method, SimLearn, has been devised to facilitate efficient knowledge discovery from simulators. Simulators are complex computer programs used in science and engineering to model diverse phenomena such as fluid flow, gravitational interactions, coupled mechanical systems, and nuclear, chemical, and biological processes. SimLearn uses active-learning techniques to efficiently address the "landscape characterization problem." In particular, SimLearn tries to determine which regions in "input space" lead to a given output from the simulator, where "input space" refers to an abstraction of all the variables going into the simulator, e.g., initial conditions, parameters, and interaction equations. Landscape characterization can be viewed as an attempt to invert the forward mapping of the simulator and recover the inputs that produce a particular output. Given that a single simulation run can take days or weeks to complete even on a large computing cluster, SimLearn attempts to reduce costs by reducing the number of simulations needed to effect discoveries. Unlike conventional data-mining methods that are applied to static predefined datasets, SimLearn involves an iterative process in which a most informative dataset is constructed dynamically by using the simulator as an oracle. On each iteration, the algorithm models the knowledge it has gained through previous simulation trials and then chooses which simulation trials to run next. Running these trials through the simulator produces new data in the form of input-output pairs. The overall process is embodied in an algorithm that combines support vector machines (SVMs) with active learning. SVMs use learning from examples (the examples are the input-output pairs generated by running the simulator) and a principle called maximum margin to derive predictors that generalize well to new inputs. In SimLearn, the SVM plays the role of modeling the knowledge that has been gained through previous simulation trials. Active learning is used to determine which new input points would be most informative if their output were known. The selected input points are run through the simulator to generate new information that can be used to refine the SVM. The process is then repeated. SimLearn carefully balances exploration (semi-randomly searching around the input space) versus exploitation (using the current state of knowledge to conduct a tightly focused search). During each iteration, SimLearn uses not one, but an ensemble of SVMs. Each SVM in the ensemble is characterized by different hyper-parameters that control various aspects of the learned predictor - for example, whether the predictor is constrained to be very smooth (nearby points in input space lead to similar output predictions) or whether the predictor is allowed to be "bumpy." The various SVMs will have different preferences about which input points they would like to run through the simulator next. SimLearn includes a formal mechanism for balancing the ensemble SVM preferences so that a single choice can be made for the next set of trials.

Burl, Michael

Graph-Based Prediction of Spatio-Temporal Vaccine Hesitancy From Insurance Claims Data

Growing vaccine hesitancy is contributing to the decline in immunization rates for highly contagious, vaccine-preventable childhood diseases. Therefore, there has been a significant interest in understanding how hesitancy is spreading at higher spatio-temporal resolutions, enabling more targeted interventions. Motivated by this, we study the problem of prediction of vaccine hesitancy at the ZIP Code level, referred to as the VaxHesitancy problem. A significant challenge for this problem is the lack of high-resolution data that indicates hesitancy. Here, we develop a hybrid VaxHesSTL framework that combines a Graph Neural Network (GNN) and a Recurrent Neural Network (RNN) to address the VaxHesitancy problem. The GNN uses a ZIP Code-level network to capture spatial signals from neighboring areas, while the RNN models the temporal dynamics present in the data. We train and evaluate VaxHesSTL using a large dataset, namely the All-Payer Claims Databases (APCD), for Virginia, consisting of insurance claims from over five million individuals for six years. We find that an aggregated contact network or graph, developed from a detailed activity-based population network, plays an important role in the performance of VaxHesSTL, compared to graph models based solely on spatial proximity. Experiments demonstrate that VaxHesSTL outperforms a range of state-of-the-art baselines, which rely solely on historical time series data without accounting for spatial relationships. Since hesitancy data at higher spatial resolution is often unavailable or hard to get, we incorporate an active learning approach with our VaxHesSTL framework to optimize the training set without compromising the prediction performance. We find that hesitancy data for only 18% of ZIP Codes selected by active learning allows us to forecast hesitancy for all the ZIP Codes in the Virginia.

60 APPLIED LIFE SCIENCES

Safe Deep Reinforcement Learning for Active Distribution System Model Predictive Control with EVs and DERs

The temporal and spatial mismatch between PV generation and electric vehicle (EV) charging and discharging may cause voltage violations in active distribution networks. Despite the widespread use of deep reinforcement learning (DRL) in power system optimization and control, it lacks guarantees on constraint satisfaction during both training and deployment. This paper proposes a Lagrangian-based safe DRL approach for model predictive control (MPC) of active distribution systems with large-scale integration of PVs, EVs, and energy storage systems (ESSs). A Transformer-LSTM time-series model is proposed to forecast EV charging demand, which is then formulated as a constraint to ensure charging requirements are met. Using this prediction, a Lagrangian-based safe soft actor-critic (SAC) framework is developed for real-time control in a three-phase unbalanced distribution system, enforcing voltage safety constraints while optimizing the cumulative net reward. By integrating the forecasting model with multi-period constraints, the proposed framework jointly coordinates PV systems, EV charging and discharging, and ESS scheduling within the MPC horizon. Numerical experiments on a modified IEEE 123-bus system with real-world data show that, under a high PV penetration scenario, the proposed method increases the net reward by 30.74% and reduces average voltage violations from 0.0011 p.u. to 0.0002 p.u. compared with standard SAC. Compared with the optimal power flow (OPF) approach, it achieves similar voltage security while yielding lower line losses. It also maintains real-time control capability, reducing operation latency to 53.21 ms per 15-minute control interval. The proposed method remains effective under varying PV/EV penetrations and load conditions.

24 POWER TRANSMISSION AND DISTRIBUTION

Distributed quantum approximate optimization algorithm on a quantum-centric supercomputing architecture

Quantum approximate optimization algorithm (QAOA) has shown promise in solving combinatorial optimization problems by providing quantum speedup on near-term gate-based quantum computing systems. However, QAOA faces challenges for high-dimensional problems due to the large number of qubits required and the complexity of deep circuits, limiting its scalability for real-world applications. In this study, we present a distributed QAOA (DQAOA), which leverages distributed computing strategies to decompose a large computational workload into smaller tasks that require fewer qubits and shallower circuits than are necessary to solve the original problem. These sub-problems are processed using a combination of high-performance and quantum computing resources. The global solution is iteratively updated by aggregating sub-solutions, allowing convergence toward the optimal solution. We demonstrate that DQAOA can handle considerably large-scale optimization problems (e.g., 1000-bit problem), achieving a high solution quality and short time-to-solution, outperforming existing strategies. Furthermore, we realize DQAOA on a quantum-centric supercomputing architecture, paving the way for practical applications of gate-based quantum computers in real-world optimization tasks. To extend DQAOA’s applicability to materials science, we further develop an active learning algorithm integrated with our DQAOA (AL-DQAOA), which involves machine learning, DQAOA, and active data production in an iterative loop. We successfully optimize photonic structures using AL-DQAOA, indicating that solving real-world optimization problems using gate-based quantum computing is feasible. We expect the proposed DQAOA to be applicable to a wide range of optimization problems and AL-DQAOA to find broader applications in material design.

Kim, Seongmin [ORNL] (ORCID:0000000159063004)

DASL-Data and Activities for Solar Learning

DASL-Data and Activities for Solar Learning Data and Activities for Solar Learning (DASL) provides a classroom learning environment based on a twenty-five year record of solar magnetograms from the National Solar Observatory (NSO) at Kitt Peak, AZ. The data, together with image processing software for Macs or PCs, can be used to learn basic facts about the Sun and astronomy at the middle school level. At the high school level, students can study properties of the Sun's magnetic cycle with classroom exercises emphasizing data and error analysis and can participate in a new scientific study, Research in Active Solar Longitudes (RASL), in collaboration with classrooms throughout the country and scientists at NSO and NASA. We present a half-day course to train teachers in the scientific content of the project and its classroom use. We will provide a compact disc with the data and software and will demonstrate software installation and use, classroom exercises, and participation in RASL with computer projection.

Jones, Harrison P.

Bayesian D‐Optimal Designs for Gaussian Process Surrogate Models

Computer experiments often employ space-filling strategies to create surrogate models with strong predictive performance. The impact of model parameter estimation for Gaussian process surrogates, however, is often overlooked. Obtaining a better initial estimate of the covariance lengthscale parameter, θ, can greatly improve the resulting Gaussian process fit through more effective sequential acquisitions during active learning. In this work, we propose a novel initial design maximizing the Bayesian D-optimality criterion of the Gaussian process lengthscale parameter. Previously published results have shown the emphasis on lengthscale estimation to be promising, but relied on an empirically driven design creation process. Our Bayesian D-optimal designs are rooted in information theory and lead to more informative sequential acquisitions by improving lengthscale estimation. In many cases, these gains eventually result in better surrogates than those seeded with space-filling initial designs. Furthermore, Bayesian D-optimal designs can be tailored to either isotropic or anisotropic covariance structures, and the Bayesian framework enables the inclusion of prior knowledge in the design process, offering greater flexibility and adaptability. Through several simulation studies, we demonstrate the advantages of Bayesian D-optimal designs in terms of both lengthscale estimation accuracy and predictive performance during active learning.

Bayesian experimental design

SIM_EXPLORE: Software for Directed Exploration of Complex Systems

Physics-based numerical simulation codes are widely used in science and engineering to model complex systems that would be infeasible to study otherwise. While such codes may provide the highest- fidelity representation of system behavior, they are often so slow to run that insight into the system is limited. Trying to understand the effects of inputs on outputs by conducting an exhaustive grid-based sweep over the input parameter space is simply too time-consuming. An alternative approach called "directed exploration" has been developed to harvest information from numerical simulators more efficiently. The basic idea is to employ active learning and supervised machine learning to choose cleverly at each step which simulation trials to run next based on the results of previous trials. SIM_EXPLORE is a new computer program that uses directed exploration to explore efficiently complex systems represented by numerical simulations. The software sequentially identifies and runs simulation trials that it believes will be most informative given the results of previous trials. The results of new trials are incorporated into the software's model of the system behavior. The updated model is then used to pick the next round of new trials. This process, implemented as a closed-loop system wrapped around existing simulation code, provides a means to improve the speed and efficiency with which a set of simulations can yield scientifically useful results. The software focuses on the case in which the feedback from the simulation trials is binary-valued, i.e., the learner is only informed of the success or failure of the simulation trial to produce a desired output. The software offers a number of choices for the supervised learning algorithm (the method used to model the system behavior given the results so far) and a number of choices for the active learning strategy (the method used to choose which new simulation trials to run given the current behavior model). The software also makes use of the LEGION distributed computing framework to leverage the power of a set of compute nodes. The approach has been demonstrated on a planetary science application in which numerical simulations are used to study the formation of asteroid families.

Burl, Michael

Machine learning-led semi-automated medium optimization reveals salt as key for flaviolin production in Pseudomonas putida

Although synthetic biology can produce valuable chemicals in a renewable manner, its progress is still hindered by a lack of predictive capabilities. Media optimization is a critical, and often overlooked, process which is essential to obtain the titers, rates and yields needed for commercial viability. Here, we present a molecule- and host-agnostic active learning process for media optimization that is enabled by a fast and highly repeatable semi-automated pipeline. Its application yielded 60% and 70% increases in titer, and 350% increase in process yield in three different campaigns for flaviolin production in Pseudomonas putida KT2440. Explainable Artificial Intelligence techniques pinpointed that, surprisingly, common salt (NaCl) is the most important component influencing production. The optimal salt concentration is very high, comparable to seawater and close to the limits that P. putida can tolerate. The availability of fast Design-Build-Test-Learn (DBTL) cycles allowed us to show that performance improvements for active learning are rarely monotonous. This work illustrates how machine learning and automation can change the paradigm of current synthetic biology research to make it more effective and informative, and suggests a cost-effective and underexploited strategy to facilitate the high titers, rates and yields essential for commercial viability.

59 BASIC BIOLOGICAL SCIENCES

Progress in Normalizing Flows for 4d Gauge Theories

Normalizing flows have arisen as a tool to accelerate Monte Carlo sampling for lattice field theories. This work reviews recent progress in applying normalizing flows to 4-dimensional nonabelian gauge theories, focusing on two advancements: an architectural improvement referred to as learned active loops, and the application of correlated ensemble methods to QCD with N f = 2 dynamical fermions.

Abbott, Ryan [Massachusetts Institute of Technolog

matsim-agents v1.0

matsim-agents is a multi-agent AI framework for atomistic materials simulation and discovery. It orchestrates large language models (LLMs), machine-learned interatomic potentials (MLIPs), and DFT codes into a single agentic loop running on laptops and DOE leadership-class supercomputers. MULTI-AGENT ORCHESTRATION A LangGraph state machine with three nodes: a Planner that converts a natural-language research objective into structured tasks; an Executor that dispatches atomistic tools and loops until the queue is empty; and an Analyst that summarizes results into a human-readable report. State is checkpointed after every step and human-in-the-loop gates can be inserted at any edge. HYPOTHESIS-DRIVEN DISCOVERY CHAT An interactive REPL (matsim-agents chat) that couples LLM dialogue with atomistic simulation. Chemical formulas are automatically detected in conversation turns and trigger a full crystal-phase exploration: structure generation → relaxation → stability scoring → result injection back into the conversation, creating a closed hypothesis-refinement loop. CRYSTAL PHASE ENUMERATION Given a composition, the phase explorer enumerates prototypes by stoichiometry: elemental (fcc/bcc/hcp/sc/diamond), binary 1:1 (rocksalt/CsCl/zincblende/ wurtzite/fluorite/rutile), ternary 1:1:3 (cubic perovskite), ternary 1:2:4 (perovskite + spinel), quaternary 1:1:2:6 (Fm-3m double perovskite). 2-D prototypes (graphene, h-BN, MoS2 2H/1T) and multilayer stacking are also supported via --include-2d and --num-layers. SUPERCELL GENERATION AND SITE DECORATION Auto-tiling to a minimum atom count (--min-atoms), explicit NxNxN tiling (--supercell), symmetry-distinct site decorations (--n-orderings), and isotropic lattice-scale sweeps (--lattice-scales) for volume bracketing. MLFF RELAXATION AND STABILITY SCORING HydraGNN (multi-headed GNN) drives structure relaxation via ASE with FIRE, BFGS, or BFGSLineSearch. Stability output: delta-E/atom ranking across phases and a max-residual-force dynamical-stability proxy. Other MLIPs (MACE, NequIP, Orb) can be plugged in through the same interface. DFT BACKENDS Quantum ESPRESSO pw.x and VASP 6.6 are first-class labellers. Both have validated GPU builds and SLURM/PBS launchers for three DOE platforms: Frontier (AMD MI250X, ROCm), Aurora (Intel PVC, oneAPI), Perlmutter (NVIDIA A100, CUDA). QE produces ~100 binaries (pw.x, ph.x, epw.x, ...). VASP supports scf, relax, vc-relax, and vc-relax-shape run types. ACTIVE-LEARNING LOOP matsim-agents al run CONFIG.yaml drives an iterative HydraGNN-DFT loop: MD generates candidates → ensemble/MC-dropout uncertainty selects the most informative → DFT labels them in parallel inside one allocation → dataset grows → HydraGNN retrains → repeat. DFT backend is a single YAML toggle (dft.backend: vasp | qe). LLM-generated seed structures are supported (no curated POSCAR library needed). Config uses ${VAR}, ${VAR:-default}, ${VAR:?msg} shell-style substitution for cross-user/cross-site portability. LLM BACKENDS Ollama (local, default), vLLM (HPC multi-GPU serving), OpenAI, Anthropic, HuggingFace Transformers+Accelerate. Selected at runtime via flag or env var with no code changes. HPC PORTABILITY Same Python entry points run on Frontier (ROCm 7.2), Aurora (oneAPI), and Perlmutter (CUDA 12). DFT and ML stacks are never co-loaded in the same shell; they couple through the scheduler and filesystem. Advanced multi-node launchers (serve, discovery-chat, single-relaxation, active-learning, QE warm-start) are provided for all three platforms. CODABENCH COMPETITION BUNDLE A self-contained benchmark: 159 atomistic test structures across 11 material classes, 5 tasks (formation energy, forces, ML relaxation, AI-DFT relaxation, phase stability ranking), public/private leaderboard split (30/70), and four ready-to-run baselines: MACE-MP-0, HydraGNN, UMA, AllScAIP.

Lupo Pasini, Massimiliano [Oak Ridge National Labo

Benchtop Autonomous Electrochemical Characterization System for Combinatorial Thin-Film Solid Oxide Electrodes

The design of materials for electrochemical energy conversion is complicated by a vast search space of candidate materials and multifaceted property requirements: multicarrier conductivity, stability, and catalytic activity are all necessary but rarely intersect. Although self-driving laboratories are rapidly rising to address such material optimization problems, the required infrastructure for integrated, large-scale robotic facilities can be cost-prohibitive. Here we develop and evaluate a closed-loop measurement system for efficient screening of proton-conducting oxide electrodes for ceramic fuel cells and electrolyzers, building on top of an existing benchtop instrument and integrating techniques for rapid impedance measurement and automated analysis. This system exemplifies a “minimum viable” self-driving implementation that can deliver substantial benefits with relatively simple infrastructure. Combinatorial thin-film microelectrode libraries are characterized with a recently developed joint time-domain and frequency-domain impedance measurement technique, which provides an order-of-magnitude acceleration relative to conventional impedance spectroscopy. The distribution of relaxation times is extracted from impedance data and analyzed without human intervention. These results feed an active learning and Bayesian optimization process that learns to predict electrochemical impedance as a function of material composition, measurement temperature, oxygen partial pressure, and electrical bias, which further reduces the screening time by tenfold with optimized experimental sequences. We apply this system to Ba⁡(Co,Fe,Zr,Y)⁢O 3−𝛿 combinatorial libraries and evaluate its effectiveness for learning material property trends and optimizing expensive-to-evaluate properties such as activation energy. This offers insights into key methodological aspects of practical autonomous experimentation, including surrogate model validation, cost-aware acquisition functions, and high-throughput data interpretation. Our results demonstrate the efficacy of the system for rapidly gathering information, but also highlight real-world experimental challenges of thin-film degradation and numerical instability in surrogate models.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

How Silica Surface Chemistry Modulates Interfacial Water: Insights from Machine Learning Molecular Dynamics

Controlling water structure and dynamics at silica interfaces are central to a wide range of technologies, including protective oxide layers for solar water splitting and nanoporous membranes. In this work, we develop a machine learning interatomic potential, trained via active learning, to achieve ab initio accuracy for water confined between hydroxylated silica surfaces over a range of silanol coverages and slit widths. We find that partially hydroxylated surfaces (50 and 75% OH) support stronger water−surface hydrogen bonding and more extended interfacial density profiles than fully hydroxylated (100% OH) surfaces, indicating that increasing OH coverage does not necessarily strengthen interfacial hydrogenbond networks. Translational diffusion decreases approximately linearly with slit width and OH coverage, whereas rotational dynamics respond nonlinearly. In particular, at the smallest slit width of 5 Å, 75% OH coverage produces an enhanced local tetrahedral ordered interfacial network that strongly suppresses reorientation, while 100% coverage yields a crowded, disordered interfacial layer that also hinders rotation. In contrast, the 50% OH coverage is sufficiently sparse that it does not markedly alter water structure or dynamics under confinement. These results show that coupled control of pore size and surface chemistry enables nonlinear tuning of interfacial water structure and transport, providing a design strategy for optimizing porous silica for either enhanced interfacial stability and controlled reactivity or rapid and selective transport.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Dataset, Code, and Models for Training Deep Learning Potentials for Low Temperature Plasma-Surface Interactions

This repository contains datasets, training scripts, and finished models, and test simulations used in the development of DeepREBO— a machine-learned interatomic potential trained to emulate the REBO2 empirical potential. The data was generated to study deep potential development for simulations of plasma-surface interactions. It uses an active learning framework, starting from a minimal dataset and iteratively expanding it. Included are those generated datasets, the trained models, and simulations used to evaluate the performance of the training process. This resource supports reproducibility and provides a reference framework for training deep potentials in plasma-surface interaction studies.

active learning

Expanding the Domain of Applicability of Machine Learning Models with Limited Data for Drug Property Prediction

Accurate machine learning models for predicting small molecule interactions with biological targets are essential for therapeutic discovery, biothreat response, and computational drug design, but their performance is often limited for understudied targets with sparse experimental data. To address this challenge, we developed and evaluated methods to improve molecular property prediction under low-data conditions, using the NimA-related kinase (NEK) family as a proof-of-concept. This work focused on two complementary goals within the ATOM Modeling PipeLine (AMPL) and the Generative Molecular Design (GMD) loop: expanding model applicability through transfer learning, representation learning, feature scaling, sampling strategies, and active-learning-inspired compound selection; and enabling efficient virtual screening to prioritize compounds that balance predicted activity, design objectives, and synthetic accessibility.

organic

Machine‐Learning‐Driven Exploration of Surface Reconstructions of Reduced Rutile TiO 2

Abstract Titanium dioxide (TiO 2 ) is widely used as a catalyst support due to its stability, tunable electronic properties, and surface oxygen vacancies, which are crucial for catalytic processes such as the reverse water‐gas shift (RWGS) reaction. Reduced TiO 2 surfaces undergo complex surface reconstructions that endow unique properties but are computationally challenging to describe. In this study, we utilize machine‐learning interatomic potentials (MLIPs) integrated with an active‐learning workflow to efficiently explore reduced rutile TiO 2 surfaces. This approach enabled the prediction of a phase diagram as a function of oxygen chemical potential, revealing a variety of reconstructed phases, including a previously unreported subsurface shear plane structure. We further investigate the electronic properties of these surfaces and validate our results by comparing experimental and theoretical high‐resolution transmission electron microscopy (HRTEM). Our findings provide new insights into how extreme surface reductions influence the structural and electronic properties of TiO 2 , with potential implications for catalyst design.

Lee, Yonghyuk [Chemistry and Biochemistry Universi

An Atomistic Study of Reactivity in Solid-State Electrolyte Interphase Formation for Li/Li7P3S11

Lithium metal batteries offer superior volumetric and gravimetric specific capacities compared to those based on traditional graphite anodes. Although advancements in solid-state electrolytes address safety concerns, challenges remain, particularly regarding interphase formation in lithium metal anodes. This work presents a computational framework based on high-throughput first-principles density functional theory and machine-learning interatomic potentials (MLIPs) including automated iterative, active learning to enable robust computational exploration of interphase formation between lithium metal anodes and an inorganic solid-state electrolyte. As a demonstration, we apply the framework to a Li/Li7P3S11 interface and find that it accurately identifies the experimentally observed, thermodynamically stable interphase products as well as their overall spatial arrangement within a heterogeneous, amorphous layered structure, with Li2S domains of nanocrystallinity. Our simulations show two stages, a fast and slow diffusion reaction regime, that corroborate the relative phase formation rate of Li x P, Li2S, and Li3P. Using the Onsager transport theory, we capture time-dependent ionic diffusion within the reacting interface, including cross-correlation effects. We found that cross-correlation effects between Li-P and P-S ionic motion significantly influence P-ion diffusion, making it highly sensitive to the local environment and potentially leading to "kinetic trapping" of Li-P phases. The passivation of the interface is shown as the ionic fluxes all approach zero, effectively halting interphase growth.

Diffusion

D–MOPH–25: diverse MOF–molecule pairs for Henry’s constants prediction

Computational methods like grand-canonical Monte Carlo simulations and machine learning (ML) have accelerated metal–organic frameworks (MOF) exploration but are typically limited to a narrow range of adsorbates due to data availability and force field constraints. In this study, we introduce a dataset of diverse MOF–molecule pairs for Henry’s constant prediction, D–MOPH–25, which systematically explores a diverse chemical space by combining 113 molecular adsorbates with over 5000 MOF structures through an active learning process. D–MOPH–25 constitutes the most diverse adsorbate dataset used in any ML study of molecular adsorption in MOFs to date. Our workflow builds a benchmark for predicting Henry’s constants at 300 K, leveraging conformal prediction for uncertainty quantification. Assessment through Shannon entropy and uniform manifold approximation and projection confirms the comprehensiveness of D–MOPH–25 while highlighting the importance of robust classification to filter out unphysical data points in regression tasks. Although future enhancements in model architecture and sampling criteria could improve predictive performance, our dataset already spans the target space using only 2.31% of total possibilities. This comprehensive dataset facilitates assessment of model generalizability across adsorbate species and can establish a foundation for high-throughput MOF screening and ML-driven separation processes.

active learning

Data‐Driven Engineering of Thermostable Collagen‐Mimetic Peptoid Triple Helices

Collagen-mimetic peptides (CMPs) are engineered molecules designed to replicate the triple-helical structure of natural collagen. A repeating x–y-Gly sequence is the defining motif of CMPs and is critical to their triple-helical structure and stability. Substitutions to the residues occupying the x and y positions present a means to modulate the CMP structure and properties. Peptoid residues—N-substituted glycine derivatives—present an attractive potential substitution due to their thermal stability, proteolytic resistance, biocompatibility, and diverse palette of non-natural side chains, but also tend to introduce a high degree of backbone flexibility that can diminish the stability of the triple helix. In this work, we report a computational active learning cycle comprising molecular dynamics simulation, Gaussian process regression, and Bayesian optimization to computationally identify a number of promising peptoid substitutions predicted to stabilize the desired quaternary structure through side chain interactions and produce stable peptoid-based collagen-like triple helices. To experimentally test the computational predictions, a top candidate identified by the screen was synthesized and imaged using scanning electron microscopy to resolve fibril-like bundles consistent with collagen-like triple helices. This work predicts a number of CMP peptoid substitutions capable of forming stable triple-helical structures, presents a generalizable design strategy for engineering desired peptoid structures, and opens new avenues for the design of peptoid-based biomimetic materials.

active learning