Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Molecular modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Molecular Modeling and Molecular Dynamics Simulation of a Packed and Intact Bacterial Microcompartment

Bacterial microcompartments (BMCs) are protein-bound organelles found in some bacteria which encapsulate enzymes for enhanced catalytic activity. These compartments spatially sequester enzymes within semipermeable shell proteins and are packed full of enzyme cargoes and metabolites as they fulfill their function. Coupling together recent SAXS and proteomics work, it is possible to develop molecular models for these microcompartments and interrogate enzyme and metabolite dynamics within. Our primary goal of this study is to quantify the permeability of metabolite glyceraldehyde-3-phosphate (G3P) and dihydroxyacetone phosphate (DHAP) across the BMC shell through classical molecular dynamics simulation. The Haliangium ochraceum model of BMC shell (PDB: 6MZX) was used to model an intact BMC of approximately 10 million atoms. Working at this scale presented its own challenges in managing large data sets, with multiple challenges and hardware advances discussed that facilitated this work. Over approximately 750 ns of aggregate simulation, we see multiple permeation events for these metabolites that were added at high concentration through the pores present within BMC shell tiles. When compared to independent permeability estimates for the same metabolites determined through replica exchange umbrella sampling simulations, the permeabilities varied by approximately 3 orders of magnitude. Regardless, the permeability coefficients for both G3P and DHAP are highly similar and very high, such that only very small concentration gradients can be maintained across the BMC shell between the cytosol and BMC interior. The large simulation systems also facilitated comparisons for molecular diffusivity in the crowded environment within the BMC shell. By our estimates, the viscosity within a packed BMC shell is at least 10-fold higher than it would be in neat solution and is the real driver for varying permeability estimates we obtained through simulation. These findings will be used as design inputs for future bioengineering efforts to make products from BMCs, highlighting how permeable BMC shells can be.

Diffusion

Building molecular model series from heterogeneous CryoEM structures using Gaussian mixture models and deep neural networks

Cryogenic electron microscopy (CryoEM) produces structures of macromolecules at near-atomic resolution. However, building molecular models with good stereochemical geometry from those structures can be challenging and time-consuming, especially when many structures are obtained from datasets with conformational heterogeneity. Here we present a model refinement protocol that automatically generates series of molecular models from CryoEM datasets, which describe the dynamics of the macromolecular system and have near-perfect geometry scores. This method makes it easier to interpret the movement of the protein complex from heterogeneity analysis and to compare the structural dynamics observed from CryoEM data with results from other experimental and simulation techniques.

59 BASIC BIOLOGICAL SCIENCES

Molecular Modeling of Surfactant Interaction on Phospholipid Bilayers Mimicking Corneal Epithelium

Surfactants found in consumer products can compromise eye corneal membrane integrity upon accidental exposure. Traditional in vitro and in vivo approaches to evaluate membrane–surfactant interaction pose experimental limitations such as species variability, reproducibility, and most often do not provide the overall picture. These limitations motivate the use of in silico models to study phenomena like cellular disruption assays caused by surfactants at the molecular scale. In this work, coarse-grained molecular dynamics simulations have been employed to investigate how nonionic alcohol ethoxylate (AE) and anionic surfactant alcohol ethoxy sulfate (AES) interact with lipid bilayer liposomes that mimic corneal epithelial cell membranes. The spherical liposome is composed of 1,2-dihexadecanoyl-sn-glycero-3-phosphocholine (DPPC), 1,2-di(9Z-octadecenoyl)-sn-glycero-3-phosphoethanolamine (DOPE), 1,2-di(9Z-octadecenoyl)-sn-glycero-3-phospho-l-serine (DOPS), and cholesterol, resembling the composition of the corneal epithelial cells’ membrane bilayer. The simulation consisted of varying degrees of representative surfactant compositions and two initial types of surfactant configurations within or outside the liposome. Our results reveal that both surfactants induce outer leaflet bulging, agreeing with membrane solubilization models. The more highly ethoxylated surfactant, AE, caused more consistent inner leaflet disruption than AES, resulting in significantly more water permeation and membrane thinning. In addition, both surfactants increase the lateral diffusion of lipids within the membrane layers, with higher ethoxylated AE showing a stronger effect than AES. This study demonstrates how surfactant structure and localization influence bilayer membrane integrity, offering mechanistic insights into the irritation potential, thus guiding the rational design of effective surfactant-based formulations.

Lipids

On-lattice kinetic Monte Carlo approaches for modeling molecular anisotropy in resveratrol crystallization

Stilbenes are a class of organic compounds with broad-ranging pharmaceutical and agricultural applications, which are typically isolated and purified through recrystallization. We are motivated by reducing experimental waste and optimizing yield via developing predictive simulations for processing-dependent crystal morphologies. Using resveratrol as a model stilbene system, we have developed an approach for simulating crystallization with molecular resolution using on-lattice kinetic Monte Carlo. In this work, we highlight modifications to the Stochastic Parallel PARticle Kinetic Simulator (SPPARKS) software package, which were essential to this application. Key enhancements include the incorporation of non-orthogonal cell shapes and monomer anisotropy approximations using bound hard spheres. This new SPPARKS application has been applied to resveratrol with attachment energy libraries obtained from density functional theory, resulting in excellent agreement with experimental morphology prediction.

crystallization

A High-Fidelity Molecular Model of the Cu(111) Repeating Unit

Dynamic processes at surfaces are central to heterogeneous catalysis, but their atomistic mechanism(s) can prove difficult to elucidate due to variations in material structure and the corresponding impact on reactivity. Moreover, disparities between reaction conditions and those employed for spectroscopic characterization at surfaces can inhibit detailed understanding of catalysis-relevant chemistries. Herein, we substantiate the so-called “cluster-surface” analogy by leveraging a low-valent tricopper architecture ( 1 ) as a model system for small molecule activation at Cu(111). Two reaction classes are explored: the adsorption of carbon monoxide (CO) and the dissociative adsorption of dihydrogen (H 2 ). These processes serve as an ideal testbed to compare the reactivity of a molecular cluster ( 1 ) to that of a heterogeneous surface, as both reactions have empirical data from measurements performed on crystalline Cu(111). Cluster 1 reversibly binds CO. Variable temperature NMR analysis with 13 CO reveals a favorable enthalpy but large negative entropy (−5.1 kcal × mol –1 and −22.9 cal × mol –1 × K –1 , respectively) for CO binding, affording a process that is marginally endergonic at room temperature (ΔG ads (298.15 K) = 1.7 ± 0.5 kcal × mol –1 ). Similarly, analogous to a Cu(111) surface, 1 is shown to oxidatively add (chemisorb) H 2 . Kinetic parameters were determined for this process and the activation enthalpy (8.4 ± 0.5 kcal × mol –1 ) closely mirrors that established for H 2 binding at the Cu(111) facet (6.0 to 12.4 kcal × mol –1 ). Together, these results showcase that a trinuclear cluster can reproduce the small molecule binding and activation energetics of a bulk crystalline surface, setting the stage for studying less-defined surface processes in an atomically precise molecular setting.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Molecular model of TFIIH recruitment to the transcription-coupled repair machinery

Transcription-coupled repair (TCR) is a vital nucleotide excision repair sub-pathway that removes DNA lesions from actively transcribed DNA strands. Binding of CSB to lesion-stalled RNA Polymerase II (Pol II) initiates TCR by triggering the recruitment of downstream repair factors. Yet it remains unknown how transcription factor IIH (TFIIH) is recruited to the intact TCR complex. Combining existing structural data with AlphaFold predictions, we build an integrative model of the initial TFIIH-bound TCR complex. We show how TFIIH can be first recruited in an open repair-inhibited conformation, which requires subsequent CAK module removal and conformational closure to process damaged DNA. In our model, CSB, CSA, UVSSA, elongation factor 1 (ELOF1), and specific Pol II and UVSSA-bound ubiquitin moieties come together to provide interaction interfaces needed for TFIIH recruitment. STK19 acts as a linchpin of the assembly, orienting the incoming TFIIH and bridging Pol II to core TCR factors and DNA. Molecular simulations of the TCR-associated CRL4CSA ubiquitin ligase complex unveil the interplay of segmental DDB1 flexibility, continuous Cullin4A flexibility, and the key role of ELOF1 for Pol II ubiquitination that enables TCR. Collectively, these findings elucidate the coordinated assembly of repair proteins in early TCR.

Paul, Tanmoy

Bridging material models across scales: An integrated approach to equation of state and molecular dynamics modeling of copper

New uncertainty-aware equation of state (EOS) and electrical conductivity (EC) models for copper have been developed. The multiphase EOS/EC models are fit to experimental solid/liquid EC isobar measurements as well as density-functional theory molecular dynamics (DFT-MD) EC calculations in both expanded and compressed regimes (0.1–16 g/ cm 3 ⁠). The liquid and solid EOS phases were fit to available experimental data along with additional DFT-MD data over the same range as the EC. Leveraging the DFT-MD data, a corresponding machine-learned interatomic potential (MLIAP) for copper was trained using genetic-algorithm optimization. The copper MLIAP was constrained by EOS shock points at high compressions. The final EOS bounded MLIAP proves to be stable over a large density range (approximately 0.1–20 g/ cm 3 ) with good agreement to an isothermal compression curve, shock Hugoniot, and liquid speed of sound measurements at high pressures (100s of GPa).

Acoustic measurements and instrumentation

In situ molecular imaging of ion clusters reveals the acid gas capture capacity and mechanism of water-lean ionic liquids

Water-lean solvents are a promising technology for capturing acid gases like carbon dioxide (CO 2 ). In situ liquid time-of-flight secondary ionization mass spectroscopy (ToF-SIMS) is used to study a representative solvent N-(2-ethoxyethyl)-3-morpholinopropan-1-amine (2-EEMPA) with different CO 2 loadings to reveal the complex solvent structure upon CO 2 capture. Characteristic peaks of 2-EEMPA, such as m/z – 215 C 11 H 23 N 2 O 2 – (deprotonated 2-EEMPA) and m/z + 217 C 11 H 25 N 2 O 2 + (protonated 2-EEMPA), are detected due to acid gas uptake. Also, solvent molecules and carboxylate ion pairs, such as m/z – 259 C 12 H 23 N 2 O 4 – [(deprotonated 2-EEMPA∙∙∙CO 2 )] and m/z + 261 C 12 H 25 N 2 O 4 + (protonated 2-EEMPA∙∙∙CO 2 ), are observed. Interestingly, more than one CO 2 molecule can be captured per each solvent molecule as evidenced in SIMS mass spectra, for example, m/z – 321 C 13 H 25 N 2 O 7 – [(deprotonated 2-EEMPA)∙∙∙2CO 2 ∙∙∙H 2 O], m/z + 305 C 13 H 25 N 2 O 4 + [(protonated 2-EEMPA)∙∙∙2CO 2 ], m/z – 389 C 17 H 29 N 2 O 8 – [(deprotonated 2-EEMPA)∙∙∙3CO 2 ∙∙∙3CH 2 ], and m/z + 373 C 16 H 25 N 2 O 8 + [(protonated 2-EEMPA)∙∙∙3CO 2 ∙∙∙2C]. However, the monomer of 2-EEMPA and CO 2 seems to be most prevalent. Furthermore, solvent clusters are detected in loaded solvents, for instance m/z + 433 C 22 H 49 N 4 O 4 + [(2-EEMPA)2∙∙∙H] and m/z + 646 [(2-EEMPA) 3 -2H], while capturing CO 2 at different amounts. Relative abundance of cluster ions provides a semi-qualitative venue to assess the free energies of gas capture energetics, indicating the relative stability trend within the same solvent system, previously impossible. These observed ion clusters are verified with molecular modeling, where dimer, trimer, and cluster ions are validated for their presence either due to weak molecular interactions or hydrogen bonds. In situ molecular imaging of ionic liquids and molecular modeling reveals that the acid gas capture mechanism by ionic liquids includes both physical adsorption and chemical bonding with multiple reaction pathways, engaging cluster formation and alteration of solvent structures.

Acid gas capture

A neural master equation framework for multiscale modeling of molecular processes: application to atomic-scale plasma processes

Plasma-surface interactions (PSI) play a crucial role in microelectronics fabrication; however, their multiscale nature and array of complex, often unknown interactions make computational modeling of PSIs extremely difficult. To this end, we propose a general neural master equation (NME) framework that uses master equations to describe the dynamics of a molecular process, wherein neural networks learned from atomistic simulations represent unknown transitions between different system states. By leveraging the physics-based structure of master equations and data-driven state transitions, the NME framework promotes generalizability and physics interpretability, and can bridge disparate length and time scales. The framework is demonstrated for multiscale modeling of Si atomic layer etching and reactive ion etching, where the learned NME-based surface kinetic models exhibit good predictive and extrapolative capabilities for predicting experimentally relevant observables as a function of process parameters. The NME-based surface kinetic models obey physical constraints, which are violated in models based on neural ordinary differential equations. The proposed NME framework for multiscale modeling of molecular processes can pave the way for the discovery of new chemistries and materials in atomic-scale plasma processes.

Chemical engineering

A Variational Autoencoder Model Toward Molecular Structure Representation Learning of Fuels

Here, in this work, a Variational Autoencoder (VAE)-based data-driven modeling framework is developed with the overarching goal of enabling fuel design. The VAE model is trained on a large dataset with several chemical species to learn a compressed latent space molecular representation. Chemical structure in the form of Simplified Molecular Input Line Entry System (SMILES) string is fed as input, encoded into the VAE latent space, and decoded back to the SMILES string using Long Short-Term Memory (LSTM) networks. Complexities of the VAE training loss function are thoroughly examined by varying the weightage (beta (𝜷) parameter) of the latent space regularization term, thereby assessing the balance between reconstruction accuracy and validity, and focusing on both accurate molecular structure reconstruction and latent space consistency. Two different strategies for 𝜷 variation are evaluated: linear annealing and cyclic annealing. In addition, the impact of total correlation adjustment and hierarchical priors is also studied with regard to the balance between reconstruction fidelity and latent space regularization, and potential issues such as posterior collapse, over-regularization, and poor disentanglement of latent variables. Overall, the best performance of the model is achieved with hierarchical priors and incrementally increasing 𝜷 from 0 to a threshold value of 0.25 over 75 epochs. The generative VAE model can be readily coupled with Quantitative Structure–Property Relationship (QSPR) analysis to develop an integrated end-to-end framework for fuel-property prediction and molecular design of novel promising fuels.

fuel design

Leveraging Natural Language Processing and Generative Models in Molecular Chemistry: Property Prediction and Novel Compound Generation

The accurate prediction of molecular properties is important for the rational design and the advancement of green chemistry and sustainable materials research. However, the predictive power of traditional computational chemistry methods is limited due to computational restrictions. Here, in this study, we examine an alternative approach to the accurate prediction of properties of organic compounds: natural language processing (NLP)-based molecular embedding. Using viscosity, partition coefficient (log P), and enthalpy of vaporization as test properties through a survey of comprehensive datasets comprising 5695 data points for viscosity, 25 870 data points for log P, and 2296 data points for enthalpy of vaporization. These are important properties for the design of greener, safer, and sustainable chemical processes. Models were trained using NLP methods such as Mol2vec and fine-tuned ChemBERTa, and results were compared with traditional input featurization techniques such as Morgan fingerprints and quantum chemistry derived sigma profiles and DFT features. Among the various machine learning models, Mol2vec demonstrated superior predictive capabilities, achieving the highest correlation coefficient (R 2 = 0.945) and lowest RMSE (0.106 mPa s) for viscosity, as well as high accuracy for log P and enthalpy of vaporization predictions. These findings establish the Mol2vec featurization technique, graph-convolutional neural networks (GCNN), and fine-tuned ChemBERTa model as powerful tools for predictive modeling of organic compounds properties, offering a significant improvement over previously used featurization techniques and opening up strategies for very-high-throughput computational screening. Finally, we integrated ML models with hybrid language-model-based generative adversarial networks (LM-GAN) to generate novel molecular sequences with desirable properties for different research applications. The ability to computationally design solvents with lower viscosity, lower log P, and lower enthalpy of vaporization offers a data-driven route to accelerating the discovery of sustainable alternatives to traditionally toxic solvents.

ChemBERTa

Shadow molecular dynamics for flexible multipole models

Shadow molecular dynamics provide an efficient and stable atomistic simulation framework for flexible charge models with long-range electrostatic interactions. Shadow molecular dynamics simulations are driven by approximate “shadow” Born–Oppenheimer potentials for which the exact charges and forces are directly accessible without relying on costly (and approximate) iterative solvers. While previous implementations have been limited to atomic monopole charge distributions, we extend this approach to flexible multipole models. We derive detailed expressions for the shadow energy functions, potentials, and force terms, explicitly incorporating monopole–monopole, dipole–monopole, and dipole–dipole interactions. In our formulation, both atomic monopoles and atomic dipoles are treated as extended dynamical variables alongside the propagation of the nuclear degrees of freedom. We demonstrate that introducing the additional dipole degrees of freedom preserves the stability and accuracy previously seen in monopole-only shadow molecular dynamics simulations. In addition, we present a shadow molecular dynamics scheme where the monopole charges are held fixed while the dipoles remain flexible. Our extended shadow dynamics provide a framework for stable, computationally efficient, and versatile molecular dynamics simulations involving long-range interactions between flexible multipoles. This is of particular current interest in combination with machine-learned interatomic potentials, including long-range electrostatic interactions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Applying deep learning methods to develop new models of molecular charge transfer, nonadiabatic dynamics, and nonlinear spectroscopy in the condensed phase

Photon- and field-induced charge transfer has central importance in the generation and storage of electricity, the novel properties of materials, photo-induced catalysis, and electro-optic activity (e.g., photovoltaic cells, fuel cells, and organic chromophores for use in optical fibers and light-emission diodes). These non-equilibrium electronic and chemical transformations are probed by ultrafast, nonlinear spectroscopies. Accurate simulations play a crucial role in our ability to understand, optimize, and control these transformations. This project applies modern deep learning and machine learning (ML) methods to dramatically improve models of electronic dynamics, electronic-nuclear dynamics, and spectroscopic measurements for improved simulations of chemistry in complex environments, far from equilibrium phenomena, and processes in extreme environments, such as materials exposed to strong or resonant fields. This project develops accurate neural net models that go beyond predictive capability to also provide new insight into the fundamental physics underlying electron and nuclear dynamics. To achieve its objectives, this project explores and develops customized versions of high-capacity deep learning algorithms/models. These techniques are developed with an emphasis on fundamental chemical insight, not just predictive accuracy, to assist the development of the next generation of quantum simulation methods.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Beyond real: alternative unitary cluster Jastrow models for molecular electronic structure calculations on near-term quantum computers

Near-term quantum devices require wavefunction ansätze that are expressive while also of shallow circuit depth in order to both accurately and efficiently simulate molecular electronic structure. While the unitary coupled cluster ansatz (e.g., UCCSD) has become a standard, the high gate count associated with the implementation of this limits its feasibility on noisy intermediate-scale quantum (NISQ) hardware. k -Fold unitary cluster Jastrow (uCJ) ansätze mitigate this challenge by providing O( kN 2 ) circuit scaling and favorable linear depth circuit implementation. Previous work has focused on the real orbitalrotation (Re-uCJ) variant of uCJ, which allows an exact (Trotter-free) implementation. Here we extend and generalize the k -fold uCJ framework by introducing two new variants, Im-uCJ and g-uCJ, which incorporate imaginary and fully complex orbital rotation operators, respectively. Similar to Re-uCJ, both of the new variants achieve quadratic gate-count scaling. Our results focus on the simplest k = 1 model, and show that the uCJ models frequently maintain energy errors within chemical accuracy (∼1 kcal mol −1 ). Both g-uCJ and Im-uCJ are more expressive in terms of capturing electron correlation and are also more accurate than the earlier Re-uCJ ansatz. We further show that Im-uCJ and g-uCJ circuits can also be implemented exactly, without any Trotter decomposition. Numerical tests using k = 1 on H 2 , H 3 + , Be 2 , C 2 H 4 , C 2 H 6 and C 6 H 6 in various basis sets confirm the practical feasibility of these shallow Jastrow-based ansätze for applications on near-term quantum hardware.

Tkachenko, Nikolay V. [University of California, B

Empirical evidence that glucan-interacting amino acid side chains within the transmembrane channel collectively facilitate cellulose synthase function

The fundamental mechanism of cellulose synthesis is widely conserved across Kingdoms and depends on cellulose synthases, which are processive, dual-function, family 2 glycosyltransferases (GT-2). These enzymes polymerize glucose on the cytoplasmic side of the plasma membrane and export the glucan chain to the cell surface through an integral transmembrane (TM) channel. Structural studies of active plant cellulose synthases (CESAs) have revealed interactions between the nascent glucan chain and the side chains of polar, charged, and aromatic amino acid residues that line the TM channel. However, the functional consequences of modifying these side chains have not been tested in vivo in CESAs or other processive GT-2s. To test this, we used an established in vivo assay based on genetic complementation of CESA5 in the moss, Physcomitrium patens. For accurate prediction of glucan-interacting amino acid residues, we generated a complete homotrimeric molecular model of PpCESA5 using a combination of homology and de novo modeling. All-atom molecular dynamics-based analyses of contact metrics and interaction energy identified 23 amino acid residues with high propensity to interact with the nascent glucan chain within the TM channel or on the apoplastic surface of PpCESA5. Mutating any one of 18 of these amino acid residues to alanine, thereby removing their side chains, abolished or impaired CESA function, with the strongest effects observed upon the loss of charged amino acid side chains. This provides direct evidence to support the hypothesis that multiple amino acid residues collectively maintain a smooth energy landscape within the TM channel to facilitate glucan translocation.

59 BASIC BIOLOGICAL SCIENCES

Generative modeling enables molecular structure retrieval from Coulomb explosion imaging

Capturing the structural changes that molecules undergo during chemical reactions in real space and time is a long-standing dream and an essential prerequisite for understanding and ultimately controlling femtochemistry. A key approach to tackle this challenging task is Coulomb explosion imaging, which benefited decisively from recently emerging high-repetition-rate X-ray free-electron laser sources. With this technique, information on the molecular structure is inferred from the momentum distributions of the ions produced by the rapid Coulomb explosion of molecules. Retrieving molecular structures from these distributions poses a highly non-linear inverse problem that remains unsolved for molecules consisting of more than a few atoms. Here, we address this challenge using a diffusion-based Transformer neural network. We show that the network reconstructs unknown molecular geometries from ion-momentum distributions with a mean absolute error below one Bohr radius, which is half the length of a typical chemical bond.

Artificial Intelligence (cs.AI)