Engineering PapersSearch

SEARCH · Engineering Papers

Results for “molecular structure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Peak2Patch: High-Fidelity Functional Group Identification through Attention-Based Fusion of Infrared and Mass Spectra

Identifying molecular structure based on spectroscopic readings is a key task in a variety of chemical and biological applications. Common spectroscopy techniques, such as Infrared (IR) Spectroscopy and Mass Spectrometry (MS), provide detailed information on the structure of molecular compounds but nonetheless require expert-level knowledge to decode. Machine learning has emerged as a potential solution for automating structure prediction from chemical spectra; however, current approaches generally focus on single sensor modalities, neglecting to leverage the complementary information contained within differing spectra. In this paper, we introduce Peak2Patch, a novel approach to fusion-enhanced prediction of functional groups from IR and mass spectra. First, we perform a detailed comparison of backbone networks for encoding both sparse mass spectra and dense IR spectra and demonstrate the superior performance of transformer neural networks over current state-of-the-art convolutional neural networks. Second, we evaluate three broad categories of fusion: early (raw feature), middle (deep feature), and late (decision) fusion, demonstrating the potential of a deep feature fusion-based approach. Lastly, we present Peak2Patch, our attention-based fusion scheme, which leverages cross-attention to mix features between encoded tokens of the two modalities. We validate our approach on a publicly available multimodal spectroscopic data set of 790k simulated molecules, demonstrating a large improvement in functional group prediction over both the previous state-of-the-art and our own strong single-modal baselines.

Jacobson, Philip [Sandia National Laboratories (SN

Beyond real: alternative unitary cluster Jastrow models for molecular electronic structure calculations on near-term quantum computers

Near-term quantum devices require wavefunction ansätze that are expressive while also of shallow circuit depth in order to both accurately and efficiently simulate molecular electronic structure. While the unitary coupled cluster ansatz (e.g., UCCSD) has become a standard, the high gate count associated with the implementation of this limits its feasibility on noisy intermediate-scale quantum (NISQ) hardware. k -Fold unitary cluster Jastrow (uCJ) ansätze mitigate this challenge by providing O( kN 2 ) circuit scaling and favorable linear depth circuit implementation. Previous work has focused on the real orbitalrotation (Re-uCJ) variant of uCJ, which allows an exact (Trotter-free) implementation. Here we extend and generalize the k -fold uCJ framework by introducing two new variants, Im-uCJ and g-uCJ, which incorporate imaginary and fully complex orbital rotation operators, respectively. Similar to Re-uCJ, both of the new variants achieve quadratic gate-count scaling. Our results focus on the simplest k = 1 model, and show that the uCJ models frequently maintain energy errors within chemical accuracy (∼1 kcal mol −1 ). Both g-uCJ and Im-uCJ are more expressive in terms of capturing electron correlation and are also more accurate than the earlier Re-uCJ ansatz. We further show that Im-uCJ and g-uCJ circuits can also be implemented exactly, without any Trotter decomposition. Numerical tests using k = 1 on H 2 , H 3 + , Be 2 , C 2 H 4 , C 2 H 6 and C 6 H 6 in various basis sets confirm the practical feasibility of these shallow Jastrow-based ansätze for applications on near-term quantum hardware.

Tkachenko, Nikolay V. [University of California, B

ezAlign: A Tool for Converting Coarse-Grained Molecular Dynamics Structures to Atomistic Resolution for Multiscale Modeling

Soft condensed matter is challenging to study due to the vast time and length scales that are necessary to accurately represent complex systems and capture their underlying physics. Multiscale simulations are necessary to study processes that have disparate time and/or length scales, which abound throughout biology and other complex systems. Herein we present ezAlign, an open-source software for converting coarse-grained molecular dynamics structures to atomistic representation, allowing multiscale modeling of biomolecular systems. The ezAlign v1.1 software package is publicly available for download at github.com/LLNL/ezAlign. Its underlying methodology is based on a simple alignment of an atomistic template molecule, followed by position-restraint energy minimization, which forces the atomistic molecule to adopt a conformation consistent with the coarse-grained molecule. The molecules are then combined, solvated, minimized, and equilibrated with position restraints. Validation of the process was conducted on a pure POPC membrane and compared with other popular methods to construct atomistic membranes. Additional examples, including surfactant self-assembly, membrane proteins, and more complex bacterial and human plasma membrane models, are also presented. By providing these examples, parameter files, code, and an easy-to-follow recipe to add new molecules, this work will aid future multiscale modeling efforts.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

Molecular Design Principles for Photosystem I-Based Biohybrid Solar Fuel Catalysts

Direct solar-to-chemical conversion offers a compelling route to clean, dispatchable energy. Photosystem I (PSI), an evolutionarily optimized light-driven oxidoreductase, can be repurposed for solar-fuel production by coupling its photochemistry to catalytic interfaces. However, the molecular determinants that govern productive electron transfer to abiotic catalysts remain poorly understood. Here, we present molecular structures of active PSI-Pt nanoparticle (PtNP) biohybrids that reveal how protein architecture controls catalyst access, binding geometry, and photocatalytic efficiency. Removal of stromal subunits exposes the electron transfer chain and enables PtNP binding proximal to the F X cluster, demonstrating that steric occlusion limits access to native acceptor regions in PSI. In contrast, in trimeric PSI, PtNPs bind at multiple sites per monomer, but only a subset are positioned within electron transfer distance of terminal cofactors, resulting in a heterogeneous population of productive and nonproductive configurations. Structural analyses and molecular dynamics simulations define the interface topology, electrostatics, and cofactor-to-nanoparticle distances that govern catalyst binding and electron transfer. These results establish that catalytic inefficiency arises not only from intrinsic electron transfer constraints but also from the distribution of binding geometries imposed by the protein scaffold. Together, these findings provide a molecular framework linking protein structure to biohybrid function and define design principles for engineering PSI-based solar fuel systems and protein-nanomaterial interfaces for light-driven catalysis.

biohybrid

Unraveling the Heterogeneous but Ordered Microstructure of the Nonionic Deep Eutectic Solvent Formed by Lauric Acid and N -Methylacetamide

The nonionic deep eutectic solvent, formed by lauric acid (LA) and N-methylacetamide (NMA), has been shown to have a heterogeneous molecular structure in which the LA and NMA form nonpolar and polar domains, respectively. Previous vibrational spectroscopy experiments demonstrated that the ability of the LA domains to solvate compounds was limited to long carbon chains, whereas other nonpolar molecules, such as W(CO) 6 , were found to be solvated by both LA and NMA. These experiments were not fully compatible with the previously proposed micelle-like structure of the nonpolar domains of the LA-NMA DES. In this work, the modeling of the DES molecular structure is pursued using classical molecular dynamics simulations. The new classical model reproduces both the SAXS structural factors and the previously experimentally derived interaction map for these LA-NMA DESs. In addition, the simulation also shows that LA-NMA DESs form highly organized LA aggregates that are difficult to disorganize. Further evidence of the correct description provided by the newly derived model is obtained using a moderately polar probe: chloroform-d. Computations using the classical model have a good agreement with the solvation behavior of the probe derived from experiments, in which the location of the probe is found to be mostly within the polar domain of the DES. The computational model also demonstrates that the probe solvation is a consequence of the tightly packed LA structure, which causes nonpolar molecules to be located at the interphase of the DES nonpolar domains.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Chemical Diversity of Oligomers in Biomass Fast Pyrolysis Oils, Part 2: Heavy Lignin-Derived Molecules and Highly Dehydrated Sugars from Dichloromethane-Insoluble Pyrolytic Lignin

Here, this paper, the second in a series, investigates the water-insoluble, dichloromethane (DCM)-insoluble residue of BTG pyrolysis oil, also known in the literature as high-molecular-weight pyrolytic lignin (HMW-PL). HMW-PL accounts for 9.2 wt % of the original bio-oil, which is known to increase upon aging and contribute to coke formation during bio-oil upgrading. Understanding its composition is crucial for managing pyrolysis oil during the processing. The HMW-PL was fractionated using a silica gel column with solvents of increasing polarity: ethyl acetate (EA), acetone (AC), isopropanol (ISO), and methanol (MeOH), yielding 49.2, 19.9, 8.2, and 8.1 wt %, respectively. Each subfraction was characterized by UV-fluorescence, FTIR, HSQC NMR, and (−)APCI-FT-Orbitrap MS. Analysis revealed that the EA subfraction primarily contained aromatic dimers (C 19 –C 20 ) and trimers (C 17 –C 13 ), while the MeOH subfraction was rich in oxygenated aliphatic monomers (C 8 –C 16 ) likely derived from sugar dehydration products. The resulting fractions were further separated via preparatory HPLC. Fifty-five candidate molecular structures were proposed based on the Orbitrap MS data, supported by UV-fluorescence, FTIR, and NMR results. This fractionation strategy defined four distinct subfractions, enabling the proposal of surrogate molecular structures and advancing the molecular-level understanding of the HMW-PL fraction in the pyrolysis bio-oil.

High-Molecular-Weight PL

Analytic Nuclear Gradients Including Oriented External Electric Fields in a Molecule-Fixed Frame

Electric-field-assisted chemistry has attracted much attention in recent years, particularly in the context of oriented external electric fields for controlling molecular structure and reactivity. Such fields have been explored in a wide range of applications, including switching materials, nanoparticles, controllable catalysts, medicines, and clinical therapies. However, the determination of fixed fields in the laboratory frame becomes ineffective for flexible molecules, as conformational changes can significantly alter the relative orientation between the applied field and molecular structure. In this work, we propose two molecular reference frames─the principal axis frame and the local reference frame─to define oriented electric fields within the molecular framework. These coordinate systems powerfully eliminate ambiguities in the relative orientation between the applied field and the molecule. Analytic nuclear gradients in the presence of external electric fields are derived and implemented, with an initial application to field-dependent geometry optimizations of cis - and trans -formanilide. Analysis of the resulting field-induced equilibrium structures reveals distinct structural responses, validating the accuracy and robustness of the proposed formalism. The analytic gradient framework enables systematic investigations of molecular properties and reactivity under arbitrarily oriented electric fields, opening new opportunities for computational modeling and rational design in electric-field-controlled chemistry.

electric fields

Influence of thermal treatment on structure and catalytic performance of ceria-zirconia supported copper oxide (CuO x /Ce y Zr 1-y O 2 ) catalysts for CO oxidation

Copper oxide (CuO x ) supported on ceria-zirconia (Ce y Zr 1-y O 2 , y = 1.0, 0.5, 0.0) catalysts were investigated to elucidate the effects of thermal treatment on their physicochemical properties and catalytic performance in carbon monoxide (CO) oxidation. Here, the catalysts were synthesized via a one-pot chemical vapor deposition (OP-CVD) method at 700˚C and 900˚C with controlled Cu loading. Characterization techniques, including synchrotron X-ray diffraction (S-XRD), Raman spectroscopy, X-ray photoelectron spectroscopy (XPS), inductively coupled plasma spectroscopy (ICP), and N 2 adsorption-desorption, were implemented to probe the crystalline structure, molecular and electronic structure, oxygen vacancies, specific surface area (SSA) and metal loading. CO oxidation was chosen as a model reaction to explore the structure-catalytic performance relationship. A ∼100% CO conversion was achieved at < 150˚C, particularly with the CuO x /CeO 2 catalyst calcined at 700˚C. In contrast, calcination at 900˚C caused a ∼90% decrease in SSA and a ∼24% increase in T 50 . Activity tests revealed that increasing ZrO 2 content lowered CO oxidation activity despite generating more defect sites. In-situ measurement of the 700 °C calcined samples revealed the presence of stable and unstable defects in CuO x /Ce 0.5 Zr 0.5 O 2 and CeO 2 respectively, which play a key role in the activity of the catalysts. The results highlight that catalytic performance is closely related to the SSA. Furthermore, an optimum calcination temperature favor significant oxygen vacancy formation with required CuO x -support interactions, enhancing redox properties and catalytic performance.

36 MATERIALS SCIENCE

Marine Algae Polysaccharides: An Overview of Characterization Techniques for Structural and Molecular Elucidation

Polysaccharides make up a large portion of the organic material from and in marine organisms. However, their structural characterization is often overlooked due to their complexity. With many high-value applications and unique bioactivities resulting from the polysaccharides’ complex and heterogeneous structures, dedicated analytical efforts become important to achieve structural elucidation. Because algae represent the largest marine resource of polysaccharides, the majority of the discussion is focused on well-known algae-based hydrocolloid polymers. The native environment of marine polysaccharides presents challenges to many conventional analytical techniques necessitating novel methodologies. We aim to deliver a review of the current state of the art in polysaccharide characterization, focused on capabilities as well as limitations in the context of marine environments. This review covers the extraction and isolation of marine polysaccharides, in addition to characterizations from monosaccharides to secondary and tertiary structures, highlighting a suite of analytical techniques.

09 BIOMASS FUELS

Twins in rotational spectroscopy: Does a rotational spectrum uniquely identify a molecule?

Rotational spectroscopy is the most accurate method for determining structures of molecules in the gas phase. It is often assumed that a rotational spectrum is a unique “fingerprint” of a molecule. The availability of large molecular databases and the development of artificial intelligence methods for spectroscopy make the testing of this assumption timely. In this paper, we pose the determination of molecular structures from rotational spectra as an inverse problem. Within this framework, we adopt a funnel-based approach to search for molecular twins, which are two or more molecules, which have similar rotational spectra but distinctly different molecular structures. Here we demonstrate that there are twins within standard levels of computational accuracy by generating rotational constants for many molecules from several large molecular databases, indicating that the inverse problem is ill-posed. However, some twins can be distinguished by increasing the accuracy of the theoretical methods or by performing additional experiments.

74 ATOMIC AND MOLECULAR PHYSICS

Navigating Large Chemical Spaces Using Graph Theory and Integer Programming

Navigating and analyzing large chemical spaces are necessary to accelerate the design and discovery of new molecules and chemical processes. In this work, we introduce a computational framework that integrates graph theory and integer programming to enable the efficient navigation of large chemical spaces. Our framework represents the chemical space as a graph, wherein nodes represent molecules and edges represent the degree of similarity or connectivity based on domain-specific information. Using the graph representation, we identify representative molecules by computing the so-called minimum dominating set (MDS), which in our context is the minimum set of molecules that is connected to all other molecules. We present a suite of solution strategies for the MDS problem including heuristic and rigorous integer programming (IP) approaches. We show that these approaches allow us to capture physicochemical properties and domain-specific logic and constraints, facilitating the identification of molecules with the target properties. We demonstrate the effectiveness of the proposed approach by navigating the chemical space of per- and polyfluoroalkyl substances (PFAS); this comprises approximately 15,000 molecular structures. We compare our framework against traditional dimensionality reduction and clustering methods such as t-SNE and K-means clustering.

Chemical structure

Two datasets are better than one: method of double moments for 3D reconstruction in cryo-EM

Cryo-electron microscopy is a powerful imaging technique for reconstructing three-dimensional molecular structures from noisy tomographic projection images of randomly oriented particles. We introduce a new data fusion framework, termed the method of double moments, which reconstructs molecular structures from two instances of the second-order moment of projection images obtained under distinct orientation distributions: one uniform, the other non-uniform and unknown. We prove that these moments generically uniquely determine the underlying structure, up to a global rotation and reflection, and we develop a convex-relaxation-based algorithm that achieves accurate recovery using only second-order statistics. Our results demonstrate the advantage of collecting and modeling multiple datasets under different experimental conditions, illustrating that leveraging dataset diversity can substantially enhance reconstruction quality in computational imaging tasks.

Kam’s method

Molecular and structural characterization of a Bacillus cereus strain producing an anthrax-like capsule

Bacillus cereus is a ubiquitous Gram-positive, spore-forming, rod-shaped saprophytic bacterium, occasionally reported to cause food-borne illnesses. However, instances of B. cereus strains harboring anthrax toxin and capsule genes have elevated certain strains as formidable pathogens and biothreats. This study focuses on the genomic analysis and the structural characterization of capsular material produced by the virulent B. cereus PATH2418 strain, isolated from the wound of a traumatic open fracture patient. The genome was sequenced using Nanopore MinION sequencing, revealing a chromosome of 5,270,283 bp and three plasmids. One plasmid, pATH1, was found to encode an operon for the biosynthesis of a bacterial capsule. This operon had sequence homology to the Bacillus anthracis capBCADE operon, which encodes the poly-γ-D-glutamate (PDGA) capsule. The capsule production in B. cereus PATH2418 was influenced by temperature and CO 2 levels. Structural analysis of the capsular material using a combined approach of nuclear magnetic resonance (NMR) and high-performance liquid chromatography (HPLC) techniques confirmed the presence of a high-molecular-weight poly-γ-glutamate capsule, with an enantiomeric composition of approximately 67% D-glutamic acid and 33% L-glutamic acid, matching that of B. anthracis.

Bacillus cereus

Quantum chemically calculated Abraham parameters for quantifying and predicting polymer hydrophobicity

The leakage and accumulation of plastic in the environment is a significant and growing problem with numerous detrimental impacts and has led to a push toward the design and development of more environmentally benign materials. To this end, we have developed a quantum chemistry-based model for predicting the mobility of polymer materials from molecular structure. Hydrophobicity is used as a surrogate for mobility given that hydrophobic interactions drive much of the partitioning of contaminants in and out of various environmentally relevant compartments. To model polymer hydrophobicity, we adjusted a previously developed Quantum Chemically Calculated Abraham Parameter model to calculate Abraham parameters of small molecules from molecular structure information. The resulting model predicted the octanol-water partition coefficient (K OW ) of polymer repeating units with a root mean square error (RMSE) of 0.48 (log scale). Additionally, the hydrophobicity of high molecular weight polymer materials was captured through solubility parameters and Nile red staining experiments from the literature and predicted with RMSEs of 1.21 (J/cc) 0.5 and 3.42 nm, respectively. Finally, to test the environmental applicability of the model, the relative adsorption capacity of three polymers was predicted and used to unify sorption isotherms across multiple sorbates and polymer sorbents.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

PySIDT: Subgraph Isomorphic Decision Trees for Molecular Property Prediction

Accurate molecular property prediction is important across all fields of chemistry. Deep neural networks (DNNs) have become increasingly popular due to their ability to train automatically, avoiding the incredibly tedious process of constructing and extending traditional property estimation schemes. However, DNNs require large amounts of training data, are challenging to interpret, require large amounts of memory to load even during inference, and have severe difficulties incorporating qualitative chemical knowledge, which are often desired for molecular property prediction tasks. Here, in this study, we present PySIDT (https://github.com/zadorlab/PySIDT), a software for training and running inference on Subgraph Isomorphic Decision Trees (SIDTs). SIDTs are graph-based decision trees made of nodes associated with molecular substructures. Inference is done by descending target molecular structures down the decision tree to nodes with matching subgraph isomorphic substructures and making predictions based on the final (most specific) nodes matched. SIDTs scale down well to dataset sizes much smaller than is feasible for DNNs. As trees of molecular substructures, SIDTs are inherently readable and easy to visualize, making them easy to analyze. They are also straightforward to extend and retrain, facilitate uncertainty estimation, and enable easy integration of expert knowledge. We demonstrate the SIDT approach discussing its application to a diverse range of molecular prediction tasks: rate coefficient estimation, diffusion coefficient estimation, thermochemistry estimation, transition state bond stretch prediction, p K a prediction, stability of molecular structures, stability of surface structures, and prediction of surface lateral interaction energetics. Additionally, we demonstrate the power of the SIDT algorithms in two direct learning curve vanilla comparisons with the popular DNN-based software Chemprop and the popular gradient boosted trees-based software XGBoost on enthalpy of formation and rate coefficient prediction tasks. In particular, in the enthalpy of formation case, vanilla PySIDT is able to outperform vanilla Chemprop and XGBoost across the full range of training/validation set sizes out to 11,560 data points.

Johnson, Matthew Sean [Sandia National Laboratorie

On the Prospect of Chemically Transferable Coarse-Grained Electronic Models for Soft Materials

Electronic coarse-graining (ECG) methods predict quantum-mechanical electronic properties directly from coarse-grained (CG) molecular configurations, enabling electronic predictions at mesoscale length scales. Here, we present a diagnostic assessment of the feasibility of chemically transferable ECG models across a broad polymer-relevant chemical space using all-atom, united-atom, and Martini-scale representations. While high-resolution ECG models achieve near-quantitative accuracy, we show that chemically transferable ECG at the Martini resolution fails because the CG force field does not sample the same configurational distribution of local molecular structure as that underlying the DFT-parameterized ECG model. We demonstrate that our proposed Element-Count-Label (ECL) representation, which augments Martini beads with explicit stoichiometric data, significantly improves chemical generalization across diverse polymer chemistries. However, we find that even with improved chemical resolution, the model cannot recover electronic property distributions that are absent from the configurational space sampled by the CG force field. These results demonstrate that chemically transferable ECG requires future Martini-like force fields to explicitly preserve quantum chemistry–compatible local molecular structure in addition to thermodynamic and structural fidelity.

Kidder, Katherine M [Department of Chemistry; Univ

Comparison of Machine Learning Approaches for Prediction of the Equivalent Alkane Carbon Number for Microemulsions Based on Molecular Properties

The chemical properties of oils are vital in the design of microemulsion systems. The hydrophilic–lipophilic difference equation used to predict microemulsions’ phase behavior expresses the oils’ physiochemical properties as the equivalent alkane carbon number (EACN). The experimental determination of EACN requires knowledge of the temperature dependence of the microemulsion system and the effects of different surfactant concentrations. Thus, the experimental determination is time-intensive and tedious, requiring days to months for proper separations. Furthermore, the experiments require high purity of chemicals because microemulsions are sensitive to impurities. Our work focuses on the quick and reliable predictions of the EACN with machine learning (ML) models. Due to the immaturity of ML chemical predictions, we compare three graph neural networks (GNNs) and a gradient-boosted tree algorithm, known as XGBoost. The GNNs use the molecular structures represented as simplified molecular-input line-entry system (SMILES) codes for the initial input, which allows us to assess whether geometry optimization is necessary for reliable results. The XGBoost model also begins with the SMILES representations of the molecules but uses molecular descriptors instead of geometry optimizations. As a result, the best model tested (crystal graph convolutional neural network with Merck molecular force field-94) has an error of 1.15 EACN units of the true EACN for unknown data with the errors skewed toward zero and an R² score of 0.9

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Visualizing the Three-Dimensional Arrangement of Hydrogen Atoms in Organic Molecules by Coulomb Explosion Imaging

Structure-sensitive methods based on femtosecond light or electron pulses are now making it possible to measure how molecular structures change during light-induced processes. Despite significant progress, high-fidelity imaging of nuclear positions remains a challenge even for relatively small molecular systems and, notably, regarding the positions of hydrogen atoms. As demonstrated in recent work, X-ray-induced Coulomb explosion imaging (CEI) may overcome this obstacle, as its sensitivity does not depend on the mass of the imaged atoms. The photoinduced ring opening of the heterocyclic molecule 2(5 H )-thiophenone has attracted recent interest. Here, in this work, we show that CEI offers a powerful route to imaging the peripheral H atoms in this molecule and thus, more generally, to tracking detailed nuclear motions (e.g., isomerizations) in organic molecules on ultrafast time scales. Specifically, we record momentum-space Coulomb explosion images that report on the three-dimensional positioning of all nuclei within the molecule, for instance, distinguishing H atoms in C–H bonds that lie within or are directed out of the plane defined by the heavy atoms. The prospect of imaging peripheral H atoms to probe photochemical dynamics is explored by coupling ab initio molecular dynamics with classical Coulomb explosion simulations, thereby differentiating potential photoproduct isomers, including those whose structures primarily differ in the position of the hydrogens.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH