Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “molecular descriptors”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

DOE BSSD Performance Management Metrics Report Q1

Microbes play key roles in our biosphere, from driving global nutrient cycling to impacting plant, animal and human health and disease. Complex data from microbial genomes, proteins, and metabolites provide a window into these tiny engines that drive life on our planet. Yet these data are dispersed among researchers’ laboratories and various repositories, making it difficult to access. This calls for new ways of managing data, improving data interoperability, advancing community standards, and creating an infrastructure where data are shared efficiently. We have built the National Microbiome Data Collaborative (NMDC) to advance how scientists create, use, and reuse data to redefine the way we understand and harness the power of microbes. The vision of the National Microbiome Data Collaborative (NMDC) is to drive a microbiome data sharing network connecting data, people, and ideas to advance microbiome innovation and discovery. The NMDC was launched in 2019 and brought together DOE National Laboratories to collaborate across resources, capabilities, and expertise. The NMDC team was strategically assembled to include software developers, microbial researchers, metadata experts, and multi-omics specialists. The diversity of the NMDC team reflects the inherently interdisciplinary nature of microbiome science, and we leverage the strengths of the DOE National Laboratory system. Towards BER’s goal of advancing an iterative systems biology approach to the understanding of microbial genomes, the NMDC serves as a foundation for infrastructure, data standards, and community building. Together with the flagship DOE User Facilities, the Joint Genome Institute (JGI) and the Environmental Molecular Sciences Laboratory (EMSL), we are developing core capabilities in metadata standards for environmental descriptors and sample handling and processing; standardized bioinformatic workflows; an interface for data search and access; and robust community engagement activities. The NMDC production platform supports long-term data infrastructure and community building for BER’s bioenergy and environmental research goals. Our approach leverages lessons learned and an ambitious framework for collaborative, interdisciplinary data infrastructure to support microbiome research. The NMDC supports data, information, and knowledge access through three defined software tools – the Submission Portal, NMDC EDGE, and the Data Portal – driven by community needs. Herein, we describe the value proposition for the microbiome research community, our overarching strategy, and challenges and opportunities for developing the NMDC as both an infrastructure and community engagement program.

59 BASIC BIOLOGICAL SCIENCES↗

Utilizing Ion-Mobility Data to Estimate Molecular Masses

A method is being developed for utilizing readings of an ion-mobility spectrometer (IMS) to estimate molecular masses of ions that have passed through the spectrometer. The method involves the use of (1) some feature-based descriptors of structures of molecules of interest and (2) reduced ion mobilities calculated from IMS readings as inputs to (3) a neural network. This development is part of a larger effort to enable the use of IMSs as relatively inexpensive, robust, lightweight instruments to identify, via molecular masses, individual compounds or groups of compounds (especially organic compounds) that may be present in specific environments or samples. Potential applications include detection of organic molecules as signs of life on remote planets, modeling and detection of biochemicals of interest in the pharmaceutical and agricultural industries, and detection of chemical and biological hazards in industrial, homeland-security, and industrial settings.

Duong, Tuan↗

Data Science-Driven Analysis of Substrate-Permissive Diketopiperazine Reverse Prenyltransferase NotF: Applications in Protein Engineering and Cascade Biocatalytic Synthesis of (-)-Eurotiumin A

Prenyltransfer is an early-stage carbon-hydrogen bond (C-H) functionalization prevalent in the biosynthesis of a diverse array of biologically active bacterial, fungal, plant, and metazoan diketopiperazine (DKP) alkaloids. Toward the development of a unified strategy for biocatalytic construction of prenylated DKP indole alkaloids, we sought to identify and characterize a substrate-permissive C2 reverse prenyltransferase (PT). As the first tailoring event within the biosynthesis of cytotoxic notoamide metabolites, PT NotF catalyzes C2 reverse prenyltransfer of brevianamide F. Solving a crystal structure of NotF (in complex with native substrate and prenyl donor mimic dimethylallyl S-thiolodiphosphate (DMSPP)) revealed a large, solvent-exposed active site, intimating NotF may possess a significantly broad substrate scope. To assess the substrate selectivity of NotF, we synthesized a panel of 30 sterically and electronically differentiated tryptophanyl DKPs, the majority of which were selectively prenylated by NotF in synthetically useful conversions (2 to > 99%). Quantitative representation of this substrate library and development of a descriptive statistical model provided insight into the molecular origins of NotF's substrate promiscuity. This approach enabled the identification of key substrate descriptors (electrophilicity, size, and flexibility) that govern the rate of NotF-catalyzed prenyltransfer, and the development of an "induced fit docking (IFD)-guided" engineering strategy for improved turnover of our largest substrates. We further demonstrated the utility of NotF in tandem with oxidative cyclization using flavin monooxygenase, BvnB. This one-pot, in vitro biocatalytic cascade enabled the first chemoenzymatic synthesis of the marine fungal natural product, (-)-eurotiumin A, in three steps and 60% overall yield.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Decoherence in Molecular Electron Spin Qubits: Insights from Quantum Many-Body Simulations

Quantum states are described by wave functions whose phases cannot be directly measured, but which play a vital role in quantum effects such as interference and entanglement. The loss of the relative phase information, termed decoherence, arises from the interactions between a quantum system and its environment. Decoherence is perhaps the biggest obstacle on the path to reliable quantum computing. Here we show that decoherence occurs even in an isolated molecule although not all phase information is lost via a theoretical study of a central electron spin qubit interacting with nearby nuclear spins in prototypical magnetic molecules. The residual coherence, which is molecule-dependent, provides a microscopic rationalization for the nuclear spin diffusion barrier proposed to explain experiments. The contribution of nearby molecules to the decoherence has a non-trivial dependence on separation, peaking at intermediate distances. Molecules that are far away only affect the long-time behavior. Because the residual coherence is simple to calculate and correlates well with the coherence time, it can be used as a descriptor for coherence in magnetic molecules. This work will help establish design principles for enhancing coherence in molecular spin qubits and serve to motivate further theoretical work

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data-Centric Development of Lignin Structure–Solubility Relationships in Deep Eutectic Solvents Using Molecular Simulations

Lignin is a natural source of aromatic chemicals with significant potential as an abundant, renewable feedstock for value-added products. Deep eutectic solvents (DES)–solvents composed of a hydrogen bond donor (HBD) and acceptor (HBA) in varying ratios–have emerged as a highly tunable class of solvents for lignin solubilization. However, the variety of possible DES compositions and limited molecular-scale understanding of lignin solubility makes solvent selection a challenge without laborious trial-and-error experimentation. To address these challenges, we use classical molecular dynamics (MD) simulations to study the interactions of lignin model compounds with various DES–water systems. Quantitative parameters (descriptors) were calculated by postprocessing the MD results and used to train a regression model that predicts experimentally determined solubilities of lignin model compounds. This approach revealed that the most important descriptors of solubility are the system temperature, solute hydrophilicity, and metrics quantifying hydrogen bonding. Maximizing the interactions between solute–HBD (hydrophobic group), water–HBD (hydrophilic group), and water–HBA molecules led to the highest model compound solubility. Our results support a hydrotropic mechanism in which extensive DES–water hydrogen bonding and favorable HBD interactions with the solute promote high solubility. We applied the regression model derived using model compounds to predict the solubility of representative lignin oligomers. The model predicted lignin oligomers’ solubilities in good agreement with experiments, indicating that the simulations of model compounds can be extended to predict the solubility of larger lignin compounds across a range of solvent compositions and temperatures. Furthermore, these findings provide new molecular-scale insight into lignin solubilization mechanisms and a new method for computationally screening potential solvent systems for lignin valorization.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

PANNA: Properties from Artificial Neural Network Architectures

We report prediction of material properties from first principles is often a computationally expensive task. Recently, artificial neural networks and other machine learning approaches have been successfully employed to obtain accurate models at a low computational cost by leveraging existing example data. Here, we present a software package “Properties from Artificial Neural Network Architectures” (PANNA) that provides a comprehensive toolkit for creating neural network models for atomistic systems following the Behler–Parrinello topology. Besides the core routines for neural network training, it includes data parser, descriptor builder for Behler–Parrinello class of symmetry functions and force-field generator suitable for integration within molecular dynamics packages. PANNA offers a variety of activation and cost functions, regularization methods, as well as the possibility of using fully-connected networks with custom size for each atomic species. PANNA benefits from the optimization and hardware-flexibility of the underlying TensorFlow engine which allows it to be used on multiple CPU/GPU/TPU systems, making it possible to develop and optimize neural network models based on large datasets.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Pushing Cu uphill of the volcano curve: Impact of a WC support on the catalytic activity of copper toward the hydrogen evolution reaction

Here, the adsorption of atomic H and H 2 on copper mono- and submonolayers supported on hexagonal WC(0001) surfaces has been investigated using density functional theory with the Perdew–Burke–Ernzerhof exchange correlation functional and D2 van der Waals corrections. Results evidence the impact of the termination of the carbide substrate on fundamental properties of Cu adatoms, and, hence, on the stability of molecular and atomic hydrogen, defining copper's catalytic activity for hydrogen evolution reaction. Using H adsorption energy as a descriptor, catalytic activity of Cu adlayers for hydrogen evolution reaction was estimated using traditional volcano curves and a curve, obtained at low hydrogen coverage. Obtained results evidence that copper adlayers supported on the WC may present a viable low-cost alternative to noble metal-based catalysts, with improved catalytic activity compared to that of copper. This, potentially, can be a useful basis for designing and developing novel functional materials with predetermined catalytic properties.

08 HYDROGEN↗

Toward Understanding and Controlling Organic Reactions on Metal Oxide Catalysts

Metal oxides have structurally complex surfaces on which a variety of adsorption site types can occur, including cation sites, anion sites, oxygen vacancy sites, and Brønsted acid sites. These sites can catalyze the catalytic transformation of organic molecules via diverse routes, thus enabling H abstraction, O abstraction, C–C bond formation, and other reactions. This Perspective provides an update on recent advances and future directions for various organic reactions on metal oxide catalyst surfaces, particularly for C–H activation of alkanes and for C–C bond formation with organic oxygenate reactants. Here, we put emphasis on the molecular scale details, on the active site structures required to enable the formation of kinetically relevant transition states, energetic descriptors, as well as contemporary ideas to enable low activation energies. This progress has been enabled by specialized experiments and the increased capabilities of modern electronic structure calculations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Bridging microscopy with molecular dynamics and quantum simulations: an atomAI based pipeline

Recent advances in (scanning) transmission electron microscopy have enabled a routine generation of large volumes of high-veracity structural data on 2D and 3D materials, naturally offering the challenge of using these as starting inputs for atomistic simulations. In this fashion, the theory will address experimentally emerging structures, as opposed to the full range of theoretically possible atomic configurations. However, this challenge is highly nontrivial due to the extreme disparity between intrinsic timescales accessible to modern simulations and microscopy, as well as latencies of microscopy and simulations per se. Addressing this issue requires as a first step bridging the instrumental data flow and physics-based simulation environment, to enable the selection of regions of interest and exploring them using physical simulations. Here we report the development of the machine learning workflow that directly bridges the instrument data stream into Python-based molecular dynamics and density functional theory environments using pre-trained neural networks to convert imaging data to physical descriptors. Additionally, the pathways to ensure structural stability and compensate for the observational biases universally present in the data are identified in the workflow. This approach is used for a graphene system to reconstruct optimized geometry and simulate temperature-dependent dynamics including adsorption of Cr as an ad-atom and graphene healing effects. However, it is universal and can be used for other material systems.

36 MATERIALS SCIENCE↗

Dataset of simulated vibrational density of states and X-ray diffraction profiles of mechanically deformed and disordered atomic structures in Gold, Iron, Magnesium, and Silicon

This dataset is comprised of a library of atomistic structure files and corresponding X-ray diffraction (XRD) profiles and vibrational density of states (VDoS) profiles for bulk single crystal silicon (Si), gold (Au), magnesium (Mg), and iron (Fe) with and without disorder introduced into the atomic structure and with and without mechanical loading. Included with the atomistic structure files are descriptor files that measure the stress state, phase fractions, and dislocation content of the microstructures. All data was generated via molecular dynamics or molecular statics simulations using the Large-scale Atomic/Molecular Massively Parallel Simulator (LAMMPS) code. This dataset can inform the understanding of how local or global changes to a materials microstructure can alter their spectroscopic and diffraction behavior across a variety of initial structure types (cubic diamond, face-centered cubic (FCC), hexagonal close-packed (HCP), and body-centered cubic (BCC) for Si, Au, Mg, and Fe, respectively) and overlapping changes to the microstructure (i.e., both disorder insertion and mechanical loading).

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Exploring the Effects of Node Topology, Connectivity, and Metal Identity on the Binding of Nerve Agents and Their Hydrolysis Products in Metal–Organic Frameworks

Recent studies have shown that metal–organic frameworks (MOFs) built from hexanuclear M(IV) oxide cluster nodes are effective catalysts for nerve agent hydrolysis, where the properties of the active sites on the nodes can strongly influence the reaction energetics. The connectivity and metal identity of these M6 nodes can be easily tuned, offering extensive opportunities for computational screening to predict promising new materials. Thus, we used density functional theory (DFT) to examine the effects of node topology, connectivity, and metal identity on the binding energies of multiple nerve agents and their corresponding hydrolysis products. By computing an optimization metric based on the relative binding strengths of key hydrolysis reaction species (water, agent, and bidentate-bound products), we predicted optimal M6 nodes for hydrolyzing specific nerve agent and simulant molecules, where our results are in qualitative agreement with observed experimental trends. This analysis highlighted the notion that no single metal or node topology is optimal for all possible organophosphates, suggesting that MOFs should be selected based on the agent of interest. Using the large amount of data generated from our DFT calculations, we then derived quantitative structure–activity relationship (QSAR) models to help explain the complex trends observed in the binding energies. Through linear regression, we identified the most important descriptors for describing the binding of nerve agents and their hydrolysis products to M6 nodes. These results suggested that both molecular and node properties, including both structural and chemical features, collectively contribute to the binding energetics. By performing a thorough statistical analysis, we showed that our QSAR models are capable of making quantitatively accurate binding energy predictions for nerve agents and their hydrolysis products in a wide variety of M(IV)-MOFs. In conclusion, the insights gained herein can be used to guide future experiments for the synthesis of MOFs with enhanced catalytic activity for organophosphate hydrolysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

DeePMD-kit v2: A software package for deep potential models

DeePMD-kit is a powerful open-source software package that facilitates molecular dynamics simulations using machine learning potentials known as Deep Potential (DP) models. This package, which was released in 2017, has been widely used in the fields of physics, chemistry, biology, and material science for studying atomistic systems. The current version of DeePMD-kit offers numerous advanced features, such as DeepPot-SE, attention-based and hybrid descriptors, the ability to fit tensile properties, type embedding, model deviation, DP-range correction, DP long range, graphics processing unit support for customized operators, model compression, non-von Neumann molecular dynamics, and improved usability, including documentation, compiled binary packages, graphical user interfaces, and application programming interfaces. This article presents an overview of the current major version of the DeePMD-kit package, highlighting its features and technical details. Additionally, this article presents a comprehensive procedure for conducting molecular dynamics as a representative application, benchmarks the accuracy and efficiency of different models, and discusses ongoing developments.

97 MATHEMATICS AND COMPUTING↗

Interpretable, extensible linear and symbolic regression models for charge density prediction using a hierarchy of many-body correlation descriptors

Here, density functional theory (DFT) is routinely used to make electronic structure predictions for high-throughput screening of materials and molecules for technologically relevant areas, like the identification of better catalysts, electronic materials, and drug discovery. However, the DFT formalism is limited by (a) its poor (quadratic-to-quartic) scaling, and (b) the need to perform repeated eigenvalue computations of the electronic Hamiltonian as part of its self-consistent field (SCF) iteration procedure to obtain the converged ground state electron density, ρ (r). Approaches that directly predict ρ (r) of a structure with high accuracy can accelerate conventional SCF calculations and can also be used in linearly scaling methods such as orbital-free DFT. To this end, we present a procedure to predict the ground state electron density of molecular and periodic three-dimensional systems directly from the atomic structure with a particular emphasis on physical interpretability. In our framework, ρ (r) is modeled using many-body correlation descriptors that accurately capture the effects of local atomic arrangements in the neighborhood of a grid point. Our use of a linear regression scheme to fit to charge density data enables transparent analysis of the relative contributions of various types of local atomic correlations. By systematically including increasingly complex correlations, our model is shown to accurately predict ρ (r) for a variety of chemically and electronically diverse systems — amorphous Ge, Al(001) slab, crystalline Ga 2 O 3 , molecular benzene, and polyethylene. We then demonstrate a symbolic regression-based protocol to construct easily computable, interpretable features from lower-order correlations that significantly improves our electron density predictions with effectively no increase in the computational cost.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Improved Rate for the Oxygen Reduction Reaction in a Sulfuric Acid Electrolyte using a Pt(111) Surface Modified with Melamine

The feasible commercialization of alkaline, phosphoric acid and polymer electrolyte membrane fuel cells depends on the development of oxygen reduction reaction (ORR) electrocatalysts with improved activity, stability, and selectivity. The rational design of surfaces to ensure these improved ORR catalytic requirements relies on the so-called "descriptors" (e.g., the role of covalent and noncovalent interactions on platinum surface active sites for ORR). Here, we demonstrate that through the molecular adsorption of melamine onto the Pt(111) surface [Pt(111)-M ad ], the activity can be improved by a factor of 20 compared to bare Pt(111) for the ORR in a strongly adsorbing sulfuric acid solution. Additionally, the M ad moieties act as "surface-blocking bodies," selectively hindering the adsorption of (bi)sulfate anions (well-known poisoning spectator of the Pt(111) active sites) while the ORR is unhindered. This modified surface is further demonstrated to exhibit improved chemical stability relative to Pt(111) patterned with cyanide species (CN ad ), previously shown by our group to have a similar ORR activity increase compared to bare Pt(111) in a sulfuric acid electrolyte, with Pt(111)-M ad retaining a greater than ninefold higher ORR activity relative to bare Pt(111) after extensive potential cycling as compared to a greater than threefold higher activity retained on a CN ad -covered Pt(111) surface. We suggest that the higher stability of the Pt(111)-M ad interface stems from melamine's ability to form intermolecular hydrogen bonds, which effectively turns the melamine molecules into larger macromolecular entities with multiple anchoring sites and thus more difficult to remove.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Physics-informed machine learning exploration of Na storage mechanisms in disordered carbon

Sodium-ion batteries are a cost-effective, sustainable alternative to lithium-ion systems for large-scale energy storage. However, optimizing sodium storage in carbon-based anodes with microstructural complexity and atomic disorder remains a major challenge. The intrinsic inhomogeneity of these materials produces diverse local environments, making it difficult for conventional methods to predict and control ion dynamics. Hard carbon (HC) anodes, composed of ranges of ordered-to-disordered graphitic and amorphous nanodomains, offer tunable ion storage and rate capacity, yet rationale design remains a challenge due to poorly understood correlation between local atomic feature and ion transport mechanism. Here, to address this challenge, we introduce a data-driven framework that integrates validated machine-learned interatomic potentials, large-scale molecular dynamics simulations, and machine learning to elucidate sodium transport mechanisms as a function of carbon and sodium loading densities. By computing per-ion structural descriptors and applying unsupervised learning, we identify distinct diffusion modes governed by microscopic features. Supervised analysis and correlation mapping then establish quantitative links between these transport regimes and processing variables such as bulk carbon density and sodium content. This physics-informed approach establishes quantitative structure–transport relationships and offers actionable design principles for engineering high-performance HC anodes.

Data-driven framework↗

Tailoring the Weight of Surface and Intralayer Edge States to Control LUMO Energies

Abstract The energies of the frontier molecular orbitals determine the optoelectronic properties in organic films, which are crucial for their application, and strongly depend on the morphology and supramolecular structure. The impact of the latter two properties on the electronic energy levels relies primarily on nearest‐neighbor interactions, which are difficult to study due to their nanoscale nature and heterogeneity. Here, an automated method is presented for fabricating thin films with a tailored ratio of surface to bulk sites and a controlled extension of domain edges, both of which are used to control nearest‐neighbor interactions. This method uses a Langmuir–Schaefer‐type rolling transfer of Langmuir layers (rtLL) to minimize flow during the deposition of rigid Langmuir layers composed of π‐conjugated molecules. Using UV–vis absorption spectroscopy, atomic force microscopy, and transmission electron microscopy, it is shown that the rtLL method advances the deposition of multi‐Langmuir layers and enables the production of films with defined morphology. The variation in nearest‐neighbor interactions is thus achieved and the resulting systematically tuned lowest unoccupied molecular orbital (LUMO) energies (determined via square‐wave voltammetry) enable the establishment of a model that functionally relates the LUMO energies to a morphological descriptor, allowing for the prediction of the range of accessible LUMO energies.

36 MATERIALS SCIENCE↗

Neural network potential from bispectrum components: A case study on crystalline silicon

In this article, we present a systematic study on developing machine learning force fields (MLFFs) for crystalline silicon. While the main-stream approach of fitting a MLFF is to use a small and localized training set from molecular dynamics simulations, it is unlikely to cover the global features of the potential energy surface. Additionally, to remedy this issue, we used randomly generated symmetrical crystal structures to train a more general Si-MLFF. Furthermore, we performed substantial benchmarks among different choices of material descriptors and regression techniques on two different sets of silicon data. Our results show that neural network potential fitting with bispectrum coefficients as descriptors is a feasible method for obtaining accurate and transferable MLFFs.

36 MATERIALS SCIENCE↗

Cluster-Graph Fingerprinting: A Framework for Quantitative Analysis of Machine-Learned Interatomic Model Training and Simulation Data

Machine-learned interatomic models represent a significant advancement in simulation methods, extending the predictive ability of first-principles methods to previously inaccessible length and time scales. However, the data-driven nature of these models can lead to difficult-to-detect errors that can compromise prediction accuracy. To address this challenge, we introduce a novel fingerprinting approach based on the Chebyshev Interaction Model for Efficient Simulation (ChIMES) ML-IAM graph-based descriptor. Our strategy enables efficient and statistically rigorous analysis of system configurations used in ML-IAM training and those generated by their application, e.g., in molecular dynamics simulations. We demonstrate that these fingerprints can effectively assess novelty of a configuration relative to an existing data set and determine dissimilarity among individual configurations, which are two key tasks in workflows for active learning-based ML-IAM training, data set curation, and on-the-fly uncertainty quantification.

36 MATERIALS SCIENCE↗