Engineering Papers⌕ Search

Engineering topics

Bhattacharjee, Himaghna

Publications and source records attributed to Bhattacharjee, Himaghna.

AIMSim : An accessible cheminformatics platform for similarity operations on chemicals datasets

The recent advances in deep learning, generative modeling, and statistical learning have ushered in a renewed interest in traditional cheminformatics tools and methods. Quantifying molecular similarity is essential in molecular generative modeling, exploratory molecular synthesis campaigns, and drug-discovery applications to assess how new molecules differ from existing ones. Further, most tools target advanced users and lack general implementations accessible to the larger community. In this work, we introduce Artificial Intelligence Molecular Similarity (AIMSim), an accessible cheminformatics platform for performing similarity operations on collections of molecules called molecular datasets. AIMSim provides a unified platform to perform similarity-based tasks on molecular datasets, such as diversity quantification, outlier and novelty analysis, clustering, dimensionality reduction, and inter-molecular comparisons. AIMSim implements all major binary similarity metrics and molecular fingerprints and is provided as a Python package that includes support for command-line use as well as a Graphical User Interface for code-free utilization with fully interactive plots.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Accurate Thermochemistry of Complex Lignin Structures via Density Functional Theory, Group Additivity, and Machine Learning

A molecular-level understanding of lignin structures and bond dissociation energies could facilitate depolymerization technologies. Still, this information is currently limited due to the lack of databases and the simplification of surrogate models. Here, substitution effects on seven common linkages in lignin polymers are systematically investigated. An automated reaction network generator is employed to create a database of structures. A new group additivity (GA) model based on principal component analysis (PCA) descriptors is introduced and trained on gas-phase density functional theory data of 4100 species at the M06-2X/6-311++G(d,p) level. Hydrogen bonds, local steric, and nonaromatic ring contributions are also incorporated. Lastly, we improve the accuracy of the group additivity model to reach the G4 theory by computing a data set of 770 species at this level and using a data fusion approach.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Thermochemical Data Fusion Using Graph Representation Learning

Large databases are required for “Big Data” applications in catalysis and materials science. Thermochemical databases can be created by combining data from various sources and by correcting low-fidelity datasets to higher accuracy with minimal computation. To achieve this “data fusion”, thermochemical quantities of interest, calculated at various levels of density functional theory (DFT), need to be mapped to the same, high levels of theory. In this work, a graph theoretical, statistical framework is proposed for such tasks. Subgraph frequencies are shown to provide a natural representation for learning these fusion maps. The maps are linear and are learnt with automated descriptor selection. Using a dataset of as few as ~1% from the QM9 database of 133,885 molecules, these models can predict multiple thermochemical quantities at a higher level of theory with an accuracy of 1 kcal/mol. Here, the method is explainable, generalizable, and provides a diagnostic tool for outlier identification

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗