Spin-free quantum chemistry. III - Bond functions and the Pauling rules.
Spin-free derivation of Pauling rules for evaluating matrix elements for spin-free Hamiltonian over anti-symmetric Slater bond functions
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Spin-free derivation of Pauling rules for evaluating matrix elements for spin-free Hamiltonian over anti-symmetric Slater bond functions
The number of materials that “bridge the gap” between single molecules and extended solids, such as metal-organic frameworks and organic semiconductors, has been increasing. Consequently, there is a growing need for modeling approaches that effectively integrate the real-space molecular perspective employed by computational chemists and the reciprocal-space dispersive perspective employed by computational physicists. Here, we propose the localized active space (LAS) approach as a promising method to successfully bridge this gap. The LAS approach extends the active space concept from multiconfigurational methods such as complete active space self-consistent field theory to multiple molecular fragments via a product-form wave function ansatz. Here, we apply this method to solid state phenomena by treating each unit cell as a fragment with different sets of local quantum numbers (e.g., charge and excitation number). State interaction between these LAS states (LASSI) thus provides a comprehensive basis for the study of charge and energy transfer, meeting and surpassing the capabilities of single-reference fragmentation approaches such as constrained density functional theory (cDFT). Most centrally, we show how combining this LASSI approach with multiconfigurational pair-density functional theory (MC-PDFT) provides an elegant and efficient method to compute band structures that capture multiconfigurational character. We apply the LASSI band structure approach to the computation of band gaps in stretched hydrogen chain, polyacetylene, and bulk nickel oxide (NiO), finding good or excellent quantitative agreement with reference values in all cases. Additionally, we use the LAS basis in one-dimensional model systems to demonstrate its ability to treat difficult solid-state phenomena such as exciton transfer and excitation at p-n junctions.
We benchmark the accuracy of Dunning correlation-consistent Gaussian basis sets for computing frequencydependent second-order hyperpolarizabilities relevant to second-harmonic generation (SHG), using multiresolution analysis (MRA) as a reference. Basis set errors are analyzed using a unit-sphere representation of the effective hyperpolarizability vector, enabling direct assessment of directional error structure. We introduce a relative RMS total error metric that integrates directional deviations over the unit sphere and complement it with signed projection errors that distinguish over- and underestimation. Unsupervised clustering based on these signed directional metrics reveals four distinct convergence behaviors across a set of 68 molecules. Unitsphere visualizations of representative systems show that basis set errors are often highly anisotropic and localized along specific bond directions, even when global error measures appear small. Doubly augmented basis sets consistently outperform singly augmented ones, and core-polarization functions are required for uniform convergence in second-row systems. Overall, this work demonstrates that directional analysis combined with clustering provides a robust framework for understanding basis set convergence in nonlinear optical response properties.
Microkinetic models for catalytic systems require estimation of many thermodynamic and kinetic parameters that can be calculated for isolated species and transition states using ab initio methods. However, the presence of nearby coadsorbates on the surface can dramatically alter these thermodynamic and kinetic parameters causing them to be dependent on species coverage fractions. As there are combinatorially many coadsorbed configurations on the surface, computing the coverage dependence of these parameters is far less straightforward. We present a framework for generating and applying machine learning models to predict coverage-dependent parameters for microkinetic models. Our toolkit enables automatic calculation and evaluation of coadsorbed configurations allowing us to sample 2,000 coadsorbed adsorbates and transition states (TSs) for a diverse set of 9 reactions on Cu(111), a challenging surface, with four possible coadsorbates. This dataset was then used to train subgraph isomorphic decision trees (SIDTs) to predict the stability and association energy of configurations. We were able to achieve mean absolute errors (MAEs) of 0.106 eV on adsorbates, 0.172 eV on TSs, and due to natural error cancellation in SIDTs for relative properties, 0.130 eV on reaction energies and 0.180 eV on activation barriers. In conclusion, we describe how to use these models to predict coverage-dependent corrections for adsorbates and TSs and demonstrate on H*, HO*, and O* comparing the generated SIDT model with an iteratively refined version.
Virtual agents and cloud computing have enabled chemists to easily access automated simulations of explicitly solvated molecules.
The combination of machine learning (ML) models with chemistry-related tasks requires the description of molecular structures in a machine-readable way. The nature of these so-called molecular descriptors has a direct and major impact on the performance of ML models and remains an open problem in the field. Structural descriptors like SMILES strings or molecular graphs lack size-independence and can be memory intensive. Machine-learned descriptors can be of low dimensionality and constant size but lack physical significance and human interpretability. Sigma profiles, which are unnormalized histograms of the surface charge distributions of solvated molecules, combine physical significance with low dimensionality and size-independence, making them a suitable candidate for a universal molecular descriptor. However, their widespread adoption in ML applications requires open access to sigma profile generation, which is currently not available. This work details the development of OpenSPGen – an open-source tool for generating sigma profiles. Also presented are studies on the effect of different settings on the efficacy of the generated sigma profiles at predicting thermophysical material properties when used as inputs to a Gaussian process as a simple surrogate ML model. We find that a higher level of theory does not translate to more accurate results. We also provide further recommendations for sigma profile calculation and use in ML models.
Future gravitational wave antennas will be approximately 100 kilogram cylinders, whose end-to-end vibrations must be measured so accurately (10 to the -19th power centimeters) that they behave quantum mechanically. Moreover, the vibration amplitude must be measured over and over again without perturbing it (quantum nondemolition measurement). This contrasts with quantum chemistry, quantum optics, or atomic, nuclear, and elementary particle physics where measurements are usually made on an ensemble of identical objects, and care is not given to whether any single object is perturbed or destroyed by the measurement. Electronic techniques required for quantum nondemolition measurements are described as well as the theory underlying them.
Simulations of quantum chemistry and quantum materials are believed to be among the most important applications of quantum information processors. However, realizing practical quantum advantage for such problems is challenging because of the prohibitive computational cost of programming typical problems into quantum hardware. Here we introduce a simulation framework for strongly correlated quantum systems represented by model spin Hamiltonians that uses reconfigurable qubit architectures to simulate real-time dynamics in a programmable way. Our approach also introduces an algorithm for extracting chemically relevant spectral properties via classical co-processing of quantum measurement results. We develop a digital–analogue simulation toolbox for efficient Hamiltonian time evolution using digital Floquet engineering and hardware-optimized multi-qubit operations to accurately realize complex spin–spin interactions. As an example, we propose an implementation based on Rydberg atom arrays. In addition, we show how detailed spectral information can be extracted from the dynamics through snapshot measurements and single-ancilla control, enabling the evaluation of excitation energies and finite-temperature susceptibilities from a single dataset. To illustrate the approach, we show how to use the method to compute key properties of a polynuclear transition-metal catalyst and two-dimensional magnetic materials.
Fragment-based quantum chemistry is a powerful strategy for calculating protein−ligand interaction energies using quantum chemistry methods. Rigorous convergence often requires hundreds of atoms in the protein binding-site model, especially if that model is constructed using distance-based criteria to select amino acid residues, while three- and four-body calculations exhibit instability related to combinatorial proliferation in the number of subsystem calculations. Here, we report an energy-based screening protocol for the many-body expansion applied to protein−ligand interactions, implemented in the open-source FRAGME∩T code. Using a combination of aggressive screening based on semiempirical quantum chemistry, with an improved graph-theoretical algorithm to eliminate unimportant subsystems, we are able to perform n-body calculations up to n = 7 using density functional theory in triple-ζ basis sets. Distance cutoffs further reduce the cost without compromising accuracy. Rapid and stable convergence of the many-body expansion is obtained by n = 4, for a pair of metalloenzymes in which a divalent ion coordinates directly to the ligand. As compared to previous results that relied solely on distance cutoffs, oscillations in the n-body corrections are reduced or eliminated, although residual errors remain in one case. This work demonstrates that benchmark-quality protein−ligand interaction energies can be systematically converged using a method with excellent parallel efficiency and scalability.
Hydrocarbon combustion involves the reaction dynamics of a tremendous number of species beginning with many-component fuel mixtures and proceeding via a complex system of intermediates to form primary and secondary products. Combustion conditions corresponding to new advanced engines and/or alternative fuels rely increasingly on autoignition and low-temperature-combustion chemistry. In these regimes various transient radical species such as HO2, ROO·, ·QOOH, HCO, NO2, HOCO, and Criegee intermediates play important roles in determining the detailed as well as more general dynamics. A clear understanding and accurate representation of these processes is needed for effective modeling. Given the difficulties associated with making reliable experimental measurements of these systems, computation can play an important role in developing these energy technologies. Accurate calculations have their own challenges since even within the simplest dynamical approximations such as transition state theory, the rates depend exponentially on critical barrier heights and these may be sensitive to the level of quantum chemistry. Moreover, it is well-known that in many cases it is necessary to go beyond statistical theories and consider the dynamics. Quantum tunneling, resonances, radiative transitions, and non-adiabatic effects governed by spin-orbit or derivative coupling can be determining factors in those dynamics. Building upon progress made during a period of prior support through the DOE Early Career Program, this project combines developments in the areas of potential energy surface (PES) fitting and multistate multireference quantum chemistry to allow spectroscopically and dynamically/kinetically accurate investigations of key molecular systems (such as those mentioned above), many of which are radicals with strong multireference character and have the possibility of multiple electronic states contributing to the observed dynamics. An ongoing area of investigation is to develop general strategies for robustly convergent electronic structure theory for global multichannel reactive surfaces including diabatization of energy and other relevant surfaces such as dipole transition. Combining advances in ab initio methods with automated interpolative PES fitting allows the construction of high-quality PESs (incorporating thousands of high-level data) to be done rapidly through parallel processing on high-performance computing (HPC) clusters. In addition, new methods and approaches to electronic structure theory will be developed and tested through applications. This project will explore limitations in traditional multireference calculations (e.g., MRCI) such as those imposed by internal contraction, lack of high-order correlation treatment and poor scaling. Methods such as DMRG-based extended active-space CASSCF and various Quantum Monte Carlo (QMC) methods will be applied (including VMC/DMC and FCIQMC). Insight into the relative significance of different orbital spaces and the robustness of application of these approaches on leadership class computing architectures will be gained. Synergy with other components of this research program such as automated PES fitting and multireference quantum chemistry will be used to address challenges encountered by the standard approaches to computational thermochemistry (those being single-reference quantum chemistry and perturbative treatments of the anharmonic vibrational energy, which break down for some cases of electronic structure or floppy strongly coupled vibrational modes).
This project convened a National Academies committee to identify opportunities and research priorities at the interface of chemistry and quantum information science (QIS). The work culminated in a consensus study report that (1) articulates three fundamental research areas to advance QIS (design and synthesis of molecular qubits; measurement and control of molecular quantum systems; and experimental and computational scaling of qubit design and function), and (2) underscores the importance of cross-disciplinary collaboration, access to facilities and instrumentation, FAIR-aligned data infrastructure, and workforce development initiatives to sustain U.S. leadership in QIS. The report and all other material associated with this project can be downloaded on the project webpage: https://www.nationalacademies.org/projects/DELS-BCST-21-01 .
Two propositions concerning quantum chemistry are proposed. First, it is proposed that the nonrelativistic Schroedinger equation, where the Hamiltonian operator is associated with an assemblage of nuclei and electrons, can never be arranged to yield specific molecules in the chemists' sense. It is argued that this result is a necessary condition if the Schroedinger has relevancy to chemistry. Second, once a system is in a particular state with regard to interactions among its components (the assemblage of nuclei and electrons), it cannot spontaneously eliminate any of those interactions. This leads to a subtle form of irreversibility.
Abstract Various photoactive molecules contain motifs built on aza-aromatic heterocycles, although a detailed understanding of the excited state photophysics and photochemistry in such systems is not fully developed. To help address this issue, the non-adiabatic dynamics operating in azanaphthalenes under hexane solvation was studied following 267 nm excitation using ultrafast transient absorption spectroscopy. Specifically, the species quinoline, isoquinoline, quinazoline, quinoxaline, 1,6-naphthyridine, and 1,8-naphthyridine were investigated, providing a systematic variation in the relative positioning of nitrogen heteroatom centres within a bicyclic aromatic structure. Our results indicate considerable differences in excited state lifetimes, and in the propensity for intersystem crossingvsinternal conversion across the molecular series. The overall pattern of behaviour can be explained in terms of potential energy barriers and spin-orbit coupling effects, as demonstrated by extensive quantum chemistry calculations undertaken at the SCS-ADC(2) level of theory. The fact that quantum chemistry calculations can achieve such detailed and nuanced agreement with experimental data across a full set of six molecules exhibiting subtle variations in their composition provides an excellent example of the current state-of-the-art and is indicative of future opportunities for rational design of photoactive molecules.
Recent advances in strong light–matter interactions have revealed a wealth of new physical phenomena in molecules embedded in optical cavities, including modified chemical reactivity, altered excitation spectra, and novel quantum correlations. To describe these effects from first-principles, the field of ab initio quantum electrodynamics (QED) has emerged as a compelling extension of quantum chemistry that treats electronic and photonic degrees of freedom on equal footing. In this Perspective, we review the growing landscape of many-body QED methods, including Hartree–Fock, density functional theory (QEDFT), time-dependent DFT (QED-TDDFT), configuration interaction (QED-CI), complete active space (QED-CASSCF), coupled cluster (QED-CC), quantum Monte Carlo (QED-QMC), and density matrix renormalization group (QED-DMRG), highlighting recent developments and implementations. We further explore real-time methods, gradient and Hessian formalisms, and the integration of nonadiabatic nuclear dynamics. Applications range from benchmark simulations of polaritonic chemistry to quantum simulations on emerging quantum hardware. We conclude by outlining future directions for theory development and interdisciplinary efforts at the interface of quantum chemistry, condensed matter, and quantum optics.
The combinatorial growth of configuration interaction (CI) has long limited this formally exact quantum chemistry method to only the smallest molecules. Here, we report a numerically exact CI calculation exceeding one quadrillion (10 15 ) determinants, made possible by a lossless categorical compression strategy within the small-tensor-product distributed active space (STP-DAS) framework. This approach overcomes the traditional memory bottlenecks of CI by a numerically exact compression of the wavefunction representation and reformulating the most computationally demanding matrix–vector operations. Using this method, we performed a fully relativistic CI calculation of the ground state of HBrTe with over 10 15 complex-valued determinants in just 34.5 h on 1000 computing nodes—the largest CI calculation ever reported. We further achieved fast computation for systems with hundreds of billions of determinants on only a few compute nodes. Extensive benchmarks confirm that the method retains full numerical exactness while cutting memory and computational cost by orders of magnitude. Compared to previous state-of-the-art CI calculations, this work achieves a 1000 times increase in CI space, a 10 6 -fold increase in floating-point operations performed, and a 10 6 -fold improvement in computational speed.
This release contains data and parameters of HydraGNN-based graph foundation models trained as a result of the work published in the pre-print "Exascale Multi-Task Graph Foundation Models for Imbalanced, Multi-Fidelity Atomistic Data" by M. Lupo Pasini et al. (https://arxiv.org/abs/2604.15380). We jointly train on 16 open first-principles datasets (544+ million structures covering 85+ elements) using a multi-task architecture with per-dataset heads and a scalable ADIOS2/DDStore data pipeline. On Frontier, we execute six large-scale DeepHyper hyperparameter optimization campaigns in FP64 and promote the top-performing message-passing models to sustained 2,048-node training, yielding a PaiNN-based lead model. The version of HydraGNN used to generate the outputs provided in this release is HydraGNN v5.0 (https://github.com/ORNL/HydraGNN/releases/tag/v5.0) The list of datasets used for the training of the graph foundation model is the following: 1) Alexandria [1] 2) ANI1x [2] 3) MPTrj [3] 4) Open Catalyst 2020 (OC20) [4] 5) Open Catalyst 2022 (OC22) [5] 6) Open Catalyst 2025 (OC25) [6] 7) Open Direct ir Capture 2023 (ODAC23) [7] 8) Open Materials 2024 (OMat24) [8] 9) Open Molecules 2025 (OMol25) [9] 10) OMol25-neutral (subset of OMol25 that contains only molecules with zero total charge) 11) OMol25-non-neutral (subset of OMol25 that contains only molecules with non-zero total charge) 12) Open Polymers 2026 (OPoly2026) [10] 13) Nabla2DFT [11] 14) QCML [12] 15) QM7X [reference 13] 16) transition1x [14] Dataset references: [1] J. Schmidt et al., “A dataset of 175k stable and metastable materials calculated with the PBEsol and SCAN functionals,” Scientific Data, vol. 9, p. 64, 2022. [2] J. S. Smith et al., “The ANI-1ccx and ANI-1x data sets, coupled-cluster and density functional theory properties for molecules,” Scientific Data, vol. 7, p. 134, 2020. [Online]. Available: https: //www.nature.com/articles/s41597-020-0473-z [3] A. Jain et al., “Commentary: The Materials Project: A materials genome approach to accelerating materials innovation,” APL Materials, vol. 1, no. 1, p. 011002, 07 2013. [Online]. Available: https://doi.org/10.1063/1.4812323 [4] L. Chanussot et al., “Open catalyst 2020 (oc20) dataset and community challenges,” ACS Catalysis, vol. 11, no. 10, pp. 6059–6072, 2021. [Online]. Available: https://doi.org/10.1021/acscatal.0c04525 [5] K. Tran et al., “Open catalyst 2022 (oc22) dataset and challenges for oxidation electrocatalysts,” ACS Catalysis, vol. 13, no. 5, pp. 3066–3084, 2023. [Online]. Available: https://doi.org/10.1021/acscatal.2c05426 [6] S. J. Sahoo et al., “The open catalyst 2025 (oc25) dataset and models for solid-liquid interfaces,” arXiv preprint arXiv:2509.17862, 2025. [Online]. Available: https://arxiv.org/abs/2509.17862 [7] A. Sriram et al., “The open DAC 2023 dataset and challenges for sorbent discovery in direct air capture,” ACS Central Science, vol. 10, no. 5, pp. 923–941, 2024. [8] L. Barroso-Luque et al., “Open materials 2024 (omat24) inorganic materials dataset and models,” 2024. [Online]. Available: https://arxiv.org/abs/2410.12771 [9] D. S. Levine et al., “The open molecules 2025 (OMol25) dataset, evaluations, and models,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08762 [10] D. S. Levine et al., The open polymers 2026 (OPoly26) dataset and evaluations,” arXiv preprint arXiv:2512.23117, 2025. [Online]. Available: https://arxiv.org/abs/2512.23117 [11] K. Khrabrov et al., “Nabla2dft: A universal quantum chemistry dataset of drug-like molecules and a benchmark for neural network potentials,” in NeurIPS 2024 Datasets and Benchmarks Track, 2024. [Online]. Available: https://openreview.net/forum?id=ElUrNM9U8c [12] S. Ganscha et al., “The QCML dataset, quantum chemistry reference data from 33.5M DFT and 14.7B semi-empirical calculations,” Scientific Data, vol. 12, p. 406, 2025. [13] J. Hoja et al., “QM7-X, a comprehensive dataset of quantum-mechanical properties spanning the chemical space of small organic molecules,” Scientific Data, vol. 8, p. 43, 2021. [Online]. Available: https://www.nature.com/articles/s41597-021-00812-2 [14] M. Schreiner et al., “Transition1x - a dataset for building generalizable reactive machine learning potentials,” Scientific Data, vol. 9, p. 779, 2022. The folder "datasets_ADIOS2_format" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "datasets_ADIOS2_format" directory contains 2 sub-directories, one for the version "v1" of the datasets and one for the version "v2" of the datasets. The version "v1" of the datasets provides values of the total energy as they are extracted from the original data as it was released by the respective institutions. The version "v2" of the datasets provides values of the energy that have been realigned. The realignment was performed by training a linear regression model that predicts the total energy as a function of the chemical composition of the atomistic structure, and then subtract such prediction from the original value of the total energy. Both folders "v1" and "v2" contain 16 sub-directories, each corresponding to an ADIOS2-formatted dataset The folder "DeepHyper-results" contains the configurational files and model's parameters for all the 186 HPO trials that were successfully completed by the scalable hyperparameter optimization (HPO) runs on Frontier. The content of the folder "DeepHyper-results" I structured as follows: 1) task-list.txt: list of mpnn name, jobid, and deephyper task id 2) gfm_${MPNN}_${JOBID}_0.${TASKID}: run directory with checkpoint files 3) gfm_${MPNN}: deephyper summary directory (*.csv) for each specific MPNN type 4) deephyper-experiment-${JOBID}: output and error logs for each job The file "deephyper-sorted.csv" contains the details of each HydraGNN model built and tested by HPO, obtained by merging the (*.csv) filed from each HPO run executed. Out of all the HPO trials, we selected 10 to continue the training of the respective HydraGNN models. Due to limited computational budget available in the LRN070 allocation we could not complete the training till convergence for all these 10 selected models. The folder "models" contains multiple sub-folders, one per each HydraGNN model trained. Each model sub-folder contains the parameters of each HydraGNN model, with multiple checkpoint-restarts. The list of sub-folders are as follows: 1) multidataset_hpo-BEST1-fp64 2) multidataset_hpo-BEST2-fp64 3) multidataset_hpo-BEST3-fp64 4) multidataset_hpo-BEST4-fp64 5) multidataset_hpo-BEST5-fp64 6) multidataset_hpo-BEST6-fp64 7) multidataset_hpo-BEST7-fp64 8) multidataset_hpo-BEST8-fp64 9) multidataset_hpo-BEST9-fp64 10) multidataset_hpo-BEST10-fp64 Within each one of these folders, additional auxiliary log files are provided with descriptions about how the training proceeded. The lead PaiNN-model is contained inside "multidataset_hpo-BEST6-fp64". The file "mlp_branch_weights" contains the parameters of the multi-layer perceptron (MLP) used to reconcile the predictions of the 16 output decoding heads of the HydragNN architectures. The MLP takes in input the chemical composition of the atomistic structure and predicts averaging weights to linearly mix the predictions of each output decoding head toward consolidating them into a single one. The folder "1.1billion-structure-inference" contains 1.1 billion atomistic structures randomly generated. Each structures is associated with energy and forces predicted with the lead-PaiNN model combined with the MLP model for reconciliation of the multi-branch predictions generated by the 16 output decoding heads. The folder "1.1billion-structure-inference" contains 9,300 (*.tar.gz) subdirectories, one per Frontier compute node used to execute the inference at exascale. Once uncompressed, each (*.tar.gz) subdirectory contains an ADIOS2 (*.bp) file container, where each atomistic structure is stored as a PyTorch-Geometric Data object. The file "export_dataset_environment_variables.sh" contains the environment variables that need to be set before running the HydraGNN code to reproduce the results provided in this dataset release. The code that can be used to load the ADIOS2 files, load HydraGNN models, and run inference is available at: https://github.com/ORNL/HydraGNN/releases/tag/v5.0