Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “tensor network algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

97 records · Page 6

Ground-state-based model reduction with unitary circuits

Here, we present a method to numerically obtain low-energy effective models based on a unitary transformation of the ground state. The algorithm finds a unitary circuit that transforms the ground state of the original model to a projected wavefunction with only the low-energy degrees of freedom. The effective model can then be derived using the unitary transformation encoded in the circuit. We test our method on the one-dimensional and two-dimensional square-lattice Hubbard model at half-filling, and obtain more accurate effective spin models than the standard perturbative approach.

Hubbard model↗

HydraGNN_Predictive_GFM_2024 - Ensemble of predictive graph foundation models for ground state atomistic materials modeling

We provide the ensemble of fifteen pre-trained graph foundation models (GFMs) for atomistic materials modeling applications. Each one of the fifteen GFMs has been trained on five open-source datasets that (once aggregated) amount to over 154 million atomistic structures, which cover over two-thirds of the natural elements of the periodic table and that comprises a broad set of organic and inorganic compounds. This vast set of atomistic structures comprises ground state configurations that are dynamically stable (i.e., equilibrated structures with atomic forces approximately close to zero values) as well as dynamically unstable structures (i.e., non-equilibrium structures with non-negligible non-zero values of atomic forces). The ensemble of datasets aggregated does NOT include excited states. The datasets have been curated to remove atomistic structures with spectral norm of the force tensor above 100 eV/angstrom. Moreover, a linear term of the energy was computed for each dataset using a linear regression model that uses the chemical concentration of each natural element as regressor. The linear term predicted by the linear regression model has been subtracted from each original energy value to perform a re-alignment of the energy values across different electronic structures approximation theories performed to generate the diverse multi-source, multi-fidelity datasets. The folder "ADIOS_files" contains the set of pre-processed datasets in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used for the development and training of GFMs in this work. The "ADIOS_files" directory contains 6 sub-directories named as follows: - ANI1x-v3.bp - MPTrj-v3.bp - OC2020-20M-v3.bp - OC2020-v3.bp - OC2022-v3.bp - qm7x-v3.bp Each sub-directory contains the pre-processed datasets converted in Adaptable I/O System (ADIOS) format (https://www.exascaleproject.org/research-project/adios/) that have been used to the development, training, and performance testing of the ensemble go predictive graph foundation models. Each GFM was developed using HydraGNN (https://github.com/ORNL/HydraGNN) as underlying graph neural network (GNN) architecture. The multi-task learning (MTL) capability of HydraGNN was used to simultaneously train the GFMs on labeled values for direct predictions of energy (a total system property of an atomistic structure that measures the chemical stability) and atomic forces (an atomic level property of an atomistic structure that measures the dynamical stability). The hyper parameters of the GFM have been tuned using scalable hyperparameter optimization (HPO) algorithms implemented in the software DeepHyper (https://github.com/deephyper/deephyper). The pre-training of each HPO trial was performed using distributed data parallelism (DDP) to scale the training across 128 compute nodes of the exascale OLCF supercomputer Frontier. Each HPO trial was trained only for 10 epochs and an early stopping was performed to avoid wasting significant computational resources on GNN architectures that were clearly underperforming. For each HPO trial, the 'omnistat' tool developed by (AMD Research - Advanced Micro Device) was used to measure the total energy consumption in kWh. The ensemble of GFMs was obtained by selecting the fifteen best performing HPO trials. Four models have been selected for their clear advantage in accuracy, and these are the GFMs with IDs 229, 156, 147, 260. Additional eleven models have been selected based on judicious balance between accuracy and energy consumption needed for training, and these are the GFMs with IDs 165, 78, 137, 1, 175, 171, 181, 67, 179, 167, 351. Each selected GFM of the ensemble was continued to cumulate a total of at most 30 epochs. In some cases, the total number of epochs actually performed was les than 30 due to two combined factors: (1) the size of the GFM (i.e., the number of model parameters to train) and (2) the total wall-clock time for which the computational resources could be allocated on OLCF-Frontier. The "Ensemble_of_models" directory contains 15 sub-directories named as follows: - gfm_0.229 - gfm_0.156 - gfm_0.147 - gfm_0.260 - gfm_0.165 - gfm_0.78 - gfm_0.137 - gfm_0.1 - gfm_0.175 - gfm_0.171 - gfm_0.181 - gfm_0.67 - gfm_0.179 - gfm_0.167 - gfm_0.351 Each one of these sub-directories refers to one of the fifteen HPO trials that have been selected to continue the pre-training with at most 30 epochs. With each sub-directory associated with a specific HPO trial, the following files can be found: - config.json: file for argument parsing to develop and train an HydraGNN architecture - gfm_0.ID_epoch_N.pk: file with model parameters for HPO ID trial after N epochs of training The ensemble of fifteen GFM architectures was used for (1) ensemble averaging to stabilize the predictions of energy and atomic forces after pre-training for post-processing analysis and (2) ensemble uncertainty quantification (UQ). The code used to develop, pre-train, and load the pre-trained models for post-processing analysis is available on the ORNL-GitHub at the following link: https://github.com/ORNL/HydraGNN/tree/Predictive_GFM_2024

36 MATERIALS SCIENCE↗

Quantum many-body simulations of the two-dimensional Fermi-Hubbard model in ultracold optical lattices

Understanding quantum many-body states of correlated electrons is one main theme in modern condensed-matter physics. Given that the Fermi-Hubbard model, the prototype of correlated electrons, was recently realized in ultracold optical lattices, it is highly desirable to have controlled numerical methodology to provide precise finite-temperature results upon doping to directly compare with experiments. Here, we demonstrate the exponential tensor renormalization group (XTRG) algorithm [Chen et al., Phys. Rev. X 8, 031082 (2018)], complemented by independent determinant quantum Monte Carlo, offers a powerful combination of tools for this purpose. XTRG provides full and accurate access to the density matrix and thus various spin and charge correlations, down to an unprecedented low temperature of a few percent of the tunneling energy. Finally, we observe excellent agreement with ultracold fermion measurements at both half filling and finite doping, including the sign-reversal behavior in spin correlations due to formation of magnetic polarons, and the attractive hole-doublon and repulsive hole-hole pairs that are responsible for the peculiar bunching and antibunching behaviors of the antimoments.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Quantum-inspired method for solving the Vlasov-Poisson equations

Kinetic simulations of collisionless (or weakly collisional) plasmas using the Vlasov equation are often infeasible due to high-resolution requirements and the exponential scaling of computational cost with respect to dimension. Recently, it has been proposed that matrix product state (MPS) methods, a quantum-inspired but classical algorithm, can be used to solve partial differential equations with exponential speed-up, provided that the solution can be compressed and efficiently represented as a MPS within some tolerable error threshold. Here, in this work, we explore the practicality of MPS methods for solving the Vlasov-Poisson equations for systems with one coordinate in space and one coordinate in velocity, and find that important features of linear and nonlinear dynamics, such as damping or growth rates and saturation amplitudes, can be captured while compressing the solution significantly. Furthermore, by comparing the performance of different mappings of the distribution functions onto the MPS, we develop an intuition of the MPS representation and its behavior in the context of solving the Vlasov-Poisson equations, which will be useful for extending these methods to higher-dimensional problems.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Performance of the rigorous renormalization group for first-order phase transitions and topological phases

Expanding and improving the repertoire of numerical methods for studying quantum lattice models is an ongoing focus in many-body physics. While the density matrix renormalization group (DMRG) has been established as a practically useful algorithm for finding the ground state in one-dimensional systems, a provably efficient and accurate algorithm remained elusive until the introduction of the rigorous renormalization group (RRG) by Landau [Nat. Phys. 11, 566 (2015)1745-247310.1038/nphys3345]. In this paper, we study the accuracy and performance of a numerical implementation of RRG at first-order phase transitions and in symmetry-protected topological phases. Our study is motivated by the question of when RRG might provide a useful complement to the more established DMRG technique. In particular, despite its general utility, DMRG can give unreliable results near first-order phase transitions and in topological phases, since its local update procedure can fail to adequately explore (near-)degenerate manifolds. Furthermore, the rigorous theoretical underpinnings of RRG, meanwhile, suggest that it should not suffer from the same difficulties. We show this optimism is justified, and that RRG indeed determines well-ordered, accurate energies even when DMRG does not. Moreover, our performance analysis indicates that in certain circumstances seeding DMRG with states determined by coarse runs of RRG may provide an advantage over simply performing DMRG.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Automatic Generation of High-Performance Convolution Kernels on ARM CPUs for Deep Learning

In this work, we present FastConv, a template-based code auto-generation open source library that can automatically generate high-performance deep learning convolution kernels of arbitrary matrices/tensors shapes. FastConv is based on the Winograd algorithm, which is reportedly the highest performing algorithm for the time-consuming convolution layers of convolutional neural networks. ARM CPUs cover a wide range designs and specifications, from embedded devices to HPC-grade CPUs. The leads to the dilemma of how to consistently optimize Winograd-based convolution solvers for convolution layers of different shapes. FastConv addresses this problem by using templates to auto-generate multiple shapes of tuned kernels variants suitable for skinny tall matrices. As a performance portable library, FastConv transparently searches for the best combination of kernel shapes, cache tiles, scheduling of loop orders, packing strategies, access patterns, and online/offline computations. Auto-tuning is used to search the parameter configuration space for the best performance for a given target architecture and problem size. The experiments with layer-wise evaluation on the VGG--16 model confirms a 1.25x performance gains is got by tuning the Winograd library. Integrated comparison results shows 1.02x to 1.40x, 1.14x to 2.17x, and 1.22x and 2.48x speedup is achieved over NNPACK, Arm NN, and FeatherCNN on the Kunpeng 920 beside few cases. Furthermore, problem size performance portability experiments with various convolution shapes shows that FastConv achieves 1.2x to 1.7x speedup and 2x to 22x speedup over NNPACK and ARM NN inference engine using Winograd on Kunpeng 920 . CPU performance portability evaluation on the VGG--16 show an average speedup over NNPACK of 1.42x, 1.21x, 1.26x, 1.37x, 2.26x, and 11.02x is observed on Kunpeng 920, Snapdragon 835, 855, 888, Apple M1, and AWS Graviton2, respectively.

97 MATHEMATICS AND COMPUTING↗

Lie-algebraic classical simulations for quantum computing

The classical simulation of quantum dynamics plays an important role in our understanding of quantum complexity and in the development of quantum technologies. Efficient techniques such as those based on the Gottesman-Knill theorem for Clifford circuits, tensor networks for low entanglement-generating circuits, or Wick's theorem for fermionic Gaussian states have become central tools in quantum computing. In this work, we contribute to this body of knowledge by presenting a framework for classical simulations, dubbed “𝔤-sim”, which is based on the underlying Lie algebraic structure of the dynamical process. When the dimension of the algebra grows at most polynomially in the system size, there exist observables for which the simulation is efficient. Indeed, we show that 𝔤-sim enables new regimes for classical simulations, is able to deal with certain forms of noise in the evolution, as well as can be used to tackle several paradigmatic variational and nonvariational quantum computing tasks. For the former, we perform Lie-algebraic simulations to train and optimize parametrized quantum circuits (thus effectively showing that some variational models can be dequantized), design enhanced parameter initialization strategies, solve tasks of quantum circuit synthesis, and train a quantum-phase classifier. For the latter, we report large-scale noiseless and noisy simulations on benchmark problems. By comparing the limitations of 𝔤-sim and certain Wick's theorem-based simulations, we find that the two methods become inefficient for different types of states or observables, hinting at the existence of distinct, nonequivalent resources for classical simulation.

97 MATHEMATICS AND COMPUTING↗