Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “matrix algebra”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Numerical solution of large scale Hartree–Fock–Bogoliubov equations

The Hartree–Fock–Bogoliubov (HFB) theory is the starting point for treating superconducting systems. However, the computational cost for solving large scale HFB equations can be much larger than that of the Hartree–Fock equations, particularly when the Hamiltonian matrix is sparse, and the number of electrons N is relatively small compared to the matrix size N b . We first provide a concise and relatively self-contained review of the HFB theory for general finite sized quantum systems, with special focus on the treatment of spin symmetries from a linear algebra perspective. We then demonstrate that the pole expansion and selected inversion (PEXSI) method can be particularly well suited for solving large scale HFB equations. For a Hubbard-type Hamiltonian, the cost of PEXSI is at most $\mathcal{O}$(N b 2 ) for both gapped and gapless systems, which can be significantly faster than the standard cubic scaling diagonalization methods. We show that PEXSI can solve a two-dimensional Hubbard-Hofstadter model with N b up to 2.88 × 10 6 , and the wall clock time is less than 100 s using 17 280 CPU cores. Finally, this enables the simulation of physical systems under experimentally realizable magnetic fields, which cannot be otherwise simulated with smaller systems.

97 MATHEMATICS AND COMPUTING↗

Computation of graph hitting time moments; Chapel code implementation.

The project that developed this is unclassified, with the mandate to produce open source code. This code computes the hitting time moments of a graph using a linear algebra configuration. The goal of this work is to explore the performance capabilities of the Chapel programming language. Toward that end, we generate random adjacency matrices which represent a random graph. The code can also read in an adjacency matrix from a file. The main computation is the Conjugate Gradient method.SAND2020-12651 M. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Barrett, Richard↗

Emergent Superconductivity and Competing Charge Orders in Hole-Doped Square-Lattice t – J Model

The square-lattice Hubbard and closely related t-J models are considered as basic paradigms for understanding strong correlation effects and unconventional superconductivity (SC). Recent large-scale density matrix renormalization group simulations on the extended t-J model have identified d-wave SC on the electron-doped side (with the next-nearest-neighbor hopping t 2 > 0) but a dominant charge density wave (CDW) order on the hole-doped side (t 2 < 0), which is inconsistent with the SC of hole-doped cuprate compounds. We re-examine the ground-state phase diagram of the extended t-J model by employing the state-of-the-art density matrix renormalization group calculations with much enhanced bond dimensions, allowing more accurate determination of the ground state. On six-leg cylinders, while different CDW phases are identified on the hole-doped side for the doping range δ = 1/16 – 1/8, a SC phase emerges at a lower doping regime, with algebraically decaying pairing correlations and d-wave symmetry. On the wider eight-leg systems, the d-wave SC also emerges on the hole-doped side at the optimal 1=8 doping, demonstrating the winning of SC over CDW by increasing the system width. Furthermore, our results not only suggest a new path to SC in general t-J model through weakening the competing charge orders, but also provide a unified understanding on the SC of both hole- and electron-doped cuprate superconductors.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Computing rank‐revealing factorizations of matrices stored out‐of‐core

This paper describes efficient algorithms for computing rank-revealing factorizations of matrices that are too large to fit in main memory (RAM), and must instead be stored on slow external memory devices such as disks (out-of-core or out-of-memory). Traditional algorithms for computing rank-revealing factorizations (such as the column pivoted QR factorization and the singular value decomposition) are very communication intensive as they require many vector-vector and matrix-vector operations, which become prohibitively expensive when data is not in RAM. Randomization allows to reformulate new methods so that large contiguous blocks of the matrix are processed in bulk. The paper describes two distinct methods. The first is a blocked version of column pivoted Householder QR, organized as a “left-looking” method to minimize the number of the expensive write operations. The second method results employs a UTV factorization. It is organized as an algorithm-by-blocks to overlap computations and I/O operations. As it incorporates power iterations, it is much better at revealing the numerical rank. Numerical experiments on several computers demonstrate that the new algorithms are almost as fast when processing data stored on slow memory devices as traditional algorithms are for data stored in RAM.

97 MATHEMATICS AND COMPUTING↗

Accelerating GNNs on GPU Sparse Tensor Cores through N:M Sparsity-Oriented Graph Reordering

Recent advancements in GPU hardware support have introduced the capability to leverage N:M sparse patterns for substantial performance gains. Graphs in Graph Neural Networks (GNNs) are typically sparse, but the sparsity is often irregular, not conforming to such sparse patterns. In this paper, we propose a novel graph reordering algorithm, the first of its kind, to reshape irregular graph data into the N:M structured sparse pattern at the tile level, allowing linear-algebra-based graph operations in GNNs to benefit from the N:M sparse hardware. The optimization is lossless, maintaining the accuracy of GNN. It can remove 98-100\% violations of the N:M sparse patterns at the vector level, and increase the proportion of conforming graphs in SuiteSparse collection from 5-9\% to 88.7-93.5\%. On A100 GPUs, the optimization accelerates Sparse Matrix Matrix (SpMM) by up to 43X (2.3X -- 7.5X on average) and speeds up the key graph operations in GNNs on real graphs by as much as 8.6X (3.5X on average).

artificial intelligence, graph neural networks↗

An Algebraic Quantum Circuit Compression Algorithm for Hamiltonian Simulation

Quantum computing is a promising technology that harnesses the peculiarities of quantum mechanics to deliver computational speedups for some problems that are intractable to solve on a classical computer. Current generation noisy intermediate-scale quantum (NISQ) computers are severely limited in terms of chip size and error rates. Shallow quantum circuits with uncomplicated topologies are essential for successful applications in the NISQ era. In this work, based on matrix analysis, we derive localized circuit transformations to efficiently compress quantum circuits for simulation of certain spin Hamiltonians known as free fermions. The depth of the compressed circuits is independent of simulation time and grows linearly with the number of spins. The proposed numerical circuit compression algorithm behaves backward stable and scales cubically in the number of spins enabling circuit synthesis beyond O(10 3 ) spins. The resulting quantum circuits have a simple nearest-neighbor topology, which makes them ideally suited for NISQ devices.

Hamiltonian simulation↗

CSPlib: A performance portable parallel software toolkit for analyzing complex kinetic mechanisms

Computational singular perturbation (CSP) is a method to analyze dynamical systems. It targets the decoupling of fast and slow dynamics using an alternate linear expansion of the right-hand side of the governing equations based on eigenanalysis of the associated Jacobian matrix. This representation facilitates diagnostic analysis, detection and control of stiffness, and the development of simplified models. For this work, we have implemented CSP in a C++ open-source library CSPlib using the Kokkos parallel programming model to address portability across diverse heterogeneous computing platforms, i.e., multi/many-core CPUs and GPUs. We describe the CSPlib implementation and present its computational performance across different computing platforms using several test problems. Specifically, we test the CSPlib performance for a constant pressure ignition reactor model on different architectures, including IBM Power 9, Intel Xeon Skylake, and NVIDIA V100 GPU. The size of the chemical kinetic mechanism is varied in these tests. As expected, the Jacobian matrix evaluation, the eigensolution of the Jacobian matrix, and matrix inversion are the most expensive computational tasks. When considering the higher throughput characteristic of GPUs, GPUs performs better for small matrices with higher occupancy rate. CPUs gain more advantages from the higher performance of well-tuned and optimized linear algebra libraries such as OpenBLAS.

97 MATHEMATICS AND COMPUTING↗

Celestial amplitudes from UV to IR

Celestial amplitudes represent 4D scattering of particles in boost, rather than the usual energy-momentum, eigenstates and hence are sensitive to both UV and IR physics. We show that known UV and IR properties of quantum gravity translate into powerful constraints on the analytic structure of celestial amplitudes. For example the soft UV behavior of quantum gravity is shown to imply that the exact four-particle scattering amplitude is meromorphic in the complex boost weight plane with poles confined to even integers on the negative real axis. Would-be poles on the positive real axis from UV asymptotics are shown to be erased by a flat space analog of the AdS resolution of the bulk point singularity. The residues of the poles on the negative axis are identified with operator coefficients in the IR effective action. Far along the real positive axis, the scattering is argued to grow exponentially according to the black hole area law. Exclusive amplitudes are shown to simply factorize into conformally hard and conformally soft factors. The soft factor contains all IR divergences and is given by a celestial current algebra correlator of Goldstone bosons from spontaneously broken asymptotic symmetries. The hard factor describes the scattering of hard particles together with the boost-eigenstate clouds of soft photons or gravitons required by asymptotic symmetries. These provide an IR safe S-matrix for the scattering of hard particles.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

PAO 1.0: A Python Library for Adversarial Optimization

PAO is a Python-based package for Adversarial Optimization. The goal of this package is to provide a general modeling and analysis capability for bilevel, trilevel and other multilevel optimization forms that express adversarial dynamics. PAO integrates two different modeling abstractions: 1. Algebraic models extend the modeling concepts in the Pyomo algebraic modeling language to express problems with an intuitive algebraic syntax. Thus, we expect that this modeling abstraction will commonly be used by PAO end-users. 2. Compact models express objective and constraints in a manner that is typically used to express the mathematical form of these problems (e.g. using vector and matrix data types). PAO denes custom Multilevel Problem Representations (MPRs) that simplify the implementation of solvers for bilevel, trilevel and other multilevel optimization problems.

97 MATHEMATICS AND COMPUTING↗

Effective rationality for local unitary invariants of mixed states of two qubits

Abstract We calculate the field of rational local unitary invariants for mixed states of two qubits, by employing methods from algebraic geometry. We prove that this field is rational (i.e. purely transcendental), and that it is generated by nine algebraically independent polynomial invariants. We do so by constructing a relative section, in the sense of invariant theory, whose Weyl group is a finite abelian group. From this construction, we are able to give explicit expressions for the generating invariants in terms of the Bloch matrix representation of mixed states of two qubits. We also prove similar rationality results for the local unitary invariants of symmetrically mixed states of two qubits. We also provide a sketch of how to generalize our results to the case of an arbitrary number of qubits. Our results apply to both complex-valued and real-valued invariants.

Physics↗

Design Considerations for GPU-based Mixed Integer Programming on Parallel Computing Platforms

Mixed Integer Programming (MIP) is a powerful abstraction in combinatorial optimization that finds real-life application across many significant sectors. The recent proliferation of graphical processing unit (GPU)-based accelerated computing architectures in large-scale parallel computing or supercomputing presents new opportunities as well as challenges in the advancement of MIP solver technology to effectively use the new accelerated computing platforms and scale to large parallel systems. Here, we recount the conventional processor-based strategies and focus on configurations where the most promising intersection lies between parallel MIP solver approaches and the specific strengths of accelerated parallel platforms. We note that the best potential lies in solving problems whose individual matrix sizes (of the linear program relaxation) fit entirely within one accelerator's memory and whose branch-and-bound (or branch-and-cut) trees cannot be fully contained within a small number of computational nodes. Additionally, we identify ideal features of computational linear algebra support on GPU accelerators that would help advance this direction of scalable parallel solution of MIP problems on GPU-based accelerated computing architectures.

Perumalla, Kalyan↗

On Compatible Transfer Operators in Nonsymmetric Algebraic Multigrid

The standard goal for an effective algebraic multigrid (AMG) algorithm is to develop relaxation and coarse-grid correction schemes that attenuate complementary error modes. In the nonsymmetric setting, coarse-grid correction &#x3A0; will almost certainly be nonorthogonal (and divergent) in any known standard product, meaning ∥&#x3A0;∥ > 1. This introduces a new consideration, that one wants coarse-grid correction to be as close to orthogonal as possible, in an appropriate norm. In addition, due to nonorthogonality, &#x3A0; may actually amplify certain error modes that are in the range of interpolation. Relaxation must then not only be complementary to interpolation, but also rapidly eliminate any error amplified by the nonorthogonal correction, or the algorithm may diverge. Here this paper develops analytic formulae on how to construct “compatible” transfer operators in nonsymmetric AMG such that ∥&#x3A0;∥ = 1 in some standard matrix-induced norm. Discussion is provided on different options for the norm in the nonsymmetric setting, the relation between “ideal” transfer operators in different norms, and insight into the convergence of nonsymmetric reduction-based AMG.

97 MATHEMATICS AND COMPUTING↗

Extended Lagrangian Born–Oppenheimer molecular dynamics using a Krylov subspace approximation

It is shown how the electronic equations of motion in extended Lagrangian Born–Oppenheimer molecular dynamics simulations can be integrated using low-rank approximations of the inverse Jacobian kernel. This kernel determines the metric tensor in the harmonic oscillator extension of the Lagrangian that drives the evolution of the electronic degrees of freedom. The proposed kernel approximation is derived from a pseudoinverse of a low-rank estimate of the Jacobian, which is expressed in terms of a generalized set of directional derivatives with directions that are given from a Krylov subspace approximation. The approach allows a tunable and adaptive approximation that can take advantage of efficient preconditioning techniques. The proposed kernel approximation for the integration of the electronic equations of motion makes it possible to apply extended Lagrangian first-principles molecular dynamics simulations to a broader range of problems, including reactive chemical systems with numerically sensitive and unsteady charge solutions. This can be achieved without requiring exact full calculations of the inverse Jacobian kernel in each time step or relying on iterative non-linear self-consistent field optimization of the electronic ground state prior to the force evaluations as in regular direct Born–Oppenheimer molecular dynamics. We note the low-rank approximation of the Jacobian is directly related to Broyden’s class of quasi-Newton algorithms and Jacobian-free Newton–Krylov methods and provides a complementary formulation for the solution of nonlinear systems of equations.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Microstate distinguishability, quantum complexity, and the eigenstate thermalization hypothesis

In this work, we use quantum complexity theory to quantify the difficulty of distinguishing eigenstates obeying the eigenstate thermalization hypothesis (ETH). After identifying simple operators with an algebra of low-energy observables and tracing out the complementary high-energy Hilbert space, the ETH leads to an exponential suppression of trace distance between the coarse grained eigenstates. Conversely, we show that an exponential hardness of distinguishing between states implies ETH-like matrix elements. Finally, the BBBV lower bound on the query complexity of Grover search then translates directly into a complexity-theoretic statement lower bounding the hardness of distinguishing these reduced states

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Efficient Integration of Algebraic Constraint for Exponential Time Integration

This work is a continuation of a previous project where the high efficiency of exponential integration methods for DRMHD systems was demonstrated. Often algebraic constraints must also be enforced for DRMHD systems of interest. Straightforward application of exponential methods to DRMHD equations with constraints leads to prohibitively computationally expensive methods. In this work, we propose new exponential schemes that allow the constraints to be removed from evaluation of exponential matrix functions which drastically reduces computational complexity by eliminating the need to perform computations enforcing constraints in exponential calculations while still preserving a high order of accuracy and allowing for a large time step even when the problem is stiff. This idea is similar to the W-methods and is achieved by carefully designing a method with the desired order of accuracy, even with an incomplete Jacobian matrix used as an argument of exponential-like functions. The constraints are accounted for by including them in the evaluation of the right-hand-side forcing function of the spatially discretized system. We study performance of the new methods on test problems and outline future research directions that this work opens.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Parametric reduced order models for graded lattice structures

Graded lattice structures, characterized by smoothly varying mechanical properties, hold significant promise for optimizing material distribution in advanced engineering applications. However, accurately modeling these structures poses substantial computational challenges due to the continuous geometric variations within their unit cells. Here, to address these challenges, this paper introduces a novel Efficient Reduced Order Model (EROM) that integrates the Matrix Discrete Empirical Interpolation Method (MDEIM) and Discrete Empirical Interpolation Method (DEIM) with polynomial regression to efficiently manage geometric parametrization in lattice structures. Unlike traditional reduced order models (ROMs) that require extensive precomputed libraries for each geometric configuration, our approach enables continuous geometric variations through a flexible algebraic formulation, significantly reducing computational costs while preserving high accuracy. The method constructs projection matrices for individual unit cells that can be efficiently assembled into global systems, leveraging the repetitive nature of lattice structures. Numerical studies demonstrate that our EROM achieves displacement errors below 1% and von Mises stress prediction errors below 4%, coupled with computational speedups exceeding two orders of magnitude compared to full-order simulations. The proposed method's modularity and scalability make it particularly suitable for design optimization and real-time simulation of functionally graded lattice structures, with applications spanning aerospace to biomedical engineering.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A supernodal all-pairs shortest path algorithm

We show how to exploit graph sparsity in the Floyd-Warshall algorithm for the all-pairs shortest path (Apsp) problem. Floyd-Warshall is an attractive choice for Apsp on high-performing systems due to its structural similarity to solving dense linear systems and matrix multiplication. However, if sparsity of the input graph is not properly exploited, Floyd-Warshall will perform unnecessary asymptotic work and thus may not be a suitable choice for many input graphs. To overcome this limitation, the key idea in our approach is to use the known algebraic relationship between Floyd-Warshall and Gaussian elimination, and import several algorithmic techniques from sparse Cholesky factorization, namely, fill-in reducing ordering, symbolic analysis, supernodal traversal, and elimination tree parallelism. When combined, these techniques reduce computation, improve locality and enhance parallelism. We implement these ideas in an efficient shared memory parallel prototype that is orders of magnitude faster than an efficient multi-threaded baseline Floyd-Warshall that does not exploit sparsity. Our experiments suggest that the Floyd-Warshall algorithm can compete with Dijkstra's algorithm (the algorithmic core of Johnson's algorithm) for several classes sparse graphs.

Sao, Piyush↗

Hierarchical off-diagonal low-rank approximation of Hessians in inverse problems, with application to ice sheet model initialization

Obtaining lightweight and accurate approximations of discretized objective functional Hessians in inverse problems governed by partial differential equations (PDEs) is essential to make both deterministic and Bayesian statistical large-scale inverse problems computationally tractable. The cubic computational complexity of dense linear algebraic tasks, such as Cholesky factorization, that provide a means to sample Gaussian distributions and determine solutions of Newton linear systems is a computational bottleneck at large-scale. These tasks can be reduced to log-linear complexity by utilizing hierarchical off-diagonal low-rank (HODLR) matrix approximations. In this work, we show that a class of Hessians that arise from inverse problems governed by PDEs are well approximated by the HODLR matrix format. In particular, we study inverse problems governed by PDEs that model the instantaneous viscous flow of ice sheets. In these problems, we seek a spatially distributed basal sliding parameter field such that the flow predicted by the ice sheet model is consistent with ice sheet surface velocity observations. Here, we demonstrate the use of HODLR Hessian approximation to efficiently sample the Laplace approximation of the posterior distribution with covariance further approximated by HODLR matrix compression. Computational studies are performed which illustrate ice sheet problem regimes for which the Gauss–Newton data-misfit Hessian is more efficiently approximated by the HODLR matrix format than the low-rank (LR) format. We then demonstrate that HODLR approximations can be favorable, when compared to global LR approximations, for large-scale problems by studying the data-misfit Hessian associated with inverse problems governed by the first-order Stokes flow model on the Humboldt glacier and Greenland ice sheet.

97 MATHEMATICS AND COMPUTING↗