Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “matrix algebra”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Solving a class of infinite-dimensional tensor eigenvalue problems by translational invariant tensor ring approximations

Here, we examine a method for solving an infinite-dimensional tensor eigenvalue problem Hx = λx, where the infinite-dimensional symmetric matrix H exhibits a translational invariant structure. We provide a formulation of this type of problem from a numerical linear algebra point of view and describe how a power method applied to e -Ht is used to obtain an approximation to the desired eigenvector. This infinite-dimensional eigenvector is represented in a compact way by a translational invariant infinite Tensor Ring (iTR). Low rank approximation is used to keep the cost of subsequent power iterations bounded while preserving the iTR structure of the approximate eigenvector. We show how the averaged Rayleigh quotient of an iTR eigenvector approximation can be efficiently computed and introduce a projected residual to monitor its convergence. In the numerical examples, we illustrate that the norm of this projected iTR residual can also be used to automatically modify the time step to ensure accurate and rapid convergence of the power method.

97 MATHEMATICS AND COMPUTING↗

Memristive linear algebra

The advent of memristive devices offers a promising avenue for efficient and scalable analog computing, particularly for linear algebra operations essential in various scientific and engineering applications. This paper investigates the potential of memristive crossbars in implementing matrix inversion algorithms. We explore both static and dynamic approaches, emphasizing the advantages of analog and in-memory computing for matrix operations beyond multiplication. In particular, we demonstrate that the electrical properties of memristive crossbars uniquely suit them for the evolution of a family of matrix exponentials, which can be exploited for the efficient computation of matrix inverses and online solutions for linear problems. Our results demonstrate that memristive arrays can reduce computational complexity. We also study power consumption and show a tradeoff between precision and energy. Furthermore, we address the challenges of device variability, precision, and scalability, providing insights into the practical implementation of these algorithms.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Topology and grid adaption for high-speed flow computations

This study investigates the effects of grid topology and grid adaptation on numerical solutions of the Navier-Stokes equations. In the first part of this study, a general procedure is presented for computation of high-speed flow over complex three-dimensional configurations. The flow field is simulated on the surface of a Butler wing in a uniform stream. Results are presented for Mach number 3.5 and a Reynolds number of 2,000,000. The O-type and H-type grids have been used for this study, and the results are compared together and with other theoretical and experimental results. The results demonstrate that while the H-type grid is suitable for the leading and trailing edges, a more accurate solution can be obtained for the middle part of the wing with an O-type grid. In the second part of this study, methods of grid adaption are reviewed and a method is developed with the capability of adapting to several variables. This method is based on a variational approach and is an algebraic method. Also, the method has been formulated in such a way that there is no need for any matrix inversion. This method is used in conjunction with the calculation of hypersonic flow over a blunt-nose body. A movie has been produced which shows simultaneously the transient behavior of the solution and the grid adaption.

Abolhassani, Jamshid S.↗

Geopotential Error Analysis from Satellite Gradiometer and Global Positioning System Observables on Parallel Architecture

The recovery of a high resolution geopotential from satellite gradiometer observations motivates the examination of high performance computational techniques. The primary subject matter addresses specifically the use of satellite gradiometer and GPS observations to form and invert the normal matrix associated with a large degree and order geopotential solution. Memory resident and out-of-core parallel linear algebra techniques along with data parallel batch algorithms form the foundation of the least squares application structure. A secondary topic includes the adoption of object oriented programming techniques to enhance modularity and reusability of code. Applications implementing the parallel and object oriented methods successfully calculate the degree variance for a degree and order 110 geopotential solution on 32 processors of the Cray T3E. The memory resident gradiometer application exhibits an overall application performance of 5.4 Gflops, and the out-of-core linear solver exhibits an overall performance of 2.4 Gflops. The combination solution derived from a sun synchronous gradiometer orbit produce average geoid height variances of 17 millimeters.

Schutz, Bob E.↗

A Flexible Power Method for Solving Infinite Dimensional Tensor Eigenvalue Problems

We propose a flexible power method for computing the leftmost, i.e., algebraically smallest, eigenvalue of an infinite dimensional tensor eigenvalue problem, $H x = \lambda x$, where the infinite dimensional symmetric matrix $H$ exhibits a translational invariant structure. We assume the smallest eigenvalue of $H$ is simple and apply a power iteration of $e^{-H}$ with the eigenvector represented in a compact way as a translational invariant infinite Tensor Ring (iTR). Hence, the infinite dimensional eigenvector can be represented by a finite number of iTR cores of finite rank. In order to implement this power iteration, we use a small parameter $t$ so that the infinite matrix-vector operation $e^{-Ht}x$ can efficiently be approximated by the Lie product formula, also known as Suzuki--Trotter splitting, and we employ a low rank approximation through a truncated singular value decomposition on the iTR cores in order to keep the cost of subsequent power iterations bounded. We also use an efficient way for computing the iTR Rayleigh quotient and introduce a finite size iTR residual which is used to monitor the convergence of the Rayleigh quotient and to modify the timestep $t$. In this paper, we discuss 2 different implementations of the flexible power algorithm and illustrate the automatic timestep adaption approach for several numerical examples.

Beeumen, Roel Van↗

Accelerating the density-functional tight-binding method using graphical processing units

Acceleration of the density-functional tight-binding (DFTB) method on single and multiple graphical processing units (GPUs) was accomplished using the MAGMA linear algebra library. Herein two major computational bottlenecks of DFTB ground-state calculations were addressed in our implementation: the Hamiltonian matrix diagonalization and the density matrix construction. The code was implemented and benchmarked on two different computer systems: (1) the SUMMIT IBM Power9 supercomputer at the Oak Ridge National Laboratory Leadership Computing Facility with 1–6 NVIDIA Volta V100 GPUs per computer node and (2) an in-house Intel Xeon computer with 1–2 NVIDIA Tesla P100 GPUs. The performance and parallel scalability were measured for three molecular models of 1-, 2-, and 3-dimensional chemical systems, represented by carbon nanotubes, covalent organic frameworks, and water clusters.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Window convolution of the galaxy clustering bispectrum

In galaxy survey analysis, the observed clustering statistics do not directly match theoretical predictions but rather have been processed by a window function that arises from the survey geometry including the sky footprint, redshift-dependent background number density and systematic weights. While window convolution of the power spectrum is well studied, for the bispectrum with a larger number of degrees of freedom, it poses a significant numerical and computational challenge. In this work, we consider the effect of the survey window in the tripolar spherical harmonic decomposition of the bispectrum and lay down a formal procedure for their convolution via a series expansion of configuration-space three-point correlation functions, which was first proposed by Sugiyama et al. (2019). We then provide a linear algebra formulation of the full window convolution, where an unwindowed bispectrum model vector can be directly premultiplied by a window matrix specific to each survey geometry. To validate the pipeline, we focus on the Dark Energy Spectroscopic Instrument (DESI) Data Release 1 (DR1) luminous red galaxy (LRG) sample in the South Galactic Cap (SGC) in the redshift bin 0.4 ≤ z ≤ 0.6. We first perform convergence checks on the measurement of the window function from discrete random catalogues, and then investigate the convergence of the window convolution series expansion truncated at a finite of number of terms as well as the performance of the window matrix. This work highlights the differences in window convolution between the power spectrum and bispectrum, and provides a streamlined pipeline for the latter for current surveys such as DESI and the Euclid mission.

79 ASTRONOMY AND ASTROPHYSICS↗

SparseLU, A Novel Algorithm and Math Library for Sparse LU Factorization

Decomposing sparse matrices into lower and upper triangular matrices (sparse LU factorization) is a key operation in many computational scientific applications. We developed SparseLU, a sparse linear algebra library that implements a new algorithm for LU factorization on general sparse matrices. The new algorithm divides the input matrix into tiles to which OpenMP tasks are created for factorization computation, where only tiles that contain nonzero elements are computed. For comparative performance analysis, we used the reference library SuperLU. Testing was performed on synthetically generated matrices which replicate the conditions of the real-world matrices. SparseLU is able to reach a mean speedup of ~29× compared to SuperLU.

Valero Lara, Pedro↗

Discretized partial differential equations - Examples of control systems defined on modules

The purpose of this paper is to show how the important problems of linear system theory can be solved concisely for a particular class of linear systems, namely block circulant systems, by exploiting the algebraic structure. This type of system arises in lumped approximations to linear partial differential equations. The computation of the transition matrix, the variation of constants formula, observability, controllability, pole allocation, realization theory, stability and quadratic optimal control are discussed. In principle, all questions which are solved here could also be solved by standard methods; the present paper clearly exposes the structure of the solution, and thus permits various savings in computational effort.

Brockett, R. W.↗

Evaluation of data driven low-rank matrix factorization for accelerated solutions of the Vlasov equation

Low-rank methods have shown success in accelerating simulations of a collisionless plasma described by the Vlasov equation, but still rely on computationally costly linear algebra every time step. We propose a data-driven factorization method using artificial neural networks, specifically with convolutional layer architecture, that trains on existing simulation data. At inference time, the model outputs a low-rank decomposition of the distribution field of the charged particles, and we demonstrate that this step is faster than the standard linear algebra technique. Numerical experiments show that the method achieves comparable reconstruction accuracy for interpolation tasks, generalizing to unseen test data in a manner beyond just memorizing training data; patterns in factorization also inherently followed the same numerical trend as those within algebraic methods (e.g., truncated singular-value decomposition). However, when training on the first 70% of a time-series data and testing on the remaining 30%, the method fails to meaningfully extrapolate. Despite this limiting result, the technique may have benefits for simulations in a statistical steady-state or otherwise showing temporal stability. These results suggest that while the model offers a computationally efficient alternative for datasets with temporal stability, its current formulation is best suited for interpolation rather than for predicting future states in time-evolving systems. This study thus lays the groundwork for further refinement of neural network-based approaches to low-rank matrix factorization in high-dimensional plasma simulations.

97 MATHEMATICS AND COMPUTING↗

Geometric invariants of quantum metrology

Here, we establish a conservation law for the Quantum Fisher Information Matrix (QFIM) expressed as follows; when the QFIM is constructed from a set of observables closed under commutation, i.e., a Lie algebra, the spectrum of the QFIM is invariant under unitary dynamics generated by these same operators. Each Lie algebra therefore endows any quantum state with a fixed “budget” of metrological sensitivity—an intrinsic resource that we show, like optical squeezing in interferometry, cannot be amplified by symmetry-preserving operations. The Uhlmann curvature tensor naturally inherits the same symmetry group, and so quantum incompatibility is similarly fixed. As a result, a metrological analog to Liouville's theorem appears; statistical distances, volumes, and curvatures are invariant under the evolution generated by the Lie algebra. We discuss this as it relates to the quantum analogs of classical optimality criteria. This enables one to efficiently classify useful classes of quantum states at the level of Lie algebras through geometric invariants.

Wilson, Christopher [University of Colorado, Bould↗

Electromagnetic Transient (EMT) Simulation Algorithms for Evaluation of Large-Scale Extreme Fast Charging Systems (T&D Models)

Simulation of high-fidelity models of extreme fast charging (XFC) systems and large-area power grids with many XFCs can be time consuming in traditional simulators. Traditional simulators use a single method of discretization for all the components that results in imposing a large computational burden of inverting a large matrix as well as increased computations related to single method of discretization (that is typically a trapezoidal method). To overcome the problem of simulating large-area power grids with many XFCs, in this paper, advanced numerical simulation algorithms are applied for the first time together to reduce the dimension of matrix inversion. Here, the algorithms include numerical stiffness-based segregation, time constant-based segregation, clustering and aggregation on differential algebraic equations (DAEs), and multi-order integration approaches. These algorithms apply multiple discretization algorithms rather than a single discretization algorithm that further reduces the computational burden. The approaches mentioned here have resulted in speed-up of up to 18x in the simulation of a single distribution system with 15 XFCs and of up to 271x in the simulation of a transmission-distribution system with 300 XFCs in multiple distribution feeders with respect to conventional simulators (like power systems computer aided design [PSCAD]).

42 ENGINEERING↗

Neutron (and other Particle) Transport at LANL: An Overview [Presentation]

For decades, Los Alamos National Laboratory has been at the forefront of neutron transport methods research and code development. One such code is PARTISN, the LANL parallel time-dependent discrete ordinate neutron transport code. In this presentation, we describe the various research efforts currently underway by the PARTISN and other code teams. Some examples of current research are a block automated mesh refinement scheme, the application of tensor trains to the discretized neutron transport equation, and GPU code porting. The block automated mesh refinement scheme uses cross section information to refine and coarsen the solution mesh to improve time to solution and reduce memory. The tensor train approach expresses discretized transport operators as tensor products of vectors and matrices to compress the size of linear systems being solved by transport codes. Rather than relying on matrix-free methods such as the transport sweep, we have access to an operator that can be inverted, reshaped, or manipulated algebraically. Finally, we describe how PARTISN is used, what problems we are looking to solve, and what the future holds for neutron transport at LANL. In addition to this, we briefly describe the various research efforts in other particle transport teams using both deterministic and Monte Carlo methods. In the presentation, we list possible opportunities for collaboration between the laboratory and faculty and students.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Large-scale computation of incompressible viscous flow by least-squares finite element method

The least-squares finite element method (LSFEM) based on the velocity-pressure-vorticity formulation is applied to large-scale/three-dimensional steady incompressible Navier-Stokes problems. This method can accommodate equal-order interpolations and results in symmetric, positive definite algebraic system which can be solved effectively by simple iterative methods. The first-order velocity-Bernoulli function-vorticity formulation for incompressible viscous flows is also tested. For three-dimensional cases, an additional compatibility equation, i.e., the divergence of the vorticity vector should be zero, is included to make the first-order system elliptic. The simple substitution of the Newton's method is employed to linearize the partial differential equations, the LSFEM is used to obtain discretized equations, and the system of algebraic equations is solved using the Jacobi preconditioned conjugate gradient method which avoids formation of either element or global matrices (matrix-free) to achieve high efficiency. To show the validity of this scheme for large-scale computation, we give numerical results for 2D driven cavity problem at Re = 10000 with 408 x 400 bilinear elements. The flow in a 3D cavity is calculated at Re = 100, 400, and 1,000 with 50 x 50 x 50 trilinear elements. The Taylor-Goertler-like vortices are observed for Re = 1,000.

Jiang, Bo-Nan↗

A least-squares finite element method for 3D incompressible Navier-Stokes equations

The least-squares finite element method (LSFEM) based on the velocity-pressure-vorticity formulation is applied to three-dimensional steady incompressible Navier-Stokes problems. This method can accommodate equal-order interpolations, and results in symmetric, positive definite algebraic system. An additional compatibility equation, i.e., the divergence of vorticity vector should be zero, is included to make the first-order system elliptic. The Newton's method is employed to linearize the partial differential equations, the LSFEM is used to obtain discretized equations, and the system of algebraic equations is solved using the Jacobi preconditioned conjugate gradient method which avoids formation of either element or global matrices (matrix-free) to achieve high efficiency. The flow in a half of 3D cubic cavity is calculated at Re = 100, 400, and 1,000 with 50 x 52 x 25 trilinear elements. The Taylor-Gortler-like vortices are observed at Re = 1,000.

Jiang, Bo-Nan↗

Fast truncated SVD of sparse and dense matrices on graphics processors

We investigate the solution of low-rank matrix approximation problems using the truncated singular value decomposition (SVD). For this purpose, we develop and optimize graphics processing unit (GPU) implementations for the randomized SVD and a blocked variant of the Lanczos approach. Our work takes advantage of the fact that the two methods are composed of very similar linear algebra building blocks, which can be assembled using numerical kernels from existing high-performance linear algebra libraries. Furthermore, the experiments with several sparse matrices arising in representative real-world applications and synthetic dense test matrices reveal a performance advantage of the block Lanczos algorithm when targeting the same approximation accuracy.

Computer Science↗

Relaxation schemes for spectral multigrid methods

The effectiveness of relaxation schemes for solving the systems of algebraic equations which arise from spectral discretizations of elliptic equations is examined. Iterative methods are an attractive alternative to direct methods because Fourier transform techniques enable the discrete matrix-vector products to be computed almost as efficiently as for corresponding but sparse finite difference discretizations. Preconditioning is found to be essential for acceptable rates of convergence. Preconditioners based on second-order finite difference methods are used. A comparison is made of the performance of different relaxation methods on model problems with a variety of conditions specified around the boundary. The investigations show that iterations based on incomplete LU decompositions provide the most efficient methods for solving these algebraic systems.

Phillips, Timothy N.↗

Light-front puzzles

Abstract Light-front formulations of quantum field theories have many advantages for computing electroweak matrix elements of strongly interacting systems and other quantities that are used to study hadronic structure. The theory can be formulated in Hamiltonian form so non-perturbative calculations of the strongly interacting initial and final states are in principle reduced to linear algebra. These states are needed for calculating parton distribution functions and other types of distribution amplitudes that are used to understand the structure of hadrons. Light-front boosts are kinematic transformations so the strongly interacting states can be computed in any frame. This is useful for computing current matrix elements involving electroweak probes where the initial and final hadronic states are in different frames related by the momentum transferred by the probe. Finally in many calculations the vacuum is trivial so the calculations can be formulated in Fock space. The advantages of light front-field theory would not be interesting if the light-front formulation was not equivalent to the covariant or canonical formulations of quantum field theory. Many of the distinguishing properties of light-front quantum field theory are difficult to reconcile with canonical or covariant formulations of quantum field theory. This paper discusses the resolution of some of the apparent inconsistencies in canonical, covariant and light-front formulations of quantum field theory. The puzzles that will be discussed are (1) the problem of inequivalent representations (2) the problem of the trivial vacuum (3) the problem of ill-posed initial value problems (4) the problem of rotational covariance (5) the problem of zero modes and (6) the problem of spontaneously broken symmetries.

Physics↗