Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “matrix decomposition”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Bringing randomized algorithms to mainstream numerical linear algebra

Numerical linear algebra (NLA) underpins huge swaths of computational science and engineering. For scientists and engineers to make the most of the DOE’s computing resources, it is essential that they have access to high-performance implementations of algorithms with best-in-class scalability and reliability. Despite this, prevailing NLA libraries have little to no support for breakthrough algorithms from the field of randomized numerical linear algebra (RandNLA) that have been developed over the past twenty years. The goal of this LDRD was to break a log-jam that had prevented broad adoption of RandNLA. Our work had two thrusts. The first was to develop RandBLAS: a trustworthy and high-performance C++ library for randomized dimension reduction (an operation widely known as sketching). The second was the development of a novel randomized algorithm for computing a challenging type of matrix decomposition known as Householder QR with column pivoting (Householder QRCP). In this one-year late-start LDRD we successfully delivered RandBLAS 1.0 and new CPU and GPU codes for Householder QRCP. RandBLAS has extensive documentation at https://randblas.readthedocs.io/en/stable/. Papers on RandBLAS and and our high-performance QRCP codes are forthcoming.

97 MATHEMATICS AND COMPUTING↗

A Data-scientific Noise-removal Method for Efficient Submillimeter Spectroscopy With Single-dish Telescopes

For submillimeter spectroscopy with ground-based single-dish telescopes, removing the noise contribution from the Earth’s atmosphere and the instrument is essential. For this purpose, here we propose a new method based on a data-scientific approach. The key technique is statistical matrix decomposition that automatically separates the signals of astronomical emission lines from the drift noise components in the fast-sampled (1–10 Hz) time-series spectra obtained by a position-switching (PSW) observation. Because the proposed method does not apply subtraction between two sets of noisy data (i.e., on-source and off-source spectra), it improves the observation sensitivity by a factor of √2. It also reduces artificial signals such as baseline ripples on a spectrum, which may also help to improve the effective sensitivity. We demonstrate this improvement by using the spectroscopic data of emission lines toward a high-redshift galaxy observed with a 2 mm receiver on the 50 m Large Millimeter Telescope. Since the proposed method is carried out offline and no additional measurements are required, it offers an instant improvement on the spectra reduced so far with the conventional method. It also enables efficient deep spectroscopy driven by the future 50 m class large submillimeter single-dish telescopes, where fast PSW observations by mechanical antenna or mirror drive are difficult to achieve.

47 OTHER INSTRUMENTATION↗

Randomized Algorithms for Low-Rank Matrix and Tensor Decompositions

This paper surveys randomized algorithms in numerical linear algebra for low-rank decompositions of matrices and tensors. The survey begins with a review of classical matrix algorithms that can be accelerated by randomized dimensionality reduction, such as the singular value decomposition (SVD) or interpolative (ID) and CUR decompositions. Recent advances in randomized dimensionality reduction are discussed, including new methods of fast matrix sketching and sampling techniques, which are incorporated into classical matrix algorithms for fast low-rank matrix approximations. The extension of randomized matrix algorithms to tensors is then explored for several low-rank tensor decompositions in the CP and Tucker formats, including the higher-order SVD, ID, and CUR decomposition.

Pearce, Katherine J. [The University of Texas at A↗

Assessment of Accelerated Stress Testing Data for Silicon Photovoltaics Using Tensor Decomposition Methods

In this work, we examine the use of high-order tensor decompositions to analyze degradation pathways emerging from accelerated stress testing of silicon photovoltaic (PV) modules. Matrix-based decompositions are powerful tools for studying two-dimensional data arrays and form the foundation of a host of classical data analysis techniques. Tensors are high-order extrapolations of matrices that are able to account for more parameter dimensions, and a variety of tensor decomposition methods have been developed that similarly seek to extend insights from matrix decompositions to higher dimensions. Applying and interpreting tensor decomposition methods to sequences of PV module image data, we seek to uncover and isolate different degradation modes occurring from accelerated stress testing procedures. Further, we consider the contributions of different modes to PV module performance degradations.

data analysis↗

Nonlinear Matrix Approximation with Radial Basis Function Components

We introduce and investigate matrix approximation by decomposition into a sum of radial basis function (RBF) components. An RBF component is a generalization of the outer product between a pair of vectors, where an RBF function replaces the scalar multiplication between individual vector elements. Even though the RBF functions are positive definite, the summation across components is not restricted to convex combinations and allows us to compute the decomposition for any real matrix that is not necessarily symmetric or positive definite. We formulate the problem of seeking such a decomposition as an optimization problem with a nonlinear and non-convex loss function. Several modern versions of the gradient descent method, including their scalable stochastic counterparts, are used to solve this problem. We provide extensive empirical evidence of the effectiveness of the RBF decomposition and that of the gradient-based fitting algorithm. While being conceptually motivated by singular value decomposition (SVD), our proposed nonlinear counterpart outperforms SVD by drastically reducing the memory required to approximate a data matrix with the same L2 error for a wide range of matrix types. For example, it leads to 2 to 6 times memory save for Gaussian noise, graph adjacency matrices, and kernel matrices. Moreover, this proximity-based decomposition can offer additional interpretability in applications that involve, e.g., capturing the inner low-dimensional structure of the data, retaining graph connectivity structure, and preserving the acutance of images.

Rebrova, Elizaveta↗

Sparsity of Radiating Characteristic Modes on Infinite Periodic Structures

Characteristic modes on infinite periodic structures are studied using spectral dyadic Green’s functions. This formulation demonstrates that, in contrast to the modal analysis of finite structures, the number of radiating characteristic modes is limited by unit cell size and incident wave vector (i.e., scan angle or phase shift per unit cell). Here, the reflection tensor is decomposed into modal contributions from radiating modes, indicating that characteristic modes are a predictably sparse basis in which to study reflection phenomena.

42 ENGINEERING↗

Observables for scattering on targets with arbitrary spin

Starting from the Weinberg formalism for fields of arbitrary spin, we discuss a method for the decomposition of matrix elements of QCD operators (local currents, quark/gluon bilinears) for targets with arbitrary spin. This procedure is advantageous for the systematic study of the structure of hadrons and nuclei, particularly in the case of spin-dependent observables. As higher spin targets exhibit new features in their hadronic structure, the investigation of these properties can enhance our understanding of the strong force. The construction allows for a unified framework to discuss spin > 1/2 very similar to the spin 1/2 case, without subsidiary conditions for the wave functions. Different types of spinors (canonical, helicity, light-front helicity) can be easily accommodated. Its numerical implementation is simple and can be entirely reduced to objects familiar from the rotation group. A natural sl(2,C) multipole decomposition emerges, enabling a physical interpretation of non-perturbative objects that multiply spinor bilinears as Generalized Form Factors. To demonstrate the efficacy of this method, we apply it to the description of a spin 1 target, such as the deuteron. We discuss extensions of the formalism to hard exclusive processes on the deuteron and beyond.

Vera, Frank↗

Tensor renormalization group for fermions

Abstract We review the basic ideas of the tensor renormalization group method and show how they can be applied for lattice field theory models involving relativistic fermions and Grassmann variables in arbitrary dimensions. We discuss recent progress for entanglement filtering, loop optimization, bond-weighting techniques and matrix product decompositions for Grassmann tensor networks. The new methods are tested with two-dimensional Wilson–Majorana fermions and multi-flavor Gross–Neveu models. We show that the methods can also be applied to the fermionic Hubbard model in 1+1 and 2+1 dimensions.

Physics↗

Tucker-1 Boolean Tensor Factorization with Quantum Annealers

Quantum annealers are an emerging computational architecture that have the potential to address some challenging computational issues that will be left unresolved as we approach the end of the Moore's Law era of computing. D-Wave quantum annealers are designed to solve a challenging set of problems - quadratic unconstrained binary optimization problems. This makes them a natural fit for solving problems with binary or Boolean variables. Here, we explore the use of a quantum annealer to solve Boolean tensor factorization. The goal of Boolean tensor factorization is to represent a high-dimensional tensor filled with Boolean values as a product of Boolean matrices and a Boolean core tensor. We show that a particular Boolean tensor factorization problem (called Tucker-1 factorization) can be decomposed into a sequence of quadratic unconstrained binary optimization problems that can be solved with a D-Wave 2000Q quantum annealer. While quantum annealers specifically and quantum computers in general are at a fairly early stage in their development, they are currently capable of solving these Boolean tensor factorization problems. Importantly, our results show that for fairly small tensors, we are frequently able to obtain an accurate (sometimes exact) factorization using quantum annealing.

97 MATHEMATICS AND COMPUTING↗

Online Voltage Event Detection Using Synchrophasor Data with Structured Sparsity-Inducing Norms

This paper develops an accurate and computationally efficient data-driven framework to detect voltage events from PMU data streams. It develops an innovative Proximal Bilateral Random Projection (PBRP) algorithm to quickly decompose the PMU data matrix into a low-rank matrix, a row-sparse event-pattern matrix and a noise matrix. Here, the row-sparse pattern matrix significantly distinguishes events from normal behavior. These matrices are then fed into a clustering algorithm to separate voltage events from normal operating conditions. Large-scale numerical study results on real-world PMU data show that the proposed algorithm is computationally more efficient and achieves higher F scores than state-of-the-art benchmarks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Analytic techniques for solving the transport equations in electroweak baryogenesis

We develop an efficient method for solving transport equations, particularly in the context of electroweak baryogenesis. It provides fully-analytical results under mild approximations and can also test semi-analytical results, which are applicable in more general cases. Key elements of our method include the reduction of the second-order differential equations to first order, representing the set of coupled equations as a block matrix of the particle densities and their derivatives, identification of zero modes, and block decomposition of the matrix. We apply our method to calculate the baryon asymmetry of the Universe (BAU) in a Standard Model effective field theory framework of complex Yukawa couplings to determine the sensitivity of the resulting BAU to modifications of various model parameters and rates, and to estimate the effect of the commonly-used thin-wall approximation.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Imposing equilibrium on experimental 3-D stress fields using Hodge decomposition and FFT-based optimization

Here, we present a methodology to impose micromechanical constraints, i.e. stress equilibrium at grain and sub-grain scale, to an arbitrary (non-equilibrated) voxelized stress field obtained, for example, by means of synchrotron X-ray diffraction techniques. The method consists in finding the equilibrated stress field closest (in L 2 -norm sense) to the measured non-equilibrated stress field, via the solution of an optimization problem. The extraction of the divergence-free (equilibrated) part of a general (non-equilibrated) field is performed using the Hodge decomposition of a symmetric matrix field, which is the generalization of the Helmholtz decomposition of a vector field into the sum of an irrotational field and a solenoidal field. The combination of: a) the Euler–Lagrange equations that solve the optimization problem, and b) the Hodge decomposition, gives a differential expression that contains the bi-harmonic operator and two times the curl operator acting on the experimental stress field. These high-order derivatives can be efficiently performed in Fourier space. The method is applied to filter the non-equilibrated parts of a synthetic piecewise constant stress fields with a known ground truth, and stress fields in Gum Metal, a beta-Ti-based alloy measured in-situ using Diffraction Contrast Tomography (DCT). In both cases, the largest corrections were obtained near grain boundaries.

36 MATERIALS SCIENCE↗

TuckerMPI: A Parallel C++/MPI Software Package for Large-scale Data Compression via the Tucker Tensor Decomposition

With this study, our goal is compression of massive-scale grid-structured data, such as the multi-terabyte output of a high-fidelity computational simulation. For such data sets, we have developed a new software package called TuckerMPI, a parallel C++/MPI software package for compressing distributed data. The approach is based on treating the data as a tensor, i.e., a multidimensional array, and computing its truncated Tucker decomposition, a higher-order analogue to the truncated singular value decomposition of a matrix. The result is a low-rank approximation of the original tensor-structured data. Compression efficiency is achieved by detecting latent global structure within the data, which we contrast to most compression methods that are focused on local structure. In this work, we describe TuckerMPI, our implementation of the truncated Tucker decomposition, including details of the data distribution and in-memory layouts, the parallel and serial implementations of the key kernels, and analysis of the storage, communication, and computational costs. We test the software on 4.5 and 6.7 terabyte data sets distributed across 100 s of nodes (1,000 s of MPI processes), achieving compression ratios between 100 and 200,000×, which equates to 99--99.999% compression (depending on the desired accuracy) in substantially less time than it would take to even read the same dataset from a parallel file system. Moreover, we show that our method also allows for reconstruction of partial or down-sampled data on a single node, without a parallel computer so long as the reconstructed portion is small enough to fit on a single machine, e.g., in the instance of reconstructing/visualizing a single down-sampled time step or computing summary statistics. The code is available at https://gitlab.com/tensors/TuckerMPI.

97 MATHEMATICS AND COMPUTING↗

BoBa

BoBa is a C++ software library for working with large matrices, tensors, and tensor decompositions. The library provides tools for dense matrix and tensor operations, tensor decompositions, and tensor decomposition methods that support modern CPU and GPU architectures. It includes portable abstractions for linear algebra, tensor algebra, and multidimensional computation. BoBa is intended for scientific computing applications that involve large multidimensional data sets or high dimensional mathematical models. Its capabilities support tasks such as data compression, linear algebra, efficient numerical computation, and the development of scalable algorithms for heterogeneous hardware. Tutorials, tests, and example applications are included to help users learn and apply the library.

Yao, Jin [Lawrence Livermore National Laboratory (↗

GSoFa: Scalable Sparse Symbolic LU Factorization on GPUs

Decomposing a matrix $\mathbf {A}$ into a lower matrix $\mathbf {L}$ and an upper matrix $\mathbf {U}$, which is also known as LU decomposition, is an essential operation in numerical linear algebra. For a sparse matrix, LU decomposition often introduces more nonzero entries in the $\mathbf {L}$ and $\mathbf {U}$ factors than in the original matrix. A symbolic factorization step is needed to identify the nonzero structures of $\mathbf {L}$ and $\mathbf {U}$ matrices. Attracted by the enormous potentials of the Graphics Processing Units (GPUs), an array of efforts have surged to deploy various LU factorization steps except for the symbolic factorization, to the best of our knowledge, on GPUs. This article introduces gSoFa, the first GPU-based symbolic factorization design with the following three optimizations to enable scalable LU symbolic factorization for nonsymmetric pattern sparse matrices on GPUs. First, here we introduce a novel fine-grained parallel symbolic factorization algorithm that is well suited for the Single Instruction Multiple Thread (SIMT) architecture of GPUs. Second, we tailor supernode detection into a SIMT friendly process and strive to balance the workload, minimize the communication and saturate the GPU computing resources during supernode detection. Third, we introduce a three-pronged optimization to reduce the excessive space consumption problem faced by multi-source concurrent symbolic factorization. Taken together, gSoFa achieves up to 31× speedup from 1 to 44 Summit nodes (6 to 264 GPUs) and outperforms the state-of-the-art CPU project, on average, by 5×. Notably, gSoFa also achieves up to 47 percent of the peak memory throughput of a V100 GPU in the Summit Supercomputer.

97 MATHEMATICS AND COMPUTING↗

PyAlbany: A Python interface to the C++ multiphysics solver Albany

Albany is a parallel C++ finite element library for solving forward and inverse problems involving partial differential equations (PDEs). In this paper we introduce PyAlbany, a newly developed Python interface to the Albany library. PyAlbany can be used to effectively drive Albany enabling fast and easy analysis and post-processing of applications based on PDEs that are pre-implemented in Albany. PyAlbany relies on the library PyBind11 to bind Python with C++ Albany code. Here we detail the implementation of PyAlbany and showcase its capabilities through a number of examples targeting a heat-diffusion problem. In particular we consider the following: (1) the generation of samples for a Monte Carlo application, (2) a scalability study, (3) a study of parameters on the performance of a linear solver, and finally (4) a tool for performing eigenvalue decompositions of matrix-free operators for a Bayesian inference application.

97 MATHEMATICS AND COMPUTING↗

Eigenmode analysis of the sheared-flow Z-pinch

Experiments have demonstrated that a Z-pinch can persist for thousands of times longer than the growth time of global magnetohydrodynamic (MHD) instabilities such as the m=0 sausage and m=1 kink modes. These modes have growth times on the order of ta=a/vi, where vi is the ion thermal speed and a is the pinch radius. Axial flows with duz/dr ≲ vi/a have been measured during the stable period, and the commonly accepted theory is that this amount of shear is sufficient to stabilize these modes as predicted by numerical studies using the ideal MHD equations. However, these studies only consider specific equilibrium profiles that typically have a modest magnitude for the logarithmic pressure gradient, qP≡d ln P/d ln r, and may not represent experimental conditions. Linear stability of the sheared-flow Z-pinch is studied here via a direct eigen-decomposition of the matrix operator obtained from the linear ideal MHD equations. Several equilibrium profiles with a large variation of qP are examined. Considering a practical range of k, 1/3 ≲ ka ≲ 10, it is shown that the shear required to stabilize m=0 modes can be expressed as duz/dr≥Cγ0/(ka)α. Here, γ0=γ0(ka) is the profile-specific growth rate in the absence of shear, which scales approximately with |qP|. Both C and α are profile-specific constants, but C is order unity and α≈1. It is further demonstrated that even a large value of shear, duz/dr=3vi/a, is not sufficient to provide linear stabilization of the m=1 kink mode for all profiles considered. This result is in contrast to the currently accepted theory predicting stabilization at much lower shear, duz/dr=0.1vi/a, and suggests that the experimentally observed stability cannot be explained within the linear ideal-MHD model.

Angus, J. R. (ORCID:0000000314740002)↗

Simplified Universal Equations for Ionic Conductivity and Transference Number

Nernst-Einstein equation can provide a reasonable estimate of the ionic conductivity of dilute solutions. For concentrated solutions, alternate methods such as Green–Kubo relations and Einstein relations are more suitable to account for ion-ion interactions. Such computations can be expensive for multicomponent systems. Simplified mathematical expressions like the Nernst-Einstein equation do not exist for concentrated multicomponent mixtures. Newman’s treatment of multicomponent concentrated solutions yields a conductivity relation in terms of species concentration and Onsager phenomenological coefficients. However, the estimation of these phenomenological coefficients is not straightforward. Here, mathematical formulations that relate the phenomenological coefficients with the friction coefficients are developed, leading to simplified, ready-to-use expressions of conductivity and transference numbers that can be used for a wide range of ionic mixtures. This approach involves spectral decomposition of the matrix of Onsager phenomenological coefficients. The general analytical expressions for conductivity and transference number are simplified for binary electrolytes, and numerical solutions are provided for ternary and quaternary mixtures with ion dissociation.

Electrochemistry↗