Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Matrix factorization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Connecting Matrix Elements to Multi-Hadron Form-Factors

We discuss developments in calculating multi-hadron form-factors and transition processes via lattice QCD. Our primary tools are finite-volume scaling relations, which map spectra and matrix elements to the corresponding multi-hadron infinite-volume amplitudes. We focus on two hadron processes probed by an external current, and provide various checks on the finite-volume formalism in the limiting cases of perturbative interactions and systems forming a bound state. By studying model-independent properties of the infinite-volume amplitudes, we are able to rigorously define form-factors of resonances.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Osteoblast fibronectin mRNA, protein synthesis, and matrix are unchanged after exposure to microgravity

The well-defined osteoblast line, MC3T3-E1 was used to examine fibronectin (FN) mRNA levels, protein synthesis, and extracellular FN matrix accumulation after growth activation in spaceflight. These osteoblasts produce FN extracellular matrix (ECM) known to regulate adhesion, differentiation, and function in adherent cells. Changes in bone ECM and osteoblast cell shape occur in spaceflight. To determine whether altered FN matrix is a factor in causing these changes in spaceflight, quiescent osteoblasts were launched into microgravity and were then sera activated with and without a 1-gravity field. Synthesis of FN mRNA, protein, and matrix were measured after activation in microgravity. FN mRNA synthesis is significantly reduced in microgravity (0-G) when compared to ground (GR) osteoblasts flown in a centrifuge simulating earth's gravity (1-G) field 2.5 h after activation. However, 27.5 h after activation there were no significant differences in mRNA synthesis. A small but significant reduction of FN protein was found in the 0-G samples 2.5 h after activation. Total FN protein 27.5 h after activation showed no significant difference between any of the gravity conditions, however, there was a fourfold increase in absolute amount of protein synthesized during the incubation. Using immunofluorescence, we found no significant differences in the amount or in the orientation of the FN matrix after 27.5 h in microgravity. These results demonstrate that FN is made by sera-activated osteoblasts even during exposure to microgravity. These data also suggest that after a total period of 43 h of spaceflight FN transcription, translation, or altered matrix assembly is not responsible for the altered cell shape or altered matrix formation of osteoblasts.

Flight Experiment↗

GPU Accelerated Sparse Cholesky Factorization

The solution of sparse symmetric positive definite linear systems is an important computational kernel in large-scale scientific and engineering modeling and simulation. We will solve the linear systems using a direct method, in which a Cholesky factorization of the coefficient matrix is performed using a right-looking approach and the resulting triangular factors are used to compute the solution. Sparse Cholesky factorization is compute intensive. In this work we investigate techniques for reducing the factorization time in sparse Cholesky factorization by offloading some of the dense matrix operations on a GPU. We will describe the techniques we have considered. We achieved up to 4x speedup compared to the CPU-only version.

Karsavuran, M Ozan↗

Quantum block encoding for one-pair semiseparable matrices

Quantum block encoding (QBE) is a crucial step in the development of most quantum algorithms, as it provides an embedding of a given matrix into a suitable larger unitary matrix. Historically, the development of efficient techniques for QBE has mostly focused on sparse matrices; less effort has been devoted to data-sparse (e.g., rank-structured) matrices. In this work we examine a particular case of rank structure, namely, one-pair semiseparable matrices. We present a new block encoding approach that relies on a suitable factorization of the given matrix as the product of triangular and diagonal factors. To encode the matrix, the algorithm needs $2\log(N)+7$ ancillary qubits. Assuming that the data input oracles can be implemented with polylogarithmic depth, or that a QRAM input model is available, our proposed method requires $\mathcal{O}({\rm polylog} (N))$ time and has an error of $\mathcal{O}(N^2)$, where $N$ is the matrix size.

Antonioli, Giacomo [Pisa U.; CERN] (ORCID:00090000↗

Form factors and spectral densities from Lightcone Conformal Truncation

We use the method of Lightcone Conformal Truncation (LCT) to obtain form factors and spectral densities of local operators $\mathcal{O}$ in $\phi^4$ theory in two dimensions. We show how to use the Hamiltonian eigenstates from LCT to obtain form factors that are matrix elements of a local operator $\mathcal{O}$ between single-particle bra and ket states, and we develop methods that significantly reduce errors resulting from the finite truncation of the Hilbert space. We extrapolate these form factors as a function of momentum to the regime where, by crossing symmetry, they are form factors of $\mathcal{O}$ between the vacuum and a two-particle asymptotic scattering state. We also compute the momentum-space time-ordered two-point functions of local operators in LCT. These converge quickly at momenta away from branch cuts, allowing us to indirectly obtain the time-ordered correlator and the spectral density at the branch cuts. We focus on the case where the local operator $\mathcal{O}$ is the trace Θ of the stress tensor.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Line strengths of N2O in the 1120-1440/cm region

Line strengths of N2O and its isotopic derivatives in the 1120-1440/cm region were measured at low pressure and high resolution (0.0054/cm). The band strength, rotationless dipole moment matrix elements, and F factor coefficients were considered. First-order nondegenerate perturbation theory was employed to derive explicit expressions for the rotationless dipole moment matrix elements and F factor coefficients. This made it possible to obtain general expressions for the F factor. The derived expressions were also applicable to CO2 bands.

Toth, R. A.↗

Implementation of thermal residual stresses in the analysis of fiber bridged matrix crack growth in titanium matrix composites

In this research, thermal residual stresses were incorporated in an analysis of fiber-bridged matrix cracks in unidirectional and cross-ply titanium matrix composites (TMC) containing center holes or center notches. Two TMC were investigated, namely, SCS-6/Timelal-21S laminates. Experimentally, matrix crack initiation and growth were monitored during tension-tension fatigue tests conducted at room temperature and at an elevated temperature of 200 C. Analytically, thermal residual stresses were included in a fiber bridging (FB) model. The local R-ratio and stress-intensity factor in the matrix due to thermal and mechanical loadings were calculated and used to evaluate the matrix crack growth behavior in the two materials studied. The frictional shear stress term, tau, assumed in this model was used as a curve-fitting parameter to matrix crack growth data. The scatter band in the values of tau used to fit the matrix crack growth data was significantly reduced when thermal residual stresses were included in the fiber bridging analysis. For a given material system, lay-up and temperature, a single value of tau was sufficient to analyze the crack growth data. It was revealed in this study that thermal residual stresses are an important factor overlooked in the original FB models.

Bakuckas, John G., Jr.↗

Efficient Implementation of an Optimal Interpolator for Large Spatial Data Sets

Scattered data interpolation is a problem of interest in numerous areas such as electronic imaging, smooth surface modeling, and computational geometry. Our motivation arises from applications in geology and mining, which often involve large scattered data sets and a demand for high accuracy. The method of choice is ordinary kriging. This is because it is a best unbiased estimator. Unfortunately, this interpolant is computationally very expensive to compute exactly. For n scattered data points, computing the value of a single interpolant involves solving a dense linear system of size roughly n x n. This is infeasible for large n. In practice, kriging is solved approximately by local approaches that are based on considering only a relatively small'number of points that lie close to the query point. There are many problems with this local approach, however. The first is that determining the proper neighborhood size is tricky, and is usually solved by ad hoc methods such as selecting a fixed number of nearest neighbors or all the points lying within a fixed radius. Such fixed neighborhood sizes may not work well for all query points, depending on local density of the point distribution. Local methods also suffer from the problem that the resulting interpolant is not continuous. Meyer showed that while kriging produces smooth continues surfaces, it has zero order continuity along its borders. Thus, at interface boundaries where the neighborhood changes, the interpolant behaves discontinuously. Therefore, it is important to consider and solve the global system for each interpolant. However, solving such large dense systems for each query point is impractical. Recently a more principled approach to approximating kriging has been proposed based on a technique called covariance tapering. The problems arise from the fact that the covariance functions that are used in kriging have global support. Our implementations combine, utilize, and enhance a number of different approaches that have been introduced in literature for solving large linear systems for interpolation of scattered data points. For very large systems, exact methods such as Gaussian elimination are impractical since they require 0(n(exp 3)) time and 0(n(exp 2)) storage. As Billings et al. suggested, we use an iterative approach. In particular, we use the SYMMLQ method, for solving the large but sparse ordinary kriging systems that result from tapering. The main technical issue that need to be overcome in our algorithmic solution is that the points' covariance matrix for kriging should be symmetric positive definite. The goal of tapering is to obtain a sparse approximate representation of the covariance matrix while maintaining its positive definiteness. Furrer et al. used tapering to obtain a sparse linear system of the form Ax = b, where A is the tapered symmetric positive definite covariance matrix. Thus, Cholesky factorization could be used to solve their linear systems. They implemented an efficient sparse Cholesky decomposition method. They also showed if these tapers are used for a limited class of covariance models, the solution of the system converges to the solution of the original system. Matrix A in the ordinary kriging system, while symmetric, is not positive definite. Thus, their approach is not applicable to the ordinary kriging system. Therefore, we use tapering only to obtain a sparse linear system. Then, we use SYMMLQ to solve the ordinary kriging system. We show that solving large kriging systems becomes practical via tapering and iterative methods, and results in lower estimation errors compared to traditional local approaches, and significant memory savings compared to the original global system. We also developed a more efficient variant of the sparse SYMMLQ method for large ordinary kriging systems. This approach adaptively finds the correct local neighborhood for each query point in the interpolation process.

Memarsadeghi, Nargess↗

The use of the QR factorization in the partial realization problem

The use of the QR factorization of the Hankel matrix in solving the partial realization problem is analyzed. Straightforward use of the QR factorization results in a realization scheme that possesses all of the computational advantages of Rissanen's realization scheme. These latter properties are computational efficiency, recursiveness, use of limited computer memory, and the realization of a system triplet having a condensed structure. Moreover, this scheme is robust when the order of the system corresponds to the rank of the Hankel matrix. When this latter condition is violated, an approximate realization could be determined via the QR factorization. In this second scheme, the given Hankel matrix is approximated by a low-rank non-Hankel matrix. Furthermore, it is demonstrated that column pivoting might be incorporated in this second scheme. The results presented are derived for a single input/single output system, but this does not seem to be a restriction.

Verhaegen, M. H.↗

Efficient Precoding for Single Carrier Modulation in Multi-User Massive MIMO Networks

Abstract—Frequency domain (FD) multi-user detection (MUD) and precoding has been shown to be an effective means of approaching the theoretical per-user capacity for single carrier modulation (SCM) schemes in massive MIMO scenarios with highly dispersive channels. When a cyclic prefix is added to the SCM waveform, the circulant structures of the resulting convolutional channel matrix allows for relatively simple expressions for the FD detection and precoding. In this paper, we further refine the computationally efficient minimum mean squared error (MMSE) FD-MUD technique in a time-division duplexing (TDD) massive MIMO setup by pre-computing the scale factor needed to produce an unbiased estimate. The resulting scale factor and the matrix inverse that was calculated for the uplink is then reused for the downlink to form a computationally efficient precoding scheme with superior performance compared to zero-forcing (ZF) precoding. Performance analysis is provided and confirmed through simulation results.

5G and Beyond Communications↗

Using a multifrontal sparse solver in a high performance, finite element code

We consider the performance of the finite element method on a vector supercomputer. The computationally intensive parts of the finite element method are typically the individual element forms and the solution of the global stiffness matrix both of which are vectorized in high performance codes. To further increase throughput, new algorithms are needed. We compare a multifrontal sparse solver to a traditional skyline solver in a finite element code on a vector supercomputer. The multifrontal solver uses the Multiple-Minimum Degree reordering heuristic to reduce the number of operations required to factor a sparse matrix and full matrix computational kernels (e.g., BLAS3) to enhance vector performance. The net result in an order-of-magnitude reduction in run time for a finite element application on one processor of a Cray X-MP.

King, Scott D.↗

Factorization of Binary Matrices: Rank Relations, Uniqueness and Model Selection of Boolean Decomposition

The application of binary matrices are numerous. Representing a matrix as a mixture of a small collection of latent vectors via low-rank decomposition is often seen as an advantageous method to interpret and analyze data. In this work, we examine the factorizations of binary matrices using standard arithmetic (real and nonnegative) and logical operations (Boolean and $\mathbb{Z}$ 2 ). We examine the relationships between the different ranks, and discuss when factorization is unique. In particular, we characterize when a Boolean factorization X = W$\land$H has a unique W, a unique H (for a fixed W), and when both W and H are unique, given a rank constraint. We introduce a method for robust Boolean model selection, called BMFk, and show on numerical examples that BMFk not only accurately determines the correct number of Boolean latent features but reconstruct the pre-determined factors accurately.

97 MATHEMATICS AND COMPUTING↗

Accelerating high-order mesh optimization using finite element partial assembly on GPUs

In this paper we present a new GPU-oriented mesh optimization method based on high order finite elements. Our approach relies on node movement with fixed topology, through the Target-Matrix Optimization Paradigm (TMOP) and uses a global nonlinear solve over the whole computational mesh, i.e., all mesh nodes are moved together. A key property of the method is that the mesh optimization process is recast in terms of finite element operations, which allows us to utilize recent advances in the field of GPU-accelerated high order finite element algorithms. For example, we reduce data motion by using tensor factorization and matrix-free methods, which have superior performance characteristics compared to traditional full finite element matrix assembly and offer advantages for GPU based HPC hardware. Furthermore, we describe the major mathematical components of the method along with their efficient GPU-oriented implementation. In addition, we propose an easily reproducible mesh optimization test that can serve as a performance benchmark for the mesh optimization community.

97 MATHEMATICS AND COMPUTING↗

Impact of Reordering on the LU Factorization Performance of Bordered Block-Diagonal Sparse Matrix

Power engineers rely on computer-based simulation tools to assess grid performance and ensure security. At the core of these tools are solvers for sparse linear equations. When transformed into a bordered block-diagonal (BBD) structure, part of the sparse linear equation solving can be parallelized. This work focuses on using the Schur-complement-based method for LU factorization on BBD matrices, specifically, Jacobian matrices from large-scale systems. Our findings show that the natural ordering method outperforms the default ordering method in computational performance for each block of the BBD matrix. This observation is validated using synthetic 25k-bus and 70k-bus cases, showing a speedup of up to 38% when using natural ordering without permutation. Additionally, the impact of the number of partitions is studied, and the result shows that computational performance improves with more, smaller partitions in the BBD matrices.

BBD matrix↗

A Study of Some Factors Affecting Rolling-Contact Fatigue Life

A series of investigations using the fatigue spin rig to study the effect of several factors contributing to rolling contact fatigue life is summarized. Ball specimens of 1/2 and 9/16 inch diameter were tested at maximum theoretical Hertz compressive stresses in the range of 600,000 to 750,000 psi. Life was found to vary inversely with the tenth power of stress. In forging fiber studies, a greater concentration of failures and poorer life were observed where the greatest angle of intersection between the fiber flow lines and the surface occurred. This effect was independent of alloy composition. Higher lubricant viscosity was found to increase fatigue life, Lubricants having the same viscosity but of different base stock produced wide differences in life that correlated with the pressure viscosity coefficient of the lubricant. Higher temperature produced lower fatigue life. Dry powder lubricants produced poor fatigue life at 450 F; failure appearance indicated that the lubricant particles probably acted as minute stress raisers. In metallographic studies, nonmetallic inclusions were found to have a deleterious effect on fatigue life the inclusion size, location, composition, and condition of the matrix being contributing factors; failures were by shear cracking in the subsurface zone of maximum shear stress and eventual propagation into a shallow surface spall. Vacuum melting improved fatigue life, although a general correlation between cleanliness and fatigue life was not found. Life results for ten different bearing materials are presented.

Carter, Thomas L.↗

Linear complexity

We present factorization and solution phases for a new linear complexity direct solver designed for concurrent batch operations on fine-grained parallel architectures, for matrices amenable to hierarchical representation. We focus on the strong-admissibility-based $\mathscr{H}^{2}$ format, where strong recursive skeletonization factorization compresses remote interactions. We build upon previous implementations of $\mathscr{H}^{2}$ matrix construction for efficient factorization and solution algorithm design, which are illustrated graphically in stepwise detail. The algorithms are ‘blackbox’ in the sense that the only inputs are the matrix and right-hand side, without analytical or geometrical information about the origin of the system. We demonstrate linear complexity scaling in both time and memory on four representative families of dense matrices up to one million in size. Parallel scaling up to 16 threads is enabled by a multi-level matrix graph coloring and avoidance of dynamic memory allocations thanks to prefix-sum memory management. An experimental backward error analysis is included. We break down the timings of different phases, identify phases that are memory-bandwidth limited, and discuss alternatives for phases that may be sensitive to the trend to employ lower precisions for performance.

Boukaram, Wajih↗

LuGo: An enhanced quantum phase estimation implementation

Quantum Phase Estimation (QPE) is a cardinal algorithm in quantum computing that plays a crucial role in various applications, including cryptography, molecular simulation, and solving systems of linear equations. However, the standard implementation of QPE faces challenges related to time complexity and circuit depth, which limit its practicality for large-scale computations. We introduce LuGo, a novel framework designed to enhance the performance of QPE by reducing circuit duplication, as well as using parallelization techniques to achieve faster generation of the QPE circuit and gate reduction. We validate the effectiveness of our framework by generating quantum linear solver circuits, which require both QPE and inverse QPE, to solve linear systems of equations. LuGo achieves significant improvements in both computational efficiency and hardware requirements without compromising on accuracy. Compared to a standard QPE implementation, LuGo reduces time consumption to generate a circuit that solves a 2 6 × 2 6 system matrix by a factor of 50.68 and over 31× reduction of quantum gates and circuit depth, with no fidelity loss on an ideal quantum simulator. Furthermore, we demonstrated the versatility and scalability of LuGo enabled HHL algorithm by simulating a canonical Hele-Shaw fluid problem using a quantum simulator. With these advantages, LuGo paves the way for more efficient implementations of QPE, enabling broader applications across several quantum computing domains.

Quantum algorithm↗

Operator dynamics in Floquet many-body systems

We study operator dynamics in many-body quantum systems, focusing on generic features of systems that are ergodic, spatially extended, and lack conserved densities. Quantum circuits of various types provide simple models for such systems. We focus on Floquet quantum circuits, comparing their behavior with what has been found previously for circuits that are random in time. Floquet circuits, which have discrete time-translation symmetry, represent an intermediate case between circuits that are random in time and lack any symmetry, and systems with a time-independent Hamiltonian and continuous time-translation invariance. By making this comparison, one of our aims is to identify signatures of time-translation symmetry in Floquet operator dynamics. To characterize behavior we examine a variety of quantities in solvable models and numerically: operator autocorrelation functions; the partial spectral form factor; the out-of-time-order correlator (OTOC); and the paths in operator space that make the dominant contributions to the ensemble-averaged autocorrelation functions. Our most striking result is that ensemble-averaged autocorrelation functions show behavior that is distinctively different in Floquet systems compared to systems in which successive time-steps are independent. Specifically, while average autocorrelation functions decay on a microscopic timescale for circuits that are random in time, in Floquet systems they have a late-time tail with a duration that grows parametrically with the size of the operator support. In the simplest models this tail is separated from the initial decay by a minimum, so that the average autocorrelation function has an intermediate-time peak. The existence of these tails provides a way to understand deviations of the spectral form factor from random matrix behavior at times shorter than the Thouless time. In contrast to this feature in autocorrelation functions, we find no new aspects to the behavior of OTOCs for Floquet models compared to random-in-time circuits. We show that this difference between averaged autocorrelation functions and OTOCs can be understood in terms of the paths in operator space that contribute to the two quantities: paths for the former retain a limited support at late times, while paths for the latter are dominated by operator spreading. Published by the American Physical Society 2025

Yoshimura, Takato (ORCID:0000000309159846)↗