Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Matrix factorization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Form factors and spectral densities from Lightcone Conformal Truncation

We use the method of Lightcone Conformal Truncation (LCT) to obtain form factors and spectral densities of local operators $\mathcal{O}$ in $\phi^4$ theory in two dimensions. We show how to use the Hamiltonian eigenstates from LCT to obtain form factors that are matrix elements of a local operator $\mathcal{O}$ between single-particle bra and ket states, and we develop methods that significantly reduce errors resulting from the finite truncation of the Hilbert space. We extrapolate these form factors as a function of momentum to the regime where, by crossing symmetry, they are form factors of $\mathcal{O}$ between the vacuum and a two-particle asymptotic scattering state. We also compute the momentum-space time-ordered two-point functions of local operators in LCT. These converge quickly at momenta away from branch cuts, allowing us to indirectly obtain the time-ordered correlator and the spectral density at the branch cuts. We focus on the case where the local operator $\mathcal{O}$ is the trace Θ of the stress tensor.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Efficient Precoding for Single Carrier Modulation in Multi-User Massive MIMO Networks

Abstract—Frequency domain (FD) multi-user detection (MUD) and precoding has been shown to be an effective means of approaching the theoretical per-user capacity for single carrier modulation (SCM) schemes in massive MIMO scenarios with highly dispersive channels. When a cyclic prefix is added to the SCM waveform, the circulant structures of the resulting convolutional channel matrix allows for relatively simple expressions for the FD detection and precoding. In this paper, we further refine the computationally efficient minimum mean squared error (MMSE) FD-MUD technique in a time-division duplexing (TDD) massive MIMO setup by pre-computing the scale factor needed to produce an unbiased estimate. The resulting scale factor and the matrix inverse that was calculated for the uplink is then reused for the downlink to form a computationally efficient precoding scheme with superior performance compared to zero-forcing (ZF) precoding. Performance analysis is provided and confirmed through simulation results.

5G and Beyond Communications↗

Factorization of Binary Matrices: Rank Relations, Uniqueness and Model Selection of Boolean Decomposition

The application of binary matrices are numerous. Representing a matrix as a mixture of a small collection of latent vectors via low-rank decomposition is often seen as an advantageous method to interpret and analyze data. In this work, we examine the factorizations of binary matrices using standard arithmetic (real and nonnegative) and logical operations (Boolean and $\mathbb{Z}$ 2 ). We examine the relationships between the different ranks, and discuss when factorization is unique. In particular, we characterize when a Boolean factorization X = W$\land$H has a unique W, a unique H (for a fixed W), and when both W and H are unique, given a rank constraint. We introduce a method for robust Boolean model selection, called BMFk, and show on numerical examples that BMFk not only accurately determines the correct number of Boolean latent features but reconstruct the pre-determined factors accurately.

97 MATHEMATICS AND COMPUTING↗

Accelerating high-order mesh optimization using finite element partial assembly on GPUs

In this paper we present a new GPU-oriented mesh optimization method based on high order finite elements. Our approach relies on node movement with fixed topology, through the Target-Matrix Optimization Paradigm (TMOP) and uses a global nonlinear solve over the whole computational mesh, i.e., all mesh nodes are moved together. A key property of the method is that the mesh optimization process is recast in terms of finite element operations, which allows us to utilize recent advances in the field of GPU-accelerated high order finite element algorithms. For example, we reduce data motion by using tensor factorization and matrix-free methods, which have superior performance characteristics compared to traditional full finite element matrix assembly and offer advantages for GPU based HPC hardware. Furthermore, we describe the major mathematical components of the method along with their efficient GPU-oriented implementation. In addition, we propose an easily reproducible mesh optimization test that can serve as a performance benchmark for the mesh optimization community.

97 MATHEMATICS AND COMPUTING↗

Impact of Reordering on the LU Factorization Performance of Bordered Block-Diagonal Sparse Matrix

Power engineers rely on computer-based simulation tools to assess grid performance and ensure security. At the core of these tools are solvers for sparse linear equations. When transformed into a bordered block-diagonal (BBD) structure, part of the sparse linear equation solving can be parallelized. This work focuses on using the Schur-complement-based method for LU factorization on BBD matrices, specifically, Jacobian matrices from large-scale systems. Our findings show that the natural ordering method outperforms the default ordering method in computational performance for each block of the BBD matrix. This observation is validated using synthetic 25k-bus and 70k-bus cases, showing a speedup of up to 38% when using natural ordering without permutation. Additionally, the impact of the number of partitions is studied, and the result shows that computational performance improves with more, smaller partitions in the BBD matrices.

BBD matrix↗

Linear complexity

We present factorization and solution phases for a new linear complexity direct solver designed for concurrent batch operations on fine-grained parallel architectures, for matrices amenable to hierarchical representation. We focus on the strong-admissibility-based $\mathscr{H}^{2}$ format, where strong recursive skeletonization factorization compresses remote interactions. We build upon previous implementations of $\mathscr{H}^{2}$ matrix construction for efficient factorization and solution algorithm design, which are illustrated graphically in stepwise detail. The algorithms are ‘blackbox’ in the sense that the only inputs are the matrix and right-hand side, without analytical or geometrical information about the origin of the system. We demonstrate linear complexity scaling in both time and memory on four representative families of dense matrices up to one million in size. Parallel scaling up to 16 threads is enabled by a multi-level matrix graph coloring and avoidance of dynamic memory allocations thanks to prefix-sum memory management. An experimental backward error analysis is included. We break down the timings of different phases, identify phases that are memory-bandwidth limited, and discuss alternatives for phases that may be sensitive to the trend to employ lower precisions for performance.

Boukaram, Wajih↗

LuGo: An enhanced quantum phase estimation implementation

Quantum Phase Estimation (QPE) is a cardinal algorithm in quantum computing that plays a crucial role in various applications, including cryptography, molecular simulation, and solving systems of linear equations. However, the standard implementation of QPE faces challenges related to time complexity and circuit depth, which limit its practicality for large-scale computations. We introduce LuGo, a novel framework designed to enhance the performance of QPE by reducing circuit duplication, as well as using parallelization techniques to achieve faster generation of the QPE circuit and gate reduction. We validate the effectiveness of our framework by generating quantum linear solver circuits, which require both QPE and inverse QPE, to solve linear systems of equations. LuGo achieves significant improvements in both computational efficiency and hardware requirements without compromising on accuracy. Compared to a standard QPE implementation, LuGo reduces time consumption to generate a circuit that solves a 2 6 × 2 6 system matrix by a factor of 50.68 and over 31× reduction of quantum gates and circuit depth, with no fidelity loss on an ideal quantum simulator. Furthermore, we demonstrated the versatility and scalability of LuGo enabled HHL algorithm by simulating a canonical Hele-Shaw fluid problem using a quantum simulator. With these advantages, LuGo paves the way for more efficient implementations of QPE, enabling broader applications across several quantum computing domains.

Quantum algorithm↗

Operator dynamics in Floquet many-body systems

We study operator dynamics in many-body quantum systems, focusing on generic features of systems that are ergodic, spatially extended, and lack conserved densities. Quantum circuits of various types provide simple models for such systems. We focus on Floquet quantum circuits, comparing their behavior with what has been found previously for circuits that are random in time. Floquet circuits, which have discrete time-translation symmetry, represent an intermediate case between circuits that are random in time and lack any symmetry, and systems with a time-independent Hamiltonian and continuous time-translation invariance. By making this comparison, one of our aims is to identify signatures of time-translation symmetry in Floquet operator dynamics. To characterize behavior we examine a variety of quantities in solvable models and numerically: operator autocorrelation functions; the partial spectral form factor; the out-of-time-order correlator (OTOC); and the paths in operator space that make the dominant contributions to the ensemble-averaged autocorrelation functions. Our most striking result is that ensemble-averaged autocorrelation functions show behavior that is distinctively different in Floquet systems compared to systems in which successive time-steps are independent. Specifically, while average autocorrelation functions decay on a microscopic timescale for circuits that are random in time, in Floquet systems they have a late-time tail with a duration that grows parametrically with the size of the operator support. In the simplest models this tail is separated from the initial decay by a minimum, so that the average autocorrelation function has an intermediate-time peak. The existence of these tails provides a way to understand deviations of the spectral form factor from random matrix behavior at times shorter than the Thouless time. In contrast to this feature in autocorrelation functions, we find no new aspects to the behavior of OTOCs for Floquet models compared to random-in-time circuits. We show that this difference between averaged autocorrelation functions and OTOCs can be understood in terms of the paths in operator space that contribute to the two quantities: paths for the former retain a limited support at late times, while paths for the latter are dominated by operator spreading. Published by the American Physical Society 2025

Yoshimura, Takato (ORCID:0000000309159846)↗

Iterated Gauss-Seidel GMRES

The GMRES algorithm of Saad and Schultz [SIAM J. Sci. Stat. Comput., 7 (1986), pp. 856-869] is an iterative method for approximately solving linear systems Ax = b, with initial guess x0 and residual r0 = b Ax0. The algorithm employs the Arnoldi process to generate the Krylov basis vectors (the columns of Vk ). It is well known that this process can be viewed as a QR factorization of the matrix Bk = [r0, AVk] at each iteration. Despite an O (..epsilon..)..kappa.. (Bk ) loss of orthogonality, for unit roundoff ..epsilon..and condition number ..kappa.. , the modified Gram-Schmidt formulation was shown to be backward stable in the seminal paper by Paige et al. [SIAM J. Matrix Anal.Appl., 28 (2006), pp. 264-284]. We present an iterated Gauss-Seidel formulation of the GMRES algorithm (IGS-GMRES) based on the ideas of Ruhe [Linear Algebra Appl., 52 (1983), pp. 591-601] and Swirydowicz et al. [Numer. Linear Algebra Appl., 28 (2020), pp. 1-20]. IGS-GMRES maintains orthogonality to the level O (..epsilon..)..kappa.. (Bk ) or O (..epsilon..), depending on the choice of one or two iterations; for two Gauss-Seidel iterations, the computed Krylov basis vectors remain orthogonal to working accuracy and the smallest singular value of Vk remains close to one. The resulting GMRES method is thus backward stable. We show that IGS-GMRES can be implemented with only a single synchronization point per iteration, making it relevant to large-scale parallel computing environments. We also demonstrate that, unlike MGS-GMRES, in IGS-GMRES the relative Arnoldi residual corresponding to the computed approximate solution no longer stagnates above machine precision even for highly nonnormal systems.

Arnoldi-QR↗

Building a Framework to Understand Transition Metals' Behavior in Euxinic Conditions (Final Technical Report)

This project focuses first and foremost on metal sulfide geochemistry and mineralogy as controlled by a complex matrix of environmental factors. The principal investigator’s group aim to illuminate the metal-sulfide reaction mechanisms, rates, and pathways through systematic experimentation and data collection and analyzing the relationships between the characteristics of the produced metal sulfide solid-phase/aqueous complexes and the environmental factors. This understanding is essential for obtaining a full picture of the complex cycling patterns of single or multi metal species in sulfidic environments ranging from deep-see basins, hydrothermal vents, inland seas, terrestrial water bodies, to engineered remediation systems. The specific goal of this past project was to illuminate the reaction mechanisms and kinetics of metal anions and sulfide in mixed metal cation-metal anion-sulfide systems under various aqueous conditions (which resembled a range of naturally occurring euxinic settings). For the period of this contract, we investigated the molybdenum-iron-sulfide system, with an emphasis on the conditions that caused solid phase formation. We focused on quantifying the mobility/sequestration of molybdenum under each experimental condition and identified the changes of valence states for each involved element (i.e., Mo, Fe, and S) in the precipitate. We also proposed pathways for the electron transfer that occurred in aqueous chemistry. The major analytical tools used for this study include UV-visible light spectroscopy, transmission electron microscopy, X-ray photoelectron spectroscopy, and synchrotron-based X-ray absorption spectroscopy (access gained through facility proposals to the Canadian Light Source). The biggest finding of this project was that besides pH, the iron-sulfur chemistry has a dominant control of the thiolation kinetics and subsequent reduction of Mo(VI), which are likely prerequisites of molybdenum sequestration in anoxic conditions. The results have been written up as manuscript by the end of this project (see Phillips et al.). The experimental results of this project may be critical for advancing our understanding of (1) basic chemistry involving transition metals and reduced sulfur species, (2) the validity of certain geochemical proxies, and (3) the stability and evolution of euxinic geochemical environments. It is noted that the basic results obtained through this project also have implications for Mo-S cluster-based catalyst development in inorganic chemistry and materials sciences.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Latent-Variable Formulation of the Poisson Canonical Polyadic Tensor Model: Maximum Likelihood Estimation and Fisher Information

We establish parameter inference for the Poisson canonical polyadic (PCP) tensor model through a latent-variable formulation. Our approach exploits the observation that any random PCP tensor can be derived by marginalizing an unobservable random tensor of one dimension larger. The loglikelihood of this larger dimensional tensor, referred to as the “complete” loglikelihood, is comprised of multiple rank one PCP loglikelihoods. Using this methodology, we first derive maximum likelihood estimators for the PCP model and demonstrate that several existing algorithms for fitting non-negative matrix and tensor factorizations are Expectation-Maximization algorithms. Next, we derive the observed and expected Fisher information matrices for the PCP model. The Fisher information provides us crucial insights into the well-posedness of the tensor model, such as the role that tensor rank plays in identifiability and indeterminacy. For the special case of rank one PCP models, we demonstrate that these results are greatly simplified.

97 MATHEMATICS AND COMPUTING↗

Measurement of differential distributions of B → D*$\ell\bar{v}_{\ell}$ and implications on | V cb |

We present a measurement of the differential shapes of exclusive B → D*$\ell\bar{v}_{\ell}$ (B = B - , $\bar{B}$ 0 and $\ell$ = e, μ) decays with hadronic tag-side reconstruction for the full 711 fb -1 Belle dataset. We extract the Caprini-Lellouch-Neubert (CLN) and Boyd-Grinstein-Lebed (BGL) form factor parameters and use an external input for the absolute branching fractions to determine the Cabibbo-Kobayashi-Maskawa matrix element and find |V cb | CLN = (40.2 ± 0.9) x 10 -3 and |V cb | BGL = (40.7 ± 1.0) x 10 -3 with the zero-recoil lattice QCD point $\mathcal{F}$(1) = 0.906 ± 0.013. We also perform a study of the impact of beyond zero-recoil lattice QCD calculations on the |V cb | determinations. Additionally, we present the lepton-flavor universality ratio R eμ = $\mathcal{B}$(B → D*e$\bar{v}$ e )/$\mathcal{B}$(B → D*μ$\bar{v}$ μ ) = 0.993 ± 0.023 ± 0.023, the electron and muon forwardbackward asymmetry and their difference ΔA FB = 0.028 ± 0.028 ± 0.008, and the electron and muon D* longitudinal polarization fraction and their difference ΔF$^{D*}_L$ = 0.030 ± 0.025 ± 0.007. The uncertainties quoted correspond to the statistical and systematic uncertainties, respectively.

79 ASTRONOMY AND ASTROPHYSICS↗

Quantum Fourier transform revisited

Summary The fast Fourier transform (FFT) is one of the most successful numerical algorithms of the 20th century and has found numerous applications in many branches of computational science and engineering. The FFT algorithm can be derived from a particular matrix decomposition of the discrete Fourier transform (DFT) matrix. In this paper, we show that the quantum Fourier transform (QFT) can be derived by further decomposing the diagonal factors of the FFT matrix decomposition into products of matrices with Kronecker product structure. We analyze the implication of this Kronecker product structure on the discrete Fourier transform of rank‐1 tensors on a classical computer. We also explain why such a structure can take advantage of an important quantum computer feature that enables the QFT algorithm to attain an exponential speedup on a quantum computer over the FFT algorithm on a classical computer. Further, the connection between the matrix decomposition of the DFT matrix and a quantum circuit is made. We also discuss a natural extension of a radix‐2 QFT decomposition to a radix‐ d QFT decomposition. No prior knowledge of quantum computing is required to understand what is presented in this paper. Yet, we believe this paper may help readers to gain some rudimentary understanding of the nature of quantum computing from a matrix computation point of view.

Camps, Daan↗

HyKKT: a hybrid direct-iterative method for solving KKT linear systems

Here, we propose a solution strategy for the large indefinite linear systems arising in interior methods for nonlinear optimization. The method is suitable for implementation on hardware accelerators such as graphical processing units (GPUs). The current gold standard for sparse indefinite systems is the LBLT factorization where L is a lower triangular matrix and B is 1×1 or 2×2 block diagonal. However, this requires pivoting, which substantially increases communication cost and degrades performance on GPUs. Our approach solves a large indefinite system by solving multiple smaller positive definite systems, using an iterative solver on the Schur complement and an inner direct solve (via Cholesky factorization) within each iteration. Cholesky is stable without pivoting, thereby reducing communication and allowing reuse of the symbolic factorization. We demonstrate the practicality of our approach on large optimal power flow problems and show that it can efficiently utilize GPUs and outperform LBL T factorization of the full system.

97 MATHEMATICS AND COMPUTING↗

SAGA1 and MITH1 produce matrix-traversing membranes in the CO2-fixing pyrenoid

Abstract Approximately one-third of global CO 2 assimilation is performed by the pyrenoid, a liquid-like organelle found in most algae and some plants. Specialized pyrenoid-traversing membranes are hypothesized to drive CO 2 assimilation in the pyrenoid by delivering concentrated CO 2 , but how these membranes are made to traverse the pyrenoid matrix remains unknown. Here we show that proteins SAGA1 and MITH1 cause membranes to traverse the pyrenoid matrix in the model alga Chlamydomonas reinhardtii . Mutants deficient in SAGA1 or MITH1 lack matrix-traversing membranes and exhibit growth defects under CO 2 -limiting conditions. Expression of SAGA1 and MITH1 together in a heterologous system, the model plant Arabidopsis thaliana , produces matrix-traversing membranes. Both proteins localize to matrix-traversing membranes. SAGA1 binds to the major matrix component, Rubisco, and is necessary to initiate matrix-traversing membranes. MITH1 binds to SAGA1 and is necessary for extension of membranes through the matrix. Our data suggest that SAGA1 and MITH1 cause membranes to traverse the matrix by creating an adhesive interaction between the membrane and matrix. Our study identifies and characterizes key factors in the biogenesis of pyrenoid matrix-traversing membranes, demonstrates the importance of these membranes to pyrenoid function and marks a key milestone toward pyrenoid engineering into crops for improving yields.

Hennacy, Jessica H.↗

Detailed analysis of excited-state systematics in a lattice QCD calculation of 𝑔 𝐴

Excited state contamination remains one of the most challenging sources of systematic uncertainty to control in lattice QCD calculations of nucleon matrix elements and form factors: early time separations are contaminated by excited states and late times suffer from an exponentially bad signal-to-noise problem. High-statistics calculations at large time separations ≳ 1 fm are commonly used to combat these issues. In this work, focusing on g A , we explore the alternative strategy of utilizing a large number of relatively low-statistics calculations at short to medium time separations (0.2–1 fm), combined with a multistate analysis. On an ensemble with a pion mass of approximately 310 MeV and a lattice spacing of approximately 0.09 fm, we find this provides a more robust and economical method of quantifying and controlling the excited state systematic uncertainty. A quantitative separation of various types of excited states enables the identification of the transition matrix elements as the dominant contamination. The excited state contamination of the Feynman-Hellmann correlation function is found to reduce to the 1% level at approximately 1 fm while, for the more standard three-point functions, this does not occur until after 2 fm. Critical to our findings is the use of a global minimization, rather than fixing the spectrum from the two-point functions and using them as input to the three-point analysis. We find that the ground state parameters determined in such a global analysis are stable against variations in the excited state model, the number of excited states, and the truncation of early-time or late-time numerical data.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A New Model for Simulating the Imbibition of a Wetting-Phase Fluid in a Matrix-Fracture Dual Connectivity System

The imbibition experiment is an effective approach for measuring petrophysical properties of porous media, with many such experiments performed over the past decade. Quite some empirical, analytical, and numerical models have been developed to simulate spontaneous imbibition of the wetting phase fluid into porous media, but limitations still exist. In previous studies, the imbibition process has been considered to give a piston-like displacement or the porous medium modeled as multiply-sized pores linked with bonds; both approaches fail to yield comprehensive results due to their neglect of the presence of irregular fractures or nonuniform flow paths through the matrix. By building a numerical model for simulating laboratory-scale experimental data, we performed imbibition tests on several fractured Barnett Shale samples having fractures either parallel ( P ) or transverse ( T ) to the bedding plane and used MATLAB to build a new numerical model by combining the imbibition process in fractures and the matrix using concepts from percolation theory. The experimental data show that the rocks with P -direction fractures have a more steady increase of imbibition rates than the case of T -direction one. As the shale matrix with low pore connectivity hampers the upward water movement, the imbibition rate of shales with T -direction fractures will decrease suddenly after the bottom layer in contact with water is saturated during the initial period. This wetting phase movement (WPM) model can simulate 3D porous media with 2D fractures. The rate of imbibition by fractured porous media is associated with physical parameters such as porosity and fracture distribution (e.g., the number and angle of fractures). Using Monte Carlo methods, we examined fracture parameters and predicted elapsed time and cumulative water imbibition, for the Barnett Shale samples. The results show that the rate of imbibed water mass is sensitive to the number of fractures directly connected to water source, and the connectivity between two neighboring grid cells is a key parameter for the wetting-front progression. The findings of this study can help to better understand the imbibition process with multiple influencing processes and factors in fractured-matrix rocks. Although the experiments, data simulation, and prediction results are based only on Barnett Shale samples, the model is readily applicable to imbibition tests of other fractured rocks to show the spatial and temporal behavior during a dynamic imbibition process that are not easily captured experimentally.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

General Bayesian Inference over the Stiefel Manifold via the Givens Transform

We introduce here an approach based on the Givens representation for posterior inference in statistical models with orthogonal matrix parameters, such as factor models and probabilistic principal component analysis (PPCA). We show how the Givens representation can be used to develop practical methods for transforming densities over the Stiefel manifold into densities over subsets of Euclidean space. We demonstrate how to deal with issues arising from the topology of the Stiefel manifold and how to inexpensively compute the change-of-measure terms. We introduce an auxiliary parameter approach that limits the impact of topological issues. We provide both analysis of our methods and numerical examples demonstrating the effectiveness of the approach. We also discuss how our Givens representation can be used to define general classes of distributions over the space of orthogonal matrices. We then give demonstrations on several examples showing how the Givens approach performs in practice in comparison with other methods.

97 MATHEMATICS AND COMPUTING↗