Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sparse matrix”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

High performance sparse multifrontal solvers on modern GPUs

Here, we have ported the numerical factorization and triangular solve phases of the sparse direct solver STRUMPACK to GPU. STRUMPACK implements sparse LU factorization using the multifrontal algorithm, which performs most of its operations in dense linear algebra operations on so-called frontal matrices of various sizes. Our GPU implementation off-loads these dense linear algebra operations, as well as the sparse scatter–gather operations between frontal matrices. For the larger frontal matrices, our GPU implementation relies on vendor libraries such as cuBLAS and cuSOLVER for NVIDIA GPUs and rocBLAS and rocSOLVER for AMD GPUs. For the smaller frontal matrices we developed custom CUDA and HIP kernels to reduce kernel launch overhead. Overall, high performance is achieved by identifying submatrix factorizations corresponding to sub-trees of the multifrontal assembly tree which fit entirely in GPU memory. The multi-GPU setting uses SLATE (Software for Linear Algebra Targeting Exascale) as a modern GPU-aware replacement for ScaLAPACK. On 4 nodes of SUMMIT the code runs ~10X faster when using all 24 V100 GPUs compared to when it only uses the 168 POWER9 cores. On 8 SUMMIT nodes, using 48 V100 GPUs, the sparse solver reaches over 50TFlop/s. Compared to SuperLU, on a single V100, for a set of 17 matrices our implementation is faster for all but one matrix, and is on average 5X (median 4X) faster

97 MATHEMATICS AND COMPUTING↗

Machine learning and shallow groundwater chemistry to identify geothermal prospects in the Great Basin, USA

This study discovers various geothermal prospects in the Great Basin, USA based on shallow groundwater chemical (geochemical) data. The geochemical data are expected to include hidden (latent) information that is a proxy for geothermal prospectivity. We processed the sparse geochemical data in the Great Basin at 14,341 locations including 18 attributes. Next, a non-negative matrix factorization with customized k-means clustering is applied to the geochemical data matrix that automatically finds three hidden geothermal signatures representing modestly, moderately, and highly confident geothermal prospects. The algorithm also evaluated the probability of occurrence of these types of resources through the studied region. There is a consistency between regional geothermal prospectivity as estimated by our ML methodology and the traditional play fairway analysis conducted over a portion of the study area. We also identify the dominant data attributes associated with each signature. Finally, our ML analyses allow us to reconstruct attributes from sparse into continuous over the study domain. The predicted continuous attributes can be used for future detailed geothermal explorations in the Great Basin.

15 GEOTHERMAL ENERGY↗

Why Is Attention Sparse In Particle Transformer?

Transformer-based models have achieved state-of-the-art performance in jet tagging at the CERN Large Hadron Collider (LHC), with the Particle Transformer (ParT) representing a leading example of such models. A striking feature of ParT is its sparse, nearly binary, attention structure, raising questions about the origin of this behavior and whether it encodes physically meaningful correlations. In this work, we investigate the source of ParT's sparse attention by comparing models trained on multiple benchmark datasets and examine the relative contributions of the attention term and the physics-inspired interaction matrix before softmax. We find that binary sparsity arises primarily from the attention mechanism itself, with the interaction matrix playing a secondary role. Moreove, we show that ParT is able to identify key jet substructure elements, such as leptons in semileptonic top decays, even without explicit particle identification inputs. These results provide new insight into the interpretability of transformer-based jet taggers and clarify the conditions under which sparse attention patterns emerge in ParT.

Legge, Timothy [UC, San Diego]↗

Sparsity-Independent Lyapunov Exponent in the Sachdev-Ye-Kitaev Model

The saturation of a recently proposed universal bound on the Lyapunov exponent has been conjectured to signal the existence of a gravity dual. This saturation occurs in the low-temperature limit of the dense Sachdev-Ye-Kitaev (SYK) model, N Majorana fermions with q body ( q > 2 ) infinite-range interactions. We calculate certain out-of-time-order correlators (OTOCs) for N ≤ 64 fermions for a highly sparse SYK model and find no significant dependence of the Lyapunov exponent on sparsity up to near the percolation limit where the Hamiltonian breaks up into blocks. This provides strong support to the saturation of the Lyapunov exponent in the low-temperature limit of the sparse SYK. A key ingredient to reaching N = 64 is the development of a novel quantum spin model simulation library that implements highly optimized matrix-free Krylov subspace methods on graphical processing units. This leads to a significantly lower simulation time as well as vastly reduced memory usage over previous approaches, while using modest computational resources. Strong sparsity-driven statistical fluctuations require both the use of a much larger number of disorder realizations with respect to the dense limit and a careful finite size scaling analysis. The saturation of the bound in the sparse SYK points to the existence of a gravity analog that would enlarge substantially the number of field theories with this feature. Published by the American Physical Society 2024

Physics↗

MFiX: Fractional-Step Method Implementation

A comprehensive, multiphase computational fluid dynamics (CFD) simulation solves several coupled transport equations including continuity, momentum, species, and energy. Chemical reactions further couple these equations through heats of reaction and rates of formation of products and rates of destruction of reactants. A fractional-step method separates changes attributed to chemical reactions from transport phenomena like convection and diffusion. When the governing equations are split into the transport and reacting components, efficient and independent methodologies can be exploited to solve the different systems. Specifically, discretization of field variable transport equations results in large, sparse matrices which are loosely coupled. These systems are solved in succession using iterative techniques that take advantage of the matrix structure. In contrast, chemical reactions tightly couple field variables locally within the domain (e.g., within a single computational cell) resulting in low dimensional but dense, nonlinear systems that are better solved using direct integration techniques.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Explicit Quantum Circuits for Block Encodings of Certain Sparse Matrices

Many standard linear algebra problems can be solved on a quantum computer by using recently developed quantum linear algebra algorithms that make use of block encodings and quantum eigenvalue/singular value transformations. A block encoding embeds a properly scaled matrix of interest A in a larger unitary transformation U that can be decomposed into a product of simpler unitaries and implemented efficiently on a quantum computer. Although quantum algorithms can potentially achieve exponential speedup in solving linear algebra problems compared to the best classical algorithm, such a gain in efficiency ultimately hinges on our ability to construct an efficient quantum circuit for the block encoding of A, which is difficult in general, and not trivial even for well structured sparse matrices. Here, in this paper, we give a few examples on how efficient quantum circuits can be explicitly constructed for some well structured sparse matrices and discuss a few strategies used in these constructions. We also provide implementations of these quantum circuits in MATLAB.

97 MATHEMATICS AND COMPUTING↗

A GPU-based compressible combustion solver for applications exhibiting disparate space and time scales

High-speed chemically active flows pose significant computational challenges due to their disparate space and time scales, with stiff chemistry often dominating simulation time. While modern scientific computing programs achieve exascale performance by leveraging graphics processing units (GPUs), existing GPU-based compressible combustion solvers face critical limitations in memory management, load balancing, and handling the highly localized nature of chemical reactions. To this end, we present a high-performance compressible reacting flow solver built on the AMReX framework and optimized for multi-GPU settings. Here, our approach addresses three GPU performance bottlenecks: memory access patterns through column-major storage optimization, computational workload variability via a bulk-sparse integration strategy for chemical kinetics, and multi-GPU load distribution for adaptive mesh refinement applications. The solver adapts existing matrix-based chemical kinetics formulations to multi-grid contexts. Using representative combustion applications, including 2D and 3D detonations and a 3D jet-in-crossflow configuration, we demonstrate 1.4–5× performance improvements over initial implementations on an in-house cluster of NVIDIA H100 GPUs, and near-ideal weak scaling on the Frontier supercomputer (Oak Ridge Leadership Computing Facility) with up to 1024 AMD Instinct MI250X GPUs. Roofline analysis reveals substantial improvements in arithmetic intensity for both convection (∼ 10 ×) and chemistry (∼ 4 ×) routines, confirming efficient utilization of GPU memory bandwidth and computational resources.

42 ENGINEERING↗

MAGMA: Enabling exascale performance with accelerated BLAS and LAPACK for diverse GPU architectures

MAGMA (Matrix Algebra for GPU and Multicore Architectures) is a pivotal open-source library in the landscape of GPU-enabled dense and sparse linear algebra computations. With a repertoire of approximately 750 numerical routines across four precisions, MAGMA is deeply ingrained in the DOE software stack, playing a crucial role in high-performance computing. Notable projects such as ExaConstit, HiOP, MARBL, and STRUMPACK, among others, directly harness the capabilities of MAGMA. In addition, the MAGMA development team has been acknowledged multiple times for contributing to the vendors’ numerical software stacks. Looking back over the time of the Exascale Computing Project (ECP), we highlight how MAGMA has adapted to recent changes in modern HPC systems, especially the growing gap between CPU and GPU compute capabilities, as well as the introduction of low precision arithmetic in modern GPUs. We also describe MAGMA’s direct impact on several ECP projects. Maintaining portable performance across NVIDIA and AMD GPUs, and with current efforts toward supporting Intel GPUs, MAGMA ensures its adaptability and relevance in the ever-evolving landscape of GPU architectures.

97 MATHEMATICS AND COMPUTING↗

Large-Scale Optimization with Linear Equality Constraints Using Reduced Compact Representation

For optimization problems with linear equality constraints, we prove that the (1,1) block of the inverse KKT matrix remains unchanged when projected onto the nullspace of the constraint matrix. In this work, we develop reduced compact representations of the limited-memory inverse BFGS Hessian to compute search directions efficiently when the constraint Jacobian is sparse. Orthogonal projections are implemented by a sparse QR factorization or a preconditioned LSQR iteration. In numerical experiments two proposed trust-region algorithms improve in computation times, often significantly, compared to previous implementations of related algorithms and compared to IPOPT.

97 MATHEMATICS AND COMPUTING↗

Exploration with Scalable Gaussian Process Reinforcement Learning

Exploration is a challenging problem in reinforcement learning (RL), especially in environments with sparse rewards. Quantifying and utilizing the parametric uncertainty has been shown to be paramount for successful exploration [Osband et al., 2018]. Bayesian, or approximately Bayesian, methods present a principled means of estimating the parametric uncertainty in RL problems. Gaussian processes, nonparametric Bayesian models, are often impractical due to poor scalability and computational bottlenecks. We introduce a scalable Gaussian process RL (GPRL) method which directly induces sparsity in the covariance matrix to facilitate faster computation. This is a departure from previous GPRL methods which instead rely on data reduction and subsampling. We compare various covariance-based exploration techniques (Thompson sampling, upper confidence bound, and probabilistic maximum variance) which leverage our scalable GP framework in sparse reward environments. Finally, we show favorable comparison against the bootstrapped deep Q-Network.

97 MATHEMATICS AND COMPUTING↗

Deployment of ADTimePix3 areaDetector Driver at Neutron and X-ray User Facilities

TimePix3 is a 65k hybrid pixel readout chip with simultaneous Time-of-Arrival (ToA) and Time-over-Threshold (ToT) recording in each pixel*. The chip operates without a trigger signal with a sparse readout where only pixels containing events are read out. The flexible architecture allows 40 MHits/s/cm² readout throughput, using simultaneous readout and acquisition by sharing readout logic with transport logic of superpixel matrix formed using 2x4 structure. The chip ToA records 1.5625 ns time resolution. The X-ray and charged particle events are counted directly. However, indirect neutron counts use 6Li fission in a scintillator matrix, such as ZnS(Ag). The fission space-charge region is limited to 5-9 um. A photon from scintillator material excites a photocathode electron, which is further multiplied in dual-stack MCP. The neutron count event is a cluster of electron events at the chip. We report on the EPICS areaDetector** ADTimePix3 driver that controls Serval*** using json commands. The driver directs data to storage and to a real-time processing pipeline and configures the chip. The time-stamped data are stored in raw .tpx3 file format and passed through a socket where the clustering software identifies individual neutron events. The conventional 2D images are available as images for each exposure frame, and a preview is useful for sample alignment. The areaDetector driver allows integration of time-enhanced capabilities of this detector into SNS beamlines controls and unprecedented time resolution.

Gofron, Kaz↗

Bit-GraphBLAS: Bit-Level Optimizations of Matrix-Centric Graph Processing on GPU

In the graph data structure like adjacency matrix, the connectivity of two nodes can be sufficiently represented using only 1 bit, but they are generally treated as 32-bit full-precision in state-of-the-art graph frameworks to adopt common sparse format such as CSR. Meanwhile, bit-level parallelism has recently be explored to have high-performance potential and low storage requirement on GPUs with dense bit-tiles. To fill the gap, our solution is a hierarchical storage format that contains the bit-indexing base and dense bit-tile units. Inherently, the granularity of the bit-tile is an essential factor in achieving both storage compression and GPU parallelism. How to find a sweet spot that trades off between avoiding sparsity and exploiting is comprehensively researched in this work. In the experiment, we evaluate the proposed storage format and algorithms on modern generation GPUs, including Pascal and Volta, to figure out critical software co-designs in conjunction with existing hardware-specific optimization.

Chen, Jou-An↗

Analysis of the ratio of ℓ 1 and ℓ 2 norms in compressed sensing

We study the ratio of ℓ 1 and ℓ 2 norms ( ℓ 1 / ℓ 2 ) as a sparsity-promoting objective in compressed sensing. We first propose a novel criterion that guarantees that an s-sparse signal is the local minimizer of the ℓ 1 / ℓ 2 objective; our criterion is interpretable and useful in practice. We also give the first uniform recovery condition using a geometric characterization of the null space of the measurement matrix, and show that this condition is satisfied for a class of random matrices. We also present analysis on the robustness of the procedure when noise pollutes data. Numerical experiments are provided that compare ℓ 1 / ℓ 2 with some other popular non-convex methods in compressed sensing. Finally, we propose a novel initialization approach to accelerate the numerical optimization procedure. We call this initialization approach support selection, and we demonstrate that it empirically improves the performance of existing ℓ 1 / ℓ 2 algorithms.

97 MATHEMATICS AND COMPUTING↗

Novel Application of Machine Learning Techniques for Rapid Source Apportionment of Aerosol Mass Spectrometer Datasets

In this work, we apply machine learning approaches sparse multinomial logistic regression to classify aerosol mass spectrometer (AMS) unit mass resolution (UMR) data followed by an ensemble regression technique for source apportionment of organic aerosols (OA). The classifier was trained on 60 well characterized laboratory and positive matrix factorization (PMF) deconvolved reference spectra to identify eight OA types. These include four laboratory-derived secondary organic aerosol (SOA) spectra, which include isoprene photooxidation SOA, isoprene epoxydiols (IEPOX) SOA, a monoterpene SOA type that includes a-pinene and ß-pinene SOA, and aromatic SOA from oxidation of naphthalene and m-xylene precursors, as well as PMF deconvolved spectra for three primary organic aerosol (POA) types, namely, hydrocarbon-like organic aerosol (HOA), biomass burning organic aerosol (BBOA), and cooking OA (COA), and a more oxidized oxygenated OA type (MO-OOA). A 5-fold cross-validation strategy, repeated 10 times, was used to assess the classifier’s performance. The classifier had high classification accuracy for COA, aromatic SOA, and isoprene SOA spectra but incorrectly classified ~9% by number of MO-OOA spectra as BBOA, 12% of BBOA spectra as HOA (and vice versa), and 18% of IEPOX-SOA spectra as aromatic SOA. Next, an ensemble regression model was trained on an artificially generated dataset consisting of mixtures of different OA types to assess its ability to predict fractional mass abundances from classification probabilities of various OA species obtained from the multinomial logistic regression classifier trained on the reference spectra. Ultimately, the proposed approach was applied for source apportionment of aircraft-based AMS measurements of OA UMR spectra during the HI-SCALE field campaign. On two representative days (May 6th and 18th, 2016), the algorithm determined that ~50-60% of OA by mass was MO-OOA, which represented a highly aged organic aerosol mixture from different sources. On both days, BBOA was determined to contribute less than 10% to OA by mass. However, on May 18th, the aromatic SOA fraction was higher compared to that on May 6th. The proposed approach is capable of rapidly analyzing AMS data in real time, making it suitable for applications where rapid source apportionment of AMS OA spectra is desirable.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Dynamics of disordered mechanical systems with large connectivity, free probability theory, and quasi-Hermitian random matrices

Disordered mechanical systems with high connectivity represent a limit opposite to the more familiar case of disordered crystals. Individual ions in a crystal are subjected essentially to nearest-neighbor interactions. In contrast, the systems studied in this paper have all their degrees of freedom coupled to each other. Thus, the problem of linearized small oscillations of such systems involves two full positive-definite and non-commuting matrices, as opposed to the sparse matrices associated with disordered crystals. Consequently, the familiar methods for determining the averaged vibrational spectra of disordered crystals, introduced many years ago by Dyson and Schmidt, are inapplicable for highly connected disordered systems. In this paper we apply random matrix theory (RMT) to calculate the averaged vibrational spectra of such systems, in the limit of infinitely large system size. At the heart of our analysis lies a calculation of the average spectrum of the product of two positive definite random matrices by means of free probability theory techniques. We also show that this problem is intimately related with quasi-hermitian random matrix theory (QHRMT), which means that the ‘hamiltonian’ matrix is hermitian with respect to a non-trivial metric. This extends ordinary hermitian matrices, for which the metric is simply the unit matrix. The analytical results we obtain for the spectrum agree well with our numerical results. The latter also exhibit oscillations at the high-frequency band edge, which fit well the Airy kernel pattern. We also compute inverse participation ratios of the corresponding amplitude eigenvectors and demonstrate that they are all extended, in contrast with conventional disordered crystals. Finally, we compute the thermodynamic properties of the system from its spectrum of vibrations. In addition to matrix model analysis, we also study the vibrational spectra of various multi-segmented disordered pendula, as concrete realizations of highly connected mechanical systems. A universal feature of the density of vibration modes, common to both pendula and the matrix model, is that it tends to a non-zero constant at vanishing frequency.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Sparsity of Radiating Characteristic Modes on Infinite Periodic Structures

Characteristic modes on infinite periodic structures are studied using spectral dyadic Green’s functions. This formulation demonstrates that, in contrast to the modal analysis of finite structures, the number of radiating characteristic modes is limited by unit cell size and incident wave vector (i.e., scan angle or phase shift per unit cell). Here, the reflection tensor is decomposed into modal contributions from radiating modes, indicating that characteristic modes are a predictably sparse basis in which to study reflection phenomena.

42 ENGINEERING↗

A Review of Advanced Test Reactor Fuel and Assessment of Its Compatibility with the ZIRCEX Chlorination Process

Advanced Test Reactor (ATR) fuel has been identified as a resource for high-assay low-enriched uranium (HALEU) production. A survey was performed on the published literature describing ATR fuel. The geometry of the fuel is complex; different parts of the fuel compact experience differing neutron flux and burnup. The literature is sparse, and access is controlled. Therefore, fundamental studies of fuel reprocessing must use a model fuel that represents the main chemical and structural features. Advanced chlorination, or chlorination with sulfur-chlorine bearing reagents is being investigated as way to separate the fuel from metal matrix alloys. A UAl x alloy will be fabricated with x = 3, 4, and 5. The potential chlorination of individual UAl x intermetallics will be assessed in the advanced chlorination process of Al-8001 and Al-6061 as well as a representative mixture. Initial studies will track the alloying elements of the Al, which are Si, Fe, Cu, Mn, Mg, Cr, Zn, and Ti, in addition to the U itself. Further studies will include fission product simulants. Because advanced chlorination solvents include sulfur, the chemistry of sulfur with major and minor constituents will also be investigated. The experimental work accompanied by neutronic calculations will allow the assessment of the feasibility of advanced chlorination to separate aluminum from uranium. If bench-scale testing appears promising, then small-scale tests in shielded facilities with irradiated cladding, lightly irradiated fuel, and spent nuclear fuel are recommended to track the complete inventory of fissile actinides, fission product impurities, and reagent solids and liquids.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Learning an Algebriac Multrigrid Interpolation Operator Using a Modified GraphNet Architecture

This work, building on previous efforts, develops a suite of new graph neural network machine learning architectures that generate data-driven prolongators for use in Algebraic Multigrid (AMG). Algebraic Multigrid is a powerful and common technique for solving large, sparse linear systems. Its effectiveness is problem dependent and heavily depends on the choice of the prolongation operator, which interpolates the coarse mesh results onto a finer mesh. Previous work has used recent developments in graph neural networks to learn a prolongation operator from a given coefficient matrix. In this paper, we expand on previous work by exploring architectural enhancements of graph neural networks. A new method for generating a training set is developed which more closely aligns to the test set. Asymptotic error reduction factors are compared on a test suite of 3-dimensional Poisson problems with varying degrees of element stretching. Results show modest improvements in asymptotic error factor over both commonly chosen baselines and learning methods from previous work.

97 MATHEMATICS AND COMPUTING↗