Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “matrix product operators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Biological removal of gaseous ammonia in biofilters: space travel and earth-based applications

Gaseous NH3 removal was studied in laboratory-scale biofilters (14-L reactor volume) containing perlite inoculated with a nitrifying enrichment culture. These biofilters received 6 L/min of airflow with inlet NH3 concentrations of 20 or 50 ppm, and removed more than 99.99% of the NH3 for the period of operation (101, 102 days). Comparison between an active reactor and an autoclaved control indicated that NH3 removal resulted from nitrification directly, as well as from enhanced absorption resulting from acidity produced by nitrification. Spatial distribution studies (20 ppm only) after 8 days of operation showed that nearly 95% of the NH3 could be accounted for in the lower 25% of the biofilter matrix, proximate to the port of entry. Periodic analysis of the biofilter material (20 and 50 ppm) showed accumulation of the nitrification product NO3- early in the operation, but later both NO2- and NO3- accumulated. Additionally, the N-mass balance accountability dropped from near 100% early in the experiments to approximately 95 and 75% for the 20- and 50-ppm biofilters, respectively. A partial contributing factor to this drop in mass balance accountability was the production of NO and N2O, which were detected in the biofilter exhaust.

NASA Discipline Life Support Systems↗

Spectrally Stabilized Interface Capturing Formulation and Implementation in Nek5000/NekRS

This report documents the formulation of a novel level-set method for incompressible two-phase flows in the continuous Galerkin (CG) high order spectral element framework. The overall method hinges on a novel implementation of the spectral vanishing viscosity (SVV) operator for the stabilization of linear/non-linear hyperbolic problems. The multidimensional SVV convolution kernels, which in essence, have a similar effect as a high pass filter applied to the derivatives, are formulated by exploiting the tensor product form, analogous to the construction of the usual stiffness matrix system. The resulting kernels are directionally decoupled and ensure a linear, symmetric positive definite, elliptic matrix operator. The SVV formulation is demonstrated to provide a robust stabilizing mechanism through challenging linear and non-linear hyperbolic problems, including problems pertinent to the level-set formulation. The two-phase framework conceptualized herein is based on the conservative level-set (CLS) method which represents the interface between the fluids by the 0.5 iso-contour of the smoothed Heaviside function. The CLS method is augmented with a preconditioning procedure for interface normals using the signed distance function which precludes the manifestation of spurious oscillations in the vicinty of the interface. Further, the existing mixed explicit-implicit approach for the solution of Navier-Stokes equations in Nek5000, as described in Tomboulides et al, is augmented with a pressure coefficient splitting approach for the Poisson equation, which greatly accelerated the convergence of pressure solver for two-phase systems with large density ratio. The robustness and accuracy of the overall two-phase method is demonstrated through canonical challenging problems involving high density and viscosity ratios, with and without surface tension. The two-phase formulation is wholly implemented in Nek5000 and the SVV stabilization method is implemented in NekRS, which is the essential precursor to the two-phase framework, undergoing active development.

97 MATHEMATICS AND COMPUTING↗

Insolubilized enzymes for food synthesis

Cellulose matrix with numerous enzyme-coated silica particles of colloidal size permanently bound at various sites within matrix was produced that has high activity and possesses requisite physical characteristics for filtration or column operations. Product also allows coupling step in synthesis of edible food to proceed under mild conditions.

Marshall, D. L.↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗

Nested Krylov methods and preserving the orthogonality

Recently the GMRESR inner-outer iteraction scheme for the solution of linear systems of equations was proposed by Van der Vorst and Vuik. Similar methods have been proposed by Axelsson and Vassilevski and Saad (FGMRES). The outer iteration is GCR, which minimizes the residual over a given set of direction vectors. The inner iteration is GMRES, which at each step computes a new direction vector by approximately solving the residual equation. However, the optimality of the approximation over the space of outer search directions is ignored in the inner GMRES iteration. This leads to suboptimal corrections to the solution in the outer iteration, as components of the outer iteration directions may reenter in the inner iteration process. Therefore we propose to preserve the orthogonality relations of GCR in the inner GMRES iteration. This gives optimal corrections; however, it involves working with a singular, non-symmetric operator. We will discuss some important properties, and we will show by experiments that, in terms of matrix vector products, this modification (almost) always leads to better convergence. However, because we do more orthogonalizations, it does not always give an improved performance in CPU-time. Furthermore, we will discuss efficient implementations as well as the truncation possibilities of the outer GCR process. The experimental results indicate that for such methods it is advantageous to preserve the orthogonality in the inner iteration. Of course we can also use iteration schemes other than GMRES as the inner method; methods with short recurrences like GICGSTAB are of interest.

Desturler, Eric↗

Exact two-body expansion of the many-particle wave function

Progress toward the solution of the strongly correlated electron problem has been stymied by the exponential complexity of the wave function. Previous work established an exact two-body exponential product expansion for the ground-state wave function. By developing a reduced density-matrix analog of Dalgarno-Lewis perturbation theory, we prove here that (i) the two-body exponential product expansion is rapidly and globally convergent with each operator representing an order of a renormalized perturbation theory, (ii) the energy of the expansion converges quadratically near the solution, and (iii) the expansion is exact for both ground and excited states. The two-body expansion offers a reduced parametrization of the many-particle wave function as well as the two-particle reduced density matrix with potential applications on both conventional and quantum computers for the study of strongly correlated quantum systems. In this work, we demonstrate the result with the exact solution of the contracted Schrödinger equation for the molecular chains H 4 and H 5 .

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Dual-unitary shadow tomography

We introduce a classical shadow tomography scheme based on dual-unitary brick-wall circuits termed "dual-unitary shadow tomography" (DUST). For this we study operator spreading and Pauli weight dynamics in one-dimensional qubit systems, evolved by random two-local dual-unitary gates arranged in a brick-wall structure, ending with a final measurement layer. We do this by deriving general constraints on the Pauli weight transfer matrix and specializing to the case of dual-unitarity. We first show that dual-unitaries must have a minimal amount of entropy production. Remarkably, we find that operator spreading in these circuits have a rich structure resembling that of relativistic quantum field theories, with massless chiral excitations that can decay or fuse into each other, which we call left- or right-movers. We develop a mean-field description of the Pauli weight in terms of $\rho(x,t)$, which represents the probability of having nontrivial support at site $x$ and depth $t$ starting from a fixed weight distribution. We develop an equation of state for $\rho(x,t)$, and simulate it numerically using Monte Carlo simulations. Lastly, we demonstrate that the fast-thermalizing properties of dual-unitary circuits make them better at predicting large operators than shallow brick-wall Clifford circuits. Our results are robust to finite-size effects due to the chirality of dual-unitary brick-wall circuits.

97 MATHEMATICS AND COMPUTING↗

Random Phase Approximation Correlation Energy Using Real-Space Density Functional Perturbation Theory

We present a real-space method for computing the random phase approximation (RPA) correlation energy within Kohn–Sham density functional theory, leveraging the low-rank nature of the frequency-dependent density response operator. In particular, we employ a cubic-scaling formalism based on density functional perturbation theory that circumvents the calculation of the response function matrix, instead relying on the ability to compute its product with a vector through the solution of the associated Sternheimer linear systems. We develop a large-scale parallel implementation of this formalism using the subspace iteration method in conjunction with the spectral quadrature method while employing the Kronecker product-based method for the application of the Coulomb operator and the conjugate orthogonal conjugate gradient method for the solution of the linear systems. We demonstrate convergence with respect to key parameters and verify the method’s accuracy by comparing with plane-wave results. We show that the framework achieves good strong scaling to many thousands of processors, reducing the time to solution for a lithium hydride system with 128 electrons to around 150 s on 4608 processors.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Long-life high performance fuel cell program

A multihundred kilowatt Regenerative Fuel Cell for use in a space station is envisioned. Three 0.508 sq ft (471.9 cm) active area multicell stacks were assembled and endurance tested. The long term performance stability of the platinum on carbon catalyst configuration suitability of the lightweight graphite electrolyte reservoir plate, the stability of the free standing butyl bonded potassium titanate matrix structure, and the long life potential of a hybrid polysulfone cell edge frame construction were demonstrated. A 18,000 hour demonstration test of multicell stack to a continuous cyclical load profile was conducted. A total of 12,000 cycles was completed, confirming the ability of the alkaline fuel cell to operate to a load profile simulating Regenerative Fuel Cell operation. An orbiter production hydrogen recirculation pump employed in support of the cyclical load profile test completed 13,000 hours of maintenance free operation. Laboratory endurance tests demonstrated the suitability of the butyl bonded potassium matrix, perforated nickel foil electrode substrates, and carbon ribbed substrate anode for use in the alkaline fuel cell. Corrosion testing of materials at 250 F (121.1 C) in 42% wgt. potassium identified ceria, zirconia, strontium titanate, strontium zirconate and lithium cobaltate as candidate matrix materials.

Martin, R. E.↗

Mapping unstructured grid computations to massively parallel computers

Investigated here is this mapping problem: assign the tasks of a parallel program to the processors of a parallel computer such that the execution time is minimized. First, a taxonomy of objective functions and heuristics used to solve the mapping problem is presented. Next, we develop a highly parallel heuristic mapping algorithm, called Cyclic Pairwise Exchange (CPE), and discuss its place in the taxonomy. CPE uses local pairwise exchanges of processor assignments to iteratively improve an initial mapping. A variety of initial mapping schemes are tested and recursive spectral bipartitioning (RSB) followed by CPE is shown to result in the best mappings. For the test cases studied here, problems arising in computational fluid dynamics and structural mechanics on unstructured triangular and tetrahedral meshes, RSB and CPE outperform methods based on simulated annealing. Much less time is required to do the mapping and the results obtained are better. Compared with random and naive mappings, RSB and CPE reduce the communication time two fold for the test problems used. Finally, we use CPE in two applications on a CM-2. The first application is a data parallel mesh-vertex upwind finite volume scheme for solving the Euler equations on 2-D triangular unstructured meshes. CPE is used to map grid points to processors. The performance of this code is compared with a similar code on a Cray-YMP and an Intel iPSC/860. The second application is parallel sparse matrix-vector multiplication used in the iterative solution of large sparse linear systems of equations. We map rows of the matrix to processors and use an inner-product based matrix-vector multiplication. We demonstrate that this method is an order of magnitude faster than methods based on scan operations for our test cases.

Hammond, Steven Warren↗

Macroscopic instructions vs microscopic operations in quantum circuits

In many experiments on microscopic quantum systems, it is implicitly assumed that when a macroscopic procedure or “instruction” is repeated many times – perhaps in different contexts – each application results in the same microscopic quantum operation. But in practice, the microscopic effect of a single macroscopic instruction can easily depend on its context. If undetected, this can lead to unexpected behavior and unreliable results. Here, we design and analyze several tests to detect context-dependence. They are based on invariants of matrix products, and while they can be as data intensive as quantum process tomography, they do not require tomographic reconstruction, and are insensitive to imperfect knowledge about the experiments. We also construct a measure of how unitary (reversible) an operation is, and show how to estimate the volume of physical states accessible by a quantum operation.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Acoustooptic linear algebra processors - Architectures, algorithms, and applications

Architectures, algorithms, and applications for systolic processors are described with attention to the realization of parallel algorithms on various optical systolic array processors. Systolic processors for matrices with special structure and matrices of general structure, and the realization of matrix-vector, matrix-matrix, and triple-matrix products and such architectures are described. Parallel algorithms for direct and indirect solutions to systems of linear algebraic equations and their implementation on optical systolic processors are detailed with attention to the pipelining and flow of data and operations. Parallel algorithms and their optical realization for LU and QR matrix decomposition are specifically detailed. These represent the fundamental operations necessary in the implementation of least squares, eigenvalue, and SVD solutions. Specific applications (e.g., the solution of partial differential equations, adaptive noise cancellation, and optimal control) are described to typify the use of matrix processors in modern advanced signal processing.

Casasent, D.↗

Pattern classification using Charge Transfer Devices

The potential uses of Charge Transfer Devices (CTDs) in pattern classification operations are explored. The needs for a hardware-based pattern classifier are established, and a matrix multiplication subsystem based upon a sum of products CTD is presented. An evaluation process for sum of products devices (particularly analog-analog correlators) is developed, and the feasibility of employing a particular device in a pattern classifier is determined. Finally, the possible impact of future trends in technology is considered.

Snyder, W. E.↗

Two decades of DOE investment lays the foundation for TRISO-fueled reactors

Tristructural isotropic (TRISO) coated particle fuel is a robust, microencapsulated fuel form developed originally for use in high-temperature gas-cooled reactors (HTGRs). The particles consist of a spherical fissile kernel surrounded by several layers of pyrocarbon and a silicon carbide (SiC) layer (Figure 1). The particles are formed into cylindrical or spherical fuel forms using a resinated graphite matrix material for insertion into an HTGR. The kernel and coating layers together act to retain fission products within the particle during normal reactor operation and during postulated accidents; TRISO particles can maintain structural integrity at extremely high temperatures, reaching as high as approximately 1,600°C in limiting HTGR accidents. This limits the fission product activity circulating in the helium coolant and the activity released to the environment during accidents. Acceptable performance of TRISO particles is therefore essential for reactor safety.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Exploring the speciation of actinide salts in the presence of contaminants related to molten salt reactors

Molten salt reactors (MSR) are up-and-coming Generation IV nuclear reactors with either coolant and/or radioactive fuel in molten salt form. Due to their improved safety, efficiency, affordability and convenient waste processing system, these reactors are considered a superior alternative to conventional ones. However, prior to commercial implementation, it is necessary to have an extensive understanding of the reactions occurring in the MSR. The safety parameters of the MSR are determined by the rheological properties such as density and viscosity of the molten salt and these properties are governed by the local structure of the molten salts. Due to the high operational temperatures of MSR's, corrosion plays an important role. Because of their influence, the container corrosion products become part of the fuel salt matrix. In this work, we are investigating how these corrosion products, e.g., Mo, Ni, Cr, would impact the speciation of the actinide metal centers. For the initial studies, lanthanides are used as a surrogate for actinides. Lanthanide chloride salts (LnCl3) are mixed with transition metals and excess of alkali or alkali earth metal salts or salt mixtures (LiCl, NaCl, KCl). These are added to an alumina crucible, placed in a quartz tube, and sealed under a vacuum. This is heated up to around 1000 °C and slowly cooled to a temperature above the melting point of the salt and at that temperature, the reaction is taken out of the furnace. The excess salt is removed, and the resulting product is analyzed using X-ray diffraction techniques, scanning electron microscopy, and spectroscopic methods such as Raman, UV-Visible, and FTIR spectroscopy.

37 - INORGANIC, ORGANIC, PHYSICAL AND ANALYTICAL C↗

Complexity of Kronecker Operations on Sparse Matrices with Applications to the Solution of Markov Models

We present a systematic discussion of algorithms to multiply a vector by a matrix expressed as the Kronecker product of sparse matrices, extending previous work in a unified notational framework. Then, we use our results to define new algorithms for the solution of large structured Markov models. In addition to a comprehensive overview of existing approaches, we give new results with respect to: (1) managing certain types of state-dependent behavior without incurring extra cost; (2) supporting both Jacobi-style and Gauss-Seidel-style methods by appropriate multiplication algorithms; (3) speeding up algorithms that consider probability vectors of size equal to the "actual" state space instead of the "potential" state space.

Buchholz, Peter↗

PANTHER: A Programmable Architecture for Neural Network Training Harnessing Energy-Efficient ReRAM

The wide adoption of deep neural networks has been accompanied by ever-increasing energy and performance demands due to the expensive nature of training them. Additionally, numerous special-purpose architectures have been proposed to accelerate training: both digital and hybrid digital-analog using resistive RAM (ReRAM) crossbars. ReRAM-based accelerators have demonstrated the effectiveness of ReRAM crossbars at performing matrix-vector multiplication operations that are prevalent in training. However, they still suffer from inefficiency due to the use of serial reads and writes for performing the weight gradient and update step. A few works have demonstrated the possibility of performing outer products in crossbars, which can be used to realize the weight gradient and update step without the use of serial reads and writes. However, these works have been limited to low precision operations which are not sufficient for typical training workloads. Moreover, they have been confined to a limited set of training algorithms for fully-connected layers only. To address these limitations, we propose a bit-slicing technique for enhancing the precision of ReRAM-based outer products, which is substantially different from bit-slicing for matrix-vector multiplication only. We incorporate this technique into a crossbar architecture with three variants catered to different training algorithms. To evaluate our design on different types of layers in neural networks (fully-connected, convolutional, etc.) and training algorithms, we develop PANTHER, an ISA-programmable training accelerator with compiler support. Our design can also be integrated into other accelerators in the literature to enhance their efficiency. Our evaluation shows that PANTHER achieves up to 8.02×, 54.21×, and 103× energy reductions as well as 7.16×, 4.02×, and 16× execution time reductions compared to digital accelerators, ReRAM-based accelerators, and GPUs, respectively.

42 ENGINEERING↗

Micromechanical response of SiC-OPyC layers in TRISO fuel particles

Tristructural isotropic (TRISO)–coated particle fuel is a proposed fuel for multiple advanced reactor concepts. The performance of the particle depends on whether the silicon carbide (SiC) layer remains intact to prevent the release of metallic and gaseous fission products. Mechanical fracture of the SiC layer is a potential failure mode under various fuel configurations and operating environments, including the potential transmission of matrix-originating cracks through TRISO particles. Furthermore, this study uses instrumented indentation techniques on cross-sectioned surrogate particles to examine the mechanical stability of the critical interface between SiC and the outer pyrolytic carbon (OPyC) layer. The observed behavior at the interface is rationalized by examining the radially dependent fracture behavior of the SiC layer and performing a numerical analysis to quantify the residual stresses that develop during the processing and cross-sectioning of the as-fabricated particle. Characterizing the SiC-OPyC interface of surrogate TRISO particles using nanoindentation provides unique insight into the interface's room-temperature residual stress and mechanical stability. The modeling efforts were used to investigate the experimental procedure further, and the results are presented herein to validate this fuel form's potential mechanical failure modes.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗