Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “matrix multiplication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Evaluating performance and portability of high-level programming models: Julia, Python/Numba, and Kokkos on exascale nodes

We explore the performance and portability of the high-level programming models: the LLVM-based Julia and Python/Numba, and Kokkos on high-performance computing (HPC) nodes: AMD Epyc CPUs and MI250X graphical processing units (GPUs) on Frontier’s test bed Crusher system and Ampere’s Arm-based CPUs and NVIDIA’s A100 GPUs on the Wombat system at the Oak Ridge Leadership Computing Facilities. We compare the default performance of a hand-rolled dense matrix multiplication algorithm on CPUs against vendor-compiled C/OpenMP implementations, and on each GPU against CUDA and HIP. Rather than focusing on the kernel optimization per-se, we select this naive approach to resemble exploratory work in science and as a lower-bound for performance to isolate the effect of each programming model. Julia and Kokkos perform comparably with C/OpenMP on CPUs, while Julia implementations are competitive with CUDA and HIP on GPUs. Performance gaps are identified on NVIDIA A100 GPUs for Julia’s single precision and Kokkos, and for Python/Numba in all scenarios. We also comment on half-precision support, productivity, performance portability metrics, and platform readiness. We expect to contribute to the understanding and direction for high-level, high-productivity languages in HPC as the first-generation exascale systems are deployed.

Godoy, William↗

A Benchmark Suite for Evaluating Scientific AI Workloads on GPUs

AI applications have been steadily increasing in the allocation portfolio among leadership computing facilities. These applications depend on deep learning frameworks with hardware acceleration and underlying software systems. With the rapid development of applications, software stacks, and hardware devices, it is essential to evaluate the performance of core operations in AI workloads for direction of optimizations and procurement of next-generation high-performance computing (HPC) infrastructures. Currently, most benchmarks lack scientific AI workloads. So, we present DeepKernelBench and the experimental results of evaluating the benchmark suite for early observations and performance comparisons on datacenter GPUs using representative workloads for scientific AI, including Attentions, General matrix multiplications, Geometrics and Fourier neural operations.

Jin, Zheming [Advanced Micro Devices (AMD)]↗

A Flexible Forwarding Scheme to Improve Latency-Bound Irregular P2P Communication in MPI

We propose an algorithm to efficiently perform latency-bound communication scenarios that consist of many small messages. In these parallel scenarios, processes typically pass around a lot of small-sized messages of a few KBs of size. Performing communication operations with P2P MPI routines or collective MPI routines (including neighborhood collectives) in such scenarios may not always yield the optimal results and may not resolve the latency bottleneck. To this end, we develop a regular structure called virtual process topology (VPT) on which the messages can be communicated in a structured and controlled manner. Using parameters of this topology, one can tune the rate of aggression in tackling the latency costs. We demonstrate that our communication algorithm is preferable to MPI P2P and collective routines for latency-bound communication and it can easily be adapted only by replacing calls to MPI routines in a parallel application. We show how to adapt existing topology-aware mapping heuristics to address the volume overhead due to communicating messages on the VPT. Moreover, we propose a novel swap-based mapping heuristic to address this overhead by optimizing the maximum volume handled by a process. Experiments on synthetic communication graphs as well as real-world applications such as parallel Canonical Polyadic sparse tensor decomposition and parallel sparse matrix-dense matrix multiplication show that our approach is a powerful way of overcoming the bottlenecks posed by sparse and latency-bound irregular communication.

communication algorithm↗

Mixed-Precision S/DGEMM Using the TF32 and TF64 Frameworks on Low-Precision AI Tensor Cores

Using NVIDIA graphics processing units (GPUs) equipped with Tensor Cores has enabled the significant acceleration of general matrix multiplication (GEMM) for applications in machine learning (ML) and artificial intelligence (AI) and in high-performance computing (HPC) generally. The use of such power-efficient, specialized accelerators can provide a performance increase between 8 × and 20 ×, albeit with a loss in precision. However, a high level of precision is required in many large scientific and HPC applications, and computing in single or double precision is still necessary for many of these applications to maintain accuracy. Fortunately, mixed-precision methods can be employed to maintain a higher level of numerical precision while also taking advantage of the performance increases from computing with lower-precision AI cores. With this in mind, we extend the state of the art by using NVIDIA’s new TF32 framework. This new framework not only burdens some constraints of the previous frameworks, such as costly 32 16-bit castings but also provides an equivalent precision and performance by using a much simpler approach. We also propose a new framework called TF64 that attempts double-precision arithmetic with low-precision Tensor Cores. Although this framework does not exist yet, we validated the correctness of this idea and achieved an equivalent of 64-bit precision on 32-bit hardware.

Valero Lara, Pedro↗

FTTN: Feature-Targeted Testing for Numerical Properties of NVIDIA & AMD Matrix Accelerators

FTTN is a test suite to evaluate the numerical behaviors of matrix accelerators of GPUs (NVIDIA Tensor Cores and AMD Matrix Cores) in a quick and simple setting. Matrix accelerators are heavily used in today's computationally intense applications to speed up matrix multiplications. This test suite provides a comprehensive study on the numerical behaviors of these accelerators, including support for subnormals, rounding modes, extra precision bits and FMA features. Is there

Laguna Peralta, Ignacio↗

A Provably Accurate Randomized Sampling Algorithm for Logistic Regression

In statistics and machine learning, logistic regression is a widely-used supervised learning technique primarily employed for binary classification tasks. When the number of observations greatly exceeds the number of predictor variables, we present a simple, randomized sampling-based algorithm for logistic regression problem that guarantees high-quality approximations to both the estimated probabilities and the overall discrepancy of the model. Our analysis builds upon two simple structural conditions that boil down to randomized matrix multiplication, a fundamental and well-understood primitive of randomized numerical linear algebra. We analyze the properties of estimated probabilities of logistic regression when leverage scores are used to sample observations, and prove that accurate approximations can be achieved with a sample whose size is much smaller than the total number of observations. To further validate our theoretical findings, we conduct comprehensive empirical evaluations. Overall, our work sheds light on the potential of using randomized sampling approaches to efficiently approximate the estimated probabilities in logistic regression, offering a practical and computationally efficient solution for large-scale datasets.

Chowdhury, Agniva↗

Accelerating Floating-Point Computations with Intel AMX

Intel AMX is a built-in component of recent Intel CPU architectures, first supported by the Intel Sapphire Rapids in 2023, that enables efficient dense matrix multiplications using mixed precision with low-precision data types. The popularity of mixed-precision algorithms has grown recently, primarily due to their use on GPUs to enhance the efficiency of HPC applications, particularly for the training of large language models. The availability of mixed precision on CPUs represents a cost-effective solution for applications where high speed is not critical. This report shows how to use the Intel AMX accelerator through examples in C++ and Python. The examples will focus on mixed-precision floating-point operations obtained by the use of bfloat16 (or BF16) to accelerate code in single precision. We employ a bottom-up methodology, starting from specific register instructions (TMUL operation) to higher-level applications in libraries such as Intel MKL, PyTorch, and TensorFlow, ensuring a comprehensive understanding of the accelerator's potential. Additionally, we provide insights into the expected performance gains when leveraging the accelerator on the Kestrel HPC machine at the National Renewable Energy Laboratory.

97 MATHEMATICS AND COMPUTING↗

Energy-Efficient Neuromorphic Architectures for Nuclear Radiation Detection Applications

A comprehensive analysis and simulation of two memristor-based neuromorphic architectures for nuclear radiation detection is presented. Both scalable architectures retrofit a locally competitive algorithm to solve overcomplete sparse approximation problems by harnessing memristor crossbar execution of vector–matrix multiplications. The proposed systems demonstrate excellent accuracy and throughput while consuming minimal energy for radionuclide detection. To ensure that the simulation results of our proposed hardware are realistic, the memristor parameters are chosen from our own fabricated memristor devices. Based on these results, we conclude that memristor-based computing is the preeminent technology for a radiation detection platform.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

TEE-ACM2

TEE-ACM2 is a library that computes matrix chain multiplication efficiently on GPUs, using a blocking strategy to load data. The library will provide an energy efficient algorithm for chain Matrix Multiplication, by minimizing both computations and off chip data transfers on the GPUs.

Lim, Hyun↗

Communication Lower Bounds and Optimal Algorithms for Symmetric Matrix Computations

In this article, we focus on the communication costs of three symmetric matrix computations: (i) multiplying a matrix with its transpose, known as a symmetric rank-k update (SYRK) (ii) adding the result of the multiplication of a matrix with the transpose of another matrix and the transpose of that result, known as a symmetric rank-2k update (SYR2K) (iii) performing matrix multiplication with a symmetric input matrix (SYMM). All three computations appear in the Level 3 Basic Linear Algebra Subroutines (BLAS) and have wide use in applications involving symmetric matrices. We establish communication lower bounds for these kernels using sequential and distributed-memory parallel computational models, and we show that our bounds are tight by presenting communication-optimal algorithms for each setting. Our lower bound proofs rely on applying a geometric inequality for symmetric computations and analytically solving constrained nonlinear optimization problems. As a result, the symmetric matrix and its corresponding computations are accessed and performed according to a triangular block partitioning scheme in the optimal algorithms.

Al Daas, Hussam [Rutherford Appleton Laboratory, D↗

239 Pu R -matrix Analysis and Neutron Multiplicities in the Neutron Energy Region up to a few keVs [Abstract]

The evaluation of 239 Pu neutron resonance parameters coupled to neutron multiplicities $\overline{v}_p$ is of particular importance to investigate the ($\mathcal{n, γf}$) reaction in which a $\mathcal{γ}$-ray emission occurs before the scission of the compound nuclear. This reaction has offered one explanation for the fluctuations in the measured values of $\overline{v}_p$ In this regard, the competition between ($\mathcal{n, γf}$) reaction and the direct fission process can be also included in the R matrix analysis of fission and capture measured data. The goal of this work is the coupled evaluation of the $\mathcal{n}$+ 239 Pu resonance parameters and related neutron multiplicities by ensuring the adoption of thermal neutron constants recently evaluated at the International Atomic Nuclear Energy as well as the recommended (thermal-neutron) induced prompt neutron fission spectrum (PFNS). Moreover, this new set of physical evaluated quantities should also guarantee the agreement for high-leakage solution benchmarks while keeping the good performance of large thermal solution assemblies.

07 ISOTOPE AND RADIATION SOURCES↗

Fast and Scalable FFT-Based GPU-Accelerated Algorithms for Block-Triangular Toeplitz Matrices with Application to Linear Inverse Problems Governed by Autonomous Dynamical Systems

In this work, we present an efficient and scalable algorithm for performing matrix-vector multiplications (matvecs) for block Toeplitz matrices. Such matrices, which are shift-invariant with respect to their blocks, arise in the context of solving inverse problems governed by autonomous systems, and time-invariant systems in particular. In this article, we consider inverse problems that infer unknown parameters from observational data of a linear time-invariant dynamical system given in the form of partial differential equations (PDEs). Matrix-free Newton-conjugate-gradient methods are often the gold standard for solving these inverse problems, but they require numerous actions of the Hessian on a vector. Matrix-free adjoint-based Hessian matvecs require solution of a pair of linearized forward/adjoint PDE solves per Hessian action, which may be prohibitive for large-scale inverse problems. Time invariance of the forward PDE problem leads to a block Toeplitz structure of the discretized parameter-to-observable (p2o) map defining the mapping from inputs (parameters) to outputs (observables) of the PDEs. This block Toeplitz structure enables us to exploit two key properties: (1) compact storage of the p2o map and its adjoint, and (2) efficient fast Fourier transform–based Hessian matvecs. The proposed algorithm is mapped onto large multi-GPU clusters and achieves more than 80% of peak bandwidth on NVIDIA A100 GPUs. Excellent weak scaling is shown for up to 48 A100 GPUs. For the targeted problems, the implementation executes Hessian matvecs within fractions of a second, which is orders of magnitude faster than can be achieved by conventional matrix-free Hessian matvecs via forward/adjoint PDE solves.

97 MATHEMATICS AND COMPUTING↗

Compute in‐Memory with Non‐Volatile Elements for Neural Networks: A Review from a Co‐Design Perspective

Abstract Deep learning has become ubiquitous, touching daily lives across the globe. Today, traditional computer architectures are stressed to their limits in efficiently executing the growing complexity of data and models. Compute‐in‐memory (CIM) can potentially play an important role in developing efficient hardware solutions that reduce data movement from compute‐unit to memory, known as the von Neumann bottleneck. At its heart is a cross‐bar architecture with nodal non‐volatile‐memory elements that performs an analog multiply‐and‐accumulate operation, enabling the matrix‐vector‐multiplications repeatedly used in all neural network workloads. The memory materials can significantly influence final system‐level characteristics and chip performance, including speed, power, and classification accuracy. With an over‐arching co‐design viewpoint, this review assesses the use of cross‐bar based CIM for neural networks, connecting the material properties and the associated design constraints and demands to application, architecture, and performance. Both digital and analog memory are considered, assessing the status for training and inference, and providing metrics for the collective set of properties non‐volatile memory materials will need to demonstrate for a successful CIM technology.

36 MATERIALS SCIENCE↗

Dynamic flux surrogate-based partitioned methods for interface problems

Loosely coupled partitioned methods for multiphysics problems treat each subproblem as a separate entity and advance them independently in time. In so doing these methods enable code reuse, increase concurrency and provide a convenient framework for plug-and-play multiphysics simulations. However, mathematically loosely coupled schemes are equivalent to a single step of an iterative solution method, which can compromise their accuracy and stability. We present a new data-driven partitioned method for coupled parametric PDEs that can improve upon the accuracy of traditional loosely coupled methods without incurring a performance penalty. To that end, we replace conventional field transfers across the interface by a surrogate for the dynamics of the interface flux exchanged between the subdomains. To develop this surrogate we apply dynamic mode decomposition to a non-standard staggered-in-time state, comprising the interface flux and small solution patches near the interface. The new approach shifts the main computational burden to an offline training phase, whereas application of the surrogate in the online phase amounts to a single matrix–vector multiplication. In conclusion, we provide stability analysis of the surrogate-based partitioned scheme and include numerical results that demonstrate its potential.

Dynamic mode decomposition (DMD)↗

Current and future federal and state sampling guidance for per- and polyfluoroalkyl substances in environmental matrices

Per- and polyfluoroalkyl substances (PFAS) are a class of emerging contaminants composed of an estimated 5000 to 10,000 human-made, fluorinated, organic chemicals. Due to the complexity of PFAS, the need for multiple environmental matrix considerations and the absence of a promulgated federal standard for environmental sampling and analysis, U.S. states have begun developing health-based regulatory and/or guidance values for a limited number of PFAS in environmental matrices. As there is a growing body of science to inform PFAS sampling guidance standard development, it is important to understand which U.S. states are implementing sampling guidelines and how they plan to handle emerging PFAS. This critical review discusses the current and impending federal and state sampling guidelines for PFAS in environmental matrices, the data gaps surrounding PFAS sampling guidance in U.S. states, and the future impacts of impending guidance documents and regulations. Ten federal guidance documents are available for PFAS sampling guidance and analysis. The maximum number of PFAS covered in these guidance documents is 25 analytes spanning across 8 unique media. While the EPA has developed several different sampling and analytical guidelines for PFAS, there is no formal regulation of PFAS or requirements of states to enforce these guidelines. Consequently, only 31 states have informally adopted sampling guidelines, while the other 19 states have no guidance documentation in place for PFAS. The introduction of new PFAS sampling guidelines by the EPA, as well as updated analytical guidelines that target more PFAS or total organofluoride, is expected to continuously shift the landscape of federal and state guidance for PFAS sampling moving forward.

54 ENVIRONMENTAL SCIENCES↗

Energy efficient photonic memory based on electrically programmable embedded III-V/Si memristors: switches and filters

Abstract Over the past few years, extensive work on optical neural networks has been investigated in hopes of achieving orders of magnitude improvement in energy efficiency and compute density via all-optical matrix-vector multiplication. However, these solutions are limited by a lack of high-speed power power-efficient phase tuners, on-chip non-volatile memory, and a proper material platform that can heterogeneously integrate all the necessary components needed onto a single chip. We address these issues by demonstrating embedded multi-layer HfO 2 /Al 2 O 3 memristors with III-V/Si photonics which facilitate non-volatile optical functionality for a variety of devices such as Mach-Zehnder Interferometers, and (de-)interleaver filters. The Mach-Zehnder optical memristor exhibits non-volatile optical phase shifts > π with ~33 dB signal extinction while consuming 0 electrical power consumption. We demonstrate 6 non-volatile states each capable of 4 Gbps modulation. (De-) interleaver filters were demonstrated to exhibit memristive non-volatile passband transformation with full set/reset states. Time duration tests were performed on all devices and indicated non-volatility up to 24 hours and beyond. We demonstrate non-volatile III-V/Si optical memristors with large electric-field driven phase shifts and reconfigurable filters with true 0 static power consumption. As a result, co-integrated photonic memristors offer a pathway for in-memory optical computing and large-scale non-volatile photonic circuits.

Cheung, Stanley (ORCID:0000000248860013)↗

Real-space solution to the electronic structure problem for nearly a million electrons

We report a Kohn–Sham density functional theory calculation of a system with more than 200 000 atoms and 800 000 electrons using a real-space high-order finite-difference method to investigate the electronic structure of large spherical silicon nanoclusters. Our system of choice was a 20 nm large spherical nanocluster with 202 617 silicon atoms and 13 836 hydrogen atoms used to passivate the dangling surface bonds. To speed up the convergence of the eigenspace, we utilized Chebyshev-filtered subspace iteration, and for sparse matrix–vector multiplications, we used blockwise Hilbert space-filling curves, implemented in the PARSEC code. For this calculation, we also replaced our orthonormalization + Rayleigh–Ritz step with a generalized eigenvalue problem step. We utilized all of the 8192 nodes (458 752 processors) on the Frontera machine at the Texas Advanced Computing Center. We achieved two Chebyshev-filtered subspace iterations, yielding a good approximation of the electronic density of states. Our work pushes the limits on the capabilities of the current electronic structure solvers to nearly 106 electrons and demonstrates the potential of the real-space approach to efficiently parallelize large calculations on modern high-performance computing platforms.

Chemistry↗

Integration of Ag-CBRAM crossbars and Mott ReLU neurons for efficient implementation of deep neural networks in hardware

In-memory computing with emerging non-volatile memory devices (eNVMs) has shown promising results in accelerating matrix-vector multiplications. However, activation function calculations are still being implemented with general processors or large and complex neuron peripheral circuits. Here, we present the integration of Ag-based conductive bridge random access memory (Ag-CBRAM) crossbar arrays with Mott rectified linear unit (ReLU) activation neurons for scalable, energy and area-efficient hardware (HW) implementation of deep neural networks. We develop Ag-CBRAM devices that can achieve a high ON/OFF ratio and multi-level programmability. Compact and energy-efficient Mott ReLU neuron devices implementing ReLU activation function are directly connected to the columns of Ag-CBRAM crossbars to compute the output from the weighted sum current. We implement convolution filters and activations for VGG-16 using our integrated HW and demonstrate the successful generation of feature maps for CIFAR-10 images in HW. Our approach paves a new way toward building a highly compact and energy-efficient eNVMs-based in-memory computing system.

Mott insulators↗